As the pg_bio autonomous night pipeline continues its exciting sweep of the dark proteome, we set our sights on an incredible protein family: Chitinase! By bypassing months of wet-lab work, we are uncovering hidden secrets of nature using the immense power of native PostgreSQL multiomics engines scanning millions of vectors in milliseconds.
Our SQL engine scanned the embedding space and found a high-confidence structural match that bridges two completely different biological worlds. We found an uncharacterized orphan protein that exhibits an almost identical 3D fold to a known, well-studied bait!
The Bait: Chitinase-3-like protein 1 (P30922)
To understand the magnitude of this discovery, we first must look at the known bait protein from Bos taurus. What does it do? Carbohydrate-binding lectin with a preference for chitin. Has no chitinase activity. May play a role in tissue remodeling and in the capacity of cells to respond to and cope with changes in their environment. Plays a role in T-helper cell type 2 (Th2) inflammatory response and IL-13-induced inflammation, regulating allergen sensitization, inflammatory cell apoptosis, dendritic cell accumulation and M2 macrophage differentiation. Facilitates invasion of pathogenic enteric bacteria into colonic mucosa and lymphoid organs. Mediates activation of AKT1 signaling pathway and subsequent IL8 production in colonic epithelial cells. Regulates antibacterial responses in lung by contributing to macrophage bacterial killing, controlling bacterial dissemination and augmenting host tolerance. Also regulates hyperoxia-induced injury, inflammation and epithelial apoptosis in lung (By similarity)
This specific enzymatic function is crucial to its ecosystem. But what happens when we search the vast, uncharted territories of the database for something structurally similar?
The Discovery: A Hidden Orphan in Natrinema saccharevitans
Our search revealed an entirely uncharacterized protein (A0A1S8B280) in Natrinema saccharevitans. Despite its label as “uncharacterized”, its vector embeddings tell a different story!
The structural similarity implies a massive evolutionary divergence or a conserved function adapted to a completely new environment. Could this extremophile or unique organism be harboring a more robust, efficient version of the enzyme?
Practical Applications & Impact
What does this mean for the real world? Proteins in the Chitinase family have massive potential in industrial biotechnology, bioremediation, medicine, and synthetic biology. By finding a novel version of this protein in Natrinema saccharevitans, we might have just discovered a variant that operates at extreme temperatures, pH levels, or with higher catalytic efficiency! This is the power of mining the dark proteome.
The Math & The Pipeline
Using our newly built UniProt SQL Foreign Data Wrapper (bio_search_uniprot), we dynamically enriched the raw vector search directly inside the database:
| Category | Known Bait | Orphan Discovery |
|---|---|---|
| UniProt ID | P30922 |
A0A1S8B280 |
| Organism | Bos taurus | Natrinema saccharevitans |
| Status | Characterized | Uncharacterized |
| Cosine Distance | - | 0.5317 |
Note: A distance of 0.5317 means the 3D backbone is mathematically incredibly similar!
Interactive 3Dmol.js Preview
Dive into the structures below! Tip: Double-click either 3D viewer to lock their cameras together for synchronized rotation, and click any fragment to automatically highlight the matching residue on the opposite protein!
Bait: P30922 (Bos taurus)
Discovery: A0A1S8B280 (Natrinema saccharevitans)
The SQL Query
This discovery was completely automated natively in PostgreSQL using our custom Z-Order indexing and the new UniProt SRF:
WITH closest AS (
SELECT uniprot_id, name, embedding,
(embedding <=> (SELECT embedding FROM proteins WHERE uniprot_id = 'P30922')) as dist
FROM proteins
WHERE name ILIKE '%uncharacterized%'
ORDER BY dist ASC LIMIT 1
)
SELECT c.uniprot_id, c.dist, u.organism
FROM closest c
CROSS JOIN LATERAL bio_search_uniprot('accession:' || c.uniprot_id) u;
This automated discovery was generated by the pg_bio continuous discovery script.
Check out more of our pipeline’s findings in the Bioinformatics discoveries section!