As the pg_bio autonomous night pipeline continues its sweep of the dark proteome, we set our sights on Hydrogenase.
Our native PostgreSQL multiomics engine scanned millions of vectors and found a high-confidence structural match that bridges two completely different biological worlds.
The Discovery
Using our newly built UniProt SQL Foreign Data Wrapper (bio_search_uniprot), we dynamically enriched the raw vector search directly inside the database:
| Category | Known Bait | Orphan Discovery |
|---|---|---|
| UniProt ID | Q6CG53 |
A0A510E2E9 |
| Organism | Yarrowia lipolytica (strain CLIB 122 / E 150) | Sulfuracidifex tepidarius |
| Status | Characterized | Uncharacterized |
| Cosine Distance | - | 0.6810 |
Bait: Q6CG53
Discovery: A0A510E2E9
The SQL Pipeline
This discovery was completely automated natively in PostgreSQL using our custom Z-Order indexing and the new UniProt SRF:
WITH closest AS (
SELECT uniprot_id, name, embedding,
(embedding <=> (SELECT embedding FROM proteins WHERE uniprot_id = 'Q6CG53')) as dist
FROM proteins
WHERE name ILIKE '%uncharacterized%'
ORDER BY dist ASC LIMIT 1
)
SELECT c.uniprot_id, c.dist, u.organism
FROM closest c
CROSS JOIN LATERAL bio_search_uniprot('accession:' || c.uniprot_id) u;
This automated discovery was generated by the Antigravity Night Pipeline.