Unearthing Cellulase: Exploring the Dark Proteome of Extreme Ecosystems

As the pg_bio autonomous night pipeline continues its sweep of the dark proteome, we set our sights on Cellulase.

Our native PostgreSQL multiomics engine scanned millions of vectors and found a high-confidence structural match that bridges two completely different biological worlds.

The Discovery

Using our newly built UniProt SQL Foreign Data Wrapper (bio_search_uniprot), we dynamically enriched the raw vector search directly inside the database:

Category Known Bait Orphan Discovery
UniProt ID A0A1G7TQG7 M0D2S5
Organism Halorientalis regularis Halosimplex carlsbadense 2-9-1
Status Characterized Uncharacterized
Cosine Distance - 0.0603

Bait: A0A1G7TQG7

Discovery: M0D2S5

The SQL Pipeline

This discovery was completely automated natively in PostgreSQL using our custom Z-Order indexing and the new UniProt SRF:

WITH closest AS (
    SELECT uniprot_id, name, embedding,
           (embedding <=> (SELECT embedding FROM proteins WHERE uniprot_id = 'A0A1G7TQG7')) as dist
    FROM proteins
    WHERE name ILIKE '%uncharacterized%'
    ORDER BY dist ASC LIMIT 1
)
SELECT c.uniprot_id, c.dist, u.organism
FROM closest c
CROSS JOIN LATERAL bio_search_uniprot('accession:' || c.uniprot_id) u;

This automated discovery was generated by the Antigravity Night Pipeline.

Jônatas Davi Paganini

Jônatas Davi Paganini

Senior developer and technical consultant with 20+ years of experience specializing in PostgreSQL, TimescaleDB, and distributed systems. Expert in database optimization, microservices architecture, and team enablement. Passionate about sharing knowledge through writing, speaking, and mentoring.

1/1