With 20+ years of experience, I specialize in database optimization, system architecture, and performance tuning large-scale applications.
Currently working as a Senior Software Engineer at Hubstaff, and former Developer Advocate at Timescale.
I'm also a community builder, a cyclist and a permaculture enthusiast.
Featured Content
Latest Talks
Explore my recent presentations at conferences and meetups around the world.
View TalksInteractive Hub & Tools
Apps & Playgrounds
Check out the interactive tools directory, featuring the Semantic Learning journey, full-screen Drawing Recorder, 7m Geodesic Dome Builder, and the new Truck Camper Bamboo Builder.
Explore AppsMandala Drawings
Explore the Mandala Synth and Playground. Draw mandalas, sonify visuals, and experiment with complex geometries mapped to audio.
Launch PlaygroundLatest Post
Unearthing PHA synthase: Exploring the Dark Proteome of Extreme Ecosystems!
September 30, 2026 bioinformatics pgvector machine-learning structural-biology pgbio
As the pg_bio autonomous night pipeline continues its exciting sweep of the dark proteome, we set our sights on an incredible protein family: PHA synthase! By bypassing months of wet-lab work, we are uncovering hidden secrets of nature using the immense power of native PostgreSQL multiomics engines scanning millions of vectors in milliseconds.
Our SQL engine scanned the embedding space and found a high-confidence structural match that bridges two completely different biological worlds. We found an uncharacterized orphan protein that exhibits an almost identical 3D fold to a known, well-studied bait!
The Bait: Unknown protein (M0G5K0)
To understand the magnitude of this discovery, we first must look at the known bait protein from Haloferax prahovense (strain DSM 18310 / JCM 13924 / TL6). What does it do? No specific function described.
This specific enzymatic function is crucial to its ecosystem. But what happens when we search the vast, uncharted territories of the database for something structurally similar?
The Discovery: A Hidden Orphan in Natrarchaeobaculum sulfurireducens
Our search revealed an entirely uncharacterized protein (A0A346PQB3) in Natrarchaeobaculum sulfurireducens. Despite its label as “uncharacterized”, its vector embeddings tell a different story!
The structural similarity implies a massive evolutionary divergence or a conserved function adapted to a completely new environment. Could this extremophile or unique organism be harboring a more robust, efficient version of the enzyme?
Practical Applications & Impact
What does this mean for the real world? Proteins in the PHA synthase family have massive potential in industrial biotechnology, bioremediation, medicine, and synthetic biology. By finding a novel version of this protein in Natrarchaeobaculum sulfurireducens, we might have just discovered a variant that operates at extreme temperatures, pH levels, or with higher catalytic efficiency! This is the power of mining the dark proteome.
The Math & The Pipeline
Using our newly built UniProt SQL Foreign Data Wrapper (bio_search_uniprot), we dynamically enriched the raw vector search directly inside the database:
| Category | Known Bait | Orphan Discovery |
|---|---|---|
| UniProt ID | M0G5K0 |
A0A346PQB3 |
| Organism | Haloferax prahovense (strain DSM 18310 / JCM 13924 / TL6) | Natrarchaeobaculum sulfurireducens |
| Status | Characterized | Uncharacterized |
| Cosine Distance | - | 0.0715 |
Note: A distance of 0.0715 means the 3D backbone is mathematically incredibly similar!
Interactive 3Dmol.js Preview
Dive into the structures below! Tip: Double-click either 3D viewer to lock their cameras together for synchronized rotation, and click any fragment to automatically highlight the matching residue on the opposite protein!
Bait: M0G5K0 (Haloferax prahovense (strain DSM 18310 / JCM 13924 / TL6))
Discovery: A0A346PQB3 (Natrarchaeobaculum sulfurireducens)
The SQL Query
This discovery was completely automated natively in PostgreSQL using our custom Z-Order indexing and the new UniProt SRF:
WITH closest AS (
SELECT uniprot_id, name, embedding,
(embedding <=> (SELECT embedding FROM proteins WHERE uniprot_id = 'M0G5K0')) as dist
FROM proteins
WHERE name ILIKE '%uncharacterized%'
ORDER BY dist ASC LIMIT 1
)
SELECT c.uniprot_id, c.dist, u.organism
FROM closest c
CROSS JOIN LATERAL bio_search_uniprot('accession:' || c.uniprot_id) u;
This automated discovery was generated by the pg_bio continuous discovery script.
Check out more of our pipeline’s findings in the Bioinformatics discoveries section!