With 20+ years of experience, I specialize in database optimization, system architecture, and performance tuning large-scale applications.

Currently working as a Senior Software Engineer at Hubstaff, and former Developer Advocate at Timescale.

I'm also a community builder, a cyclist and a permaculture enthusiast.

Check my talks or connect with me via LinkedIn or Github.




Latest Post

Validating the Dark Proteome: Genomic Synteny and PyMOL Alignment in PostgreSQL

September 28, 2026 bioinformatics pgvector machine-learning structural-biology

In our previous post, we discovered a highly promising, uncharacterized halophilic cellulase (M0D2S5). But discovering a structural match using high-dimensional vectors is just the first step. To truly trust our discovery before heading to the wet lab, we must validate its biological context and its physical 3D fold geometry.

Today, we’ve extended pg_bio with two incredibly powerful new native PostgreSQL capabilities: Genomic Synteny Fetching (direct ENA/NCBI integration) and Automated PyMOL Validation.

1. Native Genomic Synteny in SQL

A protein does not act alone; it is heavily influenced by the genes surrounding it—its Genomic Neighborhood (Synteny). If our uncharacterized halophilic enzyme (M0D2S5) is genuinely a functioning cellulase, its gene should be physically located near other biomass-degrading operons or extremophile survival genes on the genome.

We just pushed native Rust bindings for NCBI’s E-Utilities and the European Nucleotide Archive (ENA) directly into pg_bio. This allows us to fetch millions of base pairs of raw genome sequence immediately into Postgres!

Here is the exact SQL we used to enrich our triage_dashboard:

-- First, find the EMBL Nucleotide Accession from UniProt
-- Then fetch the raw 250,000+ character genome natively into Postgres!
UPDATE orphan_discoveries 
SET neighborhood_sequence = bio_fetch_ncbi_sequence('AOIU01000008', 'nuccore')
WHERE orphan_id = 'M0D2S5';

By keeping the sequence inside Postgres, we avoid massive Python memory bottlenecks and can write SQL LIKE queries or parse_vcf() operations directly over the 1D spatial data to look for characteristic operon signatures.

2. Headless PyMOL Structural Validation

We don’t just want to look at genes; we want to physically see the active sites lock together. To completely bridge the gap between database and human validation, pg_bio now generates native .pml scripts.

Using a headless PyMOL integration via uv, we executed a sequence-independent structural superposition (super) between the bait cellulase and the new orphan.

Category Bait (A0A1G7TQG7) Orphan (M0D2S5) Alignment RMSD
Organism Halorientalis regularis Halosimplex carlsbadense 1.726 Å
Fold Context Known Halophilic Cellulase Uncharacterized (Over 134 atoms)

An RMSD of 1.726 Å is a stunningly tight fit for two highly divergent species. It proves mathematically that despite genetic drift over millions of years, the core catalytic barrel shape required to break down cellulose remains physically identical.

Interactive Validation

Below is the synchronized 3D validation of our bait versus our discovery. (Hint: Double click to lock their cameras together, and click on any residue to highlight its exact structural counterpart across the alignment!)

The Conclusion

We have successfully built a closed-loop system entirely within PostgreSQL. We:

  1. Mined the dark proteome using vector similarities and Attention maps.
  2. Fetched the native genomic synteny using built-in NCBI HTTP requests in Rust.
  3. Validated the 3D active site geometry by spawning PyMOL from the database pipeline.

Tasks that traditionally required brittle multi-container ETL architectures and weeks of manual processing are now running automatically overnight, safely recorded in orphan_discoveries.

Synthetic biology just got a lot faster.


Recent Posts