Bioluminescence—the production of light by living organisms—is one of nature’s most enchanting phenomena. At the heart of this glowing magic are Luciferases, a class of oxidative enzymes that produce light. But what happens when we go hunting for uncharacterized, “dark” variants of this enzyme in the shadows of the proteome?
Using the pg_bio PostgreSQL extension, we bypassed months of tedious wet-lab work and performed a structural homology search across millions of proteins in mere seconds.
The Biological Hook: A Tale of Two Extremes
To understand the power of spatial vector search in bioinformatics, we started with a known bait: A0A060HGF9, a Luciferase-like monooxygenase family protein. We then searched our database for the closest uncharacterized orphan proteins and found A0A1H6FX56.
By utilizing the newly implemented bio_search_uniprot() SQL Set-Returning Function (SRF), we fetched their taxonomic and biological metadata directly inside the database without leaving our SQL console:
SELECT id, organism, taxonomy
FROM bio_search_uniprot('accession:A0A060HGF9');
-- Returns: Nitrososphaera viennensis EN76 (Archaea)
SELECT id, organism, taxonomy
FROM bio_search_uniprot('accession:A0A1H6FX56');
-- Returns: Natronorubrum sediminis (Archaea)
The Ecological Divergence:
- The Bait (Nitrososphaera viennensis): This archaeon was isolated from temperate soils in Vienna. It is a mesophilic ammonia oxidizer, surviving in ordinary terrestrial environments where it plays a key role in the nitrogen cycle.
- The Discovery (Natronorubrum sediminis): This uncharacterized organism lives in a radically different world. It is an extreme haloalkaliphilic archaeon isolated from the sediments of soda lakes—environments characterized by punishing hypersalinity and extreme pH levels.
Despite these wildly divergent ecosystems, the physical structure and mechanism of action for their Luciferase-like proteins have been conserved through evolution. This structural homology indicates that the uncharacterized protein in N. sediminis likely acts through a similar oxidative mechanism to emit light or manage oxidative stress in its extreme halophilic environment.
The Math & The SQL
To find this match, we used ESM (Evolutionary Scale Modeling) embeddings to index the 3D structures into our database and queried them using pgvector.
Here is the exact pg_bio SQL query used to execute the search:
SELECT
uniprot_id AS discovery_id,
name,
distance,
hybrid_score
FROM pg_bio_search_homologs('A0A060HGF9', 0.35, 50, 0.0, 5, 2)
WHERE name ILIKE '%uncharacterized%';
| Bait (Luciferase) | Discovery (Uncharacterized) | Vector Distance (<=>) |
Hybrid Score (<~>) |
|---|---|---|---|
A0A060HGF9 |
A0A1H6FX56 |
0.061 | 0.294 |
Understanding the Metrics: A vector distance of 0.061 means the 3D backbone is almost mathematically identical. Even though the protein sequence has mutated to adapt to a soda lake environment (indicated by the hybrid sequence score), the spatial folding is locked in.
Interactive 3Dmol.js Preview
Check out the structural similarity below!
Pro Tip: This blog features an interactive 3D plugin (
pg_bio_sync.js). Double-click either 3D viewer to lock their cameras together for synchronized rotation and tilt. Click any fragment on one protein, and it will automatically highlight the matching residue across all complex chains on the opposite protein by swapping colors!
Bait: Soil Luciferase
Notice the central beta-barrel core used for oxidative substrate binding.
Discovery: Soda Lake Orphan
The structural pocket is highly conserved despite extreme salinity adaptations.
The Learning
This discovery highlights a monumental shift in computational biology. By embedding high-dimensional protein structures directly into PostgreSQL using pg_bio, we can execute complex spatial and functional homology searches across billions of records using simple SQL.
What used to take months of crystalline extraction, X-ray crystallography, and wet-lab alignment now takes milliseconds. The dark proteome isn’t so dark anymore—it’s just waiting to be queried.