If there is one protein responsible for life on Earth as we know it, it is RuBisCO (Ribulose-1,5-bisphosphate carboxylase-oxygenase). It is the primary enzyme used by plants and cyanobacteria to capture $CO_2$ from the atmosphere and turn it into the sugars that fuel the global food chain.
Despite being the most abundant protein on the planet, RuBisCO is notoriously inefficient. Nature has struggled for billions of years to optimize it. But what if a better, faster, or more resilient version of RuBisCO is hiding in the unmapped genomes of extremophile bacteria?
Using the pg_bio PostgreSQL extension, we cast a structural “Bait” into the Dark Proteome to find out.
The SQL Query: Hunting for Carbon Fixers
To find unknown variants of RuBisCO, we took the 1024-dimensional AI embedding of a known cyanobacterial RuBisCO (L0JL84) and ran a spatial $K$-Nearest Neighbors search against our massive database of uncharacterized archaeal proteins.
Here is the exact query that scanned millions of proteins in milliseconds:
SELECT
b.uniprot_id AS known_bait,
o.uniprot_id AS dark_discovery,
b.struct_embedding <=> o.struct_embedding AS vector_distance
FROM
bio_bait b
CROSS JOIN LATERAL (
SELECT
uniprot_id,
struct_embedding
FROM
dark_proteome_orphans
WHERE
struct_embedding <=> b.struct_embedding < 0.1
ORDER BY
struct_embedding <=> b.struct_embedding ASC
LIMIT 1
) o
WHERE
b.name ILIKE '%rubisco%';
The Discovery: A Massive Structural Match
The batch job returned an absolutely staggering hit. Deep within the Dark Proteome, we uncovered an uncharacterized protein (A0A5B9DCV6) that is a near-perfect structural clone of our RuBisCO Bait.
| Bait (Known RuBisCO) | Discovery (Dark Proteome) | Vector Distance |
|---|---|---|
L0JL84 |
A0A5B9DCV6 |
0.0575 |
A vector distance of 0.0575 is exceptionally tight for such a massive, multi-subunit complex. While the amino acid sequence might have diverged over millions of years of evolution, the AI immediately recognized that the 3D atomic scaffolding is fundamentally identical.
The Ecosystem Context: From Salt Lakes to the Deep Ocean
What makes this discovery truly breathtaking is the ecological divergence between the two organisms.
Our Bait (L0JL84) comes from Natrinema pellirubrum, an extreme halophile that lives in hypersaline lakes—environments so salty that almost no other life can survive.
But our Discovery (A0A5B9DCV6) comes from an entirely different and legendary organism: Promethearchaeum syntrophicum.
Discovered deep in the Nankai Trough ocean trench off the coast of Japan, Promethearchaeum is an “Asgard archaeon”—the closest living relative to the ancient microbe that swallowed a bacterium to create the very first eukaryotic cell (and eventually, human life). It lives in anaerobic methane seeps, growing incredibly slowly by physically entangling itself with other bacteria in a symbiotic relationship (syntrophy).
Finding a functional RuBisCO analog inside an ancient, deep-sea Asgard archaeon completely rewrites our understanding of how early life managed carbon capture before the invention of photosynthesis!
Interactive 3D Comparison
Don’t just trust the math—explore the physical structures yourself using our interactive 3D widget.
Interactive Features:
- ⛶ Enter Fullscreen: Click the new circular button in the top right to expand the viewer and drag the center divider to resize the panels.
- Double-Click: Double-click either protein to lock their cameras together for synchronized rotation!
- Click-to-Highlight: Click any fragment (like an Alpha-helix) on the right protein, and it will instantly light up the exact matching structural components across the entire complex on the left!
Known Bait (RuBisCO)
Our Discovery (Distance: 0.0575)
(Notice the sprawling, barrel-like architectures. The identical structural topology confirms this is a novel carbon-capturing machine!)
Conclusion
Finding a novel RuBisCO structural analog is a holy grail for synthetic biology and climate technology. If this newly discovered extremophile enzyme is naturally more efficient at capturing $CO_2$ than plant RuBisCO, it could be engineered into crops to drastically increase agricultural yields, or utilized in bioreactors to pull greenhouse gases out of the atmosphere.
By leveraging AI embeddings and standard PostgreSQL pg_bio queries, we are mapping the dark corners of biology faster than ever before.