DOE Grant — Submitted
PI: David Ussery, Oklahoma State University
Co-Investigators: [Marcin Joachimiak (LBNL), Arvind Ramanathan (ANL), Bernhard Palsson (UCSD)]
Project Summary
This Phase I project advances the DOE BER mission of "Scaling the Biotechnology Revolution" through Predictive Engineering of Microbial Communities, targeting genomic "dark matter" by building on extensive knowledge from two well defined bacterial systems.
We will map LLMs for the ~500 persistent bacterial functions to the minimal-genome M. genitalium, as a tractable baseline for annotation pipelines, where currently about 60% or more of the metagenome genes have no known function. We then scale to several hundred thousand unique E. coli pan-genome families, identifying phylogroup-specific metabolic genes and predicting metabolic niche differentiation.
We benchmark ESM-3, Evo2, and ProtT5 on DOE supercomputers for protein annotation. Priority targets are dark matter genes in functional classes governing community behavior. iModulon-derived gene expression modules provide metadata as a tokenizable layer alongside sequence embeddings; dark matter genes with condition-correlated co-expression patterns are prioritized.
We systematically analyze how embeddings from each model encode secondary structure and 3D topology, using this structure-embedding correspondence to derive richer representations for unannotated proteins. KOGUT — a relation-aware graph transformer adapted from RelGT for knowledge graph reasoning — integrates these pLM sequence embeddings and tokenized expression modules with KG-Microbe's relational structure, grounding predictions in verified traits and reducing hallucinated annotations.