Research
How does protein sequence encode chemical interactions?
Recent breakthroughs with protein structure prediction highlight a successful approach: models trained on massive amounts of high quality data. Unlike structure, this data is extremely sparse for chemical biology.
The d’Oelsnitz lab focuses on scaling data generation for chemical biology by leveraging genetically encoded chemical sensors, which precisely convert small molecule concentrations into programmable genetic outputs in living cells. We largely rely on prokaryotic transcription factors due to their generalizability and ease of engineering. By linking sensor output to cell fluorescence (via GFP expression) or growth (via antibiotic marker expression), we can measure millions of protein-chemical affinities in a single experiment. Data at this scale is needed to train machine learning models with high predictive accuracy that generalize across vast chemical and protein sequence spaces.
We target two impactful protein classes: biosensors and enzymes.
Biosensing
Biosensors link small molecule concentrations into programmable gene expression outputs, powering our ability to scale biochemical measurements. In particular, we specialize in engineering prokaryotic transcription factors as chemical biosensors, which have wide-ranging applications in diagnostics, biocatalyst screens, environmental monitoring, real-time metabolite tracking, and human gut monitoring.
We have two primary aims:
- Characterize the DNA and chemical specificities of all >10 million sequenced biosensors scattered across bacterial and archaeal genomes. This would enable us to predict how microbes respond to chemistry.
- Establish a platform to engineer highly specific biosensors for virtually any biomolecule, with an initial focus on therapeutics. This would transform analytical chemistry and usher in high-throughput biochemistry.

Biocatalysis
Enzymes are ‘transforming’ the way we manufacture pharmaceuticals. However, engineering a fit-for-purpose enzyme often requires several iterative rounds of directed evolution. Using custom biosensors, we aim to generate enormous enzyme function datasets, which will be used to train models capable of efficiently navigating enzyme mutational fitness landscapes towards desired outcomes.
Our aims are two-fold:
- Develop novel biosensor-powered methods capable of quantitatively measuring the activity, rate, and stereoselectivity of >100,000 enzymes. These methods would enable us to create an enzyme database 10X larger than BRENDA in days.
- Generate models trained on our massive biochemical datasets to design enzymes with tailored properties. Initially, we will focus on industrially championed enzyme classes, such as IREDs, KREDs, and transaminases.
