Capsaicin Genomics
ESM2 Protein Language Model for Pun1 Engineering
Traditional protein engineering relies on random mutagenesis or directed evolution — generating thousands of variants and screening them for improved function. Scoville Splice replaced that brute-force approach with AI-guided protein engineering, using the ESM2 protein language model to predict which mutations would improve Pun1 catalytic efficiency before a single construct was ever assembled.
What Is ESM2?
ESM2 is a 650-million-parameter protein language model developed for predicting protein structure and function from amino acid sequences. Trained on tens of millions of protein sequences across the tree of life, ESM2 learns the statistical patterns that govern how amino acids co-evolve within and across protein families. These patterns encode deep structural and functional constraints — the model effectively learns which substitutions are tolerated at each position and which would disrupt folding, stability, or catalytic activity.
The key insight is that evolution itself has already explored an enormous space of protein variants. By training on the surviving sequences, ESM2 captures the fitness landscape of protein families without ever needing experimental data on the specific protein being engineered.
Scanning Pun1 Mutations
Scoville Splice used ESM2 to perform a saturation scan of the Pun1 protein: every possible single-amino-acid substitution at every position in the sequence was evaluated computationally. For each mutation, the model assigns a log-likelihood ratio (LLR) that quantifies how compatible the substitution is with proper protein folding and function. A higher LLR indicates the model predicts the variant will maintain or enhance structural integrity and catalytic activity. A negative LLR signals a likely deleterious mutation.
This approach is fundamentally different from random mutagenesis. Instead of screening thousands of variants in the lab, ESM2 narrows the search to a handful of high-confidence candidates — each computationally validated before construct design begins. The result is faster iteration, lower cost, and higher success rates.
Three Winning Pun1 Mutations
From the full saturation scan, three mutations emerged with the highest log-likelihood ratios, indicating strong predicted compatibility with Pun1 structure and function:
S39L — Serine → Leucine at position 39
LLR: 2.828 — highest-scoring mutation in the scan
L345G — Leucine → Glycine at position 345
LLR: 2.695 — predicted to improve active-site flexibility
C175S — Cysteine → Serine at position 175
LLR: 2.025 — removes a reactive thiol group that may cause unwanted disulfide bonds
Each of these substitutions was selected because ESM2 predicted it would improve the protein’s ability to fold correctly and maintain catalytic function under the increased substrate flux created by upstream pathway engineering. Six total enzyme variants were designed across the capsaicinoid pathway, with the Pun1 mutations specifically addressing the need for higher catalytic efficiency at the final condensation step.
Why This Matters for Capsaicin Engineering
Pun1 is the terminal enzyme in capsaicin biosynthesis. It condenses vanillylamine with a fatty acid acyl chain to produce capsaicin. When upstream enzymes like pAMT and PAL are overexpressed to increase precursor supply, Pun1 becomes the rate-limiting step. Wild-type Pun1 simply cannot process the elevated vanillylamine flux fast enough.
The ESM2-guided mutations improve Pun1’s catalytic efficiency so it can handle the increased substrate load. This is the critical connection between pathway engineering and protein engineering: it is not enough to flood the pathway with precursors if the final enzyme cannot keep up. The mutations ensure that the engineered flux translates into higher capsaicin output rather than a metabolic bottleneck.
AI-Guided vs. Traditional Approaches
Traditional directed evolution creates libraries of thousands to millions of random variants, then screens them through multiple rounds of selection. This process is powerful but expensive, slow, and requires a high-throughput screening assay. Random mutagenesis generates many neutral or deleterious mutations for every beneficial one.
The ESM2 approach inverts this workflow. Computational prediction replaces random screening. The model evaluates the entire mutational landscape in silico, ranks every possible substitution by predicted fitness, and delivers a short list of candidates worth testing. No directed evolution libraries, no high-throughput screening infrastructure, no months of iterative rounds. Each mutation is computationally validated before construct design, and only the highest-confidence variants proceed to synthesis.
Patent Pending. The ESM2-guided protein engineering approach described here, including the specific Pun1 mutations and the method of computational variant selection for capsaicinoid pathway enzymes, is proprietary to Scoville Splice.