Infrastructure
The Computational Lab
The Capsaicin Design Platform is, first and foremost, a computational operation. Every one of the 10 engineered cultivar specifications was designed entirely in silico β from RNA-seq data processing through final DNA construct assembly. Our lab is in Irvine, California 92602.
Transcriptomic Analysis Pipeline
The foundation of the platform is RNA-seq β profiling gene expression in Capsicum chinense placental tissue across five developmental time points (4, 12, 24, 36, and 54 days post-anthesis). Our pipeline uses Salmon 2.8.0 for transcript quantification against the C. annuum UCD10Xv1.1 reference genome (GCF_002878395.1), followed by PyDESeq2 0.5.4 for differential expression analysis with Wald testing and Benjamini-Hochberg correction.
This pipeline identified 4,365 differentially expressed genes (padj < 0.01, |log2FC| > 2) and mapped the complete activation sequence of the capsaicinoid biosynthesis pathway. Key gene fold-changes: Pun1/AT3 at +8.34, PAL at +6.82, pAMT at +6.41, 4CL at +5.18, COMT at +4.73.
Metabolic Modeling Environment
The flux balance analysis environment runs on COBRApy, modeling the complete capsaicinoid biosynthesis pathway as a constraint-based system of 37 metabolites and 44 reactions. This is the infrastructure that produced the central discovery: the vanillylamine branch carries 90% of total flux control, while capsaicin synthase (Pun1/AT3) carries 0%.
The modeling environment includes sensitivity analysis capabilities for parameter perturbation studies, allowing us to predict the quantitative impact of each genetic modification on SHU output before designing any constructs.
AI Protein Engineering Workstation
Our protein engineering pipeline centers on ESM2 β Metaβs 650-million-parameter protein language model. The workstation runs a comprehensive single-amino-acid substitution scan across the entire Pun1/AT3 (capsaicin synthase) sequence, scoring each mutation by log-likelihood ratio to predict which substitutions will enhance catalytic activity.
The top mutations β S39L (LLR 2.828), L345G (LLR 2.695), S39F (LLR 2.155), C175S (LLR 2.025) β were computationally identified and are incorporated into Heat 06 through Heat 10 cultivar designs. The workstation generates PyTorch-based inference results that feed directly into the construct assembly pipeline.
CRISPR Guide Design Suite
The CRISPR design suite handles guide RNA design, off-target analysis, and construct architecture. We designed 24 total guide RNAs across four genomic targets, each scored for GC content, PAM availability, and off-target risk. The primary modification β peroxidase knockout to stop capsaicinoid degradation β uses guides targeting LOC107864929 and LOC107856092.
The suite outputs complete CRISPR-SpCas9 construct specifications, including promoter sequences, guide RNA cassettes, and selection markers, formatted for direct synthesis ordering.
Blend Chemistry Calculator
Each cultivar specification requires a precise five-capsaicinoid blend ratio. The blend chemistry calculator optimizes ratios of capsaicin, dihydrocapsaicin, nordihydrocapsaicin, homodihydrocapsaicin, and nonivamide based on TRPV1 receptor binding thermodynamics and target burn characteristics (onset speed, duration, and body profile).
SHU is calculated as: SHU = Ξ£(Ci Γ Ξ±i Γ 106), where Ci is concentration in mg/g and Ξ±i is the HPLC response coefficient. The calculator produces specifications for all 10 cultivars, from Heat 01 (Verbal Warning, 3M SHU) through Heat 10 (D.N.R., 13M SHU).
DNA Construct Assembly
The final stage of the pipeline generates synthesis-ready DNA sequences. The assembly system produced 25 constructs totaling 31,911 base pairs, covering all genetic modifications across all 10 cultivar specifications. Every construct includes promoters, terminators, selection markers, and homology arms for genomic integration.
All sequences have been submitted to NCBI GenBank (SUB16548149) and archived at Zenodo (DOI: 10.5281/zenodo.23267360). The constructs are formatted for direct synthesis ordering from commercial DNA synthesis providers.