Pipeline
The Six-Stage Pipeline
The Capsaicin Design Platform integrates six computational biology disciplines into a single, end-to-end pipeline. Each stage feeds its output into the next β RNA-seq informs the metabolic model, the model identifies engineering targets, AI scores protein mutations, CRISPR designs implement the modifications, blend chemistry defines the heat profiles, and DNA assembly produces synthesis-ready constructs.
Developmental RNA-seq
Transcriptomic Profiling
We profiled gene expression in Capsicum chinense placental tissue across five developmental time points (4, 12, 24, 36, and 54 days post-anthesis) to identify exactly when and where capsaicinoid biosynthesis genes activate. Salmon 2.8.0 quantified transcripts against the C. annuum UCD10Xv1.1 reference genome, and PyDESeq2 0.5.4 performed differential expression analysis with Wald testing and Benjamini-Hochberg correction. Result: 4,365 differentially expressed genes identified, with key pathway genes showing fold-changes from +3.91 (BCAT) to +8.34 (Pun1/AT3).
Output
Gene expression atlas of capsaicinoid biosynthesis
Flux Balance Analysis
Metabolic Modeling
We built a constraint-based mathematical model of the entire capsaicinoid pathway β 37 metabolites across 44 reactions β using COBRApy. Flux balance analysis revealed the critical finding: the vanillylamine branch carries a 90% flux control coefficient, while Pun1/AT3 (capsaicin synthase) carries 0%. This overturned the prevailing assumption that capsaicin synthase was rate-limiting and redirected our engineering strategy toward vanillylamine supply.
Output
Pathway bottleneck map and engineering priorities
AI Protein Engineering
ESM2 Mutation Scanning
Using ESM2 β Meta's 650-million-parameter protein language model β we scanned every possible single-amino-acid substitution in Pun1/AT3 (capsaicin synthase). Each mutation was scored by log-likelihood ratio to predict fitness effects. The top hits: S39L (LLR 2.828), L345G (LLR 2.695), S39F (LLR 2.155), C175S (LLR 2.025). These mutations are predicted to alter substrate specificity and catalytic efficiency. Products Heat 06 through Heat 10 incorporate single, double, and triple mutant combinations.
Output
Ranked mutation library with fitness predictions
CRISPR Construct Design
Precision Gene Editing
We designed 24 CRISPR-SpCas9 guide RNAs across four genomic targets. The foundation modification β present in all 10 cultivar specifications β knocks out peroxidase enzymes (LOC107864929 and LOC107856092) that actively degrade capsaicinoids in the placenta. Higher-heat cultivars add overexpression cassettes for PAL (phenylpropanoid pathway entry) and COMT (vanillylamine branch), plus Pun1/AT3 mutant variants. Each guide is scored for GC content, PAM site availability, and off-target risk.
Output
24 guide RNAs with off-target scoring
Blend Chemistry Optimization
Five-Capsaicinoid Ratio Design
Each cultivar specification has a precise five-capsaicinoid blend ratio β capsaicin, dihydrocapsaicin, nordihydrocapsaicin, homodihydrocapsaicin, and nonivamide β optimized for its target SHU and burn profile. The capsaicin fraction climbs from 40% at Heat 01 to 75% at Heat 10, driving the shift from slow-building body burn (Heats 01-03) through aggressive mid-mouth assault (Heats 04-06) to instantaneous sinus-stripping strike (Heats 07-10). Ratios account for TRPV1 receptor binding thermodynamics.
Output
Blend specification for each of 10 cultivars
DNA Sequence Assembly
Synthesis-Ready Constructs
The final stage generates complete, synthesis-ready DNA sequences for every genetic modification across all 10 products. The pipeline produced 25 DNA constructs totaling 31,911 base pairs. Each construct includes promoters, terminators, selection markers, and homology arms for genomic integration. All sequences have been submitted to NCBI GenBank (SUB16548149) and archived at Zenodo (DOI: 10.5281/zenodo.23267360). Constructs range from single-modification peroxidase knockouts (Heats 01-04) to four-component systems with pathway overexpression and enzyme triple mutants (Heat 10).
Output
25 synthesis-ready constructs, 31,911 bp total
From Pipeline to Product
The pipelineβs modular architecture means each stage can be independently updated as new data or tools become available β a better protein language model, improved CRISPR scoring algorithms, or expanded metabolic models. The platform is not a one-time analysis. It is a reusable design system for engineering capsaicinoid production in any Capsicum species.
The entire platform and all 10 cultivar designs are documented in our preprint (BIORXIV/2026/758036) and patent pending.