Pipeline

The Six-Stage Pipeline

The Capsaicin Design Platform integrates six computational biology disciplines into a single, end-to-end pipeline. Each stage feeds its output into the next β€” RNA-seq informs the metabolic model, the model identifies engineering targets, AI scores protein mutations, CRISPR designs implement the modifications, blend chemistry defines the heat profiles, and DNA assembly produces synthesis-ready constructs.

01

Developmental RNA-seq

Transcriptomic Profiling

We profiled gene expression in Capsicum chinense placental tissue across five developmental time points (4, 12, 24, 36, and 54 days post-anthesis) to identify exactly when and where capsaicinoid biosynthesis genes activate. Salmon 2.8.0 quantified transcripts against the C. annuum UCD10Xv1.1 reference genome, and PyDESeq2 0.5.4 performed differential expression analysis with Wald testing and Benjamini-Hochberg correction. Result: 4,365 differentially expressed genes identified, with key pathway genes showing fold-changes from +3.91 (BCAT) to +8.34 (Pun1/AT3).

Output

Gene expression atlas of capsaicinoid biosynthesis

02

Flux Balance Analysis

Metabolic Modeling

We built a constraint-based mathematical model of the entire capsaicinoid pathway β€” 37 metabolites across 44 reactions β€” using COBRApy. Flux balance analysis revealed the critical finding: the vanillylamine branch carries a 90% flux control coefficient, while Pun1/AT3 (capsaicin synthase) carries 0%. This overturned the prevailing assumption that capsaicin synthase was rate-limiting and redirected our engineering strategy toward vanillylamine supply.

Output

Pathway bottleneck map and engineering priorities

03

AI Protein Engineering

ESM2 Mutation Scanning

Using ESM2 β€” Meta's 650-million-parameter protein language model β€” we scanned every possible single-amino-acid substitution in Pun1/AT3 (capsaicin synthase). Each mutation was scored by log-likelihood ratio to predict fitness effects. The top hits: S39L (LLR 2.828), L345G (LLR 2.695), S39F (LLR 2.155), C175S (LLR 2.025). These mutations are predicted to alter substrate specificity and catalytic efficiency. Products Heat 06 through Heat 10 incorporate single, double, and triple mutant combinations.

Output

Ranked mutation library with fitness predictions

04

CRISPR Construct Design

Precision Gene Editing

We designed 24 CRISPR-SpCas9 guide RNAs across four genomic targets. The foundation modification β€” present in all 10 cultivar specifications β€” knocks out peroxidase enzymes (LOC107864929 and LOC107856092) that actively degrade capsaicinoids in the placenta. Higher-heat cultivars add overexpression cassettes for PAL (phenylpropanoid pathway entry) and COMT (vanillylamine branch), plus Pun1/AT3 mutant variants. Each guide is scored for GC content, PAM site availability, and off-target risk.

Output

24 guide RNAs with off-target scoring

05

Blend Chemistry Optimization

Five-Capsaicinoid Ratio Design

Each cultivar specification has a precise five-capsaicinoid blend ratio β€” capsaicin, dihydrocapsaicin, nordihydrocapsaicin, homodihydrocapsaicin, and nonivamide β€” optimized for its target SHU and burn profile. The capsaicin fraction climbs from 40% at Heat 01 to 75% at Heat 10, driving the shift from slow-building body burn (Heats 01-03) through aggressive mid-mouth assault (Heats 04-06) to instantaneous sinus-stripping strike (Heats 07-10). Ratios account for TRPV1 receptor binding thermodynamics.

Output

Blend specification for each of 10 cultivars

06

DNA Sequence Assembly

Synthesis-Ready Constructs

The final stage generates complete, synthesis-ready DNA sequences for every genetic modification across all 10 products. The pipeline produced 25 DNA constructs totaling 31,911 base pairs. Each construct includes promoters, terminators, selection markers, and homology arms for genomic integration. All sequences have been submitted to NCBI GenBank (SUB16548149) and archived at Zenodo (DOI: 10.5281/zenodo.23267360). Constructs range from single-modification peroxidase knockouts (Heats 01-04) to four-component systems with pathway overexpression and enzyme triple mutants (Heat 10).

Output

25 synthesis-ready constructs, 31,911 bp total

From Pipeline to Product

The pipeline’s modular architecture means each stage can be independently updated as new data or tools become available β€” a better protein language model, improved CRISPR scoring algorithms, or expanded metabolic models. The platform is not a one-time analysis. It is a reusable design system for engineering capsaicinoid production in any Capsicum species.

The entire platform and all 10 cultivar designs are documented in our preprint (BIORXIV/2026/758036) and patent pending.