Capsaicin Genomics
How DNA & RNA Sequencing Works
Sequencing is the foundation of modern genomics. Whether reading the static blueprint of DNA or capturing the dynamic activity of RNA, sequencing technologies turn biological molecules into digital data that can be analyzed, compared, and used to guide engineering decisions.
DNA Sequencing
DNA sequencing reads the order of nucleotide bases — adenine (A), thymine (T), guanine (G), and cytosine (C) — in a DNA molecule. The genome of an organism is the complete set of these instructions, encoding every gene and regulatory element. DNA sequencing reveals the blueprint: what genes exist, where they are located, and what their sequences encode.
Modern short-read sequencing platforms fragment the genome into millions of small pieces (typically 100–300 base pairs each), read each fragment in parallel, and then computationally reassemble the fragments by aligning overlapping sequences to a reference genome. The result is a digital representation of the organism’s genome at single-base resolution.
RNA Sequencing (RNA-seq)
While DNA sequencing tells you what genes exist, RNA sequencing (RNA-seq) tells you which genes are actively being transcribed and at what levels. This is critical because the same genome is present in every cell, but different tissues express different genes at different levels. RNA-seq captures a snapshot of gene expression — the transcriptome — at a specific time and in a specific tissue.
The RNA-seq workflow proceeds through several stages:
- RNA extraction. Total RNA is isolated from the tissue of interest. For capsaicinoid research, this means specifically dissecting placental tissue from pepper fruit, because this is where the biosynthetic enzymes are expressed.
- Reverse transcription. RNA is converted to complementary DNA (cDNA) using reverse transcriptase. This step is necessary because sequencing platforms read DNA, not RNA directly.
- Fragmentation and library preparation. The cDNA is fragmented, adapters are ligated to each end, and the fragments are amplified by PCR to create a sequencing library.
- Sequencing. The library is loaded onto a sequencing platform, which reads millions of short fragments (reads) in parallel.
- Alignment. Reads are aligned to a reference genome or transcriptome to determine which gene each read originated from.
- Quantification. The number of reads mapping to each gene is counted. More reads means higher expression.
Transcript Quantification with Salmon
Scoville Splice uses Salmon 2.8.0 for transcript quantification. Salmon is a quasi-mapping tool, meaning it does not perform full read-to-genome alignment. Instead, it uses a lightweight index of the transcriptome to rapidly assign reads to transcripts based on k-mer matching and fragment-level compatibility. This makes Salmon dramatically faster than traditional alignment-based methods while maintaining comparable accuracy.
The output of Salmon is a table of transcript-level abundance estimates: how many fragments (reads) map to each transcript, along with normalized metrics like Transcripts Per Million (TPM) that account for transcript length and sequencing depth. These abundance estimates form the input for downstream differential expression analysis.
Differential Expression with PyDESeq2
After quantification, the next question is: which genes are significantly up-regulated or down-regulated compared to a control? This is differential expression analysis, and Scoville Splice uses PyDESeq2 — a Python implementation of the DESeq2 statistical framework — to answer it.
PyDESeq2 models read counts using a negative binomial distribution, estimates dispersion for each gene, and performs statistical tests to identify genes whose expression differs significantly between conditions (for example, engineered vs. wild-type tissue). Genes passing both a fold-change threshold and a statistical significance threshold (adjusted p-value) are called differentially expressed.
4,365 significant genes identified
In the Scoville Splice RNA-seq analysis, 4,365 genes passed the significance threshold for differential expression between engineered and wild-type placental tissue.
Why RNA-seq Matters for Pepper Engineering
RNA-seq is essential for capsaicinoid pathway engineering because it reveals which genes in the pathway are active, at what levels, and in which tissues. By profiling placental tissue — the specific site of capsaicin biosynthesis — RNA-seq shows whether engineered overexpression constructs are producing the intended transcript levels, whether compensatory changes are occurring in other pathway genes, and whether off-target effects are disrupting unrelated biological processes.
Without RNA-seq, pathway engineering would be blind. You could insert a construct, measure the final capsaicin output, and guess at what happened in between. With RNA-seq, every step of the pathway is visible at the transcript level, enabling rational iteration on construct design.
Data availability: Zenodo DOI 10.5281/zenodo.23267360 · bioRxiv BIORXIV/2026/758036