As an increasing number of RNA therapeutics progress through clinical trials and reach the clinic, the importance of accurate and thorough quality control assessments continues to rise. Next-generation and direct RNA sequencing are extremely useful technologies for validating RNA therapeutics. These approaches use basecalling to determine what nucleotide is present in each position on a DNA or RNA sequence.
Basecalling is used in all types of next-generation and direct RNA sequencing, but the mechanism varies based on the type. Different short- and long-read sequencing platforms and their basecalling approaches have different strengths and weaknesses, but regardless of the sequencing platform used, basecalling reveals critical insights into a sequence that provide information for many attributes of an RNA therapeutic.
In this eBlog, we will review basecalling for short- and long-read sequencing and discuss how they help with RNA characterization.
Basecalling with short-read sequencing
To perform basecalling during short-read sequencing, such as with Illumina sequence data, the sequence fragments must be converted into a sequencing library. When sequencing RNA, the sequence is first reverse transcribed into cDNA. Next, cDNA fragments typically undergo PCR amplification and are loaded onto the sequencer. Fragments are then amplified again into clusters. These clusters are exposed to cycles of sequencing by synthesis with fluorescent bases that show a different level of intensity or color for each nucleotide base (C, G, A, and T). Throughout the cycles, the corresponding signal is read, and a basecalling algorithm reports, or calls, the bases in order.
After basecalling, each call receives a Phred quality score, which represents how confident that specific call is. Caller confidence is determined by attributes like the purity and intensity of the signal. The algorithm presents that data in a FASTQ file.
This method of basecalling with short-read sequencing is known to be extremely accurate, but it requires fragmentation, reverse transcription, PCR, and cluster amplification to sequence the starting material. In addition, each critical quality attribute that drug developers want to characterize, such as RNA fragmentation, double-stranded RNA contamination, and capping efficiency, must be determined in its own read. These amplification steps and individual characterization reads take time and resources, which can be challenging for tight drug development timelines.
Basecalling with direct RNA sequencing
Basecalling when performing direct RNA, or long-read sequencing with nanopores, does not require preparatory steps like fragmentation or amplification. Rather, an entire full-length RNA molecule is inputted into a nanopore and each base is identified.
For this process, motor proteins feed a full-length RNA molecule through a nanopore that has an electric current moving across it. As the RNA moves through the nanopore, each type of nucleotide base blocks or alters the electric current differently. The varying blocks generate different current signals, which are read by a sensor and marked in a continuous line drawn by the signal reading machine. Each base produces a different curve or loop on the continuous line, which is called a “squiggle.”
After the signals are recorded, a basecalling algorithm decodes the squiggle, identifying what base corresponds with each signal. This algorithm uses a neural network that passes signal data between nodes (i.e. “neurons”) to determine the RNA sequence. As this computational neural network trains on basecalling data, it better recognizes patterns in the basecalling and can more accurately predict sequences.
Historically, nanopore sequencing basecalling was less accurate than basecalling with short-read sequencing. However, direct RNA sequencing has made major improvements in its technology and algorithms and is rapidly closing that gap.
In addition, direct RNA sequencing saves drugs developers time by revealing insights into their RNA sequence and quality control in a single sequencing experiment. Since it does not require PCR, reverse transcription, or other amplification, direct RNA sequencing can provide information on an entire RNA strand in a single run without extensive preparation before loading the RNA into the nanopore. Full-length reads allow drug developers to validate multiple critical quality attributes at once, further speeding up their characterization.
Efficient, multi-attribute basecalling at Eclipsebio
At Eclipsebio, we offer the benefits of direct RNA sequencing in our eSTRAND RNA QC assay. eSTRAND RNA QC reads full-length RNA sequences to provide multi-dimensional quality control insights into critical quality attributes required by regulatory organizations. The assay can reveal these multi-attribute quality control insights with a single RNA nanopore run. For example, eSTRAND RNA QC can simultaneously measure RNA identity, poly(A) tail length, capping efficiency, fragmentation hotspots, and double-stranded RNA.
By directly measuring these critical quality attributes, drug developers can identify regions of their RNA to optimize to improve quality and meet regulatory standards, helping them get their RNA therapeutic to the clinic faster.
Interested in how direct RNA sequencing basecalling can help you gain multi-attribute characterization insights into your RNA? Contact Eclipsebio today to get started.
References
Base calling. 2025. Illumina. https://support-docs.illumina.com/IN/iSeq100/Content/IN/iSeq/BaseCalling_fISQ.htm
Bouchot et al. 2014. Chapter 14 - Advances in machine learning for processing and comparison of metagenomic data. Computational Systems Biology. doi: 10.1016/B978-0-12-405926-9.00014-9
How basecalling works. 2026. Oxford Nanopore Technologies. https://nanoporetech.com/platform/technology/basecalling
Ledergerber and Dessimoz. 2011. Base-calling for next-generation sequencing platforms. Briefings in Bioinformatics. doi: 10.1093/bib/bbq077
Nanopore basecalling. 2026. CD Genomics. https://www.cd-genomics.com/longseq/resource-nanopore-basecalling.html
Latest eBlogs
Basecalling: A method to characterize multiple critical quality attributes
Basecalling directly determines what nucleotide bases are in an RNA sequence, revealing actionable characterization insights.
From promise to proof: Reflections on the 6th mRNA-Based Therapeutics Summit
This year's mRNA-Based Therapeutics Summit proved that RNA is a platform, but drug developers must close analytical gaps to see how their constructs work.