Long read sequencing for transcriptome characterization --- (how) can we quantify transcripts accurately with long reads
Long-read sequencing has been widely adopted in transcriptomics research, particularly for identifying novel and complex transcriptomic events. As the cost and throughput of long-read sequencing have been improved, the applications to quantitative transcriptome analysis are emerging. Here I will present the development of the "K-value", which is designed to: (1) quantify the influence of gene isoform structural complexity on expression estimation via data deconvolution; and (2) rigorously demonstrate the advantages of long reads over short reads in isoform quantification through mathematical modeling, simulations, experimental validation, and large-scale public datasets. I also will present the software miniQuant that integrates short reads and long reads in a gene- and data-specific manner to achieve better gene isoform quantification over long reads alone and short reads alone. The proof-of-concept applications include the discoveries of gene isoform switching during stem cell differentiation and unique pattern of zygotic genome activation of transposable elements.
