On the identifiability of the isoform deconvolution problem: application to select the proper fragment length in an RNA-seq library.

On the identifiability of the isoform deconvolution problem: application to select the proper fragment length in an RNA-seq library.
复制标题

关于异构体反卷积问题的可识别性:在 RNA-seq 文库中选择合适片段长度的应用。

DOI:
10.1093/bioinformatics/btab873
复制
发表时间:
2022
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Rubio,Angel
Rubio,Angel
中科院分区:
--
文献类型:
--
作者:
Ferrer-Bonsoms,JuanA;Morales,Xabier;Afshar,PegahT;Wong,WingH;Rubio,Angel

文献摘要

相似文献

MotivationIsoform deconvolution is an NP-hard problem. The accuracy of the proposed solutions is far from perfect. At present, it is not known if gene structure and isoform concentration can be uniquely inferred given paired-end reads, and there is no objective method to select the fragment length to improve the number of identifiable genes. Different pieces of evidence suggest that the optimal fragment length is gene-dependent, stressing the need for a method that selects the fragment length according to a reasonable trade-off across all the genes in the whole genome.ResultsA gene is considered to be identifiable if it is possible to get both the structure and concentration of its transcripts univocally. Here, we present a method to state the identifiability of this deconvolution problem. Assuming a given transcriptome and that the coverage is sufficient to interrogate all junction reads of the transcripts, this method states whether or not a gene is identifiable given the read length and fragment length distribution. Applying this method using different read and fragment length combinations, the optimal average fragment length for the human transcriptome is around 400–600 nt for coding genes and 150–200 nt for long non-coding RNAs. The optimal read length is the largest one that fits in the fragment length. It is also discussed the potential profit of combining several libraries to reconstruct the transcriptome. Combining two libraries of very different fragment lengths results in a significant improvement in gene identifiability.Availability and implementationCode is available in GitHub (https://github.com/JFerrer-B/transcriptome-identifiability).Supplementary informationSupplementary data are available atBioinformaticsonline.