Gene Model Annotations for Drosophila melanogaster: Impact of High-Throughput Data.

Gene Model Annotations for Drosophila melanogaster: Impact of High-Throughput Data.
复制标题

DOI:
10.1534/g3.115.018929
复制
发表时间:
2015-06-24
期刊:
G3 (Bethesda, Md.)
影响因子:
--
通讯作者:
FlyBase Consortium
FlyBase Consortium
中科院分区:
其他
文献类型:
--
作者:
Matthews BB;Dos Santos G;Crosby MA;Emmert DB;St Pierre SE;Gramates LS;Zhou P;Schroeder AJ;Falls K;Strelets V;Russo SM;Gelbart WM;FlyBase Consortium

文献摘要

被引文献

相似文献

我们报告了果蝇FlyBase注释基因集的现状,并强调了基于高通量数据的改进。FlyBase注释的基因集完全由手动注释的基因模型组成,但某些类别的小型非编码RNA除外。所有的基因模型都已经使用来自高通量数据集的证据进行了审查,主要来自modENCODE项目。这些数据集包括RNA-Seq覆盖数据、RNA-Seq连接数据、转录起始位点谱和翻译终止密码子通读预测。开发了新的注释指南,以考虑高通量数据的使用。我们描述了如何将这一洪水般的新数据纳入数千个新的和修订的注释。FlyBase采用了从基因模型注释中排除低置信度和低频数据的理念;我们也不试图代表复杂和模块化组织基因的所有可能排列。这使我们能够产生一个高置信度,可管理的基因注释数据集,可在FlyBase(http:flybase.org)。新注释的有趣方面包括新基因(编码,非编码和反义),许多基因具有非常长的3′ UTR(高达15-18 kb)的替代转录本,以及男性特异性基因(约占所有注释基因模型的13%)与女性特异性基因(不到1%)数量的惊人错配。在测序的菌株中鉴定的假基因和突变的数量也显著增加。我们讨论了剩余的挑战,例如,识别功能性小多肽和检测替代翻译开始。
We report the current status of the FlyBase annotated gene set for Drosophila melanogaster and highlight improvements based on high-throughput data. The FlyBase annotated gene set consists entirely of manually annotated gene models, with the exception of some classes of small non-coding RNAs. All gene models have been reviewed using evidence from high-throughput datasets, primarily from the modENCODE project. These datasets include RNA-Seq coverage data, RNA-Seq junction data, transcription start site profiles, and translation stop-codon read-through predictions. New annotation guidelines were developed to take into account the use of the high-throughput data. We describe how this flood of new data was incorporated into thousands of new and revised annotations. FlyBase has adopted a philosophy of excluding low-confidence and low-frequency data from gene model annotations; we also do not attempt to represent all possible permutations for complex and modularly organized genes. This has allowed us to produce a high-confidence, manageable gene annotation dataset that is available at FlyBase (http://flybase.org). Interesting aspects of new annotations include new genes (coding, non-coding, and antisense), many genes with alternative transcripts with very long 3′ UTRs (up to 15–18 kb), and a stunning mismatch in the number of male-specific genes (approximately 13% of all annotated gene models) vs. female-specific genes (less than 1%). The number of identified pseudogenes and mutations in the sequenced strain also increased significantly. We discuss remaining challenges, for instance, identification of functional small polypeptides and detection of alternative translation starts.