EukProt: a database of genome-scale predicted proteins across the diversity of eukaryotes

EukProt: a database of genome-scale predicted proteins across the diversity of eukaryotes
复制标题

DOI:
10.1101/2020.06.30.180687
复制
发表时间:
2020-07
期刊:
bioRxiv
影响因子:
--
通讯作者:
D. Richter;C. Berney;Jürgen F. H. Strassert;Y. Poh;Emily K. Herman;Sergio A. Muñoz-Gómez;Jeremy G. Wideman;Fabien Burki;C. de Vargas
D. Richter;C. Berney;Jürgen F. H. Strassert;Y. Poh;Emily K. Herman;Sergio A. Muñoz-Gómez;Jeremy G. Wideman;Fabien Burki;C. de Vargas
中科院分区:
其他
文献类型:
--
作者:
D. Richter;C. Berney;Jürgen F. H. Strassert;Y. Poh;Emily K. Herman;Sergio A. Muñoz-Gómez;Jeremy G. Wideman;Fabien Burki;C. de Vargas

文献摘要

相似文献

EukProt 是一个已发表和公开的预测蛋白质集的数据库,这些蛋白质集被选择来代表真核生物多样性的广度,目前包括来自所有主要超类群的 993 个物种以及孤儿类群。该数据库的目标是为整个真核生物领域的基因研究提供单一、便捷的资源,例如系统发育组学和基因家族进化。每个物种都被放置在 UniEuk 分类框架内,以便于下游分析,每个数据集都与一个唯一的、持久的标识符相关联,以促进分析之间的比较和复制。该数据库定期更新,所有版本都将永久存储并通过 FigShare 提供。当前版本有许多更新,特别是“比较集”(TCS),这是一个简化的分类集,具有较高的估计完整性,同时保持了相当大的系统发育广度,其中包括 196 个预测蛋白质组。我们邀请社区为后续版本中包含的新数据集和新注释功能提供建议,目的是建立一个协作资源,促进了解真核生物多样性和多样化的研究。
EukProt is a database of published and publicly available predicted protein sets selected to represent the breadth of eukaryotic diversity, currently including 993 species from all major supergroups as well as orphan taxa. The goal of the database is to provide a single, convenient resource for gene-based research across the spectrum of eukaryotic life, such as phylogenomics and gene family evolution. Each species is placed within the UniEuk taxonomic framework in order to facilitate downstream analyses, and each data set is associated with a unique, persistent identifier to facilitate comparison and replication among analyses. The database is regularly updated, and all versions will be permanently stored and made available via FigShare. The current version has a number of updates, notably ‘The Comparative Set’ (TCS), a reduced taxonomic set with high estimated completeness while maintaining a substantial phylogenetic breadth, which comprises 196 predicted proteomes. We invite the community to provide suggestions for new data sets and new annotation features to be included in subsequent versions, with the goal of building a collaborative resource that will promote research to understand eukaryotic diversity and diversification.