Data availability of open T-cell receptor repertoire data, a systematic assessment

Data availability of open T-cell receptor repertoire data, a systematic assessment
复制标题

DOI:
10.3389/fsysb.2022.918792
复制
发表时间:
2022-04
期刊:
bioRxiv
影响因子:
--
通讯作者:
Yu-Ning Huang;Naresh Amrat Patel;Jay Himanshu Mehta;Srishti Ginjala;P. Brodin;C. Gray;Yesha M Patel;L. Cowell;A. Burkhardt;S. Mangul
Yu-Ning Huang;Naresh Amrat Patel;Jay Himanshu Mehta;Srishti Ginjala;P. Brodin;C. Gray;Yesha M Patel;L. Cowell;A. Burkhardt;S. Mangul
中科院分区:
其他
文献类型:
--
作者:
Yu-Ning Huang;Naresh Amrat Patel;Jay Himanshu Mehta;Srishti Ginjala;P. Brodin;C. Gray;Yesha M Patel;L. Cowell;A. Burkhardt;S. Mangul

文献摘要

相似文献

新一代测序技术的进步促进了免疫遗传学领域的发展,产生了大量的免疫基因组学数据。现代数据驱动的研究有能力通过对这些数据的二次分析来促进新的生物医学发现。因此,重要的是确保数据驱动的研究具有很大的可重复性和鲁棒性,以促进免疫基因组学数据的精确和准确的二次分析。在科学研究中,需要在设计和进行实验时采取严格的行为,特别是在科学和清晰的写作、报告和解释结果方面。还必须提供原始数据,使其可供查阅,并加以充分说明或注释,以促进今后对数据的重新分析。为了评估已发表的T细胞受体(TCR)库数据的数据可用性,我们检查了2006年至2022年期间134项TCR-Seq研究对应的11,918个TCR-Seq样本。在134项研究中,只有38.1%的研究在公共存储库中共享了公开的原始TCR-Seq数据。我们还发现数据可用性声明的存在与原始数据可用性的增加之间存在统计学显著相关性(p=0.014)。然而,46.8%的数据可用性声明的研究未能共享原始TCR-Seq数据。生物医学界迫切需要提高对促进科学研究中原始数据可用性的重要性的认识,并立即采取行动改善其原始数据的可用性,使更大的科学界能够对现有的免疫基因组学数据进行具有成本效益的二次分析。
The improvement of next-generation sequencing technologies has promoted the field of immunogenetics and produced numerous immunogenomics data. Modern data-driven research has the power to promote novel biomedical discoveries through secondary analysis of such data. Therefore, it is important to ensure data-driven research with great reproducibility and robustness for promoting a precise and accurate secondary analysis of the immunogenomics data. In scientific research, rigorous conduct in designing and conducting experiments is needed, specifically in scientific and articulate writing, reporting and interpreting results. It is also crucial to make raw data available, discoverable, and well described or annotated in order to promote future re-analysis of the data. In order to assess the data availability of published T cell receptor (TCR) repertoire data, we examined 11,918 TCR-Seq samples corresponding to 134 TCR-Seq studies ranging from 2006 to 2022. Among the 134 studies, only 38.1% had publicly available raw TCR-Seq data shared in public repositories. We also found a statistically significant association between the presence of data availability statements and the increase in raw data availability (p=0.014). Yet, 46.8% of studies with data availability statements failed to share the raw TCR-Seq data. There is a pressing need for the biomedical community to increase awareness of the importance of promoting raw data availability in scientific research and take immediate action to improve its raw data availability enabling cost-effective secondary analysis of existing immunogenomics data by the larger scientific community.