EPA's DSSTox database: History of development of a curated chemistry resource supporting computational toxicology research.

EPA's DSSTox database: History of development of a curated chemistry resource supporting computational toxicology research.
复制标题

DOI:
10.1016/j.comtox.2019.100096
复制
发表时间:
2019-11-01
期刊:
Computational toxicology (Amsterdam, Netherlands)
影响因子:
--
通讯作者:
Richard AM
Richard AM
中科院分区:
其他
文献类型:
--
作者:
Grulke CM;Williams AJ;Thillanadarajah I;Richard AM

文献摘要

被引文献

相似文献

美国环境保护署(EPA)的分布式结构可搜索毒性(DSSTox)数据库于2004年公开启动,目前超过875 K种物质,涵盖了EPA和环境研究人员感兴趣的数百个列表。自成立以来,DSSTox一直致力于解决公共领域的化学标识符错误和冲突,以实现为环境研究和监管社区的重要数据和列表分配准确的化学结构的目标。准确的结构-数据关联反过来又是支持危害和风险评估的基于结构的预测模型的必要输入。2014年,传统的手工管理的DSSTox_V1内容被迁移到MySQL数据模型中,现代化学信息学工具支持手动和自动管理流程,以提高效率。随后依次自动加载三个公共数据集的过滤部分:EPA的物质注册服务(SRS),国家医学图书馆的ChemID和PubChem。此过程受到每个物质的唯一映射标识符(即CAS RN、名称和结构)的关键要求的限制,拒绝任何两个标识符在数据集中或跨数据集中冲突的内容。这些被拒绝的内容突出了公共领域中冲突的程度,不准确的物质-结构ID映射,范围从12% (EPA SRS)到49% (ChemID和PubChem)。从每次自动加载中成功添加到DSSTox的物质被分配到五个qc_level之一,传达管理员对每个数据集的信心。这一过程大大扩展了DSSTox的内容,以便更好地覆盖环境科学家感兴趣的化学景观,同时保持对物质-结构-数据关联的准确性的关注。目前,DSSTox是EPA的CompTox化学品仪表盘的核心基础[https://comptox.epa.gov/dashboard]],该仪表盘为公众提供DSSTox内容,以支持EPA内部广泛的建模和研究活动,并越来越多地跨越计算毒理学领域。
The US Environmental Protection Agency’s (EPA) Distributed Structure-Searchable Toxicity (DSSTox) database, launched publicly in 2004, currently exceeds 875 K substances spanning hundreds of lists of interest to EPA and environmental researchers. From its inception, DSSTox has focused curation efforts on resolving chemical identifier errors and conflicts in the public domain towards the goal of assigning accurate chemical structures to data and lists of importance to the environmental research and regulatory community. Accurate structure-data associations, in turn, are necessary inputs to structure-based predictive models supporting hazard and risk assessments. In 2014, the legacy, manually curated DSSTox_V1 content was migrated to a MySQL data model, with modern cheminformatics tools supporting both manual and automated curation processes to increase efficiencies. This was followed by sequential auto-loads of filtered portions of three public datasets: EPA’s Substance Registry Services (SRS), the National Library of Medicine’s ChemID, and PubChem. This process was constrained by a key requirement of uniquely mapped identifiers (i.e., CAS RN, name and structure) for each substance, rejecting content where any two identifiers were conflicted either within or across datasets. This rejected content highlighted the degree of conflicting, inaccurate substance-structure ID mappings in the public domain, ranging from 12% (within EPA SRS) to 49% (across ChemID and PubChem). Substances successfully added to DSSTox from each auto-load were assigned to one of five qc_levels, conveying curator confidence in each dataset. This process enabled a significant expansion of DSSTox content to provide better coverage of the chemical landscape of interest to environmental scientists, while retaining focus on the accuracy of substance-structure-data associations. Currently, DSSTox serves as the core foundation of EPA’s CompTox Chemicals Dashboard [https://comptox.epa.gov/dashboard], which provides public access to DSSTox content in support of a broad range of modeling and research activities within EPA and, increasingly, across the field of computational toxicology.