课题基金 / 基金详情

Linking data with Identifiers.org

Linking data with Identifiers.org
将数据与 Identifiers.org 链接
批准号:
BB/K016946/1
负责人:
Henning Hermjakob
金额:
$15.22万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2013
资助国家:
英国
项目状态:
已结题
起止时间:
2013 至 --

项目摘要

项目成果

Henning Hermjakob的其他基金

相似基金

相关文献

中文摘要
翻译
用与其他知识来源的交叉引用来注释数据、生命科学数据集一直是非常重要的。这些元数据通常将有价值的信息从堆积如山的不可用数据中区分开来。随着系统生物学的出现,数据集的大小和复杂性将平衡从直接的人类交互转移到了自动计算机处理上。如果按照标准过程并使用受控词汇表对元数据进行编码,则会极大地方便此类操作。如果这些过程和词汇表在不同类型的数据之间共享,则可以对不同的数据集进行对齐、比较和集成。任何交叉引用的一个关键部分是它所指向的资源的标识符。该标识符必须是唯一的、常年的、可解析的和免费的。大多数数据提供者为他们自己的记录创建标识符;例如,‘9606’在分类中标识‘智人’,而‘22140103’在PubMed中标识有关标识网站的最新出版物。然而,这些标识符只在给定的数据集中是唯一的,所以在更广泛的上下文中考虑记录时,它们的用处是有限的。为了实现这一目的,它使用了MIRiam注册表(http://www.ebi.ac.uk/miriam/).)中记录的信息因此,这两个项目都是最终技术解决方案的不同部分。登记处生成的识别码利用数据提供者提供的登录号,但也包含有关它们所来自的集合的信息。所有标识符都是唯一的、可解析的和健壮的。它们允许个人或软件工具通过替代提供商直接访问Web上已识别的数据片段。虽然它是一个原型,但由于它满足了他们对常年交叉引用的需求,并消除了他们以前维护和保持不断变化的Web链接(或URL)的长长列表的需要。随着越来越多的社区意识到使用Identifiers.org URI的好处,新的需求和用例也出现了。这项提议旨在加强和扩大该资源提供的服务,以回应这些新的用户请求。我们将使资源更易于在自动化过程中使用,特别是在语义Web应用程序中。这涉及到以资源描述框架(RDF)格式提供注册表的内容,并提供用于查询目的的SPARQL端点。用户将能够微调识别符的解析方式,通过创建将记录他们的偏好的“配置文件”。该资源将使社区(更具体地说是数据提供者本身)能够参与维持登记处。这将通过数据提供者对其在登记处的记录的“所有权”制度进行。虽然我们目前已经建立了自动系统来检测过时的信息,但让实际的数据提供者参与维护将确保记录信息的质量更高,这意味着提供的服务质量更高。最后,我们将改进和扩展底层计算基础设施。通过将其部署到更多的数据中心,我们将为越来越多的用户提供更可靠的服务。由此产生的资源将提供一种方式,将使用相同URI注释的所有数据无缝链接到代表相同概念的所有数据,这是迈向数据集成的关键一步。通过在这些数据集之间提供语义粘合剂,Identifiers.org将促进本地或通过语义网的数据检索、比较和集成。它还将促进对综合数据集的推理,并导致生物医学领域可能出现新的、可能自动的发现。
英文摘要
Annotating data, life science datasets with cross-references to other sources of knowledge has always be very important. These metadata are often what separate valuable information from heaps of unusable data. With the advent of systems biology, the size and complexity of datasets shifted the balance from direct human interaction to automated computer processing. Such operations are greatly facilitated if the metadata is encoded following standard procedures and using controlled vocabularies. If those procedures and vocabularies are shared between different types of data, it becomes possible to align, compare and integrate different datasets. A key part of any cross-reference is the identifier of the resource it points to. This identifier must be unique, perennial, resolvable and free. Most data providers create identifiers for their own records; for example '9606' identifies 'Homo sapiens' in the Taxonomy, and '22140103' identifies the latest publication about Identifiers.org in PubMed. However, those identifiers are only unique within a given dataset so their usefulness is limited when considering records in a wider context.Identifiers.org provides such global identifiers, and resolves them to the relevant dataset. In order to achieve this purpose, it uses the information recorded in the MIRIAM Registry (http://www.ebi.ac.uk/miriam/). Therefore both projects provide a distinct part of the final technical solution. Identifiers generated with the Registry make use of the accession numbers supplied by data providers, but also contain information about the collection they come from. All identifiers are unique, resolvable and robust. They allow persons or software tools to directly access the identified pieces of data on the web, via alternative providers. Although a prototype, Identifiers.org has been adopted by a number of communities and projects, as it fulfils their need for perennial cross-references and removes their previous need for maintaining and keeping up to date long lists of ever changing web links (or URLs).As more and more communities realise the benefits of using Identifiers.org URIs, new needs and use cases have appeared. This proposal seeks to strengthen and extend the services provided by the resource in order to respond to those new user requests. We will make the resource easier to use in automated procedures, specially for semantic web applications. This involves providing the content of the Registry in Resource Description Framework (RDF) format and supply a SPARQL endpoint for query purposes. Users will be able to fine-tuned the way identifiers are resolved, via the creation of 'profiles', that will record their preferences. The resource will allow the communities (more specially the data providers themselves) to get involved in the maintenance of the Registry. This will take place via a system of "ownership" by data providers of their record in the Registry. Although we currently have automatic systems in place to detect obsolete information, having the actual data providers contributing to the maintenance would ensure better quality of the recorded information, meaning a better quality of the services provided. Finally we will improve and extend the underlying computing infrastructure. By deploying it in more more data centres, we will provide more reliable services to an ever growing number of users.The resulting resource will provide a way to seamlessly link all data annotated with the same URI to represent the same concept, a key step towards data integration. By providing a semantic glue between those datasets, Identifiers.org will facilitate data retrieval, comparison, integration, locally or through the semantic web. It will also facilitate the reasoning on the integrated datasets and lead to new, possibly automated discovery in the biomedical domain.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1186/2041-1480-5-5
发表时间: 2014-02-05
期刊: Journal of biomedical semantics
影响因子: 1.9
作者: [Katayama T, Wilkinson MD, Aoki-Kinoshita KF, Kawashima S, Yamamoto Y, Yamaguchi A, Okamoto S, Kawano S, Kim JD, Wang Y, Wu H, Kano Y, Ono H, Bono H, Kocbek S, Aerts J, Akune Y, Antezana E, Arakawa K, Aranda B, Baran J, Bolleman J, Bonnal RJ, Buttigieg PL, Campbell MP, Chen YA, Chiba H, Cock PJ, Cohen KB, Constantin A, Duck G, Dumontier M, Fujisawa T, Fujiwara T, Goto N, Hoehndorf R, Igarashi Y, Itaya H, Ito M, Iwasaki W, Kalaš M, Katoda T, Kim T, Kokubu A, Komiyama Y, Kotera M, Laibe C, Lapp H, Lütteke T, Marshall MS, Mori T, Mori H, Morita M, Murakami K, Nakao M, Narimatsu H, Nishide H, Nishimura Y, Nystrom-Persson J, Ogishima S, Okamura Y, Okuda S, Oshita K, Packer NH, Prins P, Ranzinger R, Rocca-Serra P, Sansone S, Sawaki H, Shin SH, Splendiani A, Strozzi F, Tadaka S, Toukach P, Uchiyama I, Umezaki M, Vos R, Whetzel PL, Yamada I, Yamasaki C, Yamashita R, York WS, Zmasek CM, Kawamoto S, Takagi T]
通讯作者: Takagi T
DOI: 10.1002/psp4.3
发表时间: 2015-02
期刊: CPT-PHARMACOMETRICS & SYSTEMS PHARMACOLOGY
影响因子: 3.5
作者: [Juty, N, Ali, R, Glont, M, Keating, S, Rodriguez, N, Swat, M J, Wimalaratne, S M, Hermjakob, H, Le Novere, N, Laibe, C, Chelliah, V]
通讯作者: Chelliah, V
DOI: 10.1093/nar/gku1181
发表时间: 2015-01
期刊: Nucleic acids research
影响因子: 14.9
作者: [Chelliah V, Juty N, Ajmera I, Ali R, Dumousseau M, Glont M, Hucka M, Jalowicki G, Keating S, Knight-Schrijver V, Lloret-Villas A, Natarajan KN, Pettit JB, Rodriguez N, Schubert M, Wimalaratne SM, Zhao Y, Hermjakob H, Le Novère N, Laibe C]
通讯作者: Laibe C
DOI: 10.1186/s12918-014-0091-5
发表时间: 2014-08-15
期刊: BMC systems biology
影响因子: --
作者: [Wimalaratne SM, Grenon P, Hermjakob H, Le Novère N, Laibe C]
通讯作者: Laibe C
2021BBSRC-NSF/BIO UniPlex - Genome-Wide Protein Complex Prediction and Validation
  • 批准号:
    BB/X002179/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $55.79万
  • 财政年份:
    2023
  • 负责人:
    Henning Hermjakob
  • 依托单位:
Japan Partnering Award: Establishment of an Integrative proteomics bioinformatics platform to enable novel analysis approaches
  • 批准号:
    BB/N022440/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $3.88万
  • 财政年份:
    2016
  • 负责人:
    Henning Hermjakob
  • 依托单位:
China Partnering Award: Proteomics Data Exchange
  • 批准号:
    BB/N022432/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $3.9万
  • 财政年份:
    2016
  • 负责人:
    Henning Hermjakob
  • 依托单位:
MultiMod, flexible management for multi-scale multi-approach models in biology
  • 批准号:
    BB/N019482/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $41.67万
  • 财政年份:
    2016
  • 负责人:
    Henning Hermjakob
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
复杂数据下半参数转换模型及其在老年慢性病发展中的应用研究
  • 批准号:
    72101261
  • 项目类别:
    青年科学基金项目(C类)
  • 资助金额:
    30.0万元
  • 批准年份:
    2021
  • 负责人:
    孙韬
  • 依托单位:
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位: