课题基金 / 基金详情

ScienceLinker: A Framework for Finding, Linking, and Enriching Social Science Linked Data

ScienceLinker: A Framework for Finding, Linking, and Enriching Social Science Linked Data
ScienceLinker:查找、链接和丰富社会科学关联数据的框架
批准号:
404417453
负责人:
Dr. Benjamin Zapilko
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research data and software (Scientific Library Services and Information Systems)
财政年份:
2019
资助国家:
德国
项目状态:
已结题
起止时间:
2018-12-31 至 2022-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
应用经验研究中的科学家通常搜索数据集,特别是数据集中的度量(例如,社会科学研究中的变量),使他们能够调查其特定的研究兴趣。这些数据集用于多种目的,如回答特定的研究问题、根据不同的数据集复制特定的发现或将其与另一个数据集合并,以增加分析的可能性或减少遗漏的值。然而,找到合适的数据和衡量标准来支持自己的假设是一项具有挑战性的任务。在许多情况下,研究人员将能够在研究数据中心找到所需的数据。关于网上可用的大量数据(开放数据运动产生的),可能还有其他有趣的数据集可用,但不是由研究数据中心等有组织的基础设施提供的。此外,仍需进行手工工作,以便将发现的数据集用于相互关联,例如,以便用发现的数据和元数据中的其他内容丰富自己的数据集,也用于以后在期刊、自我存档平台或网络上发表。Science Linker项目激发了两种方法来应对这些挑战:(1)开发方法来识别在网络上以链接开放数据形式发布的数据集,这些数据集的内容与其内容兼容,并提供适当的质量;(2)将语义网技术应用于数据的使用,例如用于链接、丰富和发布。通过在可能的情况下应用广泛的自动化,这些技术将对非域用户可用。所开发的框架旨在通过以下五个步骤指导用户(例如,负责发布数据的数据提供者的雇员或正在寻找数据集以便用附加元数据完成其数据集的科学家):自动识别作为链接开放数据发布的一组相关数据集;从兼容性和质量方面评估数据集;将数据集中引用的实体与所确定的数据集相联系;通过应用一组实体类型特定规则来推断关于实体的额外信息,从而丰富数据集;以及将丰富的数据集作为链接数据或通过进一步出版的方式在自我存档平台上对出版物进行预处理。这一项目的调查和发展将保持一般性,以便能够在其他领域应用该框架。对于社会科学来说,潜在的相关链接数据源可能既不是科学的,也不是来自社会科学领域的,例如DBpedia或Geoname。为了使Science Linker框架也可以在中立的环境中执行,我们将其集成到ISI开发的已建立的数据集成平台Karma中。
英文摘要
Scientists in applied empirical research are typically searching for datasets and, in particular, measures within the datasets (e.g. variables in the case of the social science research) which allow them to investigate their specific research interest. These datasets are used for multiple purposes like for answering a particular research question, replicating a specific finding based on a different dataset or merging it with another dataset in order to increase possibilities for analysis or to reduce missing values. However, finding suitable data and measures for the support of one’s own hypothesis is a challenging task. In a lot of cases, a researcher will be able to find the desired data at a research data centre. Regarding the mass of data available on the web (resulting from the Open Data movement) additional interesting datasets are likely be available but are not provided by organized infrastructures like research data centres. Additionally, manual effort still has to be done to use the found datasets for interlinking, e.g. in order to enrich own datasets with additional content from the found data and metadata, also for a later publication in a journal, a self-archiving platform or on the web. The project ScienceLinker motivates two approaches for these challenges: (1) to develop methods to identify datasets published as Linked Open Data on the web that are compatible by their content and also provide an appropriate quality; (2) to apply Semantic Web technologies to use of the data e.g. for linking, enrichment and publishing. These techniques will be made usable for non-domain users by applying extensive automation when possible. The developed framework aims to guide the user (e.g. an employee of a data provider who is responsible for the publication of data or a scientist who is seeking datasets in order to complete his dataset with additional metadata) through the following five steps: the automatic identification of a set of related datasets published as Linked Open Data; the assessment of a dataset in terms of compatibility and quality; the linking of entities referenced in the dataset to the identified datasets; the enrichment of the dataset by applying a set of entity-type-specific rules to infer additional information about the entities also via non-identity links; and the preprocessing of the enriched dataset for a publication in self-archiving platforms, as Linked Data or via further publication ways.The investigations and developments in this project will be kept generic in order to allow an application of the framework in other domains. For the social sciences, potential related Linked Data sources may neither be scientific nor from the social science domain at all like e.g. DBpedia or Geonames. In order that the ScienceLinker framework can also be executed in a neutral environment, we will integrate it into the established data integration platform Karma which has been developed at ISI.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金