EAGER: Using Search Engines to Track Impact of Unsung Heroes of Big Data Revolution, Data Creators
EAGER: Using Search Engines to Track Impact of Unsung Heroes of Big Data Revolution, Data Creators
批准号:
1565233
负责人:
Adam Godzik
金额:
$30.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-10-01 至 2019-05-31
中文摘要
在复杂数据的洪流中,高效的数据交换机制日益成为科学(和社会)的核心。数据生产的速度及其复杂性意味着大量数据往往没有得到数据生产者的充分分析,由于其迅速传播,广大社区正在参与其分析。然而,随着注意力转移到数据集成者和分析者身上,这种模式存在忽视原始数据创建者的输入的风险。目前的信息传播和科学信用分配模式是基于同行评议的出版物和以引文形式正式承认第三方贡献的基础上的,偏向于参与知识创造和传播的后期阶段的知名科学家和研究中心。我们建议使用无偏见的互联网搜索来识别研究文献中对数据集和资源的未授权使用,允许参与数据创建早期阶段的数据创建者和研究人员为他们的工作声称自己的功劳。PI使用机器学习和文本挖掘技术,试图从通用搜索引擎的嘈杂结果中提取相关信息,并开发易于使用的界面,供公众使用这些资源,以补充官方文献计量资源。收受和引用偏差对资金和出版物中心以外的研究人员的职业生涯有重大影响,这些领域通常也是研究力量更加多样化的地方。首先,同行评议中被拒绝的可能性明显偏向不太有名的科学家和研究密集程度较低的机构的科学家。这些偏见不太可能影响数据创建,因为数据库通常在没有同行审查的情况下接受数据,并且数据的价值可以通过其使用来衡量。同样的偏见影响着引用的数量,人们倾向于引用更著名的知名科学家,或者引用经常被邀请撰写评论的评论,而只有顶尖的科学家才会被邀请撰写评论。因此,出版物和引文都严重偏向已经得到认可的科学家,这些科学家与普通科学家相比,无论是从个人角度还是从他们工作的机构来看,都代表了较少的多样性群体。这种偏见影响了年轻科学家在最佳研究机构的紧密合作网络之外工作的职业和获得拨款的能力,并创造了一个典型的富人变得越来越富,穷人变得越来越穷的循环。基于互联网的新的信息交流范式已经对研究人员产生了深远的影响?能够让他们的结果出现在显眼的公众视野中。这项建议旨在通过进一步减轻发表和引用偏见的方法,不是通过直接解决这个问题,而是通过开发更多评估对科学的贡献的开放方式。
英文摘要
Efficient mechanisms of data exchange are increasingly central to science (and society) in the midst of the deluge of complex data. The pace of data production and its complexity mean that a large amount of data is often not adequately analyzed by the data producers, and thanks to its rapid dissemination, the broad community is participating in its analysis. This model, however, carries a risk of neglecting the input of the original data creators, as attention shifts to data integrators and analyzers. The current paradigm of information dissemination and assigning credit in science, based on peer-reviewed publications and formal acknowledgment of third-party contributions in the form of citations, is biased toward high-profile, well-known scientists and research centers who participate in the latter stages of knowledge creation and dissemination. We propose to use unbiased internet searches to identify uncredited use of datasets and resources in research literature, allowing data creators and researchers participating in early stages of data creation to claim credit for their work. Using machine-learning and text-mining techniques, the PI seeks to extract relevant information from noisy results of general-purpose search engines and develop easy-to-use interfaces for public use of such resources to supplement official bibliometric resources.Acceptance and citation biases have a significant impact on careers of researchers outside the central foci of funding and publications, which are also typically places with more-diverse research forces. First, the probability of rejection in peer review is significantly biased against less-famous scientists and those at less-research-intensive institutions. These biases are less likely to affect data creation as databases typically accept data without peer review and the value of data can be measured by its use. The same biases affect number of citations, where people tend to cite more-famous, established scientists, or cite reviews that are often invited and only leading scientists would have been invited to write the review. As a result, both publications and citations are heavily biased toward already-recognized scientists that represent less-diverse populations, both in personal terms and in terms of the institutions where they work, as compared to the general population of scientists. Such biases affect careers and ability to obtaining grant funding for young scientist operating outside of the tight collaboration networks at best research institutions and creates a classical rich getting richer and poor getting poorer loop. The new internet-based information exchange paradigm has already had a profound effect on researchers? ability to getting their results in prominent, public view. This proposal aims at approaches that would further alleviate publication and citation bias, not by addressing it directly, but by developing more openways of evaluating contributions to science.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Using Search Engines to Track Impact of Unsung Heroes of Big Data Revolution, Data Creators
-
批准号:1931895
-
项目类别:Standard Grant
-
资助金额:$15.45万
-
财政年份:2018
-
负责人:Adam Godzik
-
依托单位:
I-Corps: Market Research, Customer Interviews, and Customer Discovery for Novel Cancer Biomarkers
-
批准号:1559647
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2015
-
负责人:Adam Godzik
-
依托单位:
Flexible Protein Structure Alignment Program and Server
-
批准号:0349600
-
项目类别:Standard Grant
-
资助金额:$60.26万
-
财政年份:2004
-
负责人:Adam Godzik
-
依托单位:
Conservation of Interaction Patterns in Protein Families
-
批准号:9506278
-
项目类别:Continuing Grant
-
资助金额:$30.0万
-
财政年份:1995
-
负责人:Adam Godzik
-
依托单位:
国内基金
海外基金
Capture and Release of Droplets Using Advanced Materials for High Technology Applications
-
批准号:52073127
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2020
-
负责人:Alidad Amirfazli
-
依托单位:
Molecular Interaction Reconstruction of Rheumatoid Arthritis Therapies Using Clinical Data
-
批准号:31070748
-
项目类别:面上项目
-
资助金额:34.0万元
-
批准年份:2010
-
负责人:Christine Nardini
-
依托单位: