课题基金 / 基金详情

SCISIPBIO: A data-science approach to evaluating the likelihood of fraud and error in published studies

SCISIPBIO: A data-science approach to evaluating the likelihood of fraud and error in published studies
SCISIPBIO:一种评估已发表研究中欺诈和错误可能性的数据科学方法
批准号:
1956338
负责人:
Luis Amaral
金额:
$35.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-09-01 至 2022-08-31

项目摘要

项目成果

Luis Amaral的其他基金

相似基金

相关文献

中文摘要
翻译
科学文献有几个重要的作用。在科学领域,它可以为未来的研究提供信息,可以为新发现铺平道路,并指导个别科学家的未来计划以及他们如何度过自己的时间和事业。在科学之外,科学文献也发挥着多种作用,例如为政策提供信息或指导个别司法裁决。由于所有这些原因,维护科学文献的完整性对于科学家、广大公众以及最终公众对个别科学领域的看法至关重要。然而,对于科学家和科学期刊的编辑来说,识别不可信的科学文献仍然很困难。这个项目试图在可疑的科学手稿被公开之前识别出来,使用数据科学的方法,在这种方法中,我们捕捉到科学手稿的许多明显特征及其内容,以及作者的信息。该项目的成果将包括一个程序化和基于网络的界面,允许诸如政策制定者和科学期刊之类的第三方扫描手稿,寻找科学欺诈和错误的迹象。该项目将侧重于生物医学科学。该界面下的系统包括81个不同的数据库,这些数据库已被汇总、注释(例如,包含基因的化学和生物学特性),并通过出版物元数据(例如,参考文献、作者、资助)进行链接。这些数据将与关于欺诈和错误出版物的数据库(使用retractionwatch和人工管理的数据库)相匹配。欺诈性和非欺诈性出版物的特征将取决于这些数据库,以及基于基因和作者的网络特性的附加特征。该项目将采用不同的机器学习方法,如梯度增强和自动学习器,其性能将在样本外进行评估。第四,为了提高可解释性,更好地理解科学欺诈和错误,并可能提高模型的鲁棒性,该项目将规范和简化模型,以将其预测能力降低到一小部分信息。最后,该项目将创建一个基于rest的接口,该接口将允许从定制文稿中导入。提议的工作是独特的调节手稿对手稿的高度不同的属性,包括内容和世界领先的训练数据。这将为决策者和科学编辑提供一个数据驱动的工具,以便在可疑手稿进入已发表的科学记录之前识别它们。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Scientific literature servers several important roles. Within the sciences it can inform future research, can pave the way toward new discoveries, and guide the future plans of individual scientists and how they spend their own time and careers. Outside of the sciences, scientific literature too serves several roles, such as informing policies or guiding individual judicial decisions. For all of these reasons, maintaining the integrity of the scientific literature is of uttermost importance for scientists, the broad public, and ultimately the public’s perception of individual scientific fields. Yet, identifying non-trustworthy scientific literature even remains difficult for scientists and editors of scientific journals. This project seeks to identify suspicious scientific manuscripts before they are publicized, using a data-scientific approach in which we capture many distinct traits of scientific manuscripts and their content, as well as information about the authors. The outcome of the project will include a programmatic and web-based interface that allowed third parties such as policy makers and scientific journals to scan manuscripts for signs of scientific fraud and error. The project will focus on the biomedical sciences. The system beneath this interface includes 81 distinct databases that have been aggregated, annotated (e.g., with the chemical and biological properties of included genes), and linked through publication metadata (e.g., references, authorship, funding). These data will be matched with a database on fraudulent and erroneous publications (using retractionwatch and a manually curated database). Features of fraudulent and non-fraudulent publications will be conditioned on these databases, with additional features based on network-properties of genes and authors. The project will employ distinct machine learning approaches, such as Gradient Boosting and auto-learners, whose performance will be evaluated out-of-sample. Forth, to improve interpretability, and better understand scientific fraud and error, and possibly improve the robustness of models, the project will regularize and simplify the models to reduce their predictive capabilities to a small set of the information. Lastly, the project will create a REST-based interface that will allow the import from custom manuscripts. The proposed work is unique for conditioning manuscripts on highly distinct properties of manuscripts including content and world-leading training data. This will provide a data-driven tool for policy makers and scientific editors to identify suspicious manuscripts before they enter the published scientific record.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1371/journal.pbio.3001520
发表时间: 2022-01
期刊: PLoS biology
影响因子: 9.8
作者: [Stoeger T, Nunes Amaral LA]
通讯作者: Nunes Amaral LA
A cautionary tale from the machine scientist
机器科学家的警示故事
DOI: 10.1038/s42256-022-00491-7
发表时间: 2022
期刊: Nature Machine Intelligence
影响因子: 23.8
作者: [Amaral, Luís A.]
通讯作者: Amaral, Luís A.
DOI: 10.1093/nar/gkac1139
发表时间: 2022-11-28
期刊: NUCLEIC ACIDS RESEARCH
影响因子: 14.9
作者: [Byrne, Jennifer A., Park, Yasunori, Richardson, Reese A. K., Pathmendra, Pranujan, Sun, Mengyi, Stoeger, Thomas]
通讯作者: Stoeger, Thomas
DOI: 10.7554/elife.61981
发表时间: 2020-11-24
期刊: eLife
影响因子: 7.7
作者: [Stoeger T, Nunes Amaral LA]
通讯作者: Nunes Amaral LA
A1: Systematic Content Analysis of Litigation Events (SCALES) Open Knowledge Network to Enable Transparency and Access to Court Records
  • 批准号:
    2033604
  • 项目类别:
    Cooperative Agreement
  • 资助金额:
    $499.98万
  • 财政年份:
    2020
  • 负责人:
    Luis Amaral
  • 依托单位:
Convergence Accelerator Phase I (RAISE): Northwestern Open Access to Court Records Initiative
  • 批准号:
    1937123
  • 项目类别:
    Standard Grant
  • 资助金额:
    $100.0万
  • 财政年份:
    2019
  • 负责人:
    Luis Amaral
  • 依托单位:
TLS: Early prediction of the impact of research through large-scale analysis and modeling citation dynamics
  • 批准号:
    0830388
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2008
  • 负责人:
    Luis Amaral
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
复杂数据下半参数转换模型及其在老年慢性病发展中的应用研究
  • 批准号:
    72101261
  • 项目类别:
    青年科学基金项目(C类)
  • 资助金额:
    30.0万元
  • 批准年份:
    2021
  • 负责人:
    孙韬
  • 依托单位:
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    40万元
  • 批准年份:
    2020
  • 负责人:
    Vikrant Gupta
  • 依托单位: