Deep forecasting of translational impact in medical research.

Deep forecasting of translational impact in medical research.
复制标题

DOI:
10.1016/j.patter.2022.100483
复制
发表时间:
2022-05-13
期刊:
影响因子:
6.5
通讯作者:
Nachev, Parashkev
Nachev, Parashkev
中科院分区:
其他
文献类型:
--
作者:
Nelson, Amy P. K.;Gray, Robert J.;Ruffle, James K.;Watkins, Henry C.;Herron, Daniel;Sorros, Nick;Mikhailov, Danil;Cardoso, M. Jorge;Ourselin, Sebastien;McNally, Nick;Williams, Bryan;Rees, Geraint E.;Nachev, Parashkev

文献摘要

参考文献

相似文献

生物医学研究的价值--每年1.7万亿美元的投资--最终取决于其下游的、现实世界的影响,而这种影响的可预测性,从简单的引用指标来看,仍然是无法量化的。在这里,我们试图确定未来现实世界的预防的比较可预测性,如列入专利,指南,或政策文件,从标题/摘要级内容与引用和元数据单独的复杂模型。我们使用Microsoft Academic Graph从1990年至2019年捕获的整个生物医学研究语料库,提前量化了主要领域的样本预测性能,其中包括4330万篇论文。我们发现,引文只是适度的翻译影响的预测。相比之下,标题、摘要和元数据的高维模型表现出高保真度(接受者工作曲线下面积[AUROC] > 0.9),在时间和领域上都是通用的,并且可以转移到诺贝尔奖获得者的论文识别上。我们认为,基于内容的影响模型是上级传统的,基于引用的措施,并维持一个更强有力的证据为基础的要求翻译潜力的客观测量。生物医学论文内容的深度学习模型可以准确地预测翻译深度内容模型的性能大大优于传统的引用指标在专利包含转移方面训练的模型可以预测诺贝尔奖前的论文科学政策可能更好地了解深度内容模型而不是引用科学活动与现实世界影响的关系很难描述,甚至更难量化。通过分析1990年至2019年的4330万篇生物医学论文,我们发现,出版物、标题和摘要内容的深度学习模型可以预测专利、指南或政策文件中是否包含科学论文。我们发现,这些模型中最好的,结合最丰富的信息,大大优于传统的论文成功的指标,每年的引用,并转移到预测诺贝尔奖前的论文的任务。如果对科学的转化潜力的判断是基于客观的指标,那么论文内容的复杂模型应该优先于引用。我们的方法自然可以扩展到更丰富的科学内容和不同的影响措施。其更广泛的应用可以最大限度地扩大生物医学领域内外科学活动的实际惠益。通过分析1990年至2019年的4330万篇生物医学论文,我们发现,出版物标题和摘要内容的深度学习模型可以预测专利、指南或政策文件中的内容,其准确度远远高于引用指标。如果对科学的转化潜力的判断是基于客观的指标,那么论文内容的复杂模型应该优先于引用。
The value of biomedical research—a $1.7 trillion annual investment—is ultimately determined by its downstream, real-world impact, whose predictability from simple citation metrics remains unquantified. Here we sought to determine the comparative predictability of future real-world translation—as indexed by inclusion in patents, guidelines, or policy documents—from complex models of title/abstract-level content versus citations and metadata alone. We quantify predictive performance out of sample, ahead of time, across major domains, using the entire corpus of biomedical research captured by Microsoft Academic Graph from 1990–2019, encompassing 43.3 million papers. We show that citations are only moderately predictive of translational impact. In contrast, high-dimensional models of titles, abstracts, and metadata exhibit high fidelity (area under the receiver operating curve [AUROC] > 0.9), generalize across time and domain, and transfer to recognizing papers of Nobel laureates. We argue that content-based impact models are superior to conventional, citation-based measures and sustain a stronger evidence-based claim to the objective measurement of translational potential. Deep learning models of biomedical paper content can accurately predict translation Deep content models substantially outperform traditional citation metrics Models trained on patent inclusion transfer to predicting Nobel Prize-preceding papers Science policy is potentially better informed by deep content models than by citations The relationship of scientific activity to real-world impact is hard to describe and even harder to quantify. Analyzing 43.3 million biomedical papers from 1990–2019, we show that deep learning models of publication, title, and abstract content can predict inclusion of a scientific paper in a patent, guideline, or policy document. We show that the best of these models, incorporating the richest information, substantially outperforms traditional metrics of paper success—citations per year—and transfers to the task of predicting Nobel Prize-preceding papers. If judgments of the translational potential of science are to be based on objective metrics, then complex models of paper content should be preferred over citations. Our approach is naturally extensible to richer scientific content and diverse measures of impact. Its wider application could maximize the real-world benefits of scientific activity in the biomedical realm and beyond. Analyzing 43.3 million biomedical papers from 1990–2019, we show that deep learning models of publication title and abstract content can predict inclusion in a patent, guideline, or policy document with far greater fidelity than citation metrics alone. If judgments of the translational potential of science are to be based on objective metrics, then complex models of paper content should be preferred over citations.
DOI: 10.1007/s11192-016-2237-2
发表时间: 2017
期刊: Scientometrics
影响因子: 3.9
作者:
Haunschild R;Bornmann L
通讯作者: Bornmann L
DOI: 10.1371/journal.pbio.3000416
发表时间: 2019-10-01
期刊: PLOS BIOLOGY
影响因子: 9.8
作者:
Hutchins, B. Ian;Davis, Matthew T.;Santangelo, George M.
通讯作者: Santangelo, George M.
DOI: 10.1093/bioinformatics/btz682
发表时间: 2020-02-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lee J;Yoon W;Kim S;Kim D;Kim S;So CH;Kang J
通讯作者: Kang J
DOI: 10.1007/s11192-020-03744-7
发表时间: 2021
期刊: Scientometrics
影响因子: 3.9
作者:
Ebadi A;Xi P;Tremblay S;Spencer B;Pall R;Wong A
通讯作者: Wong A
DOI: 10.2196/jmir.2177
发表时间: 2012-09-01
影响因子: 7.4
作者:
El Emam, Khaled;Arbuckle, Luk;Anderson, Kevin
通讯作者: Anderson, Kevin