Machine learning could improve innovation policy

Machine learning could improve innovation policy
复制标题

机器学习可以改善创新政策

DOI:
10.1038/s42256-020-0155-8
复制
发表时间:
2020
影响因子:
23.8
通讯作者:
Teodoridis, Florenta
Teodoridis, Florenta
中科院分区:
计算机科学1区
文献类型:
--
作者:
Furman, Jeffrey L.;Teodoridis, Florenta

文献摘要

参考文献

被引文献

相似文献

致编辑——半个多世纪以来,了解影响创新速度和方向的因素一直是科学与创新研究的中心目标。然而,在理解推动创新速度的因素方面,我们取得的进展远远多于对创新方向的理解——科学家或研究机构在给定时间点(多样性)和一段时间内(轨迹)处理的一组研究主题。经济学家早就认识到,市场对创新的激励不足,因为很难充分分配创新投资的回报,特别是当新的创新建立在旧的创新之上时1。类似的特征抑制了对创新多样性的投资2、3,从而可能影响其发展轨迹。例如,对环保技术的投资可能因化石燃料技术的初步成功而受到抑制,因此对替代技术的投资稀缺3。这些见解表明,需要更多的实证研究来揭示哪些因素影响创新方向。这样的研究具有挑战性。例如,估计创新方向的变化需要定义研究轨迹的边界。然而,这造成了一个悖论,因为研究轨迹的边界是需要估计的核心未知数的一部分。机器学习 (ML) 的最新发展有可能解决这一限制,帮助研究人员通过量化各种研究主题及其之间的距离来推断知识空间的结构。剩下的挑战是使机器学习算法适应因果关系的研究,因为创新方向的研究感兴趣的是识别研究主题的潜在分类,以便识别因果性改变这种结构的因素。这与机器学习算法的核心预测目的有很大的区别。我们呼吁研究人员加紧努力,使机器学习算法适应创新方向的研究。新兴的工作可以分为两类:(1) 由书目数据服务开发的现成的基于 ML 的分类模式(例如,国家医学图书馆的 PubMed 相关文章算法或 Microsoft Academy Graph 的分类模式 4)和 (2) 定制算法,可提供对文本语料库之间相似性的更精细数据的访问,并且可应用于所选的书目数据集。属于后一类的努力很少。在参考文献中。如图 5 所示,使用改进的分层狄利克雷过程,我们开发了一种算法,用于构建研究多样性(在时间 t 时的研究主题组合的广度)和研究轨迹(在时间 t-1 和 t 时的研究主题组合之间的知识空间距离)的度量。我们将其应用到计算机科学、电气工程和电子学领域 14 年的学术出版物和会议记录中,揭示了某些研究任务的自动化会导致研究主题多样性的增加和研究轨迹的转变,这是经济增长所希望的结果5。然而,我们的算法可用于在各个分析层面(例如个人、组织或地理区域)以及学术出版物或数据集的任何数据集开发多样性和轨迹的衡量标准。 专利。虽然我们的算法解决了现成的基于 ML 的分类模式的一些限制,但仍然存在更多限制。例如,它仅限于摘要的句法分析。扩展和未来的工作应该考虑语义分析或……
To the Editor—Understanding factors that affect the rate and direction of innovation has been a central aim of research in the study of science and innovation for more than half a century. However, substantially more progress has been achieved in understanding the factors that drive the rate of innovation than its direction—the set of research topics scientists or research institutions tackle at a given point in time (diversity) and over time (trajectory). Economists have long understood that markets provide insufficient incentives for innovation because of the difficulty of fully appropriating the returns to innovation investments, particularly when new innovations build on older ones 1. Similar features inhibit investment in the diversity of innovations 2, 3, thus potentially impacting their trajectory. For example, investment in environmentally friendly technologies may have been inhibited by initial successes using fossil-fuel technologies and hence scarce investments in alternative technologies 3. These insights show that more empirical research is needed to uncover what factors influence the direction of innovation. Such research is challenging. For example, estimating changes in the direction of innovation requires the boundaries of research trajectories to be defined. However, this creates a paradox, as boundaries of research trajectories are part of the core unknown to be estimated. Recent developments in machine learning (ML) have the potential to address this limitation by helping researchers infer the structure of the knowledge space by quantifying the various research topics and the distances between them. A remaining challenge is adapting ML algorithms to the study of causal relationships, because research on the direction of innovation is interested in identifying the latent categorization of research topics in order to then identify factors that causally change this structure.This is a non-trivial difference from the core prediction purpose of ML algorithms. We call on researchers to intensify their efforts in adapting ML algorithms to the study of the direction of innovation. Nascent efforts can be grouped in two categories:(1) off-the-shelf ML-based categorization schema developed by bibliographical data services (for example, the National Library of Medicine’s PubMed Related Articles algorithm or the Microsoft Academic Graph’s categorization schema 4) and (2) customized algorithms that provide access to more granular data on similarity between corpuses of text and that can be applied to bibliographical datasets of choice. Efforts that fall under the latter category are scarce. In ref. 5, using a modified hierarchical Dirichlet process, we developed an algorithm that constructs measures of research diversity (the breadth of one’s portfolio of research topics at time t) and research trajectory (the distance in knowledge space between one’s portfolio of research topics at times t–1 and t). We applied this to 14 years of academic publications and conference proceedings in computer science, electrical engineering and electronics to reveal that automation of certain research tasks leads to an increase in diversity of research topics and a shift in research trajectories, an outcome desirable for economic growth 5. However, our algorithm can be used to develop measures of diversity and trajectory at various levels of analysis such as individual, organization or geographic region, and for any dataset of academic publications or patents. While our algorithm addresses some of the limitations of off-the-shelf ML-based categorization schemas, more remain. For example, it is limited to a syntactic analysis of abstracts. Extensions and future work should consider a semantic analysis or one …
DOI: --
发表时间: 2011
期刊:
影响因子: --
作者:
D. Acemoglu
通讯作者: D. Acemoglu
DOI: 10.1287/orsc.2019.1308
发表时间: 2020-03-01
影响因子: 4.1
作者:
Furman, Jeffrey L.;Teodoridis, Florenta
通讯作者: Teodoridis, Florenta