CRII: III: Capturing Dynamism in Causal Relationships: A New Paradigm for Relationship Extraction from Text
CRII: III: Capturing Dynamism in Causal Relationships: A New Paradigm for Relationship Extraction from Text
批准号:
1948322
负责人:
Sunandan Chakraborty
金额:
$17.43万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-05-15 至 2023-04-30
中文摘要
文本挖掘在将海量的非结构化文本数据转化为知识的方法方面取得了重要进展。然而,当前的关系提取范式有一个主要局限性:它对信息的快照进行建模,但未能捕捉到知识的根本对话和动态性质:相互冲突的发现、不一致的发现、反驳、矛盾、增援或确认,所有这些都随着时间的推移而变化。这个项目旨在捕捉这种基本的知识动态,特别是关注因果关系。虽然许多文章,包括学术文章,提供了表达因果关系的知识和关系,但这种关系不是一成不变的,可能会随着条件的变化而变化。这个项目的目标是从文本数据中识别因果知识的线索,量化因果关系的强度,并对其在不断变化的条件下的动力学进行建模。最终,该项目的目标是对从文本中提取的知识进行更全面的建模。由于文本数据被来自不同国家重要领域的研究人员和从业者广泛使用,包括医学和卫生、经济学、公共政策、新闻学,该项目的结果试图为从业者提供新的方法来理解大型文本数据集中存在的因果关系的演变性质。具体地说,该项目开发的新方法将被应用于探索公共卫生数据,以确定气候、政治、经济条件的变化可能如何影响不同地理区域人口的心理和身体健康。此外,作为该项目的一部分,还将开展各种教育活动--该项目的新兴主题和相关主题将被纳入应用数据科学硕士项目的各种课程的课程中;促进本科生研究,特别是招募来自代表性不足和经济困难社区的学生参与该项目;组织研究研讨会,鼓励高中生参与STEM研究。项目活动包括开发一种新的因果关系提取模型,该模型利用一个结合了语义和语法线索的统一深度学习框架。这种方法将利用由名词、动词和其他词类之间的语法关系通过图形或树形模型表示的句子的关键句法特征。这项工作将确定句子是否具有表示因果关系的结构。此外,模型的顺序成分将利用语义并识别句子中某些词的影响,以表征文本中表达的因果关系的性质。这项任务将捕捉这种关系的强度(例如,使用“极有可能”、“肯定”等线索)、任何支持或反对的证据(例如,“将导致”或“不会导致”),并将识别条件线索(例如,“在存在的情况下”)等。量化这样的定性属性将导致该项目的第二次创新-因果距离。因果距离是一个时变的度量,它将表示两个实体之间的因果关系的大小,并通过随着时间的变化随着条件或新证据的变化而修改自身来捕捉关系的动态化。总体而言,这些项目所取得的进展将进一步增强我们对挖掘大型文本数据集中嵌入的因果关系线索并进行推理所需的新计算方法的理解。该项目的成果,如数据集、源代码、最终软件、成果和出版物,将通过可公开访问的URL和在线代码库共享。此外,所有的项目资源和成果将在项目网站上提供。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Text mining made important advances in methods to convert vast and unstructured text data into knowledge. However, the current paradigm of relationship extraction has one major limitation: it models snapshots of information but fails to capture the fundamentally dialogic and dynamic nature of knowledge: conflicting findings, inconsistent discoveries, refutations, contradictions, reinforcements or confirmations, all changing over time. This project aims to capture such fundamental dynamics of knowledge, specifically focusing on causal relationships. Whereas numerous articles, including academic articles, present knowledge and relationships that express causality, such relationships are not static and can change over time due to changing conditions. The objective of this project is to identify cues of causal knowledge from text data, quantify the strength of the causal relationship, and model its dynamics over changing conditions. Ultimately, the project aims at modelling a more holistic view of the knowledge extracted from text. As text data is extensively used by researchers and practitioners from different domains of national importance, including, medicine and health, economics, public policy, journalism, the results of this project seek to provide the foundation to offer practitioners new ways to understand the evolving nature of the causal relationships present in large text datasets. Specifically, the novel approaches developed in the project will be applied to explore public health data to determine how changing climatic, political, economic conditions may affect the mental and physical health of the population in different geographic areas. In addition, there will be various educational activities as part of this project - emerging and related topics from this project will be included in the curricula of various courses in the applied data science master’s program; promote undergraduate research, specifically, recruit students to work in the project who are from underrepresented and economically disadvantaged communities; organize a research workshop to encourage participation of high school students in STEM research. The project activities include the development of a novel model of causal relationship extraction that leverages a unified deep learning framework combining both semantic and syntax cues. This approach will utilize the key syntactical features of a sentence represented by the grammar relationships between noun, verbs and other parts of speech through graphical or tree-like models. This work will determine whether the sentence features a structure that signals causality. Moreover, the sequential component of the model will utilize the semantics and identify the influence of certain words in the sentence to characterize the nature of the causal relationship expressed in the text. This task will capture the strength of the relationship (e.g., using cues like "extremely likely", "definitely"), any supporting or opposing evidences (e.g., "will lead to" or "does not lead to"), and will identify conditional cues (e.g., "in the presence of") etc. Quantifying such qualitative properties will lead to the second innovation of this project – causal distance. Causal distance is a time-variant metric that will denote the magnitude of causality between two entities as well as capture the dynamism of the relationship by modifying itself over time with changing conditions or new evidences. Collectively, the advances pursued in this projects will further enhance our understanding of the novel computational approaches needed to unearth and reason on cues of causal relationships embedded in large text data sets. The outcomes of this project, such as datasets, source code, final software, results and publications will be shared via publicly accessible URLs and online code repositories. Additionally, all the project resources and outcomes will be made available on the project website.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
A Study of Extracting Causal Relationships from Text
从文本中提取因果关系的研究
DOI:
--
发表时间:
2022
期刊:
Cham.
影响因子:
--
作者:
[Gujarathi, Pranav, Reddy, Manohar, Tayade, Neha, Chakraborty, Sunandan]
通讯作者:
Chakraborty, Sunandan
DOI:
10.1109/bigdata55660.2022.10020994
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
作者:
[P. Gujarathi;Jack VanSchaik;Venkatanaidu Karri;A. Rajapuri;Biju Cheriyan;T. Thyvalikakath;Sunandan Chakraborty]
通讯作者:
P. Gujarathi;Jack VanSchaik;Venkatanaidu Karri;A. Rajapuri;Biju Cheriyan;T. Thyvalikakath;Sunandan Chakraborty
D-ISN/Collaborative Research: An Interdisciplinary Approach to the Discovery, Analysis, and Disruption of Wildlife Trafficking Networks
-
批准号:2146351
-
项目类别:Standard Grant
-
资助金额:$21.66万
-
财政年份:2022
-
负责人:Sunandan Chakraborty
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于人工智能与多组学的III期结核性脓胸CT“低密度线”形成机制及手术时机预测模型研究
-
批准号:JCZRMS202602483
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:
-
依托单位:
基于MOF–CRISPR微流控平台的雄黄As(III)/As(V)价态识别与炮制耦合机制研究
-
批准号:JCZRLH202600780
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:
-
依托单位:
白术内酯III靶向IRF4-CD36轴通过调控脂质代谢重编程提升结直肠癌奥沙利铂敏感性的机制研究
-
批准号:2026JJ82690
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:张卓
-
依托单位:
基于废水零排放的FeS-As(III)置换法从污酸中清洁脱砷处理技术研究
-
批准号:2026JJ30130
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:张二军
-
依托单位:
全钒液流电池负极V(II)/V(III)电化学氧化还原的催化机理研究
-
批准号:2025JJ50094
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:王珏
-
依托单位:
猪纤维蛋白粘合剂预防胸外科术后漏气的适应症拓展研究:一项多中心、随机对照III期临床试验
-
批准号:25SF1901800
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:赵德平
-
依托单位:
硅基III-V族亚微米线激光器的光场模式调控与耦合机理研究
-
批准号:JCZRQN202501004
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:
-
依托单位:
吡咯烷生物碱所致肝窦阻塞综合征III区肝损伤的新机制——局部氨代谢紊乱
-
批准号:JCZRYB202500652
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:
-
依托单位:
HOXC8/OPN/CD44/EGFR轴介导的奥沙利铂耐药性在III期右半结肠癌耐药进展中的研究
-
批准号:2025JJ50694
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:喻南慧
-
依托单位:
MXene/nZVI@FH材料微域层界面调控水中砷(III)氧化迁移机制
-
批准号:2025JJ50319
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:陈润华
-
依托单位: