EAGER: SaTC: Tracking Semantic Change in Medical Information
EAGER: SaTC: Tracking Semantic Change in Medical Information
批准号:
1834597
负责人:
Ritwik Banerjee
金额:
$29.88万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-05-15 至 2022-03-31
中文摘要
在信息通过网络空间时,信息含义的变化可能会误导那些访问信息的人。该项目将开发一种新的数据集和算法,以识别和分类保持原始含义或经历扭曲的医疗信息。这个项目没有给这些信息贴上外部的真/假标签,而是调查了新闻报道本身内部的一系列变化,这些变化逐渐导致了与最初的医疗索赔的偏离。识别原创医学文章和新闻故事之间的重要差异是一项具有挑战性、高风险和高回报的冒险。这项工作的更广泛影响包括通过为理解自然语言信息的时间变化作出新的贡献而给研究界带来的好处,以及以诸如问答等改进的信息工具的形式带来的社会效益。特别是在医学领域,了解时间扭曲和与实际医学结果的偏差可以减少有害健康选择的发生,例如,通过将研究结果嵌入新闻、社交媒体或搜索引擎。该项目将开发一个医学科学出版物的大型数据集,并通过设计和开发实体及其属性和关系的离散时间序列表示来记录它们在新闻中随时间变化的特征。这项任务将为设计和执行机器学习任务提供基础,这些任务利用自然语言的文体特征和时间分布来确定和分类这种变化。这项研究将超越目前局限于对个别文章进行真假分类的方法,从而能够识别和分析叙事中的信息变化,包括语义变化和细微差别,或相关信息的选择性强调。这项研究需要一种带有Bootstrapping的无监督和半监督机器学习方法,并探索一种二元标记任务来区分扭曲的信息片段和那些忠于科学发现的信息,以及一种多标签分类来学习随着时间发生的语义变化的类型。数据集将通过自然语言处理资源的存档位置传播,如语言数据联盟(https://www.ldc.upenn.edu/)),以促进其他研究人员的长期可用,并将使用BitBucket或GitHub来确保代码的开发、维护、共享和存档。该奖项反映了美国国家科学基金会的法定使命,并已通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Changes in the meaning of information as it passes through cyberspace can mislead those who access the information. This project will develop a new dataset and algorithms to identify and categorize medical information that remains true to the original meaning or undergoes distortion. Instead of imposing an external true/false label on this information, this project looks into a series of changes within the news coverage itself that gradually lead to a deviation from the original medical claims. Identifying important differences between original medical articles and news stories is a challenging, high risk-high reward venture. Broader impacts of this work include benefits to the research community by making novel contributions to understanding temporal changes in natural language information, as well as social benefits in the form of improved informational tools like question-answering. For the medical domain in particular, understanding temporal distortions and deviations from actual medical findings can reduce occurrences of harmful health choices, for instance, by embedding the research outcomes in news, social media, or search engines. This project will develop a large dataset of medical scientific publications, and record their characteristics as they change over time across news by designing and developing discrete time-series representations of entities and their attributes and relations. This task will provide the basis for designing and implementing machine learning tasks that exploit stylometric features in natural language in conjunction with temporal distributions to identify and categorize such changes. This research will go beyond current approaches limited to true/false classification of individual articles, and hence be able to identify and analyze information change in narratives, including semantic changes and nuances, or selective emphasis of related information. The research entails an unsupervised and a semi-supervised machine learning approach with bootstrapping, and exploring a binary labeling task to distinguish distorted pieces of information from those that are faithful to the scientific finding, and a multi-label categorization to learn the type of semantic change occurring through time. The dataset will be disseminated via an archival location for natural language processing resources such as the Linguistic Data Consortium (https://www.ldc.upenn.edu/) to facilitate long-term availability to other researchers, and BitBucket or GitHub will be used to ensure the development, maintenance, sharing, and archiving of code.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1007/978-3-030-28577-7_23
发表时间:
2019-09
期刊:
影响因子:
--
作者:
[Chaoyuan Zuo;A. Karakas;Ritwik Banerjee]
通讯作者:
Chaoyuan Zuo;A. Karakas;Ritwik Banerjee
DOI:
10.18653/v1/2020.emnlp-main.139
发表时间:
2020-11
期刊:
影响因子:
--
作者:
[Chaoyuan Zuo;Narayan Acharya;Ritwik Banerjee]
通讯作者:
Chaoyuan Zuo;Narayan Acharya;Ritwik Banerjee
An Empirical Assessment of the Qualitative Aspects of Misinformation in Health News
对健康新闻中错误信息定性的实证评估
DOI:
--
发表时间:
2021
期刊:
and Propaganda
影响因子:
--
作者:
[Zuo, Chaoyuan, Zhang, Qi, Banerjee, Ritwik]
通讯作者:
Banerjee, Ritwik
Diagnosis, Prevention, and Cure for Misinformation
错误信息的诊断、预防和治疗
DOI:
10.1109/cogmi52975.2021.00028
发表时间:
2021
期刊:
CogMI 2021
影响因子:
--
作者:
[Banerjee, Ritwik, Ray, Indrakshi]
通讯作者:
Ray, Indrakshi
An Investigation into the Contribution of Locally Aggregated Descriptors to Figurative Language Identification
局部聚合描述符对比喻语言识别的贡献的调查
DOI:
10.18653/v1/2021.insights-1.15
发表时间:
2021
期刊:
Proceedings of the Second Workshop on Insights from Negative Results in NLP
影响因子:
--
作者:
[Saravani, Sina Mahdipour, Banerjee, Ritwik, Ray, Indrakshi]
通讯作者:
Ray, Indrakshi
共 7 条
Collaborative Research: EAGER: MedAn: A Framework for Investigating Live Medical Data against Privacy Laws
-
批准号:2335686
-
项目类别:Continuing Grant
-
资助金额:$12.5万
-
财政年份:2023
-
负责人:Ritwik Banerjee
-
依托单位:
海外基金