CAREER: Creation, Visualization, and Mining of Domain Textual Graphs: Integrating Domain Knowledge and Human Intelligence
CAREER: Creation, Visualization, and Mining of Domain Textual Graphs: Integrating Domain Knowledge and Human Intelligence
批准号:
1739095
负责人:
Wei Jin
金额:
$39.02万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-07-01 至 2023-08-31
中文摘要
据了解,文本信息正以惊人的速度增长,这给试图发现隐藏在其中的有价值信息的分析师带来了巨大的挑战。例如,新的非平凡趋势、模式和感兴趣的实体之间的关联,如基因、蛋白质和疾病之间的关联,以及不同地方或人的共性之间的联系,都是这种形式的基础知识。本研究的目标是探索自动化解决方案,用于筛选这些广泛的文档集,以检测连接事实,命题或假设的有趣链接和隐藏信息。此外,还将编写一份深入、简明的跨文件摘要,解释每一联系的基本含义,并沿着从维基百科知识库获得的相关链接和解释,从而更全面地了解所发现的知识,这是补充或加强文本集现有信息的主要手段。该项目将影响许多领域,如国土安全,航空安全,生物医学和医疗保健应用。该技术将有可能暴露新的信息,在大型文档集合,并提供一个多视图的角度来看,发现的假设,通过整合领域知识和相关信息从维基百科获得。该项目将提供以研究为基础的教育和培训机会,使各级学生在信息分析和发现方面做好准备。本项目的重点是探索一种新颖的文本知识表示、集成和挖掘框架,包括以下几个方面:(i)自动构建用于实体关系发现的图形框架,这是一种有助于细粒度信息搜索和发现的新表示;(ii)有效整合来自多种来源的信息,包括代表性数据收集中所包含的知识、特定领域的知识(例如,领域本体),以及世界知识(例如,词汇资源,如WordNet和大规模知识库,如维基百科);(iii)新的发现算法和工具,识别实体之间的隐藏连接;(iv)通过启用自动本体驱动的场景检测和主题级建模来增强领域建模;以及(v)图形框架和发现假设的交互式可视化工具。本研究提出,下一代搜索工具需要整合来自多个相关单元的信息并结合各种证据来源的能力,这将在当前信息搜索和发现领域取得根本性进展。将探索自然语言处理(NLP)、信息提取(IE)、信息检索(IR)、数据挖掘、机器学习和语义网技术的组合来解决关键信息发现问题。欲了解更多信息,请访问项目网站:www.example.com。
英文摘要
It is understood that textual information is growing at an astounding pace, creating an enormous challenge for analysts trying to discover valuable information that is buried within. For example, new non-trivial trends, patterns, and associations among entities of interest, such as associations between genes, proteins and diseases, and the connections between different places or the commonalities of people, are such forms of underlying knowledge. The goal of this research is to explore automated solutions for sifting through these extensive document collections to detect interesting links and hidden information that connect facts, propositions or hypotheses. In addition, a more comprehensive view of discovered knowledge will be provided by generating an in-depth and concise cross-document summary explaining the underlying meaning of each connection, along with relevant links and explanations acquired from the Wikipedia knowledge base, which serves as the primary means of complementing or enhancing existing information in text collections. The project will impact many areas, such as homeland security, aviation safety, biomedical and healthcare applications. The techniques will have the potential to expose new information available in large document collections and to provide a multi-view perspective of discovered hypotheses by integrating domain knowledge and relevant information acquired from Wikipedia. Research-based education and training opportunities will be offered by this project to prepare students at all levels in information analysis and discovery. Specific attention will also be paid to promoting the participation of underrepresented groups in the research efforts.This project focuses on the exploration of a novel textual knowledge representation, integration, and mining framework that will cover the following areas: (i) automatic construction of graphical frameworks for entity relationship discovery, a new representation conducive to fine-grained information search and discovery; (ii) effective integration of information from multiple sources, including knowledge contained in representative data collections, domain-specific knowledge (e.g., domain ontologies), and world knowledge (e.g., lexical resources such as WordNet and large-scale knowledge repositories such as Wikipedia); (iii) new discovery algorithms and tools that identify hidden connections among entities; (iv) enhancement of domain modeling through enabling automatic ontology-driven scenario detection and topic-level modeling; and (v) interactive visualization tools for the graphical framework and discovered hypotheses. This research proposes that next-generation search tools require the capability of integrating information from multiple interrelated units and combining various evidence sources, which will make fundamental advances in the current state of the art for information search and discovery. A combination of techniques in Natural Language Processing (NLP), Information Extraction (IE), Information Retrieval (IR), Data Mining, Machine Learning, and Semantic Web will be explored to attack critical information discovery problems. For further information see the project web site: http://www.cs.ndsu.nodak.edu/~wjin/WSD-RelMiner.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A Holistic Approach to Improve Learning and Motivation in Introductory Programming with Automated Grading, Web-based Team Support, and Game Development
-
批准号:2345097
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2024
-
负责人:Wei Jin
-
依托单位:
CAREER: Creation, Visualization, and Mining of Domain Textual Graphs: Integrating Domain Knowledge and Human Intelligence
-
批准号:1452898
-
项目类别:Continuing Grant
-
资助金额:$49.84万
-
财政年份:2015
-
负责人:Wei Jin
-
依托单位:
A Cognitive-Apprenticeship Learning Curriculum Augmented by Cognitive Tutors (CAL-CT) for Fundamental Programming Concepts
-
批准号:0837505
-
项目类别:Standard Grant
-
资助金额:$14.97万
-
财政年份:2009
-
负责人:Wei Jin
-
依托单位:
海外基金