I-Corps: eExplorer - Unstructured Data Analyzer
I-Corps: eExplorer - Unstructured Data Analyzer
批准号:
1445411
负责人:
Manoj Pooleery
金额:
$5.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-07-01 至 2014-12-31
中文摘要
非结构化数据分布在各种媒介上——电子通信,如电子邮件、聊天记录和电话记录。非结构化文档的范围使得从非结构化文本中获取智能是一项耗时、费力的任务,有时甚至是不可能完成的任务。本提案中提出的技术解决方案将在许多不同领域产生重大影响。例如,由于非结构化数据占银行和金融业领域所有数据的80%,因此拟议的解决方案将提供一种比现有数据更快地筛选这些数据的方法,从而识别欺诈和违规行为,并为纳税人和金融公司节省成本。同样,在医疗保健领域,数字化医疗数据的更有效表示可以帮助医生根据患者病史更好地定制治疗方案,或确定健康趋势,以造福公众健康。这个团队正在利用他们在自然语言处理(NLP)、机器学习(ML)和社交网络分析(特别是使用“社交事件”的概念从叙事文本中提取社交网络)方面的研究来构建一个名为eExplorer产品套件的工具。eExplorer没有尝试将非结构化文本转换为传统的结构化数据库设计,而是在非结构化文本上创建了一种适应性强、灵活和动态的软结构。通过支持分析人员与数据之间的交互,eExplorer显著缩短了数据收集和分析之间的时间。在研究数据时,分析人员可能会提供他/她想要在数据上施加的结构类型的示例。通过使用NLP和ML技术,eExplorer可以学习分析师试图强加的结构类型,并在整个数据上添加灵活的软结构。这还有一个额外的好处,即每个分析人员都可以构建自己不同的数据视图(或软结构)。
英文摘要
Unstructured data spread across a gamut of mediums - electronic communications such as emails, chat transcripts, and telephone transcripts. The gamut of unstructured documents makes it a time consuming, labor-intensive, and at times an impossible task to derive intelligence from unstructured text. The proposed technology solutions in this proposal will have a significant impact in many different areas. For example, since unstructured data represent up to 80 percent of all data in the Banking and Finance Industry domain, the proposed solutions will provide a means of sifting through this data much faster than the existing ones, resulting in identifying frauds and compliance breaches, and cost savings for taxpayers and financial firms. Similarly, in the Healthcare domain, more efficient representations of digitized medical data can help physicians better tailor treatments by patient history, or identify health trends to benefit public health.This team is leveraging their research in natural language processing (NLP), machine learning (ML), and social network analysis (specifically for extracting social networks from narrative text using the notion of a 'social event') to build a tool called eExplorer product suite. Rather than trying to convert unstructured text into traditional structured database design, eExplorer creates an adaptable, flexible and dynamic soft-structure on the unstructured text. eExplorer significantly reduces the time between data collection and analysis by supporting an interaction between an analyst and the data. While exploring the data, an analyst may provide examples of the kind of structure he/she wants to impose on the data. Using NLP and ML techniques, eExplorer learns the type of structure an analyst is trying to impose and adds a flexible and soft-structure on the entire data. This has an additional advantage that each analyst may build a different and their own view (or soft-structure) of the data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文