Novel statistical models for text mining with applications to Chinese history and texts
Novel statistical models for text mining with applications to Chinese history and texts
批准号:
1208771
负责人:
Jun Liu
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-07-01 至 2016-06-30
中文摘要
在本项目中,研究者研究了一系列具有挑战性的中文文本信息提取问题,包括:(1)词/短语发现,(2)文本分割,(3)技术术语识别,(4)技术术语之间的关联发现。与字母语言(如英语)不同,汉语有许多特殊的属性:没有单词边界,没有明确的单词定义,传统上没有标点符号,以及独特的语法。因此,将大多数为字母语言开发的方法直接应用到汉语中是有问题的。此外,现有的中文文本分析方法也存在许多局限性。相反,研究者提出了一种先进的词字典模型(AWDM),该模型可以同时实现词发现、文本分割和技术术语识别,而这些传统上是分开研究的。其思路是首先从文本中列举出所有符合一定标准的候选词,并为每个候选词分配一个潜在的词类型标签,代表不同类型的专业术语(如姓名、地址、办公室头衔、时间标签以及背景文本)和相应的词使用频率,从而建立一个词词典。然后,给出了不同词和词类型之间的马尔可夫依赖模型,对文本的潜在语法和语义结构进行建模。在训练数据(即已知技术术语列表)的帮助下,AWDM可以从巨大的候选词空间中自动选择最有意义的词,不仅根据词的内容,而且根据词周围的上下文确定每个词的词类型,并根据语法和语义信息对文本进行分割。与文献中已有的方法相比,AWDM对语法和语义信息进行了联合建模,对词发现、文本分词和专业术语识别进行了综合分析,效率更高。结合其他文本挖掘工具,如主题模型和主题词典模型,该方法将形成一个强大的多层次(汉字级、词/短语级、主题级、话题级)的中文文本分析平台。随着互联网和数字技术的爆炸式发展,大量数字化的中文文本可以方便地收集。例如,许多用繁体中文写的中国历史文献现在都有数字形式;此外,报纸、论坛、博客和微博等公共媒体每天都在产生大量的中文文本。因此,开发文本挖掘工具来自动地从这些数据中提取信息并创建新的知识具有很大的吸引力。这个项目中的思想和方法可能会对如何研究中国历史产生重大影响。从不断增长的数字化历史文献数据库中提取信息的一种有效而可靠的方法将使研究人员能够基于大量分解的数据点分析随时间的变化,这在过去是不切实际的。此外,虽然最初是为中文设计的,但这些方法有可能应用于其他类似中文的亚洲语言,如日语和韩语,从而为亚洲历史的研究提供一个强大的多语言平台。此外,在这个项目中研究的对抗识别命名实体挑战的新方法也有可能扩展到字母语言,如英语。最后,本项目研究的思想和方法有可能被推广为一种系统的工具,它可以消化中文文本的任何数据流,并输出一个结构化的数据库,其中包含输入数据所描述的个人和组织的关键信息,从而使研究人员更容易发现我们社会生活中各种“单位”的社会网络。我们的算法发现的各种项目关联模式对公共媒体和社会学的研究也非常宝贵,可能有助于及时揭示新的重要流行病学事件和社会趋势。这些类型的信息在商业决策和政府政策制定中具有重要意义。
英文摘要
In this project, the investigators study a series of challenging problems of extracting information from Chinese text, including: (1) word/phrase discovery, (2) text segmentation, (3) technical term recognition, and (4) association discovery among technical terms. Different from alphabetical languages such as English, Chinese has many special properties: no word boundaries, no clear definition of words, traditionally no punctuation, and a unique grammar. Thus, it is problematic to apply most methods developed for alphabetical languages directly to Chinese. Moreover, the available methods for analyzing Chinese text in the literature have many limitations. Instead the investigators propose an advanced word dictionary model (AWDM) that can simultaneously achieve word discovery, text segmentation and technical term recognition, which are traditionally studied separately. The idea is to build up a word dictionary first by enumerating all word candidates satisfying a certain criterion from the texts and assign to each word candidate a latent word type label representing different types of technical terms (such as names, addresses, office titles, time labels, as well as background texts) and corresponding word usage frequencies. Then, a Markov dependence model among different words and word types is given to model the potential grammatical and semantic structure of the texts. With the help of the training data (i.e., lists of known technical terms), the AWDM can automatically select the most meaningful words from the huge space of word candidates, determine the word type for each word based on not only the content of the word but also the context around the word, and segment the texts based on both grammatical and semantic information. Compared to the existing methods in the literature, the AWDM enjoys a better efficiency due to the joint modeling of the grammatical and semantic information and the integrated analysis of word discovery, text segmentation and technical term recognition. Combined with other text mining tools, such as topic models and theme dictionary models, the proposed method will lead to a powerful multi-level (Chinese character level, word/phrase level, theme level, topic level) analysis platform for Chinese texts.With the explosive growth of the internet and digital technologies, large quantities of digitalized Chinese texts can be easily collected. For example, lots of Chinese historical documents written in traditional Chinese are now available in digital form; and, public media such as new papers, forums, blogs and microblogs, are producing huge amounts of Chinese text every day. Thus there is great appeal in developing text mining tools to automatically extract information from these data and create new knowledge. The ideas and approaches in this project may have significant impacts on how Chinese history will be studied. An efficient and reliable method for extracting information from the ever growing databases of digitized historical documents will enable researchers to analyze change over time based on large numbers of disaggregated data points, something impractical in the past. Furthermore, although originally designed for Chinese, these approaches have the potential to be applied to other Asian languages similar to Chinese, such as Japanese and Korean, and thus provide a powerful multi-language platform for the study of Asian history. In addition, the novel way of combatting the challenges in recognizing named entities studied in this project also has the potential to be extended to alphabetical languages such as English. Finally, the ideas and approaches studied in this project have the potential to be generalized into a systematic tool that digests any data flow of Chinese texts, and outputs a structured database that contains key information about the individuals and organizations described by the input data, thus making it easier for researchers to discover social network of all kinds of "units" in our social life. Various item association patterns discovered by our algorithms are also invaluable to the study of public media and sociology, and may help reveal new important epidemiological events and societal trends in a timely fashion. These types of information can have important implications in business decision making and governmental policy making.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
REU Site: Molecular Biology and Genetics of Cell Signaling
-
批准号:2349577
-
项目类别:Standard Grant
-
资助金额:$42.67万
-
财政年份:2024
-
负责人:Jun Liu
-
依托单位:
SCC-PG: Building a smart and connected rural community for improved healthcare access through the deployment of integrated mobility solutions
-
批准号:2303284
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2023
-
负责人:Jun Liu
-
依托单位:
Collaborative Research: Bayesian and Semi-Bayesian Methods for Detecting Relationships in High Dimensions
-
批准号:2015411
-
项目类别:Standard Grant
-
资助金额:$12.0万
-
财政年份:2020
-
负责人:Jun Liu
-
依托单位:
REU Site: Molecular Biology and Genetics of Cell Signaling
-
批准号:1950247
-
项目类别:Standard Grant
-
资助金额:$36.59万
-
财政年份:2020
-
负责人:Jun Liu
-
依托单位:
Domain-Engineering Enabled Thermal Switching in Ferroelectric Materials
-
批准号:2011978
-
项目类别:Continuing Grant
-
资助金额:$55.86万
-
财政年份:2020
-
负责人:Jun Liu
-
依托单位:
CAREER: Pushing the Lower Limit of Thermal Conductivity in Layered Materials
-
批准号:1943813
-
项目类别:Continuing Grant
-
资助金额:$51.88万
-
财政年份:2020
-
负责人:Jun Liu
-
依托单位:
Collaborative Research: Novel Statistical Tools for Metagenomics and Metabolomics Data
-
批准号:1903139
-
项目类别:Continuing Grant
-
资助金额:$35.0万
-
财政年份:2019
-
负责人:Jun Liu
-
依托单位:
Travel Support for Student Participation at the 2019 ASME-IMECE Micro and Nano Technology Forum; Salt Lake City, Utah; November 10-14, 2019
-
批准号:2000224
-
项目类别:Standard Grant
-
资助金额:$1.51万
-
财政年份:2019
-
负责人:Jun Liu
-
依托单位:
Collaborative Research: Theoretical and Methodological Frameworks for Causal Inference of Peer Effects
-
批准号:1712714
-
项目类别:Standard Grant
-
资助金额:$23.87万
-
财政年份:2017
-
负责人:Jun Liu
-
依托单位:
Variable Selection via Inverse Modeling for Detecting Nonlinear Relationships
-
批准号:1613035
-
项目类别:Continuing Grant
-
资助金额:$20.0万
-
财政年份:2016
-
负责人:Jun Liu
-
依托单位:
ATD: Collaborative Research: Statistical Modeling of Short-Read Counts in RNA-Seq
-
批准号:1120368
-
项目类别:Continuing Grant
-
资助金额:$30.95万
-
财政年份:2011
-
负责人:Jun Liu
-
依托单位:
Bayesian Partition Models for Detecting Influential and Interactive Variables
-
批准号:1007762
-
项目类别:Continuing Grant
-
资助金额:$34.99万
-
财政年份:2010
-
负责人:Jun Liu
-
依托单位:
CSI - An Adaptive Data Pulling Framework for Supporting Time-Critical Streaming Media Applications
-
批准号:0720809
-
项目类别:Standard Grant
-
资助金额:$12.5万
-
财政年份:2007
-
负责人:Jun Liu
-
依托单位:
RUI: Nonstationary and High Dimensioanl Nonparametric Transfer Function Models Using Polynomial Splines
-
批准号:0707082
-
项目类别:Standard Grant
-
资助金额:$10.5万
-
财政年份:2007
-
负责人:Jun Liu
-
依托单位:
A New MCMC Framework with Applications to Protein Bioinformatics
-
批准号:0706989
-
项目类别:Continuing Grant
-
资助金额:$62.92万
-
财政年份:2007
-
负责人:Jun Liu
-
依托单位:
Collaborative Research: Advanced Sequential Monte Carlo Methods and Applications
-
批准号:0244638
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Jun Liu
-
依托单位:
Statistical Problems in Hidden Markov Modeling for Biology and Chemistry
-
批准号:0204674
-
项目类别:Continuing Grant
-
资助金额:$32.58万
-
财政年份:2002
-
负责人:Jun Liu
-
依托单位:
Collaborative Research: Sequential Monte Carlo Methods and Their Applications
-
批准号:0094613
-
项目类别:Continuing Grant
-
资助金额:$33.54万
-
财政年份:2000
-
负责人:Jun Liu
-
依托单位:
New Monte Carlo Methods for Scientific and Statistical Computing
-
批准号:0196228
-
项目类别:Continuing Grant
-
资助金额:$14.23万
-
财政年份:2000
-
负责人:Jun Liu
-
依托单位:
New Monte Carlo Methods for Scientific and Statistical Computing
-
批准号:9803649
-
项目类别:Continuing Grant
-
资助金额:$14.23万
-
财政年份:1998
-
负责人:Jun Liu
-
依托单位:
国内基金
海外基金
基于随机网络演算的无线机会调度算法研究
-
批准号:60702009
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2007
-
负责人:雷蕾
-
依托单位: