课题基金 / 基金详情

Learning the morphology of complex synthetic languages

Learning the morphology of complex synthetic languages
学习复杂合成语言的形态
批准号:
EP/E010857/1
负责人:
Peter Flach
金额:
$48.03万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2006
资助国家:
英国
项目状态:
已结题
起止时间:
2006 至 --

项目摘要

项目成果

Peter Flach的其他基金

相似基金

相关文献

中文摘要
翻译
该项目旨在应用先进的机器学习技术,以学习形态学-单词是如何从成分组成-的合成(即,形态复杂的)语言。这将允许改进文本到语音系统的复杂语言,如isiZulu。词法分析是将词分解成其成分(语素),并为每个成分分配语法特征。举一个简单的例子,在英语中,单词unhappier被分解为以下成分:un(形容词性否定前缀)+happy(形容词性词干)+er(比较级后缀),同时考虑到单词成分的允许顺序和这些成分在连接时的正字法形状的变化。大多数欧洲语言中的形态学现象都可以用有限状态技术(如正则表达式)来表达。然而,这个项目关注的是结构上更复杂的合成语言(主要是非欧洲的)。这些语言具有复杂的递归形态结构,需要比有限状态自动机更强大的机制。该项目的主要研究目标是通过学习表示单词成分的允许序列的规则和改变成分的正字法形状的规则来自动将单词分解成其成分。这涉及到解决形态学学习中的一系列开放性问题,这些问题阻碍了学习整个形态学规则集。我们选择归纳逻辑编程(ILP)进行训练,因为它的逻辑基础允许表示可以通过随机特征扩展的复杂形式主义。ILP方法还可以直接从字符串等无界数据项中归纳规则,这使得注释和训练与底层语言学更自然地相关。拟议的研究将对非洲和亚洲发展中国家的文本到语音系统的生产产生巨大的好处(当地语言语音技术倡议建立的伙伴关系和联系将使实际交付成为可能,见www.llsti.org)。本项目中开发的自动形态分析工具将有助于创建可理解的文本到语音系统,这些系统需要进行形态分析,以(1)自动音调分配,这对大多数非洲语言至关重要。(2)正确的韵律,包括俄语所需的重音分配,以及大多数世界语言(包括欧洲语言)所需的短语预测。(3)印度语言印地语和泰卢固语、土耳其语和许多其他语言需要适当的字母发音规则。这项研究将为实施移动的网络提供商提供的土著和少数民族语言语音服务提供技术(如医疗保健、就业、农业、环境等信息)。这项技术的另一个重要应用将是许多亚洲和非洲国家盲人的屏幕阅读器。
英文摘要
This project aims to apply advanced machine learning techniques in order to learn the morphology -- the way words are formed from constituents -- of synthetic (i.e., morphologically complex) languages. This will allow improved text-to-speech systems for complex languages such as isiZulu. Morphological analysis is the decomposition of words into their constituents (morphemes) with the assignment of grammatical features to each of constituents. To take a simple example in English, the word unhappier is decomposed into the following components: un(adjectival negative prefix)+happy(adjectival stem)+er(comparative suffix) taking into account both the allowed sequence of word constituents and the changes of the orthographic shape of these constituents when they are concatenated. Most morphological phenomena in the majority of European languages can be expressed by finite-state techniques such as regular expressions. This project, however, is concerned with the structurally more complex synthetic languages (mostly non-European). These languages exhibit complex recursive morphological structures that require more powerful mechanisms than finite-state automata. The main research goal of the project is to automatically decompose the word into its constituents by learning the rules for representing permissible sequences of word constituents and the rules that change the orthographic shape of the constituents. This involves tackling a set of open problems in morphological learning that prevents learning the whole set of morphological rules. We have chosen Inductive Logic Programming (ILP) for training as its logical foundations allow representing complex formalisms that can be expanded by stochastic features. ILP methods can also induce rules directly from unbounded data items such as strings, which makes annotation and training more naturally related to the underlying linguistics. The proposed research will have a tremendous benefit for producing Text-to-Speech Systems in developing African and Asian countries (and the practical delivery will be enabled by the partnerships and contacts forged by the Local Language Speech Technology Initiative, see www.llsti.org). The automated morphological analysis tools developed in this project will facilitate the creation of intelligible Text-to-Speech systems that require morphological analysis for (1) Automatic tone assignment, which is essential for most African languages.(2) Proper prosody, which includes stress assignment required for Russian, and phrase prediction required for most world languages including European ones.(3) Proper letter-to-sound rules required for the Indian languages Hindi and Telugu, the Turkish language and many others. The research will provide the technology for the implementation of indigenous and minority language voice services offered by mobile network providers (such as information on healthcare, jobs, agriculture, the environment etc.) Another important application for this technology will be screen readers for blind people in many Asian and African countries.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
Ukwabelana - An open-source morphological Zulu corpus
Ukwabelana - 开源形态祖鲁语语料库
DOI: --
发表时间: 2010
期刊:
影响因子: --
作者: [Andrew Van Der Spuy]
通讯作者: Andrew Van Der Spuy
Weakly Supervised Morphology Learning for Agglutinating Languages Using Small Training Sets
使用小型训练集进行凝集语言的弱监督形态学学习
DOI: --
发表时间: 2010
期刊:
影响因子: --
作者: [Ksenia Shalonova]
通讯作者: Ksenia Shalonova
EMMA: A Novel Evaluation Metric for Morphological Analysis
EMMA:一种新的形态分析评估指标
DOI: --
发表时间: 2010
期刊:
影响因子: --
作者: [Christian Monson]
通讯作者: Christian Monson
Enhanced word decomposition by calibrating the decision threshold of probabilistic models and using a model ensemble
通过校准概率模型的决策阈值并使用模型集成来增强单词分解
DOI: --
发表时间: 2010
期刊: Proceedings of the Annual Meeting of the Association for Computational Linguistics
影响因子: --
作者: [Spiegler S.]
通讯作者: Spiegler S.
Machine learning for modelling and control of direct air capture systems
  • 批准号:
    NE/X007375/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $1.39万
  • 财政年份:
    2022
  • 负责人:
    Peter Flach
  • 依托单位:
REFRAME: Rethinking the Essence, Flexibility and Reusability of Advanced Model Exploitation
  • 批准号:
    EP/K018728/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $48.58万
  • 财政年份:
    2013
  • 负责人:
    Peter Flach
  • 依托单位:
国内基金
海外基金
量子点技术对细胞表面蛋白和受体在体内分布的研究
  • 批准号:
    30570686
  • 项目类别:
    面上项目
  • 资助金额:
    26.0万元
  • 批准年份:
    2005
  • 负责人:
    顾江
  • 依托单位: