课题基金 / 基金详情

CAREER: Metalinguistic Natural Language Understanding

CAREER: Metalinguistic Natural Language Understanding
职业:元语言自然语言理解
批准号:
2144881
负责人:
Nathan Schneider
金额:
$54.99万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-01 至 2027-06-30

项目摘要

项目成果

Nathan Schneider的其他基金

相似基金

相关文献

中文摘要
翻译
人们如何使用语言是许多领域的研究课题,包括语言学、文学、法律和语言教育。关于许多语言的大量描述已经存在于自然语言本身中。(语法教科书、写作建议和语言学文章)。(例如)英语的描述可能包括用技术术语和形式符号补充的例句。该项目将开发自然语言处理(NLP)算法来处理和挖掘这些文本资源,以便语言分析人员能够更好地检测和综合感兴趣的模式。首先,将开发算法来识别一段文本在哪里评论一个单词的意思,或者给出一个如何使用它的例子。其次,用技术描述丰富文本的算法将得到改进,以更好地评估它们自己的优缺点。最后,将开发在大型文本集合中检索歧义词的特定用法的功能,以帮助分析人员。该项目中的算法贡献预计将直接应用于对文本集合中的语言进行密切分析至关重要的各个领域,包括法律和语言学。在更广泛的范围内,这些能力有可能对人工智能(AI)产生革命性的影响,使人类和机器能够明确地相互教授语言是如何工作的,熟练地访问有关语言的学术著作,并提供和解释语言建议(例如,写作协助)。该项目开发算法和任务,着眼于技术,使人类能够更有效、更准确地对文本进行元语言查询。需要解决的主要挑战是:(1)检测文本元语言:本项目将制定任务和算法来识别文本中的元语言描述(如使用/提及区分、定义、语言示例),重点关注三个丰富的类型:法律、语言论坛和语言学。将开发一个新的基准数据集和共享任务来比较元语言标记器。(2)改进模型置信度校准,重点关注具有长尾标记集的标记器。更好的概率估计将使分析人员能够就如何平衡自动和人工处理做出明智的决定,并可以预测不同类型错误的发生率。(3)将开发逐例查询算法,用于从大型文本集合中检索歧义词或短语的特定用法。利用这种算法的工具将为语言学家、词典编纂者、语言教师和文学学者进行基于语料库的新型调查开辟道路。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
How people use language is a topic of study in a variety of fields, including linguistics, literature, law, andlanguage education. Extensive descriptions about many languages already exist in natural language itself(e.g., grammar textbooks, writing advice, and linguistics articles). Descriptions in (for example) Englishmay include example sentences supplemented with technical terminology and formal notation. Thisproject will develop natural language processing (NLP) algorithms to process and mine these textualresources so that language analysts can better detect and synthesize patterns of interest. First, algorithmswill be developed to recognize where a piece of text is commenting on the meaning of a word, or givingan example of how it could be used. Second, algorithms for enriching text with technical descriptions willbe improved to report better estimates of their own strengths and weaknesses. Finally, capabilities forretrieving specific uses of an ambiguous word in a large text collection will be developed to aid analysts.The algorithmic contributions in this project are expected to have direct application to technologies invarious fields where close analysis of language in text collections is crucial, including law and linguistics.On a wider scale, these capabilities have the potential to be transformative for artificial intelligence (AI),allowing humans and machines to teach each other explicitly about how language works, to deftly accessscholarly work about language, and to give and interpret language advice (e.g., writing assistance).This project develops algorithms and tasks with an eye toward technologies that would enable humans tomore efficiently and accurately conduct metalinguistic inquiries about text. Key challenges to beaddressed are: (1) Detecting textual metalanguage: This project will formulate tasks and algorithms torecognize metalinguistic descriptions (such as the use/mention distinction, definitions, linguisticexamples) in text, focusing on three genres where they are abundant: law, language discussion forums,and linguistics. A new benchmark dataset and shared task will be developed to compare metalinguistictaggers. (2) Improving model confidence calibration, focusing on taggers with long-tail tagsets. Betterprobability estimates will enable analysts to make informed decisions about how to balance automatic andmanual processing and can anticipate rates of different types of errors. (3) Query-by-example algorithmswill be developed for retrieving specific usages of an ambiguous word or phrase from a large textcollection. Tools leveraging such algorithms would open the way to new kinds of corpus-basedinvestigations by linguists, lexicographers, language teachers, and literary scholars.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: DASS: Transitioning open-source software projects to accountable community governance
  • 批准号:
    2217654
  • 项目类别:
    Standard Grant
  • 资助金额:
    $12.99万
  • 财政年份:
    2022
  • 负责人:
    Nathan Schneider
  • 依托单位:
NSF-BSF: RI: Small: Collaborative Research: Modeling Crosslinguistic Influences Between Language Varieties
  • 批准号:
    1812778
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $16.63万
  • 财政年份:
    2018
  • 负责人:
    Nathan Schneider
  • 依托单位:
海外基金