课题基金 / 基金详情

Putting the Clouds in Context: Statistical Machine Translation with MapReduce

Putting the Clouds in Context: Statistical Machine Translation with MapReduce
将云放在上下文中:使用 MapReduce 进行统计机器翻译
批准号:
0836560
负责人:
Jimmy Lin
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-07-01 至 2011-06-30

项目摘要

项目成果

Jimmy Lin的其他基金

相似基金

相关文献

中文摘要
翻译
统计机器翻译(SMT)有望弥合当今多元文化和多方面社会中的语言鸿沟。 能够将文本从一种语言转换为另一种语言的系统有可能改变不同的个人和组织的沟通方式。 尽管最近取得了成功,但我们看到翻译技术持续进步的两个关键障碍:(1)系统的开发依赖于对大量数据的访问,可用资源的增长远远超过了单个计算机性能的增长;(2)当前的系统在很大程度上没有考虑到它们正在翻译的内容的上下文。 除了少数例外,系统逐句翻译,并且不区分输入文本是新闻专线文章还是儿童书籍。 该项目通过解决这两个问题来推进SMT的最新技术水平。 由于在多处理器上运行的分治技术是目前大数据问题的唯一实用解决方案,因此我们必须开发可以利用大型计算机集群的可扩展算法。 MapReduce是解决这些挑战的一个有吸引力的框架,因为它隐藏了低级分布式处理问题,如同步,容错等,让研究人员可以专注于解决问题。 通过将网络分析与跨语言信息检索技术相结合,我们可以构建丰富的多语言上下文模型,这些模型将指导SMT系统翻译不同类型的文本。 我们专注于跨语言丰富的维基百科作为一个应用程序来展示这种技术。虽然维基百科已经成为人类知识的宝贵宝库,但它还没有超越语言障碍。 在大多数情况下,贡献者在语言定义的筒仓中工作,没有从其他地方积累的知识中受益。 该项目潜在的更广泛影响不亚于跨越语言界限的知识传播,这将有助于丰富世界所有公民的生活。
英文摘要
Statistical machine translation (SMT) promises to bridge the language divide in today's multi-cultural and multi-faceted society. Systems capable of converting text from one language into another have the potential to transform how diverse individuals and organizations communicate. Despite recent successes, we see two critical impediments to continued progress in translation technology: (1) the development of systems depends on access to large amounts of data, and the growth of available resources has far outpaced increases in the performance of individual computers; and (2) current systems for the most part do not take the context of what they are translating into account. With few exceptions, systems translate sentence by sentence, and do not differentiate whether the input text is a newswire article or a children's book. This project advances the state of the art in SMT by addressing both issues. Since divide-and-conquer techniques running on multiple processors are currently the only practical solutions to large-data problems, we must develop scalable algorithms that can exploit large computer clusters. MapReduce is an attractive framework for tackling these challenges since it hides low-level distributed processing issues such as synchronization, fault tolerance, etc., allowing the researcher to focus on actually solving the problem. By coupling network analysis with cross-language information retrieval techniques, we can build rich, multilingual contextual models that will guide an SMT system in translating different types of text. We focus on cross-language enrichment of Wikipedia as an application for demonstrating this technology. Although Wikipedia has emerged as a valuable repository of human knowledge, it has yet to transcend the language barrier. For the most part, contributors work in silos defined by languages, without the benefit of knowledge that is being accumulated elsewhere. The potential broader impact of this project is no less than knowledge dissemination across language boundaries, which will serve to enrich the lives of all the world's citizens.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Student Travel Support for the 2014 IEEE International Conference on Big Data
II-EN: Hadoop NextGen Infrastructure for Heterogeneous Approaches to Data-Intensive Computing
III: Small: Providing Relevant and Timely Results: Real-Time Search Architectures and Relevance Algorithms
EAGER: Learning to Efficiently Rank with Cascades
海外基金