课题基金 / 基金详情

Putting the Clouds in Context: Statistical Machine Translation with MapReduce

Putting the Clouds in Context: Statistical Machine Translation with MapReduce
将云放在上下文中:使用 MapReduce 进行统计机器翻译
批准号:
0836560
负责人:
Jimmy Lin
金额:
$20.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-07-01 至 2011-06-30

项目摘要

项目成果

Jimmy Lin的其他基金

相似基金

相关文献

中文摘要
翻译
统计机器翻译(SMT)有望在当今多元文化和多面社会中弥合语言鸿沟。能够将文本从一种语言转换为另一种语言的系统有可能改变不同个人和组织的交流方式。尽管最近取得了成功,但我们看到了翻译技术继续进步的两个关键障碍:(1)系统的开发依赖于对大量数据的访问,可用资源的增长远远超过了单个计算机性能的增长;(2)目前的系统在很大程度上没有考虑到它们所翻译的内容的背景。除了少数例外情况,系统会逐句翻译,并且不会区分输入文本是新闻专线文章还是儿童读物。该项目通过解决这两个问题,推动了SMT技术的发展。由于在多处理器上运行的分治技术是目前大数据问题的唯一实际解决方案,我们必须开发能够利用大型计算机集群的可扩展算法。MapReduce是解决这些挑战的一个有吸引力的框架,因为它隐藏了底层分布式处理问题,如同步、容错等,允许研究人员专注于实际解决问题。通过将网络分析与跨语言信息检索技术相结合,我们可以构建丰富的多语言上下文模型,从而指导SMT系统翻译不同类型的文本。我们把重点放在维基百科的跨语言丰富上,作为展示这项技术的一个应用。尽管维基百科已经成为人类知识的宝贵宝库,但它还没有超越语言障碍。在大多数情况下,贡献者在由语言定义的竖井中工作,没有从其他地方积累的知识中获益。该项目潜在的更广泛影响不亚于跨越语言界限的知识传播,这将有助于丰富世界所有公民的生活。
英文摘要
Statistical machine translation (SMT) promises to bridge the language divide in today's multi-cultural and multi-faceted society. Systems capable of converting text from one language into another have the potential to transform how diverse individuals and organizations communicate. Despite recent successes, we see two critical impediments to continued progress in translation technology: (1) the development of systems depends on access to large amounts of data, and the growth of available resources has far outpaced increases in the performance of individual computers; and (2) current systems for the most part do not take the context of what they are translating into account. With few exceptions, systems translate sentence by sentence, and do not differentiate whether the input text is a newswire article or a children's book. This project advances the state of the art in SMT by addressing both issues. Since divide-and-conquer techniques running on multiple processors are currently the only practical solutions to large-data problems, we must develop scalable algorithms that can exploit large computer clusters. MapReduce is an attractive framework for tackling these challenges since it hides low-level distributed processing issues such as synchronization, fault tolerance, etc., allowing the researcher to focus on actually solving the problem. By coupling network analysis with cross-language information retrieval techniques, we can build rich, multilingual contextual models that will guide an SMT system in translating different types of text. We focus on cross-language enrichment of Wikipedia as an application for demonstrating this technology. Although Wikipedia has emerged as a valuable repository of human knowledge, it has yet to transcend the language barrier. For the most part, contributors work in silos defined by languages, without the benefit of knowledge that is being accumulated elsewhere. The potential broader impact of this project is no less than knowledge dissemination across language boundaries, which will serve to enrich the lives of all the world's citizens.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Student Travel Support for the 2014 IEEE International Conference on Big Data
II-EN: Hadoop NextGen Infrastructure for Heterogeneous Approaches to Data-Intensive Computing
III: Small: Providing Relevant and Timely Results: Real-Time Search Architectures and Relevance Algorithms
EAGER: Learning to Efficiently Rank with Cascades
海外基金