课题基金 / 基金详情

II-NEW: Hadoop cluster acquisition, deployment and training for speech and language processing

II-NEW: Hadoop cluster acquisition, deployment and training for speech and language processing
II-NEW:用于语音和语言处理的 Hadoop 集群获取、部署和训练
批准号:
0958585
负责人:
Izhak Shafran
金额:
$40.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-04-01 至 2012-10-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目旨在教育、培训和装备研究生,使其掌握正在成为语音和NLP的关键范式:分布式算法,这是谷歌取得巨大成功的一种范式。这种分布式处理范例代表的不仅仅是计算上的便利,而是一种设计算法以优化这种环境中的性能的方法,与直接并行部署的标准算法相比,产生了巨大的改进。当前机构基础设施提案的目标是(1)获得广泛的(384核)处理器集群,用作OHSU口语理解中心(CSLU)的Hadoop集群;(2)将集群集成到现有的计算基础设施中;以及(3)开发教育资源(教程、实验室课程、课程模块和研讨会),重点关注使用Hadoop集群的“how-to”信息和分布式计算算法中的更一般主题。在CSLU-OHSU生物医学计算机科学部门的一部分-几乎所有的问题都在基础或应用NLP或语音处理研究的范围内。除了通过课程工作和研究项目工作培养研究生外,该项目中创建的基础设施将有助于语音处理和NLP以及利用这些技术的应用程序的进步,包括文本和语音挖掘中的国防应用以及生物医学应用沿着。它将使至少八个资助的NSF项目和至少五个来自其他机构的项目在数据分析中追求新的方向。
英文摘要
This project aims to educate, train and equip graduate students in what is becoming a critical paradigm in speech and NLP: distributed algorithms, a paradigm pursued by Google with great success. This distributed processing paradigm represents more than just a computational convenience, but rather an approach for designing algorithms to optimize performance within such an environment, yielding massive improvements over standard algorithms directly deployed in parallel.The objectives of the current institutional infrastructure proposal are to (1) acquire an extensive (384 core) cluster of processors for use as a Hadoop cluster at the Center for Spoken Language Understanding (CSLU) at OHSU; (2) integrate the cluster within the existing computing infrastructure; and (3) develop educational resources (tutorials, lab sessions, course modules and seminars) focused on both ``how-to'' information for using the Hadoop cluster and more general topics in algorithms for distributed computing. At CSLU -- part of the Division of Biomedical Computer Science at OHSU -- nearly all of the problems are within the scope of basic or applied NLP or speech processing research.Apart from training graduate students via both course-work and work on research projects, the infrastructure created in this project will contribute to advances in speech processing and NLP and applications that make use of these technologies, including national defense applications in text and speech mining along with biomedical applications. It will enable at least eight funded NSF projects and at least five projects from other agencies to pursue novel directions in their data analysis.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金