课题基金 / 基金详情

BIGDATA: Low-Memory Streaming Prefilters for Biological Sequencing Data

BIGDATA: Low-Memory Streaming Prefilters for Biological Sequencing Data
BIGDATA:生物测序数据的低内存流预过滤器
批准号:
8599821
负责人:
C. Titus BROWN
金额:
$24.99万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-07-19 至 2016-05-31

项目摘要

项目成果

C. Titus BROWN的其他基金

相似基金

相关文献

中文摘要
翻译
描述(申请人提供):我们很快就能对整个细菌群落以及肿瘤的每个单个细胞的DNA和RNA进行详尽的测序。这两种非常不同的测序应用都需要快速有效地对大量噪声序列数据(数十到100兆兆字节)进行排序,以将信号从噪声中分离出来,并产生生物学洞察力。然而,目前用于从这些数据中提取信息的生物信息学方法不能轻松地处理正在获得的海量数据。处理这些序列数据的主要挑战有两个:相对较高的错误率,每个碱基0.1-1\%,以及我们现在可以使用llLumina HiSeq等测序仪轻松获取的数据量。多年来,测序能力每6个月翻一番,远远快于计算能力。由于几乎所有现有的生物信息学分析方法都需要对原始数据进行多次遍历,并且许多分析算法尚未并行化,因此生物信息学分析能力继续远远落后于数据生成能力。此外,许多现有的软件包不能很容易地重新装备以利用许多核心或GPU算法,因此将不能利用计算能力和网络基础设施方面的预期进步。我们提议开发和实施用于鸟枪测序数据中的损失压缩和错误连接的新型流方法。我们的算法很少通过($<$2),不需要特定于样本的信息,可以在固定或低内存中实现;此外,它们服从并行化,可以在许多核心环境中高效运行。当作为现有分析包的预过滤器实现时,我们的方法将消除或纠正数据集中的大多数错误,大大减少使用现有包进行下游分析的计算空间和时间要求。此外,我们将通过将纠错方法扩展到mRNAseq和元基因组数据集来提供新的能力。智力优势:我们将开发一系列算法,用于对短读DNA和RNA序列数据进行空间和时间高效的压缩和纠错。这些策略将大大提高许多下游分析应用的可扩展性,从元基因组的社区分析到人类的重新测序分析。我们将提供描述空间和时间效率与敏感性之间权衡的分析,并提供我们方法的经过测试的、有文档记录的参考实现,社区可以使用这些实现进行实际评估,并将其纳入分析工具。我们的方法将通过将高效和有效的流方法引入两种最常见的短读分析类型,即映射和组装,从而显著影响短读序列分析。
英文摘要
DESCRIPTION (provided by applicant): We will soon be able to exhaustively sequence the DNA and RNA of entire communities of bacteria, as well as every individual cell of a tumor. Both of these very different applications of sequencing share in the need to rapidly and efficiently sor through large amounts of noisy sequence data (dozens to 100s of terabases) to separate signal from noise and produce biological insight. However, current bioinformatics approaches for extracting information from this data cannot easily handle the vast amounts of data being acquired. The primary challenges in processing this sequence data are twofold: the relatively high error rate of 0.1-1\%, per base, and the volume of data we can now easily acquire with sequencers such as lllumina HiSeq. For years, sequencing capacity has been doubling every 6 months -significantly faster than compute capacity. Since almost all extant bioinformatics analysis approaches require multiple passes across the primary data, and many analysis algorithms have not been parallelized, bioinformatics analysis capacity continues to lag ever further behind data generation capacity. In addition, many of the existing software packages cannot easily be retooled to take advantage of many core or GPU algorithms, and hence will not take advantage of expected advances in compute capacity and cyber infrastructure we propose to develop and implement novel streaming approaches for loss compression and error connection in shotgun sequencing data. Our algorithms are few-pass ($<$ 2), require no sample-specific information, and can be implemented in fixed or low memory; moreover, they are amenable to parallelization and can run efficiently in many core environments. When implemented as a prefilter to existing analysis packages, our approaches will eliminate or correct the majority of errors in data sets, dramatically reducing the computational space and time requirements for downstream analysis using existing packages. Moreover, we will provide novel capability by extending error correction approaches to mRNAseq and metagenomic data sets. Intellectual Merit: We will develop a range of algorithms for space- and time-efficient compression and error correction of short-read DNA and RNA sequence data. These strategies will substantially increase the scalability of many downstream analysis applications, ranging from community analysis of metagenomes to resequencing analysis of humans. We will provide analyses describing the tradeoffs between space and time efficiency and sensitivity, and deliver tested, documented reference implementations of our approaches that can be used by the community for practical evaluation and incorporation into analysis tools. Our approaches will significantly impact short-read sequence analysis by introducing efficient and effective streaming approaches to the two most common types of short-read analysis, mapping and assembly.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Tools and Workflows for Mining Genomic Data on Many Clouds
BIGDATA: Low-Memory Streaming Prefilters for Biological Sequencing Data
  • 批准号:
    8703739
  • 项目类别:
  • 资助金额:
    $20.42万
  • 财政年份:
    2013
  • 负责人:
    C. Titus BROWN
  • 依托单位:
Analyzing Next-Generation Sequencing Data
  • 批准号:
    8150859
  • 项目类别:
  • 资助金额:
    $5.4万
  • 财政年份:
    2011
  • 负责人:
    C. Titus BROWN
  • 依托单位:
Analyzing Next-Generation Sequencing Data
  • 批准号:
    8551251
  • 项目类别:
  • 资助金额:
    $5.4万
  • 财政年份:
    2011
  • 负责人:
    C. Titus BROWN
  • 依托单位:
国内基金
海外基金
Segmented Filamentous Bacteria激活宿主免疫系统抑制其拮抗菌 Enterobacteriaceae维持菌群平衡及其机制研究
  • 批准号:
    81971557
  • 项目类别:
    面上项目
  • 资助金额:
    65.0万元
  • 批准年份:
    2019
  • 负责人:
    毛开睿
  • 依托单位:
电缆细菌(Cable bacteria)对水体沉积物有机污染的响应与调控机制