POWRE: Combining Data Mining and Information Visualization Techniques with a Molecular Biology Sequence Similarity Database System
POWRE: Combining Data Mining and Information Visualization Techniques with a Molecular Biology Sequence Similarity Database System
批准号:
9753283
负责人:
Elizabeth Shoop
金额:
$7.06万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1998
资助国家:
美国
项目状态:
已结题
起止时间:
1998-01-01 至 1999-12-31
中文摘要
这个项目的主要目标是帮助基因组研究人员在大量的生物数据中阐明模式和集群的任务。对于对将基因或蛋白质序列与一个基因组内或跨基因组的序列进行比较感兴趣的基因组研究人员来说,这项任务涉及执行数十万次相似性搜索,从而产生文本输出。该项目涉及开发两个特定的软件工具,用于可视化和探索生物序列相似性结果数据库中的相似性数据。第一个工具将是交互式分类工具。该工具将在2D散点图中显示选定相似性数据库对象的属性,并允许对显示进行动态操作。这将使基因组研究人员能够探索相似性的属性,并根据这些属性对相似性进行分类。例如,基因组研究人员将能够改变用于计算每个检测到的相似性的强度的函数的输入参数,并以每个点的颜色显示每个相似性的强度,并且基于分数和统计显著性将位于2D空间中的点显示为X轴和Y轴。该工具将使基因组研究人员能够动态操纵更高级别概念或类别的生成,以检测到相似性(强、边缘和弱相似性,而不是更难比较的具有特定值和统计意义的个体相似性)。这将导致他们能够基于检测到的相似性的各种属性将命中分类为同源或同源。分数和p值不是唯一可以使用的属性--系统足够通用,可以在函数中使用其他属性,如同一性百分比、保守率和比对长度等。因此,基因组研究人员可以在基因组比较研究过程的不同阶段继续进行探索。第二个工具将是集群探索工具。使用将相似序列聚在一起的数据挖掘技术的结果,基因组研究人员将能够可视化聚类中序列之间的相似性。例如,该工具可用于发现与一组已知序列的成员相似的一组新的未知序列。新序列可以定位为二部图中左侧的节点,与其相似的已知序列可以沿右侧定位。节点之间绘制的线条根据命中的强度而有不同的颜色,这将使研究人员能够可视化集群中序列的连通性。关于聚类中每个序列和每个相似性的详细信息可以从DBMS中获得。这将使基因组研究人员能够研究一组同源或同源序列。这些工具的一个关键特征是,它们将是“瘦”客户端(通常称为小程序),通过基因组研究人员可视化地制定的查询与底层数据库管理系统进行通信。将基于Java的组件用于这些工具将使生物信息学社区和基因组研究社区能够轻松地使用和共享这些工具。这些工具的开发将证明瘦客户机方法的可行性,而瘦客户机方法是网络计算体系结构理念的标志。
英文摘要
The main objective of this project is to aid genome researchers with the task of elucidating patterns and clusters in large amounts of biological data. For genome researchers who are interested in comparing gene or protein sequences to the sequences within one genome or across genomes, this task involves executing hundreds of thousands of similarity searches that produce text output. This project involves the development of two specific software tools for visualizing and exploring the similarity data in a database of biological sequence similarity results. The first tool will be an Interactive Categorization Tool. This tool will display attributes of selected similarity database objects in a 2D scatterplot and enable dynamic manipulation of the display. This will enable the genome researcher to explore the attributes of similarities and categorize the similarities based on those attributes. For example, the genome researcher will be able to vary the input parameters of a function for computing the strength of each detected similarity and display a plot with the strength of each similarity shown as the color of each point, and the points situated in the 2D space based on score and statistical significance as the X and Y axes. The tool will enable genome researchers to dynamically manipulate the generation of higher- level concepts or categories for detected similarities (strong, marginal, and weak similarities as opposed to individual similarities with particular values of score and statistical significance that are more difficult to compare). This will lead to their ability to categorize hits as orthologous or paralogous, based on various attributes of the detected similarities. Score and p-value are not the only attributes that can be used -- the system is general enough that other attributes, such as percent identity, percent conserved, and length of alignment, among others, could be used in functions. Thus, genome researchers can cond uct exploration at different stages of the genome comparison research process. The second tool will be a Cluster Exploration Tool. Using the results from data mining techniques that cluster like sequences together, genome researchers will be able to visualize the similarities among the sequences in the clusters. For example, the tool can be used for a cluster of new unknown sequences that were found similar to members of a group of known sequences. The new sequences can be positioned as nodes on the left in a bipartite graph, and the known sequences that they are similar to can be positioned along the right. Lines drawn between the nodes, colored differently based on the strength of the hits, will enable the researcher to visualize the connectedness of the sequences in the cluster. Details about each sequence and each similarity in the cluster can be obtained from the DBMS. This will enable genome researchers to study groups of orthologous or parologous sequences. A key feature of these tools is that they will be 'thin' clients (often referred to as applets) that communicate with the underlying DBMS via queries formulated visually by the genome researchers. The use of Java- based components for these tools will enable them to be easily used and shared by the bioinformatics community and the genome research community. The development of these tools will demonstrate the feasibility of the thin-client approach that is the hallmark of the network computing architecture philosophy.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CS in Parallel: Scaling an Incremental Modular Approach to Injecting Parallel Computing Throughout CS Curricula
-
批准号:1225796
-
项目类别:Standard Grant
-
资助金额:$23.59万
-
财政年份:2012
-
负责人:Elizabeth Shoop
-
依托单位:
Collaborative Research: CCLI-Responding to manycore: A strategy for injecting parallel computing education throughout the computer science curriculum
-
批准号:0941962
-
项目类别:Standard Grant
-
资助金额:$6.92万
-
财政年份:2010
-
负责人:Elizabeth Shoop
-
依托单位:
Into the Community: Changing Perceptions and Increasing Participation in Computer Science
-
批准号:0850106
-
项目类别:Standard Grant
-
资助金额:$58.65万
-
财政年份:2009
-
负责人:Elizabeth Shoop
-
依托单位:
海外基金