Embracing new technologies to streamline improve and sustain InterPro and its contributing databases
Embracing new technologies to streamline improve and sustain InterPro and its contributing databases
批准号:
BB/F010508/1
负责人:
Rolf Apweiler
金额:
$86.45万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2008
资助国家:
英国
项目状态:
已结题
起止时间:
2008 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
New DNA sequencing technologies have led to a flood of new data in sequence databases being submitted by individual scientists, genome sequencing projects and metagenomics projects. These sequences enter the databases with little or no annotation, limiting their usefulness to the scientific community. This has inspired the development of new tools for automatic annotation of the encoded protein sequences. One of the most successful developments in this area has been in the production of so-called protein 'signatures', diagnostic methods that are able to characterise newly-determined sequences in terms of the protein families to which they belong and/or the structural or functional domains they contain. Protein signature approaches have been adopted by a number of databases, and ten of the top such resources are integrated into the InterPro database. InterPro, and its accompanying protein analysis software tool, InterProScan, is now one of the leading protein functional classification resources in the world. However, despite its success, InterPro and its partners are currently suffering from a lack of financial support. The level of funding required to maintain and improve a database of this size is often underestimated. The amount of incoming data is increasing exponentially, and databases now struggle to provide their data to the public in a timely way, while at the same time maintaining the necessary high standards of data quality. Moreover, as they become more popular, and user demands increase, these core databases endure mounting pressure not only to keep up with the expanding volume of data and growing community requirements, but also to be early adopters of newly emerging technologies. This proposal aims to resolve these issues by embracing new technologies to enhance and further develop InterPro and its source databases. It aims to streamline production processes both to provide more regular data releases and to better cope with increased volumes of data. With more formalised Consortium activities and coordination thereof, we will make more efficient use of resources and share tasks to ensure long-term sustainability of the databases. Specifically we aim to: - Streamline data production procedures to enable a faster turn-around time for releasing the data; - Develop and integrate new annotation tools and standards to make the rate-limiting annotation step quicker and easier, and share tasks, such as annotation, to remove redundancy in effort; - Work closely together to improve quality-assurance procedures for protein matches; - Coordinate the upgrade of InterProScan and other HMM-based databases to the latest HMMer version; - Improve the InterProScan protein domain-finding software; - Exploit new technologies for database linking and data exchange; and - Extend the functionality of the Web interface to better meet the needs of the user community. The planned improvements to InterProScan and the protein match procedures will improve the quality, as well as the speed of protein functional classification; streamlining the production processes will enable the databases to get new protein domains and families out to the public as soon as they become available. New technologies will facilitate easier linking between different databases, and will provide the public with access to data from different sources. They will also open the door to more complex analyses, by providing improved programmatic access to the data. In addition, these new processes and technologies will allow InterPro and its member databases to cope with the ever-increasing flood of new data and make it accessible to the public in more regular releases. Ultimately, these improvements will make InterPro and its partners easier and more efficient to maintain, paving the way to a more sustainable future and increasing their benefit and usefulness to the scientific community.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1093/bioinformatics/btu031
发表时间:
2014-05-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
[Jones P, Binns D, Chang HY, Fraser M, Li W, McAnulla C, McWilliam H, Maslen J, Mitchell A, Nuka G, Pesseat S, Quinn AF, Sangrador-Vegas A, Scheremetjew M, Yong SY, Lopez R, Hunter S]
通讯作者:
Hunter S
DOI:
10.1093/nar/gkn785
发表时间:
2009-01
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Hunter S, Apweiler R, Attwood TK, Bairoch A, Bateman A, Binns D, Bork P, Das U, Daugherty L, Duquenne L, Finn RD, Gough J, Haft D, Hulo N, Kahn D, Kelly E, Laugraud A, Letunic I, Lonsdale D, Lopez R, Madera M, Maslen J, McAnulla C, McDowall J, Mistry J, Mitchell A, Mulder N, Natale D, Orengo C, Quinn AF, Selengut JD, Sigrist CJ, Thimma M, Thomas PD, Valentin F, Wilson D, Wu CH, Yeats C]
通讯作者:
Yeats C
DOI:
10.1093/database/bar033
发表时间:
2011
期刊:
Database : the journal of biological databases and curation
影响因子:
--
作者:
[Jones P, Binns D, McMenamin C, McAnulla C, Hunter S]
通讯作者:
Hunter S
DOI:
10.1093/database/bar068
发表时间:
2012
期刊:
Database : the journal of biological databases and curation
影响因子:
--
作者:
[Burge S, Kelly E, Lonsdale D, Mutowo-Muellenet P, McAnulla C, Mitchell A, Sangrador-Vegas A, Yong SY, Mulder N, Hunter S]
通讯作者:
Hunter S
DOI:
10.1093/nar/gkr948
发表时间:
2012-01
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Hunter S, Jones P, Mitchell A, Apweiler R, Attwood TK, Bateman A, Bernard T, Binns D, Bork P, Burge S, de Castro E, Coggill P, Corbett M, Das U, Daugherty L, Duquenne L, Finn RD, Fraser M, Gough J, Haft D, Hulo N, Kahn D, Kelly E, Letunic I, Lonsdale D, Lopez R, Madera M, Maslen J, McAnulla C, McDowall J, McMenamin C, Mi H, Mutowo-Muellenet P, Mulder N, Natale D, Orengo C, Pesseat S, Punta M, Quinn AF, Rivoire C, Sangrador-Vegas A, Selengut JD, Sigrist CJ, Scheremetjew M, Tate J, Thimmajanarthanan M, Thomas PD, Wu CH, Yeats C, Yong SY]
通讯作者:
Yong SY
ARGENT: ARgentinian GEnomics for Tuberculosis
-
批准号:EP/T015446/1
-
项目类别:Research Grant
-
资助金额:$120.62万
-
财政年份:2019
-
负责人:Rolf Apweiler
-
依托单位:
Database on demand - creating customized sequence databases for efficient protein identification
-
批准号:BB/F016255/1
-
项目类别:Research Grant
-
资助金额:$6.16万
-
财政年份:2008
-
负责人:Rolf Apweiler
-
依托单位:
Further development of the QuickGO web interface for browsing and retrieving Gene Ontology Annotation data
-
批准号:BB/E023541/1
-
项目类别:Research Grant
-
资助金额:$10.83万
-
财政年份:2007
-
负责人:Rolf Apweiler
-
依托单位:
ProteomeHarvest - Excel/XML Bridge for User-friendly Proteomics Data Collection
-
批准号:BB/E00573X/1
-
项目类别:Research Grant
-
资助金额:$6.4万
-
财政年份:2006
-
负责人:Rolf Apweiler
-
依托单位:
国内基金
海外基金
登录
查看更多内容
脊髓新鉴定SNAPR神经元相关环路介导SCS电刺激抑制恶性瘙痒
-
批准号:82371478
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:焦英甫
-
依托单位:
tau轻子衰变与新物理模型唯象研究
-
批准号:11005033
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2010
-
负责人:李文君
-
依托单位:
HIV gp41的NHR区新靶点的确证及高效干预
-
批准号:81072676
-
项目类别:面上项目
-
资助金额:33.0万元
-
批准年份:2010
-
负责人:戴秋云
-
依托单位:
强子对撞机上新物理信号的多轻子末态研究
-
批准号:10675110
-
项目类别:面上项目
-
资助金额:36.0万元
-
批准年份:2006
-
负责人:蒋一
-
依托单位: