ProteomeHarvest - Excel/XML Bridge for User-friendly Proteomics Data Collection
ProteomeHarvest - Excel/XML Bridge for User-friendly Proteomics Data Collection
批准号:
BB/E00573X/1
负责人:
Rolf Apweiler
金额:
$6.4万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2006
资助国家:
英国
项目状态:
已结题
起止时间:
2006 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Today, scientific experiments in molecular biology in general and in proteomics in particular, are often done on a large scale, producing large numbers of individual data items. These large data sets are then the basis of scientific publications. Often only relatively few results are actually contributing to the final conclusions reached by the researcher, but the complete result sets can provide valuable knowledge to other researchers comparing them to their own results. However, to allow others to understand how the experiments were done, they need to be described in a very detailed manner. To avoid 'comparing apples and pears', this discription needs to be done in a systematic manner, using established rules or standards on how to describe experiments. In addition, the data needs to be easily accessible for other researchers, which can be best achieved by entering it into large databases, accessible over the internet. Overall, a lot of effort is needed to describe a large experiment in the detailed, standardised manner which allows others to understand them. In other projects, we are working on setting up common rules for the description of proteomics experiments. However, even the best rules are useless if they are not applied. As scientists, like everybody else, tend to do only the minimum amount of work to achieve their goals, their experiment description is often incomplete, focussing only on the aspects they consider relevant. And of course they tend to quickly tire of properly entering the data into databases if they have to use complicated tools they have to install on their computer, and with which they are not familiar. On the other hand, there are programs they know well, because the use them almost every day anyway to manage their data. The main purpose of this proposal is to use one such tool, Microsoft Excel, to develop forms which allow scientists to enter their results into a database in as easy a manner as possible. Biologists are used to Excel, they are familiar with its functionality, and they nearly always have it installed on their computer anyway. We plan to develop Excel forms which are as user friendly as possible, but still capture all the necessary data to appropriately describe the results of a large experiment, according to established rules and standards. While Excel is often used to store experiment results, this is often done in a very unsystematic manner, and it is usually very difficult to transfer the data into XML, a file format which is nowadays practically the standard way for entering data into databases. Also, so far it has been difficult to use and regularly update controlled vocabularies in Excel. Controlled vocabularies are lists of possible words which can be entered in a specific field in a form, to avoid typing errors, and to ensure everybody uses the same word for the same thing. In this project, we propose to develop advanced Excel forms for proteomics data harvesting. These forms should provide researchers with an easy tool to store their data in a systematic manner, ready for sending it to a database. These forms will be able to communicate with a database on the internet to provide up-to-date controlled vocabularies, and they will be able to directly send the data in the form of XML to a database on the internet. We will develop and test these forms for the existing PRIDE proteomics database, making use of the existing database for data storage, and using OLS, the ontology lookup service developed as part of PRIDE, to keep controlled vocabularies in the Excel forms up to date. By providing Excel forms as a user-friendly way to store proteomics data and send it to public databases, we hope to convince researchers to invest a little bit of extra effort to make their valuable data accessible to their collegues by sending it to public databases, and thus to maximise the use of data paid for by the tax payer anyway.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ARGENT: ARgentinian GEnomics for Tuberculosis
-
批准号:EP/T015446/1
-
项目类别:Research Grant
-
资助金额:$120.62万
-
财政年份:2019
-
负责人:Rolf Apweiler
-
依托单位:
Database on demand - creating customized sequence databases for efficient protein identification
-
批准号:BB/F016255/1
-
项目类别:Research Grant
-
资助金额:$6.16万
-
财政年份:2008
-
负责人:Rolf Apweiler
-
依托单位:
Embracing new technologies to streamline improve and sustain InterPro and its contributing databases
-
批准号:BB/F010508/1
-
项目类别:Research Grant
-
资助金额:$86.45万
-
财政年份:2008
-
负责人:Rolf Apweiler
-
依托单位:
Further development of the QuickGO web interface for browsing and retrieving Gene Ontology Annotation data
-
批准号:BB/E023541/1
-
项目类别:Research Grant
-
资助金额:$10.83万
-
财政年份:2007
-
负责人:Rolf Apweiler
-
依托单位:
海外基金