Advancing Machine Learning Methodology for New Classes of Prediction Problems
Advancing Machine Learning Methodology for New Classes of Prediction Problems
批准号:
EP/F009461/2
负责人:
Guido Sanguinetti
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2010
资助国家:
英国
项目状态:
已结题
起止时间:
2010 至 --
中文摘要
在过去的几十年里,机器学习和用于数据分类的模式识别算法的发展取得了巨大的进步。这导致了许多应用领域的相当大的进步,其中一些算法形成了无处不在的部署技术的核心。然而,有许多重要的应用,例如在生物医学中,都是高度非标准的预测问题,迫切需要为这些应用开发合适而有效的分类技术。例如,在NIPS2006上,Girolami&钟报告了对蛋白质折叠分类问题的最新预测精度,仅为62%。虽然这在一定程度上可能是由于折叠类别之间的重叠,但很明显,大多数分类算法所做的一些基本假设在这一应用中是无效的。特别是,大多数算法对数据的结构做了一些现实中不满足的假设:数据(训练和测试)是来自相同分布的独立同分布(I.I.D),标签是无偏的(即正例和反例的相对比例大致平衡),输入数据和标签上的标签噪声的存在在很大程度上可以忽略。机器学习的最新进展,如基于核的方法和贝叶斯推理的有效计算方法的可用性,为以原则性的方式解决非标准情况下的分类问题提供了巨大的希望。鉴于技术进步产生新数据集的速度令人望而生畏,开发有效的分类工具变得更加紧迫。这在生命科学中尤其如此,分子生物学和蛋白质组学的进步导致了大量数据的产生,需要开发高通量自动分析方法。提高分类精度可能会消除目前这类数据分析中的瓶颈,对进一步的生物医学研究和数百万人的生活质量产生真正的影响。目前,大多数用于生命科学应用的分类器,特别是那些被部署为生物信息学Web服务的分类器,都采用和适应了传统的机器学习方法,通常是以一种特别的方式,例如使用人工神经网络和支持向量机。然而,在现实中,这些应用中的许多都是高度非标准的分类问题,因为模式分类和决策理论的许多基本基本假设(例如,对于训练数据和测试数据的相同抽样分布、离散情况下的完美无噪声标记、可以嵌入到公共特征空间中的对象表示)被违反,并且这对可实现的性能具有直接和潜在的高度负面影响。为了在广泛的重要应用方面取得亟需的重大进展,迫切需要在一个共同框架内系统地解决相关的方法学问题,而这正是当前提议的动机所在。
英文摘要
The last few decades have seen enormous progress in the development of machine learning and pattern recognition algorithms for data classification. This has resulted in considerable advances in a number of applied fields, with some of these algorithms forming the core of ubiquitous deployed technologies. However there exist very many important applications, for example in biomedicine, which are highly non-standard prediction problems, and there is an urgent need to develop appropriate & effective classification techniques for such applications. For example, at NIPS2006 Girolami & Zhong reported state of the art prediction accuracy for a protein fold classification problem which stands at a modest 62%. While this may partly be due to overlaps between classes of fold, it is also clear that some of the fundamental assumptions made by most classification algorithms are not valid in this application. In particular, most algorithms make some assumptions on the structure of the data that are not met in reality: data (both training and test) is independent and identically distributed (i.i.d) from the same distribution, labels are unbiased (i.e. the relative proportions of positive and negative examples are approximately balanced) and the presence of labeling noise both on the input data and on the labels can be largely ignored. Recent advances in Machine Learning, such as kernel based methods and the availability of efficient computational methods for Bayesian inference, hold great promise that classification problems in non-standard situations can be addressed in a principled way. The development of effective classification tools is all the more urgent given the daunting pace at which technological advances are producing novel data sets. This is particularly true in the life sciences, where advances in molecular biology and proteomics are leading to the production of vast amounts of data, necessitating the development of methods for high-throughput automated analysis. Improving classification accuracy may lead to the removal of what is currently the bottleneck in the analysis of this type of data, leading to real impact in furthering biomedical research and in the life quality of millions of people. At present most classifiers used in life sciences applications, especially those deployed as bioinformatics web services, adopt & adapt traditional Machine Learning approaches, quite often in an ad hoc manner, e.g. employing Artificial Neural Networks & Support Vector Machines. However, in reality many of these applications are highly non-standard classification problems in the sense that a number of the fundamental underlying assumptions of pattern classification and decision theory (e.g. identical sampling distributions for 'training' and 'test' data, perfect noiseless labeling in the discrete case, object representations which can be embedded in a common feature space) are violated and this has a direct and potentially highly negative impact on achievable performance. To make much needed & significant progress on a wide range of important applications there is an urgent requirement to systematically address the associated methodological issues within a common framework and this is what motivates the current proposal.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Large scale spatio-temporal point processes: novel machine learning methodologies and application to neural multi-electrode arrays.
-
批准号:EP/L027208/1
-
项目类别:Research Grant
-
资助金额:$34.23万
-
财政年份:2014
-
负责人:Guido Sanguinetti
-
依托单位:
Computational reconstruction of stochastic regulation: from transcriptional modules to network remodelling
-
批准号:BB/I024747/1
-
项目类别:Research Grant
-
资助金额:$4.73万
-
财政年份:2011
-
负责人:Guido Sanguinetti
-
依托单位:
Systems Understanding of Microbial Oxygen-Dependent and Independent Catabolism (SUMO2)
-
批准号:BB/I004777/1
-
项目类别:Research Grant
-
资助金额:$34.07万
-
财政年份:2010
-
负责人:Guido Sanguinetti
-
依托单位:
Carbon monoxide and metal carbonyl CO-releasing molecules (CORMs) as novel antimicrobial agents - a systems approach to cellular targets and effects
-
批准号:BB/H01702X/1
-
项目类别:Research Grant
-
资助金额:$34.49万
-
财政年份:2010
-
负责人:Guido Sanguinetti
-
依托单位:
Advancing Machine Learning Methodology for New Classes of Prediction Problems
-
批准号:EP/F009461/1
-
项目类别:Research Grant
-
资助金额:$10.89万
-
财政年份:2008
-
负责人:Guido Sanguinetti
-
依托单位:
国内基金
海外基金
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位: