课题基金 / 基金详情

GPU-based Machine Learning System for fundamental biological research

GPU-based Machine Learning System for fundamental biological research
用于基础生物学研究的基于 GPU 的机器学习系统
批准号:
BB/V019805/1
负责人:
Rastko Sknepnek
金额:
$51.78万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
已结题
起止时间:
2021 至 --

项目摘要

项目成果

Rastko Sknepnek的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的十年中,生物科学见证了向数据驱动研究的重大转变。因此,高性能计算已成为生命科学的标准研究工具。然而,生物数据不仅庞大,而且非常复杂。这种复杂性需要替代传统数据分析的方法。机器学习(ML)已经成为一种强大的方法,可以成功地处理复杂生物数据的分析。试图建立大象的数学模型是一项无望的任务。然而,一个三岁的孩子可以轻松地指着照片中的大象。给孩子看一幅大象的画,并告诉他画中的物体是一头大象。换句话说,她通过看大象的照片学会了识别大象,现在她可以自己识别大象了。ML在计算机上模拟学习过程。而不是建立一个精确的模式描述,计算机被“教导”去识别它们。这是对传统计算的一种范式转变,因此也有其自身的挑战。值得注意的是,大脑非常适合通过实例学习,但在长除法方面表现不佳。另一方面,计算机的设计是为了以极高的速度和精度进行数值运算。因此,在计算机上模拟一个固有的启发式过程(如学习)需要大量的计算工作,这并不奇怪。随着图形处理单元(GPU)和固态硬盘(SSD)技术的出现,必要的计算机能力已经广泛可用。然而,传统的高性能计算设施不太适合ML应用程序,这并不奇怪。机器学习已经在生物学中成功应用了二十多年。一个很好的例子是预测癌细胞暴露于药物后的生存能力。这个想法是将一种反应(例如,癌细胞是否存活)与一组特征或特征(例如,哪些基因发生了突变,药物的化学性质是什么)联系起来。在所谓的监督学习中,向机器提供大量的训练数据,这些数据包含给定输入参数的正确响应。基于这些数据,机器学会预测对新的、以前看不见的参数的反应。一个主要的挑战是,通常不容易确定合适的特性是什么。细胞非常复杂,通常不清楚哪些是决定特定反应的最相关的特征,例如,应该考虑哪些基因的突变,等等。因此,需要专家准备适当的训练集。近年来,所谓的深度学习技术通过允许机器自动从原始数据中提取关键特征,彻底改变了学习过程。这是由一组模型神经元实现的,受生物神经细胞的启发,组织在一个分层网络中(即神经网络)。信息通过网络的各层传播,使得每一层都能捕捉到数据中越来越多的抽象特征。这大大减少了对精心定制的训练集的需求,并使机器学习适用于更广泛的问题,特别是那些专家制作的训练集不可用或制作成本过高的问题。然而,深度学习ML方法需要大量的计算资源。典型的深度学习神经网络包含数十到数百层,数千个神经元,以及它们之间数十万个链接。因此,训练它们需要以TFLOPS(每秒数万亿次操作)的速度运行的硬件,并且可以以几GB/s的速度访问数据。该提案的目的是建立一个指定的基于gpu的系统,用于将深度学习机器学习方法应用于邓迪大学的基础生物学研究。
英文摘要
In the past decade, biological sciences have witnessed a major shift towards data-driven research. Consequently, high-performance computing has become a standard research tool in life sciences. Biological data, however, is not only large but it is highly complex. This complexity requires alternative approaches to conventional data analysis. Machine learning (ML) has emerged as a powerful methodology that can successfully tackle the analysis of complex biological data. It is a hopeless task to attempt to develop a mathematical model of an elephant. Yet, a three-year-old child can with ease point at an elephant in a photo. The child was shown a picture of an elephant and told that the object in the picture was an elephant. In other words, she learnt to recognise an elephant by seeing photos of it and now she can identify it on her own. ML emulates the learning process on a computer. Instead of building a precise description of patterns, the computer is "taught" to recognise them. This is a paradigm shift from conventional computing and, thus, has its own challenges. Notably, the brain is well suited for learning by example, yet it performs poorly when it comes to long divisions. Computers, on the other hand, have been designed to perform numerical operations with great speed and precision. It is, therefore, not surprising that emulating an inherently heuristic process such as learning on a computer would require a substantial computational effort. With the recent advent in Graphics Processing Unit (GPU) and Solid State Drive (SSD) technologies, the necessary computer power has become broadly available. It is, however, not surprising that traditional High-Performance Computing facilities are not well-suited for ML applications. ML has been successfully used in biology for more than two decades. An excellent example is the prediction of the viability of cancer cells when exposed to a drug. The idea is to associate a response (e.g., whether a cancer cell survives or not) to a set of characteristics or features (e.g., which genes were mutated and what chemical properties of the drug are). In the so-called supervised learning, the machine is presented with a large set of training data that contains correct responses for given input parameters. Based on that data, the machine learns to predict the response for new, previously unseen parameters. A major challenge is that it is often not easy to identify what the appropriate features are. Cells are very complex and it is often unclear which are the most relevant features that determine a specific response, e.g., mutations of which genes one should consider, etc. An expert is, therefore, required to prepare the appropriate training set. In recent years, so-called deep learning techniques have revolutionised the learning process by allowing the machine to automatically extract the key features from raw data. This is achieved by a set of model neurons, inspired by biological neural cells, organised in a layered network (i.e., a neural network). The information propagates through layers of the network, which enables each layer to capture more and more abstract features in the data. This drastically reduces the need for carefully tailored training sets and makes the ML applicable to a wider range of problems, especially those where expert-made training sets are not available or too costly to make. Deep learning ML approaches, however, require substantial computational resources. Typical deep learning neural networks contain tens to hundreds of layers, thousands of neurons, and hundreds of thousands of links between them. Training them, therefore, requires hardware that operates at TFLOPS speeds (trillions of operations per second) and can access the data at several GB/s.The aim of this proposal is to build a designated GPU-based system for applying deep-learning ML methods in fundamental biological research at the University of Dundee.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Dry Active Matter on a Sphere
  • 批准号:
    EP/M009599/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $12.07万
  • 财政年份:
    2015
  • 负责人:
    Rastko Sknepnek
  • 依托单位:
国内基金
海外基金
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Incentive and governance schenism study of corporate green washing behavior in China: Based on an integiated view of econfiguration of environmental authority and decoupling logic
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    YU BYUNGJUN
  • 依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
  • 批准号:
    W2433169
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    HAOFEI ZHANG
  • 依托单位:
含Re、Ru先进镍基单晶高温合金中TCP相成核—生长机理的原位动态研究
  • 批准号:
    52301178
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30.00万元
  • 批准年份:
    2023
  • 负责人:
    夏万顺
  • 依托单位: