EPSRC Centre for Doctoral Training in Distributed Algorithms: the what, how and where of next-generation data science
EPSRC Centre for Doctoral Training in Distributed Algorithms: the what, how and where of next-generation data science
批准号:
EP/S023445/1
负责人:
金额:
$623.88万
依托单位:
依托单位国家:
英国
项目类别:
Training Grant
财政年份:
2019
资助国家:
英国
项目状态:
未结题
起止时间:
2019 至 --
中文摘要
该CDT将培训一批60名学生,使他们具备成为分布式算法领导者的技能和经验:利用“未来计算系统”,走向“数据驱动的未来”。这激发了当今对训练有素的数据科学家的迫切需求。这个CDT将为未来的数据科学领导者提供支持。英国(和世界)需要能够最好地利用未来计算资源的数据科学家来收获新的“石油”:数据中存在的信息。随着我们毕业生的职业发展,许多核心架构将变得越来越普遍。我们预计,未来的台式机将比现在多出数百万个内核。这一核心计数将挑战当前大数据中间假设(例如,Spark和TensorFlow)认为,未来计算系统的细节可以与数据科学工具和技术的开发解耦。更具体地说,数据科学家必须了解如何设计算法,以便在数据移动成为关键性能瓶颈的环境中有效运行。我们将提供培训,确保我们产生高度-·既了解未来计算机硬件的设计,又了解如何以及何时将算法解决方案调整到最佳状态的可雇用人员。从一开始,学生们就将被嵌入到一个计算环境中,这个计算环境预测着他们毕业后将到达他们桌子上的硬件资源,而不是今天存在的硬件。学生群体提供了激励与国际领先的超级计算中心合作的临界质量:STFC的Hartree中心是团队的一个组成部分;我们与美国IBM研究院建立的联系将为学生提供最先进的计算硬件。这种对未来计算能力的预期将确保我们的毕业生具有很高的就业能力,但也有助于激励最终用户组织参与CDT。我们已经确定了跨越两个主题的最终用户组织:国防和安全;制造业。这些主题的机构分别由表现需求和效率要求驱动。我们将根据群组,主题和个人的需要提供培训。每个奖学金将有两名学术主管(一名与“未来计算系统”保持一致,另一名与“走向数据驱动的未来”保持一致)和至少一名来自项目合作伙伴的主管。这个监督团队将共同定义每个助学金的范围。一旦高质量的学生被选中并招募,我们将与学生合作,以确定符合他们的需求和学生的具体要求的培训。我们提供的培训将包括与“未来计算系统”和“迈向数据驱动的未来”优先领域相关的培训需求。我们将使用客座讲座,例如,IBM(用于培训快速通道公务员)和加州大学伯克利分校,以确保我们最大限度地提高我们的毕业生的能力,茁壮成长,成为明天的领导者分布式算法。
英文摘要
This CDT will train a cohort of 60 students to have the skills and experience that enables them to become leaders in Distributed Algorithms: capitalising on "Future Computing Systems" to move "Towards a Data-Driven Future".Commodity Data Science is already pervasive. This motivates today's pressing need for highly-trained data scientists. This CDT will empower tomorrow's leaders of data science. The UK (and world) needs data scientists that can best exploit tomorrow's computational resources to harvest the new 'oil': the information present in data.As our graduates' careers progress, many cored architectures will become increasingly commonplace. We anticipate millions more cores in tomorrow's desktops than today's. This core count will challenge the assumption made by current Big Data middleware (e.g., Spark and TensorFlow) that the details of future computing systems can be decoupled from the development of data science tools and techniques. More specifically, it will become imperative that data scientists understand how to design algorithms that can operate effectively in environments where data movement is the key performance bottleneck.To meet this need, we will provide training that ensures we generate highly-employable individuals who have both an understanding of the design of future computer hardware as well as an understanding of how and when to flex the algorithmic solutions to best exploit the computational resources that will exist in the future.From the outset, the students will be embedded in a computing environment that anticipates the hardware resources that will arrive on their desks after they graduate, not the hardware that exists today. The cohort of students provides the critical mass that motivates engagement with internationally-leading supercomputing centres: STFC's Hartree Centre is an integral part of the team; links we have established with IBM Research in the US will provide students with access to state-of-the-art computing hardware. This anticipation of future computing capability will ensure our graduates are highly employable, but also help motivate end-user organisations to engage with the CDT.We have identified such end-user organisations that span two themes: defence and security; manufacturing. Organisations in these themes are driven by performance demands and efficiency requirements respectively.We will align the training we provide with the needs of the cohort, the theme and the individual. Each studentship will have two academic supervisors (one aligned with the "Future Computing Systems" and one aligned with moving "Towards a Data-Driven Future") and at least one supervisor from a project partner. This supervisory team will co-define the scope of each studentship. Once the high quality student has been selected and recruited, we will work with the student to define the training that aligns with their needs and the specific demands of the studentship. Our training provision will include the training needs associated with both the "Future Computing Systems" and "Towards a Data-Driven Future" priority areas. We will use guest lectures from, for example, IBM (as used to train Fast Track civil servants) and UC Berkeley to ensure we maximise our graduates' ability to thrive and to become tomorrow's leaders in Distributed Algorithms.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金