课题基金 / 基金详情

New challenges in high-dimensional statistical inference

New challenges in high-dimensional statistical inference
高维统计推断的新挑战
批准号:
EP/J017213/1
负责人:
Richard Samworth
金额:
$151.65万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2012
资助国家:
英国
项目状态:
已结题
起止时间:
2012 至 --

项目摘要

项目成果

Richard Samworth的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
As a society, more and more of the activities that we take for grantedrely on sophisticated technology, and are dependent on the fast andefficient handling of large quantities of data. Obvious examplesinclude the use of internet search engines and mobile telephones.Similarly, recent advances in healthcare are partly due to improved,highly data-intensive scanning equipment in hospitals, and thedevelopment of new, effective drug treatments, which have been theresult of extensive scientific study with data at its core.Nevertheless, such advances can only be achieved through thedevelopment of appropriate statistical models and methods which enablepractitioners to extract useful information from these vast quantitiesof data. In order to capture the complexity of the data generatingprocesses, these models are inevitably high-dimensional, and have beenthe topic of an enormous amount of research in Statistics over thelast 15 years or so.This proposal addresses some of the fundamental and important challenges in handling the huge data sets that routinely arise in the applications above, as well as many others. For instance, in high-dimensional models, sparse estimators are crucial for stability andinterpretability. But these give only a point estimate of aparameter, and typically practioners require more sophisticatedinferential statements to assess uncertainty. We will show how thisby done by proposing easy-to-use and robust p-valuesbased on these sparse estimators.One of the most important applications of sparse estimators is inbiotechnology. Indeed, we will apply our methodology described abovein a high-dimensional cancer study carried out by Danishbiostatisticians that uses microarray techniques. We will select,with an associated quantification of uncertainty, a handful of stabledistinguishing genes for diffuse large B-cell lymphomas, therebyenhancing our understanding of these cancers.Another application area facing high-dimensional challenges isneuroscience, and we will work on a study of dyscalculia that uses thebrain imaging technique of Electroencephalography (EEG). Dyscalculiais a mathematical disability that prevents normal arithmetic function.Here, existing statistical techniques used by experimentalpsychologists in this area are inadequate, and modern high-dimensionalmethods have the potential to improve dramatically our understandingof this disability.In classification problems, the challenge is to assign an observationto one of two or more classes based on its similarity to (labelled)data from each of these classes. They are some of the most frequentlyencountered high-dimensional statistical problems, particularly infields such as machine learning and areas of computer science such ascomputer vision and robotics. We will provide a simple and robustimprovement to perhaps the most popular method (the k-nearestneighbour classifier), by weighting the nearest neighbours in anoptimal fashion. We will also give a quantification of theimprovement. A related problem we will study is to quantify theuncertainty of a classifier constructed from training data. As anexample this could be used to give a doctor a measure of uncertaintyin a diagnosis.The final main issue we will address concerns model misspecification.This is a particularly important issue in high-dimensional statisticalproblems, where it is almost inevitable that our model misses someimportant effects, or does not model them in the correct way. We willprovide understanding of how statistical procedures perform in suchcircumstances and develop new ones that are robust to modelmisspecification. A particularly important application will be toIndependent Component Analysis models that are very popular instatistical signal processing for analysing data arising from multiplesources, including microarray and brain imaging data.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.17863/cam.8925
发表时间: 2017
期刊:
影响因子: --
作者: [Cannings T]
通讯作者: Cannings T
DOI: 10.1093/biomet/asz024
发表时间: 2019-09-01
期刊: BIOMETRIKA
影响因子: 2.7
作者: [Berrett, T. B., Samworth, R. J.]
通讯作者: Samworth, R. J.
DOI: 10.1214/19-aos1852
发表时间: 2020-06-01
期刊: ANNALS OF STATISTICS
影响因子: 4.5
作者: [Barber, Rina Foygel, Candes, Emmanuel J., Samworth, Richard J.]
通讯作者: Samworth, Richard J.
EFFICIENT MULTIVARIATE ENTROPY ESTIMATION VIA k-NEAREST NEIGHBOUR DISTANCES
通过 k 最近邻距离进行高效的多元熵估计
DOI: 10.17863/cam.17905
发表时间: 2019
期刊:
影响因子: --
作者: [Berrett T]
通讯作者: Berrett T
Statistical methodology and theory for the Big Data era (Ext.)
  • 批准号:
    EP/P031447/1
  • 项目类别:
    Fellowship
  • 资助金额:
    $77.84万
  • 财政年份:
    2017
  • 负责人:
    Richard Samworth
  • 依托单位:
国内基金
海外基金
Supply Chain Collaboration in addressing Grand Challenges: Socio-Technical Perspective
  • 批准号:
    --
  • 项目类别:
    外国青年学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    Lim Jia Jia
  • 依托单位:
Navigating Sustainability: Understanding Environm ent,Social and Governanc e Challenges and Solution s for Chinese Enterprises in Pakistan's CPEC Framew ork
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    Noshaba Aziz
  • 依托单位: