Theoretical guarantees on functional Lloyd's algorithm and functional k-means clustering
Theoretical guarantees on functional Lloyd's algorithm and functional k-means clustering
批准号:
EP/W003716/1
负责人:
Yi Yu
金额:
$3.85万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2022
资助国家:
英国
项目状态:
已结题
起止时间:
2022 至 --
中文摘要
功能数据分析是一个统计领域,以功能、图像、形状甚至更一般的随机对象的形式分析数据。由于现代技术的进步,越来越多的此类数据被常规收集;由于计算机器学习的快速发展,这些数据正在被分析,并在许多方面影响着我们的生活。例如,当使用摄像头捕捉你的面部特征或传感器读取你的指纹来解锁智能手机时,你的智能手机正在收集图像,检测信号并将其与预先存储的信息进行比较。另一个例子是,了解注意缺陷多动障碍(ADHD)亚型的一种方法是研究胼胝体的形状,这通常可以作为诊断的指导。尽管分析功能数据具有生命力和重要性,但在将统计方法应用于功能数据时,往往缺乏方便的理论保证。在没有理论保证的情况下,解释分析结果是有误导性的,而且可能是危险的。功能数据分析的最新理论发展,特别是功能聚类方法,受到以下问题的困扰。1、现有文献大多依赖于泛函主成分分析,将无限维协方差算子映射到低维空间,将无限维泛函空间中的分析转化为可管理的空间。然而,这种转换的成功依赖于协方差算子的非零特征值的数量存在上界的假设。这是一个强条件,因为它排除了许多标准泛函空间,例如Sobolev空间。现有的大多数理论结果都是渐近的,也就是说,这些结果说明了一些统计过程的渐近性能,而没有详细说明这些过程达到理想速率的速度有多快,或者需要多大的样本量才能达到一定的精度水平。缺乏固定样本结果也阻碍了高维数据的分析。在这个研究计划中,我将从一个具体的问题开始——为泛函劳埃德算法的收敛性提供理论保证,这是k-means聚类方法的默认值。有了这个,我将提供函数k-means聚类方法的误差控制的固定样本版本。这两个步骤的成功将为在流形学习中对更一般的对象进行聚类提供理论保证,这将是进一步计划的起点。这个议程似乎是标准的,因为k-means聚类方法在许多应用领域都是标准的和方便的。但是在函数空间中,对于收敛性没有理论保证,对于算法何时收敛没有理论认识,更不用说知道最终的聚类估计量有多好,需要多少次迭代,需要多少样本,什么样的函数采样方案是最好的。这项工作将为这些问题提供答案。
英文摘要
Functional data analysis is a statistical area analysing the data in the form of functions, images, shapes or even more general random objects. Thanks to the advance of modern technology, more and more such data are being routinely collected; and thanks to the fast improvement of computational machine learning, these data are being analysed and influencing our life in many aspects. For instance, when unlocking a smart phone using cameras capturing your facial characteristics or sensors reading your finger prints, your smart phone is collecting images, detecting the signals and comparing them to the pre-stored information. As another example, one way to understand the subtypes of the attention deficit hyperactivity disorder (ADHD) is to study the shapes of Corpus Callosum, which often serve as a guidance on diagnosing. Despite the vitality and importance of analysing functional data, the theoretical guarantees of handy statistical methods are often lacking when applying them to functional data. Without theoretical guarantees, interpreting the analysis results is misleading and can be dangerous. The state-of-the-art theoretical developments in functional data analysis, especially functional clustering methods, are suffering from the following issues.1, The majority of the exiting literature relies on the functional principal component analysis, which maps the infinite-dimensional covariance operator to a low-dimensional space, and the analysis in the infinite-dimensional functional space is transformed to a manageable space. However, the success of such transformation relies on the assumption that there is an upper bound on the number of non-zero eigenvalues of the covariance operator. This is a strong condition, since it excludes many standard functional spaces, e.g. Sobolev spaces.2, The majority of the existing theoretical results are asymptotic, in the sense that the results state the asymptotic performances of some statistical procedures, without detailing how fast these procedures reach a desirable rate, or how large the sample size needs to be in order to reach a certain accuracy level. Lacking fixed sample results also hinders the analysis of high-dimensional data. In this research proposal, I will start with a specific problem -- providing theoretical guarantees of the convergence of the functional Lloyd's algorithm, which is the default of the k-means clustering method. With this in hand, I will then provide fixed sample version of the error controls of the functional k-means clustering methods. The success of these two steps will shed light on providing theoretical guarantees on clustering more general objects in manifold learning, which will be the starting point of a further programme. The agenda seems standard, because k-means clustering method is standard and handy in many application areas. But in the functional spaces, there is no theoretical guarantee on the convergence, no theoretical understanding on when the algorithms should converge, not to mention knowing how good the final clustering estimators are, how many iterations are needed, how many samples are needed, what kind of function sampling schemes is the best. This work will provide an answer to these questions.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
--
发表时间:
2020-12
期刊:
J. Mach. Learn. Res.
影响因子:
--
作者:
[Daren Wang;Zifeng Zhao;Yi Yu;R. Willett]
通讯作者:
Daren Wang;Zifeng Zhao;Yi Yu;R. Willett
DOI:
--
发表时间:
2022-05
期刊:
影响因子:
--
作者:
[Carlos Misael Madrid Padilla;Daren Wang;Zifeng Zhao;Yi Yu]
通讯作者:
Carlos Misael Madrid Padilla;Daren Wang;Zifeng Zhao;Yi Yu
DMS-EPSRC: Change Point Detection and Localization in High-Dimensions: Theory and Methods
-
批准号:EP/V013432/1
-
项目类别:Research Grant
-
资助金额:$36.74万
-
财政年份:2021
-
负责人:Yi Yu
-
依托单位:
海外基金