Unsupervised Annotation of Complex 3D BioMedical Data.
Unsupervised Annotation of Complex 3D BioMedical Data.
批准号:
2882348
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --
中文摘要
我们的目标是为复杂的三维生物医学数据的无监督注释开发新的方法。这将允许有效地标记细胞数据的3D模型,供人类生物医学研究人员或未来的机器学习模型使用。深度学习方法越来越多地被用于分析复杂数据,然而这些方法需要大量标记的训练数据才能有效。这种训练数据通常很难获得,而且价格昂贵。最近大量3D数据的可用性给其注释带来了挑战,因为这比2D同等的任务要困难得多。这给研究带来了障碍。我们提出的解决方案是使用无监督注释。虽然从头开始标记数据非常昂贵,需要专家团队投入大量时间,但组合和操纵通过无监督方法获得的标签提供了一种更现实的方法。通过这种方法,我们可以减少为进一步研究生成标记训练数据所需的时间和金钱投入,并为未来的研究提供重要的实用价值。此外,近年来开发的虚拟现实(VR)工具提供了3D细胞模型的沉浸式视图,使医生和研究人员能够直观,详细地查看他们希望分析的细胞。特别感兴趣的是MiCellAnnGELo细胞绘画工具,它提供3D生物医学数据的手动注释,同时使用4D时间序列数据。VR的沉浸性非常适合这项任务,因为它可以让专家对细胞进行近距离的互动观察,以简化注释过程。MiCellAnnGELo利用Unity游戏引擎,这是一个专为3D环境导航设计的平台,允许最终用户轻松操纵细胞模型。这为生物医学数据的注释提供了一个易于使用的平台,可以减轻3D注释的禁止性。我们将开发新的功能来促进细胞成分的动态标记。由于完全自动化注释过程可能会影响准确性,我们的目标是将建议合并到工具中,以识别模型中可能的细胞结构和组件。通过这样做,我们可以进一步简化注释过程,增加这些可视化工具的实用性,使研究人员能够更好地分析细胞内的复杂结构,同时提供视觉辅助,以促进更好地理解细胞数据。我们将利用标准的框架和模型来实现我们的目标。对于神经网络的开发,我们打算使用PyTorch框架,因为它易于使用和灵活,这将使我们能够快速构建原型并迭代我们的研究。如果PyTorch出现问题,TensorFlow和Keras框架可以作为替代方案。对于我们的模型,我们计划基于图卷积网络(Graph convolutional Networks)来处理表面数据,基于3D卷积神经网络(3D convolutional Neural Networks)来处理体积数据。基于聚类的生成对抗网络将用于模拟额外的训练数据。目前,聚类主要应用于2D数据集,其中预期聚类的数量已经已知。因此,我们项目的一个重要组成部分是开发组合网络的策略,这些网络要么基于不同数量的集群,要么使用不依赖于固定数量的集群方法。一个重要的贡献将是在不同分辨率下使用子采样实现数据聚类,这可以保留表面网格的拓扑结构,以及体数据的超像素方法。我们必须确保我们的模型可以在不同分辨率的数据上进行训练和注释,而不限于单一大小的输入。通过对这些模型的适应和结合,为我们的研究奠定了坚实的基础,为我们自己的3d数据的分割和标注奠定了基础
英文摘要
We aim to develop new methods for the unsupervised annotation of complex 3D biomedical data. This will allow for the efficient labelling of 3D models of cellular data for use by either human biomedical researchers or future machine learning models. Deep learning methods are increasingly being used to analyse complex data however these methods require large quantities of labelled training data to be effective. Such training data is often difficult and expensive to acquire. The recent availability of large volumes of 3D data poses challenges in its annotation as this is a significantly more difficult task than the 2D equivalent. This presents a barrier to research. The solution we propose is using unsupervised annotation. While labelling data from scratch is prohibitively expensive, requiring a significant time investment from a team of experts, combining and manipulating labels obtained via unsupervised methods provides a more realistic approach. Through this method we can reduce the time and monetary investment required to produce labelled training data for further research and provide significant utility for future studies.Furthermore, Virtual Reality (VR) tools have been developed in recent years which offer an immersive view of 3D cellular models allowing doctors and researchers intuitive, detailed views of cells they wish to analyse. Of particular interest is the is the MiCellAnnGELo cell painting tool which offers manual annotation of 3D biomedical data while working with 4D timeseries data. The immersive nature of VR is well suited to this task as it can allow experts a close-up interactive view of cells to streamline the annotation process. MiCellAnnGELo makes use of the Unity gaming engine, a platform designed for the navigation of 3D environments, to allow for cell models to be easily manipulated by the end user. This provides an easy-to-use platform for the annotation of biomedical data that can alleviate much of the prohibitive nature of 3D annotation.We will develop new functionality to facilitate the dynamic labelling of cell components. As fully automating the annotation process may jeopardise accuracy, we aim to incorporate suggestions into the tool to identify likely cell structures and components within a model. In doing so we can streamline the annotation process further and increase the utility of these visualisation tools, allowing researchers to better analyse complex structures within cells while providing a visual aid to facilitate a better understanding of the cellular data.We will make use of standard frameworks and models to achieve our goals. For the development of our neural networks, we intend to use the PyTorch framework due to its ease of use and flexibility which will allow us to quickly prototype and iterate on our research. Should problems arise with PyTorch, the TensorFlow and Keras frameworks are available as alternatives. For our models, we plan to base our work upon Graph convolutional Networks for surface data and 3D Convolutional Neural Networks for volume data. Cluster based Generative Adversarial Networks will be used to simulate additional training data. Currently clustering has been primarily applied to 2D datasets where the number of expected clusters is already known. An important component of our project is therefor to develop strategies for combining networks which are either based on different numbers of clusters or using clustering methods that do not depend on a fixed number. An important contribution will then be to implement the clustering of data at different resolutions using subsampling that can preserve the topology of surface meshes, and superpixel approaches for volume data. We must ensure that our model can train on and annotate data of different resolutions and is not limited to a single size of input. By adapting and combining these models we achieve a strong foundation for our research and obtain a basis for the segmentation and annotation of our own 3Ddata
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金