Self-supervised Secondary Landmark Detection via 3D Representation Learning

Self-supervised Secondary Landmark Detection via 3D Representation Learning
复制标题

DOI:
10.1007/s11263-023-01804-y
复制
发表时间:
2023-06-01
影响因子:
19.5
通讯作者:
Hayden, Benjamin Y.
Hayden, Benjamin Y.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bala, Praneet;Zimmermann, Jan;Hayden, Benjamin Y.

文献摘要

被引文献

相似文献

最近的技术发展已经导致在包括人类在内的移动动物的关节和其他地标的计算机化跟踪方面取得了巨大进步。这种追踪有望在生物学和生物医学领域取得重大进展。现代跟踪模型严重依赖于劳动密集型的注释数据集的主要地标的非专家的人。然而,这样的注释方法可能是昂贵的和不切实际的次要地标,也就是说,那些反映动物的细粒度的几何形状,往往是特定于定制的行为任务。由于视觉和几何模糊性,非专家通常没有资格进行二级标志注释,这可能需要解剖学和动物学知识。这些障碍大大阻碍了下游行为研究,因为学习跟踪模型表现出有限的泛化能力。我们假设主要和次要地标之间存在共享表示,因为次要地标的运动范围可以近似地由主要地标的运动范围跨越。我们提出了一种方法来学习这种空间关系的主要和次要的地标在三维空间中,这可以反过来,自我监督的次要地标检测器。这种3D表示学习是通用的,因此可以应用于不同生物体的各种多视图设置,包括猕猴,苍蝇和人类。
Recent technological developments have lead to great advances in the computerized tracking of joints and other landmarks in moving animals, including humans. Such tracking promises important advances in biology and biomedicine. Modern tracking models depend critically on labor-intensive annotated datasets of primary landmarks by non-expert humans. However, such annotation approaches can be costly and impractical for secondary landmarks, that is, ones that reflect fine-grained geometry of animals, and that are often specific to customized behavioral tasks. Due to visual and geometric ambiguity, non-experts are often not qualified for secondary landmark annotation, which can require anatomical and zoological knowledge. These barriers significantly impede downstream behavioral studies because the learned tracking models exhibit limited generalizability. We hypothesize that there exists a shared representation between the primary and secondary landmarks because the range of motion of the secondary landmarks can be approximately spanned by that of the primary landmarks. We present a method to learn this spatial relationship of the primary and secondary landmarks in three dimensional space, which can, in turn, self-supervise the secondary landmark detector. This 3D representation learning is generic, and can therefore be applied to various multiview settings across diverse organisms, including macaques, flies, and humans.