Fuzzy Orders-of-Magnitude-Based Link Analysis for Qualitative Alias Detection

Fuzzy Orders-of-Magnitude-Based Link Analysis for Qualitative Alias Detection
复制标题

DOI:
10.1109/tkde.2010.255
复制
发表时间:
2012-04
影响因子:
8.9
通讯作者:
Qiang Shen;Tossapon Boongoen
Qiang Shen;Tossapon Boongoen
中科院分区:
计算机科学2区
文献类型:
--
作者:
Qiang Shen;Tossapon Boongoen

文献摘要

被引文献

相似文献

别名检测一直是多个领域应用(尤其是情报数据分析)广泛研究的重要课题。许多初步方法依赖于基于文本的措施,这对于恐怖分子姓名、出生日期和地址的虚假描述是无效的。该障碍可以通过在感兴趣的对象之间的关系中呈现的链接信息来克服。几种基于数字链接的相似性技术已被证明对于识别互联网和出版领域中的相似对象是有效的。然而,由于特殊情况的测量值过高,这些方法通常会生成不准确的相似性描述。然而,它们要么计算效率低下,要么对于基于单一属性的模型的别名检测无效。本文提出了一种新颖的基于数量级的相似性度量,它集成了多个链接属性来改进估计过程并导出语义丰富的相似性描述。该方法基于数​​量级推理,与模糊集理论相结合,提供描述符的定量语义及其明确的数学操作。通过这种解释性的形式主义,分析师可以验证生成的结果并部分解决误报问题。利用这种文字计算功能,它还可以在决策小组内进行连贯的解释和沟通。其性能是通过与恐怖主义相关的数据集进行评估的,并进一步概括了出版物和电子邮件数据收集。
Alias detection has been the significant subject being extensively studied for several domain applications, especially intelligence data analysis. Many preliminary methods rely on text-based measures, which are ineffective with false descriptions of terrorists' name, date-of-birth, and address. This barrier may be overcome through link information presented in relationships among objects of interests. Several numerical link-based similarity techniques have proven effective for identifying similar objects in the Internet and publication domains. However, as a result of exceptional cases with unduly high measure, these methods usually generate inaccurate similarity descriptions. Yet, they are either computationally inefficient or ineffective for alias detection with a single-property based model. This paper presents a novel orders-of-magnitude based similarity measure that integrates multiple link properties to refine the estimation process and derive semantic-rich similarity descriptions. The approach is based on order-of-magnitude reasoning with which the theory of fuzzy set is blended to provide quantitative semantics of descriptors and their unambiguous mathematical manipulation. With such explanatory formalism, analysts can validate the generated results and partly resolve the problem of false positives. It also allows coherent interpretation and communication within a decision-making group, using this computing-with-word capability. Its performance is evaluated over a terrorism-related data set, with further generalization over publication and email data collections.