Enriching, repairing and merging taxonomies by inducing qualitative spatial representations from the web
Enriching, repairing and merging taxonomies by inducing qualitative spatial representations from the web
批准号:
EP/K021788/1
负责人:
Steven Schockaert
金额:
$12.61万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2013
资助国家:
英国
项目状态:
已结题
起止时间:
2013 至 --
中文摘要
分类法编码来自给定域的不同术语或概念如何相互关联。它们被用来标准化词汇表(例如,生物学家使用分类法将物种组织成更广泛的类别,如科和目),并对内容进行分类,以便更容易地进行搜索(例如,图书馆员将分类法中的类别分配给书籍)。虽然分类传统上是仔细和耗时的人工过程的结果,但万维网最近的发展导致了更非正式性质的分类的激增。例如,亚马逊等在线零售商使用特别分类来组织产品,这种分类反映了客户如何使用他们的网站,而不是对潜在产品类别的语义做出任何承诺。同样,Foursquare等应用程序允许用户对地点类型进行分类。虽然这些非正式分类对组织在线内容(例如Amazon上的产品或Foursquare上的场所)很有用,但它们的质量往往很差,难以在不同的应用程序中重复使用。此外,与传统分类法一样,它们关注的是一组非常有限的语义关系;通常只考虑“是的子类别”这一关系。相反,在实践中,两个类别之间的语义关系可能不那么明确,尤其是因为存在边界情况(例如,提供食物的酒吧是否应该被归类为餐厅?)。尽管如此,如果可以使用自动化方法改进分类法,那么分类法的广泛使用可能会引起极大的兴趣。这个项目的目标是研究如何通过统计分析网络上可用的元数据,特别是来自所谓的Web 2.0网站(如Flickr)的元数据,研究如何实现这种改进。在Flickr等所谓的Web 2.0网站上,用户使用称为标签的简短文本注释来描述照片。提出的方法建立在通过统计分析此类元数据来发现类别之间的语义关系的想法之上。一方面,这些关系将编码关于典型性和相似性的信息。要了解这种关系为什么有用,请考虑一个应用程序,该应用程序允许用户搜索加的夫的餐馆。搜索引擎可以通过考虑到市中心的距离和平均评级(如果可用)等特征来对类型为“餐馆”的场所进行排名。然而,作为另一个标准,人们也会希望在早餐场所、咖啡馆或酒吧等场所之前看到“正常”的餐馆,这些场所可能被广泛地认为是餐馆,但不是用户在询问餐馆时通常会感兴趣的。类似地,当用户的查询询问“加的夫的四川餐馆”,并且不知道这样的餐馆时,可以显示最相似类别的实例(例如,广东餐馆)。另一方面,发现的关系还将编码信息,这些信息可以帮助我们查明现有分类法中可能的错误,并帮助我们合并不同的分类法,以获得给定域的单一连贯视图。特别是,这些关系将使我们能够检测现有分类中的不规则性。例如,假设相似的类别通常具有相似的属性,并且知道广东和四川餐馆非常相似,那么广东餐馆和四川餐馆都是中餐馆的子类别的分类将被认为比它们具有不同超范畴的分类更规则。我们的方法的独特之处在于,它以数据驱动的方法来丰富分类,以语义关系进行常识推理,以及建议的修复和合并现有分类的方法。在应用方面,该项目的结果将成为迈向更智能搜索引擎的重要垫脚石。
英文摘要
Taxonomies encode how different terms or concepts from a given domain are related to each other. They are used to standardise vocabularies (e.g. biologists use taxonomies to organise species into broader categories such as family and order), and to categorise content such that it can be more easily searched (e.g. librarians assigning categories from a taxonomy to books). While taxonomies are traditionally the result of a careful and time-consuming manual process, recent developments in the world wide web have led to a proliferation of taxonomies of a more informal nature. Online retailers such as Amazon, for instance, organise their products using an ad hoc taxonomy, which reflects how customers use their website, rather than any commitment on the semantics of the underlying product categories. Similarly, applications such as Foursquare allow users to contribute to a taxonomy of place types.While these informal taxonomies are useful to organise online content (e.g. products on Amazon, or venues on Foursquare), they are often of poor quality, and difficult to reuse among different applications. Moreover, like traditional taxonomies, they focus on a very limited set of semantic relations; usually only the relation "is a sub-category of" is considered. In contrast, in practice the semantic relationship between two categories may not be so clear-cut, among others because of the existence of borderline cases (e.g. should a pub which serves food be categorised as a restaurant?). Nonetheless, the widespread availability of taxonomies is of potentially great interest, provided that they can be improved using automated methods. The goal of this project is to study how such an improvement can be realised, by statistically analysing meta-data that is available on the web, and in particular from so-called Web 2.0 websites such as Flickr, where users describe photos using short textual annotations called tags.The proposed approach is built on the idea of discovering semantic relationships between categories by statistically analysing such meta-data. On the one hand, these relations will encode information about typicality and similarity. To see why such relations are useful, consider an application which allows a user to search for restaurants in Cardiff. The search engine may rank venues of type "restaurant" by taking into account features such as distance to the city centre and average ratings (if available). However, as another criterion, one would also want to see "normal" restaurants before venues such as breakfast places, coffee houses, or pubs, which may be considered as restaurants, broadly speaking, but are not what users would typically be interested in when querying about restaurants. Similarly, when the user's query asks about "Sichuan restaurants in Cardiff", and no such restaurants are known, instances of the most similar categories may be shown instead (e.g. Cantonese restaurants). On the other hand, the relations that are discovered will also encode information that can help us to pinpoint likely errors in existing taxonomies and that can help us to merge different taxonomies to get a single coherent view of a given domain. In particular, these relations will allow us to detect irregularities in existing taxonomies. For example, given the assumption that similar categories usually have similar properties, and the knowledge that Cantonese and Sichuan restaurants are very similar, a taxonomy in which Cantonese and Sichuan restaurants are both sub-categories of Chinese restaurants will be considered more regular than a taxonomy in which they have different super-categories.Our approach is unique in its data-driven approach to enrich taxonomies with semantic relations for common-sense reasoning, as well as in the proposed methods for repairing and merging existing taxonomies. Regarding applications, the results of this project will form a crucial stepping-stone towards more intelligent search engines.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1016/j.artint.2015.07.002
发表时间:
2015-11
期刊:
Artif. Intell.
影响因子:
--
作者:
[J. Derrac;Steven Schockaert]
通讯作者:
J. Derrac;Steven Schockaert
Realizing RCC8 networks using convex regions
使用凸区域实现 RCC8 网络
DOI:
10.48550/arxiv.1410.2442
发表时间:
2014
期刊:
影响因子:
--
作者:
[Schockaert S]
通讯作者:
Schockaert S
DOI:
--
发表时间:
2015
期刊:
影响因子:
--
作者:
[Steven Schockaert]
通讯作者:
Steven Schockaert
DOI:
10.3233/978-1-61499-419-0-243
发表时间:
2014-08
期刊:
影响因子:
--
作者:
[J. Derrac;Steven Schockaert]
通讯作者:
J. Derrac;Steven Schockaert
Reasoning about Structured Story Representations
-
批准号:EP/W003309/1
-
项目类别:Fellowship
-
资助金额:$163.22万
-
财政年份:2022
-
负责人:Steven Schockaert
-
依托单位:
Encyclopedic Lexical Representations for Natural Language Processing
-
批准号:EP/V025961/1
-
项目类别:Research Grant
-
资助金额:$76.1万
-
财政年份:2021
-
负责人:Steven Schockaert
-
依托单位:
海外基金