课题基金 / 基金详情

Cleaning and Analysis of Large Uncertain and Inconsistent Data Sources

Cleaning and Analysis of Large Uncertain and Inconsistent Data Sources
大量不确定且不一致的数据源的清理和分析
批准号:
RGPIN-2014-06143
负责人:
Ilyas, Ihab
金额:
$5.54万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2015
资助国家:
加拿大
项目状态:
已结题
起止时间:
2015-01-01 至 2016-12-31

项目摘要

项目成果

Ilyas, Ihab的其他基金

相似基金

相关文献

中文摘要
翻译
现代应用程序(如对象跟踪、传感器网络、健康记录管理和Web数据集成)生成的数据涉及不确定性和各种异常,如缺失值和重复。虽然许多研究工作一直专注于数据清理和处理不一致的数据库,但在真实的环境中,针对围绕从大型脏数据集中增加和提取价值的各种技术和实践挑战,采用了非常有限的研究。列举一些这些实际和技术挑战:(1)数据的保护和敏感性,其中数据保管人和监护人阻止自动修复算法改变底层数据;(2)完整性约束的异质性,这使得解决单一类型错误的建议技术在实践中不适用或无效;(3)缺乏地面实况来验证修复策略;最低限度的修复等质量指标在实践中没有显示出很好的效果;(4)缺乏数据质量方面的互动工具,使用户和专家能够对数据中有问题的部分进行推理,并解释这些错误背后的原因。 在本提案中,我们专注于在大规模不一致和脏数据库上实现数据质量分析和检索。该提案追求一系列研究方向,包括(1)非破坏性数据清洗,表示和查询可能的数据修复,而不改变底层数据;(2)整体数据清洗,解决多个异构完整性约束的违反;(3)高保真数据修复,其更多地依赖于可信的数据源和专家,而较少地依赖于启发式质量度量,例如最小限度的维修;以及(4)实用仪表板中的描述性和规定性数据质量分析,其不仅描述数据中的错误,还推荐防止未来错误的方法。 所提出的技术将在我们以前开发的系统原型中实现和测试:UClean,一个概率和质量感知数据库引擎原型,基于开源数据库管理系统;和NADEEF,一个开源的可扩展数据清洗系统。我们的目标是建立一个通用的框架,封装高效的查询处理算法,使用户能够有效地查询,分析和探索大量的不一致和不确定的数据。 开发的算法和仪表板将使研究社区和行业能够对可用数据集的质量进行推理,并就如何清理或提高目标应用程序或用例的数据质量提供指导。
英文摘要
Data generated by modern applications such as object tracking, sensor networks, health record management, and Web data integration involves uncertainty, and various anomalies such as missing values and duplication. While many research efforts have been focusing on data cleaning and dealing with inconsistent databases, very limited research has been adopted in real settings for various technical and practical challenges around increasing and extracting value from large dirty data sets. To list a few of these practical and technical challenges: (1) the protection and sensitivity of data, where data custodians and guardians prevent automatic repairing algorithms from changing the underlying data; (2) the heterogeneity of integrity constraints, which makes proposed techniques that tackle a single type of error inapplicable or ineffective in practice; (3) the lack of ground truth to validate repairing strategies; quality metrics such as minimal repairs have not been showing great results in practice; and (4) the lack of interactive tools for data quality that allow users and experts to reason about the problematic parts of the data and to explain the reasons behind these errors. In this proposal, we focus on enabling data quality analytics and retrieval on large-scale inconsistent and dirty databases. The proposal pursues a set of research directions including (1) non-destructive data cleaning that represents and queries possible data repairs without changing the underlying data; (2) holistic data cleaning, which addresses the violations of multiple heterogeneous integrity constraints; (3) high-fidelity data repairing, which depends more on trusted data sources and experts, and depends less on heuristic quality metrics, such as minimal repairs; and (4) descriptive and prescriptive data quality analytics in practical dashboards that go beyond describing errors in the data to recommending ways to prevent future errors. The proposed techniques will be implemented and tested in our previously developed system prototypes: UClean, a probabilistic and quality-aware database engine prototype, based on an open-source Database Management System; and NADEEF, an open source extensible data cleaning system. The goal is to build a generic framework that encapsulates efficient query processing algorithms to allow users to effectively query, analyze and explore large volumes of inconsistent and uncertain data. The developed algorithms and dashboard will enable both the research community and industry to reason about the quality of available data sets, and to provide guidance on how to clean or enhance the quality of this data with respect to target applications or use cases.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Scalable Cleaning, Integration and Analysis of Structured and Semi-Structured Inconsistent Data
  • 批准号:
    RGPIN-2019-04068
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2022
  • 负责人:
    Ilyas, Ihab
  • 依托单位:
Scalable Cleaning, Integration and Analysis of Structured and Semi-Structured Inconsistent Data
  • 批准号:
    RGPIN-2019-04068
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2021
  • 负责人:
    Ilyas, Ihab
  • 依托单位:
NSERC/Thomson Reuters Industrial Research Chair in Data Cleaning
  • 批准号:
    534011-2017
  • 项目类别:
    Industrial Research Chairs
  • 资助金额:
    $14.57万
  • 财政年份:
    2021
  • 负责人:
    Ilyas, Ihab
  • 依托单位:
End-to-end Extraction and Curation of Large RDF Repositories
  • 批准号:
    543961-2019
  • 项目类别:
    Collaborative Research and Development Grants
  • 资助金额:
    $11.82万
  • 财政年份:
    2020
  • 负责人:
    Ilyas, Ihab
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Intelligent Patent Analysis for Optimized Technology Stack Selection:Blockchain BusinessRegistry Case Demonstration
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    USHARANI HAREESH GOVINDARA JAN
  • 依托单位:
基于Meta-analysis的新疆棉花灌水增产模型研究
  • 批准号:
    41601604
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    22.0万元
  • 批准年份:
    2016
  • 负责人:
    赵爱琴
  • 依托单位:
大规模微阵列数据组的meta-analysis方法研究
  • 批准号:
    31100958
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2011
  • 负责人:
    赵洪雅
  • 依托单位: