课题基金 / 基金详情

III: Small: Linking and Resolving Entities in Big Data

III: Small: Linking and Resolving Entities in Big Data
III:小:大数据中实体的链接和解析
批准号:
1527536
负责人:
Sharad Mehrotra
金额:
$50.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2020-08-31

项目摘要

项目成果

Sharad Mehrotra的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This project will explore the challenge of cleaning data in the context of analysis pipelines over big data. Data cleaning has traditionally been designed to improve data quality in ETL systems where enterprise data is collected, prepared, staged, transformed, and loaded into a data warehouse to support offline data analysis. In the era of big data, such back-end processes are quickly giving way to interactive exploratory data analysis where analysts immerse themselves in data (possibly collected from heterogeneous data sources) in order to drive online (near-) real-time decision making. Existing systems do not scale to the volume, velocity, or the variability of the dynamically generated data (e.g., social media streams) and the offline architecture is unsuited for the online real-time nature of analysis. The market is abuzz with innovations in data transformation technologies, e.g., TriFacta allows analysts to visually manipulate data to generate complex analytical transformations and Data Tamer is exploring scalable data curation from diverse sources. Data quality (and hence data cleaning technologies) remain at the core of big-data analytics. Many popular media (as well as academic) articles have highlighted challenges such as entity linking and resolution as among the most important and immediate roadblocks for big data analytics. The key insight on which this project is based is that data cleaning to support analytics over big data is not simply a matter of scaling up known approaches to larger data sets by exploiting more hardware. While scale up is important, big data analytics in streaming, real-time, and interactive settings requires a paradigm shift in how data cleaning is performed. This project will significantly impact and change the modern practices of data cleaning and the way cleaning is integrated in the Big Data analysis pipeline and will explore broader impact through: (a) technology transfer opportunities with a relevant industrial partner whose existing products could benefit from the proposed research; and (b) open source effort in the context of the ongoing social media analytics system (SoDAS), currently under development, in which the proposed research algorithms will be integrated.This research will explore two new innovations that will help advance data cleaning to enable Big Data analysis. The first innovation explores a progressive approach to entity resolution to support progressive analysis. The research will explore an approach where progressiveness is pervasive spanning all the phases of the cleaning process especially in scenarios when cleaning is based on complex logic possibly requiring dynamic acquisition of additional contextual information. The second innovation is the analysis-aware data cleaning that is developed for structured queries (e.g., Hive and SQL) for both one-time and continuous query scenarios that are issued on top of static and streaming data. The project will address these methodologies at the higher conceptual level as well as implement them on modern highly-parallel computing platforms and frameworks that run on a cluster of machines. The project will exploit two concrete contexts to guide the research exploration: (a) supporting analytical queries over structured web data sources such as fusion tables; and (b) online analysis of social media data. These application contexts will serve as vehicles for testing and demonstrating the research. The planned research, system development, and educational activities (e.g., curriculum changes to incorporate projects related to big data and data quality in the CS curriculum at UCI) will significantly enhance the educational experience of students, preparing them for a brighter future in the today?s knowledge driven society. More information about the project can be found at http://sherlock.ics.uci.edu.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Travel: Request for Student Travel Support for the 48th International Conference on Very Large Databases 2022
  • 批准号:
    2230342
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.5万
  • 财政年份:
    2022
  • 负责人:
    Sharad Mehrotra
  • 依托单位:
RAPID: An Organizational Scale Approach to Privacy-Enabled Contact Tracing in COVID-19
  • 批准号:
    2032525
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2020
  • 负责人:
    Sharad Mehrotra
  • 依托单位:
III: Small: EnrichDB - Supporting Enrichment in Database Systems
  • 批准号:
    2008993
  • 项目类别:
    Standard Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2020
  • 负责人:
    Sharad Mehrotra
  • 依托单位:
Student Support for the 46th International Conference on Very Large Databases (VLDB 2020)
  • 批准号:
    2025108
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.5万
  • 财政年份:
    2020
  • 负责人:
    Sharad Mehrotra
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: