课题基金 / 基金详情

CAREER: Optimization Based Methods for Robust Pattern Recognition in Time-Series Data

CAREER: Optimization Based Methods for Robust Pattern Recognition in Time-Series Data
职业:基于优化的时间序列数据中鲁棒模式识别方法
批准号:
1454218
负责人:
Vishal Monga
金额:
$50.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-05-01 至 2020-04-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
这项研究提出了新的数学和算法工具,用于识别和有效地挖掘时间序列数据中的模式,即一系列数据点,通常在连续的时刻以均匀的时间间隔测量。大量的现实世界的数据源,如语音和音频,生物医学信号,医疗保健记录,网络流量和股票市场数据等表现为时间序列,它们的分析是政府和工业界的重要兴趣。这种数据源的爆炸只会因数字革命而加剧,即互联网上大量的音频-视频流,电子数据库中大量按时间顺序排列的医疗记录的存储,以及传感技术进步不断产生的新的时间序列数据。自动化的软件工具,可以找到一个大的时间序列序列的模式,帮助快速和可扩展的检索,并分类大的时间序列集合,因此非常可取。拟议的研究是在开发这样的软件(算法)工具,特别注重鲁棒性和可扩展性。鲁棒性的问题是指可能对人类消费者具有“相同吸引力”的时间序列(例如,相同歌曲/视频的不同版本)可能不一定在数字上相同的事实。因此,需要稳健的技术,能够承受不改变时间序列内容本质的失真。可伸缩性要求模式匹配技术快速且易于实现,以便解决方案可以部署到大型集合中。此外,为了培养下一代电气工程和计算机科学工程师,该项目包括一个强大的教育部分。这个教育组件的核心是一个寓教于乐的游戏,其中一个人类玩家,即具有不同学术准备水平的学生(高中,本科和研究生),在视频盗版挑战中与计算机算法竞争。该游戏旨在使学习过程更具互动性,特别是对于本科生。在新兴应用中挖掘时间序列数据的一个严重的实际挑战是承受扭曲的能力-也就是说,“相同的底层”时间序列的实例经常在噪声、幅度和/或时间缩放以及其他杂项操作下被扭曲。许多现有的时间序列比较技术不能实现失真鲁棒性,而那些能够实现的技术往往需要大量的计算成本。此外,现有的算法技术只能在一个直观的,通常是启发式的水平上控制时间序列特征的关键属性,如鲁棒性和唯一性。建议的研究主张明智地选择时间序列极值,旨在打破时间序列特征提取和比较中的计算效率与对失真的鲁棒性之间的经典权衡。与现有的方法,采用预处理的时间序列滤波器的“灵感”来自直觉,明确的优化thefilter提出的意义上的成本函数,捕捉关键特征属性,如鲁棒性和唯一性提取的极值。最佳极值提取将在两种不同的设置中进行研究:a.)确定性框架,其中在优化中使用示例训练时间序列,以及B.)使用时间序列随机模型的统计框架。也出现了各种相关的子问题,即:a.)与图像处理和视觉中的边缘检测问题的联系,B.)时间序列极值的编码和比较,以及c.)扩展到在时间序列的非线性操作下寻找鲁棒极值。研究计划是将算法工具的开发与两个现实世界的应用并列:1。多媒体指纹,和2.)生物医学时间序列分析此外,还将开发基于这些应用的软件工具,即寓教于乐的游戏,这将在加强PI的研究和课堂教学中发挥至关重要的作用。研究成果的传播将通过主要期刊和会议上的文章以及在线MATLAB软件工具箱进行。
英文摘要
This research proposes new mathematical and algorithmic tools for identifying patterns in and effectively mining time-series data, i.e. a sequence of data points, measured typically at successive time instantsspaced at uniform time intervals. A large variety of real-world data sources such as speech and audio,biomedical signals, health care records, network-traffic and stock market data etc. manifest as time-seriesand their analysis is of significant interest to both government and industry. The explosion of such data sources has only been exacerbated by the digital revolution, viz. the generous amount of audio-videostreams on the internet, the storage of large amounts of chronological health care records in electronicdatabases and the continuous generation of new time-series data from advances in sensing. Automated softwaretools that can find patterns in a large time-series sequence, help in fast and scalable retrieval, andcategorize large time-series collections are hence highly desirable. The proposed research is in developing such software (algorithmic)tools with a particular focus on robustness and scalability. The problem of robustness refers to the fact that time-series that may have the "same appeal" to a human consumer, e.g. different versions of the same song/video, may not necessarily be digitally identical. Hence, robust techniques are needed that can withstand distortions which do not change the essence of the time-series content. Scalability requires that the pattern-matching techniques be fast and easy to implement, so thatthe solutions can be deployed to mine large collections. Further, to prepare the next generation ofengineers in electrical engineering and computer science, the project includes a strong educationalcomponent. At the heart of this educational component is an edutainment game where a human player, i.e.students with varying levels of academic preparation (high-school, undergraduate and graduate), compete against a computer algorithm in a video piracy challenge. The game is aimed at making the learning process more interactive, particularly for undergraduate students.A serious practical challenge in mining time-series data for emerging applications is the ability towithstand distortions - that is often instances of the "same underlying" time series are observedunder noise, amplitude and/or time scaling and other miscellaneous operations. Many existingtechniques for time-series comparisons do not enable distortion robustness and the ones that do,often come at a substantial computational cost. Further, existing algorithmic techniques enablecontrol of key properties of time-series features such as robustness and uniqueness only at anintuitive, often heuristic level. The proposed research advocates judicious selection of time-seriesextrema and aims to break the classical trade-off between computational efficiency in time-seriesfeature extraction and comparison vs. enabling robustness to distortions. Unlike existing methods,which employ pre-processing time-series filters "inspired" from intuition, explicit optimization of thefilter is proposed in the sense of cost functions that capture key feature attributes such as robustnessand uniqueness of the extracted extrema. Optimal extrema extraction will be investigated in twodifferent setups: a.) a deterministic framework where example training time-series are used in theoptimization, and b.) a statistical framework where stochastic models on time-series are used. Avariety of related sub-problems also emerge, namely: a.) connections to edge detection problemsin image processing and vision, b.) encoding and comparisons of time-series extrema, and c.)extensions to finding robust extrema under non-linear operations on the time-series. The researchplan is to juxtapose the development of the algorithmic tools with two real-world applications: 1.)multimedia fingerprinting, and 2.) bio-medical time series analysis. Additionally, software toolsnamely edutainment games will be developed based on these applications which will play a crucialrole in enhancing the PI's research and classroom teaching. Dissemination of research results will bedone via articles in leading Journals and conferences, and via online MATLAB software toolboxes.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
供应链管理中的稳健型(Robust)策略分析和稳健型优化(Robust Optimization )方法研究
  • 批准号:
    70601028
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    7.0万元
  • 批准年份:
    2006
  • 负责人:
    王明征
  • 依托单位: