Automating the Integration of EPA Databases
Automating the Integration of EPA Databases
批准号:
0306899
负责人:
Eduard Hovy
金额:
$90.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-08-15 至 2007-07-31
中文摘要
由于政府必须管理的地理范围广泛和任务复杂,其数据以许多不同的方式分开,并由不同的机构在不同的时间收集。由此产生的海量数据异构性意味着人们无法有效地定位、共享或比较源中的数据,更不用说实现计算数据互操作性了。到目前为止,包装数据集合的所有方法,甚至是跨可比数据集创建映射的方法,都需要手动操作。尽管有一些有希望的工作,但这种映射的自动化创建仍处于初级阶段,因为等价性和差异性在所有级别上都很明显,从单个数据值到元数据,再到围绕整个数据收集的解释性文本。需要更通用的方法来有效地解决这个问题。将数据映射问题视为机器翻译(MT)跨语言映射问题的变体,该项目将使用自1990年以来在机器翻译社区开发的新统计算法,以发现各级可比数据集之间的对应关系。在机器翻译中,这些技术可以跨语言对齐单词和单词序列。这项研究将调整和扩展技术,不仅考虑数据值(单词的类比),而且考虑数据格式/正字法、元数据信息和相关的文本信息(元数据描述、脚注等)。在对齐过程中,并在三个级别进行对齐学习:单个数据单元格级、单元格集(列)级和多列级。在MT中以前没有尝试过多水平对齐。这些强大的学习技术从未被应用于元数据模式集成和/或数据库对齐或包装。如果这些自动学习的映射是有效的,那么数据库包装所需的手工工作量应该会显著减少。将使用两组域数据。空气质量数据将由萨克拉门托加州空气资源委员会的环保局工作人员提供,他们定期将加州约35个地区空气质量管理区的数据整合到一个加州范围的单一数据库中,并将其传递给北卡罗来纳州的联邦环保局。火灾排放数据将由不同的环保局办公室、美国农业部/林业局和内政部提供。
英文摘要
Due to the wide range of geographic scales and complex tasks the Government must administer, its data is split in many different ways and is collected at different times by different agencies. The resulting massive data heterogeneity means one cannot effectively locate, share, or compare data across sources, let alone achieve computational data interoperability. To date, all approaches to wrap data collections, or even to create mappings across comparable datasets, require manual effort. Despite some promising work, the automated creation of such mappings is still in its infancy, since equivalences and differences manifest themselves at all levels, from individual data values through metadata to the explanatory text surrounding the data collection as a whole. More general methods are required to effectively address this problem. Viewing the data mapping problem as a variant of the cross-language mapping problem of Machine Translation (MT), this project will employ the new statistical algorithms developed since 1990 in the MT community to discover correspondences across comparable datasets at all levels. In MT, the techniques align words and word sequences across languages. This research will adapt and extend the techniques to consider not only data values (the analogue of words) but also data format/orthography, metadata information, and associated textual information (metadata descriptions, footnotes, etc.) in the alignment process, and to perform alignment learning at three levels: individual data cell level, set of cells (column) level, and multi-column level. Multi-level alignment has not been attempted in MT before. These powerful learning techniques have never been applied to metadata schema integration and/or database alignment or wrapping. If these automatically learned mappings are effective, the amount of manual labor required in database wrapping should be significantly reduced. Two sets of domain data will be used. Air quality data will be provided by EPA staff at the California Air Resources Board in Sacramento, who periodically integrate data from some 35 regional Air Quality Management Districts throughout California into a single California-wide database, and pass this along to the Federal EPA in North Carolina. Fire emissions data will be provided by a different set of EPA offices, the USDA/Forest Service, and the Department of Interior.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: A Method to Retrieve Non-Textual Data from Widespread Repositories
-
批准号:1450545
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2014
-
负责人:Eduard Hovy
-
依托单位:
III: EAGER: Automatically Building Test Collections Using Implicit Relevance Signals from the Web
-
批准号:1304939
-
项目类别:Standard Grant
-
资助金额:$10.32万
-
财政年份:2012
-
负责人:Eduard Hovy
-
依托单位:
EAGER: Constructing, Indexing, and Searching Super-Enriched Document Representations in the Cloud
-
批准号:1265301
-
项目类别:Standard Grant
-
资助金额:$23.61万
-
财政年份:2012
-
负责人:Eduard Hovy
-
依托单位:
III: EAGER: Automatically Building Test Collections Using Implicit Relevance Signals from the Web
-
批准号:1147810
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2011
-
负责人:Eduard Hovy
-
依托单位:
EAGER: Constructing, Indexing, and Searching Super-Enriched Document Representations in the Cloud
-
批准号:1143703
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2011
-
负责人:Eduard Hovy
-
依托单位:
Collaborative Research III-COR: From a Pile of Documents to a Collection of Information: A Framework for Multi-Dimensional Text Analysis
-
批准号:0705091
-
项目类别:Standard Grant
-
资助金额:$32.0万
-
财政年份:2007
-
负责人:Eduard Hovy
-
依托单位:
Collaborative Research: Language Processing Technology for Electronic Rulemaking
-
批准号:0429360
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2004
-
负责人:Eduard Hovy
-
依托单位:
SGER COLLABORATIVE: A Testbed for eRulemaking Data
-
批准号:0328175
-
项目类别:Standard Grant
-
资助金额:$2.5万
-
财政年份:2003
-
负责人:Eduard Hovy
-
依托单位:
Collaborative Research:Interlingual Annotation of Multilingual Text Corporation
-
批准号:0325021
-
项目类别:Standard Grant
-
资助金额:$16.88万
-
财政年份:2003
-
负责人:Eduard Hovy
-
依托单位:
ITR: Information Discovery in Digital Government: Self-extending Topic Maps and Ontologies (GrowOnto)
-
批准号:0205111
-
项目类别:Continuing Grant
-
资助金额:$100.0万
-
财政年份:2002
-
负责人:Eduard Hovy
-
依托单位:
Digital Government: dg.o Workshop and Publicity
-
批准号:0089522
-
项目类别:Continuing Grant
-
资助金额:$74.92万
-
财政年份:2000
-
负责人:Eduard Hovy
-
依托单位:
Workshop: Support for Workshop on Multilingual Information Management
-
批准号:9807199
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:1998
-
负责人:Eduard Hovy
-
依托单位:
Workshop: MT Summit Conference Support
-
批准号:9725058
-
项目类别:Standard Grant
-
资助金额:$0.5万
-
财政年份:1997
-
负责人:Eduard Hovy
-
依托单位:
International Language Generation Workshop, June 1994, Kennebunkport, ME
-
批准号:9321870
-
项目类别:Standard Grant
-
资助金额:$1.18万
-
财政年份:1994
-
负责人:Eduard Hovy
-
依托单位:
海外基金