Utilizing Imaged-based Features in Biomedical Literature Classification
Utilizing Imaged-based Features in Biomedical Literature Classification
批准号:
8916181
负责人:
HAGIT SHATKAY
金额:
$28.0万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2017-08-31
关键词:
AddressAreaBioinformaticsBiologicalBiological PhenomenaCategoriesClassificationComparative StudyComputational BiologyComputer-Assisted Image AnalysisComputersCoupledCuesDataData SetDatabasesDelawareDevelopmentDimensionsDiseaseDrug TargetingEnsureEntropyGene ExpressionGene ProteinsGeneric DrugsGoalsGoldGray unit of radiation doseHarvestImageImage AnalysisImprove AccessIndividualInformaticsInformation ResourcesInstitutesInvestigationKnowledgeLeadLiteratureMeasuresMedicalMethodsMiningModelingMusNucleic Acid Regulatory SequencesOrganismOutcomes ResearchPaperPatientsPerformancePharmaceutical PreparationsPhysiciansProcessPropertyProteinsPubMedPublicationsPublishingReaderResearchResearch InfrastructureResearch PersonnelResource InformaticsRetrievalScanningScientistSecureSeriesShapesSourceSpeedSystemTestingTextTextbooksTextureTrainingUniversitiesVotingWorkaccurate diagnosisauthoritybasebioimagingevaluation/testingexperiencefundamental researchimage processingimprovedindexinginsightmouse genomeprotein protein interactionresearch studytext searchingtooltool development
中文摘要
描述(由申请人提供):拟议的研究旨在通过利用出版物中丰富、信息量大的图像数据以及文本,支持和改善对生物医学文献的有效访问。生物医学文献正在以大约
每年出版100万本新书。科学家和医生,作为他们日常工作的一部分,通过无数的出版物寻找相关信息。对于科学数据库管理员(生物管理员,在FlyBase或UniProt等组织中)来说,任务甚至更加艰巨,他们必须识别与数据库区域最相关的文献,在其中找到关于基因、蛋白质、生物体或疾病的高质量证据,并在数据库条目中管理相关文献的发现。值得注意的是,出版物中的许多证据都是数字。因此,图像被科学家和数据库管理员用作相关性的指标。
为了帮助和加快在文献中搜索信息,正在开发自动文本挖掘工具;尽管如此,一些共同的任务和竞争性挑战表明,需要更有效地自动识别生物医学出版物中的相关信息仍然是生物管理和科学发现的瓶颈。虽然生物医学领域内外的图像分析是一个活跃的研究领域,但目前生物医学图像处理的大多数工作都集中在检索和理解图像作为主要形式的数据。同样,生物医学文献检索和挖掘的大多数工作都只关注文本。到目前为止,在出版物中使用图像的工作很少,而图像提供了关于论文中嵌入信息的相关性的重要线索。
我们的建议的假设是,有用的信息可以直接从出版物内的图像,并与基于文本的方法相结合,从而提高识别相关出版物和其中的信息部分。拟议的研究包括广泛的比较研究图像中的高信息量的功能,开发和识别这样的图像功能,开发工具,从图像中提取这些功能和信息,并整合基于图像的信息到文本文章的分类过程,旨在确定出版物的相关性,明确定义的生物医学需求。我们将解决的基本研究任务是:A)识别和比较研究的有用功能的图像表示,专注于其实用程序为特定的生物医学需求; B)分类的生物医学图像和生物医学文件的基础上的图像数据; C)通过整合的文本和图像为基础的分类器的文件分类。为了使研究立足于真正的需求,确保获得大量图像数据,并确保结果的广泛适用性,我们将在三个不同的领域开展工作,我们已经获得了专业知识和数据:(布朗大学Cyrene项目);小鼠基因表达的证据(杰克逊实验室的GXD);蛋白质-蛋白质相互作用的实验证据(特拉华州的蛋白质信息资源)。拟议项目的成功完成将提供综合方法和工具,利用图像和文本的特点,导致更有针对性和更有效的检索和挖掘工具,从而更好地支持数据密集型生物医学发现。
英文摘要
DESCRIPTION (provided by applicant): The proposed research aims to support and improve effective access to the biomedical literature, by utilizing the rich, highly-informative image data within publications, in addition to text. The biomedical literature is expanding at a rate of about
1,000,000 new publications a year. Scientists and physicians, as part of their daily work, go through a myriad of publications searching for relevant information. The task is even more arduous for scientific database curators (bio- curators, in organizations such as FlyBase or UniProt), who have to identify the literature most relevant to the database area, locate within it high-quality evidence concerning genes, proteins, organisms, or disease, and curate the findings within a database entry, with references to the relevant literature. Notably, much of the evidence within publications lies in figures. Accordingly, images are used by scientists and database curators as indicators for relevance.
To assist and expedite the search for information within the literature, automated text-mining tools are being developed; still, several shared tasks and competitive challenges demonstrated that the need for more effective automated identification of relevant information in biomedical publications remains a bottleneck for bio-curation and for scientific discovery. While image analysis within and outside the biomedical domain is an active research area, most current work on biomedical image processing focuses on retrieval and understanding of images as a primary form of data. Likewise, most efforts on biomedical literature retrieval and mining focus on text alone. Little has been done so far to use images within publications, which provide important cues as to the relevance of information embedded in papers.
The hypothesis underlying our proposal is that useful information can be derived directly from images within publications and integrated with text-based methods, leading to improved identification of relevant publications and of informative portions within them. The proposed research comprises extensive comparative study of highly-informative features within images, development and identification of such image-features, development of tools that extract such features and information from images, and integration of image-based information into the textual articles-classification process, aiming to determine the publications' relevance to well-defined biomedical needs. The fundamental research tasks we shall address are: A) Identification and comparative study of useful features for image-representation, focusing on their utility for specific biomedical needs; B) Classification of biomedical images and biomedical documents based on image-data; C) Document classification through integration of text- and image-based classifiers. To ground the research in genuine needs, secure access to much image data, and ensure broad-applicability of the results, we shall work within three diverse areas for which we have secured access to expertise and data: Finding articles about cis-regulatory regions (Cyrene project at Brown University); Evidence for gene expression in the mouse (Jackson Lab's GXD); Experimental evidence for protein-protein interaction (Delaware's Protein Information Resource). The successful completion of the proposed project will provide integrated methods and tools, utilizing both image-based and text-based features, leading to more focused and effective retrieval and mining tools, thus better supporting data-intensive biomedical discovery.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
Corrigendum to "Text as data: Using text-based features for proteins representation and for computational prediction of their characteristics" [Methods 74 (2015) 54-64].
“文本作为数据:使用基于文本的特征进行蛋白质表示及其特征的计算预测”的勘误表 [方法 74 (2015) 54-64]。
DOI:
10.1016/j.ymeth.2016.06.011
发表时间:
2016
期刊:
Methods (San Diego, Calif.)
影响因子:
--
作者:
[Shatkay,Hagit, Brady,Scott, Wong,Andrew]
通讯作者:
Wong,Andrew
DOI:
10.1007/978-3-319-65813-1_20
发表时间:
2017-09-01
期刊:
Experimental IR meets multilinguality, multimodality, and interaction : 8th International Conference of the CLEF Association, CLEF 2017, Dublin, Ireland, September 11-14, 2017, Proceedings. Cross-Language Evaluation Forum. Conference (8...
影响因子:
--
作者:
[Li, Pengyuan, Jiang, Xiangying, Shatkay, Hagit]
通讯作者:
Shatkay, Hagit
Utilizing Imaged-based Features in Biomedical Literature Classification
-
批准号:8892560
-
项目类别:
-
资助金额:$28.0万
-
财政年份:2014
-
负责人:HAGIT SHATKAY
-
依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
-
批准号:2021JJ40433
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:孙磊
-
依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
-
批准号:32001603
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:段真珍
-
依托单位:
AREA国际经济模型的移植.改进和应用
-
批准号:18870435
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1988
-
负责人:史树中
-
依托单位: