MACE2K - Molecular And Clinical Extraction: A Natural Language Processing Tool for Personalized Medicine
MACE2K - Molecular And Clinical Extraction: A Natural Language Processing Tool for Personalized Medicine
批准号:
9146381
负责人:
Subha Madhavan
金额:
$45.71万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-22 至 2018-05-31
关键词:
AddressAlgorithmsBig DataBig Data to KnowledgeBiologicalBiomedical ResearchCancer CenterClinicalClinical Decision Support SystemsClinical TrialsComputer softwareComputing MethodologiesCrowdingDataData AggregationDatabasesDictionaryDiseaseExclusion CriteriaGene ExpressionGene MutationGenomeGoalsGoldHealthInformaticsInformation RetrievalInvestmentsLettersLiteratureMalignant NeoplasmsMapsMeta-AnalysisMethodsMolecularMolecular ProfilingMolecular TargetMutationNational Cancer InstituteNatural Language ProcessingOncologistOnline SystemsOutcomePatientsPeer ReviewPharmaceutical PreparationsPharmacotherapyPhosphorylationProcessPubMedPublicationsRecording of previous eventsReportingResearchResearch DesignResearch PersonnelSoftware ValidationSourceStructureSystemSystems BiologyTestingTherapeuticTimeUnited States National Institutes of Healthabstractingbasecrowdsourcingdata to knowledgedata wranglingdesignimprovedinclusion criteriainnovationinterestknowledge basemeetingsnovelnovel strategiespersonalized cancer carepersonalized cancer therapypersonalized medicineprogramsprotein expressionsearch enginesoftware developmentsymposiumtargeted treatmenttooluser friendly softwareverification and validation
中文摘要
英文摘要
DESCRIPTION (provided by applicant): The velocity, variety, volume and veracity of data from relevant information sources make it extremely challenging for oncologists to collect and review pertinent data that can support routine personalized treatment for their patients. There is an urgent need to develop data wrangling approaches including Natural Language Processing and information retrieval methods to extract and curate personalized-therapy related publications and clinical trials. Once curated, the structured data can be used by biomedical researchers to generate novel scientific hypotheses, design new studies, obtain a better understanding of biological mechanisms of disease, perform meta-analyses, and create clinical decision support systems. There is an urgent need to develop improved search interfaces specific to the field of personalized therapy, including ways to display, rank, and save results by
end users. While several database and web-based keyword search engine algorithms exist, there is a lack of tools that meet the unique challenges of personalized medicine. There is also an urgent need to develop software that allows for verification and validation of information extracted and ranked through computational methods using subject matter expertise to improve the gold standard corpus that can be used for biomedical research into personalized therapies. To address these issues, we will build an innovative software stack (MACE2K) to adapt and extend widely tested Biocreative natural language processing (NLP) tools to automatically retrieve and pre-process targeted therapy information from clinicaltrials.gov, PubMed abstracts as well as open access articles, and conference proceedings. We will build an entity extraction cartridge to accurately parse gene mutations, translocations, gene expression, protein expression, and protein phosphorylation. A marker disambiguation cartridge will be built to assess for trial inclusion or exclusion criteria and to determine marker-related primary endpoints. We will include a ranking cartridge that uses the disambiguated information on markers, drugs and trials to provide a rigorous scoring of trials and studies according to their relevance for personalized medicine. A novel gamification cartridge will be built to allow subject matter experts to verify and validate the information corpus. Our research leverages National Cancer Institute's investments in several programs (many of which we are involved in) including the NCI drug dictionary, National Cancer Informatics Program (NCIP), I-SPY trials, and Center for cancer systems biology (CCSB) to efficiently accomplish our aims.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
MACE2K - Molecular And Clinical Extraction: A Natural Language Processing Tool for Personalized Medicine
-
批准号:8874546
-
项目类别:
-
资助金额:$47.87万
-
财政年份:2015
-
负责人:Subha Madhavan
-
依托单位:
MACE2K - Molecular And Clinical Extraction: A Natural Language Processing Tool for Personalized Medicine
-
批准号:9282279
-
项目类别:
-
资助金额:$45.55万
-
财政年份:2015
-
负责人:Subha Madhavan
-
依托单位:
Education, Training, and Outreach Program
-
批准号:8181012
-
项目类别:
-
资助金额:$12.49万
-
财政年份:2010
-
负责人:Subha Madhavan
-
依托单位:
BIOMEDICAL (MANAGEMENT/SUPPORT)
-
批准号:8065575
-
项目类别:
-
资助金额:$94.31万
-
财政年份:2009
-
负责人:Subha Madhavan
-
依托单位:
BIOMEDICAL (MANAGEMENT/SUPPORT)
-
批准号:8317482
-
项目类别:
-
资助金额:$120.0万
-
财政年份:2009
-
负责人:Subha Madhavan
-
依托单位:
BIOMEDICAL (MANAGEMENT/SUPPORT)
-
批准号:7880289
-
项目类别:
-
资助金额:$93.93万
-
财政年份:2009
-
负责人:Subha Madhavan
-
依托单位:
BIOMEDICAL (MANAGEMENT/SUPPORT)
-
批准号:7880291
-
项目类别:
-
资助金额:$93.93万
-
财政年份:2009
-
负责人:Subha Madhavan
-
依托单位:
Informatics Support Center for the Cancer Family Registries
-
批准号:8537027
-
项目类别:
-
资助金额:$120.0万
-
财政年份:2009
-
负责人:Subha Madhavan
-
依托单位:
BIOMEDICAL (MANAGEMENT/SUPPORT)
-
批准号:8065576
-
项目类别:
-
资助金额:$94.31万
-
财政年份:2009
-
负责人:Subha Madhavan
-
依托单位:
Education, Training, and Outreach Program
-
批准号:8517451
-
项目类别:
-
资助金额:$13.43万
-
财政年份:--
-
负责人:Subha Madhavan
-
依托单位:
Education, Training, and Outreach Program
-
批准号:8377526
-
项目类别:
-
资助金额:$11.81万
-
财政年份:--
-
负责人:Subha Madhavan
-
依托单位:
Education, Training, and Outreach Program
-
批准号:8232993
-
项目类别:
-
资助金额:$12.11万
-
财政年份:--
-
负责人:Subha Madhavan
-
依托单位:
Education, Training, and Outreach Program
-
批准号:8627144
-
项目类别:
-
资助金额:$11.71万
-
财政年份:--
-
负责人:Subha Madhavan
-
依托单位:
海外基金