POET-2: High-performance computing for advanced clinical narrative preprocessing
POET-2: High-performance computing for advanced clinical narrative preprocessing
批准号:
8326648
负责人:
JOHN F. HURDLE
金额:
$31.84万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-01 至 2014-08-31
关键词:
AbbreviationsActive LearningAddressArchitectureAreaAuthorization documentationCaringCharacteristicsClinicalClinical DataComputer softwareDataElectronic Health RecordElectronicsEmployee StrikesEnsureEnvironmentEvaluationFaceGoldGrowthHealth Care ReformHealth Insurance Portability and Accountability ActHealth PersonnelHealthcareHigh Performance ComputingInformaticsInpatientsInstitutionInstitutional Review BoardsInvestmentsLaboratoriesLanguageLinguisticsLiteratureMapsMedicalMiningModelingNatural Language ProcessingNew YorkNoiseOccupationsOutpatientsPaperPathologyPathology ReportPatientsPerformanceProliferatingPublishingRadiology SpecialtyRecordsReport (document)ReportingResearchResearch DesignResearch PersonnelResearch SupportResolutionSeriesShorthandSignal TransductionSourceStructureSummary ReportsSystemTechniquesTechnologyTextTimeUniversitiesVotingWorkWritingabstractingbasecluster computingdata miningdesignimprovedmeetingsnovelphrasespressurerepairedrepositoryresearch studysuccesssymposiumtoolweb services
中文摘要
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant):
This project focuses on clinical natural language processing (cNLP), a field of emerging importance in informatics. Starting with the Linguistic String Project's Medical Language Processor (New York University) in the 1970s, researchers have made steady gains in cNLP through empirical studies and by building sophisticated high-level cNLP software applications (e.g., Columbia's MedLEE). There are no fewer than four scientific conferences devoted exclusively to biomedical/clinical NLP. The cNLP literature has been growing over the past decade, and this will gain momentum as more clinical text repositories are released, such as the MIMIC II and University of Pittsburgh BLU Lab corpora.
However, sustained success in the field of cNLP is hampered by the reality that clinical texts have a far more noise than do texts traditionally studied in NLP, such as newswire articles, biomedical abstracts, and discharge summaries. Noise in this context is defined by the parseability characteristics of the language and the linguistic structures that appear in text. Clinical texts come in a striking variety of note types, with the best studied types being discharge summaries, radiology reports, and pathology reports. These note types share an important feature: they are written to communicate care issues between healthcare providers and hence typically are well-composed, well-edited, and often are dictated. But the vast majority of notes in the electronic health record are written primarily to document care issues. They communicate as well, of course, but much less care is used in their creation than with discharge summaries and reports. As a result they are often ungrammatical; are composed of short, telegraphic phrases; are replete with misspellings and shorthand (e.g., abbreviations); are ill-formatted with templates and liberal use of white space; and are embedded with "non-prose" (e.g., strings of laboratory values). All of these sources of noise complicate otherwise straightforward NLP tasks like tokenization, sentence segmentation, and ultimately information extraction itself.
We propose a systematic study of ways to increase the signal-to-noise ratio in clinical narratives to improve cNLP. This work extends our preliminary research (under the POET project) and has the following aims:
o Develop and implement a suite of parseability improvement tools designed for all clinical note types from multiple healthcare institutions.
o Evaluate the empirical and the functional success of the parseability improvement tools.
o Design and implement a HIPAA-compliant UlMA-based pipeline cNLP framework for use in a typical high-performance, multi-processor computing environment.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
University of Utah Biomedical Informatics Training Grant Supplement
-
批准号:9380137
-
项目类别:
-
资助金额:$0.24万
-
财政年份:2016
-
负责人:JOHN F. HURDLE
-
依托单位:
POET-2: High-performance computing for advanced clinical narrative preprocessing
-
批准号:8182025
-
项目类别:
-
资助金额:$32.52万
-
财政年份:2011
-
负责人:JOHN F. HURDLE
-
依托单位:
POET: Consolidated, Comprehensive Clinical Text Preprocessing
-
批准号:7570254
-
项目类别:
-
资助金额:$16.93万
-
财政年份:2008
-
负责人:JOHN F. HURDLE
-
依托单位:
POET: Consolidated, Comprehensive Clinical Text Preprocessing
-
批准号:7689273
-
项目类别:
-
资助金额:$16.66万
-
财政年份:2008
-
负责人:JOHN F. HURDLE
-
依托单位:
POET: Consolidated, Comprehensive Clinical Text Preprocessing
-
批准号:7847940
-
项目类别:
-
资助金额:$8.47万
-
财政年份:2008
-
负责人:JOHN F. HURDLE
-
依托单位:
Statistical NLP Analysis of Cross-discipline Clinical Text
-
批准号:6836781
-
项目类别:
-
资助金额:$9.45万
-
财政年份:2004
-
负责人:JOHN F. HURDLE
-
依托单位:
Statistical NLP Analysis of Cross-discipline Clinical Text
-
批准号:6944955
-
项目类别:
-
资助金额:$3.88万
-
财政年份:2004
-
负责人:JOHN F. HURDLE
-
依托单位:
University of Utah Biomedical Informatics Training Grant
-
批准号:8681515
-
项目类别:
-
资助金额:$91.04万
-
财政年份:1997
-
负责人:JOHN F. HURDLE
-
依托单位:
University of Utah Biomedical Informatics Training Grant
-
批准号:8261299
-
项目类别:
-
资助金额:$83.68万
-
财政年份:1997
-
负责人:JOHN F. HURDLE
-
依托单位:
University of Utah Biomedical Informatics Training Grant
-
批准号:9086432
-
项目类别:
-
资助金额:$89.02万
-
财政年份:1997
-
负责人:JOHN F. HURDLE
-
依托单位:
海外基金