课题基金 / 基金详情

SIPHS: Semantic interpretation of personal health messages for generating public health summaries

SIPHS: Semantic interpretation of personal health messages for generating public health summaries
SIPHS:个人健康信息的语义解释以生成公共卫生摘要
批准号:
EP/M005089/1
负责人:
Nigel Collier
金额:
$123.85万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2015
资助国家:
英国
项目状态:
已结题
起止时间:
2015 至 --

项目摘要

项目成果

Nigel Collier的其他基金

相似基金

相关文献

中文摘要
翻译
开放的在线数据,如微博和讨论板信息,有可能成为关于人口健康的极有价值的信息来源。这类数据增长迅速,成本低,实时,似乎可能覆盖很大一部分人口。举两个例子,PatientsLikeMe增长了10%,目前拥有超过20万用户,覆盖了1500多种健康状况;普通的Twitter服务正以每年30%的速度扩张,活跃用户超过2亿。超越简单的关键字搜索并利用这些数据为公共卫生服务,对自然语言处理(NLP)来说既是机遇也是挑战。这个奖学金提案是关于帮助健康专家利用社交媒体进行自己的临床和科学研究,通过自动技术根据机器可理解的语义表示对信息进行编码。该项目寻求解决三个主要挑战:(1)知识中介:开发算法,以识别和编码对标准本体(如UMLS)的条件、治疗、药物、行为和态度的非正式描述;(2)知识管理:创建博客文本中使用的患者词汇的结构化资源,并将其链接到现有的编码系统;(3)为证据增加洞察力:与领域专家合作,利用编码信息自动生成有意义的摘要,以便后续调查。在技术层面,该奖学金旨在开拓NLP和机器学习(ML)的新方法。社交媒体仍然是NLP的一个具有挑战性的领域,原因有很多:短的去语境化的信息,高度模糊/超出词汇量的单词,俚语的使用和不断发展的词汇,以及对耸人听闻的话题的固有偏见。该奖学金旨在利用NLP迄今在商业领域的社会媒体分析方面取得的进展,并进一步发展它,以提供有意义的公共卫生证据。以前没有提到的一个关键方面是患者信息的临床编码。尽管存在用于临床和科学文本的知识代理系统(例如NLM的MetaMap),但它们在社交媒体信息上的表现很差。该研究将利用生物医学中丰富的本体论资源,并结合机器学习对注释消息数据进行歧义消除。研究还将旨在了解信息的交际功能,例如信息是否报告直接经验或与新闻,幽默或营销有关。如果成功地克服了这些问题,将消除与其他类型临床数据集成的重要障碍。为社交媒体报告提供健康编码的优势在于,它有可能研究非常大规模的人群,并对异常情况进行实时早期预警。在奖学金中,我将研究从语义编码特征中提取多变量时间序列警报的潜力,与领域专家一起评估一系列指标(例如灵敏度、及时性、错误警报率)。将探索各种方法,以跨社交媒体来源生成实时风险摘要。我们选择了两个现实世界的应用来推动这一进程:药物不良反应(adr)的早期预警和传染病监测(IDS)。项目成果将包括基础技术以及开源算法、数据集和本体。该奖学金的一个令人兴奋的方面是各级利益攸关方之间的跨学科合作:科学家、公共卫生专家和工业界。最后,将通过发布开源数据向国际社会开放参与。将邀请从事社会媒体技术工作的同事在新挑战评估讲习班上与用户进行讨论。
英文摘要
Open online data such as microblogs and discussion board messages have the potential to be an incredibly valuable source of information about health in populations. Such data has been rapidly growing, is low cost, real-time and seems likely to cover a significant proportion of the demographic. To take two examples, PatientsLikeMe has enjoyed 10% growth and now has over 200,000 users covering over 1500 health conditions; the generic Twitter service is expanding at a rate of 30% annually with over 200 million active users. Going beyond simple keyword search and harnessing this data for public health represents both an opportunity and a challenge to natural language processing (NLP). This fellowship proposal is about helping health experts leverage social media for their own clinical and scientific studies through automatic techniques that encode messages according to a machine understandable semantic representation. There are three major challenges this project seeks to address: (1) knowledge brokering: to develop algorithms to identify and code the informal descriptions of conditions, treatments, medications, behaviours and attitudes to standard ontologies such as the UMLS; (2) knowledge management: to create a structured resource of patient vocabulary used in blog texts and link it to existing coding systems; and (3) adding insight to evidence: to work with domain experts to utilize the coded information to automatically generate meaningful summaries for follow up investigation. At the technological level the fellowship seeks to pioneer new methods for NLP and machine learning (ML). Social media remains a challenging area for NLP for a variety of reasons: short de-contextualised messages, high levels of ambiguity/out of vocabulary words, use of slang and an evolving vocabulary, as well as inherent bias towards sensational topics. The fellowship seeks to harness the progress made so far in NLP for social media analysis in the commercial domain and develop it further to provide meaningful public health evidence. One key aspect not previously addressed is in the clinical coding of patient messages. Although knowledge brokering systems exist for clinical and scientific texts (e.g. the NLM's MetaMap), their performance on social media messages has been poor. The fellowship will utilise the rich availability of ontological resources in biomedicine together with ML on annotated message data to disambiguate informal language. Research will also aim to understanding the communicative function of messages, for example whether the message reports direct experience or is related to news, humour or marketing. If these problems are successfully overcome an important barrier to data integration with other types of clinical data will be removed. The advantage of providing health coding for social media reports is its potential for studying very-large scale cohorts and also in real-time early alerting of aberrations. In the fellowship I will research the potential for multi-variate time series alerting from semantically coded features, working with domain experts to evaluate across a range of metrics (e.g. sensitivity, timeliness, false alerting rates). A variety of approaches will be explored to generate real time risk summaries across social media sources. Two real-world applications have been chosen to take this forwards: early alerting for Adverse drug reactions (ADRs) and Infectious disease surveillance (IDS). Project outcomes will include fundamental technologies as well as open source algorithms, data sets and ontology. An exciting aspect of this fellowship is inter-disciplinary collaboration across stakeholders at all levels: scientists, public health experts and industry. Finally, participation will be opened up to the international community through the release of open source data. Colleagues working on social media technologies will be invited to participate in discussions with users at a new challenge evaluation workshop.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
BioReddit: Word Embeddings for User-Generation Biomedical NLP
BioReddit:用于用户生成生物医学 NLP 的词嵌入
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者: [Basaldella M]
通讯作者: Basaldella M
WSDM 2017 Workshop on Mining Online Health Reports
WSDM 2017 挖掘在线健康报告研讨会
DOI: 10.1145/3018661.3022761
发表时间: 2017
期刊:
影响因子: --
作者: [Collier N]
通讯作者: Collier N
DOI: 10.18653/v1/2020.emnlp-main.253
发表时间: 2020-10
期刊: ArXiv
影响因子: --
作者: [Marco Basaldella;Fangyu Liu;Ehsan Shareghi;Nigel Collier]
通讯作者: Marco Basaldella;Fangyu Liu;Ehsan Shareghi;Nigel Collier
DOI: 10.18653/v1/n19-1298
发表时间: 2019-06
期刊:
影响因子: --
作者: [Duy-Cat Can;Hoang-Quynh Le;Quang-Thuy Ha;Nigel Collier]
通讯作者: Duy-Cat Can;Hoang-Quynh Le;Quang-Thuy Ha;Nigel Collier
共 6 条
    EPI-AI: Automated Understanding and Alerting of Disease Outbreaks from Global News Media
    • 批准号:
      ES/T012277/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $62.61万
    • 财政年份:
      2020
    • 负责人:
      Nigel Collier
    • 依托单位:
    PheneBank: automatic extraction and validation of a database of human phenotype-disease associations in the scientific literature
    • 批准号:
      MR/M025160/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $59.12万
    • 财政年份:
      2015
    • 负责人:
      Nigel Collier
    • 依托单位:
    海外基金