A machine learning framework for discovery and enrichment of metagenomics metadata from open access publications.
A machine learning framework for discovery and enrichment of metagenomics metadata from open access publications.
复制标题
DOI:
10.1093/gigascience/giac077
复制
发表时间:
2022-08-11
期刊:
影响因子:
9.2
通讯作者:
中科院分区:
文献类型:
--
作者:
Metagenomics is a culture-independent method for studying the microbes inhabiting a particular environment. Comparing the composition of samples (functionally/taxonomically), either from a longitudinal study or cross-sectional studies, can provide clues into how the microbiota has adapted to the environment. However, a recurring challenge, especially when comparing results between independent studies, is that key metadata about the sample and molecular methods used to extract and sequence the genetic material are often missing from sequence records, making it difficult to account for confounding factors. Nevertheless, these missing metadata may be found in the narrative of publications describing the research. Here, we describe a machine learning framework that automatically extracts essential metadata for a wide range of metagenomics studies from the literature contained in Europe PMC. This framework has enabled the extraction of metadata from 114,099 publications in Europe PMC, including 19,900 publications describing metagenomics studies in European Nucleotide Archive (ENA) and MGnify. Using this framework, a new metagenomics annotations pipeline was developed and integrated into Europe PMC to regularly enrich up-to-date ENA and MGnify metagenomics studies with metadata extracted from research articles. These metadata are now available for researchers to explore and retrieve in the MGnify and Europe PMC websites, as well as Europe PMC annotations API.
登录
查看更多内容
影响因子:
1.9
作者:
Buttigieg PL;Morrison N;Smith B;Mungall CJ;Lewis SE;ENVO Consortium
通讯作者:
ENVO Consortium
影响因子:
64.8
作者:
Jumper J;Evans R;Pritzel A;Green T;Figurnov M;Ronneberger O;Tunyasuvunakool K;Bates R;Žídek A;Potapenko A;Bridgland A;Meyer C;Kohl SAA;Ballard AJ;Cowie A;Romera-Paredes B;Nikolov S;Jain R;Adler J;Back T;Petersen S;Reiman D;Clancy E;Zielinski M;Steinegger M;Pacholska M;Berghammer T;Bodenstein S;Silver D;Vinyals O;Senior AW;Kavukcuoglu K;Kohli P;Hassabis D
通讯作者:
Hassabis D
影响因子:
14.9
作者:
Harrison PW;Ahamed A;Aslam R;Alako BTF;Burgin J;Buso N;Courtot M;Fan J;Gupta D;Haseeb M;Holt S;Ibrahim T;Ivanov E;Jayathilaka S;Balavenkataraman Kadhirvelu V;Kumar M;Lopez R;Kay S;Leinonen R;Liu X;O'Cathail C;Pakseresht A;Park Y;Pesant S;Rahman N;Rajan J;Sokolov A;Vijayaraja S;Waheed Z;Zyoud A;Burdett T;Cochrane G
通讯作者:
Cochrane G
DOI:
10.1093/bioinformatics/btaa586
发表时间:
2020-09-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
Bagheri H;Severin AJ;Rajan H
通讯作者:
Rajan H
影响因子:
14.9
作者:
Chen IA;Markowitz VM;Chu K;Palaniappan K;Szeto E;Pillay M;Ratner A;Huang J;Andersen E;Huntemann M;Varghese N;Hadjithomas M;Tennessen K;Nielsen T;Ivanova NN;Kyrpides NC
通讯作者:
Kyrpides NC