Machine learning models for abstract screening task - A systematic literature review application for health economics and outcome research

Objective Systematic literature reviews (SLRs) are critical for life-science research. However, the manual selection and retrieval of relevant publications can be a time-consuming process. This study aims to (1) develop two disease-specific annotated corpora, one for human papillomavirus (HPV) assoc...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Veröffentlicht in:BMC medical research methodology Jg. 24; H. 1; S. 108 - 7
Hauptverfasser: Du, Jingcheng, Soysal, Ekin, Wang, Dong, He, Long, Lin, Bin, Wang, Jingqi, Manion, Frank J., Li, Yeran, Wu, Elise, Yao, Lixia
Format: Journal Article
Sprache:Englisch
Veröffentlicht: London BioMed Central 09.05.2024
BioMed Central Ltd
Springer Nature B.V
BMC
Schlagworte:
ISSN:1471-2288, 1471-2288
Online-Zugang:Volltext
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
Beschreibung
Zusammenfassung:Objective Systematic literature reviews (SLRs) are critical for life-science research. However, the manual selection and retrieval of relevant publications can be a time-consuming process. This study aims to (1) develop two disease-specific annotated corpora, one for human papillomavirus (HPV) associated diseases and the other for pneumococcal-associated pediatric diseases (PAPD), and (2) optimize machine- and deep-learning models to facilitate automation of the SLR abstract screening. Methods This study constructed two disease-specific SLR screening corpora for HPV and PAPD, which contained citation metadata and corresponding abstracts. Performance was evaluated using precision, recall, accuracy, and F1-score of multiple combinations of machine- and deep-learning algorithms and features such as keywords and MeSH terms. Results and conclusions The HPV corpus contained 1697 entries, with 538 relevant and 1159 irrelevant articles. The PAPD corpus included 2865 entries, with 711 relevant and 2154 irrelevant articles. Adding additional features beyond title and abstract improved the performance (measured in Accuracy) of machine learning models by 3% for HPV corpus and 2% for PAPD corpus. Transformer-based deep learning models that consistently outperformed conventional machine learning algorithms, highlighting the strength of domain-specific pre-trained language models for SLR abstract screening. This study provides a foundation for the development of more intelligent SLR systems.
Bibliographie:ObjectType-Article-1
SourceType-Scholarly Journals-1
content type line 14
ObjectType-Literature Review-2
ObjectType-Feature-3
ObjectType-Feature-2
content type line 23
ObjectType-Undefined-3
ISSN:1471-2288
1471-2288
DOI:10.1186/s12874-024-02224-3