Extending electronic medical records vector models with knowledge graphs to improve hospitalization prediction

Background Artificial intelligence methods applied to electronic medical records (EMRs) hold the potential to help physicians save time by sharpening their analysis and decisions, thereby improving the health of patients. On the one hand, machine learning algorithms have proven their effectiveness i...

Full description

Saved in:

Bibliographic Details
Published in:	Journal of biomedical semantics Vol. 13; no. 1; pp. 6 - 20
Main Authors:	Gazzotti, Raphaël, Faron, Catherine, Gandon, Fabien, Lacroix-Hugues, Virginie, Darmon, David
Format:	Journal Article
Language:	English
Published:	London BioMed Central 22.02.2022 BioMed Central Ltd BMC
Subjects:	Algorithms Artificial Intelligence Bioinformatics Combinatorial Libraries Computation and Language Computational Biology/Bioinformatics Computer Appl. in Life Sciences Computer Science Data Mining and Knowledge Discovery Document and Text Processing Electronic Health Records Electronic medical records Electronic records Feature selection Forecasts and trends Hospital care Hospitalization Humans Knowledge graph Machine Learning Mathematics Mathematics and Statistics Medical informatics Medical records Methods Pattern Recognition, Automated Predictive model Statistics Web France Feature selection Electronic medical records Predictive model Knowledge graph predictive model electronic medical records knowledge graph feature selection
ISSN:	2041-1480, 2041-1480
Online Access:	Get full text
Tags:	Add Tag No Tags, Be the first to tag this record!

Description
Summary:	Background Artificial intelligence methods applied to electronic medical records (EMRs) hold the potential to help physicians save time by sharpening their analysis and decisions, thereby improving the health of patients. On the one hand, machine learning algorithms have proven their effectiveness in extracting information and exploiting knowledge extracted from data. On the other hand, knowledge graphs capture human knowledge by relying on conceptual schemas and formalization and supporting reasoning. Leveraging knowledge graphs that are legion in the medical field, it is possible to pre-process and enrich data representation used by machine learning algorithms. Medical data standardization is an opportunity to jointly exploit the richness of knowledge graphs and the capabilities of machine learning algorithms. Methods We propose to address the problem of hospitalization prediction for patients with an approach that enriches vector representation of EMRs with information extracted from different knowledge graphs before learning and predicting. In addition, we performed an automatic selection of features resulting from knowledge graphs to distinguish noisy ones from those that can benefit the decision making. We report the results of our experiments on the PRIMEGE PACA database that contains more than 600,000 consultations carried out by 17 general practitioners (GPs). Results A statistical evaluation shows that our proposed approach improves hospitalization prediction. More precisely, injecting features extracted from cross-domain knowledge graphs in the vector representation of EMRs given as input to the prediction algorithm significantly increases the F1 score of the prediction. Conclusions By injecting knowledge from recognized reference sources into the representation of EMRs, it is possible to significantly improve the prediction of medical events. Future work would be to evaluate the impact of a feature selection step coupled with a combination of features extracted from several knowledge graphs. A possible avenue is to study more hierarchical levels and properties related to concepts, as well as to integrate more semantic annotators to exploit unstructured data.
Bibliography:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 23 PMCID: PMC8861628
ISSN:	2041-1480 2041-1480
DOI:	10.1186/s13326-022-00261-9