Natural language processing systems for extracting information from electronic health records about activities of daily living: a systematic review

Wieland-Jorna, Y., Kooten, D. van, Verheij, R.A., Man, Y. de, Francke, A.L., Oosterveld-Vlug, M.G. Natural language processing systems for extracting information from electronic health records about activities of daily living: a systematic review JAMIA Open: 2024, 7(2), p. Art. nr. ooae044.

Lees online

Objective
Natural language processing (NLP) can enhance research on activities of daily living (ADL) by extracting structured information from unstructured electronic health records (EHRs) notes. This review aims to give insight into the state-of-the-art, usability, and performance of NLP systems to extract information on ADL from EHRs.

Materials and Methods
A systematic review was conducted based on searches in Pubmed, Embase, Cinahl, Web of Science, and Scopus. Studies published between 2017 and 2022 were selected based on predefined eligibility criteria.

Results
The review identified 22 studies. Most studies (65%) used NLP for classifying unstructured EHR data on 1 or 2 ADL. Deep learning, combined with a ruled-based method or machine learning, was the approach most commonly used. NLP systems varied widely in terms of the pre-processing and algorithms. Common performance evaluation methods were cross-validation and train/test datasets, with F1, precision, and sensitivity as the most frequently reported evaluation metrics. Most studies reported relativity high overall scores on the evaluation metrics.

Discussion
NLP systems are valuable for the extraction of unstructured EHR data on ADL. However, comparing the performance of NLP systems is difficult due to the diversity of the studies and challenges related to the dataset, including restricted access to EHR data, inadequate documentation, lack of granularity, and small datasets.

Conclusion
This systematic review indicates that NLP is promising for deriving information on ADL from unstructured EHR notes. However, what the best-performing NLP system is, depends on characteristics of the dataset, research question, and type of ADL.

Contact

R.A. (Robert) Verheij

Programmaleider Zorgdata en het Lerend Zorgsysteem; bijzonder hoogleraar 'Transparantie in de zorg vanuit patiëntenperspectief', Tranzo, Tilburg University

0302729657

http://www.twitter.com/robtwit01

https://www.linkedin.com/in/robert-verheij-b5a09b15/

r.verheij@nivel.nl