Automated System to Capture Patient Symptoms From Multitype Japanese Clinical Texts: Retrospective Study

JMIR Med Inform. 2024 Sep 24:12:e58977. doi: 10.2196/58977.

Abstract

Background: Natural language processing (NLP) techniques can be used to analyze large amounts of electronic health record texts, which encompasses various types of patient information such as quality of life, effectiveness of treatments, and adverse drug event (ADE) signals. As different aspects of a patient's status are stored in different types of documents, we propose an NLP system capable of processing 6 types of documents: physician progress notes, discharge summaries, radiology reports, radioisotope reports, nursing records, and pharmacist progress notes.

Objective: This study aimed to investigate the system's performance in detecting ADEs by evaluating the results from multitype texts. The main objective is to detect adverse events accurately using an NLP system.

Methods: We used data written in Japanese from 2289 patients with breast cancer, including medication data, physician progress notes, discharge summaries, radiology reports, radioisotope reports, nursing records, and pharmacist progress notes. Our system performs 3 processes: named entity recognition, normalization of symptoms, and aggregation of multiple types of documents from multiple patients. Among all patients with breast cancer, 103 and 112 with peripheral neuropathy (PN) received paclitaxel or docetaxel, respectively. We evaluate the utility of using multiple types of documents by correlation coefficient and regression analysis to compare their performance with each single type of document. All evaluations of detection rates with our system are performed 30 days after drug administration.

Results: Our system underestimates by 13.3 percentage points (74.0%-60.7%), as the incidence of paclitaxel-induced PN was 60.7%, compared with 74.0% in the previous research based on manual extraction. The Pearson correlation coefficient between the manual extraction and system results was 0.87 Although the pharmacist progress notes had the highest detection rate among each type of document, the rate did not match the performance using all documents. The estimated median duration of PN with paclitaxel was 92 days, whereas the previously reported median duration of PN with paclitaxel was 727 days. The number of events detected in each document was highest in the physician's progress notes, followed by the pharmacist's and nursing records.

Conclusions: Considering the inherent cost that requires constant monitoring of the patient's condition, such as the treatment of PN, our system has a significant advantage in that it can immediately estimate the treatment duration without fine-tuning a new NLP model. Leveraging multitype documents is better than using single-type documents to improve detection performance. Although the onset time estimation was relatively accurate, the duration might have been influenced by the length of the data follow-up period. The results suggest that our method using various types of data can detect more ADEs from clinical documents.

Keywords: EHR; EHRs; ML; NLP; adverse; adverse drug reaction; adverse event; cancer; detect; detecting; detection; drug; drugs; machine learning; medication; medications; named entity recognition; natural language processing; neuropathy; note; notes; oncology; peripheral neuropathy; pharmaceutic; pharmaceutical; pharmaceuticals; pharmaceutics; pharmacology; pharmacotherapy; record; records; report; reports; symptom; symptoms; text; texts; textual.

MeSH terms

  • Breast Neoplasms / drug therapy
  • Breast Neoplasms / pathology
  • Drug-Related Side Effects and Adverse Reactions / diagnosis
  • Drug-Related Side Effects and Adverse Reactions / epidemiology
  • East Asian People
  • Electronic Health Records*
  • Female
  • Humans
  • Japan
  • Natural Language Processing*
  • Retrospective Studies