JMIR Preprints #22982: Extracting Family History of Patients from Clinical Narratives: Using Deep Learning Models

Current Preprint Settings

(as selected by the authors)

1. When the manuscript is submitted, allow peer review from:

(a) Anybody (open community peer review)
(b) Editor-selected reviewers (closed peer review)

2. When the manuscript is submitted, display the preprint PDF to:

(a) Anybody, anytime
(b) Logged-in users only
(c) Anybody, anytime (title and abstract only)
(d) No one

3. When the manuscript is accepted, display the accepted manuscript PDF to:

(a) Anybody, anytime
(b) Logged-in users only
(c) Anybody, anytime (title and abstract only)

Extracting Family History of Patients from Clinical Narratives: Using Deep Learning Models

Xi Yang;
Hansi Zhang;
Xing He;
Jiang Bian;
Yonghui Wu

ABSTRACT

Background:

Patients’ family history (FH) is a critical risk factor associated with numerous diseases. However, FH information is not well captured in the structured database but often documented in clinical narratives. Natural language processing (NLP) is the key technology to extract patients’ FH from clinical narratives. In 2019, the National NLP Clinical Challenge (N2C2) organized shared tasks to solicit NLP methods for FH information extraction.

Objective:

This study presents our end-to-end FH extraction system developed during the 2019 N2C2 challenge as well as the new transformer-based models that we developed after the challenge.

Methods:

We developed deep learning-based systems for FH concept extraction and relation identification. We explored deep learning models including Long-short term memory – conditional random fields (LSTM-CRFs) and Bidirectional Encoder Representations from Transformers (BERT) as well as developed ensemble models using a majority voting strategy. To further optimize performance, we systematically compared three different strategies to use BERT output representations for relation identification.

Results:

Our system was among the top-ranked systems in the challenge. Our best system submitted during this challenge achieved micro-averaged F1-scores of 0.7944 and 0.6544 for concept extraction and relation identification, respectively. After challenge, we further explored new transformer-based models and improved the performances of both subtasks to 0.8249 and 0.6775, respectively. For relation identification, our system achieved a performance comparable to the best system (0.6810) reported in the challenge.

Conclusions:

This study demonstrated the feasibility of utilizing deep learning methods to extract family history information from clinical narratives automatically.

Citation

Please cite as:

Yang X, Zhang H, He X, Bian J, Wu Y

Extracting Family History of Patients From Clinical Narratives: Exploring an End-to-End Solution With Deep Learning Models

JMIR Med Inform 2020;8(12):e22982

DOI: 10.2196/22982

PMID: 33320104

PMCID: 7772072

JMIR Publications

JMIR Preprints

Accepted for/Published in: JMIR Medical Informatics

Date Submitted: Jul 28, 2020

Date Accepted: Nov 20, 2020

Extracting Family History of Patients from Clinical Narratives: Using Deep Learning Models

ABSTRACT

Citation

JMIR Preprints

Accepted for/Published in: JMIR Medical Informatics

Date Submitted: Jul 28, 2020

Date Accepted: Nov 20, 2020

Extracting Family History of Patients from Clinical Narratives: Using Deep Learning Models

ABSTRACT

Citation

Per the author's request the PDF is not available.