Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Dec 20, 2017
Date Accepted: May 15, 2018
(closed for review but you can still tweet)

The final, peer-reviewed published version of this preprint can be found here:

Augmenting Qualitative Text Analysis with Natural Language Processing: Methodological Study

Guetterman TC, Chang T, DeJonckheere M, Basu T, Scruggs E, Vydiswaran VGV

Augmenting Qualitative Text Analysis with Natural Language Processing: Methodological Study

J Med Internet Res 2018;20(6):e231

DOI: 10.2196/jmir.9702

PMID: 29959110

PMCID: 6045788

Augmenting Qualitative Text Analysis with Natural Language Processing: Methodological Study

  • Timothy C Guetterman; 
  • Tammy Chang; 
  • Melissa DeJonckheere; 
  • Tanmay Basu; 
  • Elizabeth Scruggs; 
  • V G Vinod Vydiswaran

ABSTRACT

Background:

Qualitative research methods are increasingly being used across disciplines because of their ability to help investigators understand the perspectives of participants in their own words. However, qualitative analysis is a laborious and resource-intensive process. To achieve depth, researchers are limited to smaller sample sizes when analyzing text data. One potential method to address this concern is natural language processing (NLP). Qualitative text analysis involves researchers reading data, assigning code labels, and iteratively developing findings; NLP has the potential to automate part of this process. Unfortunately, little methodological research has been done to compare automatic coding using NLP techniques and qualitative coding, which is critical to establish the viability of NLP as a useful, rigorous analysis procedure.

Objective:

The purpose of this study was to compare the utility of a traditional qualitative text analysis, an NLP analysis, and an augmented approach that combines qualitative and NLP methods.

Methods:

We conducted a 2-arm cross-over experiment to compare qualitative and NLP approaches to analyze data generated through 2 text (short message service) message survey questions, one about prescription drugs and the other about police interactions, sent to youth aged 14-24 years. We randomly assigned a question to each of the 2 experienced qualitative analysis teams for independent coding and analysis before receiving NLP results. A third team separately conducted NLP analysis of the same 2 questions. We examined the results of our analyses to compare (1) the similarity of findings derived, (2) the quality of inferences generated, and (3) the time spent in analysis.

Results:

The qualitative-only analysis for the drug question (n=58) yielded 4 major findings, whereas the NLP analysis yielded 3 findings that missed contextual elements. The qualitative and NLP-augmented analysis was the most comprehensive. For the police question (n=68), the qualitative-only analysis yielded 4 primary findings and the NLP-only analysis yielded 4 slightly different findings. Again, the augmented qualitative and NLP analysis was the most comprehensive and produced the highest quality inferences, increasing our depth of understanding (ie, details and frequencies). In terms of time, the NLP-only approach was quicker than the qualitative-only approach for the drug (120 vs 270 minutes) and police (40 vs 270 minutes) questions. An approach beginning with qualitative analysis followed by qualitative- or NLP-augmented analysis took longer time than that beginning with NLP for both drug (450 vs 240 minutes) and police (390 vs 220 minutes) questions.

Conclusions:

NLP provides both a foundation to code qualitatively more quickly and a method to validate qualitative findings. NLP methods were able to identify major themes found with traditional qualitative analysis but were not useful in identifying nuances. Traditional qualitative text analysis added important details and context.


 Citation

Please cite as:

Guetterman TC, Chang T, DeJonckheere M, Basu T, Scruggs E, Vydiswaran VGV

Augmenting Qualitative Text Analysis with Natural Language Processing: Methodological Study

J Med Internet Res 2018;20(6):e231

DOI: 10.2196/jmir.9702

PMID: 29959110

PMCID: 6045788

Per the author's request the PDF is not available.

Advertisement