TY - GEN
T1 - Systematic Evaluation and Enhancement of Speech Recognition in Operational Medical Environments
AU - Kar, Snigdhaswin
AU - Mishra, Prabodh
AU - Lin, Ju
AU - Woo, Min Jae
AU - Deas, Nicholas
AU - Linduff, Caleb
AU - Niu, Sufeng
AU - Yang, Yuzhe
AU - McClendon, Jerome
AU - Smith, D. Hudson
AU - Smith, Melissa C.
AU - Gimbel, Ronald W.
AU - Wang, Kuang Ching
N1 - Publisher Copyright: © 2021 IEEE.
PY - 2021/7/18
Y1 - 2021/7/18
N2 - Operational medical environments require reliable hands-free solutions to extract data from audio captured under noisy scenarios during rescue missions and provide timely information. However, approaches using automatic speech recognition (ASR) and natural language processing (NLP) techniques are complex as these conversations have a wide range of noise, involve medical terms from multiple speakers, and occur in high-stress environments, among others. These are further complicated by the lack of large training datasets for operational medical scenarios. To address these issues, we developed a platform that enables resilient hands-free data collection, preserves complete documentation through stages of care, and presents the information in near real-time, critical for the medical operation. Our work uniquely focused on systematic evaluation and improvement of a deep neural network-based ASR system by leveraging realistic testing data obtained from medical simulations of battlefield scenarios, which to our knowledge have not been addressed in any prior work. The system performance is shown to improve significantly using multi-style training, language model adaptation for the medical domain, speech enhancement, and NLP techniques.
AB - Operational medical environments require reliable hands-free solutions to extract data from audio captured under noisy scenarios during rescue missions and provide timely information. However, approaches using automatic speech recognition (ASR) and natural language processing (NLP) techniques are complex as these conversations have a wide range of noise, involve medical terms from multiple speakers, and occur in high-stress environments, among others. These are further complicated by the lack of large training datasets for operational medical scenarios. To address these issues, we developed a platform that enables resilient hands-free data collection, preserves complete documentation through stages of care, and presents the information in near real-time, critical for the medical operation. Our work uniquely focused on systematic evaluation and improvement of a deep neural network-based ASR system by leveraging realistic testing data obtained from medical simulations of battlefield scenarios, which to our knowledge have not been addressed in any prior work. The system performance is shown to improve significantly using multi-style training, language model adaptation for the medical domain, speech enhancement, and NLP techniques.
KW - Automatic Speech Recognition
KW - Multi-style training
KW - Natural Language Processing
KW - Operational Medical Environments
KW - Prehospital documentation
KW - Speech enhancement
UR - https://www.scopus.com/pages/publications/85116414742
U2 - 10.1109/IJCNN52387.2021.9533607
DO - 10.1109/IJCNN52387.2021.9533607
M3 - Conference contribution
T3 - Proceedings of the International Joint Conference on Neural Networks
BT - IJCNN 2021 - International Joint Conference on Neural Networks, Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2021 International Joint Conference on Neural Networks, IJCNN 2021
Y2 - 18 July 2021 through 22 July 2021
ER -