A recent Nature Machine Intelligence study investigated the efficacy of audio-based artificial intelligence (AI) classifiers in predicting severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) infection status. SARS-CoV-2 is the causal organism of the coronavirus disease 2019 (COVID-19) pandemic.

Background
Since SARS-CoV-2 infection could cause both symptomatic and asymptomatic manifestations, it is important to develop accurate tests to avoid general population quarantine. Previous studies have revealed that AI-based classifiers trained with respiratory audio data could identify SARS-CoV-2 status.
Although these studies indicated the effectiveness of AI-based classifiers, many challenges surfaced while applying them in real-world settings. Some factors that withheld AI-based classifier applications were sampling biases, unvalidated data on participants’ COVID-19 status, and delay between infection and audio recording. It is imperative to determine whether the audio biomarkers of COVID-19 are unique to SARS-CoV-2 infection or are inappropriate confounding signals.
About the Study
The current study focussed on determining whether audio-based classifiers can be accurately used for COVID-19 screening. A large-scale polymerase chain reaction (PCR) dataset linked to audio-based COVID-19 screening (ABCS) was used. For this study, participants of the Real-time Assessment of Community Transmission (REACT) program and the National Health Service (NHS) Test-and-Trace (T+T) service were invited. All relevant demographic data was extracted from T+T/REACT records.
Participants were asked to complete survey questions and record four audio clips. For audio recordings, they were asked to read a specific sentence, followed by three successive exhalations, making a “ha” sound. Furthermore, the participants were asked to record forced coughs once and three times in succession. All recordings were documented in .wav format. The quality of the audio recordings was assessed, and 5,157 records were removed for quality-related issues.
Human figures represent study participants and their corresponding COVID-19 infection status, with the different colours portraying different demographic or symptomatic features. When participants are randomly split into training and test sets, the randomized split models perform well at COVID-19 detection, achieving AUCs in excess of 0.8; however, matched test set performance is seen to drop to estimated AUC between 0.60 and 0.65, with an AUC of 0.5 representing random classification. Inflated classification performance is also seen in engineered out of distribution test sets such as: the designed test set, in which a select set of demographic groups appear solely in the testing set, and the longitudinal test set, in which there is no overlap in the time of submission between train and test instances. The 95% confidence intervals calculated via the normal approximation method are shown, along with the corresponding n numbers of the train and test sets.