The Psychiatric Record | Research Methods
A speech model may appear to recognize depression or anxiety from an interview. Before deciding what that means clinically, ask a more basic question: whose speech, and which parts of the interaction, were in the audio the model learned from?
The recording was the diagnostic interview
A study published online September 9, 2026 in the Journal of Clinical Psychiatry (Bayrampour and colleagues) developed speech-based prediction models for major depressive disorder and generalized anxiety disorder during pregnancy. Data came from 146 pregnant participants in British Columbia, recorded between July 2019 and May 2020: 19 with MDD, 28 with GAD, 3 of whom had both, and 102 controls with no known current or past psychiatric disorder.
Diagnoses were established with the Structured Clinical Interview for DSM-5. The audio fed to the models was the recording of that same interview. The authors ran two versions: one on the full recording, containing both interviewer and patient, and one on manually verified segments of patient speech only.
The two versions did not perform alike. On the full interview, F1 scores were 74 to 77 percent for MDD and 76 to 80 percent for GAD. On patient speech alone, F1 fell to 42 to 45 percent for MDD, which the authors call poor, and 51 to 65 percent for GAD, which they call acceptable. Adding pregnancy and sociodemographic variables did not improve the best GAD models and slightly reduced performance for MDD.
An accompanying commentary by Mark Opler presents the study as a promising case for speech as a biomarker. The difference between full-interview and patient-only performance makes the recording conditions central to assessing that promise.
What the gap tells us
Separating the speakers materially changed the reported performance. That makes the composition of the recording central to interpreting the result.
The comparison does not establish why performance fell. Removing the interviewer changes both the amount and composition of the audio; differences in processing may also matter. The MDD results also come from a sample containing just 19 participants with depression, making replication in a larger sample important.
The interview's structure offers another possible explanation. The SCID does not take every participant through an identical sequence of questions. Columbia University's SCID-5 documentation gives a concrete example: assessment of past GAD is completed only if criteria for current GAD are not met. Additional questions can follow either the presence or the absence of diagnostic criteria, depending on the branch. The implication is that interviews can differ in shape, not that cases or controls must have longer recordings. If those differences survive in the features supplied to a model, it could learn aspects of the diagnostic process alongside characteristics of the patient's voice. That is a possible mechanism, not a demonstrated explanation for this study's results.
Patient-only segments also remain responses elicited during a diagnostic interview. Removing the interviewer's voice does not remove the influence of the questions on those responses. Their performance matters for a tool intended to receive only patient speech, but it still does not establish performance outside that interview setting.
The practical question is therefore what recording a proposed clinical tool would receive. Evidence from a complete diagnostic interview cannot automatically support a tool intended to analyze a patient speaking under different conditions.
To explain the gap, readers would need the full methods and relevant analyses: how much audio each condition contained, how recordings were processed, how development and evaluation data were separated, and whether the models were evaluated in an independent population. The abstract alone does not answer those questions.
What the pregnancy sample can tell us
The authors conclude that the most influential speech features for prediction in pregnancy aligned with those reported in the general population.
That is a useful observation, but it does not establish that a model will perform equally well across populations. All participants in this study were pregnant, including the controls. Simply distinguishing pregnancy from non-pregnancy therefore cannot explain the classification of cases and controls in this sample.
For a perinatal reader, the next question is whether the model performs reliably in other pregnant populations and in the recording conditions of the intended clinical setting. Similar influential features provide a reason for further study, not an answer to that question.
From a metric to a decision
An F1 score combines precision and recall for a set of classification results. It is not the proportion of patients correctly classified, and it carries no information about what a missed case or a false positive would cost.
The question that follows is what a model output would change. Flagging a patient for a fuller assessment, ordering an existing queue, and assigning a diagnosis are three different jobs, and each requires evidence gathered under the conditions of that job. Performance on recorded SCID interviews does not establish performance on a two-minute intake recording, a telephone call, or a waiting-room kiosk. Each may elicit different speech and provide different conversational context.
Before applying either set of results, identify whose speech the proposed tool would receive, how that speech would be elicited, and what decision the output would inform. The reported performance needs to match those conditions before it can support that use.
Sources and reading notes
Bayrampour H, Amorim J, Pimentel J, Tamana SK, Janssen P, Bacon A, Rudzicz F. Using Speech to Develop Multivariable Prediction Models for Major Depressive Disorder and Generalized Anxiety Disorder During Pregnancy. J Clin Psychiatry. 2026;87(4):26m16487. Published online September 9, 2026. Study abstract. DOI: 10.4088/JCP.26m16487. Abstract reviewed September 15, 2026; full text is subscriber-only and was not opened. Sample size, diagnostic method, F1 ranges by condition, and the feature-alignment conclusion are taken from the abstract. Data separation, audio duration by condition, and the presence of any external evaluation remain unverified.
Opler M. Voice of the Patient. J Clin Psychiatry. 2026;87(4):26com16630. Published online September 9, 2026. Commentary. DOI: 10.4088/JCP.26com16630. Cited for its framing of speech as a promising biomarker.
Columbia University Department of Psychiatry. Changes to SCID-5. This page summarizes changes from SCID-IV. See Changes to Module F for the conditional assessment of past GAD. This documents a feature of the interview instrument, not how the study configured its interviews or which information its models used.
For professional education. This article does not endorse a diagnostic product or provide individual medical advice.
