Svara analyses a short, structured voice sample and reports quantitative speech markers that the literature associates with mood, motor-speech, and cognitive-linguistic function. It is intended to serve as a pre-consultation triage aid — a low-burden, repeatable prompt that may help prioritise which patients receive a full standardised assessment sooner. It supplements, and never replaces, clinical interview and validated instruments.
Total core protocol: approximately 90 seconds. Each task is chosen because different functions surface in different speech elicitations; a marker is only ever computed from a task that is valid for it.
| Task | Duration | Function assessed |
|---|---|---|
| Reading a standardised passage | ~25 s | Prosody, speaking rate, rhythm |
| Sustained vowel (/a/) | ~6 s | Phonatory stability, voice quality |
| Picture description (spontaneous) | ~40 s | Connected language, word-finding |
Optional add-on tasks: diadochokinesis (pa-ta-ka) for motor-speech timing;
semantic and phonemic verbal fluency; pitch glide; maximum phonation time; s/z ratio; story recall.
Features are extracted with openSMILE eGeMAPSv02 (the INTERSPEECH standard parameter set), Praat voice-quality algorithms, and transcript-derived measures. Each marker is expressed as a z-score against a healthy reference distribution and reported as a percentile, so an "elevated" reading means N standard deviations from the reference rather than an arbitrary threshold.
Marker directions and relative weights are taken from the published literature rather than chosen by hand:
Outputs are aligned to instruments already in routine use, as a reference frame:
| Indicator | Aligned instrument | Screen threshold |
|---|---|---|
| Mood / depression markers | PHQ-8 (also DASS-21 Depression) | ≥ 10 |
| Tension / anxiety markers | GAD-7 (also DASS-21 Anxiety, PSS-10) | ≥ 10 |
Published voice-screening accuracies above 90% should be treated with scepticism. A 2026 controlled study demonstrated that such figures are largely produced by speaker leakage: when recordings from the same individual appear in both training and test sets, the model learns to recognise the speaker rather than the condition. In that study a model reporting 97.65% depression accuracy also achieved 90.95% speaker-identification accuracy; a comparable model without speaker information scored 62–67%. All figures below are strictly speaker-independent.
Trained on a public sustained-phonation corpus (195 phonations, 32 speakers) and evaluated by leave-one-speaker-out cross-validation, with a 95% confidence interval bootstrapped by resampling speakers (rows within a speaker are correlated, so resampling rows would understate uncertainty):
| Metric | Per clip | Per subject |
|---|---|---|
| AUC | 0.75 | 0.75 |
| 95% CI | 0.56 – 0.92 | 0.51 – 0.93 |
| PR-AUC | 0.87 | — |
| Brier score (calibration) | 0.163 | — |
| At screening threshold | Sens 0.95 / Spec 0.44 | Sens 0.92 / Spec 0.50 |
An AUC in the 0.7–0.8 range reads as poor only if compared against diagnostic tests. Screening instruments in routine use operate in a similar range:
| Instrument | Reported performance | Source |
|---|---|---|
| PHQ-9 (cut-off 10) | Sens 0.78 (0.70–0.84) · Spec 0.87 (0.84–0.90) | Meta-analysis, 36 studies |
| GAD-7 | Sens 0.81 · Spec 0.78 · sROC AUC 0.87 | Meta-analysis, 43 studies |
| Screening mammography | Sens 55–91% · Spec 84–97%; AUC ≈ 0.73 in one analysis | Systematic reviews |
| Svara phonatory model | AUC 0.75 (0.56–0.92) | This system, LOSO |
The system is configured for high sensitivity at the cost of specificity (0.95 / 0.44 per clip), which is the appropriate posture for a triage aid: the objective is to avoid missing someone who warrants assessment, accepting that roughly half of flagged individuals will not have the condition. A positive result should therefore be expected to be a false positive much of the time, and carries no weight on its own — it indicates only that a standard assessment is worthwhile.
| Domain | Realistic speaker-independent performance |
|---|---|
| Depression | AUC ≈ 0.62–0.72 |
| Anxiety | ≈ 0.60 UAR — weak; interpret with caution |
| Motor-speech / phonatory | AUC ≈ 0.80–0.90 (strongest signal) |
| Cognitive-linguistic | AUC ≈ 0.85 in properly-controlled challenge data |
The system administers PHQ-8 and GAD-7 after the voice tasks. These are validated, public-domain instruments and are reported alongside — not merged into — the voice indicators. Where the two disagree, the interface directs the patient according to the questionnaire, since it is the better-established measurement.
Audio is processed on the server and used to compute the reported markers. All analysis runs locally on CPU; no audio is sent to any third-party service. Voice recordings constitute personal health information — the deploying institution is responsible for consent, retention, and applicable data protection obligations before any use with identifiable patients.
Kroenke K et al. The PHQ-8 as a measure of current depression. J Affect Disord, 2009.
Spitzer RL et al. A brief measure for assessing generalized anxiety disorder: the GAD-7.
Arch Intern Med, 2006.
Lovibond SH & Lovibond PF. Manual for the Depression Anxiety Stress Scales (DASS), 1995.
Cummins N et al. A review of depression and suicide risk assessment using speech analysis.
Speech Communication, 2015.
The voice of depression: speech features as biomarkers for major depressive disorder.
BMC Psychiatry, 2024.
Acoustic and linguistic features of impromptu speech and their association with anxiety.
JMIR Mental Health, 2022.
Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based
Depression Detection. arXiv, 2026.
Eyben F et al. The Geneva Minimalistic Acoustic Parameter Set (eGeMAPS).
IEEE Trans Affective Computing, 2016.
Bridge2AI-Voice Consortium. An ethically-sourced, diverse voice dataset linked to health
information. PhysioNet, 2025–26.