Svara·Voice Screening
Clinical Reference Note for Mental Health Practitioners · v1.0 · health.d2c.in
Regulatory and clinical status Svara is not a medical device. It has no regulatory clearance (FDA, CE, CDSCO) and has not undergone clinical validation on a patient population. It is a research and triage-support instrument that produces screening indicators only. It must not be used to make, confirm, or exclude a diagnosis, nor to direct treatment.

1. Purpose

Svara analyses a short, structured voice sample and reports quantitative speech markers that the literature associates with mood, motor-speech, and cognitive-linguistic function. It is intended to serve as a pre-consultation triage aid — a low-burden, repeatable prompt that may help prioritise which patients receive a full standardised assessment sooner. It supplements, and never replaces, clinical interview and validated instruments.

2. Assessment protocol

Total core protocol: approximately 90 seconds. Each task is chosen because different functions surface in different speech elicitations; a marker is only ever computed from a task that is valid for it.

TaskDurationFunction assessed
Reading a standardised passage~25 sProsody, speaking rate, rhythm
Sustained vowel (/a/)~6 sPhonatory stability, voice quality
Picture description (spontaneous)~40 sConnected language, word-finding

Optional add-on tasks: diadochokinesis (pa-ta-ka) for motor-speech timing; semantic and phonemic verbal fluency; pitch glide; maximum phonation time; s/z ratio; story recall.

3. What is measured

Acoustic / prosodic

  • F0 mean and variability (monotonicity)
  • Speaking rate, pause frequency and duration
  • Loudness mean and dynamics
  • Jitter, shimmer, harmonics-to-noise ratio

Linguistic / timing

  • Words per minute, lexical diversity
  • Filler and disfluency rate
  • Verbal fluency output counts
  • Syllable rate and rhythm regularity (DDK)

Features are extracted with openSMILE eGeMAPSv02 (the INTERSPEECH standard parameter set), Praat voice-quality algorithms, and transcript-derived measures. Each marker is expressed as a z-score against a healthy reference distribution and reported as a percentile, so an "elevated" reading means N standard deviations from the reference rather than an arbitrary threshold.

4. Evidence base

Marker directions and relative weights are taken from the published literature rather than chosen by hand:

5. Reference frame and reporting

Outputs are aligned to instruments already in routine use, as a reference frame:

IndicatorAligned instrumentScreen threshold
Mood / depression markersPHQ-8 (also DASS-21 Depression)≥ 10
Tension / anxiety markersGAD-7 (also DASS-21 Anxiety, PSS-10)≥ 10
Important distinction In its present configuration Svara does not output a measured PHQ-8 or GAD-7 score. It reports a speech-derived indicator aligned to that instrument's frame of reference. A calibrated numeric estimate becomes available only after a regression head is trained and validated on a labelled patient cohort at the deploying site.

6. Expected performance — read in context

6.1 Why published voice-AI figures mislead

Published voice-screening accuracies above 90% should be treated with scepticism. A 2026 controlled study demonstrated that such figures are largely produced by speaker leakage: when recordings from the same individual appear in both training and test sets, the model learns to recognise the speaker rather than the condition. In that study a model reporting 97.65% depression accuracy also achieved 90.95% speaker-identification accuracy; a comparable model without speaker information scored 62–67%. All figures below are strictly speaker-independent.

6.2 Measured performance of the phonatory model

Trained on a public sustained-phonation corpus (195 phonations, 32 speakers) and evaluated by leave-one-speaker-out cross-validation, with a 95% confidence interval bootstrapped by resampling speakers (rows within a speaker are correlated, so resampling rows would understate uncertainty):

MetricPer clipPer subject
AUC0.750.75
95% CI0.56 – 0.920.51 – 0.93
PR-AUC0.87—
Brier score (calibration)0.163—
At screening thresholdSens 0.95 / Spec 0.44Sens 0.92 / Spec 0.50
Interpret the confidence interval, not just the point estimate The wide interval (0.56–0.92) is driven by sample size — 32 speakers — not by instability in the model. It means this corpus is too small to establish performance precisely. A larger, local validation cohort would narrow it substantially. Reporting the interval rather than a single flattering number is deliberate.

6.3 What these numbers mean next to established instruments

An AUC in the 0.7–0.8 range reads as poor only if compared against diagnostic tests. Screening instruments in routine use operate in a similar range:

InstrumentReported performanceSource
PHQ-9 (cut-off 10)Sens 0.78 (0.70–0.84) · Spec 0.87 (0.84–0.90) Meta-analysis, 36 studies
GAD-7Sens 0.81 · Spec 0.78 · sROC AUC 0.87Meta-analysis, 43 studies
Screening mammographySens 55–91% · Spec 84–97%; AUC ≈ 0.73 in one analysis Systematic reviews
Svara phonatory modelAUC 0.75 (0.56–0.92)This system, LOSO
This comparison is about magnitude, not equivalence PHQ-9, GAD-7 and mammography are extensively validated instruments with decades of evidence and, where applicable, regulatory approval. Svara has none of that. The table is included solely to show that screening-level discrimination is normally in this range — not to imply Svara is interchangeable with a validated instrument. Real-world PHQ-9 performance itself varies widely across settings (sensitivity 0.37–0.98), which is a reminder that context governs the usefulness of any screen.

6.4 Operating point

The system is configured for high sensitivity at the cost of specificity (0.95 / 0.44 per clip), which is the appropriate posture for a triage aid: the objective is to avoid missing someone who warrants assessment, accepting that roughly half of flagged individuals will not have the condition. A positive result should therefore be expected to be a false positive much of the time, and carries no weight on its own — it indicates only that a standard assessment is worthwhile.

6.5 Expected performance in other domains

DomainRealistic speaker-independent performance
DepressionAUC ≈ 0.62–0.72
Anxiety≈ 0.60 UAR — weak; interpret with caution
Motor-speech / phonatoryAUC ≈ 0.80–0.90 (strongest signal)
Cognitive-linguisticAUC ≈ 0.85 in properly-controlled challenge data

6.6 Self-report module (PHQ-8 and GAD-7)

The system administers PHQ-8 and GAD-7 after the voice tasks. These are validated, public-domain instruments and are reported alongside — not merged into — the voice indicators. Where the two disagree, the interface directs the patient according to the questionnaire, since it is the better-established measurement.

PHQ-8 is used rather than PHQ-9, deliberately PHQ-9 item 9 assesses thoughts of self-harm. PHQ-8 omits it and was designed for settings where immediate clinical follow-up on a positive response cannot be guaranteed — which describes unsupervised self-administration. Consequently this tool does not screen for suicidal ideation. Any severe result (PHQ-8 ≥ 20, GAD-7 ≥ 15) displays crisis contact information, but risk assessment remains entirely a clinical responsibility and must not be assumed to have been performed here.

7. Limitations

8. Suggested use in practice

  1. Record in a quiet room; the system reports signal quality and rejects unusable audio.
  2. Read the summary as a prompt, not a result.
  3. Where an indicator is elevated, administer the aligned standardised instrument (PHQ-9 / GAD-7) and proceed with normal clinical assessment.
  4. Where all indicators are low, this is not evidence of absence — a patient reporting distress should be assessed regardless of what the speech markers show.
  5. Consider repeat recordings over time; trajectory is more meaningful than a single reading.
Clinical override Clinical judgement always takes precedence. No screening output should delay, override, or substitute for assessment of a patient who reports distress, hopelessness, or risk of self-harm.

9. Data handling

Audio is processed on the server and used to compute the reported markers. All analysis runs locally on CPU; no audio is sent to any third-party service. Voice recordings constitute personal health information — the deploying institution is responsible for consent, retention, and applicable data protection obligations before any use with identifiable patients.

10. Selected references

Kroenke K et al. The PHQ-8 as a measure of current depression. J Affect Disord, 2009.
Spitzer RL et al. A brief measure for assessing generalized anxiety disorder: the GAD-7. Arch Intern Med, 2006.
Lovibond SH & Lovibond PF. Manual for the Depression Anxiety Stress Scales (DASS), 1995.
Cummins N et al. A review of depression and suicide risk assessment using speech analysis. Speech Communication, 2015.
The voice of depression: speech features as biomarkers for major depressive disorder. BMC Psychiatry, 2024.
Acoustic and linguistic features of impromptu speech and their association with anxiety. JMIR Mental Health, 2022.
Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based Depression Detection. arXiv, 2026.
Eyben F et al. The Geneva Minimalistic Acoustic Parameter Set (eGeMAPS). IEEE Trans Affective Computing, 2016.
Bridge2AI-Voice Consortium. An ethically-sourced, diverse voice dataset linked to health information. PhysioNet, 2025–26.