Benchmark dashboard

Vocal Affect Bench

A vocal emotion benchmark for evaluating whether emotion detection or omni models can identify expressed affect from raw speech without transcripts or metadata.

Top accuracy44.3%

Gemini 3.5 Flash

Emotion classes7

angry, disgusted, fearful, happy, neutral, sad, surprised

Models evaluated7

emotion detection and omni models

Clips280

40 per emotion class

Avg. accuracy34.7%

across all 7 models

Random baseline14.3%

uniform 7-class chance

Detecting the correct vocal emotion is critical to having the correct response. Vocal Affect Bench evaluates emotion detection accuracy in the latest SOTA models.

words spoken"that's fine"conceding, joking, or frustrated — the transcript is identical
what ASR capturestext onlytone, pace, pauses, and intensity are lost after transcription
what voice agents needexpressed affectwhich emotion was delivered, not inferred from the words

Leaderboard

Model results

Ranked by seven-way accuracy. Best model reaches 44.3% — roughly 3× the 14.3% random baseline, but still misses more than half of all clips. Inkling is a post-publication baseline; the paper remains a snapshot of the original release.

#1Gemini 3.5 Flash
44.3%
#2Hume Prosody
38.0%
#3Qwen3.5 Omni+
37.9%
#4Voxtral Small
34.3%
#5Thinking Machines Inkling
32.1%
#6Inworld Voice
28.6%
#7OpenAI Realtime
27.9%
RankModelAcc.AngryDisgustedFearfulHappyNeutralSadSurprised
#1Gemini 3.5 Flash44.3%57.5%27.5%32.5%45.0%75.0%70.0%2.5%
#2Hume Prosody38.0%50.0%5.1%0.0%62.5%85.0%22.5%40.0%
#3Qwen3.5 Omni+37.9%35.0%25.0%17.5%47.5%65.0%65.0%10.0%
#4Voxtral Small34.3%17.5%72.5%17.5%35.0%47.5%27.5%22.5%
#5Thinking Machines Inkling32.1%32.5%2.5%12.5%35.0%85.0%55.0%2.5%
#6Inworld Voice28.6%37.5%0.0%12.5%30.0%97.5%17.5%5.0%
#7OpenAI Realtime27.9%32.5%2.5%10.0%25.0%87.5%27.5%10.0%

Class difficulty

Per-class recall and precision

Most models can detect neutral emotion correctly at 77.5% recall. The recall for detecting fearful and surprised is only 14.6% and 13.2%. A model that looks useful on neutral-heavy traffic can still fail on the states that matter most for escalation.

Recall — what fraction of true clips were correctly identified

neutral
77.5%
sad
40.7%
happy
40.0%
angry
37.5%
disgusted
19.3%
fearful
14.6%
surprised
13.2%

Precision — when predicted, how often was it correct

neutral
23.4%
sad
38.3%
happy
41.8%
angry
61.8%
disgusted
40.3%
fearful
77.4%
surprised
33.9%

Neutral bias

Top confusions

710 of 1,279 incorrect predictions collapse to neutral. The dominant error is over-predicting neutral rather than mis-labeling between non-neutral emotions.

fearfulneutral
138
disgustedneutral
133
sadneutral
124
surprisedneutral
110
happyneutral
105
angryneutral
100
surprisedhappy
62
fearfulsad
54

Valence analysis

Coarse positive / neutral / negative accuracy

Instead of predicting the distinct seven classes of emotions, it is easier for models to detect positive, neutral and negative. The accuracy now sits at 48.9% on average, but still not great.

#1Gemini 3.5 Flash
66.2%
#2Voxtral Small
55.0%
#3Qwen3.5 Omni+
52.5%
#4Hume Prosody
47.7%
#5Thinking Machines Inkling
47.1%
#6Inworld Voice
38.3%
#7OpenAI Realtime
35.4%
ModelNeg. out.Pos. out.Neutral out.Ambig. out.Output skew
Gemini 3.5 Flash51.4%10.7%36.8%1.1%Negative
Hume Prosody24.4%27.6%36.2%11.8%Neutral
Qwen3.5 Omni+42.9%12.9%41.8%2.5%Negative
Voxtral Small50.7%16.4%20.0%12.9%Negative
Thinking Machines Inkling30.0%7.5%59.3%3.2%Neutral
Inworld Voice16.8%14.6%67.5%1.1%Neutral
OpenAI Realtime17.9%6.1%69.6%6.4%Neutral

Dataset composition

Clip and class distribution

VocalAffectBench is a test-only evaluation dataset, not a training corpus. 280 clips across 7 balanced classes, sourced from acted and naturalistic recordings.

EmotionClipsMinutesMean (sec)
angry4017.125.7
disgusted4019.028.5
fearful4017.526.3
happy4022.834.2
neutral4015.623.3
sad4026.038.9
surprised4021.031.6

Scoring method

Evaluation protocol

1

Audio input

The model receives raw audio only — no transcript, label hint, or context metadata.

2

Emotion label output

The model must return one of seven canonical labels: angry, disgusted, fearful, happy, neutral, sad, or surprised.

3

Exact-match scoring

Seven-way accuracy, per-class precision/recall, and valence-bucket accuracy are computed against gold labels.