Research

Evaluating voice AI with real human data

We strive to find where voice AI breaks today using real human data.

Library

All research

We regularly publish and open-source audio data for research purpose. We focus on where voice and audio models break today. If you are interested in collaborating, please email partnership@besimple.ai.

BenchmarkReleased

September 10, 2026

Duplex Cue

A full-duplex voice benchmark for measuring whether an ongoing speaker continues, adapts, or yields when a listener contributes during the turn.

Problem

Can a voice agent incorporate a listener's correction or completion while it keeps speaking, rather than ignoring the cue or yielding the floor?

MethodReleased

August 28, 2026

Inkling Post-Training

A controlled scaling study of targeted human speech data for improving exact structured-value transcription in Inkling.

Problem

How much can targeted human speech data move an already capable speech model on the exact values that voice workflows depend on?

BenchmarkReleased

July 7, 2026

Vocal Affect Bench

A vocal emotion benchmark for evaluating whether emotion detection or omni models can identify expressed affect from raw speech without transcripts or metadata.

Problem

Voice agents often rely on a separate model for emotion detection, but how accurate are they in detection human emotions?

BenchmarkReleased

May 1, 2026

Voice Code Bench

A speech-to-text benchmark for exact structured values in English workplace speech, including emails, command-line flags, file paths, URLs, account identifiers, dates, and measurements.

Problem

Voice agents can transcribe something incorrectly but still get to the right outcome, but what if it's on critical tokens that cannot be recovered later?