Open benchmark

Measure the words, the speakers, and the failures.

Daisy will not publish a magic accuracy percentage. The benchmark is an executable protocol that keeps raw output, versions, references, and scoring rules together so every row can be checked.

Current status

Neutral scorerReadyWER, CER, DER, JER, speaker count, capture completeness and RTF.
Daisy product runnerReadyRuns the same archive decoder, final Whisper profile and FluidAudio diarizer as the app.
Synthetic pipeline smokePassedValidates the harness only. TTS numbers are never product accuracy evidence.
Public AMI baselinePublishedOne real 17-minute, four-speaker case with raw output, reference, hashes and environment.
Daisy vs Humla vs OpenWhisprMeasuringPublic results stay hidden until the shared real/public dataset is complete.

First public baseline

AMI ES2004a Mix-Headset · 17:29 · English · 4 speakers · automatic speaker count. This is one reproducible diarization case, not a claim about every meeting.

SystemDER ↓JER ↓SpeakersCaptureRTF ↓
Daisy 1.0.7.5915.68%20.28%4 / 4100%0.122×
HumlaNot measuredNot measuredNot measuredNot measuredNot measured
OpenWhisprNot measuredNot measuredNot measuredNot measuredNot measured

Words-only RTTM reference, 0.25 s collar, overlap scored. RTF is the median of three warm runs (0.139× / 0.122× / 0.107×). WER/CER stay blank until the transcript reference is normalized. Competitors stay blank until their uncorrected output exists for this exact WAV.

Open raw evidence and report

What gets measured

WER / CER

Word and character errors after one shared multilingual normalization policy.

DER / JER

Speaker confusion, missed speech and false alarms. Labels are permutation-invariant; overlap is scored.

Capture

Whether microphone and system audio survived for the complete expected duration.

Time

Processing RTF plus warm p50/p95 where interactive latency matters.

Recovery

Sleep, lid close, route changes, force-quit and whether a usable archive remains.

Ownership / MCP

What exports without a vendor service and whether current clients receive usable tools.

The shared test matrix

Duration
15 / 60 / 180 minutes
Language
English / Russian / mixed RU↔EN
Speakers
2 / 4 / 6, with automatic and hinted count recorded separately
Audio
Clean / noise / overlap and cross-talk
Capture
Microphone + system audio, including failure and recovery paths
Runs
Three warm measured runs; exact Mac, OS, app, model and settings pinned

Publication gate

A result appears here only with a reference transcript, RTTM where relevant, raw uncorrected hypothesis, SHA-256, exact build and model settings, environment, and scorer output. Weak cases and unavailable features remain in the table.

Inspect the harness on GitHubProtocol last updated 17 August 2026.