Open benchmark
Measure the words, the speakers, and the failures.
Daisy will not publish a magic accuracy percentage. The benchmark is an executable protocol that keeps raw output, versions, references, and scoring rules together so every row can be checked.
Current status
| Neutral scorer | Ready | WER, CER, DER, JER, speaker count, capture completeness and RTF. |
|---|---|---|
| Daisy product runner | Ready | Runs the same archive decoder, final Whisper profile and FluidAudio diarizer as the app. |
| Synthetic pipeline smoke | Passed | Validates the harness only. TTS numbers are never product accuracy evidence. |
| Public AMI baseline | Published | One real 17-minute, four-speaker case with raw output, reference, hashes and environment. |
| Daisy vs Humla vs OpenWhispr | Measuring | Public results stay hidden until the shared real/public dataset is complete. |
First public baseline
AMI ES2004a Mix-Headset · 17:29 · English · 4 speakers · automatic speaker count. This is one reproducible diarization case, not a claim about every meeting.
| System | DER ↓ | JER ↓ | Speakers | Capture | RTF ↓ |
|---|---|---|---|---|---|
| Daisy 1.0.7.59 | 15.68% | 20.28% | 4 / 4 | 100% | 0.122× |
| Humla | Not measured | Not measured | Not measured | Not measured | Not measured |
| OpenWhispr | Not measured | Not measured | Not measured | Not measured | Not measured |
Words-only RTTM reference, 0.25 s collar, overlap scored. RTF is the median of three warm runs (0.139× / 0.122× / 0.107×). WER/CER stay blank until the transcript reference is normalized. Competitors stay blank until their uncorrected output exists for this exact WAV.
Open raw evidence and report ↗What gets measured
WER / CER
Word and character errors after one shared multilingual normalization policy.
DER / JER
Speaker confusion, missed speech and false alarms. Labels are permutation-invariant; overlap is scored.
Capture
Whether microphone and system audio survived for the complete expected duration.
Time
Processing RTF plus warm p50/p95 where interactive latency matters.
Recovery
Sleep, lid close, route changes, force-quit and whether a usable archive remains.
Ownership / MCP
What exports without a vendor service and whether current clients receive usable tools.
The shared test matrix
Publication gate
A result appears here only with a reference transcript, RTTM where relevant, raw uncorrected hypothesis, SHA-256, exact build and model settings, environment, and scorer output. Weak cases and unavailable features remain in the table.