Reproducible corpus
Public licensed speech and consented scripts, using the same audio bytes where supported to isolate recognition and post-processing.
Methodology · version 0.1.0
A preregistration-style protocol for making favourable and unfavourable outcomes equally inspectable. The current release is still a pilot; these rules define the gate for a future benchmark.
Every system is a tuple: public product version, operating-system version, engine or model revision, settings that affect recognition and cleanup, device class, and capture path. Results from different tuples are never silently pooled.
If a product cannot accept the same audio path, it is reported in a separate stratum. A live microphone capture is not presented as directly comparable with a file replay.
| Platform | Baseline class | Current class | Purpose |
|---|---|---|---|
| macOS | Entry Apple Silicon | Current Apple Silicon | Local desktop range |
| Windows | Mainstream x86 laptop | Current performance laptop | CPU and GPU variance |
| Android | Supported mid-range device | Current flagship | Mobile variance |
Public licensed speech and consented scripts, using the same audio bytes where supported to isolate recognition and post-processing.
Calibrated playback through documented rooms and microphones to expose denoise, pre-roll, silence-gate, and tail-flush behaviour.
Newly consented speakers reading preregistered prompts. Private support captures are never promoted into the public corpus.
Language, regional variety, duration, noise, room conditions, names, numbers, technical terms, pauses, and spoken corrections.
Word error rate is calculated as substitutions plus insertions plus deletions, divided by reference words. The study also reports each operation separately because a single fluent transcript can hide a missing phrase.
Aggregate rates are micro-word-weighted: operation counts and reference words are summed before the rate is calculated. Latency publishes sample count, median, p95, and maximum. The completed pilot has mixed thermal state; controlled cold and warm runs and live UI delivery remain future measurements.
Processing time includes recognition, pipeline finalization and local cleanup when enabled. It excludes initial runtime construction, fixture loading, live recording, shortcut dispatch and cursor insertion. All lanes score final evaluator output; cleanup-off lanes can still apply vocabulary and formatting changes.
Latency uses 60 equally weighted decode samples per lane and R-7 linear quantiles. The 20 fixtures contain 491 reference-word positions per pass, counted three times in the 1,473-word denominator. Repeats are not independent utterances. Quality-threshold failures remain included: 3 with Core ML cleanup on, 3 with Core ML cleanup off and 6 with Whisper.
Interpretation corrected 6 September 2026. Original numbers are unchanged. Read the evidence audit, external references, remaining limitations and reproduction instructions.
The protocol records whether a workflow completes offline after setup, observed destinations and bytes, and whether audio or text payloads were observed leaving the device. Packet capture has blind spots, so the defensible statement is “no audio/text payload was observed in this test.” Never claim “nothing ever leaves.”
The current harness mainly tests the transcript pipeline after capture. It does not fully replay fixture audio through the live microphone stack, so capture clipping, denoise artefacts, pre-roll, silence gating, and tail flush require Phase B instrumentation.
Network observation cannot prove the absence of encrypted content or predict future behaviour. A public corpus cannot represent every accent, disability, microphone, room, or spontaneous speaking style. Every result is tied to a specific version and date.