Skip to content

Methodology · version 0.1.0

The rules come before the results.

A preregistration-style protocol for making favourable and unfavourable outcomes equally inspectable. The current release is still a pilot; these rules define the gate for a future benchmark.

01 · System definition

A product name is not a test condition.

Every system is a tuple: public product version, operating-system version, engine or model revision, settings that affect recognition and cleanup, device class, and capture path. Results from different tuples are never silently pooled.

If a product cannot accept the same audio path, it is reported in a separate stratum. A live microphone capture is not presented as directly comparable with a file replay.

Proposed device matrix for the first complete benchmark release.
PlatformBaseline classCurrent classPurpose
macOSEntry Apple SiliconCurrent Apple SiliconLocal desktop range
WindowsMainstream x86 laptopCurrent performance laptopCPU and GPU variance
AndroidSupported mid-range deviceCurrent flagshipMobile variance
02 · Three phases

Lab repeatability, then real-device reality.

Phase A

Reproducible corpus

Public licensed speech and consented scripts, using the same audio bytes where supported to isolate recognition and post-processing.

Phase B

Real-device capture

Calibrated playback through documented rooms and microphones to expose denoise, pre-roll, silence-gate, and tail-flush behaviour.

Phase C

Natural dictation

Newly consented speakers reading preregistered prompts. Private support captures are never promoted into the public corpus.

Across all phases

Visible strata

Language, regional variety, duration, noise, room conditions, names, numbers, technical terms, pauses, and spoken corrections.

03 · Primary metrics

WER is the start, not the whole story.

Word error rate is calculated as substitutions plus insertions plus deletions, divided by reference words. The study also reports each operation separately because a single fluent transcript can hide a missing phrase.

Aggregate rates are micro-word-weighted: operation counts and reference words are summed before the rate is calculated. Latency publishes sample count, median, p95, and maximum. The completed pilot has mixed thermal state; controlled cold and warm runs and live UI delivery remain future measurements.

Processing time includes recognition, pipeline finalization and local cleanup when enabled. It excludes initial runtime construction, fixture loading, live recording, shortcut dispatch and cursor insertion. All lanes score final evaluator output; cleanup-off lanes can still apply vocabulary and formatting changes.

Latency uses 60 equally weighted decode samples per lane and R-7 linear quantiles. The 20 fixtures contain 491 reference-word positions per pass, counted three times in the 1,473-word denominator. Repeats are not independent utterances. Quality-threshold failures remain included: 3 with Core ML cleanup on, 3 with Core ML cleanup off and 6 with Whisper.

Interpretation corrected 6 September 2026. Original numbers are unchanged. Read the evidence audit, external references, remaining limitations and reproduction instructions.

Network observation

The protocol records whether a workflow completes offline after setup, observed destinations and bytes, and whether audio or text payloads were observed leaving the device. Packet capture has blind spots, so the defensible statement is “no audio/text payload was observed in this test.” Never claim “nothing ever leaves.”

04 · Publication gates

No leaderboard until every gate passes.

Method frozen firstVersion and analysis commit recorded before benchmark collection.
Corpus pinnedImmutable revisions and SHA-256 digests for every source and emitted file.
References reviewedTwo-person reference review and one normalization policy for every system.
Inputs comparableSame bytes or an explicit, separately reported capture stratum.
Provenance completeProduct, model, device, settings, power state, and capture path recorded.
Failures retainedTimeouts and missing transcripts remain in the denominator.
Operations reconcileSubstitutions, insertions, and deletions add up to edit distance.
Privacy reviewSanitizer passes and a human inspects the public artifact boundary.
Conflict disclosedYaps authorship, exclusions, deviations, and limitations remain prominent.
Data ships with claimsSanitized rows, aggregates, manifests, and correction channel publish together.
05 · Known limitations

What this protocol still cannot see.

The current harness mainly tests the transcript pipeline after capture. It does not fully replay fixture audio through the live microphone stack, so capture clipping, denoise artefacts, pre-roll, silence gating, and tail flush require Phase B instrumentation.

Network observation cannot prove the absence of encrypted content or predict future behaviour. A public corpus cannot represent every accent, disability, microphone, room, or spontaneous speaking style. Every result is tied to a specific version and date.

State of Private Dictation 2026Method version 0.1.0Pilot evidence onlyDesigned and published by Yaps