Skip to content

Pilot results · interpretation corrected 6 September 2026

Three test categories. One limited pilot study.

The same 20 clean synthetic fixtures, decoded three times per lane on one M1 Max MacBook Pro using Yaps 2.3.2129. Core ML Parakeet and Whisper Base are model configurations inside Yaps, not competitor products.

01 · Aggregate record

Recognition with local cleanup, measured together.

Yaps synthetic fixture-replay pilot. Lower WER, deletion rate, and latency values indicate fewer errors or shorter processing time within this exact run setup; they are not product rankings.
Yaps output laneRoleRunsWords, with repeatsMicro WERDeletion rateMedian processingp95 processingMaximum processing
Core ML Parakeet + local cleanupCleanup on601,4733.46%0.41%1,124 ms2,807 ms2,958 ms
Core ML Parakeet, cleanup offCleanup off601,4733.67%0.20%250 ms722 ms1,863 ms
Whisper Base, cleanup offCleanup off601,4736.52%0.41%844 ms1,669 ms2,189 ms

Processing time includes recognition, pipeline finalization and local cleanup when enabled. It excludes initial runtime construction, fixture loading, live recording, shortcut dispatch and cursor insertion. All lanes score final evaluator output; cleanup-off lanes can still apply vocabulary and formatting changes.

Latency uses 60 equally weighted decode samples per lane and R-7 linear quantiles. The 20 fixtures contain 491 reference-word positions per pass, counted three times in the 1,473-word denominator. Repeats are not independent utterances. Quality-threshold failures remain included: 3 with Core ML cleanup on, 3 with Core ML cleanup off and 6 with Whisper.

Interpretation corrected 6 September 2026. Original numbers are unchanged. Read the evidence audit, external references, remaining limitations and reproduction instructions.

“Micro” means edit-operation counts and reference words were summed before calculating each rate. It is not the macro average of per-clip percentages.

02 · Limited reading

What these three rows can support.

Cleanup off

Matched cleanup-off configurations

Whisper Base with cleanup off had 2.85 percentage points higher WER, 0.20 points higher deletion rate, a 595 ms higher median, and a 947 ms higher p95 than Core ML Parakeet with cleanup off in these matched pilot runs.

Cleanup-enabled lane

Recognition and cleanup together

The Core ML cleanup-enabled lane includes local semantic cleanup. Its 1,124 ms median and 2,807 ms p95 describe recognition and finalization together, before insertion into another application.

What worked

Publication pipeline

The sanitizer retained counts and latency while excluding transcripts, references, prompts, and application context. Public file integrity is recorded with SHA-256 digests.

What remains

Benchmark evidence

Human speech, multiple accents and languages, live microphone capture, controlled thermal order, network observation, more devices, and third-party systems are still required.

03 · Explicit exclusions

What this pilot does not show.

  • Comparative performance against any other product.
  • Performance on human speech, accents, or multilingual speech.
  • Live microphone, denoise, silence-gate, first-word, or tail-flush behaviour.
  • Cold-versus-warm performance under a controlled thermal protocol.
  • Offline or privacy behaviour established through network observation.
  • A general claim that one raw or cleaned configuration is better.

The corpus, device matrix, order, and thermal procedure were not preregistered. These observations describe these sequential runs, on this corpus, device, software version, and date.

Download the sanitized rows, aggregates, manifests, and checksums.

State of Private Dictation 2026Method version 0.1.0Pilot evidence onlyDesigned and published by Yaps