Pilot results · interpretation corrected 6 September 2026
Three test categories. One limited pilot study.
The same 20 clean synthetic fixtures, decoded three times per lane on one M1 Max MacBook Pro using Yaps 2.3.2129. Core ML Parakeet and Whisper Base are model configurations inside Yaps, not competitor products.
01 · Aggregate record
Recognition with local cleanup, measured together.
Yaps synthetic fixture-replay pilot. Lower WER, deletion rate, and latency values indicate fewer errors or shorter processing time within this exact run setup; they are not product rankings.
Yaps output lane
Role
Runs
Words, with repeats
Micro WER
Deletion rate
Median processing
p95 processing
Maximum processing
Core ML Parakeet + local cleanup
Cleanup on
60
1,473
3.46%
0.41%
1,124 ms
2,807 ms
2,958 ms
Core ML Parakeet, cleanup off
Cleanup off
60
1,473
3.67%
0.20%
250 ms
722 ms
1,863 ms
Whisper Base, cleanup off
Cleanup off
60
1,473
6.52%
0.41%
844 ms
1,669 ms
2,189 ms
Processing time includes recognition, pipeline finalization and local cleanup when enabled. It excludes initial runtime construction, fixture loading, live recording, shortcut dispatch and cursor insertion. All lanes score final evaluator output; cleanup-off lanes can still apply vocabulary and formatting changes.
Latency uses 60 equally weighted decode samples per lane and R-7 linear quantiles. The 20 fixtures contain 491 reference-word positions per pass, counted three times in the 1,473-word denominator. Repeats are not independent utterances. Quality-threshold failures remain included: 3 with Core ML cleanup on, 3 with Core ML cleanup off and 6 with Whisper.
“Micro” means edit-operation counts and reference words were summed before calculating each rate. It is not the macro average of per-clip percentages.
02 · Limited reading
What these three rows can support.
Cleanup off
Matched cleanup-off configurations
Whisper Base with cleanup off had 2.85 percentage points higher WER, 0.20 points higher deletion rate, a 595 ms higher median, and a 947 ms higher p95 than Core ML Parakeet with cleanup off in these matched pilot runs.
Cleanup-enabled lane
Recognition and cleanup together
The Core ML cleanup-enabled lane includes local semantic cleanup. Its 1,124 ms median and 2,807 ms p95 describe recognition and finalization together, before insertion into another application.
What worked
Publication pipeline
The sanitizer retained counts and latency while excluding transcripts, references, prompts, and application context. Public file integrity is recorded with SHA-256 digests.
What remains
Benchmark evidence
Human speech, multiple accents and languages, live microphone capture, controlled thermal order, network observation, more devices, and third-party systems are still required.
03 · Explicit exclusions
What this pilot does not show.
Comparative performance against any other product.
Performance on human speech, accents, or multilingual speech.
Live microphone, denoise, silence-gate, first-word, or tail-flush behaviour.
Cold-versus-warm performance under a controlled thermal protocol.
Offline or privacy behaviour established through network observation.
A general claim that one raw or cleaned configuration is better.
The corpus, device matrix, order, and thermal procedure were not preregistered. These observations describe these sequential runs, on this corpus, device, software version, and date.