Scoring Japanese Pronunciation on the Phone: DTW, Rust, and the Bugs That Made It Real

Most pronunciation apps send your voice to a server. For Kaiwa Coach we decided the scoring had to run on the phone: per-minute cloud speech recognition is expensive, voice data is sensitive, and learners practice on planes. This post is about the engine that makes that possible — and the bugs that taught us where the real risks live.
The pipeline is deliberately boring
There is no neural scorer anywhere. The engine is traditional signal processing, chosen because every stage is testable in isolation:
- Voice activity detection cuts the utterance.
- Mora segmentation aligns the audio to the expected reading from the lesson.
- YIN pitch tracking produces an F0 contour.
- Dynamic time warping (DTW) aligns the learner contour against a native reference contour.
Then per-mora scoring reads out timing, pitch, and intelligibility. When a verdict looks wrong, we can point at the stage that produced it. A neural scorer would have hidden that.
The bug that was really a benchmark
The first version split the VAD span evenly across morae — simple, wrong. Long vowels came out at 20% detection recall against an >80% gate, and geminates at 0%. The vowels were fine; the segmentation was treating speaking rate as vowel length.
The fix was structural: segment from the expected reading, measure real duration distributions (median mora ≈ 130 ms — not the 200 ms we had assumed), and gate on content match before scoring at all. If the recording doesn't match the exercise, the engine refuses to score rather than guessing.
Reference contours must be regridded
DTW step penalties assume near-diagonal paths. If your reference contour is produced at a different frame rate than the engine hop (ours is 16 ms), you silently bias every alignment. Regrid the reference before it ships, and verify with a known-good recording:
pipeline F0 vs engine YIN: 100% voicing agreement
median pitch error: 0.018 cents
That check lives in the pipeline tests so content authors can't ship a contour the engine can't align.
FFI is a build problem first
The engine worked on Android immediately and exported zero of nine symbols on iOS. Two days went into "the Rust engine is broken" before the real cause showed up: LTO was stripping extern "C" symbols out of the iOS static library.
The fix is two linker flags and one rule:
[profile.release]
lto = false # for the FFI build
Plus -u linker flags to force symbol retention. If you ship Rust behind FFI on more than one platform, test symbol export in CI before you write engine logic.
Calibrate every gate against real speech
A pitch power gate of 5.0 looked reasonable on synthetic tones. Real speech sits around 1.7, so the gate rejected every frame. This is the general shape of the mistake: a threshold tuned on generated data, applied to the physical world. Our gates now have recorded-human regression tests.
Signed content, honest limits
Lessons ship as versioned packs signed with Ed25519 and verified by SHA-256, with rollback. 107 lessons came out of a reproducible pipeline over Common Voice ja 26.0: 584,920 clips → 7,048 N5-friendly candidates → 87 pronunciation + 20 shadowing lessons. Contours-only packs cut app assets from 9.6 MB to 988 KB before we optimized audio delivery.
And the honest part: our calibration suite currently runs on synthetic audio. It reports 100/100 reference frames and 10/10 long vowels — with tiny denominators and no geminate coverage. Technically true, practically misleading. Real-speaker validation is now a release gate, not a nice-to-have.
What we'd tell anyone building this
- Traditional DSP is a feature. Explainable stages beat a black box you cannot debug at 2 a.m.
- A benchmark is only as honest as its denominator. Count your samples.
- FFI failures are usually build failures. Check symbols before logic.
- Test gates with real-world inputs, or they will reject the real world.
The engine is 74 Rust tests, 43 Flutter tests, and 62 pipeline tests — and zero bytes of audio leaving the device.
