Skip to content
← all writing

Kartochki: MFCC + DTW pronunciation scoring for Russian

  • kartochki
  • speech
  • mfcc
  • dtw
  • pronunciation
Kartochki: MFCC + DTW pronunciation scoring for Russian

Kartochki is a mobile web flashcard app. Russian cards now get a real pronunciation score instead of a yes/no from the browser speech API.

Kartochki scoring ужин (dinner) at 90 PASS
Live pass on ужин - score 90, threshold 50.

What we ship

Tap the mic on a Russian word. The app records a short clip, converts it to features, and compares it to stored references for that exact word. You get a 0-100 score and a pass/fail badge. Pass is 50 or higher.

The "training" data

There is no neural model trained on graded attempts. The data is a closed set of reference templates.

MFCC (Mel-Frequency Cepstral Coefficients) turns a short audio clip into a sequence of small vectors. Each ~10 ms frame becomes 12 spectrum-shape numbers plus first-order deltas (24-D total). Those numbers track how the voice spectrum is shaped - vowel and consonant color - more than raw volume or pitch. A template is that sequence for a known good take of one word, stored ahead of time.

DTW (Dynamic Time Warping) is the compare step. Think of Levenshtein distance on strings: count inserts, deletes, and substitutions to turn one sequence into another. DTW does the same job for time series. Instead of characters, the units are MFCC frames. Instead of edit ops, the path stretches and squeezes time so a slow speaker can still line up with a faster reference. Frame cost is Euclidean distance in MFCC space, not "wrong letter."

Levenshtein DTW on MFCC
Objects discrete symbols continuous feature vectors over time
Match unit character / token ~10 ms audio frame
Flexibility insert / delete / substitute time warping (speaking rate)
Output edit count path cost → we map to 0-100

For each of 545 Russian deck words we synthesized two native-style takes (dmitry, svetlana). Offline, each clip is turned into an MFCC matrix with the same recipe the browser uses at score time. Those matrices live in mfcc_cache.json (~8 MB) and load once per language - about 1,090 templates total, not labeled human grades.

How scoring works

  1. Capture mono audio and resample to 16 kHz.
  2. Pre-emphasize, energy-trim silence, frame at 25 ms with a 10 ms hop.
  3. 40-mel log filterbank → DCT → MFCC coefficients 1-12, plus first-order deltas (24-D).
  4. Mean-center the utterance (no variance norm - CMVN hurts short words).
  5. Banded DTW (Euclidean frame cost) against each reference for that word; keep the best distance.
  6. Map distance through a logistic curve to 0-100.
score = 100 / (1 + exp(0.10 * (distance - 58)))

Same word across voices lands low distance / high score. A different word lands high distance / low score. The pass line is calibrated on that split, then loosened slightly for real mic noise (center=58, scale=0.10, pass ≥ 50).

Capture notes

Mic reliability mattered as much as the scorer. Chrome will keep a MediaStream marked live after a few MediaRecorder cycles while the samples go silent. Fix: open a fresh stream on every tap (permission stays granted), keep one AudioContext for the page, and auto-retry once on empty takes. Start getUserMedia on the click itself - not after the cache fetch - so the user gesture is still valid.

Try it

Russian deck: kartochki.app/russian. Hard-refresh, allow the mic, speak during Listening.

Termagotchi
_

Ryan Underdown

Autodidact. Rarely listens to advice.

Follow on X @catamarammed or GitHub @underdown