Calibration locks the Chao 1–5 pitch scale to your own voice range so your tone scores are accurate. Do it once; redo it if your microphone or recording setup changes.
How the record button works. “Hold to talk” records while you hold the button (or the Space key on a computer) and analyses when you let go. “Tap to start / stop” records on one tap and stops on the next - easier on phones.
Which characters to display. Traditional (繁) is standard for Cantonese; Simplified (简) is used in Mainland China. Pronunciation and tones are identical - only the writing changes.
The colour theme for the app - pick whichever is easiest on your eyes.
What the Enter key plays for the current word. When a word has no native recording, Enter plays your selected TTS voice automatically.
Which reference contours to overlay on your pitch chart. The TTS contour loads on demand from the voice you picked next to “Hear TTS”.
Sign in to save your progress and (with Pro) unlock the full library.
Unlock the full word library, every TTS voice, Story mode and Quiz. Cancel anytime.
Loading plans…
Questions? hello@wing.tools
Read this short Cantonese phrase aloud in your normal speaking voice, with natural intonation. ~3–5 seconds is enough. We use it to lock the Chao 1–5 scale to your actual pitch range so your tones stay consistent across loud / quiet recordings.
Pick a mode, pick a deck, pick a count. Score at the end.
Speak each item. Score is contour distance to the native.
-
Hide for a harder mode - force yourself to recall without hints.
Average contour distance over your last 50 attempts per tone. Lower is better; under 1.0 is good, under 0.7 is great.
| Tone | Attempts | Avg dist | Trend | Diagnosis |
|---|
| Item | Tone | Last dist | SRS next |
|---|
Tap any question to expand. New here? Open the 👋 Welcome panel from the header for a quick intro.
Not exactly wrong. It's a real habit called voice onset rise. Your vocal folds engage from rest in a brief "creaky" register before settling into the target pitch. Native recordings don't show this because the speakers rehearsed, and the silence is pre-clipped.
Three things that fix it, in order:
Match the green. It's a real native speaker reading the word. Blue is the textbook contour. Useful as a reference for the tone's shape, but native speakers rarely hit it exactly. They glide T2 from ~35 to ~45 instead of the textbook 25, and keep T1 high-flat at ~45 instead of 55. Matching the green is the truer goal.
Low tones (T4 low-falling, T6 low-level) are the hardest for non-native learners because they sit below most people's relaxed speaking pitch. Two techniques:
Check the Stats tab. The per-tone diagnostic surfaces patterns like "consistently too flat on T2" once you have 50+ attempts on that tone.
Calibration locks the Chao 1-5 scale to your actual pitch range. Without it, the analyzer guesses your range from each recording, which makes results drift between recordings.
Recalibrate when:
Reading 一二三四五六七八九十 (one through ten) covers most of your natural pitch range across 10 syllables of varied tones. We extract your low end (10th percentile of F0) and high end (90th percentile). Those become the anchors for the Chao 1-5 scale.
Try to read it with natural intonation, not monotone. The wider your range during calibration, the more accurate the scale will be on future recordings.
Every recording in Words, Story, or Quiz mode is counted. The streak requires 3 attempts in a single day to advance. Stats refresh in real-time as you practice.
Your data lives in your browser's localStorage (wingtone_progress_v1). Clearing site data wipes it. Once accounts ship, this will sync to the cloud.
SRS = Spaced Repetition System, like Anki. Items you score poorly on come back to you in 1 day. Items you nail go to a longer interval (3, 7, 14, 30+ days). When items are due, the Stats tab shows a "Due for review" count. Click "Start review" to launch a quiz with only those items.
This is how you turn occasional practice into actual retention.
Most native recordings come from Lingua Libre, a Wikimedia project where native speakers contribute single-word pronunciations under CC BY-SA. Main contributors for Cantonese: Luilui6666, Justinrleung, Jonashtand. Around 7,400 words covered.
Story sentences come from the same project's longer recordings.
For items Lingua Libre doesn't cover - including the Cantonese Names category (surnames and given-name characters) and longer phrases - we use synthetic TTS Cantonese voices. Pick one from the TTS voice dropdown next to Hear TTS: Gentle Lady, Pro Host (F), Pro Host (M), or Playful Man. They're synthetic but considerably more natural on contour-heavy tones than the older voice they replaced.
Lingua Libre's coverage is uneven. Common words usually have a recording. Rarer ones may not. When native audio is missing, the chart shows only the blue target line. You still get scored against the canonical contour, just without a green reference.
The app evaluates your tone with an autocorrelation pitch tracker - the same kind phoneticians use. Steps:
Backend tone evaluation runs in Python; your live recording runs the same algorithm ported to JavaScript in the browser - around 90% tone-classification match against the reference, with identical voicing thresholds, normalization, and resampling.
No. Unvoiced and silent frames are explicitly excluded. They get marked as NaN in the contour and dropped before resampling. The first plotted point is your first voiced sample, not silence.
So if your red line starts low at t=0, that's your actual voice onset (see the voice-onset-rise question above), not a silence artifact.
RMS (root-mean-square) distance between your contour and the reference, in Chao units. Lower is better:
Missing something? The Welcome panel (header, 👋 button) has the visual intro. For pitch analysis internals, see the README in the repo.
Cantonese has 6 tones. Same syllable, different tone, different word. Tap each below to hear it.
Three different words. Only the pitch contour changes.
This app draws your pitch contour against the canonical shape of each tone so you can see exactly where you deviate.
Before you start, we'll calibrate the Chao 1–5 scale to your voice. Takes 5 seconds: read "一二三四五六七八九十" aloud.
Skip if you're just browsing - you can calibrate any time from the header.