WingTones

-
jyutping
meaning
TTS
Pick a word, then hold the button or Space to talk.
Tone Evaluation
Detected: -
Target: -
Voiced: -
Duration: -
Your range: -
Match to ideal tone: -
Match to native: -
-
-
jyutping
meaning
Click a sentence, then hold Space to repeat it.
Voiced: -
Your duration: -
Native duration: -
Your range: -
Distance to native: -

    Test mode

    Pick a mode, pick a deck, pick a count. Score at the end.

    Speak each item. Score is contour distance to the native.

    -

    Hide for a harder mode - force yourself to recall without hints.

    Your progress

    0
    Day streak 🔥
    Practice today to start.
    0
    Attempts today
    3 daily attempts maintain a streak.
    0
    Total attempts
    -
    0
    Due for review

    Per-tone diagnostic

    Average contour distance over your last 50 attempts per tone. Lower is better; under 1.0 is good, under 0.7 is great.

    ToneAttemptsAvg distTrendDiagnosis

    Recently practiced

    ItemToneLast distSRS next

    Help & tips

    Tap any question to expand. New here? Open the 👋 Welcome panel from the header for a quick intro.

    Recording technique

    My red line always starts low and rises. Am I doing it wrong?

    Not exactly wrong. It's a real habit called voice onset rise. Your vocal folds engage from rest in a brief "creaky" register before settling into the target pitch. Native recordings don't show this because the speakers rehearsed, and the silence is pre-clipped.

    Three things that fix it, in order:

    • Attack the pitch. Don't ramp up. Start the syllable already at the target pitch, like a singer hitting a note.
    • Speak louder. A weak attack stretches the warm-up. A confident attack collapses it under 30ms and edge-trim removes it.
    • Settle, then attack. Wait about half a second after you press record before you speak - the mic's auto-gain needs a moment to stabilise and it keeps the button tap out of your clip. (WingTones also auto-trims the first ~100 ms.) The pause is just for the mic; when you do speak, still start the syllable immediately at the target pitch rather than sliding up to it.
    Should I match the BLUE line or the GREEN line?

    Match the green. It's a real native speaker reading the word. Blue is the textbook contour. Useful as a reference for the tone's shape, but native speakers rarely hit it exactly. They glide T2 from ~35 to ~45 instead of the textbook 25, and keep T1 high-flat at ~45 instead of 55. Matching the green is the truer goal.

    I'm struggling with a specific tone (T4 or T6 low ones)

    Low tones (T4 low-falling, T6 low-level) are the hardest for non-native learners because they sit below most people's relaxed speaking pitch. Two techniques:

    • Drop your jaw and chest, not your throat. Reaching for low tones with throat constriction creates creaky voice that the analyzer reads as unstable.
    • Slow it down. Use the 0.5× playback button on the native, hum along, then speak it at full speed.

    Check the Stats tab. The per-tone diagnostic surfaces patterns like "consistently too flat on T2" once you have 50+ attempts on that tone.

    Calibration

    What does calibration do, and when should I redo it?

    Calibration locks the Chao 1-5 scale to your actual pitch range. Without it, the analyzer guesses your range from each recording, which makes results drift between recordings.

    Recalibrate when:

    • You first set up the app (mandatory; the welcome flow handles this).
    • You haven't used the app for more than a day (vocal cords change daily).
    • You've been talking continuously for over 30 minutes (warmed-up voice differs from cold).
    • You notice contours sitting consistently too high or too low on the chart.
    What does the calibration phrase do?

    Reading 一二三四五六七八九十 (one through ten) covers most of your natural pitch range across 10 syllables of varied tones. We extract your low end (10th percentile of F0) and high end (90th percentile). Those become the anchors for the Chao 1-5 scale.

    Try to read it with natural intonation, not monotone. The wider your range during calibration, the more accurate the scale will be on future recordings.

    The tabs

    Words vs Story vs Quiz vs Stats. When do I use which?
    • Words. Practice single characters and short compounds against the canonical tone contour. Best for drilling specific tones.
    • Story. Practice full sentences with native audio. Best for connected speech and tone transitions.
    • Quiz. Score yourself on 10/20/30 random items, or pick a custom count. Two modes: pronunciation (speak each, scored by contour distance) and listening (hear audio, pick the tone).
    • Stats. Streak, per-tone diagnostic, SRS review queue (items you struggled with come back later).
    Why does the Stats tab show "0"? When do my attempts count?

    Every recording in Words, Story, or Quiz mode is counted. The streak requires 3 attempts in a single day to advance. Stats refresh in real-time as you practice.

    Your data lives in your browser's localStorage (wingtone_progress_v1). Clearing site data wipes it. Once accounts ship, this will sync to the cloud.

    What's the SRS review queue?

    SRS = Spaced Repetition System, like Anki. Items you score poorly on come back to you in 1 day. Items you nail go to a longer interval (3, 7, 14, 30+ days). When items are due, the Stats tab shows a "Due for review" count. Click "Start review" to launch a quiz with only those items.

    This is how you turn occasional practice into actual retention.

    Audio sources

    Where does the native audio come from?

    Most native recordings come from Lingua Libre, a Wikimedia project where native speakers contribute single-word pronunciations under CC BY-SA. Main contributors for Cantonese: Luilui6666, Justinrleung, Jonashtand. Around 7,400 words covered.

    Story sentences come from the same project's longer recordings.

    What about words without a native recording?

    For items Lingua Libre doesn't cover - including the Cantonese Names category (surnames and given-name characters) and longer phrases - we use synthetic TTS Cantonese voices. Pick one from the TTS voice dropdown next to Hear TTS: Gentle Lady, Pro Host (F), Pro Host (M), or Playful Man. They're synthetic but considerably more natural on contour-heavy tones than the older voice they replaced.

    Some words have no audio. Why?

    Lingua Libre's coverage is uneven. Common words usually have a recording. Rarer ones may not. When native audio is missing, the chart shows only the blue target line. You still get scored against the canonical contour, just without a green reference.

    Pipeline details (for the curious)

    How does the pitch analysis work?

    The app evaluates your tone with an autocorrelation pitch tracker - the same kind phoneticians use. Steps:

    1. Your raw audio is run through 75–500 Hz pitch detection per 10ms frame.
    2. Edge-trim silence at start and end. Interior pauses are kept.
    3. Map Hz to the Chao 1–5 scale using your calibrated voice range.
    4. Resample to 30 evenly-spaced points across the voiced span.
    5. Compare against canonical or native contours via RMS distance.

    Backend tone evaluation runs in Python; your live recording runs the same algorithm ported to JavaScript in the browser - around 90% tone-classification match against the reference, with identical voicing thresholds, normalization, and resampling.

    Is silence treated as the lowest tone?

    No. Unvoiced and silent frames are explicitly excluded. They get marked as NaN in the contour and dropped before resampling. The first plotted point is your first voiced sample, not silence.

    So if your red line starts low at t=0, that's your actual voice onset (see the voice-onset-rise question above), not a silence artifact.

    What does the "distance" number mean?

    RMS (root-mean-square) distance between your contour and the reference, in Chao units. Lower is better:

    • ≤ 0.7. Great match. Your contour tracks the reference closely.
    • 0.7 to 1.0. Good. Direction is right. Tighten the peaks and valleys.
    • 1.0 to 1.5. Off by a meaningful margin. Check whether you're catching the rise or fall.
    • Above 1.5. The shape is wrong. Possibly a different tone. Listen again.

    Missing something? The Welcome panel (header, 👋 button) has the visual intro. For pitch analysis internals, see the README in the repo.