Use AI models on your local device to analyze your pronunciation.
Currently, focused on German beginners/kids.
Two language settings:
- ui-lang: language of the UI (auto-detected from browser, changeable)
- study-lang: language being practised (must be chosen explicitly; en-GB or de)
Demo: https://relaxandplay.de/phoneme-party/
Build on:
Github: https://github.com/guettli/pp
Built as static web-site. No server required!
First Steps:
- Application has a set of well known phrases and the desired pronounciation.
- User sees a random phrase, and the corresponding picture.
- User records his/her voice.
- The distance between desired and actual phonemes gets shown.
My goal is to help children in developing countries to learn reading and speaking.
Currently, learning English and German is worked on.
I was inspired by the book Moral Ambition.
Most developing countries in Africa and South Asia use British English in their schools and official exams. By using en-GB, my web app helps children pass their classes and succeed in the system they already live in. British spelling and grammar are still seen as the "gold standard" for professional jobs and higher education in these regions. Choosing this version ensures the web app is a practical tool for a child's academic and future career growth.
Phoneme Party leverages browser-based AI models to provide real-time pronunciation feedback without sending data to external servers.
-
Text to Phonemes (Target)
- User inputs a German phrase or sentence
- A phonetic dictionary or text-to-phoneme model converts the text into the expected phoneme sequence
- For German, this uses IPA (International Phonetic Alphabet) or similar phonetic representation
-
Audio Recording
- User's voice is captured via the browser's Web Audio API
-
Speech to Phonemes (Actual)
- An automatic speech recognition (ASR) model runs locally in the browser
- The model outputs IPA symbols
-
Phoneme Comparison
- The target phoneme sequence is aligned with the actual phoneme sequence
- Distance metrics calculate pronunciation accuracy:
- Phoneme Error Rate (PER): percentage of incorrect phonemes
- Edit distance: insertions, deletions, and substitutions needed
- Per-phoneme scoring: identifies which specific sounds were mispronounced
-
Visual Feedback
- Results displayed in a user-friendly format
- Color-coded phonemes (green = correct, yellow = close, red = incorrect)
- Suggestions for improvement on specific sounds
- Node.js (v20+)
- pnpm
- ffmpeg (for audio processing in tests)
pnpm install# Start development server
pnpm dev
# Build for production
pnpm build
# Type checking
pnpm typecheck
# Linting
pnpm lint
# Run Playwright browser tests
pnpm testSee SCRIPTS.md for all available helper scripts.
Test IPA extraction on all FLAC files in tests/data/:
./run tsx scripts/test-all-flac.ts # test all
./run tsx scripts/test-all-flac.ts --update # update YAML files with new IPA
./run tsx scripts/test-all-flac.ts Wasser # test specific phraseAudio test files are stored in git under tests/data/:
tests/data/
├── de-DE/
│ ├── Apfel/
│ │ ├── Apfel-edge-tts-conrad.flac
│ │ └── Apfel-edge-tts-conrad.flac.yaml
│ └── ...
└── en-GB/
├── Apple/
│ ├── Apple-edge-tts-guy.flac
│ └── Apple-edge-tts-guy.flac.yaml
└── ...
Each phrase can have multiple audio files from different sources. The YAML metadata records the audio provenance (TTS voice, creation date, format).
For a good offline experience small data size is important.
Example: 👗 has size of ~500 Bytes. A jpeg image has usualy a much bigger size.
IPA API: https://www.dwds.de/api/ipa/?q=Haus
List of Words (A1, A2, B1): https://www.dwds.de/d/api#wb-list-goethe
MIT License - See LICENSE file for details.