This is a web app that allows practising irregular English verbs. It is pre-seeded with irregular verbs in various tenses. It then continuously generates exercises via an LLM as requested by the user, and tracks their progress answering these exercises; when the answer is incorrect feedback is also given via the LLM to help the user improve.
When generating exercises, the user can choose the difficulty, which are currently configured as follows:
- Easy - only generate exercises testing the most common 20 irregular verbs in the simple past tense (i.e. 20 verb forms can be used).
- Medium - generate exercises testing the most common 20 irregular verbs in any other tense, and the next most common 80 verbs in the simple past tense (i.e. 140 verb forms).
- Hard - generate exercises testing the next most common 80 verbs in any other tense (i.e. 240 verb forms).
They can also choose the exercise type; currently only one exercise type (multiple choice) is available, but the design accounts for how the others would be implemented.
Install:
To run all needed dev setup steps:
task setupThe app can be run in stub mode just by not setting an API key - this will still fully run and generate exercises, but this is done in a naive deterministic way rather than via an LLM, and the exercises and feedback will be less interesting or useful:
task startOr it can be run against a real LLM by setting a key in the environment.
Currently the app is set up to use OpenAI's gpt-oss-120b, hosted on
https://groq.com. To run against the real LLM:
export GROQ_API_KEY=<your API key>
task startEither way, you can then visit the URL output by the frontend server - likely http://127.0.0.1:5173.
To run the tests:
task testIn particular, the explicitly requested tests are covered by:
- Progress calculation logic -
verbs/tests/api/test_progress.py::test_progress_reflects_seeded_data - An end-to-end API flow using AI stub mode -
verbs/tests/api/test_attempt.py, both of which generate a session and then submit an answer, checking the result and persisted attempt:test_attempt_with_correct_answer_saves_and_returns_correct- a correct answer.test_attempt_with_incorrect_answer_saves_and_returns_feedback- an incorrect answer, additionally checking the stub-generated feedback.
- Session generation itself is covered in more detail by
verbs/tests/api/test_session.py::test_session_generates_deterministic_exercise_via_stub, which asserts the exact deterministic exercise (verb, sentence, choices) stub mode produces.
- Main tools - Python, TypeScript, Django,
django-ninja, React - were largely selected due to my familiarity with them and because they are powerful and in common-use, allowing me to focus on the problem at hand over any extra setup or plumbing. - Sqlite as the database - usually I would use Postgres for most production apps, but since this is just an example app Sqlite is a capable database and this simplifies the quick start setup for you above.
llmlibrary for interacting with LLMs, rather than a provider-specific library - this gives a largely generic interface to different LLM APIs and local models, so the LLM used can easily be swapped out; this was useful during development to switch from Gemini 2.5 Flash togpt-oss-120bafter quickly hitting Gemini's low free-tier rate limit.gpt-oss-120bwas selected as it's free to run with fairly high rate limits I haven't yet hit, gives reasonable output, and can generate structured outputs.
This was designed to support all current and likely future features. The greyed
out exercise types have not yet been implemented, but have been planned for -
all exercise types inherit from a central Exercise model via multi-table
inheritance, and it is this which other models which work for any exercise,
like ExerciseAttempt, will relate to; this way these can keep being treated
generically without needing variants for different exercise types, and with a
normal, database-enforced foreign key rather than a Django-managed polymorphic
foreign key.
-
Only the
multiple_choiceexercise type was implemented - chosen as this seemed one of the more interesting exercise types to me to integrate with an LLM and validate the response for; implementing the other types should slot into the existing data model as above, and with their own corresponding exercise and feedback generation classes similar to the current ones. -
As suggested, no user authentication or per-user tracking was added - since this was not really the focus of the exercise, and would be easy to add if needed via the built-in Django behaviour and adding a relation between the user model and
ExerciseAttempt, and then also filtering the progress to only the current user. -
Difficulty modeled as an enum (easy/medium/hard) rather than the suggested ints between 1-3 - I generally prefer using text enums where possible over int codes as this makes the meaning explicit to anything reading the database, not requiring a particular app to interpret the meaning of each code.
-
I added full backend test coverage (except for a single line in
verbs/ai/llm_model.py, where we actually call the model) as this was useful to ensure correct behaviour during refactoring; I did not add frontend testing as this seemed out-of-scope from the brief and it is also a simple UI that is easy to exhaustively test directly (but this would of course be useful in a larger, production app).
-
All verbs and their tenses to be tested are pre-seeded into the database, as described above - this is factual, deterministic data and we want to have a broad range of verbs, so fits much better as static data like this rather than generating from the LLM (where it could be incorrect/some things missing/get a less broad range of verbs etc.).
-
Verb rarity is also seeded based on real verb frequency, and then along with the different tenses is used to determine the difficulty - what should count as what difficulty was not specified, but this seemed a reasonable interpretation that resulted in a few common verb forms being tested for "easy", through to many uncommon verb forms being tested for "hard". For a real app, it might be better to have a user choose the specific rarity or type of the verbs and the tense they want to practice, rather than a less granular and more opaque trinary difficulty field.
-
We use the LLM to generate the exercises and feedback on incorrect answers, but explicitly do not trust the responses and pass it through several checks that the responses look valid. All LLM interactions go through a subclass of the
LLMGeneratorclass, which calls the LLM with a particular schema (some LLM APIs actually enforce this, others treat it more as a suggestion). The response is then checked that it is valid JSON, that it conforms to the schema, and that it conforms to other validations specific to the particular use case. If any validations fail we ask the LLM to generate a response again (up to 3 times total), with the history available and the validation error passed in. -
/api/session's response never identifies the correct answer - this is only checked on later submission to/api/attempt, preventing someone "cheating" by inspecting the response directly. -
Stub mode is fully deterministic as requested, but still generates a range of exercises rather than repeating the same one: it tracks the last verb ID used and advances through seeded verbs in a fixed order each time a new exercise is needed, wrapping around once exhausted.
