An Urdu voice agent that checks money, medicine and documents for blind and low-literacy Pakistanis.
Baseerat means insight — the kind that doesn't come from eyes.
npm install
npm run devOpen http://localhost:3000. It works immediately, with no API key.
Existing accessibility apps — Be My Eyes, Seeing AI, Envision — are genuinely excellent, and they describe the world in English. A blind person in Lahore who doesn't read English gets nothing from them.
They're also read-only and low-stakes. They'll tell you there's a paper on the table. They won't tell you it's a loan agreement at 20% monthly interest that you're about to put your thumbprint on.
Baseerat is built for the moments where not knowing costs you money, health, or legal standing:
| 💵 Money | Which note is this? Is this the right change? |
| 💊 Medicine | Is this the right strip? The right dose? Has it expired? |
| 📄 Documents | What am I about to agree to? |
It speaks Urdu. Not localised menus — the agent thinks and answers in spoken Pakistani Urdu, in Urdu script, with proper ur-PK neural speech.
There is no shutter button. A blind user cannot know when the picture is good, so we never ask them to take one. The camera simply keeps looking, and the agent decides for itself when it has a frame worth answering from — talking the user into position along the way: "تھوڑا قریب لائیں… ہاتھ روکیں… مل گیا۔" Every competitor says "point your phone at the document," which quietly assumes you can see whether the document is in frame.
It checks against your prescription. Save what the doctor prescribed once, and every strip you hold up afterwards gets checked against it. Holding up 500mg when you were prescribed 250mg triggers a spoken warning, not a description.
It warns before you commit. Document mode doesn't just read a contract aloud — it tells you what you'd be agreeing to, and flags the parts that bind you.
Next.js 16 · React 19 · TypeScript · Tailwind v4 · OpenRouter vision · Edge TTS
src/
app/
page.tsx the whole agent — one screen, one big control
setup/ prescription, voice settings, history
api/see/ one frame in, a structured Urdu reading out
api/speak/ Urdu text-to-speech
lib/
frames.ts local sharpness/brightness triage
use-vision.ts the camera loop
voice.ts speaking and listening, with fallbacks
vision/prompts.ts what the model is told, per mode
modes.ts money / medicine / document
urdu.ts every spoken string, in one file
store.ts local-only storage
Sending every frame to a vision model would burn a free tier in minutes and make the agent feel slow. So each frame is triaged in the browser first — a Laplacian-variance sharpness measure and a brightness check on a downscaled greyscale copy. Blurry or dark frames never leave the phone; the user just hears "ہاتھ روکیں" immediately, which is faster than any network round trip could be.
Only a frame that passes gets a network call. All three modes share one model — money used to run on a cheaper tier on the assumption that a note's large printed digits are an easy case, but the actual risk the money prompt warns about is a folded or worn note, and a misread note is as irreversible as a misread dose. Not worth the discount.
The model is reached through OpenRouter, which is deliberate: the biggest unknown in this project is which model actually reads Urdu Nastaliq well, and routing means that gets settled by changing an environment variable and re-testing rather than by swapping a dependency.
Speaking has three layers, tried in order: the server's Edge TTS route, then the browser's own speechSynthesis, then on-screen text. A blind user standing in a shop holding up a banknote needs to hear something.
For the same reason, /api/see returns a spoken hint rather than an HTTP error when the model is unreachable — the camera loop stays alive and a transient blip self-heals on the next frame.
Every prompt says never guess. A wrong answer about a medicine dose is worse than no answer. The model is told to set usable: false and ask the user to move the camera whenever it isn't sure, and to mark low confidence freely — which the agent then says aloud as "مجھے پورا یقین نہیں".
With no OPENROUTER_API_KEY set, Baseerat starts in demo mode: scripted readings that walk through the full interaction — two aiming hints, then an answer, complete with a prescription-mismatch warning. Both the screen and the voice say it's demo mode, so nobody is misled into trusting a canned answer about their medicine.
This exists so the app is never dead on arrival, and so a demo survives venue wifi falling over.
To switch the real vision on, get a key from OpenRouter:
cp .env.example .env.local
# add: OPENROUTER_API_KEY=sk-or-v1-...Point it at a :free model slug and it costs nothing at all — see .env.example.
Urdu speech needs no key at all. It uses Microsoft Edge's ur-PK neural voices (Asad and Uzma) through msedge-tts — free, no signup. Worth knowing: that rides an undocumented Microsoft endpoint, which is fine for a hackathon but should become the paid Azure Speech API before anyone depends on it.
These aren't garnish; they're the product.
- No account. Asking a blind user to complete a sign-up form before they can identify a banknote is exactly the barrier this app exists to remove. Everything is stored on the device.
- The whole screen is the button. Tap anywhere to talk. Double-tap repeats the last answer — the most requested thing in any spoken interface, because people miss what was said.
- Swipe to change mode, with spoken confirmation. Keys 1/2/3 do the same, for switch-access users and for testing.
- Light on near-black, very high contrast, nothing below 18px, 64px minimum touch targets.
- Amber warnings, not red — legible with red-green colour blindness, and readable on a dark ground.
- Pinch-zoom stays enabled. Disabling it is a common and avoidable accessibility failure.
- Urdu is set in Nastaliq where the device has it, falling back to Naskh — never to a Latin font, which would make the script unreadable.
- Answers are spoken in a deliberate order: headline, then warning, then detail. A listener cannot skim.
Vercel, and it needs no configuration — Next.js is detected automatically.
- Import the repo at vercel.com/new
- Add one environment variable:
OPENROUTER_API_KEY - Deploy
Then open /api/health on the deployed URL. It reports whether the vision
key is set and whether Edge TTS can hold its websocket out to Microsoft — the
one thing that can silently break in a serverless environment. If the voice
check fails the app still works; speech falls back to the browser's own Urdu
voice, which most Android phones have.
Set the project's function region to somewhere near Pakistan (Mumbai) in project settings — the default US region adds noticeable latency to every answer, and this is a conversation.
Any host that runs a normal Node process — Railway, Fly, Replit — works without that caveat, since the websocket is unquestionably fine there.
The camera and the Urdu voice can only really be judged on a phone, and browsers block both on any non-HTTPS address. So don't use the LAN address — build and serve, then tunnel:
npm run build && npm start
npx cloudflared tunnel --url http://localhost:3000Open the https://…trycloudflare.com URL it prints on the phone.
Use the production build for this, not npm run dev. The dev server refuses
to serve its JavaScript to an unrecognised origin, so a tunnelled dev server
renders the page but loads no JS — every button silently does nothing, which
looks exactly like a broken app.
npm run dev # development
npm run build # production build
npm start # serve the build
npm test # 40 tests
npm run lint # type-check
npm run check # bothTests cover mode detection from Urdu, roman-Urdu and English speech; the spoken string catalogue (script, length, English mirrors); the prompt contracts; and demo mode.
- Nastaliq OCR is hard. Urdu's cursive, context-dependent script is much harder to read than Naskh. Modern vision models do better than you'd expect, but test with real documents before trusting it — and compare a few models through OpenRouter rather than assuming the default is best.
- Speech input is Chrome-only and needs a connection. Every screen therefore works by touch as well.
- PKR note recognition is unverified against worn, folded, real-world notes. The prompt refuses to guess at partly visible notes, which is the safe failure — but it needs field testing.
- No blind user has tested this yet. That's the most important gap, and it's the next thing to fix — not another feature.
MIT licensed.