codewithfourtix/baseerat

★ 1Forks 0TypeScriptGitHub ↗Compare

Project website ↗

README

Baseerat · بصیرت

An Urdu voice agent that checks money, medicine and documents for blind and low-literacy Pakistanis.

Baseerat means insight — the kind that doesn't come from eyes.

npm install
npm run dev

Open http://localhost:3000. It works immediately, with no API key.


The problem

Existing accessibility apps — Be My Eyes, Seeing AI, Envision — are genuinely excellent, and they describe the world in English. A blind person in Lahore who doesn't read English gets nothing from them.

They're also read-only and low-stakes. They'll tell you there's a paper on the table. They won't tell you it's a loan agreement at 20% monthly interest that you're about to put your thumbprint on.

Baseerat is built for the moments where not knowing costs you money, health, or legal standing:

💵 Money Which note is this? Is this the right change?
💊 Medicine Is this the right strip? The right dose? Has it expired?
📄 Documents What am I about to agree to?

Four things that make it different

It speaks Urdu. Not localised menus — the agent thinks and answers in spoken Pakistani Urdu, in Urdu script, with proper ur-PK neural speech.

There is no shutter button. A blind user cannot know when the picture is good, so we never ask them to take one. The camera simply keeps looking, and the agent decides for itself when it has a frame worth answering from — talking the user into position along the way: "تھوڑا قریب لائیں… ہاتھ روکیں… مل گیا۔" Every competitor says "point your phone at the document," which quietly assumes you can see whether the document is in frame.

It checks against your prescription. Save what the doctor prescribed once, and every strip you hold up afterwards gets checked against it. Holding up 500mg when you were prescribed 250mg triggers a spoken warning, not a description.

It warns before you commit. Document mode doesn't just read a contract aloud — it tells you what you'd be agreeing to, and flags the parts that bind you.


How it's built

Next.js 16 · React 19 · TypeScript · Tailwind v4 · OpenRouter vision · Edge TTS

src/
  app/
    page.tsx              the whole agent — one screen, one big control
    setup/                prescription, voice settings, history
    api/see/              one frame in, a structured Urdu reading out
    api/speak/            Urdu text-to-speech
  lib/
    frames.ts             local sharpness/brightness triage
    use-vision.ts         the camera loop
    voice.ts              speaking and listening, with fallbacks
    vision/prompts.ts     what the model is told, per mode
    modes.ts              money / medicine / document
    urdu.ts               every spoken string, in one file
    store.ts              local-only storage

The camera loop is the interesting part

Sending every frame to a vision model would burn a free tier in minutes and make the agent feel slow. So each frame is triaged in the browser first — a Laplacian-variance sharpness measure and a brightness check on a downscaled greyscale copy. Blurry or dark frames never leave the phone; the user just hears "ہاتھ روکیں" immediately, which is faster than any network round trip could be.

Only a frame that passes gets a network call. All three modes share one model — money used to run on a cheaper tier on the assumption that a note's large printed digits are an easy case, but the actual risk the money prompt warns about is a folded or worn note, and a misread note is as irreversible as a misread dose. Not worth the discount.

The model is reached through OpenRouter, which is deliberate: the biggest unknown in this project is which model actually reads Urdu Nastaliq well, and routing means that gets settled by changing an environment variable and re-testing rather than by swapping a dependency.

Nothing is allowed to leave the user in silence

Speaking has three layers, tried in order: the server's Edge TTS route, then the browser's own speechSynthesis, then on-screen text. A blind user standing in a shop holding up a banknote needs to hear something.

For the same reason, /api/see returns a spoken hint rather than an HTTP error when the model is unreachable — the camera loop stays alive and a transient blip self-heals on the next frame.

Safety before helpfulness

Every prompt says never guess. A wrong answer about a medicine dose is worse than no answer. The model is told to set usable: false and ask the user to move the camera whenever it isn't sure, and to mark low confidence freely — which the agent then says aloud as "مجھے پورا یقین نہیں".


No API key? It still works

With no OPENROUTER_API_KEY set, Baseerat starts in demo mode: scripted readings that walk through the full interaction — two aiming hints, then an answer, complete with a prescription-mismatch warning. Both the screen and the voice say it's demo mode, so nobody is misled into trusting a canned answer about their medicine.

This exists so the app is never dead on arrival, and so a demo survives venue wifi falling over.

To switch the real vision on, get a key from OpenRouter:

cp .env.example .env.local
# add: OPENROUTER_API_KEY=sk-or-v1-...

Point it at a :free model slug and it costs nothing at all — see .env.example.

Urdu speech needs no key at all. It uses Microsoft Edge's ur-PK neural voices (Asad and Uzma) through msedge-tts — free, no signup. Worth knowing: that rides an undocumented Microsoft endpoint, which is fine for a hackathon but should become the paid Azure Speech API before anyone depends on it.


Accessibility decisions

These aren't garnish; they're the product.

  • No account. Asking a blind user to complete a sign-up form before they can identify a banknote is exactly the barrier this app exists to remove. Everything is stored on the device.
  • The whole screen is the button. Tap anywhere to talk. Double-tap repeats the last answer — the most requested thing in any spoken interface, because people miss what was said.
  • Swipe to change mode, with spoken confirmation. Keys 1/2/3 do the same, for switch-access users and for testing.
  • Light on near-black, very high contrast, nothing below 18px, 64px minimum touch targets.
  • Amber warnings, not red — legible with red-green colour blindness, and readable on a dark ground.
  • Pinch-zoom stays enabled. Disabling it is a common and avoidable accessibility failure.
  • Urdu is set in Nastaliq where the device has it, falling back to Naskh — never to a Latin font, which would make the script unreadable.
  • Answers are spoken in a deliberate order: headline, then warning, then detail. A listener cannot skim.

Deploying

Vercel, and it needs no configuration — Next.js is detected automatically.

  1. Import the repo at vercel.com/new
  2. Add one environment variable: OPENROUTER_API_KEY
  3. Deploy

Then open /api/health on the deployed URL. It reports whether the vision key is set and whether Edge TTS can hold its websocket out to Microsoft — the one thing that can silently break in a serverless environment. If the voice check fails the app still works; speech falls back to the browser's own Urdu voice, which most Android phones have.

Set the project's function region to somewhere near Pakistan (Mumbai) in project settings — the default US region adds noticeable latency to every answer, and this is a conversation.

Any host that runs a normal Node process — Railway, Fly, Replit — works without that caveat, since the websocket is unquestionably fine there.


Testing on a phone

The camera and the Urdu voice can only really be judged on a phone, and browsers block both on any non-HTTPS address. So don't use the LAN address — build and serve, then tunnel:

npm run build && npm start
npx cloudflared tunnel --url http://localhost:3000

Open the https://…trycloudflare.com URL it prints on the phone.

Use the production build for this, not npm run dev. The dev server refuses to serve its JavaScript to an unrecognised origin, so a tunnelled dev server renders the page but loads no JS — every button silently does nothing, which looks exactly like a broken app.


Commands

npm run dev     # development
npm run build   # production build
npm start       # serve the build
npm test        # 40 tests
npm run lint    # type-check
npm run check   # both

Tests cover mode detection from Urdu, roman-Urdu and English speech; the spoken string catalogue (script, length, English mirrors); the prompt contracts; and demo mode.


Honest limitations

  • Nastaliq OCR is hard. Urdu's cursive, context-dependent script is much harder to read than Naskh. Modern vision models do better than you'd expect, but test with real documents before trusting it — and compare a few models through OpenRouter rather than assuming the default is best.
  • Speech input is Chrome-only and needs a connection. Every screen therefore works by touch as well.
  • PKR note recognition is unverified against worn, folded, real-world notes. The prompt refuses to guess at partly visible notes, which is the safe failure — but it needs field testing.
  • No blind user has tested this yet. That's the most important gap, and it's the next thing to fix — not another feature.

MIT licensed.

Contributors

codewithfourtixMuhammadAnasTahirthesocialobaid

Issues