Pulkit7070/antiglasses

★ 0Forks 0TypeScriptGitHub ↗Compare

Project website ↗

README

Meian · See what's real.

Smart glasses with an on-device detection brain that catch deepfake scam voices and faces in real time, and warn you before you can be fooled.

Built for India's voice-clone / "digital arrest" / deepfake fraud epidemic.

Next.js landing + live demo · FastAPI detection backend · 4 real on-device AI models · working end-to-end


The problem

India lost Rs 19,813 crore to cyber fraud in a single year (2025), across 21.77 lakh complaints, one every fourteen seconds. The fastest-growing attacks are AI voice-clones and deepfake faces: a scammer needs just 3 seconds of your voice to clone it.

And it's almost never recoverable, only ~6% of stolen money is ever returned. Once the victim transfers it, it's gone. Every existing defence (caller-ID, awareness campaigns, bank OTP warnings, post-hoc fraud analytics) fires too late or assumes a human attacker. Prevention, in the moment of the call, is the only thing that works.

Source: India Cyber Crime Coordination Centre (I4C) / NCRP, 2025.

What Meian does

Meian is a detection brain you wear. It detects (analyses live audio + video on-device), decides (fuses multiple AI signals into one verdict), and warns (a spoken vernacular alert + heads-up display) , all in under a second, while the call is still happening, without anything leaving the device.

This repo contains a working version of that brain plus a live demo you can run today.

It works (verified, real models)

These are real outputs from the running system, not mockups:

Input Verdict Confidence
Synthetic scam call SYNTHETIC VOICE · fake 100%
Real human voice safe 0%
Deepfake face video flagged 4/4 frames 69%
Real face video safe 8%

It catches the fakes and never cries wolf on the real thing, that second property is the whole product thesis (a guardian that false-alarms gets switched off).

Model confidence: synthetic scam 100%, real voice 0%, deepfake face 69%, real face 8%

How it works

VOICE  ─┬─►  Whisper ASR          (speech → text)
        ├─►  AST / ASVspoof5      (synthetic voice?)     ─►  verdict fusion  ─►  spoken warning + HUD
        └─►  mDeBERTa (zero-shot) (scam-script tactics)

VIDEO  ────► frame sampler ─► dima806 ViT deepfake detector ─► frame vote ─► verdict

The voice verdict fuses three signals: an anti-spoof score, a scam-script tactic score (fake-authority, isolation, payment-demand, etc.), and the transcript. Audio and video are analysed live and discarded, never stored, never uploaded.

Models (open, on-device, downloaded on first use)

Purpose Model
Speech-to-text faster-whisper base
Synthetic-voice (anti-spoof) MattyB95/AST-ASVspoof5-Synthetic-Voice-Detection
Scam-script (zero-shot NLI) MoritzLaurer/mDeBERTa-v3-base-mnli-xnli
Video deepfake (per frame) dima806/deepfake_vs_real_image_detection

Live demo

# 1) Backend (the detection brain)
cd backend
python -m venv .venv
.venv/Scripts/python -m pip install -r requirements.txt        # macOS/Linux: .venv/bin/python
.venv/Scripts/python -m uvicorn app.main:app --port 8000

# 2) Frontend (landing page + /live demo)  — in a second terminal, from repo root
npm install
npm run dev

Open http://localhost:3000 for the landing page, or http://localhost:3000/live for the demo.

On /live, hit a one-click sample (Scam call, Genuine call, Deepfake face, Real face) , no audio of your own needed , and watch the verdict render with a plain-language explanation and hear the spoken warning. You can also record/upload your own clip. The first analysis of each type downloads the model (one-time, needs internet); after that it runs offline. Needs ~3 GB free disk.

Repo structure

glasses/
├── src/                  # Next.js landing page + /live demo (viewfinder/HUD design)
│   ├── app/live/         # the interactive demo page
│   ├── components/live/  # VoicePanel, VideoPanel, Verdict, Analyzing (skeleton)
│   └── lib/              # api client, content, alert (spoken warning)
├── backend/              # FastAPI detection brain
│   ├── app/              # main, schemas, verdict fusion, asr, voice_spoof, scam_script, video_deepfake
│   └── tests/            # unit tests + verified demo fixtures
├── deck/                 # case-competition pitch deck
│   ├── Meian_Pitch.pptx # 10-slide deck with speaker notes
│   ├── build_deck.py     # reproducible deck builder (python-pptx + matplotlib)
│   └── assets/           # generated charts
└── public/samples/       # the verified clips used by the live demo

Pitch deck

A 10-slide, judge-ready deck (with speaker notes) lives at deck/Meian_Pitch.pptx. It is fully reproducible: python deck/build_deck.py regenerates the charts and the PPTX.

Honest limitations

We don't oversell. Deepfake detection is an active research problem and we're transparent about the edges:

  • Anti-spoof generalisation: the AST model reliably flags the synthetic voices we tested (and passes real human speech), but no off-the-shelf detector catches every voice generator. The product's "continual model-update service" exists precisely because scams evolve.
  • Video: the pipeline discriminates real vs. fake on our samples, but we have not yet validated on a large set of true face-swap video deepfakes (vs. GAN stills). That's the next validation step.
  • Hardware: the detection brain runs today; the glasses themselves are on the roadmap. We ship a phone app first to prove demand.

Roadmap

When Milestone
Now Working detection brain + live demo, verified on real models
0-6 mo Mobile-app pilot with an NGO/bank for elderly-fraud protection
6-18 mo Edge device + glasses prototype; paid B2B pilots with telcos/banks
18 mo + Consumer glasses at scale; continual model-update service

Tech stack

Next.js · React · TypeScript · Tailwind · framer-motion · three.js / R3F · FastAPI · faster-whisper · transformers · PyTorch · OpenCV


"When the voice on the line lies, the glasses tell the truth, before the money moves."

Contributors

Pulkit7070

Issues