mjamiv/vox2txt

★ 0Forks 0JavaScriptGitHub ↗Compare

README

northstar.LM

Transform meetings into actionable insights with AI

Live Demo: https://mjamiv.github.io/vox2txt/


Overview

northstar.LM is a client-side web application that uses OpenAI's AI models to analyze meeting content and enable intelligent cross-meeting queries. All processing happens in your browser with your own API key—no server-side data handling.

Application Purpose
Agent Builder Process recordings, videos, documents, images, or text into AI agents with summaries, key points, action items, and sentiment analysis
Agent Orchestrator Combine multiple agents for cross-meeting analysis using the RLM pipeline

Key Capabilities

Agent Builder

  • Multi-format Input: Audio (MP3, WAV, M4A), Video (MP4, WebM), PDF, Images, Text, URLs
  • AI Analysis: Summaries, key points, action items, sentiment analysis via GPT-5.2
  • Auto-Generated Agenda: Next meeting agenda created automatically after analysis
  • Voice Chat: Two modes for voice interaction with meeting content
    • Push-to-Talk: Hold mic → Whisper transcription → Chat → TTS response (~$0.02/exchange)
    • Real-time: Continuous voice conversation via OpenAI Realtime API (~$0.30/min)
  • RLM-Powered Chat: Direct/RLM toggle for intelligent context handling
  • Custom Audio Player: Styled player with play/pause, progress bar, volume control, and download
  • Meeting Infographic: 4 style presets (Executive, Dashboard, Action Board, Timeline) with black/gold theme
  • Collapsible Results: Key Points, Action Items, Agenda, and Infographic in expandable cards
  • Chat Reminder: Tooltip appears after analysis to encourage interaction
  • Professional Export: DOCX reports and portable agent files (.md)

Agent Orchestrator

  • Simplified Settings: Three preset modes replace complex configuration
    • Quick: Fast responses, lower cost (Direct Chat)
    • Balanced: Recommended for most queries (RLM + Signal-Weighted Memory)
    • Deep: Thorough multi-agent reasoning (RLM + Hybrid Focus)
  • Active Agent Chips: Visual indicators showing which agents are active, click to toggle
  • Knowledge Base Views: Toggle between interactive Canvas and sortable List view
  • Multi-Agent Analysis: Load and query multiple meeting agents simultaneously
  • RLM Pipeline: Intelligent query decomposition with source attribution
  • Cross-Meeting Insights: Collapsible cards for themes, trends, risks, recommendations, and actions
    • Color-coded borders by category (gold/blue/red/purple/green)
    • Click headers to expand/collapse individual sections
  • Executive Audio Briefing: ~3 minute podcast-style update with upbeat, engaging delivery via TTS
  • Insights Infographic: Visual summary using DALL-E with premium black/gold executive theme
  • Weekly Agenda: RLM-powered agenda generation with:
    • Methodology narrative explaining multi-agent Societies of Thought analysis
    • Per-person task lists for core standup team members only
    • External dependencies assigned to core team owners
  • Advanced Options: Collapsible drawer with model settings, context gauge, and test prompting
  • Export Options:
    • DOCX report with insights and weekly agenda
    • Full session export (.md) with embedded JSON for later import
    • Chat history export
  • Session Import: Restore agents, groups, insights, chat, and settings from exported sessions
  • Comprehensive Metrics: Token usage, costs, response times, and CSV export
  • Toast Notifications: Non-intrusive feedback for actions and errors

RLM Validation Results (January 2026)

Comprehensive stress testing validates RLM's ability to maintain conversation context across extended sessions.

Response Capability

Test Direct Chat RLM Improvement
25-Question 80% 100% +25%
50-Question 35% 96% +174%
100-Question 18% 96% +433%

Critical Finding: Direct Chat loses access to earlier conversation turns by Turn 7-10. RLM maintains 95-96% response capability through 100+ turns by re-querying source agents.

Cost & Performance

Mode Avg Cost/Prompt Token Reduction Latency
Direct Chat $0.128 — 13.4s
RLM + SWM $0.055 77-81% 27.9s
RLM + Hybrid $0.054 77-81% 33.6s

Recommendation: Default to RLM + Signal-Weighted Memory for conversations exceeding 5-7 turns. Use Direct Chat for quick, single-turn queries.

Full Report: Testing/RLM-Validation-Study-Final-Report.html


Processing Modes

Agent Builder

Toggle between Direct and RLM modes for chat and agenda generation. RLM is enabled by default, providing enhanced context handling via signal-weighted memory.

Agent Orchestrator

Three preset modes simplify configuration:

Preset Mode Best For
Quick Direct Chat Fast, single-turn queries
Balanced RLM + Signal-Weighted Memory Most queries (recommended)
Deep RLM + Hybrid Focus Complex multi-agent analysis

Advanced users can access detailed settings (model selection, effort level, individual toggles) via the Advanced Options drawer.


Application Workflow

flowchart TB
    subgraph Input["INPUT SOURCES"]
        A1[Audio/Video]
        A2[PDF/Image]
        A3[Text/URL]
    end

    subgraph Process["PROCESSING"]
        B1[Whisper API]
        B2[PDF.js / Vision AI]
        B3[Text Parser]
    end

    subgraph Analysis["AI ANALYSIS"]
        C1[GPT-5.2 Analysis]
        C2[Summary + Key Points]
        C3[Actions + Sentiment]
    end

    subgraph Builder["AGENT BUILDER"]
        D1[KPI Dashboard]
        D2[RLM Chat Interface]
        D3[Export Agent .md]
    end

    subgraph Orchestrator["AGENT ORCHESTRATOR"]
        E1[Load Multiple Agents]
        E2[RLM Pipeline]
        E3[Cross-Meeting Insights]
        E4[Audio / Infographic / Agenda]
        E5[Export DOCX / Session]
    end

    A1 --> B1 --> C1
    A2 --> B2 --> C1
    A3 --> B3 --> C1
    C1 --> C2 --> C3
    C3 --> D1 & D2 & D3
    D3 --> E1 --> E2 --> E3 --> E4 --> E5

    style Input fill:#1a1f2e,stroke:#d4a853,color:#fff
    style Process fill:#1a2a1a,stroke:#4ade80,color:#fff
    style Analysis fill:#2a1a2a,stroke:#a855f7,color:#fff
    style Builder fill:#1a2a3a,stroke:#60a5fa,color:#fff
    style Orchestrator fill:#2a2a1a,stroke:#fbbf24,color:#fff
Loading

Technology Stack

Category Technologies
Frontend Vanilla HTML, CSS, JavaScript (ES Modules), PWA
AI Models Whisper, GPT-5.2, GPT-5.2 Vision, GPT-4o-mini-TTS, GPT-Image-1.5, GPT-4o-Realtime
Libraries docx.js, PDF.js, marked.js, Pyodide
Deployment GitHub Pages

Pricing Reference

Model Input Output
GPT-5.2 $1.75/1M tokens $14.00/1M tokens
GPT-5-mini $0.25/1M tokens $2.00/1M tokens
Whisper $0.006/minute —
GPT-4o-mini-TTS — $0.015/1K chars
Realtime API $0.06/min (audio) $0.24/min (audio)

Getting Started

Agent Builder

  1. Visit https://mjamiv.github.io/vox2txt/
  2. Enter your OpenAI API key (stored locally)
  3. Upload audio, video, PDF, image, or paste text
  4. Click Analyze Meeting
  5. Review KPI dashboard, summary, and auto-generated agenda
  6. Expand Key Points, Action Items, or Infographic sections as needed
  7. Use the floating chat widget to ask follow-up questions
  8. Generate Audio Briefing or Infographic from the Generate menu
  9. Export as DOCX report or Agent file

Agent Orchestrator

  1. Export meetings as Agent files from the Builder
  2. Visit the Orchestrator
  3. Load agent files into the Knowledge Base (drag & drop or click to upload)
  4. Select a preset mode: Quick, Balanced (default), or Deep
  5. Use the chat to ask cross-meeting questions
  6. Click agent chips to toggle agents on/off for focused queries
  7. Generate Cross-Meeting Insights (collapsible cards for each category)
  8. Generate deliverables:
    • Audio Update: Podcast-style executive briefing
    • Infographic: Visual insights summary
    • Weekly Agenda: Per-person task lists with methodology narrative
  9. Export session or insights as DOCX/Markdown

Privacy & Security

  • Local Storage: API key stored in browser localStorage
  • Direct API Calls: Requests go directly to OpenAI
  • No Third Parties: No data sent to external servers
  • Client-Side Only: All processing in your browser

Local Development

git clone https://github.com/mjamiv/vox2txt.git
cd vox2txt

# Serve locally
npx http-server -p 3000
# or
python -m http.server 3000

# Open http://localhost:3000

Documentation

Document Description
CLAUDE.md Development guide and architecture reference
RLM_STATUS.md RLM implementation details and API reference
Voice Chat Guide Voice conversation implementation (Push-to-Talk & Real-time)
Testing/ Validation test data and reports

License

MIT License

Contributors

mjamivclaudecursoragent

Issues