darshangavate/memora

โ˜… 0Forks 0PythonGitHub โ†—Compare

README

๐Ÿง  Memora โ€” Organizational Reasoning Engine

Trace how your organization's decisions were made. Ask one question and reconstruct the full decision journey across Slack, Gmail, meetings, and final documentation.

Status Python License


๐ŸŽฏ What is Memora?

Memora is an organizational memory and decision intelligence platform that combines real-time communication data with structured documentation to answer "Why did we decide this?" questions.

Instead of scrolling through months of Slack threads and email chains, Memora:

  1. Ingests Slack conversations, Gmail threads, meeting notes, and final decisions
  2. Embeds everything into a semantic vector database (ChromaDB)
  3. Retrieves relevant evidence using semantic search
  4. Ranks sources by organizational weight (Final Document > Meeting > Email > Slack)
  5. Reasons with Gemini AI to synthesize a grounded explanation
  6. Traces back to the original sources with full provenance

The Problem It Solves

  • ๐Ÿ” Lost institutional knowledge โ€” "Why did we choose FastAPI?"
  • โฐ Decision archaeology โ€” Decisions disappear from Slack after 30 days
  • ๐Ÿค Onboarding friction โ€” New team members can't find decision rationale
  • ๐Ÿ“Š Audit trails โ€” No clear evidence of how decisions were made
  • ๐Ÿงฉ Scattered context โ€” Decision history lives in 5+ different places

โœจ Key Features

๐Ÿ”„ Incremental Sync

  • Only fetches new messages since last sync (no full resets)
  • Deduplicates automatically
  • Tracks sync state in data/sync_state.json
  • No wasted API calls

๐Ÿ“š Multi-Source Ingestion

  • Slack โ€” Real-time channel discussions
  • Gmail โ€” Email thread context and formal decisions
  • Meeting Notes โ€” Structured discussion points
  • Final Documents โ€” Authoritative decision records

๐ŸŽ“ Intelligent Ranking

Final Document (strongest authority)
    โ†“
Meeting Notes (consensus discussion)
    โ†“
Gmail Threads (formal reasoning)
    โ†“
Slack (informal context)

๐Ÿงฌ Semantic Search & Reasoning

  • Powered by Google Gemini Embedding-001 for semantic understanding
  • Gemini 2.5 Flash for reasoning and synthesis
  • ChromaDB for sub-millisecond vector search

๐Ÿ“Š Decision Dashboard

  • Evidence cards with source provenance
  • Reasoning timeline showing how decisions evolved
  • Confidence scoring based on source diversity
  • Full traceability โ€” click any source to see original context

๐Ÿ”Œ Sync Control Panel

  • Manual "Sync Now" button for on-demand updates
  • Auto-sync every 60-120 seconds (configurable)
  • Last sync timestamp visible in sidebar
  • Real-time status feedback

๐Ÿ” Evidence Explorer

  • Unified view โ€” Single "All" tab shows combined records from all sources (Slack, Gmail, meetings, documents)
  • Per-source filtering โ€” Switch to individual source tabs for focused browsing
  • Full-text reading โ€” Always show complete source content in the detail pane
  • Quick identification โ€” Count badges on tabs show results per source
  • Clear selection โ€” Highlighted cards with visual accent bar make selection obvious
  • Internal scrolling โ€” Clean, compact record list with smooth pagination
  • Source badges โ€” Quick visual distinction of record origin (Slack, Gmail, Meeting, Document)
  • Responsive layout โ€” Reader pane dominates the right side for better scanning

๐Ÿ—๏ธ Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    Memora Application                        โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                               โ”‚
โ”‚  Frontend (Streamlit)                                        โ”‚
โ”‚  โ”œโ”€ Query Interface                                          โ”‚
โ”‚  โ”œโ”€ Evidence Cards                                           โ”‚
โ”‚  โ”œโ”€ Decision Timeline                                        โ”‚
โ”‚  โ””โ”€ Sync Control Panel                                       โ”‚
โ”‚                                                               โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                               โ”‚
โ”‚  Ingestion Pipeline                                          โ”‚
โ”‚  โ”œโ”€ fetch_slack.py โ”€โ”€โ”€โ”€โ”€โ”€โ–บ Slack API                        โ”‚
โ”‚  โ”œโ”€ fetch_gmail.py โ”€โ”€โ”€โ”€โ”€โ”€โ–บ Gmail API                        โ”‚
โ”‚  โ””โ”€ ingest.py                                                โ”‚
โ”‚      โ”œโ”€ Semantic Chunking                                    โ”‚
โ”‚      โ”œโ”€ Gemini Embeddings                                    โ”‚
โ”‚      โ””โ”€ ChromaDB Upsert                                      โ”‚
โ”‚                                                               โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                               โ”‚
โ”‚  Vector Store (ChromaDB)                                     โ”‚
โ”‚  โ”œโ”€ slack_messages                                           โ”‚
โ”‚  โ”œโ”€ gmail_messages                                           โ”‚
โ”‚  โ”œโ”€ meeting_notes (chunked)                                  โ”‚
โ”‚  โ””โ”€ final_documents (chunked)                                โ”‚
โ”‚                                                               โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                               โ”‚
โ”‚  Reasoning Engine                                            โ”‚
โ”‚  โ”œโ”€ Retrieval (semantic search)                              โ”‚
โ”‚  โ”œโ”€ Ranking (source priority)                                โ”‚
โ”‚  โ””โ”€ Synthesis (Gemini reasoning)                             โ”‚
โ”‚                                                               โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿš€ Quick Start

Prerequisites

  • Python 3.8+
  • Google Cloud credentials (Gmail API, Gemini API)
  • Slack Bot Token
  • Internet connection

1. Clone & Install

git clone https://github.com/yourusername/memora.git
cd memora
pip install -r requirements.txt

2. Set Up Environment Variables

Create a .env file in the root directory:

# Google APIs
GEMINI_API_KEY=your_gemini_api_key_here
GMAIL_GROUP=[email protected]

# Slack
SLACK_TOKEN=xoxb-your-slack-bot-token
SLACK_CHANNEL=all-memora-labs

# Authentication (optional but recommended)
ADMIN_EMAIL=[email protected]

3. Set Up Google Cloud Credentials

Place credentials.json in the root directory (from Google Cloud Console OAuth 2.0 setup):

# First run creates token.json after browser auth
python fetch_gmail.py

4. Initialize the Vector Store

# Fetches current Slack messages, Gmail threads, and ingests static docs
python ingest.py

This creates:

  • chroma_data/ โ€” Vector store with embeddings
  • data/raw/ โ€” Raw JSON from Slack/Gmail
  • data/sync_state.json โ€” Tracks last sync timestamp

5. Run the App

streamlit run app.py

Visit http://localhost:8501 in your browser.


๐Ÿ’ก How Memora Works

Example: "Why did we choose FastAPI over MERN/Node?"

Step 1: Retrieval

Query: "Why did we choose FastAPI over MERN/Node?"
โ†“
ChromaDB semantic search returns:
  โœ“ Slack: @alice "FastAPI is async-first"
  โœ“ Meeting Notes: "Team consensus on FastAPI"
  โœ“ Gmail: "Final decision approved FastAPI"
  โœ“ Final Document: "Tech Stack: FastAPI + React"

Step 2: Ranking

Sources prioritized by weight:
  1st โญโญโญ Final Document (authoritative)
  2nd โญโญ  Meeting Notes (consensus)
  3rd โญ   Gmail (formal reasoning)
  4th      Slack (informal context)

Step 3: Reasoning

Gemini synthesizes:
  "We chose FastAPI because:
   - Async performance (Slack discussion)
   - Team consensus in meeting
   - Formally approved in final document
   - Confidence: High"

Step 4: Traceability

User clicks "Final Document" โ†’
  Reads full context immediately
User clicks "Meeting Notes" โ†’
  Sees exact discussion points
User clicks "Gmail" โ†’
  Reads formal approval email

๐Ÿ“ Project Structure

memora/
โ”œโ”€โ”€ app.py                      # Main Streamlit application
โ”œโ”€โ”€ ingest.py                   # Incremental ingestion pipeline
โ”œโ”€โ”€ fetch_slack.py              # Slack API integration
โ”œโ”€โ”€ fetch_gmail.py              # Gmail API integration
โ”œโ”€โ”€ requirements.txt            # Python dependencies
โ”œโ”€โ”€ credentials.json            # Google OAuth (โš ๏ธ gitignore)
โ”œโ”€โ”€ token.json                  # Gmail token (โš ๏ธ gitignore)
โ”‚
โ”œโ”€โ”€ data/
โ”‚   โ”œโ”€โ”€ raw/
โ”‚   โ”‚   โ”œโ”€โ”€ slack_messages.json
โ”‚   โ”‚   โ”œโ”€โ”€ slack_users.json
โ”‚   โ”‚   โ”œโ”€โ”€ gmail_messages.json
โ”‚   โ”‚   โ”œโ”€โ”€ gmail_threads.json
โ”‚   โ”‚   โ”œโ”€โ”€ meeting_notes.txt
โ”‚   โ”‚   โ””โ”€โ”€ final_document.txt
โ”‚   โ”œโ”€โ”€ sync_state.json         # Tracks last sync timestamps
โ”‚   โ””โ”€โ”€ config.json
โ”‚
โ”œโ”€โ”€ chroma_data/                # Vector store (persistent)
โ”‚   โ””โ”€โ”€ [auto-created]
โ”‚
โ”œโ”€โ”€ pages/
โ”‚   โ”œโ”€โ”€ 1_Login.py             # Authentication page
โ”‚   โ”œโ”€โ”€ 2_My_Organization.py   # Org settings page
โ”‚   โ””โ”€โ”€ source_explorer.py     # Browse sources
โ”‚
โ”œโ”€โ”€ utils/
โ”‚   โ””โ”€โ”€ auth.py                # Session management
โ”‚
โ””โ”€โ”€ README.md

๐Ÿ”ง Configuration

Auto-Sync Interval

In the Streamlit sidebar, select sync frequency:

  • 60 sec โ€” Check for new messages very frequently
  • 90 sec โ€” Balanced (default)
  • 120 sec โ€” Less frequent network calls

Static Documents

Place these in data/raw/ for one-time ingestion:

  • meeting_notes.txt โ€” Structured meeting minutes
  • final_document.txt โ€” Decision records

Slack Channel

Set in .env:

SLACK_CHANNEL=important-decisions

Gmail Group

Set in .env:

GMAIL_GROUP=[email protected]

๐Ÿงฌ Technical Deep Dive

Incremental Ingestion Strategy

Before (โŒ Full Reset)

# Old approach: wasteful
reset_chroma()  # Delete everything
fetch_all_slack()
fetch_all_gmail()
re-embed everything

After (โœ… Incremental)

# New approach: efficient
last_ts = load_sync_state()["slack"]["last_ts"]
new_messages = fetch_slack(oldest=last_ts)
upsert(new_messages)  # Only add new docs
save_sync_state(last_ts)

Benefits:

  • โšก 10-100x faster (only new items)
  • ๐Ÿ’ฐ Fewer API calls
  • ๐Ÿ”„ No downtime for ingestion
  • ๐Ÿ“Š Full history preserved

Vector Database (ChromaDB)

Why ChromaDB?

  • โœ… Persistent local storage
  • โœ… Semantic search (cosine similarity)
  • โœ… Metadata filtering
  • โœ… No external dependencies (SQLite)

Schema:

{
  "id": "slack_1712742600.123456",
  "document": "Complete message text...",
  "metadata": {
    "source": "slack",
    "channel": "general",
    "user": "U123456",
    "user_name": "Alice",
    "ts": "1712742600.123456"
  }
}

Embedding & Ranking

Embeddings:

  • Model: gemini-embedding-001
  • Dimension: 768
  • Update: Only on new documents (incremental)

Ranking Function:

def source_priority(source: str) -> int:
    return {
        "final_document": 1,      # Strongest
        "meeting": 2,
        "gmail": 3,
        "slack": 4,               # Weakest
    }.get(source, 99)

๐Ÿ› ๏ธ Advanced Usage

Manual Data Ingestion

Ingest static files manually:

from ingest import run_incremental_ingestion

result = run_incremental_ingestion(include_static_docs=True)
# Returns: {"new_slack": 5, "new_gmail": 3, "meeting_chunks": 12, ...}

Direct Vector Search

from chromadb import PersistentClient

client = PersistentClient(path="./chroma_data")
collection = client.get_collection("org_memory")

results = collection.query(
    query_texts=["Why did we choose FastAPI?"],
    n_results=5
)

Custom Reasoning Prompt

Edit the prompt in app.py (around line 700) to change reasoning behavior:

prompt = f"""
You are Memora, an organizational reasoning engine.

[Customize instructions here]

Evidence:
{context}
"""

๐Ÿ” Security & Privacy

Protected Data

  • Slack messages & emails are sensitive
  • Store credentials.json and token.json in .gitignore
  • Use environment variables for API keys
  • Implement row-level access control in utils/auth.py

Authentication

  • Built-in session management
  • Role-based access (Admin, Member, Viewer)
  • User profile stored in st.session_state

Audit Trail

Every decision answer includes:

  • Source documents
  • Retrieval confidence
  • Timestamp of analysis

๐Ÿšง Roadmap

Phase 1 โœ… (Current)

  • Slack + Gmail integration
  • Incremental sync
  • Vector search & ranking
  • Gemini reasoning
  • Decision traceability

Phase 2 (Planned)

  • Calendar integration (calendar events as context)
  • Jira/Linear issue tracking
  • Slack thread replies (nested discussions)
  • Multi-org support
  • Advanced filtering (date range, users, channels)

Phase 3 (Future)

  • Decision workflow automation
  • Confidence scoring ML model
  • Custom LLM fine-tuning
  • GraphQL API for external tools
  • Slack bot for inline answers

๐Ÿ“š Examples

Query 1: Tech Stack Decision

Q: "Why did we choose React over Vue?" A: Best explanation sourced from Final Document + Meeting Notes + Gmail + Slack

Query 2: Process Changes

Q: "How did we decide on our deployment pipeline?" A: Timeline shows Slack discussions โ†’ Email approvals โ†’ Meeting consensus โ†’ Final doc

Query 3: Architecture Decisions

Q: "What was the rationale for microservices?" A: Ranks sources by authority; explains evolution of thinking


๐Ÿค Contributing

Contributions welcome! To add features:

  1. Fork the repo
  2. Create a feature branch (git checkout -b feature/amazing-feature)
  3. Make changes (update tests if needed)
  4. Commit (git commit -m 'Add amazing feature')
  5. Push (git push origin feature/amazing-feature)
  6. Open a Pull Request

๐Ÿ“„ License

This project is licensed under the MIT License โ€” see LICENSE.md for details.


๐Ÿ™‹ Support & Questions

  • Documentation: See DESIGN.md for UI/UX specs
  • Issues: Open a GitHub issue for bugs
  • Discussions: Start a discussion for feature requests
  • Email: [email protected]

๐ŸŽ“ Credits

Built with โค๏ธ for organizational transparency and decision intelligence.

Technologies:


๐Ÿ“Š Project Stats

  • Lines of Code: ~2000
  • API Integrations: 3 (Slack, Gmail, Gemini)
  • Vector Dimensions: 768
  • Supported Data Sources: 2 (+ static)
  • Reasoning Model: Gemini 2.5 Flash
  • Database: ChromaDB + SQLite

Made with ๐Ÿง  for better organizational memory.

โฌ† Back to Top

Contributors

darshangavatekalyaniugale

Issues