emrsyah/OpenTA

The intelligent open directory for Telkom University research. Search, ask, and discover insights from alumni papers using advanced AI.

★ 10Forks 0TypeScriptGitHub ↗Compare

Project website ↗

README

image

OpenTA - AI-Powered Co-Researcher for Telkom University

Next.js TypeScript Tailwind CSS PostgreSQL DSPy LangChain LangGraph Voyage AI Exa AI License

An agent-native research workspace that helps you discover papers, find lecturers, and synthesize knowledge from Telkom University's academic community.

Features • Quick Start • Architecture • Contributing


📖 Overview

OpenTA is an AI-powered co-researcher designed to accelerate academic research. Unlike traditional paper repositories, this is an agent-native workspace where AI agents actively help you:

  • 🔍 Discover relevant papers from Telkom University's vast research database
  • 👨‍🏫 Find Lecturers with AI-powered semantic search and web-enriched profiles
  • 🧠 Synthesize knowledge across multiple papers and sources
  • 🔬 Run deep research tasks with autonomous agents that can perform multi-step investigations
  • 📊 Generate insights through context engineering and agent harness patterns

Built with LangChain/LangGraph for agent orchestration, the system uses session-based memory for conversation context and custom tools for research operations.

Vision

To create an AI research assistant that doesn't just retrieve papers, but actively collaborates in the research process—helping researchers find connections, synthesize knowledge, and accelerate discovery at Telkom University.

✨ Features

🚀 Current Features

Feature Description Status
AI Research Assistant LangChain-powered agents for research queries ✅ Implemented
Paper Discovery Search Tel-U alumni papers with semantic search ✅ Implemented
Cari Dosen AI-powered lecturer search with Exa web enrichment ✅ Implemented
Research Filtering Metadata filters for refined research results ✅ Implemented
User Feedback Widget Collect feedback on AI responses ✅ Implemented
Conversation Management Persistent research sessions with history ✅ Implemented
Source Citations Inline citations with paper metadata ✅ Implemented
Saved Papers Save papers to collections for later reference ✅ Implemented
Collections Create and manage personal paper collections ✅ Implemented
JWT Backend Auth Secure auth for DSPy backend service ✅ Implemented

📋 Planned Features

  • Deep Research Agent - Autonomous agents that run long-form research tasks
  • Experiment Simulation - Agents that can propose and validate hypotheses
  • Literature Review Agent - Automated systematic reviews
  • Ideas Exploration - Agents that can explore ideas and concepts
  • Citation Network Analysis - Visualize paper relationships
  • Multi-Agent Collaboration - Specialized agents working together
  • Research Task Queuing - Schedule and track long-running research
  • Export Research Reports - Generate comprehensive research summaries

🏗️ Architecture

System Overview

┌─────────────────────────────────────────────────────────────────┐
│                    Frontend (Next.js)                           │
│  ┌──────────────┐  ┌──────────────┐  ┌──────────────────────┐  │
│  │   Pages      │  │  Components  │  │   Hooks & Utils     │  │
│  │  (App Router)│  │   (UI + Chat)│  │  (State Management) │  │
│  └──────────────┘  └──────────────┘  └──────────────────────┘  │
└─────────────────────────────────────────────────────────────────┘
                            │
                    ┌───────┴────────┐
                    │                │
            ┌───────▼────────┐  ┌───▼────────────┐
            │  API Routes    │  │ better-auth    │
            │  (Next.js)     │  │   Sessions     │
            └───────┬────────┘  └────────────────┘
                    │
        ┌───────────┼────────────┐
        │           │            │
┌───────▼─────┐ ┌──▼──────────┐ └───┐
│   PostgreSQL │ │ LangChain    │     │
│   Database   │ │   Agent      │     │
│  (Drizzle)   │ │ (Next.js)    │     │
│              │ │  + Tools     │     │
└──────────────┘ └─────────────┘     │
      │                                │
      └────────────────────────────────┘
           Context + Research Flow

┌─────────────────────────────────────────────────────────────────┐ │ Frontend (Next.js) │ │ ┌──────────────┐ ┌──────────────┐ ┌──────────────────────┐ │ │ │ Pages │ │ Components │ │ Hooks & Utils │ │ │ │ (App Router)│ │ (UI + Chat)│ │ (State Management) │ │ │ └──────────────┘ └──────────────┘ └──────────────────────┘ │ └─────────────────────────────────────────────────────────────────┘ │ ┌───────┴────────┐ │ │ ┌───────▼────────┐ ┌───▼────────────┐ │ API Routes │ │ better-auth │ │ (Next.js) │ │ Sessions │ └───────┬────────┘ └────────────────┘ │ ┌───────────┼────────────┐ │ │ │ ┌───────▼─────┐ ┌──▼──────────┐ └───┐ │ PostgreSQL │ │ DSPy │ │ │ Database │ │ Backend │ │ │ (Drizzle) │ │ (FastAPI) │ │ │ │ │ + Agents │ │ └──────────────┘ └─────────────┘ │ │ │ └────────────────────────────────┘ Context + Research Flow


### Agent Architecture

#### Research Agent Flow

1. **User Query** → Research question or task
2. **Session Memory** → Load conversation history and context
3. **LangGraph Agent** → Route to appropriate tools (search, retrieve, analyze)
4. **Tool Execution** → Agent runs reasoning chain with custom tools
5. **Response Generation** → Structured research output with citations

1. **User Query** → Research question or task
2. **Context Engineering** → Gather relevant papers, history, and domain knowledge
3. **Agent Harness** → Route to appropriate DSPy agent (search, synthesize, analyze)
4. **DSPy Execution** → Agent runs reasoning chain with tools
5. **Response Generation** → Structured research output with citations

┌─────────────┐ Research Task ┌──────────────────┐ │ User │────────────────────▶│ Agent Harness │ └─────────────┘ │ (Orchestrator) │ └────────┬─────────┘ │ ┌──────────────┼──────────────┐ │ │ │ ┌──────▼─────┐ ┌────▼─────┐ ┌────▼────────┐ │ Search │ │ Analyze │ │ Synthesize │ │ Agent │ │ Agent │ │ Agent │ └──────┬─────┘ └────┬─────┘ └────┬────────┘ │ │ │ └────────────┼────────────┘ │ ┌────────────▼────────────┐ │ Context Engineering │ │ (Paper DB + History) │ └────────────────────────┘


### LangChain/LangGraph Agent System

**Search Tool**: Find relevant papers using semantic search  
**Retrieve Tool**: Fetch paper details and metadata  
**Analysis Tool**: Extract key insights, methodologies, findings  
**Synthesis Tool**: Combine multiple papers into coherent answer  
**Deep Research Agent** (Planned): Run multi-step investigations with subtasks


**Analysis Agent**: Extract key insights, methodologies, findings  
**Synthesis Agent**: Combine multiple papers into coherent answer  
**Deep Research Agent** (Planned): Run multi-step investigations with subtasks

### Tech Stack

#### Frontend (Next.js)
- **Framework**: [Next.js 16.1.6](https://nextjs.org/) (App Router, React Server Components)
- **Language**: [TypeScript 5](https://www.typescriptlang.org/)
- **Styling**: [Tailwind CSS 4](https://tailwindcss.com/)
- **UI Components**: [Radix UI](https://www.radix-ui.com/) + shadcn
- **Animations**: [Motion](https://motion.dev/)
- **Icons**: [Lucide React](https://lucide.dev/)
- **Streamdown**: [Streamdown](https://streamdown.dev/) (code, math, mermaid, CJK support)

#### Backend (LangChain/LangGraph Agent)
- **Agent Framework**: [LangChain](https://langchain.com/) + [LangGraph](https://langgraph.ai/) (Graph-based agent orchestration)
- **LLM**: OpenAI GPT models
- **Embedding**: [Voyage AI](https://voyageai.com/) for semantic search
- **Tools**: Custom tools for paper search, retrieval, and analysis
- **Session Memory**: Conversation history with semantic retrieval
- **Agent Framework**: [DSPy](https://github.com/stanfordnlp/dspy) (Declarative agent programming)
- **FastAPI**: REST API for agent endpoints
- **Principle - Agent Harness** : [Agent Harness](https://www.philschmid.de/agent-harness-2026)

#### Database & ORM
- **Database**: [PostgreSQL](https://www.postgresql.org/) with vector search
- **ORM**: [Drizzle ORM](https://orm.drizzle.team/)
- **Migrations**: [Drizzle Kit](https://kit.drizzle.team/)
- **Vector Embeddings**: pgvector for semantic search with [Voyage AI](https://voyageai.com/)

#### External APIs
- **Voyage AI**: Embedding generation for semantic search (lecturer matching)
- **Exa AI**: Web search for lecturer profiles and contact information

#### Authentication
- **Auth Library**: [better-auth 1.4.18](https://www.better-auth.com/)
- **OAuth**: Google SSO
- **Session Management**: JWT-based stateless sessions
- **Agent Security**: Session-based authentication for internal LangChain agent
- **Backend Security**: JWT token validation for DSPy service

#### Development Tools
- **Package Manager**: [Bun](https://bun.sh/)
- **Linting**: [Biome](https://biomejs.dev/)
- **Type Checking**: TypeScript 5

## 🚀 Quick Start

### Prerequisites

Ensure you have the following installed:

- [Node.js 20+](https://nodejs.org/) or [Bun](https://bun.sh/)
- [PostgreSQL 14+](https://www.postgresql.org/download/) with pgvector
- [OpenAI API Key](https://platform.openai.com/) (for LLM)
- Google Cloud Project (for OAuth)
- [PostgreSQL 14+](https://www.postgresql.org/download/) with pgvector
- [Python 3.10+](https://www.python.org/downloads/) (for DSPy backend)
- Google Cloud Project (for OAuth)

### 1. Clone the Repository

```bash
git clone https://github.com/yourusername/open-ta-telyu.git
cd open-ta-telyu

2. Install Dependencies

bun install

3. Environment Setup

Create a .env file in the root directory:

cp .env.example .env

Configure your environment variables:

# Application
NEXT_PUBLIC_APP_URL=http://localhost:3000

# Database
# Application
NEXT_PUBLIC_APP_URL=http://localhost:3000
NEXT_PUBLIC_BACKEND_URL=http://localhost:8000

# Database
DATABASE_URL=postgresql://postgres:password@localhost:5432/openta

# Better Auth
BETTER_AUTH_SECRET=your-super-secret-key-at-least-32-chars-long
BETTER_AUTH_URL=http://localhost:3000

# Google OAuth
BETTER_AUTH_SECRET=your-super-secret-key-at-least-32-chars-long
BETTER_AUTH_URL=http://localhost:3000

# Backend API Shared Secret (for DSPy service)
BACKEND_API_SECRET=your-backend-api-secret-min-32-chars

# Google OAuth
GOOGLE_CLIENT_ID=your-google-client-id.apps.googleusercontent.com
GOOGLE_CLIENT_SECRET=your-google-client-secret

# Voyage AI (for vector embeddings)
VOYAGE_API_KEY=your-voyage-api-key

# OpenAI (for LLM)
OPENAI_API_KEY=sk-...

# Exa AI (for web search - lecturer profiles)
EXA_API_KEY=your-exa-api-key

VOYAGE_API_KEY=your-voyage-api-key

Exa AI (for web search - lecturer profiles)

EXA_API_KEY=your-exa-api-key


Generate secrets with:
```bash
openssl rand -base64 32

4. Database Setup

# Push database schema
bun run db:push

# (Optional) Open Drizzle Studio to inspect database
bun run db:studio
bun run dev

Open http://localhost:3000 in your browser.

📁 Project Structure

# Start the development server
bun run dev

Open http://localhost:3000 in your browser.

# Frontend
bun run dev

# Backend (separate repository)
cd open-ta-backend
python -m uvicorn main:app --reload

Open http://localhost:3000 in your browser.

📁 Project Structure

open-ta-telyu/ (Frontend)
├── src/
│   ├── app/                    # Next.js App Router pages
│   │   ├── page.tsx           # Home page (research interface)
│   │   ├── browse/            # Paper browse page
│   │   ├── cari-dosen/        # Lecturer search page (Cari Dosen)
│   │   │   ├── page.tsx       # Main lecturer search
│   │   │   └── [name]/        # Lecturer detail page
│   │   ├── [id]/              # Research session page
│   │   ├── api/               # API routes
│   │   │   ├── auth/          # better-auth endpoints
│   │   │   ├── chat/          # LangChain agent endpoint
│   │   │   ├── conversations/ # Session CRUD
│   │   │   ├── conversations/ # Session CRUD
│   │   │   ├── catalog/       # Paper search
│   │   │   ├── lecturers/     # Lecturer search & details
│   │   │   └── feedback/      # User feedback submission
│   │   ├── layout.tsx         # Root layout
│   │   └── globals.css        # Global styles
│   ├── components/            # React components
│   │   ├── ui/                # shadcn/ui components
│   │   ├── chat/              # Chat/research components
│   │   ├── browse/            # Browse page components
│   │   ├── auth/              # Authentication components
│   │   ├── ai-elements/       # AI response elements
│   │   └── lecturer-card.tsx  # Lecturer display card
│   ├── hooks/                 # Custom React hooks
│   │   ├── lib/                   # Utility libraries
│   │   │   ├── auth/              # Auth utilities
│   │   │   ├── db/                # Database functions
│   │   │   ├── voyage.ts          # Voyage AI embedding client
│   │   │   ├── lecturer-utils.ts  # Lecturer data utilities
│   │   │   └── ai/                # LangChain agent implementation
│   │   │       ├── agent.ts       # Main agent definition
│   │   │       ├── tools.ts       # Custom agent tools
│   │   │       ├── retriever.ts   # Document retriever
│   │   │       ├── session-memory.ts # Session memory management
│   │   │       ├── types.ts       # TypeScript types
│   │   │       ├── stream.ts      # Streaming utilities
│   │   │       ├── prompts/       # Agent prompts
│   │   │       └── citation-audit.ts # Citation verification
│   │   ├── auth/              # Auth utilities (JWT generation)
│   │   ├── db/                # Database functions
│   │   ├── voyage.ts          # Voyage AI embedding client
│   │   └── lecturer-utils.ts  # Lecturer data utilities
│   └── db/                    # Database schema
│       ├── schema/            # Drizzle schema definitions
│       └── migrations/        # SQL migrations
├── public/                    # Static assets
├── scripts/                   # Utility scripts
├── drizzle.config.ts          # Drizzle ORM config
├── biome.json                 # Biome linter config
├── next.config.ts             # Next.js configuration
├── tailwind.config.ts         # Tailwind CSS config
├── tsconfig.json              # TypeScript config
└── package.json               # Dependencies

open-ta-backend/ (DSPy Agents - Separate Repo)
├── agents/                    # DSPy agent definitions
├── context/                   # Context engineering modules
├── harness/                   # Agent orchestration patterns
├── tools/                     # Agent tools (search, retrieve, etc.)
└── main.py                    # FastAPI application

🔌 API Endpoints

Authentication (better-auth)

Endpoint Method Description
/api/auth/sign-in/google GET Initiate Google OAuth
/api/auth/sign-out POST Sign out user
/api/auth/session GET Get current session

Research Sessions (Conversations)

Endpoint Method Description Auth Required
/api/conversations GET List user research sessions ✅
/api/conversations POST Create new research session ✅
/api/conversations/[id] DELETE Delete research session ✅
/api/conversations/[id]/messages GET Get session history ✅

Research (DSPy Backend)

Endpoint Method Description Auth Required
/api/chat POST Stream agent research response ✅

Paper Catalog

Endpoint Method Description Auth Required
/api/catalog POST Search research papers ❌

Lecturer Search (Cari Dosen)

Endpoint Method Description Auth Required
/api/lecturers/list GET List all lecturers ❌
/api/lecturers/search POST Semantic search lecturers ❌
/api/lecturers/detail GET Get lecturer details ❌
/api/lecturers/web-search GET Exa web search for lecturer ❌

| /api/feedback | POST | Submit user feedback | ✅ |

Saved Papers & Collections

Endpoint Method Description Auth Required
/api/saved-papers GET List saved papers ✅
/api/saved-papers POST Save a paper to collection ✅
/api/saved-papers/[id] DELETE Remove saved paper ✅
/api/saved-papers/[id] PATCH Update saved paper (note, collection) ✅
/api/saved-papers/status/[catalogId] GET Check if paper is saved ✅
/api/collections GET List user collections ✅
/api/collections POST Create new collection ✅
/api/collections/[id] DELETE Delete collection ✅

🔒 Authentication & Agent Security

Session Flow

  1. User Sign-In: Redirects to Google OAuth
  2. Session Creation: better-auth creates session in database
  3. JWT Generation: Frontend generates short-lived JWT for DSPy backend
  4. Agent Verification: DSPy service validates JWT signature
  5. Request Processing: User ID extracted from verified JWT
┌─────────────┐     OAuth      ┌──────────────┐
│   User      │───────────────▶│  Google OAuth │
└─────────────┘                 └──────┬───────┘
                                      │
                                      │ callback
                                      ▼
                               ┌──────────────┐
                               │  better-auth │
                               │   Session    │
                               └──────┬───────┘
                                      │
                                      │ JWT Generation
                                      ▼
┌─────────────┐   Bearer JWT   ┌──────────────┐
│   Next.js   │───────────────▶│ DSPy Backend  │
│  Frontend   │                │  (Verified)   │
└─────────────┘                └───────────────┘

🗄️ Database Schema

Core Tables

conversations (Research Sessions)

- id: varchar(128) PK (nanoid)
- user_id: text FK → user.id
- title: text
- is_incognito: boolean
- research_context: jsonb  -- Agent context state
- created_at: timestamp
- updated_at: timestamp

messages (Research Interactions)

- id: serial PK
- conversation_id: varchar(128) FK → conversations.id
- question: text
- answer: text
- sources: jsonb  -- Paper citations and references
- agent_reasoning: jsonb  -- DSPy trace (optional)
- search_query: text
- created_at: timestamp

catalog (Tel-U Research Papers)

- id: serial PK
- title: text
- catalog_number: varchar(100)
- catalog_type: enum
- author: text
- abstract: text
- embedding: vector(1024)  -- For semantic search
- publication_year: smallint

feedback (User Feedback)

- id: serial PK
- user_id: text FK → user.id
- conversation_id: varchar(128) FK → conversations.id
- message_id: integer FK → messages.id
- rating: smallint  -- 1-5 rating
- comment: text  -- Optional feedback comment
- created_at: timestamp

🚢 Deployment

Environment Variables (Production)

# Production URLs
NEXT_PUBLIC_APP_URL=https://your-domain.com
NEXT_PUBLIC_BACKEND_URL=https://api.your-domain.com

# Production Database (Supabase/Neon/Railway with pgvector)
DATABASE_URL=postgresql://user:pass@host:5432/dbname

# Auth (Use strong secrets in production!)
BETTER_AUTH_SECRET=production-secret-min-32-chars
BETTER_AUTH_URL=https://your-domain.com
BACKEND_API_SECRET=backend-api-secret-min-32-chars

# Google OAuth (Production)
GOOGLE_CLIENT_ID=production-client-id.apps.googleusercontent.com
GOOGLE_CLIENT_SECRET=production-client-secret

# External APIs (for Cari Dosen)
VOYAGE_API_KEY=your-voyage-api-key
EXA_API_KEY=your-exa-api-key

Deployment Platforms

Vercel (Recommended for Frontend)

# Install Vercel CLI
bun install -g vercel

# Deploy
vercel --prod

Environment Variables: Set in Vercel Dashboard → Settings → Environment Variables

Docker Deployment

# Dockerfile (example)
FROM node:20-alpine AS base
WORKDIR /app
COPY package.json bun.lockb ./
RUN bun install
COPY . .
RUN bun run build
EXPOSE 3000
CMD ["bun", "start"]
docker build -t open-ta-telyu .
docker run -p 3000:3000 --env-file .env open-ta-telyu

Database Migration

# Run migrations on production
bun run db:push

# Or use Drizzle migrate
bun run db:migrate

🧪 Testing

# Run linter
bun run lint

# Format code
bun run format

# Type check (if using tsc)
tsc --noEmit

🤝 Contributing

We welcome contributions! Please follow these guidelines:

Development Workflow

  1. Fork the repository
  2. Clone your fork: git clone https://github.com/emrsyah/open-ta-telyu.git
  3. Create a branch: git checkout -b feature/your-feature-name
  4. Make your changes
  5. Test thoroughly
  6. Commit: git commit -m "feat: add your feature"
  7. Push: git push origin feature/your-feature-name
  8. Open a Pull Request

Commit Convention

Follow Conventional Commits:

  • feat: New feature
  • fix: Bug fix
  • docs: Documentation changes
  • style: Code style changes (formatting, etc.)
  • refactor: Code refactoring
  • test: Adding or updating tests
  • chore: Maintenance tasks

Code Style

  • Use Biome for linting and formatting
  • Follow TypeScript best practices
  • Write meaningful commit messages
  • Add comments for complex logic
  • Update documentation for new features

Pull Request Guidelines

  • Describe what you changed and why
  • Link to related issues
  • Ensure all checks pass
  • Request review from maintainers
  • Keep PRs focused and atomic

📝 License

This project is licensed under the MIT License - see the LICENSE file for details.

🙏 Acknowledgments

🔗 Related Repositories

📧 Contact


Built with ❤️ for the Telkom University academic community

Accelerating research through AI collaboration

⬆ Back to Top

Contributors

emrsyah

Issues