liano3/LongTutor

★ 4Forks 0PythonGitHub ↗Compare

README

LongTutor

Code and data utilities for the paper “LongTutor: Benchmarking Large Language Models for Long-term Personalized Tutoring”.

This repository contains:

  • JSONL datasets (student interaction sequences, question banks, and evaluation annotations)
  • Scripts to build history features, generate benchmark test cases, run AI-tutor inference, and compute evaluation metrics

Repository Layout

  • data/
    • MOOCRadar/ and XES3G5M/: processed datasets and intermediate artifacts
    • Common files you will see:
      • sequences.jsonl / sequences_long.jsonl: per-student interaction sequences
      • questions.jsonl: question metadata
      • history_features_lastq*.jsonl: per-sample “current question + history + features” inputs for LLMs
      • pipeline_an*.jsonl: benchmark test cases / annotations
      • human_an_updated.jsonl: human reference annotations (gold)
  • scripts/: runnable scripts (Python)

All JSONL files are one JSON object per line.

Setup

Recommended: Python 3.10+.

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

API Credentials

LLM calls use the OpenAI Python SDK, and can work with any OpenAI-compatible endpoint.

  • Required: OPENAI_API_KEY
  • Optional: OPENAI_BASE_URL (for non-OpenAI endpoints)
export OPENAI_API_KEY="..."
# export OPENAI_BASE_URL="https://your-openai-compatible-endpoint/v1"

Some scripts also support separate credentials for an “optimizer” model (see scripts/gpt_memory_diagnose.py --help).

Quickstart

1) Build history features

scripts/compute_history_stats.py can build the history_features_lastq*.jsonl inputs.

By default it runs with hard-coded paths in its __main__ block, e.g.:

  • input: data/MOOCRadar/sequences_long.jsonl
  • input: data/MOOCRadar/questions.jsonl
  • output: data/MOOCRadar/history_features_lastq.jsonl

If you want to run it on a different dataset, edit the paths in the __main__ block.

python scripts/compute_history_stats.py

For XES3G5M, please use the XES3G5M-specific concept segmentation function shown at the top of scripts/compute_history_stats.py (currently commented out). The active implementation is for MOOCRadar.

For XES3G5M, history_features_lastq.jsonl is the full history-input file. Its first 1,000 rows correspond to human_an_updated.jsonl (LongTutor-Gold), while history_features_lastq_scale.jsonl is the remaining 2,437 rows used for the scaled/synthetic evaluation.

2) Generate benchmark test cases (synthetic pipeline)

scripts/gpt_memory_diagnose.py generates benchmark items (memory queries + diagnosis/strategy + teaching content) from history features.

python scripts/gpt_memory_diagnose.py \
	--input data/MOOCRadar/history_features_lastq.jsonl \
	--output data/MOOCRadar/pipeline_an.jsonl \
	--model gpt-5

Notes:

  • --resume True (default) appends and skips finished keys.
  • --optimize enables an additional “optimizer” pass.

3) Run AI Tutor inference

scripts/eval_ai_tutor.py runs a tutor model against the test cases.

For the 1,000 human-annotated LongTutor-Gold samples:

python scripts/eval_ai_tutor.py \
	--input data/XES3G5M/history_features_lastq.jsonl \
	--tests data/XES3G5M/human_an_updated.jsonl \
	--output data/XES3G5M/preds.jsonl \
	--model gpt-5 \
	--temperature 0 \
	--workers 4

4) Compute evaluation metrics

scripts/compute_ai_tutor_eval_metrics.py compares predictions vs. gold annotations.

python scripts/compute_ai_tutor_eval_metrics.py \
	--pred data/XES3G5M/preds.jsonl \
	--gold data/XES3G5M/human_an_updated.jsonl \
	--out data/XES3G5M/metrics.jsonl \
	--model gpt-5 \
	--temperature 0

The script prints a JSON summary including:

  • Memory accuracy (overall and by memory-query type)
  • Diagnosis accuracy and macro-F1
  • Teaching quality scores (LLM-graded) and ROUGE-L

Other Utilities

  • scripts/data_processing.py: dataset conversion / filtering helpers (see examples in its __main__ block)
  • scripts/openai_helper.py: OpenAI-compatible chat helper with retries and JSON extraction

Reproducibility Notes

  • Many scripts default to --resume True and append to existing output JSONL files.
  • When re-running, delete the existing output file if you want a clean run.
  • XES3G5M uses sequences of length 200: the final item is the current question and the previous 199 items are the history input. The L=100 statement in the paper is a typo/oversight; please follow the released scripts as the authoritative setting.

License

See the paper/project for licensing and dataset terms.

Contributors

liano3

Issues