Code and data utilities for the paper “LongTutor: Benchmarking Large Language Models for Long-term Personalized Tutoring”.
This repository contains:
- JSONL datasets (student interaction sequences, question banks, and evaluation annotations)
- Scripts to build history features, generate benchmark test cases, run AI-tutor inference, and compute evaluation metrics
data/MOOCRadar/andXES3G5M/: processed datasets and intermediate artifacts- Common files you will see:
sequences.jsonl/sequences_long.jsonl: per-student interaction sequencesquestions.jsonl: question metadatahistory_features_lastq*.jsonl: per-sample “current question + history + features” inputs for LLMspipeline_an*.jsonl: benchmark test cases / annotationshuman_an_updated.jsonl: human reference annotations (gold)
scripts/: runnable scripts (Python)
All JSONL files are one JSON object per line.
Recommended: Python 3.10+.
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtLLM calls use the OpenAI Python SDK, and can work with any OpenAI-compatible endpoint.
- Required:
OPENAI_API_KEY - Optional:
OPENAI_BASE_URL(for non-OpenAI endpoints)
export OPENAI_API_KEY="..."
# export OPENAI_BASE_URL="https://your-openai-compatible-endpoint/v1"Some scripts also support separate credentials for an “optimizer” model (see scripts/gpt_memory_diagnose.py --help).
scripts/compute_history_stats.py can build the history_features_lastq*.jsonl inputs.
By default it runs with hard-coded paths in its __main__ block, e.g.:
- input:
data/MOOCRadar/sequences_long.jsonl - input:
data/MOOCRadar/questions.jsonl - output:
data/MOOCRadar/history_features_lastq.jsonl
If you want to run it on a different dataset, edit the paths in the __main__ block.
python scripts/compute_history_stats.pyFor XES3G5M, please use the XES3G5M-specific concept segmentation function shown at the top of scripts/compute_history_stats.py (currently commented out). The active implementation is for MOOCRadar.
For XES3G5M, history_features_lastq.jsonl is the full history-input file. Its first 1,000 rows correspond to human_an_updated.jsonl (LongTutor-Gold), while history_features_lastq_scale.jsonl is the remaining 2,437 rows used for the scaled/synthetic evaluation.
scripts/gpt_memory_diagnose.py generates benchmark items (memory queries + diagnosis/strategy + teaching content) from history features.
python scripts/gpt_memory_diagnose.py \
--input data/MOOCRadar/history_features_lastq.jsonl \
--output data/MOOCRadar/pipeline_an.jsonl \
--model gpt-5Notes:
--resume True(default) appends and skips finished keys.--optimizeenables an additional “optimizer” pass.
scripts/eval_ai_tutor.py runs a tutor model against the test cases.
For the 1,000 human-annotated LongTutor-Gold samples:
python scripts/eval_ai_tutor.py \
--input data/XES3G5M/history_features_lastq.jsonl \
--tests data/XES3G5M/human_an_updated.jsonl \
--output data/XES3G5M/preds.jsonl \
--model gpt-5 \
--temperature 0 \
--workers 4scripts/compute_ai_tutor_eval_metrics.py compares predictions vs. gold annotations.
python scripts/compute_ai_tutor_eval_metrics.py \
--pred data/XES3G5M/preds.jsonl \
--gold data/XES3G5M/human_an_updated.jsonl \
--out data/XES3G5M/metrics.jsonl \
--model gpt-5 \
--temperature 0The script prints a JSON summary including:
- Memory accuracy (overall and by memory-query type)
- Diagnosis accuracy and macro-F1
- Teaching quality scores (LLM-graded) and ROUGE-L
scripts/data_processing.py: dataset conversion / filtering helpers (see examples in its__main__block)scripts/openai_helper.py: OpenAI-compatible chat helper with retries and JSON extraction
- Many scripts default to
--resume Trueand append to existing output JSONL files. - When re-running, delete the existing output file if you want a clean run.
- XES3G5M uses sequences of length 200: the final item is the current question and the previous 199 items are the history input. The
L=100statement in the paper is a typo/oversight; please follow the released scripts as the authoritative setting.
See the paper/project for licensing and dataset terms.