Extending the moral-emotion classification framework of Kim et al. (2024) to Modern Hebrew parliamentary speech, using the IsraParlTweet corpus of Knesset transcripts.
Final project for the Natural Language Processing course, Reichman University — Tom Ron & Liron Zarhay.
The full report (LaTeX + PDF) is not part of this public repo, since its title page carries course-required student ID numbers; this README summarizes the methodology and results.
- RQ1 — Can Haidt's six-category moral-emotion taxonomy be reliably applied to Hebrew parliamentary text by a few-shot LLM annotator, and does agreement with human raters depend on whether the LLM is Hebrew-sovereign or a Western multilingual generalist?
- RQ2 — Can a Hebrew transformer, fine-tuned on LLM-generated silver labels, classify moral emotion in Hebrew parliamentary sentences well enough to match the LLM teacher it was distilled from?
| Human inter-annotator agreement (multiclass Cohen's κ) | 0.760 (vs. Kim et al.'s Korean 0.72, English 0.42) |
| LLM–human agreement: DictaLM (Hebrew-sovereign) vs. Gemma 4 (Western generalist) | κ = 0.417 vs. 0.350 |
| Fine-tuned classifier macro-F1 on gold set (best student vs. LLM teacher) | 0.349 vs. 0.426 |
DictaLM tracks human judgment on Hebrew moral-emotion labeling more closely than Gemma 4 does — an empirical demonstration of the Western-centric LLM bias that Kim et al. only raise as a caveat. None of the three fine-tuned Hebrew transformers (DictaBERT, AlephBERT, Knesset-DictaBERT) surpass the unfine-tuned DictaLM teacher, a genuine negative result traced to sparse silver-label coverage of the rarer categories rather than to the fine-tuning procedure itself. Full tables, figures, and discussion are in the project report (not published in this repo — see License below).
notebooks/ 01–05, run in order (see "Reproducing" below)
src/ shared library code imported by every notebook
labels.py category taxonomy + label parsing/normalization
io.py canonical data loaders
iaa.py inter-annotator agreement metrics (kappa, alpha, pi)
metrics.py classifier evaluation metrics (F1, bootstrap CIs)
annotate.py few-shot LLM annotation via Ollama (prompt + caching)
scripts/
build_sample.py builds the sentence pool + gold set from raw IsraParlTweet data
run_llm_annotation.py CLI driver for src.annotate, used for the background LLM labeling runs
data/ CSV/XLSX data products and figures (raw corpus is gitignored, see below)
(report/, the LaTeX source and compiled PDF of the final report, exists locally but is gitignored --
see License below.)
Requires Python ≥3.13 and uv.
uv syncLLM silver-annotation runs two local models via Ollama:
ollama pull hf.co/dicta-il/DictaLM-3.0-Nemotron-12B-Instruct-GGUF:Q8_0
ollama pull gemma4:31bFine-tuning (notebook 04) uses PyTorch/Transformers on whatever accelerator is available (MPS on Apple Silicon, CUDA, or CPU).
data/raw/ (the IsraParlTweet Knesset-speech corpus: knesset_speeches_v1.1.csv,
knesset_sentences.json, metadata.csv) is gitignored — it's a large, separately-licensed
(CC-BY-4.0) dataset, not redistributed in this repo. Download it from
IsraParlTweet on HuggingFace and place it
under data/raw/ before running scripts/build_sample.py.
Everything downstream of that — the sampled pools, the human-annotated gold set, LLM silver labels, and
all figures — is checked into data/ and is what the notebooks actually run against, so the full analysis
(notebooks 01–05) is reproducible without re-downloading the raw corpus.
scripts/build_sample.py— samplessample_pool_3000.csv/sample_pool_3250.csvandgold_250.csvfromdata/raw/(already done; outputs are committed).notebooks/01_data_and_iaa.ipynb— human inter-annotator agreement, adjudication, final gold set.notebooks/02_llm_annotation.ipynb— few-shot LLM silver labeling (both model arms).notebooks/03_llm_vs_human.ipynb— RQ1: LLM-vs-human agreement.notebooks/04_finetune.ipynb— fine-tunes three Hebrew BERT variants on the silver labels.notebooks/05_eval_and_figures.ipynb— RQ2: evaluation on the gold set, report figures/tables.
All notebooks are self-contained (sys.path + from src import ...) and read/write ../data/. LLM
responses and model checkpoints are cache-backed (cache/, models/, both gitignored) so re-running is
cheap once populated locally.
Code is provided as-is for academic purposes. The underlying IsraParlTweet corpus is CC-BY-4.0
(Mor-Lan et al., 2024) and is not redistributed here — see the Data section above. The final report
(report/) is gitignored rather than published in this repo, since its title page carries student ID
numbers required for course submission.