tomron87/NLP_Final_Project

★ 0Forks 0Jupyter NotebookGitHub ↗Compare

README

Moral Emotion Classification in Hebrew Parliamentary Discourse

Extending the moral-emotion classification framework of Kim et al. (2024) to Modern Hebrew parliamentary speech, using the IsraParlTweet corpus of Knesset transcripts.

Final project for the Natural Language Processing course, Reichman University — Tom Ron & Liron Zarhay.

The full report (LaTeX + PDF) is not part of this public repo, since its title page carries course-required student ID numbers; this README summarizes the methodology and results.

Research questions

  • RQ1 — Can Haidt's six-category moral-emotion taxonomy be reliably applied to Hebrew parliamentary text by a few-shot LLM annotator, and does agreement with human raters depend on whether the LLM is Hebrew-sovereign or a Western multilingual generalist?
  • RQ2 — Can a Hebrew transformer, fine-tuned on LLM-generated silver labels, classify moral emotion in Hebrew parliamentary sentences well enough to match the LLM teacher it was distilled from?

Headline results

Human inter-annotator agreement (multiclass Cohen's κ) 0.760 (vs. Kim et al.'s Korean 0.72, English 0.42)
LLM–human agreement: DictaLM (Hebrew-sovereign) vs. Gemma 4 (Western generalist) κ = 0.417 vs. 0.350
Fine-tuned classifier macro-F1 on gold set (best student vs. LLM teacher) 0.349 vs. 0.426

DictaLM tracks human judgment on Hebrew moral-emotion labeling more closely than Gemma 4 does — an empirical demonstration of the Western-centric LLM bias that Kim et al. only raise as a caveat. None of the three fine-tuned Hebrew transformers (DictaBERT, AlephBERT, Knesset-DictaBERT) surpass the unfine-tuned DictaLM teacher, a genuine negative result traced to sparse silver-label coverage of the rarer categories rather than to the fine-tuning procedure itself. Full tables, figures, and discussion are in the project report (not published in this repo — see License below).

Repository layout

notebooks/          01–05, run in order (see "Reproducing" below)
src/                 shared library code imported by every notebook
  labels.py            category taxonomy + label parsing/normalization
  io.py                canonical data loaders
  iaa.py               inter-annotator agreement metrics (kappa, alpha, pi)
  metrics.py           classifier evaluation metrics (F1, bootstrap CIs)
  annotate.py          few-shot LLM annotation via Ollama (prompt + caching)
scripts/
  build_sample.py       builds the sentence pool + gold set from raw IsraParlTweet data
  run_llm_annotation.py CLI driver for src.annotate, used for the background LLM labeling runs
data/                CSV/XLSX data products and figures (raw corpus is gitignored, see below)

(report/, the LaTeX source and compiled PDF of the final report, exists locally but is gitignored -- see License below.)

Setup

Requires Python ≥3.13 and uv.

uv sync

LLM silver-annotation runs two local models via Ollama:

ollama pull hf.co/dicta-il/DictaLM-3.0-Nemotron-12B-Instruct-GGUF:Q8_0
ollama pull gemma4:31b

Fine-tuning (notebook 04) uses PyTorch/Transformers on whatever accelerator is available (MPS on Apple Silicon, CUDA, or CPU).

Data

data/raw/ (the IsraParlTweet Knesset-speech corpus: knesset_speeches_v1.1.csv, knesset_sentences.json, metadata.csv) is gitignored — it's a large, separately-licensed (CC-BY-4.0) dataset, not redistributed in this repo. Download it from IsraParlTweet on HuggingFace and place it under data/raw/ before running scripts/build_sample.py.

Everything downstream of that — the sampled pools, the human-annotated gold set, LLM silver labels, and all figures — is checked into data/ and is what the notebooks actually run against, so the full analysis (notebooks 01–05) is reproducible without re-downloading the raw corpus.

Reproducing the pipeline

  1. scripts/build_sample.py — samples sample_pool_3000.csv / sample_pool_3250.csv and gold_250.csv from data/raw/ (already done; outputs are committed).
  2. notebooks/01_data_and_iaa.ipynb — human inter-annotator agreement, adjudication, final gold set.
  3. notebooks/02_llm_annotation.ipynb — few-shot LLM silver labeling (both model arms).
  4. notebooks/03_llm_vs_human.ipynb — RQ1: LLM-vs-human agreement.
  5. notebooks/04_finetune.ipynb — fine-tunes three Hebrew BERT variants on the silver labels.
  6. notebooks/05_eval_and_figures.ipynb — RQ2: evaluation on the gold set, report figures/tables.

All notebooks are self-contained (sys.path + from src import ...) and read/write ../data/. LLM responses and model checkpoints are cache-backed (cache/, models/, both gitignored) so re-running is cheap once populated locally.

License

Code is provided as-is for academic purposes. The underlying IsraParlTweet corpus is CC-BY-4.0 (Mor-Lan et al., 2024) and is not redistributed here — see the Data section above. The final report (report/) is gitignored rather than published in this repo, since its title page carries student ID numbers required for course submission.

Contributors

tomron87

Issues