-
GERTuraX is a series of pretrained encoder-only language models for German.
-
The models are ELECTRA-based and pretrained with the TEAMS approach on the CulturaX corpus.
-
In total, three different models were trained and released with pretraining corpus sizes ranging from 147GB to 1.1TB.
This repository hosts all necessary code to conduct the GERTuraX fine-tuning experiments on various downstream tasks using the awesome Flair library.
- 18.07.2025: Add new results for BarNER dataset.
- 12.07.2025: Add new results for German CoNLL-2003 Original.
- 20.06.2025: Add new results for ModernGBERT and GeistBERT.
- 09.02.2025: Add initial version of this repository.
First, Flair and other dependencies must be installed:
$ pip3 install -r requirements.txtThe following environment variables can be set:
| Variable | Required | Description |
|---|---|---|
CONFIG |
✔️ | Path to JSON-based configuration file, e.g. configs/germeval14/gbert_base.json. |
HUB_ORG_NAME |
✖️ | Organization/User name on Hugging Face Model Hub. Must be set for model upload. |
HF_UPLOAD |
✖️ | Defines if model should be uploaded to Model Hub or not. Disabled by default. |
Here's an example for the used JSON-based configuration format:
{
"batch_sizes": [
32,
16
],
"learning_rates": [
1e-05,
2e-05,
3e-05,
4e-05
],
"epochs": [
20
],
"context_sizes": [
0
],
"seeds": [
1,
2,
3,
4,
5
],
"layers": "-1",
"subword_poolings": [
"first"
],
"use_crf": false,
"use_tensorboard": true,
"hf_model": "deepset/gbert-base",
"model_short_name": "gbert_base",
"task": "ner/germeval14",
"cuda": "0"
}Hyper-parameter searches are possible, e.g. different batch sizes, learning rates, epochs or seeds can be set. The CUDA device id can be set via cuda (expecting the id as string).
After environment variables are set, the fine-tuning can be started with:
$ python3 script.pyThe flair-log-parser.py script can be used to get an overview of best configurations and their correspondig F1-Scores.
GERTuraX and other German Language Models were fine-tuned on GermEval 2014 (NER), GermEval 2018 (Sentiment analysis), CoNLL-2003 (NER) and BarNER (NER).
We use the same hyper-parameters for GermEval 2014, GermEval 2018, CoNLL-2003 and BarNER as used in the GeBERTa paper (cf. Table 5) using 5 runs with different seed and report the averaged score, conducted with the awesome Flair library.
| Model Name | Avg. Development F1-Score | Avg. Test F1-Score |
|---|---|---|
| GBERT Base | 87.53 ± 0.22 | 86.81 ± 0.16 |
| GERTuraX-1 (147GB) | 88.32 ± 0.21 | 87.18 ± 0.12 |
| GERTuraX-2 (486GB) | 88.58 ± 0.32 | 87.58 ± 0.15 |
| GERTuraX-3 (1.1TB) | 88.90 ± 0.06 | 87.84 ± 0.18 |
| GeBERTa Base | 88.79 ± 0.16 | 88.03 ± 0.16 |
| ModernGBERT 134M | 87.86 ± 0.29 | 86.79 ± 0.29 |
| GeistBERT Base | 88.48 ± 0.34 | 87.67 ± 0.26 |
GermEval 2014 - Without Wikipedia
| Model Name | Avg. Development F1-Score | Avg. Test F1-Score |
|---|---|---|
| GBERT Base | 90.48 ± 0.34 | 89.05 ± 0.21 |
| GERTuraX-1 (147GB) | 91.27 ± 0.11 | 89.73 ± 0.27 |
| GERTuraX-2 (486GB) | 91.70 ± 0.28 | 89.98 ± 0.22 |
| GERTuraX-3 (1.1TB) | 91.75 ± 0.17 | 90.24 ± 0.27 |
| GeBERTa Base | 91.74 ± 0.23 | 90.28 ± 0.21 |
| ModernGBERT 134M | 90.64 ± 0.21 | 89.13 ± 0.31 |
| GeistBERT Base | 90.88 ± 0.31 | 90.14 ± 0.31 |
| Model Name | Avg. Development F1-Score | Avg. Test F1-Score |
|---|---|---|
| GBERT Base | 63.66 ± 4.08 | 51.86 ± 1.31 |
| GERTuraX-1 (147GB) | 62.87 ± 1.95 | 50.61 ± 0.36 |
| GERTuraX-2 (486GB) | 64.37 ± 1.31 | 51.02 ± 0.90 |
| GERTuraX-3 (1.1TB) | 66.39 ± 0.85 | 49.94 ± 2.06 |
| GeBERTa Base | 65.81 ± 3.29 | 52.45 ± 0.57 |
| ModernGBERT 134M | 59.69 ± 2.12 | 48.75 ± 3.33 |
| GeistBERT Base | 64.84 ± 1.59 | 53.47 ± 1.12 |
| Model Name | Avg. Development F1-Score | Avg. Test F1-Score |
|---|---|---|
| GBERT Base | 83.15 ± 1.83 | 76.39 ± 0.64 |
| GERTuraX-1 (147GB) | 83.72 ± 0.68 | 77.11 ± 0.59 |
| GERTuraX-2 (486GB) | 84.51 ± 0.88 | 78.07 ± 0.91 |
| GERTuraX-3 (1.1TB) | 84.33 ± 1.48 | 78.44 ± 0.74 |
| GeBERTa Base | 83.54 ± 1.27 | 78.36 ± 0.79 |
| ModernGBERT 134M | 83.16 ± 2.05 | 76.01 ± 0.89 |
| GeistBERT Base | 83.77 ± 0.89 | 77.81 ± 0.98 |
| Model Name | Avg. Development F1-Score | Avg. Test F1-Score |
|---|---|---|
| GBERT Base | 88.33 ± 0.24 | 85.13 ± 0.36 |
| GERTuraX-1 (147GB) | 88.56 ± 0.07 | 86.00 ± 0.51 |
| GERTuraX-2 (486GB) | 88.81 ± 0.12 | 86.28 ± 0.29 |
| GERTuraX-3 (1.1TB) | 88.88 ± 0.13 | 86.03 ± 0.06 |
| GeBERTa Base | 88.51 ± 0.06 | 86.21 ± 0.75 |
| ModernGBERT 134M | 87.42 ± 0.18 | 84.37 ± 0.37 |
| GeistBERT Base | 88.35 ± 0.14 | 85.94 ± 0.25 |
| Model Name | Avg. Development F1-Score | Avg. Test F1-Score |
|---|---|---|
| GBERT Base | 92.15 ± 0.10 | 88.73 ± 0.21 |
| GERTuraX-1 (147GB) | 92.32 ± 0.14 | 90.09 ± 0.12 |
| GERTuraX-2 (486GB) | 92.75 ± 0.20 | 90.15 ± 0.14 |
| GERTuraX-3 (1.1TB) | 92.77 ± 0.28 | 90.83 ± 0.16 |
| GeBERTa Base | 92.87 ± 0.21 | 90.94 ± 0.24 |
| ModernGBERT 134M | 91.49 ± 0.15 | 89.64 ± 0.29 |
| GeistBERT Base | 92.55 ± 0.11 | 90.33 ± 0.20 |
| Model Name | Avg. Development F1-Score | Avg. Test F1-Score |
|---|---|---|
| GBERT Base | 69.32 ± 1.60 | 69.62 ± 2.81 |
| GERTuraX-1 (147GB) | 70.04 ± 0.87 | 71.81 ± 1.14 |
| GERTuraX-2 (486GB) | 71.57 ± 0.66 | 68.08 ± 1.84 |
| GERTuraX-3 (1.1TB) | 72.00 ± 1.01 | 69.94 ± 1.50 |
| GeBERTa Base | 72.25 ± 0.63 | 69.07 ± 2.06 |
| ModernGBERT 134M | 50.50 ± 2.39 | 60.62 ± 1.30 |
| GeistBERT Base | 71.78 ± 1.46 | 71.12 ± 1.20 |
GERTuraX is the outcome of the last 12 months of working with TPUs from the awesome TRC program and the TensorFlow Model Garden library.
Many thanks for providing TPUs!
Made from Bavarian Oberland with ❤️ and 🥨.