Instructions to use arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER
How to use arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1 with GLiNER:
from gliner import GLiNER model = GLiNER.from_pretrained("arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1") text = "Cristiano Ronaldo dos Santos Aveiro was born on 5 February 1985 in Funchal, Madeira, Portugal." labels = ["person", "date", "location"] entities = model.predict_entities(text, labels) for entity in entities: print(entity["text"], "=>", entity["label"]) - Notebooks
- Google Colab
- Kaggle
gliner-ettin-68m-ptbr-pii-full-3x-v1
GLiNER fine-tune of jhu-clsp/ettin-encoder-68m (~68M params) for Brazilian-Portuguese PII detection (CPF, RG, PIS, names, addresses, LGPD-sensitive descriptions), with cross-source generalization tested on English/multilingual PII and spam corpora. Best ettin model of the family: 0.7076 strict / 0.7979 partial F1 on the PT-BR holdout source at step 41,400. The mmBERT-small full-3x model is stronger on every holdout source under the same protocol.
It belongs to a sweep of GLiNER models trained on the same MI300X GPU with the same loss recipe, differing in encoder, training subset, batch size and step budget; the family comparison below puts them side by side.
Related repos
| Repo | Type | Role |
|---|---|---|
| arthrod/gliner-ettin-32m-ptbr-pii-full-3x-v1 | model | GLiNER on jhu-clsp/ettin-encoder-32m, full ~984k-row train set, 69,000 steps. Smallest model in the family. |
| arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1 (this repo) | model | GLiNER on jhu-clsp/ettin-encoder-68m, full ~984k-row train set, 69,000 steps (trainer state uploaded through step 58,650). |
| arthrod/gliner-ettin-68m-ptbr-pii-top50k-v1 | model | GLiNER on jhu-clsp/ettin-encoder-68m, 50k curated subset, 11,700 steps. Ablation: shows the curated subset under-generalizes. |
| arthrod/gliner-mmbert-small-ptbr-pii-full-3x-v1 | model | GLiNER on jhu-clsp/mmBERT-small, full train set + 2k Ritápolis rows, 69,000 steps. Strongest GLiNER model of the family. |
| arthrod/gliner-opf-ptbr-pii-v1 | model | Fine-tune of OpenAI Privacy Filter (not GLiNER despite the name). Best partial F1 on the in-distribution val set and by far the fewest false positives on no-PII text. |
| arthrod/gliner-opf-ptbr-pii-bench-v1 | dataset | Benchmark: predictions and metrics for the mmBERT-small and Privacy Filter models (plus two unreleased ettin variants) on the PT-BR val set, held-out sources and no-PII spam/phishing text. |
| arthrod/gliner-opf-ptbr-pii-demo | space | Gradio demo running the Privacy Filter fine-tune. |
arthrod/gliner_eval_folder (private) |
dataset | Earlier (Feb 2026) sweep: 18 GLiNER models x 9 eval sets. |
| arthrod/gliner_review_comparison | dataset | Earlier GLiNER PII review/comparison data. |
Quick start
The repo root holds checkpoint-41400 (the released, best checkpoint; its pytorch_model.bin is byte-identical to checkpoint-41400/pytorch_model.bin), so the plain repo id loads it:
# pip install "gliner==0.2.25" # version the model was trained with
from gliner import GLiNER
model = GLiNER.from_pretrained("arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1")
text = (
"Sou Maria Silva, moro na Rua das Flores, 123, em São Paulo. "
"Meu CPF é 123.456.789-09 e meu telefone é (11) 91234-5678."
)
labels = [
"first name", "last name", "middle name",
"cpf document number", "rg document number", "pis document number",
"phone number", "email address", "credit card",
"location street", "location building number", "location neighborhood",
"location city", "location state", "location state abbreviation",
"location zip", "location full address",
"dob",
]
for ent in model.predict_entities(text, labels, threshold=0.3):
print(ent["text"], "→", ent["label"], f"({ent['score']:.2f})")
To load any other checkpoint, download just that folder (skip the optimizer state) and pass the local path:
from huggingface_hub import snapshot_download
path = snapshot_download("arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1", allow_patterns=["checkpoint-13800/*"],
ignore_patterns=["*/optimizer.pt", "*/scheduler.pt", "*/rng_state.pth"])
model = GLiNER.from_pretrained(f"{path}/checkpoint-13800")
Files and checkpoint layout
| path | content |
|---|---|
gliner_config.json, pytorch_model.bin, tokenizer.json, tokenizer_config.json |
released model = checkpoint-41400 |
checkpoint-N/ (17 folders) |
full Trainer checkpoint every 3,450 steps: gliner_config.json, pytorch_model.bin, tokenizer, trainer_state.json, plus optimizer.pt, scheduler.pt, rng_state.pth for resuming |
eval_holdout_strict_f1.parquet |
per-evaluation strict F1 on all 11 holdout sources, extracted from checkpoint-58650/trainer_state.json (see Evaluation) |
Uploaded checkpoint steps: 3450, 6900, 10350, 13800, 17250, 20700, 24150, 27600, 31050, 34500, 37950, 41400, 44850, 48300, 51750, 55200, 58650.
The run went to step 69,000, but checkpoints 62100, 65550 and 69000 were not uploaded; the last trainer state in the repo is checkpoint-58650.
Training data
- Train:
data/splits/train.jsonl(~984k samples, full train set). Not published; it is assembled from the same upstream PII and spam sources as the evaluation holdout, with the holdout sliced off before the train split was built (no leakage). - Validation (eval loss):
data/splits/val_5k.jsonl, 5,000 rows. - Holdout for model selection:
data/holdout_5k_original, 5,000 rows, proportional across 11 sources (list under Evaluation). Public upstream sources include nvidia/Nemotron-PII, ai4privacy/open-pii-masking-500k-ai4privacy and ai4privacy/pii-masking-400k; the PT-BR source (gliner2_pii_ptbr_reward_split) is synthetic and private. - Benchmark on the shared PT-BR val set with predictions: arthrod/gliner-opf-ptbr-pii-bench-v1 (this model is not in it).
Training recipe
From gliner_config.json and checkpoint-58650/trainer_state.json, plus the original run notes:
| setting | value |
|---|---|
| base encoder | jhu-clsp/ettin-encoder-68m (ModernBERT, 19 layers, hidden 512, 8 heads, vocab 50,371) |
| GLiNER head | span_mode: token_level, max_len: 1024, max_width: 100, max_types: 100, subtoken_pooling: first, dropout: 0.3, fine_tune: true |
| steps | 69,000 (max_steps in trainer state), eval and save every 3,450 steps |
| batch size | train 128, eval 128 |
| epochs over the train set | ~9 (7.62 at step 58,650, the last uploaded state) |
| learning rate | encoder 1.5e-05, other params 5e-5 (peak logged LR 5e-5) |
| schedule | cosine, warmup ratio 0.1 (logged LR peaks at step 6925) |
| weight decay | 0.01 (encoder and others), max grad norm 1.0 |
| loss | focal, alpha 0.75, gamma 2.0, reduction mean |
| precision | bf16, HIPBLASLT_ALLOW_TF32=0 (avoids MI300X NaN intermittency) |
| hardware | single AMD MI300X |
| library | gliner==0.2.25 plus a local fix for the upstream token_level regression that landed in 0.2.26 |
Evaluation
Protocol
Every 3,450 steps the trainer re-scores a 5,000-sample, proportional, source-mixed holdout (data/holdout_5k_original). For each row it builds a per-source label superset (the 24 canonical PT-BR labels ∪ that row's gold labels, capped at 100, lowercased) and calls model.inference(..., flat_ner=True, multi_label=False, threshold=0.3). Per-source F1 is computed by nervaluate (strict / exact / partial / ent_type). Model selection: best PT-BR partial F1.
Eleven sources cover PT-BR PII, EN PII, and negative-evidence spam/phishing:
gliner2_pii_ptbr_reward_split: PT-BR PII (the target)nemotron_pii: nvidia/Nemotron-PII, ENopen_pii_masking_500k,pii_masking_400k: multilingual PII (ai4privacy)enron_spam_bvk,enron_spam_setfit,phishing_darkknight,phishing_zefang,sms_spam_multilingual,spam_messages_mshenoda,spamassassin: negative evidence (no gold PII). Their F1 is 0 by construction; it says nothing about false positives. For false-positive counts on no-PII text see the benchmark dataset (covers the mmBERT and Privacy Filter models only).
Provenance. Strict F1 values are in this repo (checkpoint-58650/trainer_state.json, extracted to eval_holdout_strict_f1.parquet) and were checked against the tables below. Partial F1 and the per-entity breakdown come from the original run's evaluation logs, which are not stored in this repo; they are reproduced from the April 2026 card unchanged.
Released checkpoint (checkpoint-41400): per-source F1
| eval source | strict F1 | partial F1 |
|---|---|---|
| PT-BR (gliner2 reward-split) | 0.7076 | 0.7979 |
| EN (nvidia/Nemotron-PII) | 0.6149 | 0.6799 |
| open-pii-masking-500k | 0.4902 | 0.6464 |
| pii-masking-400k | 0.4701 | 0.6024 |
| enron-spam (bvk) | 0.0000 | 0.0000 |
| enron-spam (setfit) | 0.0000 | 0.0000 |
| phishing (darkknight) | 0.0000 | 0.0000 |
| phishing (zefang) | 0.0000 | 0.0000 |
| SMS spam (multilingual) | 0.0000 | 0.0000 |
| spam-messages (mshenoda) | 0.0000 | 0.0000 |
| spamassassin | 0.0000 | 0.0000 |
Per-entity breakdown (4 PII sources)
Averaged only across sources that contain the label (sources where the label is absent in gold are excluded — including them would falsely deflate the aggregate, e.g. personal description of ethnicity averaged with two sources that lack it would drop from ~0.58 to ~0.29 partial F1). n_src is the number of contributing sources. Sorted by partial F1.
| entity | n_src | strict P | strict R | strict F1 | partial P | partial R | partial F1 |
|---|---|---|---|---|---|---|---|
| ipv4 | 1 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
| email address | 4 | 0.975 | 0.970 | 0.972 | 0.982 | 0.977 | 0.980 |
| device identifier | 1 | 1.000 | 0.955 | 0.977 | 1.000 | 0.955 | 0.977 |
| url | 1 | 0.915 | 0.896 | 0.905 | 0.957 | 0.938 | 0.947 |
| pis document number | 1 | 0.974 | 0.881 | 0.925 | 0.974 | 0.881 | 0.925 |
| rg document number | 1 | 0.869 | 0.960 | 0.912 | 0.869 | 0.960 | 0.912 |
| phone number | 4 | 0.864 | 0.926 | 0.892 | 0.873 | 0.936 | 0.901 |
| date time | 1 | 0.952 | 0.855 | 0.901 | 0.952 | 0.855 | 0.901 |
| biometric identifier | 1 | 0.944 | 0.838 | 0.887 | 0.944 | 0.838 | 0.887 |
| date | 2 | 0.817 | 0.924 | 0.860 | 0.833 | 0.941 | 0.876 |
| mac address | 1 | 0.960 | 0.800 | 0.873 | 0.960 | 0.800 | 0.873 |
| coordinate | 1 | 0.746 | 0.800 | 0.772 | 0.831 | 0.891 | 0.860 |
| bank routing number | 1 | 0.870 | 0.839 | 0.855 | 0.870 | 0.839 | 0.855 |
| medical record number | 1 | 0.760 | 0.973 | 0.854 | 0.760 | 0.973 | 0.854 |
| cpf document number | 1 | 0.828 | 0.837 | 0.832 | 0.828 | 0.837 | 0.832 |
| county | 1 | 0.727 | 0.899 | 0.804 | 0.745 | 0.921 | 0.824 |
| vehicle identifier | 1 | 0.698 | 1.000 | 0.822 | 0.698 | 1.000 | 0.822 |
| location full address | 1 | 0.736 | 0.709 | 0.722 | 0.835 | 0.805 | 0.819 |
| health plan beneficiary number | 1 | 0.716 | 0.914 | 0.803 | 0.730 | 0.931 | 0.818 |
| fax number | 1 | 0.964 | 0.692 | 0.806 | 0.964 | 0.692 | 0.806 |
| country | 1 | 0.691 | 0.938 | 0.795 | 0.694 | 0.942 | 0.799 |
| time | 2 | 0.692 | 0.893 | 0.768 | 0.708 | 0.918 | 0.787 |
| swift bic | 1 | 0.625 | 1.000 | 0.769 | 0.625 | 1.000 | 0.769 |
| credit card | 4 | 0.818 | 0.733 | 0.760 | 0.824 | 0.738 | 0.765 |
| customer id | 1 | 0.663 | 0.887 | 0.759 | 0.663 | 0.887 | 0.759 |
| dob | 3 | 0.811 | 0.761 | 0.746 | 0.819 | 0.774 | 0.756 |
| pin | 1 | 0.719 | 0.793 | 0.754 | 0.719 | 0.793 | 0.754 |
| cvv | 1 | 0.632 | 0.857 | 0.727 | 0.632 | 0.857 | 0.727 |
| certificate license number | 1 | 0.652 | 0.811 | 0.723 | 0.652 | 0.811 | 0.723 |
| language | 1 | 0.571 | 0.973 | 0.720 | 0.571 | 0.973 | 0.720 |
| account number | 2 | 0.707 | 0.695 | 0.699 | 0.707 | 0.695 | 0.699 |
| blood type | 1 | 0.561 | 0.719 | 0.630 | 0.622 | 0.797 | 0.699 |
| license plate | 1 | 0.553 | 0.750 | 0.636 | 0.605 | 0.821 | 0.697 |
| personal description of organizational affiliation | 1 | 0.463 | 0.500 | 0.481 | 0.659 | 0.711 | 0.684 |
| location street | 3 | 0.566 | 0.653 | 0.603 | 0.639 | 0.740 | 0.682 |
| personal description of sexual information | 1 | 0.450 | 0.494 | 0.471 | 0.648 | 0.711 | 0.678 |
| location building number | 3 | 0.535 | 0.917 | 0.671 | 0.536 | 0.920 | 0.673 |
| employee id | 1 | 0.510 | 0.974 | 0.670 | 0.510 | 0.974 | 0.670 |
| location zip | 4 | 0.570 | 0.832 | 0.659 | 0.576 | 0.843 | 0.667 |
| location state | 2 | 0.494 | 0.905 | 0.637 | 0.496 | 0.909 | 0.639 |
| ipv6 | 1 | 0.643 | 0.409 | 0.500 | 0.821 | 0.523 | 0.639 |
| religious belief | 1 | 0.526 | 0.714 | 0.606 | 0.553 | 0.750 | 0.636 |
| last name | 4 | 0.526 | 0.726 | 0.603 | 0.552 | 0.765 | 0.635 |
| personal description of political opinion | 1 | 0.407 | 0.495 | 0.447 | 0.574 | 0.698 | 0.630 |
| company name | 1 | 0.466 | 0.672 | 0.551 | 0.533 | 0.768 | 0.629 |
| tax id number | 3 | 0.651 | 0.568 | 0.601 | 0.679 | 0.584 | 0.621 |
| personal description of religious convictions | 1 | 0.414 | 0.538 | 0.468 | 0.539 | 0.701 | 0.610 |
| sex or gender | 2 | 0.444 | 0.875 | 0.587 | 0.444 | 0.875 | 0.587 |
| personal description of ethnicity | 2 | 0.422 | 0.726 | 0.527 | 0.468 | 0.795 | 0.582 |
| first name | 4 | 0.438 | 0.830 | 0.555 | 0.456 | 0.868 | 0.579 |
| location neighborhood | 1 | 0.435 | 0.702 | 0.537 | 0.462 | 0.746 | 0.570 |
| http cookie | 1 | 0.304 | 0.241 | 0.269 | 0.630 | 0.500 | 0.558 |
| location city | 4 | 0.410 | 0.737 | 0.525 | 0.428 | 0.771 | 0.548 |
| api key | 1 | 0.387 | 0.763 | 0.513 | 0.407 | 0.803 | 0.540 |
| id card number | 2 | 0.448 | 0.660 | 0.529 | 0.453 | 0.668 | 0.535 |
| social security number | 3 | 0.665 | 0.496 | 0.530 | 0.669 | 0.498 | 0.533 |
| user name | 2 | 0.413 | 0.658 | 0.507 | 0.421 | 0.672 | 0.517 |
| age | 2 | 0.413 | 0.791 | 0.514 | 0.413 | 0.791 | 0.514 |
| personal description of medical conditions | 1 | 0.279 | 0.439 | 0.341 | 0.413 | 0.652 | 0.506 |
| password | 2 | 0.282 | 0.321 | 0.299 | 0.453 | 0.531 | 0.487 |
| middle name | 1 | 0.304 | 0.789 | 0.439 | 0.304 | 0.789 | 0.439 |
| employment status | 1 | 0.304 | 0.562 | 0.394 | 0.315 | 0.582 | 0.409 |
| driver license number | 2 | 0.605 | 0.323 | 0.368 | 0.612 | 0.330 | 0.375 |
| location state abbreviation | 1 | 0.212 | 1.000 | 0.349 | 0.212 | 1.000 | 0.349 |
| education level | 1 | 0.124 | 0.245 | 0.165 | 0.248 | 0.491 | 0.329 |
| unique id | 1 | 0.190 | 0.571 | 0.286 | 0.190 | 0.571 | 0.286 |
| passport number | 1 | 0.400 | 0.200 | 0.267 | 0.400 | 0.200 | 0.267 |
| sexuality | 1 | 0.153 | 0.667 | 0.248 | 0.153 | 0.667 | 0.248 |
| political view | 1 | 0.107 | 0.692 | 0.186 | 0.119 | 0.769 | 0.206 |
| title | 1 | 0.111 | 0.582 | 0.187 | 0.117 | 0.609 | 0.196 |
| occupation | 1 | 0.040 | 0.232 | 0.068 | 0.075 | 0.435 | 0.128 |
Reading guide. Structured tokens (email, phone, doc numbers, IPs, MAC) approach
strict ≈ partialbecause their boundaries are unambiguous. Long natural-language spans (full address, the personal description of … labels) carry a meaningful strict-vs-partial gap because exact boundaries are inherently fuzzy — even two human annotators would disagree. For those labels, partial F1 is the operationally meaningful metric.
Checkpoint sweep
"in repo": root = weights at the repo root, folder = checkpoint-N/ exists. Strict columns come from eval_holdout_strict_f1.parquet; partial columns from the original logs. Best row in bold. Steps marked * were not uploaded; their PT-BR strict F1 is from the original card only.
| step | in repo | PT-BR strict | Nemotron strict | open-pii strict | pii-masking strict | mean-4 strict | PT-BR partial | Nemotron partial | open-pii partial | pii-masking partial |
|---|---|---|---|---|---|---|---|---|---|---|
| 3450 | folder | 0.0435 | 0.0289 | 0.0724 | 0.0417 | 0.0466 | 0.3185 | 0.1783 | 0.3494 | 0.3067 |
| 6900 | folder | 0.1928 | 0.0624 | 0.1621 | 0.1041 | 0.1303 | 0.4170 | 0.2571 | 0.4157 | 0.3310 |
| 10350 | folder | 0.2774 | 0.1741 | 0.1668 | 0.1336 | 0.1880 | 0.4642 | 0.3307 | 0.4036 | 0.3302 |
| 13800 | folder | 0.3107 | 0.2034 | 0.1704 | 0.1593 | 0.2110 | 0.4872 | 0.3371 | 0.3994 | 0.3343 |
| 17250 | folder | 0.4126 | 0.2584 | 0.2046 | 0.2070 | 0.2707 | 0.6059 | 0.3688 | 0.3910 | 0.3869 |
| 20700 | folder | 0.4838 | 0.3571 | 0.2318 | 0.2295 | 0.3255 | 0.6549 | 0.4791 | 0.4424 | 0.4135 |
| 24150 | folder | 0.4802 | 0.5093 | 0.2638 | 0.2840 | 0.3843 | 0.6348 | 0.6219 | 0.4534 | 0.4592 |
| 27600 | folder | 0.5864 | 0.4761 | 0.2911 | 0.2948 | 0.4121 | 0.7220 | 0.5603 | 0.4536 | 0.4289 |
| 31050 | folder | 0.5480 | 0.4819 | 0.3188 | 0.3575 | 0.4265 | 0.6731 | 0.5524 | 0.4509 | 0.4876 |
| 34500 | folder | 0.5921 | 0.5541 | 0.3647 | 0.3709 | 0.4704 | 0.6957 | 0.6209 | 0.4843 | 0.5082 |
| 37950 | folder | 0.6329 | 0.5491 | 0.4265 | 0.4273 | 0.5090 | 0.7441 | 0.6189 | 0.5580 | 0.5634 |
| 41400 | root + folder | 0.7076 | 0.6149 | 0.4902 | 0.4701 | 0.5707 | 0.7979 | 0.6799 | 0.6464 | 0.6024 |
| 44850 | folder | 0.6747 | 0.5928 | 0.4939 | 0.4543 | 0.5539 | 0.7687 | 0.6522 | 0.6310 | 0.5855 |
| 48300 | folder | 0.6994 | 0.5679 | 0.4940 | 0.4725 | 0.5584 | 0.7859 | 0.6205 | 0.6253 | 0.5899 |
| 51750 | folder | 0.6312 | 0.6189 | 0.4918 | 0.4651 | 0.5518 | 0.7073 | 0.6697 | 0.6222 | 0.5806 |
| 55200 | folder | 0.6658 | 0.5639 | 0.5013 | 0.4783 | 0.5523 | 0.7415 | 0.6109 | 0.6391 | 0.6015 |
| 58650 | folder | 0.6408 | 0.5562 | 0.4930 | 0.4341 | 0.5310 | 0.7147 | 0.6031 | 0.6187 | 0.5514 |
| 62100 | no | 0.6765* | — | — | — | — | 0.7504 | 0.6239 | 0.6448 | 0.5808 |
| 65550 | no | 0.6724* | — | — | — | — | 0.7475 | 0.6172 | 0.6419 | 0.5827 |
| 69000 | no | 0.6701* | — | — | — | — | 0.7448 | 0.6147 | 0.6415 | 0.5845 |
eval_holdout_strict_f1.parquet
One row per evaluation (17 rows), extracted from checkpoint-58650/trainer_state.json (log_history). Load with pd.read_parquet("hf://arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1/eval_holdout_strict_f1.parquet") (pandas + huggingface_hub; polars hf:// only supports dataset/Space repos, so use hf_hub_download first there).
| column | type | description |
|---|---|---|
step |
Int32 | global training step of the evaluation |
epoch |
Float64 | training epoch at that step (fractional) |
checkpoint_dir |
String | checkpoint-N/ folder holding that step's weights in this repo, null if not uploaded |
is_repo_root |
Boolean | true for the step whose weights sit at the repo root (the released model) |
eval_loss |
Float64 | validation loss on val_5k |
pii4_mean_strict_f1 |
Float64 | mean strict F1 over the 4 PII sources (computed from the four columns below) |
avg11_strict_f1 |
Float64 | mean strict F1 over all 11 sources as logged by the trainer (the 7 spam/phishing sources are always 0, so this is not a quality metric) |
gliner2_pii_ptbr_reward_split_strict_f1 |
Float64 | nervaluate strict F1 on the gliner2_pii_ptbr_reward_split holdout source |
nemotron_pii_strict_f1 |
Float64 | nervaluate strict F1 on the nemotron_pii holdout source |
open_pii_masking_500k_strict_f1 |
Float64 | nervaluate strict F1 on the open_pii_masking_500k holdout source |
pii_masking_400k_strict_f1 |
Float64 | nervaluate strict F1 on the pii_masking_400k holdout source |
enron_spam_bvk_strict_f1 |
Float64 | nervaluate strict F1 on the enron_spam_bvk holdout source (no gold entities, always 0) |
enron_spam_setfit_strict_f1 |
Float64 | nervaluate strict F1 on the enron_spam_setfit holdout source (no gold entities, always 0) |
phishing_darkknight_strict_f1 |
Float64 | nervaluate strict F1 on the phishing_darkknight holdout source (no gold entities, always 0) |
phishing_zefang_strict_f1 |
Float64 | nervaluate strict F1 on the phishing_zefang holdout source (no gold entities, always 0) |
sms_spam_multilingual_strict_f1 |
Float64 | nervaluate strict F1 on the sms_spam_multilingual holdout source (no gold entities, always 0) |
spam_messages_mshenoda_strict_f1 |
Float64 | nervaluate strict F1 on the spam_messages_mshenoda holdout source (no gold entities, always 0) |
spamassassin_strict_f1 |
Float64 | nervaluate strict F1 on the spamassassin holdout source (no gold entities, always 0) |
Example rows (first evaluation and the released step):
[
{
"step": 3450,
"epoch": 0.448518,
"checkpoint_dir": "checkpoint-3450",
"is_repo_root": false,
"eval_loss": 0.001429,
"pii4_mean_strict_f1": 0.046608,
"avg11_strict_f1": 0.016948,
"gliner2_pii_ptbr_reward_split_strict_f1": 0.043478,
"nemotron_pii_strict_f1": 0.02889,
"open_pii_masking_500k_strict_f1": 0.072408,
"pii_masking_400k_strict_f1": 0.041657,
"enron_spam_bvk_strict_f1": 0.0,
"enron_spam_setfit_strict_f1": 0.0,
"phishing_darkknight_strict_f1": 0.0,
"phishing_zefang_strict_f1": 0.0,
"sms_spam_multilingual_strict_f1": 0.0,
"spam_messages_mshenoda_strict_f1": 0.0,
"spamassassin_strict_f1": 0.0
},
{
"step": 41400,
"epoch": 5.382215,
"checkpoint_dir": "checkpoint-41400",
"is_repo_root": true,
"eval_loss": 0.000242,
"pii4_mean_strict_f1": 0.570687,
"avg11_strict_f1": 0.207523,
"gliner2_pii_ptbr_reward_split_strict_f1": 0.707593,
"nemotron_pii_strict_f1": 0.614901,
"open_pii_masking_500k_strict_f1": 0.490177,
"pii_masking_400k_strict_f1": 0.470078,
"enron_spam_bvk_strict_f1": 0.0,
"enron_spam_setfit_strict_f1": 0.0,
"phishing_darkknight_strict_f1": 0.0,
"phishing_zefang_strict_f1": 0.0,
"sms_spam_multilingual_strict_f1": 0.0,
"spam_messages_mshenoda_strict_f1": 0.0,
"spamassassin_strict_f1": 0.0
}
]
Comparison across the family
Same holdout protocol (strict F1, released checkpoint of each repo). Source: each repo's eval_holdout_strict_f1.parquet.
| model | released step | PT-BR | Nemotron (EN) | open-pii-500k | pii-masking-400k | mean of 4 |
|---|---|---|---|---|---|---|
| ettin-32m full-3x | 65550 | 0.6326 | 0.4945 | 0.4081 | 0.4193 | 0.4886 |
| ettin-68m full-3x (this) | 41400 | 0.7076 | 0.6149 | 0.4902 | 0.4701 | 0.5707 |
| ettin-68m top-50k | 9945 | 0.2732 | 0.0345 | 0.0609 | 0.0755 | 0.1110 |
| mmBERT-small full-3x | 41400 | 0.7678 | 0.8158 | 0.7217 | 0.7076 | 0.7532 |
| Privacy Filter fine-tune | e3_c3 | not evaluated with this protocol |
Shared PT-BR val benchmark (24 canonical labels). None of the ettin repos in this family were run on it.
| model | partial P | partial R | partial F1 | exact P | exact R | exact F1 |
|---|---|---|---|---|---|---|
| Privacy Filter fine-tune e3_c3 | 0.917 | 0.879 | 0.897 | 0.871 | 0.835 | 0.853 |
| mmBERT-small full-3x (41400) | 0.951 | 0.832 | 0.888 | 0.932 | 0.815 | 0.870 |
| ettin-68m easter-egg (not public) | 0.893 | 0.761 | 0.822 | 0.857 | 0.731 | 0.789 |
| ettin-32m easter-egg (not public) | 0.905 | 0.729 | 0.808 | 0.862 | 0.694 | 0.769 |
Source: comparison_report.md in the Privacy Filter repo. 5,000-row PT-BR val set, restricted to the 24 canonical labels, nervaluate; GLiNER predictions at threshold 0.5, Privacy Filter with its default Viterbi decoding. The ettin models in this table are separate easter-egg fine-tunes, not the ettin repos of this family, which were not run on this benchmark.
Original sweep log (best by PT-BR partial F1 of each run). Reproduced from the April 2026 card; most of these runs (32m/68m top-100k, the 23k-step "full" runs, the easter-egg fine-tunes) are not published.
ettin-32m top-50k : PT-BR strict 0.1411 partial 0.3710 Nemotron partial 0.1710 cross-avg 0.2808 (step 8500)
ettin-32m top-100k : PT-BR strict 0.0946 partial 0.3165 Nemotron partial 0.1455 cross-avg 0.2428 (step 5200)
ettin-68m top-50k : PT-BR strict 0.2732 partial 0.4802 Nemotron partial 0.2397 cross-avg 0.3117 (step 9945)
ettin-68m top-100k : PT-BR strict 0.4109 partial 0.6191 Nemotron partial 0.3046 cross-avg 0.4497 (step 8580)
ettin-68m full : PT-BR strict 0.3663 partial 0.5416 Nemotron partial 0.3563 cross-avg 0.4239 (step 19550)
ettin-32m full : PT-BR strict 0.1429 partial 0.3536 Nemotron partial 0.2640 cross-avg 0.2867 (step 19550)
ettin-32m full-3x : PT-BR strict 0.6326 partial 0.7167 Nemotron partial 0.5721 cross-avg 0.6031 (step 65550)
ettin-68m full-3x : PT-BR strict 0.7076 partial 0.7979 Nemotron partial 0.6799 cross-avg 0.6817 (step 41400)
gliner-ettin-32m-ptbr-pii-easter-egg: PT-BR strict 0.6445 partial 0.7344 Nemotron partial 0.6019 cross-avg 0.6243 (step 100)
gliner-ettin-68m-ptbr-pii-easter-egg: PT-BR strict 0.7109 partial 0.8073 Nemotron partial 0.7182 cross-avg 0.7167 (step 100) *** PT-BR best *** cross-source best
Take-away. The 3x re-runs (69 000 steps) blew past the original 23 000-step sweep — the original 32M variants were straightforwardly under-trained, and the longer cosine schedule unlocks 1.5–2× higher F1 on every PII source. Curated subsets (top-50k / top-100k) overfit to PT-BR but trail badly on cross-source generalization; the full-data 3x runs dominate on every metric we care about. Among the ettin models, prefer gliner-ettin-{32m,68m}-ptbr-pii-full-3x-v1 — they sit at the Pareto frontier of in-domain F1 and cross-source generalization.
Since then, mmBERT-small full-3x beat both ettin full-3x models on every holdout source (table above).
Earlier result: ettin-68m full (23k-step run) vs prior baselines on Nemotron-PII
For the ettin-68m full checkpoint at step 17250, evaluated independently with eval_nemotron_prior.py (5000 samples, threshold=0.5, gold-label set, single-label flat NER):
| model | strict F1 | partial F1 |
|---|---|---|
| GLiNER multitask-large (prior) | 0.5834 | 0.6490 |
| mmbert-teacher (prior) | 0.4574 | 0.5418 |
| Albertina ckpt-12k (prior) | 0.1801 | 0.3399 |
| ettin-68m (prior) | 0.1493 | 0.3210 |
| ettin-68m full v1 step-17250 (this sweep) | 0.2747 | 0.4306 |
i.e. +0.125 strict F1 / +0.110 partial F1 over the prior ettin-68m baseline at the same protocol — with this sweep's recipe (per-source label superset + bf16 mean-reduction + token_level fix).
Limitations
- A research checkpoint trained and evaluated on synthetic or public PII datasets; there is no human-annotated real-world test set. Do not use it as the sole source of truth for redaction without human review.
- PT-BR performance was the optimization target. The spam/phishing sources score 0 F1 by construction and do not measure abstention; this model's false-positive rate on text with no PII was not measured. The one GLiNER model of the family that was measured (mmBERT-small full-3x) flagged at least one false positive, mostly names, in 63.8% of 2,419 spam/phishing rows (benchmark dataset); expect similar behavior and raise the threshold when false positives are costly.
- The partial-F1 and per-entity numbers cannot be re-derived from files in this repo (see Provenance).
- The original sweep was limited to ~23k steps for 68M and ~12k–16k for the curated runs by single-MI300X budget; this 3x run (69,000 steps) is the follow-up.
License
Apache-2.0, same as the base encoder. Evaluation sources keep their own licenses (see their dataset cards).
- Downloads last month
- 24
Model tree for arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1
Base model
jhu-clsp/ettin-encoder-68m