gliner-ettin-32m-ptbr-pii-full-3x-v1

GLiNER fine-tune of jhu-clsp/ettin-encoder-32m (~32M params) for Brazilian-Portuguese PII detection (CPF, RG, PIS, names, addresses, LGPD-sensitive descriptions), with cross-source generalization tested on English/multilingual PII and spam corpora. Smallest member of the family. It reaches 0.6326 strict F1 on the PT-BR holdout source, behind ettin-68m full-3x (0.7076) and mmBERT-small full-3x (0.7678) on the same protocol; pick it only when model size matters more than accuracy.

It belongs to a sweep of GLiNER models trained on the same MI300X GPU with the same loss recipe, differing in encoder, training subset, batch size and step budget; the family comparison below puts them side by side.

Related repos

Repo Type Role
arthrod/gliner-ettin-32m-ptbr-pii-full-3x-v1 (this repo) model GLiNER on jhu-clsp/ettin-encoder-32m, full ~984k-row train set, 69,000 steps. Smallest model in the family.
arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1 model GLiNER on jhu-clsp/ettin-encoder-68m, full ~984k-row train set, 69,000 steps (trainer state uploaded through step 58,650).
arthrod/gliner-ettin-68m-ptbr-pii-top50k-v1 model GLiNER on jhu-clsp/ettin-encoder-68m, 50k curated subset, 11,700 steps. Ablation: shows the curated subset under-generalizes.
arthrod/gliner-mmbert-small-ptbr-pii-full-3x-v1 model GLiNER on jhu-clsp/mmBERT-small, full train set + 2k Ritápolis rows, 69,000 steps. Strongest GLiNER model of the family.
arthrod/gliner-opf-ptbr-pii-v1 model Fine-tune of OpenAI Privacy Filter (not GLiNER despite the name). Best partial F1 on the in-distribution val set and by far the fewest false positives on no-PII text.
arthrod/gliner-opf-ptbr-pii-bench-v1 dataset Benchmark: predictions and metrics for the mmBERT-small and Privacy Filter models (plus two unreleased ettin variants) on the PT-BR val set, held-out sources and no-PII spam/phishing text.
arthrod/gliner-opf-ptbr-pii-demo space Gradio demo running the Privacy Filter fine-tune.
arthrod/gliner_eval_folder (private) dataset Earlier (Feb 2026) sweep: 18 GLiNER models x 9 eval sets.
arthrod/gliner_review_comparison dataset Earlier GLiNER PII review/comparison data.

Quick start

The repo root holds checkpoint-65550 (the released, best checkpoint; its pytorch_model.bin is byte-identical to checkpoint-65550/pytorch_model.bin), so the plain repo id loads it:

# pip install "gliner==0.2.25"   # version the model was trained with
from gliner import GLiNER

model = GLiNER.from_pretrained("arthrod/gliner-ettin-32m-ptbr-pii-full-3x-v1")

text = (
    "Sou Maria Silva, moro na Rua das Flores, 123, em São Paulo. "
    "Meu CPF é 123.456.789-09 e meu telefone é (11) 91234-5678."
)

labels = [
    "first name", "last name", "middle name",
    "cpf document number", "rg document number", "pis document number",
    "phone number", "email address", "credit card",
    "location street", "location building number", "location neighborhood",
    "location city", "location state", "location state abbreviation",
    "location zip", "location full address",
    "dob",
]

for ent in model.predict_entities(text, labels, threshold=0.3):
    print(ent["text"], "→", ent["label"], f"({ent['score']:.2f})")

To load any other checkpoint, download just that folder (skip the optimizer state) and pass the local path:

from huggingface_hub import snapshot_download
path = snapshot_download("arthrod/gliner-ettin-32m-ptbr-pii-full-3x-v1", allow_patterns=["checkpoint-13800/*"],
                         ignore_patterns=["*/optimizer.pt", "*/scheduler.pt", "*/rng_state.pth"])
model = GLiNER.from_pretrained(f"{path}/checkpoint-13800")

Files and checkpoint layout

path content
gliner_config.json, pytorch_model.bin, tokenizer.json, tokenizer_config.json released model = checkpoint-65550
checkpoint-N/ (20 folders) full Trainer checkpoint every 3,450 steps: gliner_config.json, pytorch_model.bin, tokenizer, trainer_state.json, plus optimizer.pt, scheduler.pt, rng_state.pth for resuming
eval_holdout_strict_f1.parquet per-evaluation strict F1 on all 11 holdout sources, extracted from checkpoint-69000/trainer_state.json (see Evaluation)

Uploaded checkpoint steps: 3450, 6900, 10350, 13800, 17250, 20700, 24150, 27600, 31050, 34500, 37950, 41400, 44850, 48300, 51750, 55200, 58650, 62100, 65550, 69000.

Training data

  • Train: data/splits/train.jsonl (~984k samples, full train set). Not published; it is assembled from the same upstream PII and spam sources as the evaluation holdout, with the holdout sliced off before the train split was built (no leakage).
  • Validation (eval loss): data/splits/val_5k.jsonl, 5,000 rows.
  • Holdout for model selection: data/holdout_5k_original, 5,000 rows, proportional across 11 sources (list under Evaluation). Public upstream sources include nvidia/Nemotron-PII, ai4privacy/open-pii-masking-500k-ai4privacy and ai4privacy/pii-masking-400k; the PT-BR source (gliner2_pii_ptbr_reward_split) is synthetic and private.
  • Benchmark on the shared PT-BR val set with predictions: arthrod/gliner-opf-ptbr-pii-bench-v1 (this model is not in it).

Training recipe

From gliner_config.json and checkpoint-69000/trainer_state.json, plus the original run notes:

setting value
base encoder jhu-clsp/ettin-encoder-32m (ModernBERT, 10 layers, hidden 384, 6 heads, vocab 50,371)
GLiNER head span_mode: token_level, max_len: 1024, max_width: 100, max_types: 100, subtoken_pooling: first, dropout: 0.3, fine_tune: true
steps 69,000 (max_steps in trainer state), eval and save every 3,450 steps
batch size train 256, eval 192
epochs over the train set ~17.9 (18 configured)
learning rate encoder 2e-5, other params 5e-5 (peak logged LR 5e-5)
schedule cosine, warmup ratio 0.1 (logged LR peaks at step 6925)
weight decay 0.01 (encoder and others), max grad norm 1.0
loss focal, alpha 0.75, gamma 2.0, reduction mean
precision bf16, HIPBLASLT_ALLOW_TF32=0 (avoids MI300X NaN intermittency)
hardware single AMD MI300X
library gliner==0.2.25 plus a local fix for the upstream token_level regression that landed in 0.2.26

Evaluation

Protocol

Every 3,450 steps the trainer re-scores a 5,000-sample, proportional, source-mixed holdout (data/holdout_5k_original). For each row it builds a per-source label superset (the 24 canonical PT-BR labels ∪ that row's gold labels, capped at 100, lowercased) and calls model.inference(..., flat_ner=True, multi_label=False, threshold=0.3). Per-source F1 is computed by nervaluate (strict / exact / partial / ent_type). Model selection: best PT-BR partial F1.

Eleven sources cover PT-BR PII, EN PII, and negative-evidence spam/phishing:

  • gliner2_pii_ptbr_reward_split: PT-BR PII (the target)
  • nemotron_pii: nvidia/Nemotron-PII, EN
  • open_pii_masking_500k, pii_masking_400k: multilingual PII (ai4privacy)
  • enron_spam_bvk, enron_spam_setfit, phishing_darkknight, phishing_zefang, sms_spam_multilingual, spam_messages_mshenoda, spamassassin: negative evidence (no gold PII). Their F1 is 0 by construction; it says nothing about false positives. For false-positive counts on no-PII text see the benchmark dataset (covers the mmBERT and Privacy Filter models only).

Provenance. Strict F1 values are in this repo (checkpoint-69000/trainer_state.json, extracted to eval_holdout_strict_f1.parquet) and were checked against the tables below. Partial F1 and the per-entity breakdown come from the original run's evaluation logs, which are not stored in this repo; they are reproduced from the April 2026 card unchanged.

Released checkpoint (checkpoint-65550): per-source F1

eval source strict F1 partial F1
PT-BR (gliner2 reward-split) 0.6326 0.7167
EN (nvidia/Nemotron-PII) 0.4945 0.5721
open-pii-masking-500k 0.4081 0.5584
pii-masking-400k 0.4193 0.5654
enron-spam (bvk) 0.0000 0.0000
enron-spam (setfit) 0.0000 0.0000
phishing (darkknight) 0.0000 0.0000
phishing (zefang) 0.0000 0.0000
SMS spam (multilingual) 0.0000 0.0000
spam-messages (mshenoda) 0.0000 0.0000
spamassassin 0.0000 0.0000

Per-entity breakdown (4 PII sources)

Averaged only across sources that contain the label (sources where the label is absent in gold are excluded — including them would falsely deflate the aggregate, e.g. personal description of ethnicity averaged with two sources that lack it would drop from ~0.58 to ~0.29 partial F1). n_src is the number of contributing sources. Sorted by partial F1.

entity n_src strict P strict R strict F1 partial P partial R partial F1
rg document number 1 0.929 0.944 0.936 0.956 0.972 0.964
pis document number 1 0.950 0.897 0.922 0.971 0.917 0.943
cpf document number 1 0.906 0.895 0.901 0.947 0.936 0.942
email address 4 0.849 0.848 0.848 0.914 0.913 0.913
coordinate 1 0.738 0.818 0.776 0.820 0.909 0.862
phone number 4 0.778 0.881 0.825 0.809 0.916 0.858
bank routing number 1 0.768 0.946 0.848 0.768 0.946 0.848
medical record number 1 0.769 0.933 0.843 0.769 0.933 0.843
dob 3 0.720 0.869 0.779 0.752 0.917 0.817
date 2 0.693 0.761 0.724 0.777 0.850 0.811
customer id 1 0.712 0.910 0.799 0.712 0.910 0.799
url 1 0.598 0.583 0.591 0.793 0.773 0.783
county 1 0.667 0.809 0.731 0.713 0.865 0.782
swift bic 1 0.606 1.000 0.755 0.606 1.000 0.755
device identifier 1 0.645 0.909 0.755 0.645 0.909 0.755
cvv 1 0.641 0.893 0.746 0.641 0.893 0.746
biometric identifier 1 0.622 0.925 0.744 0.622 0.925 0.744
employee id 1 0.621 0.923 0.742 0.621 0.923 0.742
health plan beneficiary number 1 0.596 0.914 0.721 0.607 0.931 0.735
location state abbreviation 1 0.580 1.000 0.734 0.580 1.000 0.734
fax number 1 0.737 0.718 0.727 0.737 0.718 0.727
ipv4 1 0.500 0.564 0.530 0.682 0.769 0.723
credit card 4 0.665 0.785 0.686 0.697 0.821 0.719
mac address 1 0.438 0.467 0.452 0.688 0.733 0.710
location full address 1 0.491 0.518 0.504 0.690 0.727 0.708
middle name 1 0.534 0.965 0.687 0.534 0.965 0.687
location zip 4 0.544 0.911 0.661 0.554 0.936 0.676
account number 2 0.627 0.730 0.674 0.627 0.730 0.674
certificate license number 1 0.524 0.892 0.660 0.524 0.892 0.660
vehicle identifier 1 0.492 1.000 0.659 0.492 1.000 0.659
pin 1 0.545 0.828 0.658 0.545 0.828 0.658
location building number 3 0.573 0.772 0.647 0.582 0.782 0.656
date time 1 0.544 0.623 0.581 0.608 0.696 0.649
country 1 0.474 0.991 0.642 0.474 0.991 0.642
time 2 0.505 0.737 0.584 0.546 0.822 0.638
personal description of organizational affiliation 1 0.381 0.447 0.411 0.582 0.684 0.629
location street 3 0.463 0.597 0.518 0.553 0.726 0.623
company name 1 0.463 0.562 0.507 0.557 0.677 0.611
license plate 1 0.435 0.714 0.541 0.489 0.804 0.608
first name 4 0.433 0.819 0.557 0.456 0.855 0.584
location state 2 0.422 0.824 0.556 0.439 0.858 0.579
api key 1 0.392 0.816 0.530 0.424 0.882 0.573
last name 4 0.441 0.757 0.536 0.464 0.809 0.565
personal description of ethnicity 2 0.360 0.673 0.468 0.419 0.778 0.544
religious belief 1 0.353 0.857 0.500 0.382 0.929 0.542
location neighborhood 1 0.370 0.702 0.485 0.398 0.754 0.521
social security number 3 0.482 0.567 0.505 0.491 0.576 0.514
personal description of sexual information 1 0.226 0.357 0.277 0.397 0.627 0.486
language 1 0.314 1.000 0.477 0.314 1.000 0.477
sex or gender 2 0.296 0.875 0.443 0.301 0.887 0.449
ipv6 1 0.111 0.182 0.138 0.361 0.591 0.448
age 2 0.295 0.953 0.445 0.295 0.953 0.445
passport number 1 0.353 0.600 0.444 0.353 0.600 0.444
location city 4 0.288 0.710 0.407 0.313 0.769 0.442
personal description of religious convictions 1 0.212 0.444 0.287 0.314 0.658 0.425
tax id number 3 0.325 0.535 0.379 0.378 0.581 0.422
id card number 2 0.324 0.642 0.420 0.324 0.642 0.420
user name 2 0.239 0.516 0.321 0.293 0.650 0.396
personal description of medical conditions 1 0.158 0.333 0.215 0.291 0.614 0.395
http cookie 1 0.060 0.103 0.076 0.310 0.534 0.392
sexuality 1 0.253 0.741 0.377 0.259 0.759 0.387
password 2 0.119 0.249 0.161 0.274 0.573 0.371
employment status 1 0.223 0.671 0.334 0.227 0.685 0.341
personal description of political opinion 1 0.107 0.342 0.163 0.197 0.631 0.300
driver license number 2 0.396 0.266 0.283 0.410 0.279 0.297
education level 1 0.102 0.358 0.159 0.177 0.623 0.276
blood type 1 0.035 0.125 0.055 0.150 0.531 0.234
political view 1 0.110 0.769 0.192 0.126 0.885 0.221
unique id 1 0.122 0.857 0.214 0.122 0.857 0.214
title 1 0.089 0.636 0.156 0.097 0.700 0.171
occupation 1 0.018 0.158 0.032 0.051 0.449 0.092

Reading guide. Structured tokens (email, phone, doc numbers, IPs, MAC) approach strict ≈ partial because their boundaries are unambiguous. Long natural-language spans (full address, the personal description of … labels) carry a meaningful strict-vs-partial gap because exact boundaries are inherently fuzzy — even two human annotators would disagree. For those labels, partial F1 is the operationally meaningful metric.

Checkpoint sweep

"in repo": root = weights at the repo root, folder = checkpoint-N/ exists. Strict columns come from eval_holdout_strict_f1.parquet; partial columns from the original logs. Best row in bold.

step in repo PT-BR strict Nemotron strict open-pii strict pii-masking strict mean-4 strict PT-BR partial Nemotron partial open-pii partial pii-masking partial
3450 folder 0.0418 0.0458 0.0748 0.0553 0.0544 0.2343 0.2430 0.2433 0.2756
6900 folder 0.0382 0.0295 0.0601 0.0441 0.0430 0.2149 0.1812 0.2302 0.2328
10350 folder 0.0734 0.0595 0.0830 0.0549 0.0677 0.2940 0.2590 0.2771 0.2248
13800 folder 0.1170 0.1087 0.1154 0.0721 0.1033 0.3049 0.3213 0.3507 0.2452
17250 folder 0.1967 0.1372 0.1237 0.1010 0.1397 0.3828 0.3123 0.3514 0.2705
20700 folder 0.2919 0.1860 0.2000 0.1651 0.2107 0.4961 0.3535 0.3929 0.3474
24150 folder 0.3328 0.2044 0.2345 0.1817 0.2383 0.5115 0.3537 0.4192 0.3535
27600 folder 0.3604 0.2490 0.2510 0.2071 0.2669 0.5117 0.3845 0.4342 0.3690
31050 folder 0.4700 0.3228 0.2935 0.2809 0.3418 0.6083 0.4437 0.4714 0.4472
34500 folder 0.5147 0.2946 0.2891 0.2562 0.3386 0.6458 0.4121 0.4639 0.3983
37950 folder 0.4858 0.3393 0.3280 0.2803 0.3584 0.5946 0.4575 0.5052 0.4288
41400 folder 0.5340 0.3897 0.3542 0.3373 0.4038 0.6477 0.4892 0.5191 0.4856
44850 folder 0.5758 0.4577 0.3983 0.3749 0.4517 0.6812 0.5484 0.5658 0.5332
48300 folder 0.5670 0.4002 0.3779 0.3425 0.4219 0.6703 0.4890 0.5218 0.4836
51750 folder 0.5812 0.4624 0.3307 0.3698 0.4360 0.6716 0.5480 0.4660 0.5083
55200 folder 0.6080 0.4595 0.3754 0.3891 0.4580 0.6959 0.5392 0.5234 0.5287
58650 folder 0.6276 0.4874 0.4110 0.4144 0.4851 0.7160 0.5665 0.5622 0.5599
62100 folder 0.6248 0.4981 0.4006 0.4205 0.4860 0.7076 0.5764 0.5494 0.5655
65550 root + folder 0.6326 0.4945 0.4081 0.4193 0.4886 0.7167 0.5721 0.5584 0.5654
69000 folder 0.6291 0.4925 0.4066 0.4189 0.4868 0.7135 0.5694 0.5566 0.5636

eval_holdout_strict_f1.parquet

One row per evaluation (20 rows), extracted from checkpoint-69000/trainer_state.json (log_history). Load with pd.read_parquet("hf://arthrod/gliner-ettin-32m-ptbr-pii-full-3x-v1/eval_holdout_strict_f1.parquet") (pandas + huggingface_hub; polars hf:// only supports dataset/Space repos, so use hf_hub_download first there).

column type description
step Int32 global training step of the evaluation
epoch Float64 training epoch at that step (fractional)
checkpoint_dir String checkpoint-N/ folder holding that step's weights in this repo, null if not uploaded
is_repo_root Boolean true for the step whose weights sit at the repo root (the released model)
eval_loss Float64 validation loss on val_5k
pii4_mean_strict_f1 Float64 mean strict F1 over the 4 PII sources (computed from the four columns below)
avg11_strict_f1 Float64 mean strict F1 over all 11 sources as logged by the trainer (the 7 spam/phishing sources are always 0, so this is not a quality metric)
gliner2_pii_ptbr_reward_split_strict_f1 Float64 nervaluate strict F1 on the gliner2_pii_ptbr_reward_split holdout source
nemotron_pii_strict_f1 Float64 nervaluate strict F1 on the nemotron_pii holdout source
open_pii_masking_500k_strict_f1 Float64 nervaluate strict F1 on the open_pii_masking_500k holdout source
pii_masking_400k_strict_f1 Float64 nervaluate strict F1 on the pii_masking_400k holdout source
enron_spam_bvk_strict_f1 Float64 nervaluate strict F1 on the enron_spam_bvk holdout source (no gold entities, always 0)
enron_spam_setfit_strict_f1 Float64 nervaluate strict F1 on the enron_spam_setfit holdout source (no gold entities, always 0)
phishing_darkknight_strict_f1 Float64 nervaluate strict F1 on the phishing_darkknight holdout source (no gold entities, always 0)
phishing_zefang_strict_f1 Float64 nervaluate strict F1 on the phishing_zefang holdout source (no gold entities, always 0)
sms_spam_multilingual_strict_f1 Float64 nervaluate strict F1 on the sms_spam_multilingual holdout source (no gold entities, always 0)
spam_messages_mshenoda_strict_f1 Float64 nervaluate strict F1 on the spam_messages_mshenoda holdout source (no gold entities, always 0)
spamassassin_strict_f1 Float64 nervaluate strict F1 on the spamassassin holdout source (no gold entities, always 0)

Example rows (first evaluation and the released step):

[
 {
  "step": 3450,
  "epoch": 0.897036,
  "checkpoint_dir": "checkpoint-3450",
  "is_repo_root": false,
  "eval_loss": 0.001316,
  "pii4_mean_strict_f1": 0.05445,
  "avg11_strict_f1": 0.0198,
  "gliner2_pii_ptbr_reward_split_strict_f1": 0.041816,
  "nemotron_pii_strict_f1": 0.045839,
  "open_pii_masking_500k_strict_f1": 0.074838,
  "pii_masking_400k_strict_f1": 0.055308,
  "enron_spam_bvk_strict_f1": 0.0,
  "enron_spam_setfit_strict_f1": 0.0,
  "phishing_darkknight_strict_f1": 0.0,
  "phishing_zefang_strict_f1": 0.0,
  "sms_spam_multilingual_strict_f1": 0.0,
  "spam_messages_mshenoda_strict_f1": 0.0,
  "spamassassin_strict_f1": 0.0
 },
 {
  "step": 65550,
  "epoch": 17.043682,
  "checkpoint_dir": "checkpoint-65550",
  "is_repo_root": true,
  "eval_loss": 0.000228,
  "pii4_mean_strict_f1": 0.488643,
  "avg11_strict_f1": 0.177688,
  "gliner2_pii_ptbr_reward_split_strict_f1": 0.632647,
  "nemotron_pii_strict_f1": 0.494539,
  "open_pii_masking_500k_strict_f1": 0.40806,
  "pii_masking_400k_strict_f1": 0.419325,
  "enron_spam_bvk_strict_f1": 0.0,
  "enron_spam_setfit_strict_f1": 0.0,
  "phishing_darkknight_strict_f1": 0.0,
  "phishing_zefang_strict_f1": 0.0,
  "sms_spam_multilingual_strict_f1": 0.0,
  "spam_messages_mshenoda_strict_f1": 0.0,
  "spamassassin_strict_f1": 0.0
 }
]

Comparison across the family

Same holdout protocol (strict F1, released checkpoint of each repo). Source: each repo's eval_holdout_strict_f1.parquet.

model released step PT-BR Nemotron (EN) open-pii-500k pii-masking-400k mean of 4
ettin-32m full-3x (this) 65550 0.6326 0.4945 0.4081 0.4193 0.4886
ettin-68m full-3x 41400 0.7076 0.6149 0.4902 0.4701 0.5707
ettin-68m top-50k 9945 0.2732 0.0345 0.0609 0.0755 0.1110
mmBERT-small full-3x 41400 0.7678 0.8158 0.7217 0.7076 0.7532
Privacy Filter fine-tune e3_c3 not evaluated with this protocol

Shared PT-BR val benchmark (24 canonical labels). None of the ettin repos in this family were run on it.

model partial P partial R partial F1 exact P exact R exact F1
Privacy Filter fine-tune e3_c3 0.917 0.879 0.897 0.871 0.835 0.853
mmBERT-small full-3x (41400) 0.951 0.832 0.888 0.932 0.815 0.870
ettin-68m easter-egg (not public) 0.893 0.761 0.822 0.857 0.731 0.789
ettin-32m easter-egg (not public) 0.905 0.729 0.808 0.862 0.694 0.769

Source: comparison_report.md in the Privacy Filter repo. 5,000-row PT-BR val set, restricted to the 24 canonical labels, nervaluate; GLiNER predictions at threshold 0.5, Privacy Filter with its default Viterbi decoding. The ettin models in this table are separate easter-egg fine-tunes, not the ettin repos of this family, which were not run on this benchmark.

Original sweep log (best by PT-BR partial F1 of each run). Reproduced from the April 2026 card; most of these runs (32m/68m top-100k, the 23k-step "full" runs, the easter-egg fine-tunes) are not published.

ettin-32m top-50k       : PT-BR strict 0.1411  partial 0.3710  Nemotron partial 0.1710  cross-avg 0.2808  (step 8500)
ettin-32m top-100k      : PT-BR strict 0.0946  partial 0.3165  Nemotron partial 0.1455  cross-avg 0.2428  (step 5200)
ettin-68m top-50k       : PT-BR strict 0.2732  partial 0.4802  Nemotron partial 0.2397  cross-avg 0.3117  (step 9945)
ettin-68m top-100k      : PT-BR strict 0.4109  partial 0.6191  Nemotron partial 0.3046  cross-avg 0.4497  (step 8580)
ettin-68m full          : PT-BR strict 0.3663  partial 0.5416  Nemotron partial 0.3563  cross-avg 0.4239  (step 19550)
ettin-32m full          : PT-BR strict 0.1429  partial 0.3536  Nemotron partial 0.2640  cross-avg 0.2867  (step 19550)
ettin-32m full-3x       : PT-BR strict 0.6326  partial 0.7167  Nemotron partial 0.5721  cross-avg 0.6031  (step 65550)
ettin-68m full-3x       : PT-BR strict 0.7076  partial 0.7979  Nemotron partial 0.6799  cross-avg 0.6817  (step 41400)
gliner-ettin-32m-ptbr-pii-easter-egg: PT-BR strict 0.6445  partial 0.7344  Nemotron partial 0.6019  cross-avg 0.6243  (step 100)
gliner-ettin-68m-ptbr-pii-easter-egg: PT-BR strict 0.7109  partial 0.8073  Nemotron partial 0.7182  cross-avg 0.7167  (step 100)  *** PT-BR best  *** cross-source best

Take-away. The 3x re-runs (69 000 steps) blew past the original 23 000-step sweep — the original 32M variants were straightforwardly under-trained, and the longer cosine schedule unlocks 1.5–2× higher F1 on every PII source. Curated subsets (top-50k / top-100k) overfit to PT-BR but trail badly on cross-source generalization; the full-data 3x runs dominate on every metric we care about. Among the ettin models, prefer gliner-ettin-{32m,68m}-ptbr-pii-full-3x-v1 — they sit at the Pareto frontier of in-domain F1 and cross-source generalization.

Since then, mmBERT-small full-3x beat both ettin full-3x models on every holdout source (table above).

Earlier result: ettin-68m full (23k-step run) vs prior baselines on Nemotron-PII

For the ettin-68m full checkpoint at step 17250, evaluated independently with eval_nemotron_prior.py (5000 samples, threshold=0.5, gold-label set, single-label flat NER):

model strict F1 partial F1
GLiNER multitask-large (prior) 0.5834 0.6490
mmbert-teacher (prior) 0.4574 0.5418
Albertina ckpt-12k (prior) 0.1801 0.3399
ettin-68m (prior) 0.1493 0.3210
ettin-68m full v1 step-17250 (this sweep) 0.2747 0.4306

i.e. +0.125 strict F1 / +0.110 partial F1 over the prior ettin-68m baseline at the same protocol — with this sweep's recipe (per-source label superset + bf16 mean-reduction + token_level fix).

Limitations

  • A research checkpoint trained and evaluated on synthetic or public PII datasets; there is no human-annotated real-world test set. Do not use it as the sole source of truth for redaction without human review.
  • PT-BR performance was the optimization target. The spam/phishing sources score 0 F1 by construction and do not measure abstention; this model's false-positive rate on text with no PII was not measured. The one GLiNER model of the family that was measured (mmBERT-small full-3x) flagged at least one false positive, mostly names, in 63.8% of 2,419 spam/phishing rows (benchmark dataset); expect similar behavior and raise the threshold when false positives are costly.
  • The partial-F1 and per-entity numbers cannot be re-derived from files in this repo (see Provenance).
  • The original sweep was limited to ~23k steps for 68M and ~12k–16k for the curated runs by single-MI300X budget; this 3x run (69,000 steps) is the follow-up.

License

Apache-2.0, same as the base encoder. Evaluation sources keep their own licenses (see their dataset cards).

Downloads last month
20
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for arthrod/gliner-ettin-32m-ptbr-pii-full-3x-v1

Finetuned
(32)
this model

Collection including arthrod/gliner-ettin-32m-ptbr-pii-full-3x-v1