gliner-ettin-68m-ptbr-pii-full-3x-v1

GLiNER fine-tune of jhu-clsp/ettin-encoder-68m (~68M params) for Brazilian-Portuguese PII detection (CPF, RG, PIS, names, addresses, LGPD-sensitive descriptions), with cross-source generalization tested on English/multilingual PII and spam corpora. Best ettin model of the family: 0.7076 strict / 0.7979 partial F1 on the PT-BR holdout source at step 41,400. The mmBERT-small full-3x model is stronger on every holdout source under the same protocol.

It belongs to a sweep of GLiNER models trained on the same MI300X GPU with the same loss recipe, differing in encoder, training subset, batch size and step budget; the family comparison below puts them side by side.

Related repos

Repo Type Role
arthrod/gliner-ettin-32m-ptbr-pii-full-3x-v1 model GLiNER on jhu-clsp/ettin-encoder-32m, full ~984k-row train set, 69,000 steps. Smallest model in the family.
arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1 (this repo) model GLiNER on jhu-clsp/ettin-encoder-68m, full ~984k-row train set, 69,000 steps (trainer state uploaded through step 58,650).
arthrod/gliner-ettin-68m-ptbr-pii-top50k-v1 model GLiNER on jhu-clsp/ettin-encoder-68m, 50k curated subset, 11,700 steps. Ablation: shows the curated subset under-generalizes.
arthrod/gliner-mmbert-small-ptbr-pii-full-3x-v1 model GLiNER on jhu-clsp/mmBERT-small, full train set + 2k Ritápolis rows, 69,000 steps. Strongest GLiNER model of the family.
arthrod/gliner-opf-ptbr-pii-v1 model Fine-tune of OpenAI Privacy Filter (not GLiNER despite the name). Best partial F1 on the in-distribution val set and by far the fewest false positives on no-PII text.
arthrod/gliner-opf-ptbr-pii-bench-v1 dataset Benchmark: predictions and metrics for the mmBERT-small and Privacy Filter models (plus two unreleased ettin variants) on the PT-BR val set, held-out sources and no-PII spam/phishing text.
arthrod/gliner-opf-ptbr-pii-demo space Gradio demo running the Privacy Filter fine-tune.
arthrod/gliner_eval_folder (private) dataset Earlier (Feb 2026) sweep: 18 GLiNER models x 9 eval sets.
arthrod/gliner_review_comparison dataset Earlier GLiNER PII review/comparison data.

Quick start

The repo root holds checkpoint-41400 (the released, best checkpoint; its pytorch_model.bin is byte-identical to checkpoint-41400/pytorch_model.bin), so the plain repo id loads it:

# pip install "gliner==0.2.25"   # version the model was trained with
from gliner import GLiNER

model = GLiNER.from_pretrained("arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1")

text = (
    "Sou Maria Silva, moro na Rua das Flores, 123, em São Paulo. "
    "Meu CPF é 123.456.789-09 e meu telefone é (11) 91234-5678."
)

labels = [
    "first name", "last name", "middle name",
    "cpf document number", "rg document number", "pis document number",
    "phone number", "email address", "credit card",
    "location street", "location building number", "location neighborhood",
    "location city", "location state", "location state abbreviation",
    "location zip", "location full address",
    "dob",
]

for ent in model.predict_entities(text, labels, threshold=0.3):
    print(ent["text"], "→", ent["label"], f"({ent['score']:.2f})")

To load any other checkpoint, download just that folder (skip the optimizer state) and pass the local path:

from huggingface_hub import snapshot_download
path = snapshot_download("arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1", allow_patterns=["checkpoint-13800/*"],
                         ignore_patterns=["*/optimizer.pt", "*/scheduler.pt", "*/rng_state.pth"])
model = GLiNER.from_pretrained(f"{path}/checkpoint-13800")

Files and checkpoint layout

path content
gliner_config.json, pytorch_model.bin, tokenizer.json, tokenizer_config.json released model = checkpoint-41400
checkpoint-N/ (17 folders) full Trainer checkpoint every 3,450 steps: gliner_config.json, pytorch_model.bin, tokenizer, trainer_state.json, plus optimizer.pt, scheduler.pt, rng_state.pth for resuming
eval_holdout_strict_f1.parquet per-evaluation strict F1 on all 11 holdout sources, extracted from checkpoint-58650/trainer_state.json (see Evaluation)

Uploaded checkpoint steps: 3450, 6900, 10350, 13800, 17250, 20700, 24150, 27600, 31050, 34500, 37950, 41400, 44850, 48300, 51750, 55200, 58650.

The run went to step 69,000, but checkpoints 62100, 65550 and 69000 were not uploaded; the last trainer state in the repo is checkpoint-58650.

Training data

  • Train: data/splits/train.jsonl (~984k samples, full train set). Not published; it is assembled from the same upstream PII and spam sources as the evaluation holdout, with the holdout sliced off before the train split was built (no leakage).
  • Validation (eval loss): data/splits/val_5k.jsonl, 5,000 rows.
  • Holdout for model selection: data/holdout_5k_original, 5,000 rows, proportional across 11 sources (list under Evaluation). Public upstream sources include nvidia/Nemotron-PII, ai4privacy/open-pii-masking-500k-ai4privacy and ai4privacy/pii-masking-400k; the PT-BR source (gliner2_pii_ptbr_reward_split) is synthetic and private.
  • Benchmark on the shared PT-BR val set with predictions: arthrod/gliner-opf-ptbr-pii-bench-v1 (this model is not in it).

Training recipe

From gliner_config.json and checkpoint-58650/trainer_state.json, plus the original run notes:

setting value
base encoder jhu-clsp/ettin-encoder-68m (ModernBERT, 19 layers, hidden 512, 8 heads, vocab 50,371)
GLiNER head span_mode: token_level, max_len: 1024, max_width: 100, max_types: 100, subtoken_pooling: first, dropout: 0.3, fine_tune: true
steps 69,000 (max_steps in trainer state), eval and save every 3,450 steps
batch size train 128, eval 128
epochs over the train set ~9 (7.62 at step 58,650, the last uploaded state)
learning rate encoder 1.5e-05, other params 5e-5 (peak logged LR 5e-5)
schedule cosine, warmup ratio 0.1 (logged LR peaks at step 6925)
weight decay 0.01 (encoder and others), max grad norm 1.0
loss focal, alpha 0.75, gamma 2.0, reduction mean
precision bf16, HIPBLASLT_ALLOW_TF32=0 (avoids MI300X NaN intermittency)
hardware single AMD MI300X
library gliner==0.2.25 plus a local fix for the upstream token_level regression that landed in 0.2.26

Evaluation

Protocol

Every 3,450 steps the trainer re-scores a 5,000-sample, proportional, source-mixed holdout (data/holdout_5k_original). For each row it builds a per-source label superset (the 24 canonical PT-BR labels ∪ that row's gold labels, capped at 100, lowercased) and calls model.inference(..., flat_ner=True, multi_label=False, threshold=0.3). Per-source F1 is computed by nervaluate (strict / exact / partial / ent_type). Model selection: best PT-BR partial F1.

Eleven sources cover PT-BR PII, EN PII, and negative-evidence spam/phishing:

  • gliner2_pii_ptbr_reward_split: PT-BR PII (the target)
  • nemotron_pii: nvidia/Nemotron-PII, EN
  • open_pii_masking_500k, pii_masking_400k: multilingual PII (ai4privacy)
  • enron_spam_bvk, enron_spam_setfit, phishing_darkknight, phishing_zefang, sms_spam_multilingual, spam_messages_mshenoda, spamassassin: negative evidence (no gold PII). Their F1 is 0 by construction; it says nothing about false positives. For false-positive counts on no-PII text see the benchmark dataset (covers the mmBERT and Privacy Filter models only).

Provenance. Strict F1 values are in this repo (checkpoint-58650/trainer_state.json, extracted to eval_holdout_strict_f1.parquet) and were checked against the tables below. Partial F1 and the per-entity breakdown come from the original run's evaluation logs, which are not stored in this repo; they are reproduced from the April 2026 card unchanged.

Released checkpoint (checkpoint-41400): per-source F1

eval source strict F1 partial F1
PT-BR (gliner2 reward-split) 0.7076 0.7979
EN (nvidia/Nemotron-PII) 0.6149 0.6799
open-pii-masking-500k 0.4902 0.6464
pii-masking-400k 0.4701 0.6024
enron-spam (bvk) 0.0000 0.0000
enron-spam (setfit) 0.0000 0.0000
phishing (darkknight) 0.0000 0.0000
phishing (zefang) 0.0000 0.0000
SMS spam (multilingual) 0.0000 0.0000
spam-messages (mshenoda) 0.0000 0.0000
spamassassin 0.0000 0.0000

Per-entity breakdown (4 PII sources)

Averaged only across sources that contain the label (sources where the label is absent in gold are excluded — including them would falsely deflate the aggregate, e.g. personal description of ethnicity averaged with two sources that lack it would drop from ~0.58 to ~0.29 partial F1). n_src is the number of contributing sources. Sorted by partial F1.

entity n_src strict P strict R strict F1 partial P partial R partial F1
ipv4 1 1.000 1.000 1.000 1.000 1.000 1.000
email address 4 0.975 0.970 0.972 0.982 0.977 0.980
device identifier 1 1.000 0.955 0.977 1.000 0.955 0.977
url 1 0.915 0.896 0.905 0.957 0.938 0.947
pis document number 1 0.974 0.881 0.925 0.974 0.881 0.925
rg document number 1 0.869 0.960 0.912 0.869 0.960 0.912
phone number 4 0.864 0.926 0.892 0.873 0.936 0.901
date time 1 0.952 0.855 0.901 0.952 0.855 0.901
biometric identifier 1 0.944 0.838 0.887 0.944 0.838 0.887
date 2 0.817 0.924 0.860 0.833 0.941 0.876
mac address 1 0.960 0.800 0.873 0.960 0.800 0.873
coordinate 1 0.746 0.800 0.772 0.831 0.891 0.860
bank routing number 1 0.870 0.839 0.855 0.870 0.839 0.855
medical record number 1 0.760 0.973 0.854 0.760 0.973 0.854
cpf document number 1 0.828 0.837 0.832 0.828 0.837 0.832
county 1 0.727 0.899 0.804 0.745 0.921 0.824
vehicle identifier 1 0.698 1.000 0.822 0.698 1.000 0.822
location full address 1 0.736 0.709 0.722 0.835 0.805 0.819
health plan beneficiary number 1 0.716 0.914 0.803 0.730 0.931 0.818
fax number 1 0.964 0.692 0.806 0.964 0.692 0.806
country 1 0.691 0.938 0.795 0.694 0.942 0.799
time 2 0.692 0.893 0.768 0.708 0.918 0.787
swift bic 1 0.625 1.000 0.769 0.625 1.000 0.769
credit card 4 0.818 0.733 0.760 0.824 0.738 0.765
customer id 1 0.663 0.887 0.759 0.663 0.887 0.759
dob 3 0.811 0.761 0.746 0.819 0.774 0.756
pin 1 0.719 0.793 0.754 0.719 0.793 0.754
cvv 1 0.632 0.857 0.727 0.632 0.857 0.727
certificate license number 1 0.652 0.811 0.723 0.652 0.811 0.723
language 1 0.571 0.973 0.720 0.571 0.973 0.720
account number 2 0.707 0.695 0.699 0.707 0.695 0.699
blood type 1 0.561 0.719 0.630 0.622 0.797 0.699
license plate 1 0.553 0.750 0.636 0.605 0.821 0.697
personal description of organizational affiliation 1 0.463 0.500 0.481 0.659 0.711 0.684
location street 3 0.566 0.653 0.603 0.639 0.740 0.682
personal description of sexual information 1 0.450 0.494 0.471 0.648 0.711 0.678
location building number 3 0.535 0.917 0.671 0.536 0.920 0.673
employee id 1 0.510 0.974 0.670 0.510 0.974 0.670
location zip 4 0.570 0.832 0.659 0.576 0.843 0.667
location state 2 0.494 0.905 0.637 0.496 0.909 0.639
ipv6 1 0.643 0.409 0.500 0.821 0.523 0.639
religious belief 1 0.526 0.714 0.606 0.553 0.750 0.636
last name 4 0.526 0.726 0.603 0.552 0.765 0.635
personal description of political opinion 1 0.407 0.495 0.447 0.574 0.698 0.630
company name 1 0.466 0.672 0.551 0.533 0.768 0.629
tax id number 3 0.651 0.568 0.601 0.679 0.584 0.621
personal description of religious convictions 1 0.414 0.538 0.468 0.539 0.701 0.610
sex or gender 2 0.444 0.875 0.587 0.444 0.875 0.587
personal description of ethnicity 2 0.422 0.726 0.527 0.468 0.795 0.582
first name 4 0.438 0.830 0.555 0.456 0.868 0.579
location neighborhood 1 0.435 0.702 0.537 0.462 0.746 0.570
http cookie 1 0.304 0.241 0.269 0.630 0.500 0.558
location city 4 0.410 0.737 0.525 0.428 0.771 0.548
api key 1 0.387 0.763 0.513 0.407 0.803 0.540
id card number 2 0.448 0.660 0.529 0.453 0.668 0.535
social security number 3 0.665 0.496 0.530 0.669 0.498 0.533
user name 2 0.413 0.658 0.507 0.421 0.672 0.517
age 2 0.413 0.791 0.514 0.413 0.791 0.514
personal description of medical conditions 1 0.279 0.439 0.341 0.413 0.652 0.506
password 2 0.282 0.321 0.299 0.453 0.531 0.487
middle name 1 0.304 0.789 0.439 0.304 0.789 0.439
employment status 1 0.304 0.562 0.394 0.315 0.582 0.409
driver license number 2 0.605 0.323 0.368 0.612 0.330 0.375
location state abbreviation 1 0.212 1.000 0.349 0.212 1.000 0.349
education level 1 0.124 0.245 0.165 0.248 0.491 0.329
unique id 1 0.190 0.571 0.286 0.190 0.571 0.286
passport number 1 0.400 0.200 0.267 0.400 0.200 0.267
sexuality 1 0.153 0.667 0.248 0.153 0.667 0.248
political view 1 0.107 0.692 0.186 0.119 0.769 0.206
title 1 0.111 0.582 0.187 0.117 0.609 0.196
occupation 1 0.040 0.232 0.068 0.075 0.435 0.128

Reading guide. Structured tokens (email, phone, doc numbers, IPs, MAC) approach strict ≈ partial because their boundaries are unambiguous. Long natural-language spans (full address, the personal description of … labels) carry a meaningful strict-vs-partial gap because exact boundaries are inherently fuzzy — even two human annotators would disagree. For those labels, partial F1 is the operationally meaningful metric.

Checkpoint sweep

"in repo": root = weights at the repo root, folder = checkpoint-N/ exists. Strict columns come from eval_holdout_strict_f1.parquet; partial columns from the original logs. Best row in bold. Steps marked * were not uploaded; their PT-BR strict F1 is from the original card only.

step in repo PT-BR strict Nemotron strict open-pii strict pii-masking strict mean-4 strict PT-BR partial Nemotron partial open-pii partial pii-masking partial
3450 folder 0.0435 0.0289 0.0724 0.0417 0.0466 0.3185 0.1783 0.3494 0.3067
6900 folder 0.1928 0.0624 0.1621 0.1041 0.1303 0.4170 0.2571 0.4157 0.3310
10350 folder 0.2774 0.1741 0.1668 0.1336 0.1880 0.4642 0.3307 0.4036 0.3302
13800 folder 0.3107 0.2034 0.1704 0.1593 0.2110 0.4872 0.3371 0.3994 0.3343
17250 folder 0.4126 0.2584 0.2046 0.2070 0.2707 0.6059 0.3688 0.3910 0.3869
20700 folder 0.4838 0.3571 0.2318 0.2295 0.3255 0.6549 0.4791 0.4424 0.4135
24150 folder 0.4802 0.5093 0.2638 0.2840 0.3843 0.6348 0.6219 0.4534 0.4592
27600 folder 0.5864 0.4761 0.2911 0.2948 0.4121 0.7220 0.5603 0.4536 0.4289
31050 folder 0.5480 0.4819 0.3188 0.3575 0.4265 0.6731 0.5524 0.4509 0.4876
34500 folder 0.5921 0.5541 0.3647 0.3709 0.4704 0.6957 0.6209 0.4843 0.5082
37950 folder 0.6329 0.5491 0.4265 0.4273 0.5090 0.7441 0.6189 0.5580 0.5634
41400 root + folder 0.7076 0.6149 0.4902 0.4701 0.5707 0.7979 0.6799 0.6464 0.6024
44850 folder 0.6747 0.5928 0.4939 0.4543 0.5539 0.7687 0.6522 0.6310 0.5855
48300 folder 0.6994 0.5679 0.4940 0.4725 0.5584 0.7859 0.6205 0.6253 0.5899
51750 folder 0.6312 0.6189 0.4918 0.4651 0.5518 0.7073 0.6697 0.6222 0.5806
55200 folder 0.6658 0.5639 0.5013 0.4783 0.5523 0.7415 0.6109 0.6391 0.6015
58650 folder 0.6408 0.5562 0.4930 0.4341 0.5310 0.7147 0.6031 0.6187 0.5514
62100 no 0.6765* — — — — 0.7504 0.6239 0.6448 0.5808
65550 no 0.6724* — — — — 0.7475 0.6172 0.6419 0.5827
69000 no 0.6701* — — — — 0.7448 0.6147 0.6415 0.5845

eval_holdout_strict_f1.parquet

One row per evaluation (17 rows), extracted from checkpoint-58650/trainer_state.json (log_history). Load with pd.read_parquet("hf://arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1/eval_holdout_strict_f1.parquet") (pandas + huggingface_hub; polars hf:// only supports dataset/Space repos, so use hf_hub_download first there).

column type description
step Int32 global training step of the evaluation
epoch Float64 training epoch at that step (fractional)
checkpoint_dir String checkpoint-N/ folder holding that step's weights in this repo, null if not uploaded
is_repo_root Boolean true for the step whose weights sit at the repo root (the released model)
eval_loss Float64 validation loss on val_5k
pii4_mean_strict_f1 Float64 mean strict F1 over the 4 PII sources (computed from the four columns below)
avg11_strict_f1 Float64 mean strict F1 over all 11 sources as logged by the trainer (the 7 spam/phishing sources are always 0, so this is not a quality metric)
gliner2_pii_ptbr_reward_split_strict_f1 Float64 nervaluate strict F1 on the gliner2_pii_ptbr_reward_split holdout source
nemotron_pii_strict_f1 Float64 nervaluate strict F1 on the nemotron_pii holdout source
open_pii_masking_500k_strict_f1 Float64 nervaluate strict F1 on the open_pii_masking_500k holdout source
pii_masking_400k_strict_f1 Float64 nervaluate strict F1 on the pii_masking_400k holdout source
enron_spam_bvk_strict_f1 Float64 nervaluate strict F1 on the enron_spam_bvk holdout source (no gold entities, always 0)
enron_spam_setfit_strict_f1 Float64 nervaluate strict F1 on the enron_spam_setfit holdout source (no gold entities, always 0)
phishing_darkknight_strict_f1 Float64 nervaluate strict F1 on the phishing_darkknight holdout source (no gold entities, always 0)
phishing_zefang_strict_f1 Float64 nervaluate strict F1 on the phishing_zefang holdout source (no gold entities, always 0)
sms_spam_multilingual_strict_f1 Float64 nervaluate strict F1 on the sms_spam_multilingual holdout source (no gold entities, always 0)
spam_messages_mshenoda_strict_f1 Float64 nervaluate strict F1 on the spam_messages_mshenoda holdout source (no gold entities, always 0)
spamassassin_strict_f1 Float64 nervaluate strict F1 on the spamassassin holdout source (no gold entities, always 0)

Example rows (first evaluation and the released step):

[
 {
  "step": 3450,
  "epoch": 0.448518,
  "checkpoint_dir": "checkpoint-3450",
  "is_repo_root": false,
  "eval_loss": 0.001429,
  "pii4_mean_strict_f1": 0.046608,
  "avg11_strict_f1": 0.016948,
  "gliner2_pii_ptbr_reward_split_strict_f1": 0.043478,
  "nemotron_pii_strict_f1": 0.02889,
  "open_pii_masking_500k_strict_f1": 0.072408,
  "pii_masking_400k_strict_f1": 0.041657,
  "enron_spam_bvk_strict_f1": 0.0,
  "enron_spam_setfit_strict_f1": 0.0,
  "phishing_darkknight_strict_f1": 0.0,
  "phishing_zefang_strict_f1": 0.0,
  "sms_spam_multilingual_strict_f1": 0.0,
  "spam_messages_mshenoda_strict_f1": 0.0,
  "spamassassin_strict_f1": 0.0
 },
 {
  "step": 41400,
  "epoch": 5.382215,
  "checkpoint_dir": "checkpoint-41400",
  "is_repo_root": true,
  "eval_loss": 0.000242,
  "pii4_mean_strict_f1": 0.570687,
  "avg11_strict_f1": 0.207523,
  "gliner2_pii_ptbr_reward_split_strict_f1": 0.707593,
  "nemotron_pii_strict_f1": 0.614901,
  "open_pii_masking_500k_strict_f1": 0.490177,
  "pii_masking_400k_strict_f1": 0.470078,
  "enron_spam_bvk_strict_f1": 0.0,
  "enron_spam_setfit_strict_f1": 0.0,
  "phishing_darkknight_strict_f1": 0.0,
  "phishing_zefang_strict_f1": 0.0,
  "sms_spam_multilingual_strict_f1": 0.0,
  "spam_messages_mshenoda_strict_f1": 0.0,
  "spamassassin_strict_f1": 0.0
 }
]

Comparison across the family

Same holdout protocol (strict F1, released checkpoint of each repo). Source: each repo's eval_holdout_strict_f1.parquet.

model released step PT-BR Nemotron (EN) open-pii-500k pii-masking-400k mean of 4
ettin-32m full-3x 65550 0.6326 0.4945 0.4081 0.4193 0.4886
ettin-68m full-3x (this) 41400 0.7076 0.6149 0.4902 0.4701 0.5707
ettin-68m top-50k 9945 0.2732 0.0345 0.0609 0.0755 0.1110
mmBERT-small full-3x 41400 0.7678 0.8158 0.7217 0.7076 0.7532
Privacy Filter fine-tune e3_c3 not evaluated with this protocol

Shared PT-BR val benchmark (24 canonical labels). None of the ettin repos in this family were run on it.

model partial P partial R partial F1 exact P exact R exact F1
Privacy Filter fine-tune e3_c3 0.917 0.879 0.897 0.871 0.835 0.853
mmBERT-small full-3x (41400) 0.951 0.832 0.888 0.932 0.815 0.870
ettin-68m easter-egg (not public) 0.893 0.761 0.822 0.857 0.731 0.789
ettin-32m easter-egg (not public) 0.905 0.729 0.808 0.862 0.694 0.769

Source: comparison_report.md in the Privacy Filter repo. 5,000-row PT-BR val set, restricted to the 24 canonical labels, nervaluate; GLiNER predictions at threshold 0.5, Privacy Filter with its default Viterbi decoding. The ettin models in this table are separate easter-egg fine-tunes, not the ettin repos of this family, which were not run on this benchmark.

Original sweep log (best by PT-BR partial F1 of each run). Reproduced from the April 2026 card; most of these runs (32m/68m top-100k, the 23k-step "full" runs, the easter-egg fine-tunes) are not published.

ettin-32m top-50k       : PT-BR strict 0.1411  partial 0.3710  Nemotron partial 0.1710  cross-avg 0.2808  (step 8500)
ettin-32m top-100k      : PT-BR strict 0.0946  partial 0.3165  Nemotron partial 0.1455  cross-avg 0.2428  (step 5200)
ettin-68m top-50k       : PT-BR strict 0.2732  partial 0.4802  Nemotron partial 0.2397  cross-avg 0.3117  (step 9945)
ettin-68m top-100k      : PT-BR strict 0.4109  partial 0.6191  Nemotron partial 0.3046  cross-avg 0.4497  (step 8580)
ettin-68m full          : PT-BR strict 0.3663  partial 0.5416  Nemotron partial 0.3563  cross-avg 0.4239  (step 19550)
ettin-32m full          : PT-BR strict 0.1429  partial 0.3536  Nemotron partial 0.2640  cross-avg 0.2867  (step 19550)
ettin-32m full-3x       : PT-BR strict 0.6326  partial 0.7167  Nemotron partial 0.5721  cross-avg 0.6031  (step 65550)
ettin-68m full-3x       : PT-BR strict 0.7076  partial 0.7979  Nemotron partial 0.6799  cross-avg 0.6817  (step 41400)
gliner-ettin-32m-ptbr-pii-easter-egg: PT-BR strict 0.6445  partial 0.7344  Nemotron partial 0.6019  cross-avg 0.6243  (step 100)
gliner-ettin-68m-ptbr-pii-easter-egg: PT-BR strict 0.7109  partial 0.8073  Nemotron partial 0.7182  cross-avg 0.7167  (step 100)  *** PT-BR best  *** cross-source best

Take-away. The 3x re-runs (69 000 steps) blew past the original 23 000-step sweep — the original 32M variants were straightforwardly under-trained, and the longer cosine schedule unlocks 1.5–2× higher F1 on every PII source. Curated subsets (top-50k / top-100k) overfit to PT-BR but trail badly on cross-source generalization; the full-data 3x runs dominate on every metric we care about. Among the ettin models, prefer gliner-ettin-{32m,68m}-ptbr-pii-full-3x-v1 — they sit at the Pareto frontier of in-domain F1 and cross-source generalization.

Since then, mmBERT-small full-3x beat both ettin full-3x models on every holdout source (table above).

Earlier result: ettin-68m full (23k-step run) vs prior baselines on Nemotron-PII

For the ettin-68m full checkpoint at step 17250, evaluated independently with eval_nemotron_prior.py (5000 samples, threshold=0.5, gold-label set, single-label flat NER):

model strict F1 partial F1
GLiNER multitask-large (prior) 0.5834 0.6490
mmbert-teacher (prior) 0.4574 0.5418
Albertina ckpt-12k (prior) 0.1801 0.3399
ettin-68m (prior) 0.1493 0.3210
ettin-68m full v1 step-17250 (this sweep) 0.2747 0.4306

i.e. +0.125 strict F1 / +0.110 partial F1 over the prior ettin-68m baseline at the same protocol — with this sweep's recipe (per-source label superset + bf16 mean-reduction + token_level fix).

Limitations

  • A research checkpoint trained and evaluated on synthetic or public PII datasets; there is no human-annotated real-world test set. Do not use it as the sole source of truth for redaction without human review.
  • PT-BR performance was the optimization target. The spam/phishing sources score 0 F1 by construction and do not measure abstention; this model's false-positive rate on text with no PII was not measured. The one GLiNER model of the family that was measured (mmBERT-small full-3x) flagged at least one false positive, mostly names, in 63.8% of 2,419 spam/phishing rows (benchmark dataset); expect similar behavior and raise the threshold when false positives are costly.
  • The partial-F1 and per-entity numbers cannot be re-derived from files in this repo (see Provenance).
  • The original sweep was limited to ~23k steps for 68M and ~12k–16k for the curated runs by single-MI300X budget; this 3x run (69,000 steps) is the follow-up.

License

Apache-2.0, same as the base encoder. Evaluation sources keep their own licenses (see their dataset cards).

Downloads last month
24
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1

Finetuned
(25)
this model

Collection including arthrod/gliner-ettin-68m-ptbr-pii-full-3x-v1