Instructions to use runelab-ai/kenning-turkish-legal-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use runelab-ai/kenning-turkish-legal-v1 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("runelab-ai/kenning-turkish-legal-v1") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
Kenning (kenning-turkish-legal-v1)
A key for Turkish legal texts: Kenning finds the right provision, clause or decision in a stack of documents.
Turkish legal retrieval embeddings, fine-tuned from EmbeddingGemma 2 by RuneLab.ai. Türkçe özet için: Türkçe
Summary
Kenning is an open-weight embedding model for retrieval over Turkish legal text: contracts, case law and regulation. It is a fine-tune of Google's EmbeddingGemma 2 (271M-parameter text model) and the first Turkish legal fine-tune of that model, released within a day of the base model.
- +2.50 nDCG@10 over the base model on the Mizan Turkish legal test sets (legal mean 48.44 → 50.94).
- On par with bge-m3 on legal retrieval (legal mean 50.94 vs 51.14), and ahead of it on Regulation (41.04 vs 38.04) and Caselaw (30.51 vs 29.30).
- General Turkish, English and multilingual performance on par with the base model.
- Ships with the official EmbeddingGemma 2 prompts and loads directly in Sentence Transformers.
Why the name
A kenning is an Old Norse poetic device that names a thing through a figurative compound, such as "whale-road" for the sea. An embedding model does something similar: it restates the meaning of a text in a new, compressed form.
Intended use
- Semantic search over Turkish contracts, court decisions and legislation or regulation.
- First-stage retriever in RAG pipelines over Turkish legal documents.
- Candidate retrieval for legal research and document review tools, with results checked by people.
Out-of-scope use
- Using retrieved results as legal advice or as the basis for automated legal decisions without human review.
- General-purpose Turkish embedding work; Kenning performs on par with the base model there, so expect no gain.
- Semantic similarity scoring (STS) as the main objective.
- Legal corpora in other languages or jurisdictions, which were not evaluated for this release.
Quickstart
pip install "transformers==5.19.0" "sentence-transformers==6.1.0" torchvision pillow # torchvision and pillow are required to load EmbeddingGemma 2 models
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("runelab-ai/kenning-turkish-legal-v1", truncate_dim=768) # 768d matches the evaluation setting
query_prompt = "task: search result | query: " # SearchQuery prompt: the prompt used for every score below. For natural-language questions, "task: question answering | query: " also works.
queries = [
"Kiracı kira bedelini ödemezse kiraya veren sözleşmeyi feshedebilir mi?",
]
documents = [
"Kiracı, kira bedelini veya yan giderleri ödeme borcunu ifa etmezse, kiraya veren kiracıya yazılı olarak bir süre verip bu süre içinde de ödeme yapılmazsa sözleşmeyi feshedeceğini bildirebilir.",
"Belirsiz süreli iş sözleşmesinin feshinden önce durumun diğer tarafa bildirilmesi gerekir.",
"Anonim şirketlerde genel kurul, yönetim kurulu tarafından toplantıya çağrılır.",
]
query_embeddings = model.encode(queries, prompt=query_prompt)
document_embeddings = model.encode_document(documents) # applies the shipped document prompt
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
Evaluation
Methodology. Mizan Turkish legal test sets; metric nDCG@10; Kenning and base embeddings at 768 dimensions; one uniform prompt (SearchQuery, task: search result | query: ) across all tasks; bge-m3 at its native dimension.
Turkish legal retrieval (nDCG@10)
| Task | EmbeddingGemma 2 base | Kenning | Δ vs base | bge-m3 (reference, native dim) |
|---|---|---|---|---|
| Legal mean | 48.44 | 50.94 | +2.50 | 51.14 |
| Contracts¹ | 76.73 | 81.27 | +4.54 | 86.08 |
| Caselaw | 29.67 | 30.51 | +0.84 | 29.30 |
| Regulation | 38.91 | 41.04 | +2.13 | 38.04 |
¹ The Contracts test set is built from a small number of source contracts, so this score is less representative than the other tasks.
General benchmarks (Kenning vs base)
| Benchmark | Base | Kenning | Δ |
|---|---|---|---|
| STSbTR | 78.89 | 77.73 | −1.16 |
| Turkish 14-task mean | 66.59 | 66.75 | +0.17 |
| English 10-task mean | 67.23 | 66.98 | −0.25 |
| Multilingual (6 tasks) | 75.93 | 76.21 | +0.27 |
Limitations
- Kenning is specialised for Turkish legal retrieval. General Turkish performance stays on par with the base model; it is not a general Turkish upgrade.
- The Contracts test set is built from a small number of source contracts, so the Contracts score is less representative than the other tasks.
- Similarity-style tasks such as STSbTR are not the target; Kenning is tuned for retrieval.
- bge-m3 scores are reported at its native dimension, while Kenning is reported at 768d.
Training details
- Base model: EmbeddingGemma 2 (271M-parameter text model), released October 6, 2026.
- Training data and recipe: Contrastive fine-tuning with Sentence Transformers (CachedMultipleNegativesRanking inside Matryoshka loss, batch 256, bf16, learning rate 2e-5, 1 epoch, best checkpoint chosen on internal dev sets), on about 17k pairs from: (1) Turkish Court of Cassation chamber decisions (catchword headings to decision text; public-domain
mrfg/turkish-court-decisions, CC0), (2) current consolidated Turkish laws and regulations from the Ministry of Justice legislation system, as heading-to-article pairs and as LLM-written questions with same-law sibling articles as hard negatives, (3) 160 LLM-drafted generic Turkish contracts (synthetic; no benchmark contract family) with clause-level questions and sibling-clause negatives, and (4) Turkish news article-summary pairs (bltlab/lr-sum, CC BY 4.0). Generated questions and hard negatives were filtered by automatic judges (a reranker for all of them, plus a second independent LLM judge for most). A self-anchor loss keeps embeddings close to the base model on multilingual sentences. The released weights are the fine-tuned weights averaged 50/50 with the base model weights. No benchmark test split, Mizan corpus, or ruling from the benchmark source families was used in training; exact-match contamination checks against the evaluation sets found no overlap. - Release date: October 7, 2026.
- License: Apache 2.0
Citation
@misc{runelab2026kenning,
title = {Kenning: A Turkish Legal Retrieval Fine-Tune of EmbeddingGemma 2},
author = {{RuneLab.ai}},
year = {2026},
howpublished = {\url{https://huggingface.co/runelab-ai/kenning-turkish-legal-v1}}
}
Please also cite EmbeddingGemma 2 and the Mizan benchmark as their authors request.
About RuneLab.ai NexTech
RuneLab.ai is a Turkey-based AI company.
- NexTech: enterprise AI strategy and end-to-end projects.
- RuneLab Academy: AI training for individuals and teams.
- Products: RuneDesk (document assistant), RuneBox (on-premise AI infrastructure), RuneGate (model gateway) and Vitki (adaptive learning, early access).
More at huggingface.co/runelab-ai · info@runelab.ai · www.runelab.ai
Türkçe
Özet
Kenning, Google'ın EmbeddingGemma 2 modelinin Türkçe hukuki metinlerde arama (retrieval) için fine-tune edilmiş, açık ağırlıklı sürümüdür. Temel modelin yayımlanmasından bir gün içinde çıkan Kenning, EmbeddingGemma 2'nin Türk hukuku için yapılmış ilk fine-tune'udur.
- Mizan Türkçe hukuk test setlerinde temel modele göre ortalama +2.50 nDCG@10 iyileşme sağlar.
- Hukuki retrieval'da bge-m3 ile başa baştır; mevzuat ve içtihat görevlerinde bge-m3'ün önündedir.
- Genel Türkçe, İngilizce ve çok dilli performansı temel modelle aynı seviyededir.
Kullanım alanları
- Türkçe sözleşme, mahkeme kararı ve mevzuat metinleri üzerinde anlamsal arama
- Hukuki RAG sistemlerinde ilk aşama retriever olarak kullanım
- Hukuk araştırma ve doküman inceleme araçlarında aday belge getirme (sonuçlar insan kontrolünden geçmelidir)
Kapsam dışı
- Arama sonuçlarının insan incelemesi olmadan hukuki tavsiye ya da otomatik hukuki karar için kullanılması
- Genel amaçlı Türkçe embedding işleri: bu görevlerde Kenning temel modelle aynı seviyededir, ek kazanç beklenmemelidir
- Anlamsal benzerlik (STS) skorlamasını ana hedef alan kullanımlar
- Başka dillerdeki veya başka hukuk sistemlerine ait metinler (bu sürümde değerlendirilmedi)
Sonuçlar
Görev bazında tüm skorlar, değerlendirme kurulumu ve bge-m3 karşılaştırması yukarıdaki Evaluation bölümündedir.
Sınırlamalar
- Kenning, Türkçe hukuki retrieval için uzmanlaşmıştır. Genel Türkçe performansı temel modelle aynı seviyededir; genel bir Türkçe iyileştirmesi değildir.
- Sözleşmeler (Contracts) test seti az sayıda kaynak sözleşmeden oluşturulmuştur; bu nedenle bu görevin skoru diğer görevlere göre daha az temsil edicidir.
- STSbTR gibi benzerlik görevleri hedef değildir; Kenning retrieval için ayarlanmıştır.
- bge-m3 skorları kendi doğal boyutunda, Kenning skorları ise 768 boyutta raporlanmıştır.
Hızlı başlangıç
Kod örneği yukarıdaki Quickstart bölümündedir; örnek sorgu ve belgeler Türkçedir.
RuneLab.ai NexTech hakkında
RuneLab.ai, Türkiye merkezli bir yapay zeka şirketidir.
- NexTech: kurumsal yapay zeka stratejisi ve uçtan uca projeler.
- RuneLab Akademi: bireylere ve ekiplere yapay zeka eÄŸitimi.
- Ürünler: RuneDesk (doküman asistanı), RuneBox (kurum içi yapay zeka altyapısı), RuneGate (model geçidi) ve Vitki (adaptif öğrenme, erken erişim).
Daha fazlası: huggingface.co/runelab-ai · info@runelab.ai · www.runelab.ai
- Downloads last month
- 48
Model tree for runelab-ai/kenning-turkish-legal-v1
Base model
google/embeddinggemma-2Evaluation results
- nDCG@10 (768d) on Mizan Legal (mean of Contracts, Caselaw, Regulation)test set self-reported50.940
- nDCG@10 (768d) on Mizan Contractstest set self-reported81.270
- nDCG@10 (768d) on Mizan Caselawtest set self-reported30.510
- nDCG@10 (768d) on Mizan Regulationtest set self-reported41.040