Instructions to use AhiskaAI/AhiskaAI-v0.5-440m-IT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AhiskaAI/AhiskaAI-v0.5-440m-IT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AhiskaAI/AhiskaAI-v0.5-440m-IT") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AhiskaAI/AhiskaAI-v0.5-440m-IT") model = AutoModelForCausalLM.from_pretrained("AhiskaAI/AhiskaAI-v0.5-440m-IT", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AhiskaAI/AhiskaAI-v0.5-440m-IT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AhiskaAI/AhiskaAI-v0.5-440m-IT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AhiskaAI/AhiskaAI-v0.5-440m-IT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AhiskaAI/AhiskaAI-v0.5-440m-IT
- SGLang
How to use AhiskaAI/AhiskaAI-v0.5-440m-IT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AhiskaAI/AhiskaAI-v0.5-440m-IT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AhiskaAI/AhiskaAI-v0.5-440m-IT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AhiskaAI/AhiskaAI-v0.5-440m-IT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AhiskaAI/AhiskaAI-v0.5-440m-IT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AhiskaAI/AhiskaAI-v0.5-440m-IT with Docker Model Runner:
docker model run hf.co/AhiskaAI/AhiskaAI-v0.5-440m-IT
- AhıskaAI v0.5 440M IT
- Türkçe
AhıskaAI v0.5 440M IT
(Türkçe Açıklama Aşağıda)
AhıskaAI v0.5 Series
AhıskaAI is an independent open-source AI research project focused on small language models, custom tokenization, Turkish and Ahıska Turkish, and reproducible low-resource AI experiments.
The AhıskaAI v0.5 series continues this work with larger pretraining data and a cleaner, standardized training pipeline, while keeping the models lightweight and accessible.
This release is the instruction-tuned (IT) version of AhıskaAI v0.5 440M Base — the largest model in the AhıskaAI v0.5 series so far.
AhıskaAI v0.5 440M IT
AhıskaAI v0.5 440M IT is an instruction-tuned language model designed for Turkish question answering, instruction following, text transformation, and conversational experiments.
The underlying AhıskaAI v0.5 440M Base model was trained completely from scratch. This version was then instruction-tuned using a ChatML-style conversation format (system, user, assistant roles).
This is an experimental small language model developed as part of the AhıskaAI research project.
What's New Compared to v0.4
- More pretraining data: The Base model was trained on ~5.1B tokens of FineWeb-2 HQ Turkish, compared to ~1.7B tokens for AhıskaAI v0.4 235M Base.
- Larger model: With ~440M parameters, it is nearly twice the size of v0.4 235M IT, the largest model of the v0.4 series.
- Better benchmark results than v0.4: It scores higher overall on the AhıskaAI v0.1 Lite Benchmark than every v0.4 model, and has the highest Correctness score among all AhıskaAI models on this benchmark (24.24%). Its overall score is slightly below that of the smaller v0.5 120M IT.
- New tokenizer: A 24K BPE tokenizer trained separately for v0.5 on its own data. It is not interchangeable with the v0.4 tokenizer.
- Cleaner SFT data: Instruction-tuned on the current, cleaned, standard AhıskaAI SFT dataset, instead of the shared SFT mixture used in v0.4 (which included part of Ethosoft/nedo-turkish-sft-mixtures).
- Trade-off: The maximum context length is 1024 tokens (2048 in v0.4 235M).
| Benchmark metric | v0.4 145M IT (New) | v0.4 235M IT | v0.5 120M IT | v0.5 440M IT |
|---|---|---|---|---|
| Overall | 54.88% | 51.18% | 57.91% | 57.58% |
| Turkish Fluency | 97.98% | 96.97% | 98.99% | 97.98% |
| Topic Relevance | 44.44% | 38.38% | 55.56% | 50.51% |
| Correctness | 22.22% | 18.18% | 19.19% | 24.24% |
Why Is the Performance Close to the Smaller Models?
Although AhıskaAI v0.5 440M IT has roughly 3.7 times the parameters of v0.5 120M IT, their benchmark scores are very close (57.58% vs 57.91% overall), and it also stays close to the larger v0.4 models. We attribute this mainly to the following:
- Limited pretraining data: The entire v0.5 series was pretrained on ~5.1B tokens of FineWeb-2 HQ Turkish. For the 440M model this is only ~12 tokens per parameter (about 43 for the 120M model), so the larger model remained relatively undertrained and could not fully use its additional capacity.
- Shallow architecture by design: Both v0.5 models use 16 layers. The layer count was deliberately kept low to increase training MFU (Model FLOPs Utilization); the 440M model gets its size from width (hidden size 1536, intermediate size 4096) rather than depth.
- Limited training budget: The 440M Base model was pretrained on a Kaggle TPU v5e-8 in about 15 hours 13 minutes (the 120M Base took about 7 hours 21 minutes).
- Small benchmark: The benchmark has only 99 questions, so differences of a fraction of a point can come down to a single question.
Model Specifications
| Specification | Value |
|---|---|
| Model name | AhıskaAI v0.5 440M IT |
| Parameters | ~440M |
| Architecture | LlamaForCausalLM |
| Hidden size | 1536 |
| Hidden layers | 16 |
| Attention heads | 12 |
| Key/Value heads | 4 |
| Intermediate size | 4096 |
| Head dimension | 128 |
| Maximum context length | 1024 |
| Vocabulary size | 24,000 |
| Activation | SiLU |
| Data type | bfloat16 |
| Position encoding | RoPE (theta 10,000) |
| Word embeddings | Tied |
| Attention dropout | 0.0 |
| Attention bias | None |
| MLP bias | None |
The model uses the LlamaForCausalLM architecture, but the underlying Base model was trained completely from scratch rather than initialized from the weights of an existing Llama model.
Training
Pretraining
The underlying AhıskaAI v0.5 440M Base model was trained completely from scratch on FineWeb-2 HQ Turkish, using approximately 5.1B tokens. It was not initialized from, or fine-tuned from, any other pretrained model.
- Pretraining hardware: Kaggle TPU v5e-8
- Pretraining time: ~15 hours 13 minutes
The number of layers was deliberately kept low to increase training MFU (Model FLOPs Utilization).
Instruction-Tuning
After Base model pretraining, this version received an instruction-tuning stage using the current, cleaned, standard AhıskaAI SFT dataset. This dataset is not publicly released.
- Hardware: TBD
- Training time: TBD
- Per-device batch size: TBD
- Gradient accumulation steps: TBD
The instruction-tuning data uses a ChatML-style conversation format consisting of system, user, and assistant roles:
<|im_start|>system
Sen kibar, sorulan soruları tam cümlelerle yanıtlayan Türkçe bir asistansın.<|im_end|>
<|im_start|>user
Aşağıdaki cümleyi daha okunaklı hale getirmek için metin biçimlendirmesini değiştirin. Öğrencinin tamamlaması gereken karmaşık bir problemi vardı.<|im_end|>
<|im_start|>assistant
Öğrencinin tamamlaması gereken karmaşık bir problem vardı.<|im_end|>
This format was used to teach the model to respond to user instructions while maintaining a conversational structure. The instruction-tuning stage focused on improving instruction following, question answering, text transformation, conversational responses, and Turkish response quality.
Tokenizer
AhıskaAI v0.5 uses a custom 24K vocabulary BPE tokenizer focused on Turkish, Ahıska Turkish, and Turkish morphology. This tokenizer was trained separately for the v0.5 series on its own training data. Although its vocabulary size matches the v0.4 tokenizer, it is a different tokenizer and is not interchangeable with v0.4 models.
Intended Use
AhıskaAI v0.5 440M IT can be used for:
- Turkish question answering
- Instruction following
- Text transformation
- Simple conversational applications
- Turkish text generation
- Small language model research
- Low-resource language experiments
- Instruction-tuning research on small models
The model is particularly intended for experimentation and research rather than production-critical applications.
How to Run
pip install -U transformers torch
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "AhiskaAI/AhiskaAI-v0.5-440M-IT"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=dtype
).to(device)
messages = [
{
"role": "system",
"content": "Sen kibar, sorulan soruları tam cümlelerle yanıtlayan Türkçe bir asistansın."
},
{
"role": "user",
"content": "Ahıska Türkleri hakkında kısa bir bilgi ver."
}
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
inputs = tokenizer(prompt, return_tensors="pt").to(device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=100,
temperature=0.8,
top_p=0.9,
do_sample=True,
repetition_penalty=1.1,
use_cache=True
)
generated_tokens = outputs[0][inputs["input_ids"].shape[1]:]
print(tokenizer.decode(generated_tokens, skip_special_tokens=True))
Manual Prompt Format
<|im_start|>system
Sen kibar, sorulan soruları tam cümlelerle yanıtlayan Türkçe bir asistansın.<|im_end|>
<|im_start|>user
Kullanıcı sorusu buraya gelir.<|im_end|>
<|im_start|>assistant
Generation should begin after the assistant header.
AhıskaAI v0.1 Lite Benchmark
AhıskaAI v0.1 Lite Benchmark is an experimental benchmark evaluating Turkish quality, topic relevance, and factual/logical correctness (binary scoring, averaged).
Results for AhıskaAI v0.5 440M IT:
- Turkish Fluency: 97.98%
- Topic Relevance: 50.51%
- Correctness: 24.24%
- Overall Average: 57.58%
- Total Questions: 99
Note: AhıskaAI v0.1 Lite Benchmark is an experimental benchmark and should not be considered a standardized benchmark.
Overall Benchmark
| Rank | Model | Score |
|---|---|---|
| 1 | Gemma 3 1B | 85.86% |
| 2 | AhıskaAI v0.5 120M IT | 57.91% |
| 3 | AhıskaAI v0.5 440M IT | 57.58% |
| 4 | Qwen 3 0.6B | 55.22% |
| 5 | AhıskaAI v0.4 145M IT (New) | 54.88% |
| 6 | AhıskaAI v0.4 235M IT | 51.18% |
| 7 | AhıskaAI v0.4 145M IT (Old) | 50.17% |
| 8 | Gemma 3 270M | 46.46% |
| 9 | Qwen 2.5 0.5B | 43.77% |
| 10 | AhıskaAI v0.4 35M | 43.43% |
| 11 | Llama 3.2 1B | 35.35% |
| 12 | SmolLM2 1.7B | 22.89% |
| 13 | SmolLM2 360M | 12.12% |
| 14 | SmolLM2 135M | 1.01% |
Turkish Performance
| Rank | Model | Score |
|---|---|---|
| 1 | AhıskaAI v0.4 145M IT (Old) | 100.00% |
| 1 | Gemma 3 1B | 100.00% |
| 3 | AhıskaAI v0.5 120M IT | 98.99% |
| 4 | AhıskaAI v0.4 145M IT (New) | 97.98% |
| 4 | AhıskaAI v0.5 440M IT | 97.98% |
| 6 | AhıskaAI v0.4 235M IT | 96.97% |
| 6 | Gemma 3 270M | 96.97% |
| 8 | AhıskaAI v0.4 35M | 87.88% |
| 9 | Qwen 3 0.6B | 86.87% |
| 10 | Qwen 2.5 0.5B | 81.82% |
| 11 | Llama 3.2 1B | 48.48% |
| 12 | SmolLM2 1.7B | 46.46% |
| 13 | SmolLM2 360M | 28.28% |
| 14 | SmolLM2 135M | 3.03% |
Topic Relevance
| Rank | Model | Score |
|---|---|---|
| 1 | Gemma 3 1B | 84.85% |
| 2 | AhıskaAI v0.5 120M IT | 55.56% |
| 3 | AhıskaAI v0.5 440M IT | 50.51% |
| 4 | Qwen 3 0.6B | 47.47% |
| 5 | AhıskaAI v0.4 145M IT (New) | 44.44% |
| 6 | Llama 3.2 1B | 40.40% |
| 7 | AhıskaAI v0.4 235M IT | 38.38% |
| 8 | AhıskaAI v0.4 145M IT (Old) | 36.36% |
| 9 | Qwen 2.5 0.5B | 33.33% |
| 10 | AhıskaAI v0.4 35M | 32.32% |
| 11 | Gemma 3 270M | 26.26% |
| 12 | SmolLM2 1.7B | 13.13% |
| 13 | SmolLM2 360M | 5.05% |
| 14 | SmolLM2 135M | 0.00% |
Correctness
| Rank | Model | Score |
|---|---|---|
| 1 | Gemma 3 1B | 72.73% |
| 2 | Qwen 3 0.6B | 31.31% |
| 3 | AhıskaAI v0.5 440M IT | 24.24% |
| 4 | AhıskaAI v0.4 145M IT (New) | 22.22% |
| 5 | AhıskaAI v0.5 120M IT | 19.19% |
| 6 | AhıskaAI v0.4 235M IT | 18.18% |
| 7 | Llama 3.2 1B | 17.17% |
| 8 | Qwen 2.5 0.5B | 16.16% |
| 8 | Gemma 3 270M | 16.16% |
| 10 | AhıskaAI v0.4 145M IT (Old) | 14.14% |
| 11 | AhıskaAI v0.4 35M | 10.10% |
| 12 | SmolLM2 1.7B | 9.09% |
| 13 | SmolLM2 360M | 3.03% |
| 14 | SmolLM2 135M | 0.00% |
Limitations
This model is intentionally small and experimental.
Because of its size, it may:
- Produce factually incorrect information
- Generate meaningless or incomplete text
- Fail to follow complex instructions
- Lose context in longer prompts (maximum context is 1024 tokens)
- Repeat words or phrases
- Produce off-topic responses
- Struggle with reasoning and multi-step tasks
- Hallucinate information
Instruction tuning improves the model's usability and instruction-following behavior, but it does not eliminate these limitations. In particular, the underlying Base model was pretrained on ~5.1B tokens, which is relatively few for its parameter count (roughly 12 tokens per parameter). This may affect knowledge coverage and fluency relative to what its size alone would suggest.
Do not rely on this model for medical, legal, financial, safety-critical, or other high-stakes decisions.
Base Model
This model is instruction-tuned from:
The Base model was trained completely from scratch (~5.1B tokens of FineWeb-2 HQ Turkish) before the instruction-tuning stage.
Other Versions
- AhıskaAI v0.5 440M Base — Base language model underlying this release
- AhıskaAI v0.5 440M IT — This model
- AhıskaAI v0.5 120M Base — Smaller Base model
- AhıskaAI v0.5 120M IT — Instruction-tuned 120M model
- AhıskaAI v0.4 Series — Previous generation (35M, 145M, 235M)
License
This model is released under the Apache 2.0 License.
If you use this model in a project, publication, application, or other public work, attribution to AhıskaAI is appreciated.
Suggested Citation
AhıskaAI v0.5 440M IT — An instruction-tuned small language model for Turkish and Ahıska Turkish research.
Project: AhıskaAI Model: AhıskaAI v0.5 440M IT
Türkçe
AhıskaAI v0.5 Serisi
AhıskaAI; küçük dil modelleri, özel tokenizasyon, Türkçe ve Ahıska Türkçesi ve düşük kaynaklı yapay zekâ deneyleri üzerine çalışan bağımsız, açık kaynaklı bir AI araştırma projesidir.
AhıskaAI v0.5 serisi, daha büyük ön eğitim verisi ve daha temiz, standartlaştırılmış bir eğitim süreciyle bu çalışmayı sürdürürken modelleri hafif ve erişilebilir tutmayı amaçlamaktadır.
Bu sürüm, AhıskaAI v0.5 serisinin şu ana kadarki en büyük modeli olan AhıskaAI v0.5 440M Base modelinin instruction-tuned (IT) sürümüdür.
AhıskaAI v0.5 440M IT
AhıskaAI v0.5 440M IT, Türkçe soru-cevap, talimat takibi, metin dönüştürme ve sohbet deneyleri için geliştirilmiş instruction-tuned bir dil modelidir.
Modelin temelini oluşturan AhıskaAI v0.5 440M Base modeli tamamen sıfırdan eğitilmiştir. Bu sürüm ise daha sonra ChatML tarzı bir konuşma formatı (system, user, assistant rolleri) ile instruction tuning sürecinden geçirilmiştir.
Bu model, AhıskaAI araştırma projesinin bir parçası olarak geliştirilmiş deneysel bir küçük dil modelidir.
v0.4'e Göre Yenilikler
- Daha fazla ön eğitim verisi: Base model, ~5.1B token FineWeb-2 HQ Turkish verisiyle eğitilmiştir. AhıskaAI v0.4 235M Base'de bu miktar ~1.7B token idi.
- Daha büyük model: ~440M parametre ile, v0.4 serisinin en büyük modeli olan v0.4 235M IT'nin yaklaşık iki katı büyüklüktedir.
- v0.4'ten daha iyi benchmark sonuçları: AhıskaAI v0.1 Lite Benchmark'ta genel skorda tüm v0.4 modellerini geride bırakmıştır ve bu benchmark'ta tüm AhıskaAI modelleri arasında en yüksek Doğruluk skoruna (%24.24) sahiptir. Genel skoru, daha küçük olan v0.5 120M IT'nin biraz altındadır.
- Yeni tokenizer: v0.5 için kendi verisi üzerinde ayrıca eğitilmiş 24K BPE tokenizer. v0.4 tokenizer'ı ile birbirinin yerine kullanılamaz.
- Daha temiz SFT verisi: v0.4'te kullanılan ortak SFT karışımı (Ethosoft/nedo-turkish-sft-mixtures'ın bir kısmını da içeriyordu) yerine güncel, temizlenmiş ve standartlaştırılmış AhıskaAI SFT veri seti ile instruction tuning uygulanmıştır.
- Takas: Maksimum context uzunluğu 1024 token'dır (v0.4 235M'de 2048).
| Benchmark metriği | v0.4 145M IT (New) | v0.4 235M IT | v0.5 120M IT | v0.5 440M IT |
|---|---|---|---|---|
| Genel | 54.88% | 51.18% | 57.91% | 57.58% |
| Türkçe Düzgünlüğü | 97.98% | 96.97% | 98.99% | 97.98% |
| Konu Alakası | 44.44% | 38.38% | 55.56% | 50.51% |
| Doğruluk | 22.22% | 18.18% | 19.19% | 24.24% |
Performans Neden Daha Küçük Modellere Yakın?
AhıskaAI v0.5 440M IT, v0.5 120M IT'nin yaklaşık 3.7 katı parametreye sahip olmasına rağmen benchmark skorları birbirine çok yakındır (genel skorda %57.58 ve %57.91) ve daha büyük v0.4 modellerine de yakın kalmaktadır. Bunu temel olarak şu nedenlere bağlıyoruz:
- Sınırlı ön eğitim verisi: v0.5 serisinin tamamı ~5.1B token FineWeb-2 HQ Turkish verisiyle ön eğitim görmüştür. 440M model için bu, parametre başına yalnızca ~12 token'dır (120M model için yaklaşık 43). Bu nedenle büyük model görece az eğitilmiş kalmış ve ek kapasitesini tam olarak kullanamamıştır.
- Bilinçli olarak sığ mimari: Her iki v0.5 modeli de 16 katman kullanmaktadır. Katman sayısı, eğitim MFU'sunu (Model FLOPs Utilization) artırmak amacıyla bilinçli olarak düşük tutulmuştur; 440M model boyutunu derinlikten değil genişlikten (hidden size 1536, intermediate size 4096) almaktadır.
- Sınırlı eğitim bütçesi: 440M Base model, Kaggle TPU v5e-8 üzerinde yaklaşık 15 saat 13 dakikada ön eğitimden geçirilmiştir (120M Base için yaklaşık 7 saat 21 dakika).
- Küçük benchmark: Benchmark yalnızca 99 sorudan oluşmaktadır; bu yüzden bir puanın küçük bir kısmı kadar farklar tek bir soruya bağlı olabilir.
Model Özellikleri
| Özellik | Değer |
|---|---|
| Model adı | AhıskaAI v0.5 440M IT |
| Parametre sayısı | ~440M |
| Mimari | LlamaForCausalLM |
| Hidden size | 1536 |
| Hidden layers | 16 |
| Attention heads | 12 |
| Key/Value heads | 4 |
| Intermediate size | 4096 |
| Head dimension | 128 |
| Maksimum context | 1024 |
| Vocabulary | 24,000 |
| Aktivasyon | SiLU |
| Veri tipi | bfloat16 |
| Position encoding | RoPE (theta 10.000) |
| Word embeddings | Tied |
| Attention dropout | 0.0 |
| Attention bias | Yok |
| MLP bias | Yok |
Model LlamaForCausalLM mimarisini kullanmaktadır ancak temel model mevcut bir Llama modelinin ağırlıkları kullanılarak oluşturulmamıştır. Base model tamamen sıfırdan eğitilmiştir.
Eğitim
Ön Eğitim
Modelin temelini oluşturan AhıskaAI v0.5 440M Base, FineWeb-2 HQ Turkish verisi üzerinde yaklaşık 5.1B token kullanılarak tamamen sıfırdan eğitilmiştir. Başka bir pretrained modelden başlatılmamış veya fine-tune edilmemiştir.
- Ön eğitim donanımı: Kaggle TPU v5e-8
- Ön eğitim süresi: ~15 saat 13 dakika
Katman sayısı, eğitim sırasındaki MFU (Model FLOPs Utilization) değerini artırmak amacıyla bilinçli olarak düşük tutulmuştur.
Instruction-Tuning
Base model ön eğitiminin ardından bu sürüme, güncel, temizlenmiş ve standartlaştırılmış AhıskaAI SFT veri seti ile instruction tuning uygulanmıştır. Bu veri seti herkese açık olarak yayınlanmamaktadır.
- Donanım: TBD
- Eğitim süresi: TBD
- Per-device batch size: TBD
- Gradient accumulation steps: TBD
Instruction tuning verileri, system, user ve assistant rollerinden oluşan ChatML tarzı bir konuşma formatı kullanmaktadır:
<|im_start|>system
Sen kibar, sorulan soruları tam cümlelerle yanıtlayan Türkçe bir asistansın.<|im_end|>
<|im_start|>user
Aşağıdaki cümleyi daha okunaklı hale getirmek için metin biçimlendirmesini değiştirin. Öğrencinin tamamlaması gereken karmaşık bir problemi vardı.<|im_end|>
<|im_start|>assistant
Öğrencinin tamamlaması gereken karmaşık bir problem vardı.<|im_end|>
Bu format kullanılarak modelin kullanıcı talimatlarına uygun cevaplar üretmesi ve konuşma yapısını takip etmesi hedeflenmiştir. Instruction tuning aşaması özellikle talimat takibi, soru-cevap, metin dönüştürme, sohbet tarzı cevaplar ve Türkçe cevap kalitesi üzerine odaklanmıştır.
Tokenizer
AhıskaAI v0.5 serisinde, Türkçe/Ahıska Türkçesi ve Türkçe morfolojisine odaklı özel bir 24K vocabulary BPE tokenizer kullanılmaktadır. Bu tokenizer, v0.5 serisi için kendi eğitim verisi üzerinde ayrıca eğitilmiştir. Vocabulary boyutu v0.4 tokenizer'ı ile aynı olsa da farklı bir tokenizer'dır ve v0.4 modelleriyle birbirinin yerine kullanılamaz.
Kullanım Alanları
AhıskaAI v0.5 440M IT şu alanlarda kullanılabilir:
- Türkçe soru-cevap
- Talimat takibi
- Metin dönüştürme
- Basit sohbet uygulamaları
- Türkçe metin üretimi
- Küçük dil modeli araştırmaları
- Düşük kaynaklı dil deneyleri
- Küçük modeller üzerinde instruction tuning araştırmaları
Model öncelikli olarak deneysel ve araştırma amaçlıdır.
Çalıştırma
pip install -U transformers torch
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "AhiskaAI/AhiskaAI-v0.5-440M-IT"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=dtype
).to(device)
messages = [
{
"role": "system",
"content": "Sen kibar, sorulan soruları tam cümlelerle yanıtlayan Türkçe bir asistansın."
},
{
"role": "user",
"content": "Ahıska Türkleri hakkında kısa bir bilgi ver."
}
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
inputs = tokenizer(prompt, return_tensors="pt").to(device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=100,
temperature=0.8,
top_p=0.9,
do_sample=True,
repetition_penalty=1.1,
use_cache=True
)
generated_tokens = outputs[0][inputs["input_ids"].shape[1]:]
print(tokenizer.decode(generated_tokens, skip_special_tokens=True))
Manuel Prompt Formatı
<|im_start|>system
Sen kibar, sorulan soruları tam cümlelerle yanıtlayan Türkçe bir asistansın.<|im_end|>
<|im_start|>user
Kullanıcı sorusu buraya gelir.<|im_end|>
<|im_start|>assistant
Üretim işlemi assistant başlığından sonra başlatılmalıdır.
AhıskaAI v0.1 Lite Benchmark
Türkçe kalitesi, konu ilgisi ve bilgisel/mantıksal doğruluğu değerlendiren deneysel bir benchmark (binary scoring, ortalama).
AhıskaAI v0.5 440M IT sonuçları:
- Türkçe Düzgünlüğü: %97.98
- Konu Alakası: %50.51
- Doğruluk: %24.24
- Genel Ortalama: %57.58
- Toplam Soru: 99
Not: Bu deneysel bir benchmark'tır, standartlaştırılmış bir benchmark olarak değerlendirilmemelidir.
Genel Benchmark
| Sıra | Model | Skor |
|---|---|---|
| 1 | Gemma 3 1B | 85.86% |
| 2 | AhıskaAI v0.5 120M IT | 57.91% |
| 3 | AhıskaAI v0.5 440M IT | 57.58% |
| 4 | Qwen 3 0.6B | 55.22% |
| 5 | AhıskaAI v0.4 145M IT (New) | 54.88% |
| 6 | AhıskaAI v0.4 235M IT | 51.18% |
| 7 | AhıskaAI v0.4 145M IT (Old) | 50.17% |
| 8 | Gemma 3 270M | 46.46% |
| 9 | Qwen 2.5 0.5B | 43.77% |
| 10 | AhıskaAI v0.4 35M | 43.43% |
| 11 | Llama 3.2 1B | 35.35% |
| 12 | SmolLM2 1.7B | 22.89% |
| 13 | SmolLM2 360M | 12.12% |
| 14 | SmolLM2 135M | 1.01% |
Türkçe Performansı
| Sıra | Model | Skor |
|---|---|---|
| 1 | AhıskaAI v0.4 145M IT (Old) | 100.00% |
| 1 | Gemma 3 1B | 100.00% |
| 3 | AhıskaAI v0.5 120M IT | 98.99% |
| 4 | AhıskaAI v0.4 145M IT (New) | 97.98% |
| 4 | AhıskaAI v0.5 440M IT | 97.98% |
| 6 | AhıskaAI v0.4 235M IT | 96.97% |
| 6 | Gemma 3 270M | 96.97% |
| 8 | AhıskaAI v0.4 35M | 87.88% |
| 9 | Qwen 3 0.6B | 86.87% |
| 10 | Qwen 2.5 0.5B | 81.82% |
| 11 | Llama 3.2 1B | 48.48% |
| 12 | SmolLM2 1.7B | 46.46% |
| 13 | SmolLM2 360M | 28.28% |
| 14 | SmolLM2 135M | 3.03% |
Konu İlgisi
| Sıra | Model | Skor |
|---|---|---|
| 1 | Gemma 3 1B | 84.85% |
| 2 | AhıskaAI v0.5 120M IT | 55.56% |
| 3 | AhıskaAI v0.5 440M IT | 50.51% |
| 4 | Qwen 3 0.6B | 47.47% |
| 5 | AhıskaAI v0.4 145M IT (New) | 44.44% |
| 6 | Llama 3.2 1B | 40.40% |
| 7 | AhıskaAI v0.4 235M IT | 38.38% |
| 8 | AhıskaAI v0.4 145M IT (Old) | 36.36% |
| 9 | Qwen 2.5 0.5B | 33.33% |
| 10 | AhıskaAI v0.4 35M | 32.32% |
| 11 | Gemma 3 270M | 26.26% |
| 12 | SmolLM2 1.7B | 13.13% |
| 13 | SmolLM2 360M | 5.05% |
| 14 | SmolLM2 135M | 0.00% |
Doğruluk
| Sıra | Model | Skor |
|---|---|---|
| 1 | Gemma 3 1B | 72.73% |
| 2 | Qwen 3 0.6B | 31.31% |
| 3 | AhıskaAI v0.5 440M IT | 24.24% |
| 4 | AhıskaAI v0.4 145M IT (New) | 22.22% |
| 5 | AhıskaAI v0.5 120M IT | 19.19% |
| 6 | AhıskaAI v0.4 235M IT | 18.18% |
| 7 | Llama 3.2 1B | 17.17% |
| 8 | Qwen 2.5 0.5B | 16.16% |
| 8 | Gemma 3 270M | 16.16% |
| 10 | AhıskaAI v0.4 145M IT (Old) | 14.14% |
| 11 | AhıskaAI v0.4 35M | 10.10% |
| 12 | SmolLM2 1.7B | 9.09% |
| 13 | SmolLM2 360M | 3.03% |
| 14 | SmolLM2 135M | 0.00% |
Sınırlamalar
Bu model özellikle küçük ve deneysel bir modeldir.
Model:
- Yanlış bilgiler üretebilir
- Anlamsız veya eksik metinler oluşturabilir
- Karmaşık talimatları takip etmekte zorlanabilir
- Uzun promptlarda bağlamı kaybedebilir (maksimum context 1024 token'dır)
- Kelime veya ifadeleri tekrarlayabilir
- Konu dışı cevaplar verebilir
- Mantıksal çıkarım ve çok adımlı görevlerde zorlanabilir
- Halüsinasyon üretebilir
Instruction tuning modelin kullanılabilirliğini ve talimat takip yeteneğini geliştirmeyi amaçlamaktadır ancak bu sınırlamaları tamamen ortadan kaldırmaz. Özellikle, temelindeki Base model parametre sayısına göre görece az token ile (~5.1B token, parametre başına yaklaşık 12 token) ön eğitim görmüştür. Bu durum, bilgi kapsamını ve akıcılığı, boyutunun tek başına işaret ettiği seviyeye göre etkileyebilir.
Model tıbbi, hukuki, finansal, güvenlik açısından kritik veya yüksek riskli kararlar için kullanılmamalıdır.
Base Model
Bu model şu temel model üzerine instruction tuning uygulanarak oluşturulmuştur:
Base model, instruction tuning aşamasından önce tamamen sıfırdan eğitilmiştir (FineWeb-2 HQ Turkish üzerinde ~5.1B token).
Diğer Sürümler
- AhıskaAI v0.5 440M Base — Bu sürümün temelini oluşturan Base dil modeli
- AhıskaAI v0.5 440M IT — Bu model
- AhıskaAI v0.5 120M Base — Daha küçük Base model
- AhıskaAI v0.5 120M IT — 120M instruction-tuned model
- AhıskaAI v0.4 Serisi — Önceki nesil (35M, 145M, 235M)
Lisans
Bu model Apache 2.0 License altında yayınlanmıştır.
Bu modeli bir proje, yayın, uygulama veya başka bir kamuya açık çalışmada kullanıyorsanız AhıskaAI'ye atıfta bulunmanız rica edilir.
Önerilen Atıf
AhıskaAI v0.5 440M IT — Türkçe ve Ahıska Türkçesi araştırmaları için geliştirilmiş instruction-tuned küçük dil modeli.
Proje: AhıskaAI Model: AhıskaAI v0.5 440M IT
- Downloads last month
- -
Model tree for AhiskaAI/AhiskaAI-v0.5-440m-IT
Base model
AhiskaAI/AhiskaAI-v0.5-440M-Base