WSKAN logo

wskan-1m-wikitext

985,956 parameters · WSKAN-11 · byte-level LM trained on WikiText-103 (local-coherence tier)

Introduction

wskan-1m-wikitext is a checkpoint of WSKAN-11 (Wavelet-like State-Space KAN, version 11), a Kolmogorov-Arnold Network in which every edge function is the impulse response of a stable state-space model — a wavelet-like damped oscillator. One parameterization supports two modes: a closed-form wavelet edge (static, WavKAN-style) and a recurrent SSM edge (Mamba-style selective scan); this checkpoint runs the recurrent mode through a fused Triton associative-scan kernel.

This model is one cell of the 3-epoch release matrix (5 sizes x 3 datasets, seed 42) trained for the WSKAN architecture study. It is released to make the study's qualitative claims directly inspectable: the generation results below are real, unedited outputs of these exact weights.

Features

  • Kolmogorov-Arnold Network: learnable univariate edge functions replace scalar weights.
  • Wavelet-like SSM edges: each edge function is the impulse response of a stable state-space model — a damped oscillator (decay sigma, frequency omega) that acts as a wavelet-like kernel.
  • Fully selective recurrence: input-dependent step size dt and input/output projections (B, C), so the wavelet transform is content-warped per token (Mamba-style selectivity).
  • Fused Triton scan: the associative scan is a single fused GPU kernel (log-depth), verified against the reference scan (fwd err 7e-7, grads <= 4e-6).
  • Byte-level: vocabulary is the 256 byte values — no tokenizer, no BPE, no special tokens.
  • Interpretable by construction: every edge exposes its learned (sigma, omega) spectrum; see the interpretation reports in the GitHub repo.

Model collection

Model Params Dataset Final eval loss Byte PPL
wskan-1k-tinystories 3,610 TinyStories 2.1416 8.51
wskan-10k-tinystories 12,546 TinyStories 1.4401 4.22
wskan-100k-tinystories 113,612 TinyStories 0.8224 2.28
wskan-1m-tinystories 985,956 TinyStories 0.6135 1.85
wskan-10m-tinystories 10,153,996 TinyStories 0.5453 1.73
wskan-1k-ultrachat 3,610 UltraChat 200k 2.4298 11.36
wskan-10k-ultrachat 12,546 UltraChat 200k 1.8724 6.5
wskan-100k-ultrachat 113,612 UltraChat 200k 1.2316 3.43
wskan-1m-ultrachat 985,956 UltraChat 200k 0.9854 2.68
wskan-10m-ultrachat 10,153,996 UltraChat 200k 0.8472 2.33
wskan-1k-wikitext 3,610 WikiText-103 2.3627 10.62
wskan-10k-wikitext 12,546 WikiText-103 1.9053 6.72
wskan-100k-wikitext 113,612 WikiText-103 1.3508 3.86
wskan-1m-wikitext (this model) 985,956 WikiText-103 1.1541 3.17
wskan-10m-wikitext 10,153,996 WikiText-103 1.0266 2.79

This model is trained on encyclopedic prose (Salesforce/wikitext wikitext-103-raw-v1); the other rows cover the remaining two domains and sizes. All rows use the identical WSKAN-11 architecture, differing only in depth/width.

Generation results (real outputs)

The following are actual outputs of this checkpoint, generated with experiments/RELEASE_infer.py from the released weights (multinomial sampling, temperature 0.8, seed 1234, 400 new bytes). Nothing is cherry-picked or edited — these are the first and only samples drawn for this card.

Prompt:

The history of the city began

Output (seed 1234, temperature 0.8, 400 new bytes):

The history of the city began accusing the landing eleven 197 yards , six players peaked only four times , carrying 28 @,@ 000 points against the Kansas Publications Hotel . All of the feudscannial changes were worth the discovery of the need for sound . This came to equate the film 's character , pointing with a couple of the birth to produce all the most commonly increased share of the substance for a special entrance to th

Prompt:

In the early 20th century, scientists discovered

Output (seed 1234, temperature 0.8, 400 new bytes):

In the early 20th century, scientists discovered accusations of the laughs in humans . Also , the first herbal pressure did not want many churches or absence of a formal clouditian in which the integer is expected . 




 = = = Counsel = = = 




 Illmatic national sources prove advanced more efforts to depose the idea of Green was confident to suit . At 19 : 19 years of the commercial residents are reduced . The variety of a court to shop and 

Prompt:

The film was directed by

Output (seed 1234, temperature 0.8, 400 new bytes):

The film was directed by Ben Daniel , leading element of religious captain . The breakdown of the Jagdca Grosman were presented to the project as they are required to become popular and authoritating the reform . 


 As with lemuringness , Alexander publicly went to the Panic 's business and appointed Beatles — the standards of corporation ( or implemented by his involvement ) . Assumina varied and a court to scream to

Evaluation of the outputs

Article-like structure with parenthetical asides, em-dashes, and consistently formatted invented entities ("the Jagdca Grosman", "the Panic 's business"). Sentences are grammatical and the first sentence after the prompt stays on-register ("Ben Daniel , leading element of religious captain"). Topical consistency degrades over a paragraph, and every fact is fabricated. The model is a reliable WikiText stylist, an unreliable encyclopedia.

Objective metrics for this checkpoint: final eval loss 1.1541 (byte-level perplexity 3.17) on a held-out slice of WikiText-103, at training step 66,000.

Configuration

The released config.yaml (verbatim):

model_name: wskan-1m-wikitext
architecture_family: WSKAN-11 (Wavelet-like State-Space KAN, fused Triton scan)
architecture:
  vocab_size: 256
  d_model: 80
  n_layers: 6
  n_states: 6
  use_feature_bc: true
  bc_rank: 32
  wz_diag: false
  g_rank: null
  oscillatory: true
  bf16_scan: true
tokenizer:
  type: byte-level
  description: raw UTF-8 bytes; vocab ids 0-255; no BPE, no special tokens
training:
  dataset: Salesforce/wikitext wikitext-103-raw-v1
  steps: 66000
  checkpoint_step: 66000
  batch_size: 64
  block_size: 256
  lr: 0.001
  lr_schedule: cosine
  optimizer: AdamW (weight_decay=0.0, grad clip 1.0)
  seed: 42
  epochs: ~3
  precision: fp32 weights, bf16 scan
evaluation:
  final_eval_loss: 1.1541
  final_eval_byte_ppl: 3.17
generation_defaults:
  max_new: 400
  temperature: 0.8
  seed: 1234

Architecture quick reference: d_model=80, n_layers=6, n_states=6, bc_rank=32, vocab=256 (raw bytes). Training: 66,000 steps, batch 64, block 256, lr 0.001 (cosine), AdamW, seed 42, ~3 epochs, bf16 scan.

Local run

Inference requires an NVIDIA GPU with Triton (the scan kernel is fused Triton; there is no CPU path) and Python 3.12+.

git clone https://github.com/llaa33219/wskan.git
cd wskan
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements.txt
uv pip install --python .venv/bin/python safetensors pyyaml huggingface_hub

# generate from this model (downloads config.yaml + model.safetensors from the Hub)
.venv/bin/python experiments/RELEASE_infer.py \
    --model llaa33219/wskan-1m-wikitext \
    --prompt "The history of the city began" \
    --max-new 400 --temperature 0.8 --seed 1234

--model also accepts a local directory containing config.yaml and model.safetensors. Weights are fp32; set bf16_scan: false in config.yaml for an fp32 scan (outputs will differ slightly from the samples above, which used the training-time bf16 scan).

To retrain from scratch under the identical protocol:

uv pip install --python .venv/bin/python datasets transformers
.venv/bin/python experiments/V1_train_tinystories_lm.py \
    --model wskan11 --scale 1m --dataset wikitext --seed 42 \
    --batch 64 --block 256 \
    --steps 66000 --lr 0.001 --lr-schedule cosine \
    --ckpt-every 25000 --eval-every 1000 --out-tag 3ep --compile --bf16

Limitations

  • Research artifact. These are 3.6k-10M parameter proof-of-concept models from an architecture study, not production LMs.
  • No factual reliability. All models confabulate names, dates, and facts; the larger ones merely do so more fluently. Never use outputs as a source of truth.
  • Not instruction-tuned. Even the UltraChat models only imitate conversational format; they do not follow instructions.
  • Short-range memory. The interpretation reports show the models' effective memory is local (dozens of bytes); long-range consistency is not to be expected.
  • English-only, single-domain. Each model knows only its training slice.
  • CUDA GPU required for inference. The recurrent scan is a fused Triton kernel; there is no CPU path. (config.yaml: architecture.bf16_scan: false switches the scan to fp32 on GPU.)
  • Sampling only. Generation is multinomial sampling; outputs vary with seed and temperature.

Provenance

  • Checkpoint: checkpoints/wskan11_wikitext_1m_3ep_s42/latest.pt (step 66,000) of the WSKAN repository, converted to safetensors (fp32) without any weight modification.
  • Training data: WikiText-103, byte-encoded UTF-8.
  • Seed 42 is the canonical seed of the 5-seed campaign; the released sample outputs above used sampling seed 1234.
Downloads last month
22
Safetensors
Model size
1.01M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including llaa33219/wskan-1m-wikitext