wskan-1m-wikitext
985,956 parameters · WSKAN-11 · byte-level LM trained on WikiText-103 (local-coherence tier)
Introduction
wskan-1m-wikitext is a checkpoint of WSKAN-11 (Wavelet-like State-Space KAN, version 11), a
Kolmogorov-Arnold Network in which every edge function is the impulse response of a stable
state-space model — a wavelet-like damped oscillator. One parameterization supports two modes: a
closed-form wavelet edge (static, WavKAN-style) and a recurrent SSM edge (Mamba-style selective
scan); this checkpoint runs the recurrent mode through a fused Triton associative-scan kernel.
This model is one cell of the 3-epoch release matrix (5 sizes x 3 datasets, seed 42) trained for the WSKAN architecture study. It is released to make the study's qualitative claims directly inspectable: the generation results below are real, unedited outputs of these exact weights.
- GitHub repository (code, reports, training scripts): https://github.com/llaa33219/wskan
- Architecture document: models/V11_README.md
- Definitive interpretation report: experiments/W11_FINAL_INTERPRETATION_REPORT.md
Features
- Kolmogorov-Arnold Network: learnable univariate edge functions replace scalar weights.
- Wavelet-like SSM edges: each edge function is the impulse response of a stable state-space model — a damped oscillator (decay sigma, frequency omega) that acts as a wavelet-like kernel.
- Fully selective recurrence: input-dependent step size dt and input/output projections (B, C), so the wavelet transform is content-warped per token (Mamba-style selectivity).
- Fused Triton scan: the associative scan is a single fused GPU kernel (log-depth), verified against the reference scan (fwd err 7e-7, grads <= 4e-6).
- Byte-level: vocabulary is the 256 byte values — no tokenizer, no BPE, no special tokens.
- Interpretable by construction: every edge exposes its learned (sigma, omega) spectrum; see the interpretation reports in the GitHub repo.
Model collection
| Model | Params | Dataset | Final eval loss | Byte PPL |
|---|---|---|---|---|
| wskan-1k-tinystories | 3,610 | TinyStories | 2.1416 | 8.51 |
| wskan-10k-tinystories | 12,546 | TinyStories | 1.4401 | 4.22 |
| wskan-100k-tinystories | 113,612 | TinyStories | 0.8224 | 2.28 |
| wskan-1m-tinystories | 985,956 | TinyStories | 0.6135 | 1.85 |
| wskan-10m-tinystories | 10,153,996 | TinyStories | 0.5453 | 1.73 |
| wskan-1k-ultrachat | 3,610 | UltraChat 200k | 2.4298 | 11.36 |
| wskan-10k-ultrachat | 12,546 | UltraChat 200k | 1.8724 | 6.5 |
| wskan-100k-ultrachat | 113,612 | UltraChat 200k | 1.2316 | 3.43 |
| wskan-1m-ultrachat | 985,956 | UltraChat 200k | 0.9854 | 2.68 |
| wskan-10m-ultrachat | 10,153,996 | UltraChat 200k | 0.8472 | 2.33 |
| wskan-1k-wikitext | 3,610 | WikiText-103 | 2.3627 | 10.62 |
| wskan-10k-wikitext | 12,546 | WikiText-103 | 1.9053 | 6.72 |
| wskan-100k-wikitext | 113,612 | WikiText-103 | 1.3508 | 3.86 |
| wskan-1m-wikitext (this model) | 985,956 | WikiText-103 | 1.1541 | 3.17 |
| wskan-10m-wikitext | 10,153,996 | WikiText-103 | 1.0266 | 2.79 |
This model is trained on encyclopedic prose (Salesforce/wikitext wikitext-103-raw-v1); the other rows cover the remaining two domains and sizes. All rows use the identical WSKAN-11 architecture, differing only in depth/width.
Generation results (real outputs)
The following are actual outputs of this checkpoint, generated with
experiments/RELEASE_infer.py from the released weights (multinomial sampling, temperature 0.8,
seed 1234, 400 new bytes). Nothing is cherry-picked or edited — these are the first and only samples
drawn for this card.
Prompt:
The history of the city began
Output (seed 1234, temperature 0.8, 400 new bytes):
The history of the city began accusing the landing eleven 197 yards , six players peaked only four times , carrying 28 @,@ 000 points against the Kansas Publications Hotel . All of the feudscannial changes were worth the discovery of the need for sound . This came to equate the film 's character , pointing with a couple of the birth to produce all the most commonly increased share of the substance for a special entrance to th
Prompt:
In the early 20th century, scientists discovered
Output (seed 1234, temperature 0.8, 400 new bytes):
In the early 20th century, scientists discovered accusations of the laughs in humans . Also , the first herbal pressure did not want many churches or absence of a formal clouditian in which the integer is expected .
= = = Counsel = = =
Illmatic national sources prove advanced more efforts to depose the idea of Green was confident to suit . At 19 : 19 years of the commercial residents are reduced . The variety of a court to shop and
Prompt:
The film was directed by
Output (seed 1234, temperature 0.8, 400 new bytes):
The film was directed by Ben Daniel , leading element of religious captain . The breakdown of the Jagdca Grosman were presented to the project as they are required to become popular and authoritating the reform .
As with lemuringness , Alexander publicly went to the Panic 's business and appointed Beatles — the standards of corporation ( or implemented by his involvement ) . Assumina varied and a court to scream to
Evaluation of the outputs
Article-like structure with parenthetical asides, em-dashes, and consistently formatted invented entities ("the Jagdca Grosman", "the Panic 's business"). Sentences are grammatical and the first sentence after the prompt stays on-register ("Ben Daniel , leading element of religious captain"). Topical consistency degrades over a paragraph, and every fact is fabricated. The model is a reliable WikiText stylist, an unreliable encyclopedia.
Objective metrics for this checkpoint: final eval loss 1.1541 (byte-level perplexity 3.17) on a held-out slice of WikiText-103, at training step 66,000.
Configuration
The released config.yaml (verbatim):
model_name: wskan-1m-wikitext
architecture_family: WSKAN-11 (Wavelet-like State-Space KAN, fused Triton scan)
architecture:
vocab_size: 256
d_model: 80
n_layers: 6
n_states: 6
use_feature_bc: true
bc_rank: 32
wz_diag: false
g_rank: null
oscillatory: true
bf16_scan: true
tokenizer:
type: byte-level
description: raw UTF-8 bytes; vocab ids 0-255; no BPE, no special tokens
training:
dataset: Salesforce/wikitext wikitext-103-raw-v1
steps: 66000
checkpoint_step: 66000
batch_size: 64
block_size: 256
lr: 0.001
lr_schedule: cosine
optimizer: AdamW (weight_decay=0.0, grad clip 1.0)
seed: 42
epochs: ~3
precision: fp32 weights, bf16 scan
evaluation:
final_eval_loss: 1.1541
final_eval_byte_ppl: 3.17
generation_defaults:
max_new: 400
temperature: 0.8
seed: 1234
Architecture quick reference: d_model=80, n_layers=6, n_states=6, bc_rank=32, vocab=256 (raw bytes). Training: 66,000 steps, batch 64, block 256, lr 0.001 (cosine), AdamW, seed 42, ~3 epochs, bf16 scan.
Local run
Inference requires an NVIDIA GPU with Triton (the scan kernel is fused Triton; there is no CPU path) and Python 3.12+.
git clone https://github.com/llaa33219/wskan.git
cd wskan
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -r requirements.txt
uv pip install --python .venv/bin/python safetensors pyyaml huggingface_hub
# generate from this model (downloads config.yaml + model.safetensors from the Hub)
.venv/bin/python experiments/RELEASE_infer.py \
--model llaa33219/wskan-1m-wikitext \
--prompt "The history of the city began" \
--max-new 400 --temperature 0.8 --seed 1234
--model also accepts a local directory containing config.yaml and model.safetensors.
Weights are fp32; set bf16_scan: false in config.yaml for an fp32 scan (outputs will differ
slightly from the samples above, which used the training-time bf16 scan).
To retrain from scratch under the identical protocol:
uv pip install --python .venv/bin/python datasets transformers
.venv/bin/python experiments/V1_train_tinystories_lm.py \
--model wskan11 --scale 1m --dataset wikitext --seed 42 \
--batch 64 --block 256 \
--steps 66000 --lr 0.001 --lr-schedule cosine \
--ckpt-every 25000 --eval-every 1000 --out-tag 3ep --compile --bf16
Limitations
- Research artifact. These are 3.6k-10M parameter proof-of-concept models from an architecture study, not production LMs.
- No factual reliability. All models confabulate names, dates, and facts; the larger ones merely do so more fluently. Never use outputs as a source of truth.
- Not instruction-tuned. Even the UltraChat models only imitate conversational format; they do not follow instructions.
- Short-range memory. The interpretation reports show the models' effective memory is local (dozens of bytes); long-range consistency is not to be expected.
- English-only, single-domain. Each model knows only its training slice.
- CUDA GPU required for inference. The recurrent scan is a fused Triton kernel; there is no CPU
path. (
config.yaml: architecture.bf16_scan: falseswitches the scan to fp32 on GPU.) - Sampling only. Generation is multinomial sampling; outputs vary with seed and temperature.
Provenance
- Checkpoint:
checkpoints/wskan11_wikitext_1m_3ep_s42/latest.pt(step 66,000) of the WSKAN repository, converted to safetensors (fp32) without any weight modification. - Training data: WikiText-103, byte-encoded UTF-8.
- Seed 42 is the canonical seed of the 5-seed campaign; the released sample outputs above used sampling seed 1234.
- Downloads last month
- 22