Open source · AI research
An open-source Turkish language model built and trained from scratch: its own architecture, its own tokenizer, its own weights.
Role
Creator and maintainer
Context
Open-source research project
Period
2026 — ongoing
Platform
Python · PyTorch · CUDA, Apple Silicon (MPS), CPU
Links
The challenge
Most Turkish language models are English-centred models fine-tuned on Turkish text. They inherit a tokenizer and an architecture that were never designed for an agglutinative language with vowel harmony and rich morphology. Toprak asks a different question: what does a model look like when Turkish is the starting point?
The approach
Nothing is fine-tuned from an existing model. The project contains the full pipeline: data collection and cleaning with provenance and licence metadata, a 32,000-token Turkish BPE tokenizer, a modern decoder-only transformer, the training loop, inference and an evaluation suite. Training objectives specific to Turkish are implemented as optional auxiliary losses, so their effect can be measured in ablations.
What was built
- 01
Modern decoder-only architecture
RMSNorm, SwiGLU, rotary position embeddings and grouped-query attention with a KV cache.
- 02
Four model sizes
Configurations from roughly 80 million to 941 million parameters.
- 03
Turkish tokenizer
A 32K SentencePiece BPE vocabulary with coverage, fertility and morphology analysis.
- 04
Turkish-specific objectives
Optional auxiliary losses for vowel harmony, consonant assimilation, morphology, and syllable metre and rhyme.
- 05
Data governance & evaluation
Quality, PII and deduplication filters, provenance records, and a deterministic multi-dimensional evaluation suite.
The outcome
Toprak is published under the Apache 2.0 licence and is in active development. Code, architecture and training process are fully open, with guides for reproducing experiments, comparing tokenizers and running ablations.
Next project