An open-source Turkish language model built and trained from scratch: its own architecture, its own tokenizer, its own weights.

Role

Creator and maintainer

Context

Open-source research project

Period

2026 — ongoing

Platform

Python · PyTorch · CUDA, Apple Silicon (MPS), CPU

Links

Most Turkish language models are English-centred models fine-tuned on Turkish text. They inherit a tokenizer and an architecture that were never designed for an agglutinative language with vowel harmony and rich morphology. Toprak asks a different question: what does a model look like when Turkish is the starting point?

Nothing is fine-tuned from an existing model. The project contains the full pipeline: data collection and cleaning with provenance and licence metadata, a 32,000-token Turkish BPE tokenizer, a modern decoder-only transformer, the training loop, inference and an evaluation suite. Training objectives specific to Turkish are implemented as optional auxiliary losses, so their effect can be measured in ablations.

  1. 01

    Modern decoder-only architecture

    RMSNorm, SwiGLU, rotary position embeddings and grouped-query attention with a KV cache.

  2. 02

    Four model sizes

    Configurations from roughly 80 million to 941 million parameters.

  3. 03

    Turkish tokenizer

    A 32K SentencePiece BPE vocabulary with coverage, fertility and morphology analysis.

  4. 04

    Turkish-specific objectives

    Optional auxiliary losses for vowel harmony, consonant assimilation, morphology, and syllable metre and rhyme.

  5. 05

    Data governance & evaluation

    Quality, PII and deduplication filters, provenance records, and a deterministic multi-dimensional evaluation suite.

Toprak is published under the Apache 2.0 licence and is in active development. Code, architecture and training process are fully open, with guides for reproducing experiments, comparing tokenizers and running ablations.

AI-powered vehicle marketplace

AI · Automotive