
A detailed walkthrough of the Transformer architecture behind modern LLMs, and the roadmap for building GPT, from scratch.
We cover tokens and embeddings, positional information, self-attention and multi-head attention, feed-forward layers, residual connections and layer normalization, and how a decoder-only model learns to predict the next token. The lecture slides are attached as a downloadable PDF.
Meet the data behind Jabarti. We explore the jabarti-llm-dataset on Hugging Face. The pretrain subset has about bilingual Wikipedia records, mixing general Wikipedia with a curated Egyptian-history collection.
- Github: https://github.com/bakrianoo/jabarti-llm-from-scratch
- Dataset: https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset/
The finetune subset has Arabic and English question–answer pairs. We look at the columns, the Arabic/English mix, how the eval split holds out whole articles to avoid leakage, and the data-quality issues that shape the cleaning and tokenizer choices in the next section.
dataset : https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset/
..
Before we train a tokenizer, the text has to be clean.
In this lecture, we download the Jabarti dataset and clean the raw Wikipedia text. We remove leftover citation templates, reference sections, and very short sections.
Then we build a normalizer for Arabic text. It removes diacritics and tatweel, and it unifies the different forms of alef and ya. We also explain why we keep ta marbuta as it is.
The normalizer is saved inside the tokenizer. This way, the same rules run during training and during inference.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
Now we train the tokenizer itself.
We use Byte Pair Encoding (BPE) with a vocabulary of 32,000 tokens. It learns from a sample of Arabic and English documents.
We add special tokens for padding, start and end of text, language, and chat roles. Then we test the tokenizer on real Arabic and English sentences.
Finally, we tokenize the whole corpus once and save it as a binary file. Each document gets [BOS] and [EOS] markers, so the model knows where one document ends and the next one starts.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
A practical PyTorch lesson with only what we need to build the model.
We start with tensors. We create them, do math with them, and change their shape with view, reshape, and transpose. We also move tensors to the GPU.
Next, we use autograd to compute gradients. We build a tiny model with nn.Module and train it step by step: forward pass, loss, backward pass, and weight update.
At the end, we see how a language model turns token IDs into embeddings, and then into predictions for the next token.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
Our model can't read token IDs directly. It needs vectors.
In this lecture, we build the input embedding layer. Each token ID gets a learned vector from a token embedding table.
Then we add position embeddings, so the model knows the order of the tokens. We add the two together and apply dropout.
We also check the shapes at every step, from (B, T) token IDs to (B, T, d_model) vectors.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
This is the core of the transformer.
We build attention one step at a time on a small example. First, we compare tokens with dot products. Then we give each token a query, a key, and a value.
We scale the scores and add a causal mask, so no token can look at the future. Then softmax turns the scores into weights, and we use them to mix the values.
After that, we see why one head is not enough. We split the work across many heads and run them all together with one matrix multiplication.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
When the model writes text, it adds one token at a time. Without a cache, it computes the keys and values for the whole sequence again at every step.
In this lecture, we add a KV cache to our attention layer. We keep the keys and values from earlier steps, and only compute them for the new token.
We also handle the position offset and the causal mask. This way, the cached version gives the same result as before, only much faster.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
Now we put the pieces together into one transformer block.
We add the feed-forward network. It works on each token by itself: it makes the vector 4 times wider, applies GELU, and brings it back to its original size.
Then we add residual connections and layer normalization. We compare pre-norm and post-norm, and explain why we use pre-norm.
By the end, we have a block that we can stack as many times as we need.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
This is where the parts become a full language model.
We stack the transformer blocks, add a final layer norm, and add an output layer that gives a score to every token in the vocabulary.
We share one weight matrix between the input embeddings and the output layer. We also set up the weight initialization used in GPT-2.
Then we compute the loss. The targets are the inputs shifted by one token, and padding positions are ignored.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
We have a model. Now we need to feed it data.
In this lecture, we go back to the Jabarti dataset and build a PyTorch Dataset for it. It reads the token file we packed in the tokenizer section.
The file is opened with a memory map, so we never load the whole corpus into RAM. We cut the token stream into fixed windows, and each window becomes one training example.
We also look at why each window has one extra token, and how the [EOS] markers show where documents end inside a window.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
Training examples come one by one. The model needs them as batches.
In this lecture, we write the collate function that builds each batch. It pads every example to the longest one in the batch, not to the maximum length, so we don't waste compute on padding.
Then it shifts the tokens by one to make the inputs and the targets. Each position learns to predict the token that comes after it.
We also explain why padding at the end doesn't need its own attention mask. The causal mask already takes care of it.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
The model gives us a score for every token in the vocabulary. How do we pick the next one?
We start with greedy decoding, which always picks the most likely token. It is simple, but the text quickly becomes repetitive.
Then we add sampling. Temperature controls how random the choice is. Top-k keeps only the k best tokens. Top-p keeps the smallest group of tokens that covers most of the probability.
Finally, we write the generation loop. It uses the KV cache and stops at the [EOS] token.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
Before we start training, we build every piece the training loop needs.
We set up the AdamW optimizer, with weight decay only on the weight matrices. We add a learning rate schedule with a short warmup and a cosine decay.
Then we write the trainer. It supports gradient accumulation and gradient clipping, and it runs evaluation every few steps. It also generates sample text from fixed Arabic and English prompts, so we can see the model improve.
We finish with checkpoints, so a run can be saved and resumed, and with Trackio to log the metrics.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
Time to train Jabarti.
We run the pretraining script on a GPU. First we do a short test run on a small part of the data, to make sure everything works. Then we start the full run over the whole corpus.
We go through the main settings: batch size, gradient accumulation, learning rate, warmup, and how often to evaluate and save.
While it trains, we follow the loss and the sample outputs in the Trackio dashboard. We also open the dashboard from a remote machine using a Cloudflare tunnel.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
A pretrained model can continue text, but it can't answer questions yet. In this lecture, we teach it to.
We turn the Arabic and English question-answer pairs from the Jabarti dataset into chat examples, using the system, user, and assistant tokens. The loss is computed only on the answer, not on the question.
Then we build LoRA from scratch. We freeze the whole model and add two small matrices next to each attention projection. Only these matrices are trained, which is a tiny part of the model's parameters.
After training, we merge the LoRA weights back into the model. The result is a normal model again, with no extra cost at inference time.
Source code: https://github.com/bakrianoo/jabarti-llm-from-scratch/
في هذا الكورس سنبني نموذجًا لغويًا كبيرًا
(LLM)
من الصفر باستخدام
PyTorch
خطوة بخطوة.
النموذج اسمه جبرتي (Jabarti). هو نموذج GPT يفهم العربية والإنجليزية، وسندربه على بيانات حقيقية من ويكيبيديا، مع تركيز خاص على تاريخ مصر.
لن نستخدم نماذج جاهزة. سنكتب كل جزء بأنفسنا، ونفهم لماذا يعمل بهذه الطريقة.
الشرح باللغة العربية، والكود والمصطلحات التقنية بالإنجليزية، كما ستجدها في أي مشروع حقيقي.
ماذا ستبني في هذا الكورس؟
- Tokenizer يدعم العربية والإنجليزية، مع تنظيف وتوحيد النصوص العربية
- طبقات Embeddings وMulti-Head Attention وKV Cache
- نموذج GPT كامل من Transformer Blocks
- توليد النصوص بطرق Temperature وTop-k وTop-p
- تدريب النموذج (Pretraining) على GPU ومتابعة النتائج
Build a Large Language Model from scratch in PyTorch, taught in Arabic
In this course, you build Jabarti, a bilingual Arabic-English GPT model, and train it on real Wikipedia data. You write every part yourself, so you understand how an LLM works, not just how to call one.
What you will build:
- A BPE tokenizer for Arabic and English, with Arabic text normalization
- Token and position embeddings
- Multi-head self-attention with a causal mask and a KV cache
- Transformer blocks and a complete GPT model
- A data pipeline that streams a large corpus without loading it into memory
- Text generation with greedy decoding, temperature, top-k, and top-p
- A full pretraining loop with AdamW, a cosine learning rate schedule, checkpoints, and Trackio
The course starts with the Transformer architecture and a PyTorch refresher, so you don't need deep learning experience. You only need to know Python.
All the code is open source on GitHub, and the dataset is public on Hugging Face.