Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
[Arabic] Build an LLM from Scratch with Pytorch
Bestseller
Hot & New
Rating: 5.0 out of 5(14 ratings)
316 students

[Arabic] Build an LLM from Scratch with Pytorch

ابنِ نموذج GPT عربي-إنجليزي بنفسك خطوة بخطوة: Tokenizer و Attention و KV Cache والتدريب وتوليد النصوص
Last updated 9/2026
Arabic
Arabic [Auto],

What you'll learn

  • Build a GPT-style large language model (LLM) from scratch in PyTorch, one component at a time
  • Train a bilingual Arabic-English BPE tokenizer and prepare a real Wikipedia dataset for training
  • Understand and code self-attention, multi-head attention, and the KV cache
  • Pretrain your own LLM on a GPU with AdamW, learning rate scheduling, checkpoints, and experiment tracking
  • Assemble transformer blocks into a full GPT model, with weight tying, GPT-2 initialization, and the training loss
  • Generate text with greedy decoding, temperature, top-k, and top-p sampling
  • بناء نموذج لغوي كبير (LLM) من الصفر باستخدام PyTorch، وفهم كل جزء فيه من خلال الكود
  • تجهيز النصوص العربية للتدريب: التنظيف والتوحيد وبناء Tokenizer يدعم العربية والإنجليزية

Course content

7 sections • 16 lectures • 16h 24m total length
  • Introduction to Transformers59:32

    A detailed walkthrough of the Transformer architecture behind modern LLMs, and the roadmap for building GPT, from scratch.

    We cover tokens and embeddings, positional information, self-attention and multi-head attention, feed-forward layers, residual connections and layer normalization, and how a decoder-only model learns to predict the next token. The lecture slides are attached as a downloadable PDF.

  • Dataset Exploration1:00:53

    Meet the data behind Jabarti. We explore the jabarti-llm-dataset on Hugging Face. The pretrain subset has about  bilingual Wikipedia records, mixing general Wikipedia with a curated Egyptian-history collection.

    - Github: https://github.com/bakrianoo/jabarti-llm-from-scratch
    - Dataset: https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset/

    The finetune subset has Arabic and English question–answer pairs. We look at the columns, the Arabic/English mix, how the eval split holds out whole articles to avoid leakage, and the data-quality issues that shape the cleaning and tokenizer choices in the next section.

    dataset : https://huggingface.co/datasets/bakrianoo/jabarti-llm-dataset/
    ..

Requirements

  • Python programming: functions, classes, lists, and dictionaries
  • Basic math: matrices
  • No PyTorch experience needed. The course has a full PyTorch section before we build the model
  • A GPU (local or cloud) for the full training run. You can follow most lectures on a normal laptop
  • معرفة أساسية بلغة Python، ولا يشترط وجود خبرة سابقة في PyTorch أو الذكاء الاصطناعي

Description

في هذا الكورس سنبني نموذجًا لغويًا كبيرًا
(LLM)
من الصفر باستخدام
PyTorch
خطوة بخطوة.

النموذج اسمه جبرتي (Jabarti). هو نموذج GPT يفهم العربية والإنجليزية، وسندربه على بيانات حقيقية من ويكيبيديا، مع تركيز خاص على تاريخ مصر.

لن نستخدم نماذج جاهزة. سنكتب كل جزء بأنفسنا، ونفهم لماذا يعمل بهذه الطريقة.

الشرح باللغة العربية، والكود والمصطلحات التقنية بالإنجليزية، كما ستجدها في أي مشروع حقيقي.

ماذا ستبني في هذا الكورس؟
- Tokenizer يدعم العربية والإنجليزية، مع تنظيف وتوحيد النصوص العربية
- طبقات Embeddings وMulti-Head Attention وKV Cache
- نموذج GPT كامل من Transformer Blocks
- توليد النصوص بطرق Temperature وTop-k وTop-p
- تدريب النموذج (Pretraining) على GPU ومتابعة النتائج

Build a Large Language Model from scratch in PyTorch, taught in Arabic

In this course, you build Jabarti, a bilingual Arabic-English GPT model, and train it on real Wikipedia data. You write every part yourself, so you understand how an LLM works, not just how to call one.

What you will build:
- A BPE tokenizer for Arabic and English, with Arabic text normalization
- Token and position embeddings
- Multi-head self-attention with a causal mask and a KV cache
- Transformer blocks and a complete GPT model
- A data pipeline that streams a large corpus without loading it into memory
- Text generation with greedy decoding, temperature, top-k, and top-p
- A full pretraining loop with AdamW, a cosine learning rate schedule, checkpoints, and Trackio

The course starts with the Transformer architecture and a PyTorch refresher, so you don't need deep learning experience. You only need to know Python.

All the code is open source on GitHub, and the dataset is public on Hugging Face.

Who this course is for:

  • Python developers who want to understand how LLMs like GPT work from the inside
  • Machine learning students and engineers who want hands-on experience training a language model
  • Anyone working on Arabic NLP who wants to build models that handle Arabic text well
  • المبرمجون العرب الذين يريدون تعلم بناء نماذج الذكاء الاصطناعي اللغوية خطوة بخطوة، بشرح عربي واضح