
Discover ai engineering fundamentals by balancing science and engineering, from temperature and tokenization to transformer concepts and cost-aware pipelines, culminating in chat-with-docs and rag applications.
Manage api keys with environment variables to keep secrets out of code and across environments. Avoid committing .env files, use secret stores or managers, and rotate keys with least privilege.
Learn secrets hygiene to keep passwords, API keys, and tokens out of code and logs, rotate them, use secret managers, redact authorization headers, and rely on Gitleaks to block leaks.
Configure a secure local environment by creating a dot env file for API keys, adding it to gitignore, and loading keys from the environment with Cursor, git, and uv.
Apply determinism controls to make LLM outputs repeatable for the same prompt by stabilizing inputs and enforcing request hygiene, then tune temperature and use low settings for debugging and testing.
Capture structured, timestamped logs to debug and monitor LLM apps without exposing secrets. Track request IDs, model, latency, and errors to reproduce issues, monitor costs, and maintain audits.
Compare open versus closed models, detailing weights access, self-hosting, fine-tuning, and api-based inference. Weigh cost, privacy, customization, and ops burden to guide the right choice for various workloads.
discover how to run and manage large language models locally with Ollama, a runtime and manager exposing a local API at localhost:11434 for private on-device prototyping.
Install ollama, pull a model, and run a local OpenAI-compatible API. Switch between local and cloud with the same code, exploring privacy, offline use, and cost savings.
Build a real LLM app that runs on OpenAI in the cloud and Ollama locally, with secure config, deterministic outputs, safe logging, and multi-provider support.
Discover why AI engineering is the fastest-growing tech career, with high salaries and a clear path from Python to production AI systems, through rag, agents, and production apis.
Discover the transformer architecture, the attention-based, parallelizable foundation powering modern L-L-Ms, including encoder-decoder and decoder-only designs. Learn how multi-head attention, MLPs, embeddings, and positional info enable long-range context.
Master self-attention, the transformer mechanism that lets each token attend to all others, compute compatibility scores, and capture long-range dependencies for coherent language processing.
Explain the q-k-v mechanism of attention, showing how queries, keys, and values compute similarity, weight values into a weighted sum, and produce a soft lookup in self- and cross-attention.
Learn how attention masking controls what a transformer can attend to, using causal masks for generation and padding masks for batched sequences, applied before softmax.
Enforce left-to-right autoregressive generation with causal masking, preventing future-token access by masking the upper triangle of the attention matrix while allowing the current token on the diagonal.
Understand scaled dot-product attention in transformers by scaling dot products with the square root of the key dimension to stabilize training and prevent softmax saturation.
Visualize self-attention by computing raw QK^T scores, applying softmax to form attention weights, and interpreting heatmaps of full and causal masked attention across tokens.
Explore multi-head attention, where multiple attention mechanisms run in parallel, each learning distinct relationships like syntax, coreference, and formatting patterns, to enrich transformer representations for n-l-p tasks.
Learn how positional encoding injects sequence information into transformers, using sine-cosine or learned embeddings, and when to rely on simple encodings for text and sequential data.
Understand why transformers are order blind and how sinusoidal positional encoding gives each position a unique fingerprint that is added to token embeddings to guide attention.
Explore how feed-forward networks complement attention in transformers by applying a two-layer mlp to each token, expanding representations and enabling non-linear feature extraction.
Layer normalization stabilizes transformer training by normalizing each token across features to zero mean and unit variance, with learnable gamma and beta for expressiveness.
Understand residual connections in transformers, where x plus F(x) creates a skip path around attention and feed-forward sublayers, preserving information and enabling gradient flow for deep stacking.
Explore encoder-decoder models, a classic sequence-to-sequence architecture that encodes input bidirectionally and decodes with cross-attention for translation, summarization, and rewriting.
Explore decoder-only models that read left to right, predict the next token, and use causal masking. Learn how prompt-in, continuation-out interfaces power chat, code completion, and summarization.
Explore llm inference as runtime that turns prompts into outputs by tokenizing input, performing a forward pass to obtain next-token probabilities, and applying decoding strategies for latency, safety, and creativity.
Understand parameter scaling and its trade-offs, including increased parameters, layers, attention heads, and MLP size, with scaling laws guiding when to scale, optimize compute, and use R-A-G.
Build a mini transformer from scratch, implementing self-attention, multi-head attention, positional encodings, a feed-forward network, layer norm, and residuals, then train for autoregressive next-token character generation with causal masking.
This module unpacks transformer machinery: attention, QKV projections, and positional encoding for GPT-style decoding. See how multi-head attention, residual connections, and layer normalization enable debugging and scalable LLMs.
Explore the differences between data scientists, ML engineers, and AI engineers, and learn how foundation models shifted from training to composing, enabling rapid production of AI products in months.
Master tokenization, token IDs, and context windows; and experiment with temperature, top-k, top-p, stop sequences, and streaming in an interactive LLM playground to balance cost and quality.
Explore tokenization in AI engineering fundamentals. Learn how it converts text into tokens that LLMs process, affecting prompts and API costs, and discover language-specific tokenizers, subword pieces, and BPE.
Explore Byte-Pair Encoding (BPE) as a frequency-based tokenization method used by GPT, merging frequent character pairs to build a compact, efficient vocabulary.
Explore tokenizer vocabularies, the model’s dictionary of token pieces, including token IDs and subwords, and learn how fixed vocabularies shape cost, context window, and you can't add words at runtime.
Learn how token IDs convert text into numeric labels for LLMs, how embeddings map them, and how token counts affect context windows and costs.
Explore how text is tokenized into tokens and IDs, and compare BPE, WordPiece, and SentencePiece. Stress test with emojis and multilingual text, and build a cost estimator to optimize prompts.
Master the context window by budgeting tokens for instructions, retrieved context, and the answer, while balancing input size, history, and tool outputs to ensure reliable LLM performance.
Master token limits, tokenization, and vocabularies to manage input and output tokens within a model's context window, using trimming, retrieval, and stepwise tasks to prevent truncation and errors.
Budget your context window by counting tokens, preventing lost in the middle with strategies, compare model contexts, and reserve a safety margin, while using chunking with overlap or extract-then-synthesize.
Explore autoregressive generation, the token-by-token loop powering LLM text. Learn how predictions, stop conditions, and drift impact streaming responses, and how constraints like temperature, outlines, and validation improve generation.
Temperature governs randomness in next-token probabilities, making outputs deterministic at low values and varied at high values, with low temperature ideal for consistent summaries and extractions.
Explore top-k sampling: keep the top K tokens, renormalize, and sample to achieve coherent yet varied outputs for tasks like customer support drafting and marketing content.
Explore top-p (nucleus) sampling, an adaptive method that keeps the smallest set of tokens whose cumulative probability reaches p, balancing creativity and coherence.
Beam search maintains B candidate sequences, expands them at each step by cumulative log probability, and keeps the top beams to maximize likelihood and avoid premature greedy choices.
Explore log probabilities (logprobs) to reveal how confidently a model forecasts each next word, using log space to add token likelihoods instead of multiplying tiny probabilities.
Explore token probabilities through logprobs and top logprobs to measure model confidence, detect uncertainty, and build simple classifiers. Use perplexity and hallucination signals for production-ready filtering.
Select the right LLMs to balance latency, cost, reliability, and risk, routing tasks to cheap or strong models as needed. Measure prompts, tokens, and performance with task-specific testing and monitoring.
Compare multiple models with identical prompts to evaluate quality, speed, and cost, then apply a repeatable framework for model selection using OpenAI, Anthropic, and Ollama.
Explore tokenization, temperature, top-p and top-k controls in a hands-on llm playground, streaming outputs, and cross-model comparisons to optimize prompts, costs, and production decisions.
Turn prompts into an engineered interface with system and user messages, templates, and JSON mode. Stress-test against adversarial inputs and apply few-shot and chain-of-thought reasoning, plus injection defense.
Master prompt engineering by clearly specifying goals, context, constraints, and output formats for accuracy. Treat prompts like API designs, using explicit intent and examples to guide LLMs toward structured outputs.
Design prompts as an interface contract with LLMs, detailing goals, context, constraints, inputs, and output formats like JSON or tables to ensure reliable results.
Craft precise user prompts that specify task, inputs, and output formats, while respecting the system prompt and history; avoid vagueness, label untrusted data, and include goals, constraints, and success criteria.
Master role prompting to steer LLMs by setting a role and decision rules within the prompt context, shaping behavior across production tasks from polite support to secure code review.
Learn how system prompts shape model behavior in prompt engineering. Define identity, behavior rules, constraints, and output format, and explore role prompting for targeted perspectives.
Explore context injection in ai apps, how untrusted retrieved content can influence models, and strategies to label untrusted data, delimit content, and mitigate risks in rag and tool-augmented agents.
Explore zero-shot prompting, giving clear instructions without examples to power fast baselines in classification, summarization, rewriting, with JSON output schemas for precise results.
Learn chain-of-thought prompting for multi-step tasks, and its security trade-offs in production. Use concise final outputs with limited justification in real-world scenarios like support triage and invoice matching.
Explore chain-of-thought prompting and self-consistency to reveal verifiable reasoning in multi-step problems, compare direct prompting with chain-of-thought prompting, and apply majority voting for higher accuracy in ai engineering fundamentals.
Learn self-consistency, a prompting technique that runs the same prompt multiple times with randomness to derive a consensus answer for reasoning-heavy tasks.
Explore how ReAct prompting enables an LLM to reason, act, and observe with external data, while using guardrails for untrusted tool outputs to deliver grounded, multi-step answers.
Unlock the ReAct pattern—reason, act, and observe—building a multi-step agent that uses tools like a calculator and lookup. Learn prompt engineering, tool orchestration, and adaptive reasoning for complex questions.
Define a strict JSON schema to guide LLM responses, enabling reliable parsing, validation, and production-ready automation in downstream systems like support queues and fintech records.
Develop reliable, structured data from model output using output parsing, including JSON fields, typing, and validation to prevent errors and prompt injection.
Turn noisy llm outputs into reliable data by using json mode, schema prompts, and pydantic validation. Build a reusable production pattern for structured data extraction from llms.
Leverage prompt chaining to break complex tasks into sequential prompts, where each step handles a piece of the task and passes results forward to improve outputs.
Explore prompt chaining: decompose tasks into a three-step pipeline—extract key points, classify sentiment with a score, and deliver a summary and recommendations.
Iteratively optimize prompts with explicit constraints and testing to produce consistent, structured outputs, minimize drift, and apply in customer support and legal ops.
Apply version control to prompts, treating them as production code, tracking changes, versions, and the full bundle to reproduce outputs and roll back safely.
This document data extractor lab builds a robust pipeline that outputs validated JSON from messy unstructured text, across invoices, using a defined schema, JSON mode, and Pydantic validation.
Learn how prompt injection undermines system prompts and test five attack types—direct override, persona hijack, goal hijack, indirect extraction, and completion tricks, and explore defense in depth.
Defend ai prompts from adversarial inputs by validating inputs, setting context boundaries, and monitoring interactions to ensure secure, reliable outputs. Apply prompt defense in security-sensitive and public-facing ai systems.
Build a defense-in-depth chatbot template that resists prompt injection and jailbreak attempts through input sanitization, hardened system prompts, and output filtering, with logging and validation against adversarial prompts.
Apply reliable prompt engineering with roles, system prompts, templates, and schema-first outputs to ensure secure, production-ready prompts. Leverage embeddings, document processing, and guarded retrieval to manage untrusted input and injection.
Master the RAG pipeline from embeddings and document preparation to chunking and indexing with HNSW, enabling fast retrieval via cosine similarity in a vector database.
Embedding models turn text into vectors, so meaning-based similarity guides semantic search and recommendations. Avoid exact matches; chunk long documents for better recall.
Understand embedding dimensions as the length of the embedding vector and how dimension choice trades detail for storage and search cost, ensuring matched dimensions across your vector index.
Understand semantic similarity as a meaning-based match using embedding vectors and cosine similarity to find paraphrased queries, synonyms, and near-duplicate content in search, retrieval, and recommendations.
Turn language into math with a vector space where text becomes vectors, enabling semantic search and nearest-neighbor ranking. Use a single embedding model for consistency to avoid cross-space comparisons.
Turn text into numerical embeddings with OpenAI's embedding model, compare their directions in vector space with cosine similarity, and enable semantic search and retrieval in RAG systems.
Explore how vector databases store embeddings and enable fast nearest-neighbor search for semantic search and rag pipelines, complementing relational databases and guiding chunking quality for reliable retrieval.
Store embedding vectors with metadata and enable fast similarity searches for semantic search and RAG pipelines. Use HNSW and other approximate methods, and avoid mixing vectors across models.
Discover how vector indexing turns embedding vectors into fast, scalable search for semantic queries across millions of documents, using consistent models and index types like HNSW and IVF.
Understand how nearest neighbor search uses embeddings and distances in embedding space, via cosine similarity or dot product, to retrieve top k semantically similar items for semantic search and rag.
Explore hnsw, a hierarchical navigable small-world graph for fast, approximate nearest neighbor search in embeddings. Tune ef search, M, and ef construction to balance speed and recall in vector databases.
Learn how approximate nearest neighbors enables fast, scalable vector search over millions of embeddings by trading accuracy for speed, with HNSW and IVF and recall-latency tuning.
Master document loaders that convert messy content such as PDFs, web pages, and emails into clean, metadata rich text chunks for embeddings and search, boosting retrieval quality in RAG pipelines.
Learn how chunk overlap preserves boundary context in RAG, improving retrieval for contracts, manuals, and policies by repeating surrounding text across chunks.
Explore how chunk size and overlap impact retrieval quality in rag pipelines, using a hands-on mini-lab to compare character-based chunking with overlapping chunks and observe effects on embeddings.
Split long documents by meaning rather than length, cutting at topic changes to keep each chunk coherent, improving embeddings and retrieval in RAG across policies and knowledge bases.
Extract structured metadata from messy documents to create clean json records with fields like date, author, and topics before embedding and indexing, enabling filtering, routing, and secure retrieval alongside embeddings.
Attach structured metadata to each chunk and use metadata filtering to sharpen semantic search results. Compare filtered and unfiltered queries to improve retrieval relevance and recall in production rag systems.
Turn your built document pipeline into a hire-ready portfolio with four connected demos: module 5 semantic search, module 6 RAG system, module 7 tool-using agent, module 10 production API.
Learn the end-to-end rag pipeline, from retrieval to citations, with context injection and stuffing within the token budget.
Learn how retrieval augmented generation grounds LLM answers in your knowledge base by fetching relevant documents and citing sources, while prioritizing high-quality chunks.
Learn how context retrieval powers RAG systems by selecting relevant private data to ground LLM answers. Explore hybrid retrieval with vector search and BM25 to avoid noise and missed matches.
Learn how context injection shapes AI outputs in RAG, and how to label untrusted content, delimit it, and apply mitigations to prevent misuse.
Context stuffing packs retrieved chunks into prompts, causing noise, higher costs, and slower responses. In production, curate to two to four relevant passages and enforce a strict context budget.
Explore how source citation grounds rag answers by attaching evidence from documents, urls, page numbers, and text snippets to verify and audit every response.
Build an end-to-end RAG pipeline that retrieves relevant chunks, embeds and chunks documents, stores them in Chroma DB, injects context, and generates grounded answers with citations.
Develop a production-grade chat-with-documents pipeline: ingest, chunk, embed into a vector store, retrieve with RAG, apply lost-in-the-middle mitigation, and cite sources via a Gradio UI.
Build an end-to-end rag pipeline with retrieval, context formatting, and citation-based generation. Emphasize grounded results, proper source citations, and the agent pattern for multi-step tasks.
Define tool definitions as contracts that tell the model how to request capabilities, with strict schemas, validated inputs and outputs, enums, and server-side authorization to ensure reliable, safe agent actions.
Learn function calling to bridge language models and real actions by defining tools with JSON schema and executing Python functions in a round trip, while the LLM never executes code.
Tool execution lets an LLM act by running external actions like API calls or queries, while the application validates arguments with schemas and requires human confirmation for destructive actions.
Build a framework-free python calculator tool that demonstrates parameter extraction, tool execution, and response parsing by dispatching to a registered function and returning a string result.
Execute multiple independent tool calls in parallel to reduce latency, then synthesize results into an answer. Learn when to run calls concurrently, set concurrency limits and backoff, and avoid over-parallelizing.
Convert a model’s free-form output into structured data like JSON or typed objects, extracting fields, tool calls, and actions with strict validation and retries.
Design end-to-end agent architectures that orchestrate input perception, reasoning, tool calls, memory and state management, and safety checks to reach a goal with clear stopping criteria.
Build an agent loop where the LLM acts as brain, tools provide capabilities, and memory guides observe, act, and repeat for multi-step tasks within a max iterations cap.
Plan and execute with an agent pattern that first creates an explicit step-by-step plan, then executes each step with tools, separating planning from doing for traceability.
Learn the ReAct agent loop—think, act, observe—and compare its transparent reasoning with plan-and-execute, mastering debugging and tool-driven multi-step workflows.
Understand how agent memory stores and reuses information across steps and sessions. Learn how it distinguishes short-term, working, and long-term memory to support multi-step tasks and personalization.
Long term memory provides selective, persistent storage for intelligent agents to remember user preferences and project state across sessions, enabling continuity and personalization.
Engineer stateless LLMs by adding short-term memory for session context and long-term memory for cross-session facts, using memory tools and a JSON file to persist knowledge.
Master task decomposition by breaking big goals into small, verifiable steps aligned to tools and actions with verification checkpoints to catch errors early.
Learn how stop conditions bound cost, time, and loops by defining exit rules for AI agents, including goal checks, progress cues, and tool-failure handling, with practical examples.
Coordinate multi-agent workflows with structured messages, clear schemas, and explicit state to avoid contradictions. Use typed intents, task IDs, and acceptance criteria to ensure reliable handoffs and aligned goals.
Explore strict tool schemas, observe-act-verify cycles, and parallel tool calls to build reliable agents with memory, guardrails, and structured extraction.
Define the exact output structure upfront with schema first prompting to turn open tasks into a bounded form for reliable JSON extractions. Remember, structure isn't truth and requires validation.
Learn validation error recovery to patch structured outputs that fail schema checks: detect the exact wrong fields, re-prompt only those parts, merge fixes, and validate with capped retries.
Process large collections of documents with a structured extraction pipeline, applying the same schema, prompts, and validation rules. Include retries, idempotency, recovery, and per-item logging for auditability and reliable throughput.
Build a production-ready document extractor that classifies documents, extracts data into typed pydantic schemas, scores confidence with logprobs, and processes batches asynchronously with validation and recovery.
Define and fill a tiny JSON schema to extract structured data from unstructured documents like invoices, then validate and compare zero-shot and few-shot extraction using system prompts and JSON outputs.
Define a Pydantic schema as a contract and enforce JSON mode to extract and validate invoice data with type-safe, nested structures, handling optional fields and validation errors with retries.
Classify documents to identify type, then select the matching schema and extract structured data across invoices, receipts, resumes, and contracts. Build a universal extractor using schema-first prompts.
Batch process folders of documents using async concurrency with a bounded limit and progress tracking. Track per-document status, handle errors gracefully, and generate a report with JSON exports and analytics.
Develop a repeatable evaluation mindset that measures every prompt tweak, model swap, and retrieval adjustment with reference-based metrics like Bleu and Rouge, plus LLM-as-judge and rubric scoring.
Evaluate LLM development by measuring evaluation metrics like quality, safety, latency, and cost via offline tests, online tests, and human review. Use AB experiments and multi-metric dashboards to prevent regressions.
Compare large language model outputs to a trusted reference or gold answer using deterministic, heuristic, or judge scoring to yield repeatable correctness checks across tasks.
Explore reference-based evaluation using BLEU and ROUGE to measure precision and recall in LLM outputs, compare against a gold standard, and choose metrics by task.
Use an llm-based judge to score another model's outputs against a rubric with structured json, enabling fast pairwise comparisons for ci, rag, and customer support.
Design an auto-grader that applies a weighted rubric to evaluate LLM outputs consistently. Build domain-specific criteria, define score levels, and parse results to produce reproducible, interpretable scores.
Agent evaluation measures an AI agent's performance across multi-step tasks, tool calls, safety, and cost, using full trajectory analysis rather than final answers alone.
Develop an end-to-end evaluation pipeline that blends BLEU, ROUGE, LLM-as-judge, and rubric scoring into a single pass or fail decision for Q&A outputs.
Build a two-layer agent evaluation pipeline that analyzes traces with deterministic checks and an llm rubric judge to diagnose tool selection, parameter accuracy, and answer quality, plus failure patterns.
Wrap up module 9 by outlining practical evaluation techniques: measure regressions, use BLEU and ROUGE for references, deploy judge prompts with rubrics, and automate evals to gate production.
Bridge a demo to a production API using fast api, async concurrency, streaming, caching, retries with backoff and jitter, and observability for cost tracking and diagnosis.
Expose typed, validated ML and LLM endpoints with FastAPI, a modern Python web framework that auto-generates docs, supports sync and async handlers, and centralizes microservices.
Turn a notebook llm into a robust api endpoint with request handling. Manage parsing, auth, rate limits, timeouts, and model calls to deliver a consistent response envelope for production reliability.
Define a stable, structured response format for AI APIs to ensure consistent, versioned, machine readable outputs with clear data or error envelopes, citations, and warnings.
Learn to run multiple LLM calls concurrently with async and asyncio gather using the OpenAI async client, compare sequential versus concurrent performance, and implement per-task error handling for production throughput.
Learn caching strategies for AI apps to reduce repeated LLM calls, embeddings, and retrieved documents with TTL, versioned keys, stale-while-revalidate patterns, spike protection, avoiding pitfalls.
Master retry patterns to build resilient AI apps, applying exponential backoff with jitter, cap attempts, and idempotency keys to handle transient 429, 503, and timeouts while avoiding permanent errors.
Explore observability for LLM apps using logs, metrics, and traces to reveal what happened, why, and how well, with practical guidance on request IDs and privacy.
Learn how structured logs with correlation IDs enable end-to-end tracing, debugging, and performance monitoring for AI apps. Prioritize privacy by redacting data and logging targeted events only.
Learn to track LLM costs with a repeatable per-request calculator using prompt and completion tokens, model pricing, and labeling to identify top cost drivers and observability insights.
Ship production-grade LLM endpoints with FastAPI, async streaming, caching, and retries, guided by robust API contracts and observability, ensuring predictable behavior under load and failure.
Stop watching AI tutorials. Start engineering AI systems, the scientific way.
Most AI courses teach you to copy notebooks without understanding what happens underneath. This course takes a different approach. You’ll learn AI engineering through first principles, controlled experiments, measurable results, and honest analysis of what works, what fails, and why.
Using our Bricks → Walls → Castles model, you’ll progress from foundational concepts to focused mini-labs and complete project builds. Every concept is explained before you apply it, every lab supports a larger goal, and every project produces something worth showcasing on GitHub, your résumé, or your portfolio.
WHAT MAKES THIS COURSE DIFFERENT
Scientist-led teaching. TechBricks brings together scientists and engineers with more than 45 years of combined academic and industry experience. We teach both intuition and mathematics, without black boxes or hand-waving.
Concept first, code second. You’ll understand why each technique works before implementing it. Small ideas become practical components, and those components become complete AI systems.
Framework-light learning. You’ll work directly with the OpenAI SDK, without LangChain, LlamaIndex, or unnecessary abstractions. By the end, you’ll understand what AI frameworks do internally and when you should or shouldn’t use them.
Real projects, not toy demos. You’ll build systems for language modeling, structured extraction, multi-step reasoning, prompt-injection defense, retrieval-augmented generation, tool use, automated evaluation, and production deployment.
Cloud and local models. Work with the OpenAI API and free local models through Ollama, making the course accessible with or without an API budget.
PROJECTS YOU’LL ADD TO YOUR PORTFOLIO
• Mini-Transformer Built from Scratch
• Interactive LLM Playground
• Intelligent Document Data Extractor
• Multi-Step Reasoning Engine
• Injection-Resistant Secure Chatbot
• Chat-with-Docs RAG Application
• Framework-Free Tool-Using AI Agent
• Automated Evaluation Pipeline
• Production-Grade FastAPI LLM Service
37 mini-labs + 13 project labs = 50 hands-on exercises.
Enroll today and start building AI applications you actually understand.
See you in Module 1.
The TechBricks Team