
In this lecture, we lay out what the course covers and, just as important, what it doesn't: how modern generative AI actually works under the hood, not how to prompt a chatbot. You'll get the map for everything ahead, including large language models (LLMs), tokens, RAG, agents, the Model Context Protocol (MCP), and AI security. By the end you'll know what you'll be able to explain and evaluate for yourself, and why the people who understand these foundations are the ones who get to adopt and secure AI at work.
In this lightboard lecture, we trace how AI got from hand-written rules to the models behind today's tools, and the one shift that made it happen: we stopped writing the rules and let machines learn the patterns from data. You'll see why rule-based systems kept breaking, how machine learning and neural networks changed the approach, and what the 2017 transformer and its attention mechanism made possible, the large language model (LLM). The idea to carry out is simple: nobody wrote the rules anymore, the machine learned them.
In this lightboard lecture, we map out the four layers of the modern AI stack, model, harness, agent, and tools, and how they turn a text-predicting model into something that gets real work done. By the end, you'll know what separates a model from an agent, and where each of these pieces fits.
Download the course slides and lightboard images here.
In this lightboard lecture, we break down what a large language model (LLM) actually is: one big file of trained numbers called parameters, or weights. You'll see how training differs from inference, what a size like "70B" means, and why an LLM answers one token at a time by predicting the next word instead of looking anything up.
Learn how AI tokens and tokenization work through a hands-on tokenizer demo. You’ll see how text is broken into tokens, why different models can produce different token counts, and how tokens affect AI model pricing, context windows, prompts, and overall LLM usage.
In this lightboard lecture, we break down the context window: the fixed working memory, measured in tokens, that holds everything a large language model (LLM) can see for a single request. You'll learn what fills the window, why the oldest parts of a long conversation quietly fall out, and why that explains both the forgetting and the rising cost of long chats.
In this lecture, we look at why the same prompt can give you a different answer every time, which comes down to how the model turns a list of token probabilities into one actual word. You'll see how sampling makes that pick, and how three settings, temperature, top-p, and top-k, reshape the odds before each choice. By the end you'll understand why turning temperature down makes a model more repeatable but not more correct, and when you'd want it low for consistent output versus higher for more variety.
In this lightboard lecture, we break down AI hallucinations: why a large language model (LLM) confidently invents facts and citations that don't exist. You'll see how the same next-token prediction loop produces both right answers and made-up ones, why the problem can be reduced but never removed, and which prompts to double-check.
In this lecture, we map out the AI model landscape, from the closed, hosted families like Claude, GPT, and Gemini to the open-weight models like Llama, Mistral, and DeepSeek that you can download and run yourself. You'll learn how to read any model name by breaking it into its family, size tier, and version, and why the biggest, most capable tier isn't always the right pick. By the end you'll be able to tell which model your everyday tools, like GitHub Copilot or Cursor, are actually built on, and why that matters.
In this demo, we work out what an AI workload actually costs by sending one real request to a model and reading the token usage the provider bills on. You'll see why input tokens and output tokens are priced separately, how a fraction of a cent per call turns into a monthly number once you multiply by request volume and calls per task, and why context size and output length move that number more than the rate card does. We close with prompt caching, what it discounts, and when it changes the math.
In this lecture, we compare calling a hosted API against running your own model, and the real tradeoffs between them. You'll see how the two differ on cost, data control, capability, and operational burden, where the cost crossover makes self-hosting cheaper at steady volume, and why owning your own model carries hidden costs like GPU capacity and MLOps staffing. By the end you'll know which workloads belong on a hosted API and which belong on your own infrastructure, and why most enterprises end up routing between both.
In this demo, you'll watch an open-weight large language model (LLM) run locally on a MacBook Pro with Ollama and answer a real support request with the Wi-Fi turned off, so nothing leaves the machine and nothing lands on an API bill. You'll see how much memory the model takes while it runs and why a bigger model answers more slowly on the same hardware, which is what decides whether a model fits on a laptop or needs a data center GPU. We finish with Unsloth Desktop, a second way to run models locally, so you can weigh a hosted API against running a model yourself.
In this lecture, we cover prompt engineering, the practice of structuring your input on purpose to get a reliable answer back instead of a different one every time you ask. You'll see the parts that make up a strong prompt, the difference between a system prompt and a user prompt, and two techniques that change output quality the most, few-shot examples and chain of thought reasoning. In the demo, you'll watch a weak prompt get refactored step by step into one you can actually trust.
In this lightboard lecture, we break down the retrieval step in retrieval-augmented generation (RAG) and show how it finds relevant source material before a large language model (LLM) answers. You'll see how documents are split into chunks, converted into embeddings, stored in a vector database, and compared with an embedded query through vector search. By the end, you'll understand how a RAG system can retrieve useful passages even when the user's words differ from the source documentation.
In this lightboard lecture, we complete the retrieval-augmented generation (RAG) pipeline by showing how an application builds the input sent to a large language model (LLM). You'll see how model instructions, the user's question, and retrieved source text are combined as context for a single request. The model then uses that temporary context to generate a grounded answer without being retrained.
Learn how to choose between RAG, fine-tuning, and long context based on the problem you’re trying to solve. We’ll compare how each approach works, where each fits best, and the tradeoffs around freshness, cost, maintenance, and model behavior. By the end, you’ll have a practical decision framework for knowing when to retrieve information, provide it directly in context, fine-tune the model, or combine multiple approaches.
In this lightboard lecture, we draw what actually makes something an AI agent: the same LLM as always, no smarter than before, now run inside a loop that keeps acting until it reaches a goal instead of answering once and stopping. You'll follow the agent loop one step at a time, decide, act, observe, and repeat, and see how the model picks each action while something outside it does the real work, and how the loop itself decides when the job is done. It's the foundation for everything else in agentic AI, and it comes down to one line: a model answers a question, but a loop finishes the job.
In this lightboard lecture, we name the software wrapped around the model that turns a large language model (LLM) into a working agent: the harness. You'll see its four jobs: running the agent loop, executing the actions, feeding the model its context every turn, and enforcing the limits, which is why the model is only half the product.
In this lightboard lecture, we look at what happens when one agent loop isn't enough and you move to a multi-agent system: instead of a bigger, smarter agent, you run more of them and coordinate the work. You'll see how a subagent is just another agent called like a tool, and what an orchestrator does to split a big goal into pieces and combine the results that come back. By the end you'll recognize the same agent loop repeating underneath any multi-agent system, with an orchestrator on top.
In this lecture, we look at how to build custom agents and subagents with their own system prompt, context window, tools, model, and permissions. You'll see how a main agent delegates work to a specialist and what goes into a Markdown agent definition with YAML frontmatter. The final slides cover Claude Agent SDK configuration, least privilege, model selection, version control, and tests for both correct routing and non-trigger cases.
In this lecture, we look at human in the loop (HITL), the practice of letting an agent do the work on its own while pausing for a person at the points where the consequences justify it. You'll learn the difference between approvals that gate a specific action before it runs, checkpoints that review an agent's progress partway through a workflow, and guardrails, the system-enforced limits an agent can't cross even when no one is watching. By the end you'll know how to match the level of oversight to how risky and how reversible an action is, and why approval fatigue can quietly undo the whole point.
In this lightboard lecture, we look at why connecting AI assistants to real tools was broken before a common standard existed, a problem known as N×M. You'll see what happens when three assistants each need their own custom connector to GitHub, Slack, Jira, and a database, why every new tool or app multiplied the work, and how a single shared standard turns N×M into N+M. This is the problem the Model Context Protocol (MCP) was built to solve, and it sets up everything in the next lecture.
In this demo, you'll watch an AI agent connect to a Model Context Protocol (MCP) server and go from answering questions to doing work. We plug a small Python agent into GitHub's official MCP server, let it discover the server's tools instead of hand-coding them, and then watch it read a support ticket and a log file from a repository, post its triage as a comment, and open an escalation issue on its own. Along the way you'll see the MCP host, client, and server as real code and a real process, and what a tool catalog costs you in context window space.
In this lecture, we break down orchestration frameworks, the libraries that run the agent loop for you, and compare LangChain, LangGraph, and CrewAI with a minimal code example for each. By the end you'll know which one matches the shape of your workflow, and when a plain loop is all you need.
In this lecture, you'll learn how Agent Skills give AI agents reusable instructions they can discover and load when a task calls for them. We break down the SKILL.md format, YAML frontmatter, skill descriptions, and progressive disclosure, then show how references, scripts, and assets keep detailed workflows out of the context window until they are needed. You'll also learn when to build a skill, how to find existing options on skills.sh, and how to turn a repeated prompt into a workflow your team can version and reuse.
In this demo lecture, a support ticket carrying hidden instructions turns an AI agent against itself, making it leak internal data it was never supposed to share, without anyone touching the agent's code. You'll see the difference between direct and indirect prompt injection, why a language model can't reliably tell its instructions apart from the data it reads, and how the same attack slips through disguised as a routine footer. Then we cut the risk two ways on camera, with least-privilege tool access and a hardened system prompt, and set up the governance, identity, and observability topics that close out the security section.
In this lecture, we bring the components of an AI agent together by following a VPN support request from prompt to verified result. You'll see how the model and harness use skills and retrieval-augmented generation (RAG), connect to tools through the Model Context Protocol (MCP), and require human approval before sensitive actions. We then revisit the same architecture to examine token costs and security risks, giving you a practical way to understand how these systems work.
In this lecture, we wrap up the AI Foundations course with practical next steps for building hands-on experience. I'll encourage you to tackle one small task each week and work through problems in your own lab to build confidence. We'll also discuss keeping up with AI developments and revisiting earlier challenges as tools improve.
Check out bonus content here!
Join the Inner Circle. Build the skills and confidence to pass the cert and deliver in production.
Somebody in a meeting asks why the new AI assistant keeps inventing ticket numbers. Everyone looks at the most technical person in the room. If that person is you, and you don't have an answer, then this course is for you.
I built this course for the technical professionals who suddenly have AI in every product they touch: solution engineers, DevOps and platform engineers, IT pros, architects, and anyone else expected to have real answers. Why do models hallucinate? What does a token actually cost? Should your team call a hosted API or run something local? What is MCP, and why does every vendor slide deck suddenly mention it?
Most AI training today falls into two buckets. The beginner courses that teach you to draft emails with ChatGPT, or the engineering bootcamps assume you want to spend 30 hours writing Python. Almost nothing exists for the person in between, who needs a working understanding of the whole system more than another tool tutorial. That gap is the reason this course exists.
***
How the course works
We spend about five and a half hours building up one complete picture of the modern AI stack. I teach the way I always have on Udemy: short, focused lectures, with a lightboard for the concepts that need drawing out and live demos so you can watch everything actually run.
The whole course follows a single running example, a fictional internal support assistant called HelpDesk AI. Every section adds one layer to it. First we look inside the model it runs on. Then we price that model out and decide where it should live. We give it company knowledge with RAG, turn it into an agent, wire it to real systems over MCP, and then, in my favorite section, we attack it and watch it leak data. By the final lecture, the entire architecture is on the lightboard, and every piece of it is something you understand.
Along the way you'll learn:
How large language models actually work: parameters, training vs inference, and why the same prompt gives different answers
Tokens and context windows, and how they drive both model behavior and your bill
Token economics: estimating what an AI workload really costs, and how prompt caching changes the math
Choosing a model: the tradeoffs between performance, cost, and latency, plus LLM vs SLM and open vs closed weights
Hosted APIs vs self-hosted models, including a demo running a local model with Ollama
Prompt engineering that goes past the basics, and where context engineering takes over
Retrieval-augmented generation (RAG): embeddings, vector search, and when to use RAG vs fine-tuning vs a bigger context window
AI agents and agentic AI: the agent loop, tool calling, subagents, orchestrators, and human-in-the-loop controls
What a harness is, and why the model is only half the product
Model Context Protocol (MCP): hosts, clients, servers, and a live demo connecting one
Orchestration frameworks like LangChain, LangGraph, and CrewAI, explained in plain terms so you know when you'd care
AI security fundamentals: prompt injection, over-privileged agents, and why observability matters
Every section ends with a quiz so you can validate these ideas yourself.
***
Before you enroll
You will see code in this course, because pretending AI systems involve no code would be silly. An agent created using Python. A tool call is JSON. An MCP config is a file. I put them on screen and walk through what they mean, and you will never be asked to write or debug anything. The only prerequisite is general technical literacy. If you can explain what an API is, you're technical enough for this course.
Fair warning on scope: we cover some of the basic security fundamentals here. The deep material on securing AI systems belongs to a separate, dedicated course. There's just too much there to include it.
***
About the Instructor
I'm Bryan Krausen. I've spent years teaching Terraform, Vault, and GitHub certification courses here on Udemy, and I work with this technology hands-on as a consultant. Every course I make follows the same pattern: draw the concept, demo the real thing, then hand you an exercise.
Words like agents, RAG, MCP, and context windows are showing up in vendor pitches and architecture reviews right now, whether you feel ready for them or not. After this course, you'll know exactly what sits behind each one, and you'll be the person in the room who can explain it.
Enroll and let's get started.