
This course contains the use of artificial intelligence.
Building an LLM application is easy. Knowing whether it actually works is much harder.
This course teaches you how to evaluate, test, debug, and improve AI systems built with large language models, Retrieval-Augmented Generation (RAG), and AI agents. Instead of relying on a few hand-picked prompts and deciding that an output “looks good,” you will learn how to approach AI evaluation like an engineer: define measurable criteria, build evaluation datasets, detect regressions, analyze failures, and make evidence-based decisions about your system.
You will explore why evaluating generative AI is fundamentally different from traditional software testing. LLM outputs are probabilistic, multiple answers can be valid, and a response can sound convincing while still being incorrect. Throughout the course, you will learn how to turn these challenges into structured and repeatable evaluation workflows.
A major part of the course focuses on LLM-as-a-Judge. You will learn how model-based evaluators work, where they are useful, and the biases and failure modes that can make evaluation scores misleading. You will design evaluation rubrics, compare model outputs, measure qualitative characteristics, and understand when automated evaluation should be combined with deterministic checks and human judgment.
For RAG systems, you will go beyond evaluating only the final generated answer. You will examine retrieval quality, context relevance, answer correctness, groundedness, and faithfulness so you can determine whether a failure originated from the retriever, the supplied context, the generation step, or the underlying knowledge base.
You will also learn how evaluation changes when working with AI agents. Agentic systems introduce tool calls, intermediate decisions, multi-step workflows, and trajectories that cannot be properly evaluated by checking only the final response. You will learn how to evaluate task completion, tool-use correctness, intermediate behavior, and overall agent reliability.
The course also covers evaluation in CI and production workflows, helping you understand how to detect regressions when prompts, models, retrieval pipelines, tools, datasets, or application logic change.
By the end of this course, you will be able to:
Design meaningful evaluation datasets and test cases
Define metrics and evaluation criteria for LLM applications
Use LLM-as-a-Judge effectively and understand its limitations
Identify evaluator bias and unreliable scoring
Evaluate RAG retrieval quality, groundedness, faithfulness, and answer quality
Test AI agents and multi-step workflows
Perform systematic error and failure analysis
Compare prompts, models, and system configurations
Detect regressions as AI applications evolve
Integrate AI evaluations into CI and engineering workflows
This course is designed for AI engineers, software engineers, ML engineers, developers, and technical practitioners who want to move beyond AI demos and build LLM-powered applications that can be tested, measured, monitored, and improved systematically.
If you can build an AI system, the next skill is learning how to prove that it works.