Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
AI Agents, RAG & LLM Evals for Beginners: DeepEval & RAGAS
Bestseller
Role Play
Rating: 4.3 out of 5(776 ratings)
7,510 students

AI Agents, RAG & LLM Evals for Beginners: DeepEval & RAGAS

2026 - Path to AI QA Engineer to test LLMs and AI Apps using DeepEval, RAGAs and HF Evaluate with Local LLMs like Ollama
Created byKarthik KK
Last updated 6/2026
English
Arabic [Auto],German [Auto],

What you'll learn

  • Understand the purpose of Testing LLM and LLM based Application
  • Understand DeepEval and RAGAs in detail from complete ground up
  • Understand different metrics and evaluations to evaluate LLMs and LLM based app using DeepEval and RAGAs
  • Understand the advanced concepts of DeepEval and RAGAs
  • Testing RAG based application using DeepEval and RAGAs
  • Testing AI Agents using DeepEval to understand how tool callings can be tested

Course content

17 sections127 lectures14h 40m total length
  • Why Testing LLM Apps Is Completely Different from Traditional Testing9:11

    Develop testing and evaluation of large language models using traditional and non traditional metrics, ground truth comparisons, and tools like Ollama, DeepEval, and RAGs for local and hosted models.

  • Chatbots, AI Agents & RAG: The 3 Types of AI Apps You'll Be Testing13:09

    Learn the types of AI applications, including chatbots and AI agents, and how retrieval augmented generation and vector stores power real-world tasks with LLMs.

  • LLM Evaluation 101: The Different Ways to Measure AI Quality10:03

    Explore the basics of llm evaluation with prompts and compare human-based, code-based, and llm-based approaches using tools like DeepEval, RAGAs, and Hugging Face Evaluate to optimize prompts and deployment.

  • Human Graded Evaluation: Using the Anthropic Workbench to Score AI17:17

    Explore human graded evaluation for large language models using the Anthropic workbench, translating prompts and code across languages, and generating test cases for robust QA automation.

  • Every Key Metric You Need to Evaluate AI Apps (Relevancy, Bias & More)5:48

    Explore common evaluation metrics for AI applications, from answer relevancy and contextual precision to tool selection accuracy and bias detection, and learn how to assess function arguments.

  • DeepEval vs RAGAs vs OpenAI Evals vs HuggingFace: Which Tool Wins?8:03

    Compare major llm evaluation libraries—deep evals, ragas, OpenAI evals, Galileo, and Huggingface evaluate—highlighting observability, evaluation matrices, online monitoring, and addressing hallucination to test llm applications.

  • Check your knowledge!

Requirements

  • Basics of working with LLM like using ChatGPT
  • Basics of any programing language like Java or Javascript
  • Basics of python will be a plus

Description

Testing AI & LLM App with DeepEval, RAGAs & more using Ollama and Local Large Language Models (LLMs)

Master the essential skills for testing and evaluating AI applications, particularly Large Language Models (LLMs). This hands-on course equips QA, AI QA, Developers, data scientists, and AI practitioners with cutting-edge techniques to assess AI performance, identify biases, and ensure robust application development.



Topics Covered:

  • Section 1: Foundations of AI Application Testing (Introduction to LLM testing, AI application types, evaluation metrics, LLM evaluation libraries).

  • Section 2: Local LLM Deployment with Ollama (Local LLM deployment, AI models, running LLMs locally, Ollama implementation, GUI/CLI, setting up Ollama as API).

  • Section 3: Environment Setup (Jupyter Notebook for tests, setting up Confident AI).

  • Section 4: DeepEval Basics (Traditional LLM testing, first DeepEval code for AnswerRelevance, Context Precision, evaluating in Confident AI, testing with local LLM, understanding LLMTestCases and Goldens).

  • Section 5: Advanced LLM Evaluation (LangChain for LLMs, evaluating Answer Relevancy, Context Precision, bias detection, custom criteria with GEval, advanced bias testing).

  • Section 6: RAG Testing with DeepEval (Introduction to RAG, understanding RAG apps, demo, creating GEval for RAG, testing for conciseness & completeness).

  • Section 7: Advanced RAG Testing with DeepEval (Creating multiple test data, Goldens in Confident AI, actual output and retrieval context, LLMTestCases from dataset, running evaluation for RAG).

  • Section 8: Testing AI Agents and Tool Callings (Understanding AI Agents, working with agents, testing agents with and without actual systems, testing with multiple datasets).

  • Section 9: Evaluating LLMs using RAGAS (Introduction to RAGAS, Context Recall, Noise Sensitivity, MultiTurnSample, general purpose metrics for summaries and harmfulness).

  • Section 10: Testing RAG applications with RAGAS (Introduction and setup, creating retrievers and vector stores, MultiTurnSample dataset for RAG, evaluating RAG with RAGAS).



Who this course is for:

  • QA Engineers
  • AI QA Test Engineers
  • Business Analyst
  • AI Engineers