
If you're on the fence about taking this course, this introduction is for you! You will learn:
What Grafana MCP is and why an AI coding assistant cannot see production data without it
About your instructor, a senior SRE with ~17,000 learners on various platforms
What you will run locally, on Grafana Cloud, and on AWS
The AI clients we connect to
The complete course code is in Chapter 3, video 1.
This course is for developers who want production errors in the editor, SREs who want faster incident triage, and platform engineers who need to run MCP on AWS.
If you're STILL on the fence, this lecture goes into more detail about what you'll learn.
At the end of this course, you will be able to:
Explain MCP at a fundamental level, including architecture and security
Run Grafana MCP locally and on Grafana Cloud
Deploy Prometheus, Loki, Grafana, and a demo app on AWS with Terraform
Implement production-grade security hardening of MCP on ECS and EKS.
Query live data through an agent from at least one of Cursor, Claude Code, Devin Desktop, or OpenCode.
If you want to learn:
What the Model Context Protocol is and why teams stopped writing one-off Grafana plugins for AI
The difference between an MCP host, client, and server
What tools, resources, and prompts are
How Grafana MCP is tool-heavy
Then this lecture is for you.
MCP is an open standard for Cursor or Claude to discover tools on a server instead of pasting JSON into chat.
This lecture breaks down MCP Architecture in a way that's easy to understand. The MCP protocol speaks JSON-RPC on the data layer and transports it over stdio, SSE, or streamable HTTP.
We set the foundation for important decisions later about what transport types to use locally vs. on AWS.
This lecture explains why MCP tools are ideal for observability workflows. In a time-sensitive incident, you will appreciate an agent that can quickly gather relevant information and correlate root causes. And when building new systems, you will appreciate an agentic code editor that understands the current state of production.
This lecture introduces you to Grafana Labs' official MCP server. We discuss how to set the Grafana URL, create a service account token, and how different tools work.
The official documentation and GitHub repository are attached to this lecture as resources.
You will be tempted to run the official Docker image on port 8000 and connect your AI client straight away. Once that container is up, it listens with no authentication of its own. Anyone who can reach the port can list dashboards, run PromQL, read logs, and change alerts without stealing a password.
By the end of this lecture, you will be able to take each attack in the MCP specification and describe how it affects the real Grafana MCP server. In Chapter 5, we will close these gaps on AWS.
If you want to learn:
How to run Grafana MCP locally in Docker
When Docker should use stdio versus SSE versus streamable HTTP
Then this lecture is for you.
We spin up Grafana on port 3000, create a service account token, and test our solution. The complete course code is attached to this lecture as a downloadable ZIP, so keep it handy for the AWS and client chapters.
If you want to learn:
Why Grafana MCP should use a service account, not your user login
What role and token expiry keep the assistant from being a Grafana admin
How to test the MCP server with MCP Inspector
Then this lecture is for you.
Grafana Cloud is hosted Grafana plus managed Mimir, Loki, and Tempo. Using Grafana Cloud means you don't have to manage your own MCP server and Grafana instances in production. This lecture explains everything in case that's your preferred route.
If you want to learn:
How to create a Grafana Cloud stack
How to create a Cloud service account token for MCP
How to point your MCP server at your Grafana Cloud instance
Then this lecture is for you.
This lecture is the payoff for Chapter 3. You learn the following:
How to wire Grafana MCP into Cursor for the first time
How to list dashboards, inspect metrics, and query logs in plain English
What the tool-call panel looks like when the model runs PromQL or LogQL for you
We will connect Cursor to Grafana Cloud and watch how the model picks tools. In Chapter 6, we'll go into more detail about various clients.
So far, you have been running Grafana on your laptop. This lecture is where you get ready to build the rest of the environment on AWS with Terraform, so you can describe the infrastructure you want and apply the same project again later.
You do not need to have used Terraform before. We go through providers, state, variables, outputs, and modules, and through the commands you will actually run: init, plan, apply, and destroy. By the end, you should be able to read the files in this project and know what each one is for.
If you want to learn:
Why a Grafana stack like this should not sit in the account’s default VPC
Basic subnet and Availability Zone design
How to create public and private subnets across AZs with the official AWS VPC module
Which NAT and subnet tags EKS load balancers will need later
Then this lecture is for you. We will create the network foundation for the rest of the videos in this chapter.
If you want to learn:
How to provision Amazon EKS with Terraform’s official module and a managed node group
What AWS runs (control plane) versus what you run (worker nodes)
How to install the Kubernetes provider and talk to the cluster with kubectl
Then this lecture is for you.
In this lecture, you'll learn the following:
How to deploy Prometheus, Loki, and Grafana on EKS with Terraform’s Helm provider
What kube-prometheus-stack gives you, including the ServiceMonitor CRD
How this Grafana becomes the backend Grafana MCP will query later
In this lecture, we'll ship the full demo app on EKS (Deployment, Service, ConfigMap) with the Kubernetes provider. By the end, you'll understand the ServiceMonitor better and how logs can reach Loki via Promtail.
The observability stack from the last chapter is on AWS, but the MCP server your assistant uses to query it has only ever run on your laptop. This chapter puts that server in production.
By the end of this chapter, you will be able to run Grafana MCP on AWS, either on ECS Fargate or EKS. We use ECS Fargate as the default, because there is very little to operate.
But if you already run Kubernetes and Grafana lives in the cluster, the official Helm chart is the better path, since MCP can reach Grafana over the internal network.
In this lecture, you will learn:
How to run the Grafana MCP server on ECS Fargate
How to put the Grafana token in Secrets Manager instead of the task definition
How to validate the health of the MCP server after deployment
In this lecture, you will learn:
How to put the MCP task in private subnets behind an ALB
How to allow port 8000 only from the load balancer’s security group
The role of VPC endpoints
In this lecture, you will learn:
The difference between the Execution role and the Task role on the MCP Fargate task
How to scope IAM to Secrets Manager and nothing extra
How to rotate the Grafana service account token without taking MCP down
In this lecture, you will learn:
How to install Grafana’s grafana-mcp Helm chart on EKS
How MCP reaches Grafana on the in-cluster network, never the public internet
Where the token lives as a Kubernetes Secret versus ECS Secrets Manager
In this lecture, you will learn:
Why TLS on the ALB is not the same as authenticating callers of Grafana MCP
How to put OIDC or Cognito on the ALB before traffic reaches the task
How to use WAF, private endpoints, and audit logging
In this lecture, you will learn how to monitor the MCP server you've deployed.
How to turn on mcp-grafana’s Prometheus metrics (--metrics, /metrics)
How to scrape it with a ServiceMonitor
What metrics to look out for
In this lecture, you will learn:
How to add Grafana MCP with claude mcp add
Local versus project scope
stdio Docker versus HTTPS to Grafana Cloud’s hosted MCP
In this lecture, you will learn:
How to add Grafana MCP in ~/.cursor/mcp.json or project .cursor/mcp.json
How to toggle tools so the agent is not dumped with forty Grafana tools at once
How to debug from the editor: code, dashboards, PromQL, and Loki in one window
In this lecture, you will learn:
How to configure Grafana MCP for Devin Local (not Cascade)
Where .devin/mcp_config.json lives
In this lecture, you will learn:
How OpenCode’s opencode.json declares Grafana MCP with an explicit type: local or type: remote
How to verify with built-in opencode mcp commands
When you started, you probably understood what Grafana MCP was, but your assistant had no way to see your dashboards, metrics, or logs. This recap walks through the system you can now build and run yourself.
You should leave this recap satisfied with everything you've learnt, from the MCP clients to a server you control on AWS to the metrics and logs. The source code is yours to reuse on your own account. If the lab environment is still running when you finish, please destroy so it does not keep billing you.
I have taught ~17000 students about monitoring with Prometheus and Grafana. Recently, I've been receiving several questions about how to deploy Grafana MCP in production. This course contains an extensive library of information to help you get up and running with MCP from the ground up. You will gain practical experience that you can immediately apply in your production systems.
At the end of the course, you'll have experience with the following:
Running Grafana MCP locally with Docker and connecting it to a real Grafana instance.
Connecting Grafana MCP to Grafana Cloud and query Cloud-backed metrics, logs, and dashboards.
Deploying Prometheus, Loki, Grafana, and a demo app on AWS with Terraform.
Deploying Grafana MCP on Amazon ECS Fargate and Amazon EKS with TLS, IAM, and Secrets Manager.
Querying live dashboards, metrics, logs, and alerts from Cursor, Claude Code, Devin Desktop, and OpenCode.
Explaining MCP architecture: clients, servers, tools, and transports (stdio, SSE, and streamable HTTP).
Scoping Grafana service accounts and MCP tools so the assistant runs with least privilege, not admin.
Identifying MCP security risks and apply the controls you actually need on AWS.
Choosing an ECS or EKS architecture for Grafana MCP and monitoring the MCP server itself.
Investigating incidents from your editor using live Prometheus and Loki data instead of screenshots.