
Build an analytics data pipeline from Confluent Kafka to Google Cloud, ingesting JSON data, syncing hourly to Cloud Storage, transforming with Cloud Data Fusion, loading BigQuery, and visualizing in Looker.
Create a confluent account, set up a Google Cloud Platform cluster, and generate an API key to connect to the cluster, with guidance on endpoints and key management.
Create a fully managed Datagen source connector in Confluent Kafka, generating sample transactions data in JSON and pushing it into a topic.
Create a cloud storage bucket on Google Cloud Platform for the confluent kafka data pipeline, named confluent_gcp_data, with a hierarchical namespace, single region us east one, and public access blocked.
Create a Google Cloud service account and download a JSON key to enable the Confluent Kafka connector to access a Cloud Storage bucket for the streaming data pipeline.
Create a fully managed Google Cloud Storage sink connector to move data from a confluent topic to a GCP bucket, using JSON output and hourly path formatting.
Create a BigQuery dataset and table to support a streaming data pipeline with Confluent Kafka on Google Cloud, and connect Looker for dynamic operational reports using SQL queries.
Demonstrate validating a Confluent to Google Cloud data pipeline, from source connector to Confluent topic and storage bucket, enabling Cloud Data Fusion, BigQuery, and Looker Studio analytics.
Learn to create a Cloud Data Fusion instance on Google Cloud to run ETL for a streaming pipeline, including enabling APIs, authorization, and monitoring instance creation.
Build a streaming data pipeline from Google Cloud Storage to BigQuery using Confluent and Cloud Data Fusion Studio, configure GCS source, JSON schema, and insert mode, then deploy and schedule.
Develop and deploy an upsert-enabled ETL pipeline in Google Cloud Data Fusion that ingests JSON from Google Cloud Storage buckets via Confluent Kafka into BigQuery for analytics with Looker Studio.
Build and execute a streaming data pipeline using Confluent Kafka and Google Cloud, moving data from Cloud Data Fusion to Google Cloud Storage and into BigQuery for analytics in Looker.
Verify end-to-end data flow from a source connector and confluent topic through the sync connector into Google Cloud Storage, run Cloud Data Fusion ETL to BigQuery, and refresh Looker reports.
Review end-to-end analytics pipeline from Confluent Kafka to Google Cloud, detailing data ingestion, cloud storage, ETL with Cloud Data Fusion, and Looker reporting to BigQuery.
Course Overview
In this hands-on course, participants will follow a step-by-step approach to build a real-time streaming analytics data pipeline that integrates Confluent Kafka with Google Cloud Platform services. Learners will design a pipeline that streams data from Kafka into Google Cloud Storage, processes it with Dataflow, stores it in BigQuery, and finally visualizes operational insights using Looker Studio. This practical journey ensures participants not only grasp the concepts but also apply them directly in a cloud-native environment.
Google Cloud Platform Experience
On the GCP side, participants will gain exposure to key services including Cloud Storage, Dataflow, BigQuery, and Looker Studio. They will also learn essential skills in logging, monitoring, troubleshooting, and configuration, which are critical for building secure, scalable, and reliable streaming applications in the cloud.
Confluent Kafka Experience
On the Kafka side, learners will set up a fully managed Kafka cluster, create topics for message streaming, configure a Datagen Source Connector to simulate real-world data ingestion, and integrate with GCP Storage Sink Connector. This provides valuable hands-on experience with enterprise-grade Kafka infrastructure—without the complexity of managing it manually.
Learning Outcomes
By the end of the course, participants will have:
Built a working, scalable streaming analytics pipeline.
Gained insights into cloud-native architectures and streaming design patterns.
Acquired hands-on skills directly applicable to real-world projects and professional roles.
Who Should Attend
This course is ideal for Cloud Engineers, Data Engineers, Product Owners, Product Managers, Scrum Masters, and Technology Leaders seeking practical experience in building real-time streaming analytics pipelines with Confluent Kafka and Google Cloud.