
Review the exam guide to understand the data engineer exam and take the assessment test, then build hands-on Google Cloud skills with $300 credit and study guides to prepare.
Understand what object storage is used for
Use gsutil, Transfer Service and other methods to upload data.
Use IAM and access control lists to limit access to data in Cloud Storage
Use policies to manage objects in Cloud Storage
Use GCP console to manage Cloud Storage
Check your understanding of Cloud Storage
Know the solution to the exercise.
Effectively use GCP's managed relational database options
Use Cloud SQL for regional and zonal relational databases
Create Cloud SQL databases
Learn to monitor Cloud SQL databases using the cloud console, track metrics like CPU, memory, storage, and ingress/egress, adjust time spans, and spot performance trends.
Learn how to securely connect to Cloud SQL with the Cloud SQL Auth Proxy, Google's recommended solution that uses service account credentials and automatic TLS for public or private endpoints.
Install the Cloud SQL Auth proxy on a Compute Engine VM by installing the Postgres client, downloading the proxy, and configuring it with your instance name to enable connections.
Check your knowledge of Cloud SQL
Review the correct way to deploy a Cloud SQL database
When to use Cloud Spanner for multi-regional and global database applicaitons
Create a Cloud Spanner instance.
Optimized Cloud Spanner I/O performance
Learn how Google Cloud Datastore Firestore uses entities and kinds to model data with flexible properties, JSON-like key-value structures, and support for atomic values, arrays, and other entities.
Understand how indexing controls queries in Cloud Firestore, with automatic single-field indexing for maps and arrays and manual composite indexes for multi-attribute filters.
Explore cloud data store basics by creating and viewing entities and kinds, using namespaces, and leveraging flexible properties across departments and products.
Cloud Firestore supports atomic transactions with serializable isolation, ensuring all operations succeed or fail together and reads see a consistent snapshot.
Practice creating a Cloud Firestore kind named store and defining an entity with an auto-generated ID, a store name, and the city where the store is located.
Create a kind named store and add entities with properties like name and city in the default namespace using numeric key identifiers, after completing initial setup and selecting a mode.
Design row keys in Bigtable to evenly distribute reads and writes across nodes and Colossus, avoid hotspots, and use high cardinality prefixes with time-reversed components.
Learn design patterns for time series data in Bigtable, including time buckets, row keys, and single event, serialized and unserialized storage strategies, with tradeoffs in performance and storage.
Discover BigQuery, Google's serverless analytical data warehouse that uses SQL for analytics. Leverage data sets, tables, views, and materialized views with federated queries across storage systems.
Explore BigQuery's sql interface with simple selects and unnest arrays and nested structures. See how to reference records, arrays, and BigQuery's analytical nature, not a relational database.
Run a hands-on BigQuery query on the Census Bureau International and Midyear Population public dataset to list countries with a 2021 population over 100 million.
Query the census bureau international mid-year population table in BigQuery public data using standard SQL to list countries with a mid-year population over 100 million in 2021.
Learn how to apply access controls in BigQuery with IAM at organization, project, dataset, and table levels, use column level security with Data Catalog tags, and understand BigQuery roles.
BigQuery partitions to boost query performance and reduce costs by scanning less data, using ingestion time, timestamp, null, and integer range partitions, with partition filters; partitioning is preferred over charting.
Load data into BigQuery from Google Cloud storage, using streaming inserts, bulk loads, Avro or Parque formats, and Cloud Dataflow for transformations with manual, API, or bq cli loading.
Learn how to control BigQuery access with IAM by assigning roles to identities, understanding permissions, and applying predefined, custom, and basic roles at organization, project, dataset, or table levels.
Define policies and taxonomy tags to control column level access to sensitive data, apply policy tags to columns, and enforce taxonomy-based access for highly critical information.
Prepare and discover data warehouse workloads, assess the current state, plan the migration, execute offload or full migrations, and validate results in BigQuery.
Data pipelines drive data warehousing by processing data through extraction, transformation, and load, enabling data enrichment and realtime analysis, then loading into the data warehouse.
Explore reporting and analysis for data warehouse migrations, featuring descriptive, predictive, and prescriptive analytics, and leverage Data Studio, Looker, and BigQuery for OLAP querying and BigQuery Geo Viz.
Learn how Cloud Memory Store, a managed caching service for Redis and Memcached, delivers low-latency caching, high availability, and seamless integration with App Engine, Compute Engine, Cloud Functions, and Kubernetes.
Create a Redis cache in Cloud Memorystore, selecting an instance, region, and basic tier, set capacity and parameters, note costs, configure networking and IP/port, then delete the cache.
learn how Cloud Composer, a managed Apache Airflow service, defines and runs DAG-based workflows as Python scripts stored in cloud storage, with logging via the Airflow web interface.
Discover cloud data fusion, a managed, code-free ETL tool built on CDAP that connects AWS, GCP, and Azure with 150-plus connectors via a drag-and-drop interface.
The need for data engineers is constantly growing and certified data engineers are some of the top paid certified professionals. Data engineers have a wide range of skills including the ability to design systems to ingest large volumes of data, store data cost-effectively, and efficiently process and analyze data with tools ranging from reporting and visualization to machine learning. Earning a Google Cloud Professional Data Engineer certification demonstrates you have the knowledge and skills to build, tune, and monitor high performance data engineering systems.
This course is designed and developed by the author of the official Google Cloud Professional Data Engineer exam guide and a data architect with over 20 years of experience in databases, data architecture, and machine learning. This course combines lectures with quizzes and hands-on practical sessions to ensure you understand how to ingest data, create a data processing pipelines in Cloud Dataflow, deploy relational databases, design highly performant Bigtable, BigQuery, and Cloud Spanner databases, query Firestore databases, and create a Spark and Hadoop cluster using Cloud Dataproc.
The final portion of the course is dedicated to the most challenging part of the exam: machine learning. If you are not familiar with concepts like backpropagation, stochastic gradient descent, overfitting, underfitting, and feature engineering then you are not ready to take the exam. Fortunately, this course is designed for you. In this course we start from the beginning with machine learning, introducing basic concepts, like the difference between supervised and unsupervised learning. We’ll build on the basics to understand how to design, train, and evaluate machine learning models. In the process, we’ll explain essential concepts you will need to understand to pass the Professional Data Engineer exam. We'll also review Google Cloud machine learning services and infrastructure, such as BigQuery ML and tensor processing units.
The course includes a 50 question practice exam that will test your knowledge of data engineering concepts and help you identify areas you may need to study more.
By the end of this course, you will be ready to use Google Cloud Data Engineering services to design, deploy and monitor data pipelines, deploy advanced database systems, build data analysis platforms, and support production machine learning environments.
ARE YOU READY TO PASS THE EXAM? Join me and I'll show you how!