Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
Batch Processing with Apache Beam in Python
Rating: 4.2 out of 5(63 ratings)
286 students

Batch Processing with Apache Beam in Python

Easy to follow, hands-on introduction to batch data processing in Python
Created byAlexandra Abbas
Last updated 9/2020
English
English [Auto],

What you'll learn

  • Core concepts of the Apache Beam framework
  • How to design a pipeline in Apache Beam
  • How to install Apache Beam locally
  • How to build a real-world ETL pipeline in Apache Beam
  • How to read and write CSV data from Apache Beam
  • How to apply built-in and custom transformations on a dataset
  • How to deploy your pipeline to Cloud Dataflow on Google Cloud

Course content

3 sections19 lectures1h 9m total length
  • Welcome2:17

    Design, build, and deploy batch data pipelines with Apache Beam in Python, leveraging extract, transform, load concepts and Beam as a unifying framework with Cloud Dataflow on Google Cloud Platform.

  • What is Apache Beam2:14

    Beam unifies parallel data processing across APIs and translates code to Spark, Flink, or Hadoop, enabling write-once pipelines; the course uses the Python SDK for batch and streaming pipelines.

  • Apache Beam concepts2:06

    Explore beam concepts: a pipeline defines a beam program with command-line options, while P collections enable parallel data processing and B transforms read, write, and process bounded or unbounded data.

  • Design a pipeline2:24

    Design a pipeline that reads a CSV web analytics dataset, computes session duration from timestamps, maps IPs to countries, and averages duration by country, then writes results into fire.

  • Install Apache Beam2:51

    Create and activate a Python 3.7 virtual environment named Apache Beam Tutorial, install Apache Beam and the Google Cloud Dataflow package to deploy pipelines to Dataflow.

Requirements

  • Python programming experience
  • Having an idea of distributed data processing e.g. You have used Spark before
  • Having Conda (or other Virtual Environment Manager) installed on your machine

Description

Apache Beam is an open-source programming model for defining large scale ETL, batch and streaming data processing pipelines. It is used by companies like Google, Discord and PayPal.

In this course you will learn Apache Beam in a practical manner, with every lecture comes a full coding screencast. By the end of the course you'll be able to build your own custom batch data processing pipeline in Apache Beam.

This course includes 20 concise bite-size lectures and a real-life coding project that you can add to your Github portfolio! You're expected to follow the instructor and code along with her.

You will learn:

  • How to install Apache Beam on your machine

  • Basic and advanced Apache Beam concepts

  • How to develop a real-world batch processing pipeline

  • How to define custom transformation steps

  • How to deploy your pipeline on Cloud Dataflow

This course is for all levels. You do not need any previous knowledge of Apache Beam or Cloud Dataflow.

Who this course is for:

  • Data Engineers
  • Aspiring Data Engineers
  • Python developers interested in Apache Beam