
Learn to harness data streaming with AWS Kinesis to build real-time analytics and event-driven architectures using serverless, fully managed, scalable cloud technologies in AWS.
Explore Kinesis data streams, Kinesis data firehose, and Kinesis data analytics, and learn how immutable logs decouple producers from consumers for real-time processing.
Prepare by having programming experience in Java or Python, a basic understanding of the AWS cloud, and familiarity with the AWS CLI, S3, IAM, and CloudWatch for real-time streaming analytics.
Explore how pub/sub messaging underpins real-time data streaming with AWS Kinesis, enabling decoupled producers and consumers, fault tolerance, and scalable, event-driven architectures.
Learn to create a Kinesis data stream using the AWS CLI, including setting a unique stream name and shard count, then describe, delete, and verify the stream.
Produce individual records to a Kinesis data stream with the Java SDK and the put record API, using JSON orders. The demo uses a UUID partition key and shard distribution.
Demonstrates producing single, ordered records to a Kinesis data stream with the Java SDK put record API, using seller id as the partition key and sequence numbers for ordering.
Demonstrates producing individual records to a Kinesis stream with the Python boto3 SDK using put_record, including a unique order payload and partition keys for shard distribution.
Learn how to write individual records to a Kinesis stream using Boto3 with guaranteed ordering via a sequence number and the seller_id partition key.
Provide a Java 11 walkthrough using the AWS SDK to batch 20 records with PutRecords for Kinesis data stream, including JSON order serialization, partition keys, and per record error handling.
Use the Python boto3 sdk to batch and publish json-encoded records to a kinesis data stream via put records api, with 20-item batches and order id as partition key.
Discover the Java Kinesis Producer Library for high throughput data ingestion, which batches and aggregates records, uses a C++ daemon, and provides intelligent retries with CloudWatch metrics.
Learn to consume records from a Kinesis data stream with python and boto3, listing shards, creating shard iterators, and using get_records in a continuous poll loop.
Inspect the DynamoDB lease table used by the Kinesis client library to checkpoint for each shard, showing one record per shard and the last processed sequence number.
Explore how serverless lambda acts as a kinesis consumer, with standard pole-based and enhanced fan-out delivery, and tune batch size, windowing, starting position, and split batch on air.
Set up a java-based AWS SAM project to deploy a Lambda as a Kinesis data stream consumer, using CloudFormation templates and the SAM CLI for packaging, testing, and deployment.
Demonstrates building a kinesis lambda consumer in java using the serverless application model with the java 11 runtime, configuring maven dependencies, deploying via sam, and validating with cloudwatch logs.
Demonstrates using AWS Lambda as an enhanced fan out consumer for a Kinesis stream, delivering push-based, high-throughput processing with a Java Lambda consumer at two megabytes per second.
Monitor kinesis data streams with cloudwatch metrics like incoming bytes and records, read and write throughput, latency, and iterator age to keep producers and consumers in sync.
Understand Kinesis data firehose capacity by region: regions handle up to 5000 records per second and 5 MB per second, others reach 1000 records per second and 1 MB max.
Create a Kinesis data stream, feed ordered data, and connect it to a Kinesis Data Firehose delivery stream to export raw orders to Amazon S3 for long-term storage.
Demonstrates a Kinesis data firehose transformation using a lambda to compute order subtotals from order items and write totals to S3 via a SAM deployment.
Discover how Kinesis Data Analytics offers a managed serverless platform for real-time streaming data processing, including streaming SQL and a managed Apache Flink runtime with Zeppelin notebooks and Kafka.
Explore Kinesis data analytics streaming SQL, its SQL-like interface, and how streams, pumps, and external CSV reference data enable continuous analytics with sinks like Kinesis Data Streams, Firehose, and S3.
Explore Kinesis data analytics streaming SQL architecture, treating streams and S3 data as tables, performing continuous stateful aggregations with SQL, and routing results to Kinesis streams or Firehose.
Explore tumbling windows in Kinesis Data Analytics streaming SQL, using fixed-length, non-overlapping 60-second windows to count orders and refresh results every interval.
Learn sliding windows in aws kinesis data analytics for streaming sql, where each new event updates last 60 seconds and emits results, contrasting with fixed-length tumbling windows and double counting.
Master staggered windows in Kinesis data analytics to compute event-time based order counts by order creation time, handling late arrivals and out-of-order events, truncating to minute intervals.
Learn to build a Kinesis Data Analytics streaming SQL application to compute seller revenue over 30-second tumbling windows, using boto3 to create streams and a SQL script to output trend.
Demonstrates building an interactive Apache Flink SQL app in the Kinesis Data Analytics Studio notebook, creating input/output streams, AWS Glue metadata, and a 30-second tumbling window for seller revenue.
Real-time streaming technologies are growing in popularity among the many technological drivers of business innovation because users are increasingly demanding personalized experiences which adapt and respond to them based on their journey through digital products and services. The AWS Kinesis suite of stream persistence and processing services have come to be recognized as first class choice for achieving the kinds of event driven architectures feeding into real-time analytics.
In this course students learn to harness the power of Kinesis Data Streams (KDS) and Kinesis Data Firehose (KDF) to construct high-throughput, low latency, pipelines of data across a variety of architectural components leading to scalable and loosely coupled systems. Additional focus is placed on how these stream persistence technologies are used in conjunction with Kinesis Data Analytics to perform advanced, real-time, computations which drive informed business actions and insights.
The course goes beyond the theory of what these services are, making heavy use of demonstrations and code walkthroughs to give examples of how these technologies are used in practice. Most code examples are demonstrated in parallel using both the Python and Java programming languages in an effort to reach the largest audience of developers. However, some examples are presented only in one language in cases where either one language doesn’t support a particular functionality or is significantly less complex to demonstrate.