
Introduction to the course and Sagemaker Unified Data studio
Discover why SageMaker Unified Studio is a must-learn tool for careers in data and AI. This course offers beginners and intermediates a comprehensive overview of key concepts, analytical workflows, and AI development on AWS. Gain essential skills to prepare for interviews and build data-driven AI products. Stay ahead with evolving content that addresses the most relevant topics in the data and AI landscape.
The Amazon SageMaker Unified Studio Administrator Guide provides comprehensive instructions for setting up and managing the Unified Studio environment. It covers topics such as creating domains, managing user access, associating AWS accounts, and configuring project profiles for data analytics, AI/ML model development, and generative AI applications. This guide is essential for administrators to effectively deploy and maintain a collaborative workspace that integrates AWS data, analytics, AI, and machine learning services.
The Amazon SageMaker Unified Studio Administrator Guide provides comprehensive instructions for setting up and managing the Unified Studio environment. It covers topics such as creating domains, managing user access, associating AWS accounts, and configuring project profiles for data analytics, AI/ML model development, and generative AI applications. This guide is essential for administrators to effectively deploy and maintain a collaborative workspace that integrates AWS data, analytics, AI, and machine learning services.
This interactive session provides participants with hands-on experience in managing SSO Users, SSO Groups, and IAM Roles within SageMaker Unified Data Studio. Learn to configure user access, assign group-based permissions, and integrate SSO for seamless authentication.
This interactive, hands-on session guides participants through SDS Domains, Domain Units, and Projects in SageMaker Unified Data Studio. Participants will create, configure, and manage domains, set up project workflows, and practice real-world scenarios. Gain practical experience to enhance collaboration, resource management, and project execution for your Data/AI team.
In this session, administrators will learn how to leverage blueprints within Amazon SageMaker Unified Studio to standardize project configurations and streamline resource provisioning. Participants will gain insights into enabling or disabling specific blueprints, managing blueprint authorizations, and customizing project profiles to align with organizational requirements. By mastering these tools, administrators can enhance team productivity and maintain consistent development environments across projects.
In this interactive session, participants will gain hands-on experience with Account Association in SageMaker Unified Studio. Learn to publish data from external AWS accounts into the SageMaker catalog, request and manage account associations, and enable cross-account project collaboration. Participants will practice setting up IAM permissions, authorizing domain accounts, and leveraging shared resources across multiple AWS accounts to drive unified data workflows.
This session equips administrators with the skills to implement data mesh architectures using Amazon SageMaker Unified Data Studio. Participants will learn to establish a centralized data domain and catalog across multiple AWS accounts, promoting decentralized data ownership and scalability. Key topics include configuring cross-account data sharing, setting up unified data access , and implementing robust data governance and security measures. By the end of the course, attendees will be proficient in leveraging SageMaker Unified Data Studio to create a cohesive, federated data ecosystem within their organizations.
Identify SageMaker unified studio quotas, including 2500 JupyterLab instances, 2500 project members, and 200 microenvironments. Learn where to find admin-focused quota details in the documentation.
In this session, participants will explore the hierarchical organization of data within Amazon SageMaker Unified Studio, focusing on the structure of data domains, data products, and data assets. They will learn how data domains represent broad business areas, such as customer information or sales data, and how within these domains, data products are curated datasets or services tailored to specific business needs. Each data product comprises individual data assets, including databases, tables, or files containing the actual data. Understanding this hierarchy is essential for efficient data governance, discovery, and utilization across various projects and teams.
In this session, participants will learn how to effectively discover and subscribe to data products and assets within Amazon SageMaker Unified Studio . The session will cover navigating the data catalog, submitting subscription requests with appropriate justifications, and understanding the approval workflows that grant access to the desired data. By mastering these processes, attendees will enhance their ability to efficiently access and utilize data resources, fostering improved collaboration and data-driven decision-making within their organizations.
Restrict access to specific row/columns in a table, ensuring sensitive information (e.g., SSN, Salary) is visible only to authorized users.
In this session, participants will learn how to explore and query data stored in local or Amazon S3 location using Amazon SageMaker Data Studio. Exploration is facilitated by the compute options Amazon Athena and Amazon Redshift serverless to perform SQL queries directly within the unified interface.
In this session, participants will gain hands-on experience with JupyterLab notebooks, focusing on creating, editing, and executing code within this versatile environment. The session will guide attendees through the process of launching JupyterLab, creating new notebooks, and utilizing the "Getting Started" notebook to familiarize themselves with essential features and functionalities. By the end of the session, participants will be equipped to effectively use JupyterLab for interactive computing, data analysis, and visualization tasks.
JupyterLab
In this session, participants will learn how to utilize Amazon SDS JupyterLab notebooks to run PySpark workloads on various compute services, including Amazon EMR Serverless, Amazon EMR on EC2, and AWS Glue. Attendees will gain hands-on experience in connecting their notebooks to these services, enabling scalable data processing and analysis without the need to manage underlying infrastructure. This integration streamlines big data workflows, allowing data scientists and engineers to focus on developing insights and models efficiently
In this session, participants will learn how to create and manage Apache Iceberg tables using PySpark within a JupyterLab notebook. The session will cover configuring the Spark environment to integrate with Iceberg, establishing a catalog, and executing Data Definition Language (DDL) commands to create and manipulate Iceberg tables. Attendees will gain hands-on experience in defining schemas and performing data operations, enabling efficient data lakehouse management and analytics.
In this session, participants will learn how to utilize Amazon SDS JupyterLab notebooks to read from and write to Amazon Redshift tables. Attendees will gain hands-on experience establishing secure connections between SageMaker and Redshift, enabling seamless data ingestion, analysis, and storage. The session will cover best practices for data integration, empowering users to efficiently manage and manipulate data across these platforms.
In this session, participants will learn how to connect Amazon SDS to both first-party and third-party databases, including Snowflake and Google BigQuery. Attendees will gain hands-on experience establishing secure connections, enabling seamless data exploration and analysis within the SageMaker environment. The session will cover best practices for configuring these integrations, empowering users to efficiently manage and manipulate diverse data sources for their machine learning workflows.
In this hands-on session, participants will learn to design and implement a visual ETL (Extract, Transform, Load) workflow using Amazon SageMaker Unified Studio. The focus will be on creating a simple data pipeline that extracts data from an Amazon S3 source, applies transformations, and loads the processed data back into an Amazon S3 destination. This exercise will demonstrate the capabilities of SageMaker's visual ETL tools in streamlining data integration tasks without the need for extensive coding.
In this introductory session, participants will explore the comprehensive AI and machine learning (ML) capabilities of Amazon SDS. The session will cover the platform's unified environment for data exploration, preparation, and integration, as well as tools for big data processing, SQL analytics, and generative AI application development. Attendees will gain insights into how SageMaker streamlines the end-to-end AI/ML workflow, enabling efficient model development, training, and deployment. This foundational understanding will equip participants to leverage SageMaker's integrated features for their AI and ML projects.
In this hands-on session, participants will delve into the development of generative AI chat agents using Amazon SageMaker Unified Studio. Leveraging the Amazon Bedrock Integrated Development Environment (IDE), attendees will learn to build and customize chat agents without the need for extensive coding. The session will cover creating knowledge bases, implementing guardrails, and integrating functions to enhance chat agent capabilities, providing a comprehensive understanding of deploying generative AI solutions within SageMaker.
In this hands-on session, participants will utilize Amazon SDS JupyterLab notebooks to generate synthetic datasets and apply the Linear Learner algorithm for training and inference tasks. Attendees will gain practical experience in creating synthetic data, configuring the Linear Learner model, and executing training and inference workflows within the SageMaker SDS environment. This session will cover best practices for data generation, model training, and deployment, enabling participants to effectively develop and operationalize machine learning models using SageMaker's integrated tools.
In this hands-on session, participants will utilize Amazon SageMaker Studio's JupyterLab notebooks to generate synthetic datasets, apply the Linear Learner algorithm for training and inference, and integrate MLflow for experiment tracking and model management. Attendees will gain practical experience in creating synthetic data, configuring the Linear Learner model, and executing training and inference workflows within the SageMaker environment. Additionally, they will learn to set up and run an MLflow server to capture metrics, parameters, and models, enhancing reproducibility and collaboration in machine learning projects.
Introduction to Sagemaker Lakehouse, Catalogs and talks about interoperability and performance
Explaining AWS Glue catalog within Sagemaker Lakehouse
Managed Catalog for Interoperable or Performance
Managing data catalogs across multiple cloud platforms is challenging due to siloed metadata, inconsistent governance, and fragmented access. This solution unifies all catalogs under one roof, enabling a parent-child relationship that harmonizes metadata, streamlines access, and ensures governance consistency. By centralizing catalog management, organizations gain better data discoverability, lineage tracking, and policy enforcement across all sources.
Amazon SageMaker JumpStart is a machine learning hub that provides access to a wide range of publicly available and proprietary foundation models, built-in algorithms, and prebuilt solutions for common ML tasks1.
You can quickly deploy, fine-tune, and evaluate these models for tasks like text summarization, image generation, and fraud detection, all within an easy-to-use interface or via SDK.
JumpStart supports customizing models with your own data and sharing ML artifacts like models and notebooks across your organization to accelerate development.
All data remains private and encrypted within your AWS environment, ensuring security and compliance throughout the ML workflow
Unified Studio inference endpoints are fully managed endpoints that let you deploy machine learning models and get real-time or batch predictions without managing infrastructure.
You can choose from several deployment options, including real-time endpoints for low-latency predictions, serverless endpoints for automatic scaling, and asynchronous endpoints for large or long-running requests.
Endpoints can host single or multiple models, support autoscaling, and offer features like A/B testing, shadow deployments, and inference pipelines for combining preprocessing, prediction, and post-processing tasks.
Once deployed, you can invoke endpoints via REST API, AWS SDKs, or the SageMaker Python SDK, making it easy to integrate predictions into your applications.
Apache Airflow can be integrated to orchestrate and automate complex workflows within SageMaker Unified Studio. It allows you to schedule, monitor, and manage tasks like data preprocessing, model training, and deployment, ensuring streamlined and repeatable ML pipelines. Airflow's flexibility with DAGs (Directed Acyclic Graphs) provides robust control over dependencies and execution order.
Introduction to AWS Strands on Sagemaker Unified Studio
Watch how to build an AI-powered agent to handle restaurant bookings, showcasing live reservation capabilities and seamless workflow integration. See how AWS resources like Bedrock, Knowledgebase, and DynamoDB are used to connect agents with real-world data and booking actions.Discover the architecture and essential AWS tools that let your agent retrieve restaurant info, validate availability, and automate reservations—all in one dynamic demo
Explore the transformative capabilities of Amazon SageMaker Unified Studio in this introductory course, crafted as one of the first comprehensive learning experiences for beginners and intermediate-level enthusiasts in the fields of data and AI. This course dives into the essentials of SageMaker Unified Studio, a user-friendly tool designed to simplify building data pipelines, executing analytical workflows, and creating AI-driven products. With a focus on accessibility, it provides the foundational knowledge needed to start leveraging this platform effectively.
Participants will gain hands-on experience in implementing exploratory data analysis (EDA), data analytics, AI, and generative AI practices using SageMaker Unified Studio. Starting with administrative best practices, the course includes constructing visual ETL workflows, such as AI development pipelines, enabling participants to ETL using Spark data seamlessly and create impactful data products. Additionally, learners will explore how to manage and govern data effectively within a unified analytics environment, applying AWS best practices for real-world scenarios. With an emphasis on practical applications and evolving industry standards, this course empowers attendees to confidently embark on their data and AI journey.
Whether you are a developer, an engineering manager, or a decision-maker, this beginner-friendly course provides the foundational tools to succeed in building innovative data solutions and preparing for a career in the dynamic world of data and AI. As SageMaker Unified Studio moves from preview to general availability (GA), this course will continue to evolve, ensuring it remains aligned with the latest features and industry developments.