
Introduces a refresher Python big data course with quiz-based format, covering Git, Jenkins, HDFS commands, task workflows, testing, release processes, and managing pushes, simulation runs, and golden copy comparisons.
Explore an intermediate python big data refresher via a 100-question quiz, part two, outlining goals and topics for what you will learn.
Improve productivity with this intermediate Python big data refresher, using ticket and quiz based tasks, final project assessment, and detailed wiki notes on what worked and what didn't.
Engage in quiz-based learning with over 100 questions to reach intermediate levels, covering full stack framework, back-end engines, git commands, Jenkins, remote working best practices, and Python.
Prepare to reinforce Python basics, framework testing, and Hadoop admin through quizzes, notebooks, csv samples, and full stack grid runs with debugging to validate progress.
Develop intermediate skills by debugging Git commands, fixing Jenkins pipelines, and using HDFS commands. Perform data matching with golden copies and manage release processes on JFrog for pip installation.
Compare back end and full stack runs, with back end engines using parquet or csv and full stack runs relying on hive for stability, noting input file issues and permissions.
Compare backend and full stack runs: backend uses fargate with csv inputs, while full stack uses hive tables for input, output, highlighting development versus production, schema considerations, and permission issues.
Learn how to match the golden copy across back-end and full-stack runs, diagnose data conversion and schema changes, and test inputs with local grid and YAML-driven workflows.
Explore the backend computational engine that uses YAML input and a UI front end, emphasizing proper indentation to prevent driver or submission failures; activate the virtual environment and complete authentication.
Run the backend computation engine with csv and yaml inputs, configure spark, and test the pocket output, then suffix yaml and shell scripts and package for the full stack engine.
Learn how comparison suits line-by-line compare outputs, files, rows, and schemas for Spark data frames in a Jupyter Notebook. Run from source and comparison tool folders, and manage versioned outputs.
Explore comparison tools for big data Spark data frames, clone the repository, run comparisons with specified columns and tolerance, and review HTML reports for mismatches.
Master git commands, enforce jira-style commit messages, and fix errors with soft resets or amends; build eggs per repository instructions and prefer PyCharm’s git GUI for pull, fetch, and stash.
Resolve git push failures by checking commit messages and JIRA references, using soft reset to undo commits, and using PyCharm for git tasks with correct software package numbers.
Learn Winscp to connect to driver nodes and transfer files with get and put, and use Winmerge for recursive comparisons; avoid downloading CSV from workbenches and prefer notebooks.
Explore essential big data tools on Windows: Winscp for copying between machines and drivers, HDFS get/put commands, workbench for Hive or Oracle via ODBC, and Winmerge for file comparisons.
Explore essential shell commands and hdfs operations, including diff for byte-level file comparisons, hdfs file management, and file name searching, with wiki pages as quick references.
This refresher covers releasing code to artifactory, enabling pip installs via Bitbucket releases, choosing wheel files over agg, and storing release notes and documentation in SharePoint.
Learn the release process from code packaging in Jfrog Artifactory and egg files to publishing releases via CI/CD pipelines with Jenkins and YAML, including release branches and notes.
Compare full stack and backend runs; backend developers independently build code, while full stack includes front end and data processing, with a notebook to compare outputs in hive and spark.
Match the back end or full stack with its old golden copy, then run full stack with the UI framework and release code, or back end via the Bitbucket branch.
Explore golden copy matching across full stack and backend, using yaml inputs, running the grid locally, and debugging with print logger and notebook-driven checks.
Explore how the full stack uses the Jfrog artifactory to store and deliver backend code, while the frontend fetches it, and spark submit is used for backend testing.
Learn how Jenkins builds full stack and backend projects, manage Bitbucket repos, and use Winscp and PyCharm git workflows to deploy and maintain code.
This refresher course revisits Python for big data at an intermediate level with quiz-based tickets, covering Git, Jenkins, HDFS, full stack and backend matching, and release workflows.
Explore a refresher quiz on git, hdfs, and testing, focusing on data matching across back-end and full-stack pipelines with golden copies.
Refresher Python Big Data Intermediate Quiz based Course 102
Learn git, jenkins, golden copy matching, testing at intermediate levels with 100 quiz questions
Exploring the below:
Advanced Git commands
Advanced Jenkins commands
Advanced HDFS commands
How to write task and workflows
Advanced Testing
Release Process
Content:
Once you know basics of python computational framework, testing, hadoop admin then we can jump to more advanced topics in this course.
This course starts with quizzes to make you more advanced developer for the role.
We generally start with very simple notebook with sample csv data but at the end we do the full stack grid runs and save data in hive using a fairly complex code.
Passing and re-taking quiz is very important
Write up based questions will be evaluated by me to gauge your progress.
Remote working requires extensive training and this course will help progress in your remote role that has minimum supervision and guidance.
Enables independent problem solving my learning about different issues we might get into. This is important to work with minimum supervision.
The course is not coding heavy but prepare you with the issues you will face.
The course focus more on debugging why things are not working.