Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
Modern Computer Vision AI with Vision Transformers and LLMs
New
9 students

Modern Computer Vision AI with Vision Transformers and LLMs

Master Recognition, Detection, Segmentation, Vision-Language Models, Reasoning, Video Intelligence and Generative Vision
Last updated 7/2026
English

What you'll learn

  • Master Modern Computer Vision using Vision Transformers and Large Language Models (LLMs)
  • Understand the fundamentals of Image Recognition, Object Detection, and Image Segmentation
  • Build Computer Vision applications using Python, PyTorch, and OpenCV.
  • Learn Vision Transformer (ViT) architecture and self-attention for image classification.
  • Perform real-time Object Detection using YOLO, DETR, DINO, Grounding DINO
  • Perform Image Segmentation using YOLO Segmentation and Segment Anything Model (SAM)
  • Learn Vision-Language Models using CLIP for image-text understanding
  • Build AI-powered Image Captioning and Visual Question Answering using BLIP
  • Understand Multimodal AI by combining Computer Vision with Large Language Models
  • Learn Depth Estimation and Human Pose Estimation for 3D scene understanding
  • Analyze videos using TimeSformer and Video Transformers
  • Understand Image Generation using modern Generative AI techniques. Learn Diffusion Models for AI-powered Image Editing and Inpainting.
  • Understand the latest advancements in Computer Vision, Vision AI, and Generative AI
  • Gain practical skills to build end-to-end AI applications using state-of-the-art Vision AI models

Course content

8 sections47 lectures6h 55m total length
  • Introduction8:06
  • Course Documentation2:00

Requirements

  • Basic knowledge of Python programming. Basic understanding of Deep Learning is recommended.
  • No prior Computer Vision experience is required.
  • No prior knowledge of Vision Transformers or Large Language Models (LLMs) is required.
  • Willingness to learn through hands-on coding and practical projects.
  • Curiosity to explore modern Computer Vision, Vision AI, and Multimodal AI technologies.

Description

Modern Computer Vision AI with Vision Transformers and LLMs

Modern Computer Vision has evolved far beyond traditional image classification. Today, Vision AI systems can recognize objects, detect multiple objects, segment images, understand natural language, reason about visual scenes, analyze videos, generate realistic images, and intelligently edit existing images. These capabilities are transforming industries such as healthcare, autonomous driving, robotics, manufacturing, retail, surveillance, and smart automation.

This course provides a structured learning journey through the world of Modern Computer Vision, Vision Transformers, Vision-Language Models, and Large Language Models (LLMs). Rather than learning individual models in isolation, you'll understand how they connect together to build intelligent Vision AI systems capable of solving real-world problems.

Your learning journey includes:

  • Building a strong foundation in Modern Computer Vision and Vision AI

  • Understanding how Computer Vision has evolved from CNNs to Vision Transformers and Foundation Models

  • Learning the core Computer Vision tasks:

    • Image Recognition

    • Object Detection

    • Image Segmentation

    • Vision-Language Models

    • Vision Reasoning

    • Depth & Pose Estimation

    • Video Intelligence

    • Image Generation

    • Image Editing

  • Exploring state-of-the-art Vision AI models, including:

    • ResNet-50

    • Vision Transformer (ViT)

    • YOLO

    • DETR

    • DINO

    • Grounding DINO

    • Segment Anything Model (SAM)

    • CLIP

    • BLIP

    • TimeSformer

    • Diffusion Models

  • Building practical Vision AI applications using Python, PyTorch, Hugging Face Transformers, OpenCV, Ultralytics YOLO, and Streamlit

The course follows a simple and consistent learning approach for every major model:

  1. Understand the problem the model solves

  2. Learn the core intuition behind the architecture

  3. Explore how the model works

  4. Implement the model using modern AI frameworks

  5. Apply it to real-world Computer Vision applications

This structured approach helps you develop both conceptual understanding and practical implementation skills, making advanced Computer Vision topics easier to learn, understand, and apply.

By the end of this course, you'll be able to:

  • Understand the complete Modern Computer Vision and Vision AI ecosystem

  • Explain how Vision Transformers and Large Language Models (LLMs) are transforming Computer Vision

  • Apply state-of-the-art Computer Vision and Vision AI models to real-world problems

  • Build practical AI-powered Computer Vision applications

  • Develop a strong foundation for advanced Computer Vision, Multimodal AI, and Generative AI systems

Whether you're building intelligent AI applications, exploring the latest advances in Vision AI, or expanding your expertise in Computer Vision, this course provides a clear, practical, and comprehensive path to mastering the technologies that power the next generation of AI systems.

Who this course is for:

  • Anyone looking to stay up to date with the latest advancements in Modern Computer Vision, Vision Transformers, and Large Language Models (LLMs).
  • Researchers and graduate students working in Computer Vision and AI
  • Anyone interested in Vision Transformers and Multimodal AI