Udemy
    •  
    •  
    •  
    •  
    •  
    •  
    •  
    •  
Turn what you know into an opportunity and reach millions around the world.
Learn More
Your cart is empty.
Keep shopping
(Ken Cen出品)Generative AI第 31 部 多模態融合:視覺+語言模型深入解析 (上)
Rating: 5.0 out of 5(5 ratings)
81 students

(Ken Cen出品)Generative AI第 31 部 多模態融合:視覺+語言模型深入解析 (上)

關於多模態模型,Siglilp,Vision Transformer,Batch Norm,Layer Norm,注意力機制
Created byKen Cen
Last updated 10/2025
Chinese (Traditional)

What you'll learn

  • 深入瞭解什麼是對比學習
  • 深入瞭解為什麼需要SigLIP & 什麼是Vision Transformer
  • 學會如何使用 Pytorch 編寫Siglilp Vision Transformer
  • 深入瞭解什麼是 Batch Norm 和 Layer Norm
  • 學會如何編寫多模態模型的注意力機制代碼

Course content

4 sections • 9 lectures • 4h 42m total length
  • 加入Udemy 全球最大的中文 AI 課程2:16
  • 課程工具準備2:41

    學員將為課程工具準備

  • 如何使用uv 作為包管理器和項目管理工具10:53

    學員個將瞭解如何使用uv 作為包管理器和項目管理工具

Requirements

  • 一臺電腦

Description

一般 AI 模型都是只能處理某個功能。例如,語言模型 & 圖像模型。

而隨著 AI 時代的發展,多模態模型誕生了。

它能够同時處理多種不同的輸入信息。


本課程將介紹如下將 用戶的 Prompt & 圖像輸入,分別用 Vison Transformer 和 Tokenizer 捕捉當中的內容,同時,結合兩種的Embedding ,並輸入到 Gemma 模型當中處理。


課程內容如下:

  1. 什麼是對比學習 & 為什麼需要SigLIP & 什麼是Vision Transformer

  2. 如何 Pytorch 編寫Siglilp Vision Transformer

  3. 如何處理 Embedding 到Patches & 什麼是 Batch Norm 和 Layer Norm

  4. 如何編寫多模態模型的注意力機制代碼

  5. 如何處理輸入 Image 和輸入 Prompt 並合併在一起

  6. 如何製作 PaliGemma 多模態模型並導入權重為推理做準備

Who this course is for:

  • AI/機器學習工程師 (AI/ML Engineers)
  • 數據科學家 (Data Scientists)
  • 深度學習研究員 (Deep Learning Researchers)
  • 對生成式 AI 有濃厚興趣的開發者 (Developers Interested in Generative AI)
  • 學生 (Students)