
The instructor brings twenty years in product and quality engineering, leads large teams, and shares lessons from over 100 projects, staying hands-on and up-to-date with the industry.
Run a five-minute AI benchmark to check a model's bias and performance, from downloading Roberta base to tokenizing data and running an evaluation script that measures disparity for production.
Explore non-functional testing for large language models, including ethics, adversarial testing such as prompt injection and data poisoning, and evaluating human-like behavior, consistency, and robustness in conversation.
Learn how to tune foundation models by modifying data and algorithms, running pre-trained tests, retraining on new data, validating metrics, and deploying with ongoing monitoring.
Debunk the myth of a universal ai model; emphasize task-specific models, proper benchmarking, and fine-tuning with relevant data to avoid garbage in, garbage out.
Explore what intelligence means and how artificial intelligence uses machine learning, natural language processing, computer vision, and speech to perform tasks that require understanding.
Explore how natural language processing enables computers to understand and generate human language across text, voice, and video, covering natural language understanding and generation with foundation models.
Explore the four core machine learning types—supervised, unsupervised, reinforced learning, and deep learning—through neural networks, unstructured data, and the role of large language models.
Learn how supervised learning uses labeled data to train a model, test with unlabeled data, and refine through retraining to improve dog or not dog predictions.
Reinforcement learning enables an artificial intelligence to learn from feedback by observing environment state, taking actions, and earning rewards or penalties through trial and error, as in a self-driving car.
Explain what token is in large language models as a unit of text that can be characters, words, or spaces, and show how tokenization affects prompt length and token counts.
Understand the difference between Narrow Task AI vs General Purpose AI vs Artificial General Intelligence (AGI)
https://code.visualstudio.com/download
Install python on Windows via the Microsoft Store to automatically resolve dependencies, then verify the installation in the terminal and confirm python is in PATH for benchmarking tasks.
Install Python dependencies with pip by downloading get-pip.py and running it to install pip. Then use pip install to add packages like open AI for the hands on demo.
Install miniconda to create isolated python environments, activate and deactivate them, manage dependencies with conda and yaml files, and validate setup across Windows or Mac.
install node.js and npm from the installer at nodejs.org, then verify with node -v and npm -v and adjust PATH so node and npm bin are accessible.
Link to repository: danteachqe/LLMs: a comprehensive code repo for testing LLMs
Discover how to access ChatGPT via API by securing a subscription with prepaid credit, adding a credit/debit card in billing, and creating an API key.
Explore the Hugging Face community page to discover open-source models (over 1,000,000), datasets, and spaces, learn to filter and run models, and preview Transformers with evaluate.
Explore the Hugging Face Transformers library, install PyTorch and Transformers, and use pipelines to benchmark and compare pre-trained models for sentiment analysis.
Explore how to load and use any hugging face model with transformers and pipelines, run local benchmarks (BLEU, TER, GLUE), and test diffusion and text-to-speech workflows.
Learn to use the Hugging Face evaluate package to benchmark machine learning models with metrics like accuracy, precision, recall, F1, and perplexity, and to run BLEU, TER, and GLUE benchmarks.
Accuracy measures the proportion of correct predictions, calculated as correct divided by total predictions. Imbalanced data and small tests can distort results, underscoring the need for relevant, high-quality training data.
Learn how to measure precision of LLMs by identifying true positives and false positives. Explore how false positives and false negatives impact model quality, illustrated with a spam detection example.
Recall measures the proportion of true positives among actual positives, highlighting false negatives; illustrated by cancer screening and fraud detection, with a case showing 80% recall on 100 positives.
Learn to calculate perplexity for a text using PyTorch and Transformers with a GPT-2 model and its tokenizer from Hugging Face, including setup and benchmarking.
Repo ->https://github.com/danteachqe/LLMs/tree/main/LLM/Data_Splitting
Explore k-fold cross validation, splitting data into five folds with four training and one testing set, and compare basic and stratified folding to prevent bias in imbalanced data.
Repo -> https://github.com/danteachqe/LLMs/tree/main/LLM/Data_Splitting
https://gluebenchmark.com
Dateset : https://huggingface.co/datasets/nyu-mll/glue
Tasks: https://gluebenchmark.com/tasks
Learn to run a Glue benchmark by preparing an open model with a tokenizer, loading libraries, tokenizing data, fine-tuning on the training set, and submitting predictions to the Glue leaderboard.
Run and visualize a GLUE benchmark on a BERT model from Hugging Face using evaluate_glue.py, showing accuracy around 50% and the impact of training on a downstream task.
Benchmark ChatGPT against the SST-2 sentiment task using the OpenAI API, then evaluate accuracy, precision, recall, and F1 on the first 50 items to compare models.
Explore retrieval augmented generation, known as rag, and how external databases, vector embeddings, and memory extension reduce drift and hallucinations, with retriever, documentation, and generation steps.
Explore four popular rag techniques: standard rag, corrective rag, speculative rag, and graph rag. Understand how document scoring and fact checking with multi-document search improve accuracy.
Explore ragas retrieval validation framework to evaluate large language models using context precision, context recall, and relevance metrics, comparing prompt responses to ground-truth references with sample documents.
The Rag framework evaluation emphasizes fluency, relevance, coherence, and concision, using an LLM as a judge within a deep eval framework to benchmark Rag pipelines.
Analyze how training time drives cost and performance by examining data size, neural network and training hyperparameters, including epochs, learning rate, and loss functions.
Discover how memory and token limits affect chat models, explaining context, input and output tokens, and manual memory management when testing apis.
This comprehensive course delves into the essential practices, tools, and datasets for AI model benchmarking. Designed for AI practitioners, researchers, and developers, this course provides hands-on experience and practical insights into evaluating and comparing model performance across tasks like Natural Language Processing (NLP) and Computer Vision.
What You’ll Learn:
Fundamentals of Benchmarking:
Understanding AI benchmarking and its significance.
Differences between NLP and CV benchmarks.
Key metrics for effective evaluation.
Setting Up Your Environment:
Installing tools and frameworks like Hugging Face, Python, and CIFAR-10 datasets.
Building reusable benchmarking pipelines.
Working with Datasets:
Utilizing popular datasets like CIFAR-10 for Computer Vision.
Preprocessing and preparing data for NLP tasks.
Model Performance Evaluation:
Comparing performance of various AI models.
Fine-tuning and evaluating results across benchmarks.
Interpreting scores for actionable insights.
Tooling for Benchmarking:
Leveraging Hugging Face and OpenAI GPT tools.
Python-based approaches to automate benchmarking tasks.
Utilizing real-world platforms to track performance.
Advanced Benchmarking Techniques:
Multi-modal benchmarks for NLP and CV tasks.
Hands-on tutorials for improving model generalization and accuracy.
Optimization and Deployment:
Translating benchmarking results into practical AI solutions.
Ensuring robustness, scalability, and fairness in AI models.
Benchmark RAG implementations
RAGAS
Coherence
Confident AI - Deepeval
Hands-On Modules:
Implementing end-to-end benchmarking pipelines.
Exploring CIFAR-10 for image recognition tasks.
Comparing supervised, unsupervised, and fine-tuned model performance.
Leveraging industry tools for state-of-the-art benchmarking