
understand how sound, as a longitudinal wave, travels through a medium, how vocal cords produce it, and how the vocal tract shapes speech, with fundamentals and harmonics.
Explore how sound is a wave and amplitude measures pressure variation. Note amplitude relates to decibels from 30 dB to 80 dB, with vowels higher than consonants.
Explore how frequency shapes voice pitch, contrasting vowels and consonants, and explain how time-domain waveforms differ from frequency-domain spectra, with practical FFT window trade-offs.
Represent and visualize sound through time-domain waveforms and frequency-domain spectra, explaining amplitude and loudness, plosives and vowels, and silence; discuss fast Fourier transform windows, granularity, pitch, formants, and harmonics.
Explore spectrograms as visual representations of sound, with time on the x-axis, frequency on the y-axis, and amplitude by color, highlighting wideband and narrowband analyses for ASR and DTS.
Explore how raw audio becomes discrete tokens via spectrograms and MFCCs, enabling asr and tts with CNNs and transformers, and self-attention for context, while highlighting noise, speaker variability, and accents.
Navigate phase loss, aliasing, and quantization that degrade whispered speech. Identify error drivers, including noise, accents, emotion, and speed, and address discrete unit learning and robustness for capturing prosody.
Explore how SpeechLM tackles challenges with self-supervised learning and unified modelling. Leverage masked audio modelling, speech-text alignment, and phoneme-subword hybrids to improve performance and robustness.
Learn how sound, a wave with amplitude and frequency, is represented by time-based and frequency-based methods and spectrograms, enabling deep learning and self-supervised, unified speech models.
Record live audio, render waveform and spectrogram, transcribe with whisper, and play back the result, illustrating code-driven steps with SciPy and Librosa in module 2.1.
Explore the source-filter model of speech production, detailing how source and filter shape speech output and phoneme distinctions. Understand its relevance to speech processing and practical challenges.
Explore the source-filter model of speech production, linking articulation to vocal cords and vocal tract. Learn how speech equals source plus filter, with implications for synthesis, recognition, and analysis.
Explore the source filter model’s three source types—voice sounds, unvoiced sounds, and mixed sounds—and their processing: pitch and formants for voiced, spectral energy for unvoiced, and adaptive hybrids for mixed.
Explore the vocal tract as the filter that shapes sound from the vocal cords, via jaw, tongue, lips, and vellum, producing formant frequencies f1, f2, and f3.
Learn how speech is produced by the lungs, vocal cords, and vocal tract through instrument analogies, where the lungs power, cords generate, the tract shapes, and lip rounding creates sounds.
Explore key concepts of speech: voiced sounds, about 65% of speech and periodic, and unvoiced sounds, about 30% and aperiodic; formants and spectrum peaks drive voice synthesis, recognition, and analysis.
Explore the source-filter model's role in speech processing, detailing formant synthesis with f1, f2, f3, mfcc analysis separating source and filter, and lpc coding for articulation, with phones using tts.
Explore challenges in generative voice AI, including filter complexity, source issues, and output artifacts, and learn how deep learning, WaveNet, and voice conversion GANs address articulation, coupling, and variability.
Explore how vocal cords generate voice via excitation and vibration, how the vocal tract shapes sound into vowels and consonants, and how speech language models tackle challenges in speech processing.
Explore the building blocks of speech by examining phones, phonemes, and allophones; identify how phones are any sound, phonemes define language-specific sound sets, and allophones are phoneme variants.
Explore phonetics as the science of speech sounds, covering the articulatory, acoustic, and auditory branches, plus the source-filter model and IPA rules with their challenges.
Explore phonetic features as the dna of speech, including voicing and voiceless contrasts, place and manner of articulation, and their impact on ASR and word error rate.
Explore phonology, the study of phonemes and rules that organize sound in language, covering phonotactics, prosody, and morphonology, and their impact on tts and asr.
Map sounds to phonemes and phonetic features using acoustic cues and spectrograms, from HMM/GMM to neural nets, with lexicon-based word mapping.
Explore how phonemes link audio to language models in deep learning. Delve into the text-to-speech pipeline, co-articulation challenges, dialect variations, low-resource languages, and multilingual phoneme embeddings.
Explore phonetic and phonology challenges in generative voice AI, including variability, ambiguity, rare phonemes, and prosody, and review deep-learning hurdles like sound-to-token alignment with self-supervised and adversarial accent adaptation solutions.
Explore a hands-on workflow in phonetics and phonology: record audio, save as wav, perform automatic speech recognition, map words to phonemes and phonetic features, and plot spectrogram with CMU dict.
Explore why audio feature extraction matters for speech and sound samples, review traditional methods (waveform, spectrum, spectrogram, MFCC) and modern approaches, compare them, and discuss universal and method-specific challenges.
Explore traditional MFCC feature extraction and its seven steps, including pre-emphasis, framing, and mel-filtering, and contrast it with end-to-end, self-supervised speech models shaping noisy ASR and emotion detection.
Explore traditional MFCC features and the shift to raw-audio end-to-end learning, converting waveforms into tokens for transformers, with WaveNet and SyncNet as pioneers and gains in speech and music tasks.
Explore modern speech-language models from raw audio to cnn encoder, discrete tokens, and transformer with self-supervised training, featuring Wave2vec2, Hubert, Whisper, and mask prediction plus M4T and Valley deployment.
Compare traditional speech pipelines with modern speech-language models like Wave2, Vect2, and Whisper. Highlight the shift from MFCC and spectrogram inputs to raw waveforms and end-to-end learning.
Navigate universal and modern speech-language model challenges, including ambiguity, word error rate in noisy accents, and information loss, plus compute costs and interpretability hurdles.
Compare traditional MFCC-based acoustic feature extraction with modern raw-audio methods, analyze challenges, and highlight the 2020-to-2024 shift toward speech language models, including a noise-focused coding demo.
Examine how white noise affects audio representations such as MFCC and Mell's spectrogram, and learn to compute, visualize, and transcribe noisy versus clean audio with Librosa, Plotly, and Whisper.
Explore the paradigm shift from traditional cascaded acoustic models to single speech LMs, detailing discrete tokenization, end-to-end generation, and the rise of audio LM tech like Volley.
Explore the shift from cascaded TTS to end-to-end speech LMs using token-based generation. See how residual vector quantization and encodec enable zero-shot voice cloning via tokens with autoregressive transformers.
The objective in tts is to map text x to audio y, capturing emotion and speaker identity, while cascaded mel spectrogram workflows contrast with end-to-end speech language models.
Explore the speech lm paradigm that converts text to discrete acoustic tokens, trains end-to-end with a transformer decoder, and disentangles content, prosody, and speaker identity for scalable, expressive tts.
Convert 44 kHz raw audio to discrete tokens via vector quantization, reducing to 50–75 Hz token rates for transformer processing, enabling end-to-end autoregressive token prediction conditioned on text.
Analyze the cascade architecture bottlenecks, including mean square error driven averaging, entanglement of voice identity with content, and missing phase, and learn how SpeechLM’s discrete token space improves expressivity.
SpeechLM moves from continuous mel space to discrete tokens via a neural codec, enabling transformer-based autoregressive generation and cross-entropy training for more expressive, varied voice outputs.
Contrasts cascade regression with speech LM training, where cross entropy prevents mean averaging, reducing robot voice and enabling expressive, end-to-end transformer based voices like AudioLM, Volley, and SpeechGPT.
Explain the four-step cascade tts pipeline from text to audio, using linguistic encoding to phonemes with bidirectional LSTM and attention, then Tacotron 2 to spectrograms and WaveGlow to audio.
Explore SpeechLM's four-step workflow for speech generation: tokenization with a neural audio codec, contextual prompting, autoregressive token prediction with a transformer, and decoding to raw audio in an end-to-end model.
Explore the classic cascade Tacotron 2 and WaveGlow pipeline, converting text to mel spectrograms and to audio via a vocoder, using bi-LSTM location-sensitive attention and addressing flat prosody.
Contrast the cascaded mel spectrogram approach with acoustic model and vocoder. Explore the token-space speech language model era enabling zero-shot voice cloning from a short audio prompt.
Explore quantization noise trade-offs in speech language models and the sequence length explosion, highlighting codebooks and vocabulary sizes from 512 to 4096, and their impact on compute and fidelity.
Trace the evolution of speech synthesis from canned, pre-recorded words to neural end-to-end transformer models, and explore the discrete versus continuous trade-off shaping future unified token-space architectures.
Explore three future directions for bridging discrete and continuous in voice ai: latent diffusion for speech, unified text-audio tokens for a multi-model transformer, and efficient state-space architectures for longer audio.
Compare cascade tts with speech lm tts, detailing SpeechT5 mel spectrograms and HiFiGAN vocoder. Contrast encoder-based token space in speech lm with mel-based cascade, using 24 kHz encodec and tokens.
Explore how self-supervised learning enables audio representation through wav2vec 2.0 and HuBERT, leveraging pretext tasks like contrastive learning and mask prediction to derive semantic tokens.
Explore why self-supervised learning unlocks rich audio representations from unlabeled speech with pretext tasks, enabling fine-tuning with only a few labeled hours through models like wave2wake2 and Hubert.
Contrast resource-rich and resource-poor languages and explain why languages lack labeled data. Leverage self-supervised learning to reduce reliance on labeled data and enable fine tuning with little data.
Explore how self-supervised learning reformulates speech tasks into pretext tasks, using masking in latent space to compress raw audio and predict hidden units with transformer models for efficient ASR.
Explains pretext tasks in self-supervised speech learning, focusing on mask prediction and contrastive learning, their impact on phonetic representations, temporal dependencies, and speaker invariance.
Explore how contrastive learning trains self-supervised audio models by distinguishing positives from distractors using infoNCE loss, maximizing similarity to masked inputs and minimizing negatives.
Learn how masked prediction trains generative voice ai by predicting masked audio frames from context, capturing phonetics, prosody, and speaker consistency, aided by offline clustering with MFCC features using k-means.
Turn continuous audio into discrete targets with pseudo labeling by clustering with k-means or Gaussian mixture models, generating pseudo labels that resemble phoneme identities. Leverage self-supervised learning pipelines.
Describe a four-step audio pipeline turning raw waveforms into representations. Show how CNN feature extraction, masking, bidirectional transformers, and contrastive or predictive losses capture phonetic understanding, prosody, and speaker identity.
Compress raw waveform with a CNN encoder into latent tokens, apply masking, train a bidirectional transformer to predict masked tokens, and optimize with task-specific losses like contrastive or cross-entropy.
Pass raw audio through a CNN encoder to form a latent space. Quantize with online discrete code books via differentiable Gumbel softmax, then train with masking, bidirectional transformer, and infoNCE.
Compare Wave2-Vec2 with HuBERT: online versus offline clustering and different loss functions—contrastive using InfoNC versus cross entropy—and note how HuBERT’s iterative refinement of pseudolabels improves phonetic alignment, highlighting SSL challenges.
Explore three SSL challenges in audio: information collapse, acoustic entanglement, and domain shift; contrastive and info-nc losses help mitigate collapse.
Explore the evolution of audio representation from handcrafted MFCC features to bidirectional self-supervised speech models like Wave2Vec2, Hubert, and WaveLM, and the shift from unidirectional to bidirectional SSL.
Compare HuBERT and wav2vec 2.0 SSL approaches, implement masking with pseudo labels from mel spectrograms, and apply contrastive learning using true future and distractor frames with cosine similarity.
Convert raw speech to discrete semantic tokens for transformer models using self-supervised learning with hubert or wav2vec2, followed by quantization, enabling zero-shot voice cloning and token-based ASR and TTS.
Explore how vector quantization maps continuous latent space from an SSL encoder into k means clusters, using centroids to maximize intra-cluster homogeneity and inter-cluster heterogeneity for robust semantic tokens.
Apply k-means clustering to discretize continuous latents into an audio dictionary of centroids, mapping new latents to the nearest centroid by Euclidean distance and assigning discrete IDs.
The lecture explains how models vary cluster counts and create static centroids, and offers solutions such as Gumbel Softmax and iterative centroid updates for latent space.
Contrast semantic and acoustic tokens for voice synthesis, showing semantic tokens encode meaning via Hubert-derived latent spaces, while acoustic tokens preserve timbre, emotion, and prosody with residual vector quantization.
Convert raw audio to a 50 samples per second latent space of 768-dim vectors and quantize with k-means or gumbel softmax to map to the nearest centroid.
Iteratively update the centroid with new vectors, map unseen audio to centroids for semantic tokenization, and optionally collapse IDs to reduce compute by up to 80% under the manifold hypothesis.
Explore how wav2vec 2.0 compresses raw audio into continuous latents and uses Gumbel Softmax to create discrete codes for online quantization, masking, and InfoNCE training.
Explore HuBERT's offline clustering for semantic tokenization, using MFCC or SSL features, k-means to create pseudo labels, and cross-entropy training that iteratively updates cluster IDs.
Trace the evolution from gmm-hmm to self-supervised and generative speech models, highlighting discrete token spaces and zero-shot voice cloning with AudioLM and Volley.
Advance end-to-end online codebooks using differentiable vector quantization and RVQ to learn discrete tokens, and align text and audio embeddings for cross-modal, zero-shot, unified generation.
Explore neural audio codecs and acoustic tokenization, from raw waveforms to EnCodec, SoundStream, DAC, and RVQ, covering prosody, emotion, and speaker identity, with theory, history, and future directions.
Explain how acoustic tokens compress 48 kHz audio into 50–100 discrete tokens via neural codecs, using CNN/transformer encoders and a neural decoder to produce high-fidelity, clonable speech.
Move from high sample rates to a latent space, quantize latents with discrete codebooks using l1 distance, and decode integer IDs into audio, highlighting neural audio codec trade-offs.
Explain RVQ by converting raw audio to latent space, then tokenize into hierarchical code books of discrete tokens, moving from coarse structure to fine sound details before reconstructing the waveform.
Explore how the latent distribution Z is quantized through multiple codebooks using nearest-neighbor selection. Apply residual updates across codebooks to reconstruct the original vector via residual vector quantization.
Compresses a raw audio sample from 8 samples per second to a 50 hertz latent representation using convolution, kernel, and stride, forming a feature map with filters and dot products.
Explains the hierarchical property of RVQ, showing coarse-to-fine coding across C1 to C4 from melody and vowel identity to high-frequency details, with discrete tokens for autoregressive generation.
RVQ uses multiple codebooks, typically four, reducing search complexity and boosting fidelity, with an encoder producing acoustic tokens decoded back to audio by DAC, Encodex, and Soundstream.
Explore Google Soundstream, the first neural audio codec with end-to-end training of the encoder, RVQ quantizer, and decoder, featuring quantizer dropout for variable bitrate.
The Descript DAC introduces snake activation to improve high-frequency audio, replaces ReLU, uses factorized codebooks and L2 normalization to prevent collapse, with adaptive RVQ layers for balanced bitrate and fidelity.
The lecture explores two theoretical challenges of rvq-based neural codecs for generative voice, the autoregressive burden and the hierarchical dependency, highlighting token counts and GPU compute limits.
Explore discrete tokenization via residual vector quantization and continuous flow tokenization for audio, using encodec 24kHz, multiple code books, and velocity-based flow with ODEs to evaluate reconstruction quality.
Explore neural audio codecs, including Sound Stream, Encodec by Meta, and DAC, and examine vector quantizers, residual quantizers, and codebooks for variable bit rate and audio compression.
Decode and vocode acoustic tokens back to audio using neural decoders and transposed convolution. Compare token-space decoding with mel-spectrogram pipelines, and explore models like HiFiGAN, Encodec, and DAC.
Explore how RVQ indices from discrete acoustic tokens are decoded by neural codecs into high-fidelity audio, contrasting traditional vocoders with modern token-based speech LMs for streaming and zero-shot cloning.
Explain transport convolution for neural upsampling from 50 hertz latent frames to 44 kilohertz audio, using transposed convolutions with stride and padding, and examine phase reconstruction in speech lm-based approaches.
Explore the phase reconstruction problem, contrasting old mel-spectrogram pipelines that discard phase with Griffin-Lim and artifacts, to RVQ token–based speech LM decoders that preserve phase for clearer waveforms.
map rvq tokens to 768-dimensional vectors via embedding lookup, sum codebook vectors to form the continuous latent sequence, and apply transpose convolution guided by adversarial losses to generate audio.
Explore how temporal dilation and upsampling convert latent space to waveform, and how dilated residual blocks and MPD and MSD adversarial losses preserve long-range acoustic structure in neural audio decoding.
Explore the legacy HiFiGAN vocoder that converts mel spectrograms from Tachotron 2 into crisp raw audio. Examine multi receptive field fusion with varied kernels, phase learning, and input bottlenecks.
Learn how Encodec's decoder uses token-based upsampling via transposed convolution and LSTM layers to achieve high temporal coherence and artifact-free audio, preserving long-range dependencies.
Explore how the Descript Audio Codec DAC uses snake activation to enhance high-frequency details and temporal coherence, enabling real-time generation without LSTM.
Trace the evolution from parametric vocoders and Griffin-Lim to token-based models like SoundStream, Encodec, and DEC, highlighting GAN training and transport convolution upsampling as future directions.
Survey diffusion-based neural decoders as a promising alternative to GANs for voice generation, highlighting latent diffusion and flow matching for stable, artifact-minimized, token-to-audio synthesis.
Translate RVQ tokens back to audio by encoding tokens into a latent space, upsampling from 75 fps to 24,000, then decoding with MPD in GAN training to produce the waveform.
Compare flattening and interleaving RVQ codebooks for multi-stream tokens, converting 2D caustic token grids into 1D sequences for transformers, with notes on models like audio LM, Volley, and MusicGen.
Explore multistream token strategies for RVQ-based discrete tokenization, comparing flattening, interleaving, and autoregressive versus non-autoregressive codebook approaches.
Explore three strategies for 2D token grids—flattening, multistage, and interleaving—covering sequence lengths T and K, AR vs NAR, and masking-guided single-model interleaving.
Describe converting 2D audio tokens, organized as frames and layers, into a 1D sequence for transformers using interleaving, improving attention efficiency from O(n^2) to O(T+K).
Flatten the rvq grid with raster scan to serialize frames for autoregressive token prediction in a transformer. Assess autoregressive vs non-autoregressive generation, causal masking, and search strategies.
Explore the raster scan flattening approach for token serialization in generative voice AI, noting its simplicity and transformer compatibility that preserves full dependencies, and its memory and autoregressive drawbacks.
Explore the delayed interleaving pattern for token generation, using a 2d time-frame matrix and code books to generate tokens diagonally, highlighting the t plus k minus 1 steps.
Explore how the modern Speech LM uses a delayed pattern with a custom causal masking to reduce steps from t to t+k, improving speed and efficiency over flattening.
Flatten sequences with a raster scan, then apply masking innovations like a delayed pattern to feed the transformer; compare flattening and interleaving for sequence length and steps, noting faster parallelism.
Explore audio lm, a pioneer in multi-stream tokenization that separates semantic tokens from coarse and fine acoustic tokens. See how a three-stage hierarchy with three transformers generates semantic, coarse, and fine details, and why flattening, training complexity, and non end-to-end training shape its limits.
Explore music gen’s interleaved delayed pattern that replaces brute-force flattening with diagonal generation, lowering steps from t×k to t+k-1 and enabling low-compute, autoregressive and conditional independence considerations.
Compare brute-force raster-scan flattening of 2d rvq sequences with a delayed pattern masking approach that preserves causality while reducing compute from t times k to t plus k minus 1.
Discover autoregressive decoding for audio, covering sampling methods like greedy, top-k, top-p, temperature scaling, and classifier-free guidance, from tokenization to conditional and unconditional latents.
Explore autoregressive decoding for speech synthesis by turning transformer logits into probability distributions via softmax, and compare greedy, temperature, top-k, top-p, and CFG to balance phonetic accuracy with acoustic diversity.
Explain top k sampling by selecting the largest softmax probabilities from a small vocabulary, showing stability vs diversity, and note a range around 40–100 before top p is discussed next.
Apply top-p sampling with a cumulative probability threshold to select acoustic tokens, compare with top-k, and tailor settings for zero-shot voice cloning, expressive narrative, or conservative outputs.
Explore the end-to-end speech decoding cycle from history tokens to autoregressive token generation, using CFG, conditional/unconditional runs, Z logits, softmax, temperature, top P, and stochastic sampling to shape prosody.
Traces the evolution of decoding from beam search to dynamic sampling and cfg-based generation, highlighting bland repetition, discrete logits, and the static w choice.
Explore speculative decoding with a draft model and a larger verifier to reduce CFG latency, and develop adaptive temperature for phoneme boundaries, verbal nuclei, and silences.
Explore autoregressive decoding strategies by coding temperature scaling, top-k, and top-p sampling, implementing their functions, visualizing results, and generating expressive voices with bark tts, comparing low and high temperatures.
Explore speaker conditioning, emotional control, and prosody injection for zero-shot voice cloning with reference audio and text prompts, using classifier-free guidance and models like Volley, Spear TTS, and GST.
Explore how speaker conditioning, emotional control, and prosody injection create natural, expressive voice from text or cloned audio using reference audio and emotion labels.
Learn how transformers autoregressively generate token distributions conditioned on text, speaker, emotion, and prosody. Explore disentanglement with orthogonal subspaces and continuous vs discrete speaker embeddings.
Compare continuous and discrete speaker embeddings, using RVQ indices for discrete tokens to enable zero short cloning with end-to-end training and preserved timbre, emotion, and prosody.
Explore the first three steps of speaker and emotion conditioning for natural sounding cloned TTS, including text sequence inputs, a reference clip, text tokenizer, discrete tokens, and acoustic prompt prepending.
See how transformer self-attention and cross-attention condition autoregressive generation on acoustic tokens and prompts to preserve speaker identity, emotion, and prosody across audio, with prompts guiding voice to remain cloned.
Discover how the Volley model achieves zero-shot voice cloning from a three-second reference. It uses discrete acoustic tokens with cross-attention to text prompts and requires no fine-tuning or speaker encoder.
Spear TTS separates semantics from acoustics to prevent speaker leakage, using a two-stage, disentangled approach with semantic and acoustic models, improving pronunciation, robustness, and control.
GST legacy uses a bank of 512 embeddings, each 256 dimensions, to encode emotion and prosody from a reference audio without labeling. It uses cross-attention to enable smoother emotion interpolation.
The lecture outlines challenges in expressive speech control—timbre-prosody entanglement, acoustic leakage, and out of distribution prosody—and presents directions like disentangled code books, environment invariant conditioning, adaptive guidance, and contrastive decoding.
Trace the evolution of voice conditioning from 2000s statistical parametric methods and two-model cascades to today's discrete-token generative speech LLMs, enabling zero-shot voice cloning with Volley, AudioLM, SpearTTS.
Explore factorized codebooks to disentangle semantic content, timber, and emotion in RVQ-based tokenization, enabling targeted voice cloning and independent control of emotion and environment through unsupervised discovery.
Explore speaker conditioning through discrete volley token-based generation and continuous GST style tokens, comparing bi-directional and causal attention, and mastering emotion control and prosody injection in agentic TTS.
Extend LLM vocabularies to accommodate acoustic tokens by unifying text and audio tokenization. Compare semantic tokens with acoustic tokens and build a joint vocabulary space with cross-modal self-attention.
Explore expanding llm vocabularies with semantic and acoustic tokens, including shared embeddings and separate ranges, to enable unified ASR and TTS.
Explore how transformer-based speech-language models like AudioLM, Volley, SpeechGPT, and SpiritLM evolve from separate text and acoustic vocabularies to a unified shared embedding space for cross-modal reasoning.
Explore semantic versus acoustic audio tokens, showing semantic tokens at about 50 samples per second with one token per frame, while acoustic tokens use 4–8 per frame at higher rates.
Unify text and audio tokens by encoding text with byte pair encoding and processing audio through a CNN encoder and residual vector quantizer to form joint token space for transformers.
Expands the model vocabulary to include audio tokens alongside text tokens, preserving a fixed dimensional space, enabling cross-model interleaving and auto regression switching for ASR and TTS tasks.
This lecture contrasts cascaded systems with ASR, LLM, and TTS to unified speech language models that operate in a shared token space, preserving tone and prosody.
Discover how Audiopalm enables speech-to-speech translation with a joint tokenizer, a frozen palm LLM, trainable audio embeddings, and projection adapters for end-to-end audio output.
Leverage a unified model with joint text and acoustic token space for speech to text, text to speech, and zero short voice cloning, reducing error propagation versus cascaded systems.
Examine the theoretical challenges of joint text-audio tokenization, including modality mismatch, quadratic attention limits for long audio, and codebook collapse, with remedies like k-means and Gumbel softmax.
Trace the evolution from hidden Markov models and Gaussian mixture models to deep learning and speech language models with a unified token space enabling audio generation and zero-shot voice cloning.
Discuss the shift from discrete tokenization to continuous latent space for audio, using diffusion or flow models to reduce quantization noise and approach fully continuous multimodal generation.
Expand the LLM vocabulary by adding 1024 acoustic tokens and four control tokens, apply cross-modal interleaving with GPT-2, train embeddings and LM head via transfer learning, illustrated with AudioPalm.
Maximize mutual information to align audio and text in a shared latent space, using the MI formulation and info NCE with text and speech encoders producing embeddings.
Apply multi-layer vector quantization to continuous speech, using codebooks and centroids to produce discrete tokens and residuals across layers, enabling autoregressive and contrastive learning in speech.
Explore connectionist temporal classification (CTC) for text–audio alignment, using blank tokens and monotonic order to sum path probabilities and train ASR models with RNNs or transformer architectures.
Apply infoNCE to maximize similarity between positive speech-text pairs and push apart negatives. Combine this with joint mask prediction to recover masked tokens from context, enabling global and local understanding.
Compare early fusion and late fusion strategies for cross-modal alignment of text and speech, highlighting token-level, fine-grained temporal alignment and the trade-offs in compute cost.
Compare AudioLM's two-stage process with semantic tokens and Volley's single-stage approach using delayed attention to generate acoustic tokens and final audio, guided by a joint text-audio vocabulary.
Explore continuous flow matching and diffusion models that generate speech from Gaussian noise via ode solvers, guided by text embeddings, and outperform discrete token autoregression with higher fidelity.
Examine the modality gap and information asymmetry in generative voice ai, and address entanglement of identity and semantics to enable disentangled, controllable manipulation of speaker, emotion, and noise.
Trace the four generations of cross-modal speech alignment from cascaded HMMs to end-to-end models, and then to bidirectional SSL with joint pre-training. Look ahead to diffusion-based approaches and frontier research.
Explore paralinguistics—laughter, breathing, disfluency, and affect—and how to embed them in TTS with tokens and embeddings, addressing disentanglement and control.
Explore the theoretical foundations of paralinguistic modeling by introducing a latent variable Z that captures emotion, breathing, and style, enabling one-to-many TTS generation and addressing ill-posedness.
Explore orthogonal latent subspaces that disentangle content, speaker identity, and paralinguistics such as prosody and emotion in TTS. Prevent entanglement to avoid regression to the mean and enable deeper controllability.
Explore capturing paralinguistics like laughter and breath from continuous audio with RVQ vector quantization, creating tokens for a unified speech LM to generate natural, human-like voice.
Explore joint autoregressive training for interleaved text and paralinguistic tokens, enabling affect-aware generation with CFG guidance through unconditional and conditional decoding to produce emotion-aware speech tokens.
Explain extracting paralinguistics from raw speech into a latent space and applying residual vector quantization across layers. Then interleave paralinguistic tokens with text tokens for joint autoregressive training in SpeechLMs.
Explore the ELBO method and the re-parameterization trick within hierarchical VAE, detailing encoder-decoder structure, multi-scale latent hierarchies, KL divergence, and reconstruction terms.
Examine continuous flow matching and diffusion for paralinguistic audio generation, replacing discrete tokens with a continuous latent space guided by time-dependent velocity fields to reduce noise and improve controllability.
Explore the entanglement of paralinguistic and speaker identity, the long-tail distribution of rare cues, and the lack of objective metrics, and seek contrastive disentanglement, data augmentation, and learned discriminator scoring.
Trace the evolution of paralinguistic modeling from rule-based and HMM systems to neural, joint-generation with discrete tokens, and the shift toward continuous latent representations in generation.
Model microprosody and macroprosody to shape paralinguistics in speech, inject breathing and disfluencies using entropy thresholds for natural voice generation.
Explore DDPM, score-based, and latent diffusion models for speech generation, from forward and reverse processes to latent-space diffusion and latency reduction in mel-spectrograms.
Explain probabilistic diffusion mechanics in stable diffusion models by contrasting the forward noise-adding process from image to gaussian noise with the reverse denoising path back to the original image.
Explain how forward diffusion adds Gaussian noise to x0 to xt and how the reverse process removes it from xt to x0, using alpha_t, beta_t, and the learned noise prediction.
Describe the forward diffusion process by adding noise from x0 to xt using the fixed distribution q, gaussian noise, and a beta_t scheduler, highlighting the Markov chain nature.
Explore the forward diffusion from original image to pure noise across 1000 steps, with beta increasing from 0.001 to 0.02, and the model learns the reverse by predicting epsilon theta.
Explain the reparameterization trick by modeling p_theta of x_t given x_{t-1} with mu and sigma, not exact distributions, using a fixed forward Markov process with a variance scheduler.
Explore the reverse diffusion process that denoises from pure gaussian noise at step T to a clear image by predicting x_{t-1} from x_t in a Markov chain.
Explore the U-Net architecture as the learning engine, featuring down blocks, bottleneck, up sampling with transpose convolutions, and skip connections to preserve spatial information, enabling noise prediction across steps.
Demonstrates training a diffusion model by predicting noise in the reverse process using a unit model, minimizing the MSC loss, and generating images by denoising gaussian noise.
Demonstrate forward noising and reverse denoising of a synthetic image using diffusion, alpha-beta schedules, and noise prediction via the reparameterization trick, and compare linear, cosine, and sigmoid schedulers.
Explore a practical diffusion-based coding exercise that uses a lightweight unit model with BEFT and diffusers to simulate forward noising and reverse denoising, visualize steps, and compare VRAM requirements.
Analyze how stable diffusion splits generation into forward noise addition and reverse denoising, and explain the compute bottleneck and latent diffusion with conditioning for query-based image generation.
Assess the latency challenge in generative diffusion, where a single A100 GPU yields 100 seconds per image and real-time requires 0.1 second. Propose latent-space compression to bypass pixel-space limits.
Compress images into a latent space with a variational autoencoder to solve the compute bottleneck, then train latent diffusion models on latent representations before decoding.
Explain spatial compression by reducing height and width while increasing channels, achieving 48x compression, and describe how vae emphasizes high-level representations at the cost of low-level noise.
Train uses a pre-trained vae to map images to 64×64×4 latent representations. Then diffusion operates in this latent space, and a separate vae decoder recovers the image.
Discover how latent diffusion models (LDM) dramatically reduce compute and latency compared to pixel diffusion, using 20–50 denoising steps and consumer GPUs for accessible, fast generation.
Convert the text prompt into embeddings using a clip encoder to create a shared text–image space. Steer the diffusion-based generation with cross-modal attention between image queries and text embeddings.
Explore cross attention integration in the U-net, applying text embeddings at the bottleneck and upsampling blocks while avoiding downsampling, and compare compute needs of pixel-based and latent diffusion.
Learn how latent diffusion reduces compute and memory by mapping 1024x1024 images to 64x64 latent space, with forward and reverse denoising in a vae and sd 1.5 workflow.
convert input text to phonemes via phonemizer and transformer encoder, build a conditioning matrix, then diffuse from Gaussian noise to sculpt mel spectrograms for text-to-speech.
Explore iterative denoising in acoustic diffusion, from Gaussian noise to a mel spectrogram via a diffusion transformer and text conditioning, then convert to audio with a vocoder like HiFiGAN.
Understand denoising diffusion probabilistic models (DDPMs) and their forward–reverse steps, with a scheduler adding noise to mel spectrograms and the model predicting epsilon for denoising, acknowledging slow inference.
Explore latent diffusion models for speech, encoding audio into a latent space with a VAE to speed up denoising, enabling fast, high-quality zero-shot voice cloning and pitch control.
Investigate the diffusion-based speech AI challenges—sampling speed versus quality, audio-native architecture for male spectrograms, and error accumulation in multi-step denoising that impacts pitch—and their impact on real-time, high-fidelity speech.
Explore flow matching and rectified flows to enable linear denoising with fewer steps, and investigate discrete diffusion with diffusion transformers for token-space, fast audio generation.
Analyze diffusion based voice architectures from DDPM to latent diffusion, and learn how audio LDM uses clap, unit, and VAE to convert MEL spectrograms to audio.
Explore bridging discrete LLM semantic tokens with continuous acoustic diffusion to condition diffusion models for TTS, using cross-modal attention, theoretical formulation, and discrete-to-continuous conditioning.
This lecture explains hybrid speech synthesis by merging discrete LLM semantic tokens with continuous diffusion-based acoustic modeling to achieve high-fidelity, well-structured long-form speech with natural prosody.
Bridge discrete llm tokens with continuous diffusion to transform text into acoustic outputs using autoregressive tokens, semantic latent spaces, and a denoising diffusion process.
Explore how semantic tokens reduce speech data from 44 kHz to 50 Hz, preserving phonetic content and macro prosody with pacing rhythm, while discarding noise and guiding diffusion for speech.
Compress audio to a latent acoustic manifold with an acoustic encoder, apply diffusion in the latent space guided by semantic tokens, and decode back to high fidelity speech.
The lecture explains cross attention guiding denoising with a softmax weight on V, using Q from the noisy latent and K from semantic tokens, linking tokens to diffusion.
Outline a three-stage pipeline that converts text to semantic tokens with an LLM, performs temporal alignment and upsampling to the acoustic latent space, and applies diffusion-based conditioning for waveform generation.
Demonstrate bridging discrete LLM tokens with a diffusion model: produce a denoise latent via cross-attention and reverse diffusion, then upsample with a Hi-Fi GAN decoder to generate the final waveform.
Explore Audio LDM 2, a latent diffusion hybrid that pairs a T5/BART LLM semantic tokenizer with Audio MAE to derive phonetic semantics, enabling zero-shot, high-fidelity audio generation via hard conditioning.
Explore natural speech 2 and 3's disentangled factorized latents, with pitch, duration, speaker, and semantics, plus phoneme and semantic tokenization guiding continuous diffusion for high-fidelity speech.
Trace the evolution from hidden Markov models to neural cascades and continuous mel spectrograms. Explore the shift to a hybrid diffusion approach guided by LLM tokens for high-fidelity, coherent audio.
Explore future directions beyond cascaded hybrids: joint probability flow in a shared latent space, discrete-token diffusion with CTMCs, and a unified text-audio manifold via flow matching for end-to-end generation.
Bridge discrete semantic tokens to continuous acoustic diffusion by using a length regulator, a diffusion model predicting mel spectrogram noise, and a vocoder to synthesize output wav.
Explore the shift from discrete to continuous time steps in diffusion models, from DDPM to rectified flow, and learn how SDEs and ODEs enable efficient speech synthesis.
Explore continuous time diffusion and ode solvers, compare discrete ddpm steps with a continuous path, and learn how adaptive steps enable faster generation for video and 3d high-resolution images.
In discrete-time diffusion, a fixed forward process adds Gaussian noise across thousands of steps, while a learned reverse denoises to recover X0, whereas continuous-time SDE/ODE formulations decouple training from sampling.
Convert discrete steps into a continuous flow by introducing a time variable t in [0,1], moving from a forward process with signal and noise to a reverse stochastic differential equation.
Compare discrete DDPM with continuous SDE or ODE approaches, highlighting decoupled training and sampling, resolution independence, and distillation-free, faster sampling via flow matching.
Compare continuous time diffusion with ordinary differential equation solvers to see how 20-step dpm performs against 1000-step ddpm, highlighting drift and diffusion.
Compare Euler, Heun, and DPM ODE solvers for diffusion models; show probabilistic flow ODE converts noisy SDE to a smooth vector field, enabling fast sampling and quality-based solver choice.
Move from fixed-step solvers like Euler and Heun to adaptive dpm solvers and dopri, achieving faster, higher-accuracy diffusion with cfg and dauperry 5 and 8 variants.
This lecture compares continuous and discrete diffusion, showing continuous flow methods are fast and flexible for high-resolution vision and video, with flow matching and rectified flow as preferred approaches.
Compare continuous and discrete time steps using rectified flow, showing DDPM’s curved noise path versus a linear trajectory, guided by a velocity predictor in an ODE solver for HiFiGAN audio.
Compare flow matching and diffusion models, highlighting how discrete DDPM steps create latency, while continuous-time SDEs and ODEs enable direct paths from noise to final speech using velocity predictions.
Discover how diffusion geometry shapes model training by contrasting the curved, thousands-of-steps path with a direct straight-line path to Gaussian noise, reducing compute demands.
Explore how flow matching and rectified flow turn diffusion from thousands of curved steps into near straight paths, reducing cumulative error and enabling one-step generation in forward and reverse processes.
Explore why curved paths in diffusion models require many small steps to prevent error accumulation and drift from the ideal path toward pure noise, achieving lower FID and higher quality.
Explore knowledge distillation for diffusion models, comparing teacher and student models to balance small steps, latency, fid, and image quality amid nonlinear curve trajectories.
Explore diffusion step-count trade-offs, showing how 50 to 100,000 steps drive compute and latency. Lowering steps degrades FID, making real-time generation unacceptable; see a notebook-based, application-focused example.
Explore flow matching to guide Gaussian noise toward images using a trainable vector field, and apply rectified flow to achieve a straight, deterministic path from noise to the final image.
Rectified flow trains a vector field to map x1 to x0 along a straight path, using supervised training on sample pairs with a time parameter t balancing noise and signal.
Rectified flow cuts latency dramatically by replacing 1000 diffusion steps with far fewer steps, 200x faster, delivering good quality even with one step and enabling GPU parallelization.
Industry has embraced flow-based models like rectified flow and flow matching to achieve real-time performance, evolving from diffusion 20–50 steps to 4–10 steps and stable diffusion 3.
Explain how rectified flow uses a time parameter t to blend data and noise, guiding the model along a straight line from pure noise to a final image.
Apply Euler sampling to solve linear differential equations with a straight path, reducing steps and latency while preserving image quality and FID.
Explore diffusion versus flow matching in audio generation by comparing Grade TTS, Voicebox, and Mach R TTS, showing step reductions from thousands to four via optimal transport flow.
Explains flow matching for tts by mapping gaussian noise to Mell spectrogram via a vector field and ODEs, using linear interpolation and a text-guided velocity field.
Explore continuous flow matching in generative voice AI, guiding Gaussian noise through a vector field along a probability density path to form a Mell spectrogram from conditioning.
Break down a two-phase conditional flow matching approach for tts, detailing text encoding, duration alignment, and condition matrix creation, then noise-to-MEL spectrogram evolution via optimal transport and a target velocity.
Learn how Matcher TTS uses conditional flow matching to map noise to MEL spectrograms via an optimal transport path, enabling sub-10 step speech generation guided by text conditioning.
Explore voice box, a model for masked and conditional generation in continuous flow components, enabling infilling, editing, and zero short style transfer in speech synthesis.
Explore simulation-free training and one-step generation via rectified flow, using iterative retraining, and assess direct-to-waveform flow matching to bypass mel spectrogram and vocoder.
Learn conditional flow matching for text-to-speech by predicting velocity along a noise-to-mel spectrogram path with text conditioning, Euler integration, and classifier-free guidance to generate audio via Hi-Fi GAN.
Adopt end-to-end multi-modal architectures that replace cascade bottleneck, converting audio to acoustic tokens for a transformer and neural decoder to produce speech, tracing GSLM and GPT-4.0 progress and open challenges.
End-to-end multimodal architectures replace the cascaded ASR–LLM–TTS pipeline with a unified transformer that maps audio to audio, preserving paralinguistic cues and delivering real-time, low-latency output.
Explore a joint distribution approach for end-to-end multimodal voice synthesis that preserves emotion and sarcasm, with a single-transformer model lowering latency and mitigating cascade bottlenecks and error compounding.
Convert raw audio to discrete acoustic tokens using neural codec encoders with rvq, combine them with text tokens in a latent space, and apply self-attention on the interleaved multi-modal sequence.
Explore autoregressive audio generation with a unified transformer over interleaved audio and text tokens, converting generated tokens to waveforms via codec encoder/decoder for end-to-end, low-latency, emotionally coherent text-to-speech.
Discover how native multimodal transformers unify text and audio tokens in a shared embedding space, enabling end-to-end training, interleaved attention, low latency, and prosody-preserving, semantically rich TTS and ASR.
The lecture traces the shift from separate ASR and TTS to a unified transformer with text, audio, and vision in a shared embedding space, enabled by neural audio codecs.
Explores three frontiers in end-to-end multimodal architectures: zero-shot modality transfer via a shared semantic space, full-duplex speech with continuous streaming and back channeling, and ultra low bit-rate audio.
Explore zero-shot voice cloning and dynamic emotion transfer in native agents, guided by disentanglement theory and orthogonal latent spaces to control linguistics, speaker identity, and paralinguistics.
Master zero-shot voice cloning and dynamic emotion transfer for native agents using a 3 to 5 second reference to capture who and how, plus text for what to say.
Explore how a speech language model generates audio from text using three inputs—linguistic content, speaker identity, and emotion—facilitating controllable voice through orthogonal subspaces.
Using the information bottleneck, compress the reference audio into a dense speaker embedding that strips linguistic content and noise, while the source-filter model separates identity and emotion.
Explore how expressive native agents use speaker embeddings, frame-by-frame emotion latents (pitch, energy, duration), and orthogonal acoustic subspaces to enable zero-shot voice cloning via in-context learning.
Explore a six-step zero-shot voice cloning and emotion transfer framework that combines speaker identity, semantics, and emotion inputs to generate audio with a decoder.
Explore conditional fusion to inject speaker and emotion into semantic content via adaptive normalization or cross attention, acoustic realization with autoregressive or diffusion models, followed by vocoding to raw audio.
Compare traditional speaker adaptation, which required fine-tuning, with speech-language models enabling zero-shot voice cloning from 3–5 seconds of reference audio in context, with disentangled speaker identity and emotion control.
Explore global style tokens and variational autoencoders for voice cloning, comparing continuous GST and VAE approaches with discrete tokens, micro-prosody control, and state-of-the-art speech LMs.
Explore autoregressive and diffusion-based cloning for speech, modeling with acoustic tokens versus continuous latent spaces, guided by semantics, speaker, and emotion, and refined by vocoders.
Explore the evolution from 1990s gmm-based voice models to zero-shot cloning with speech lms, focusing on timbre-prosody entanglement, auto-distribution generalization, and the 3-second information limit.
Advance two futures in expressive voice cloning: an emotion manifold enabling unsupervised and self-supervised emotion discovery with interpolation and fine-grained micro prosody, and zero-shot multi-speaker synthesis for overlapping, full-duplex dialogues.
Apply chunked inference and speculative decoding to reduce synthesis latency to 200 milliseconds. Coordinate a fast draft model with a large target model for parallel verification to enable streaming audio.
Autoregressive transformers generate tokens sequentially, creating a chain rule bottleneck with high latency; chunked inferencing and speculative decoding reduce time to first byte and boost throughput.
Explore speculative decoding by pairing a draft model for generation with a fast target model for parallel verification, dramatically lowering latency and compute costs.
Learn how autoregressive models use causal masking and self-attention, moving from quadratic T^2 complexity to linear C×T with chunked attention, enabling low-latency streaming and memory efficiency.
Explore latency concepts in speech generation, focusing on time to first byte and streaming bitrate, and how draft and target models use speculative decoding to improve latency and throughput.
Learn how rejection sampling aligns draft and target model distributions with an acceptance alpha, using PX and QX to accept or replace tokens and reduce latency with KV caching.
Explore how key-value caching in streaming frames preserves acoustic continuity by reusing KV history to avoid recomputing tokens, reduce latency, and smooth pitch contours.
Explore state caching and stitching to smooth transitions between streaming TTS chunks, using KV cache, overlap-add, and latent-space smoothing to ensure continuous, artifact-free audio around 200 ms latency.
Compare traditional monolithic TTS with streaming SpeechLMs, showing how chunked input and causal masking deliver real-time latency around 200 ms while capturing global and local prosody.
Learn streaming latent diffusion with flow matching and sliding-window denoising for real-time voice generation, using overlapping 2–3 second windows to cut latency to around 200 ms while balancing quality.
Explore three latency challenges in streaming real-time speech synthesis—acoustic discontinuity, boundary issues, and draft target divergence—plus proposed fixes like overlap-add, KVCache, and adaptive K.
Trace the three eras of real-time speech synthesis, from pre-deep learning canned speech to modern streaming. See how cascaded models and WaveNet enable streaming via speculative decoding and chunked inference.
Minimize synthesis latency in real-time streaming tts by applying speculative decoding and chunked inference with overlap add, using draft and target models to balance speed and quality.
Explore deployment patterns for real-time voice interaction with WebSocket streaming, VAD, and turn-taking. Examine interruptions, barging, flow-yielding, open problems, and the evolution from half-duplex to full-duplex.
Explore deployment patterns for real-time voice agents, focusing on WebSocket streaming, interruption handling, and turn taking to maintain stateful, natural speech interactions with low latency.
Master full duplex voice agents with WebSocket streaming. Use finite state machines for idle, listening, processing, speaking, interrupted states, and turn-taking via energy, pitch, phonetics, ASR, VAD, and TTS.
Explore persistent bidirectional streaming for real-time voice agents, streaming audio via WebSocket with generation and playback in parallel, masking latency with Opus and near-instant time to first byte.
Learn how voice activity detection and neural endpointing enable robust turn-taking in real-time voice agents, using acoustic, ASR, and prosodic features to predict end-of-turn.
Explore barge-in and interruption handling in generative voice AI, where user energy detected by VAD triggers the agent to flush the audio buffer, stop DTS generation, reset context, and respond.
Explore turn-taking dynamics in dialogue with agents, distinguishing interruptions from back-channeling using energy, pitch, duration, and a turn-taking classifier to manage barging in real-time.
Describe end-to-end deployment of a speech agent, from initialization with a bidirectional web socket and mic streaming to endpoint triggering and fast stream generation with a speech LM and vocoder.
Explain how vad detects interruptions in a full duplex voice agent, trigger signal collision handling, stop tts, flush audio, apply semantic reversion, and yield the floor to the user.
Explore half-duplex systems where the user speaks while the agent listens, and vice versa, illustrating why real-time interaction struggles when acoustic activity is zero and wake words precede responses.
Explore how a full duplex native agent enables fluid conversation through parallel perception and generation in a shared vector space, delivering real-time dialogue with VAD, ASR, and TTS.
Identify three real-time deployment challenges: false margins from non-speech noises breaching acoustic energy thresholds, resource constraints for WAD and streaming TTS, and echo cancellation latency.
Trace the evolution of voice deployment architectures from legacy IVR to full-duplex WebSocket and WebRTC technologies, highlighting bidirectional flow, latency, and natural conversational dynamics.
This lecture explains moving from a reactive to an anticipatory turn-taking approach by embedding predictive end-pointing directly into the speech LM, enabling zero-latency generation and human-like naturalness.
Master deployment patterns for real-time voice agents using WebSocket streaming, interruption handling, and semantic reversion to manage tokens and latency during turn-taking.
Generative voice AI has moved far beyond simple text-to-speech — and this course takes you from the physics of sound all the way to building production-grade, agentic voice systems.
Most TTS courses stop at basic vocoders or off-the-shelf APIs. This one goes deeper. You'll start with the fundamentals of human speech — acoustics, phonetics, and prosody — before diving into the architectures actually powering today's state-of-the-art voice models: self-supervised representation learning (wav2vec 2.0, HuBERT), neural audio codecs (EnCodec, SoundStream, DAC), and the tokenization strategies that let LLMs "speak."
From there, you'll master the two dominant modern paradigms — autoregressive codec-based TTS and latent diffusion / conditional flow matching — understanding exactly when and why each is used in real systems. You'll also explore unified speech-text models, paralinguistic modeling (laughter, breathing, affect), and zero-shot voice cloning.
By the final module, you'll understand how to build low-latency, streaming, agentic voice pipelines — the same techniques behind real-time conversational AI agents — covering chunked inference, speculative decoding, WebSocket streaming, and turn-taking.
What you'll learn:
The science of speech production and acoustic feature extraction
How neural audio codecs and semantic tokenization work
Autoregressive and diffusion/flow-based TTS architectures
Cross-modal speech-text alignment techniques
Building low-latency, interruption-aware conversational voice agents
Whether you're an ML engineer, researcher, or voice-tech founder, this course gives you the complete architectural picture — from tokens to agents.