Posts
All the articles I've posted.
Deep Learning Audio Note 4: Language, Knowledge, and Deployment in ASR
Connect the five dimensions of ASR through language model fusion, knowledge sources, and the constraints of offline, streaming, and on-device deployment.
Deep Learning Audio Note 3: Alignment in ASR
From DTW and HMMs to CTC, LAS, and RNN-T: how continuous speech maps to discrete text, and the tradeoffs between alignment constraints, output dependencies, and streaming deployment.
Deep Learning Audio Note 2: ASR and Speech Representations
Updated:From ASR applications, WER/CER, and real-world challenges to MFCCs, CNNs, Transformers, Conformers, and self-supervised speech representations.
Deep Learning Audio Note 1: Speech Tasks and Audio Representations
An introduction to speech tasks, sound, sampling rates, Fourier transforms, STFT, Mel spectrograms, and MFCCs.
Omni Model Architectures: One Framework, Two Families
A visual comparison of omni model architectures across audio encoding, Thinker, Talker, vocoder, discrete codecs, continuous latents, and ASR.
Encode Before You Admit
How Fun-ASR moved full-context audio encoding ahead of LM admission with batching, single-flight deduplication, and LRU caching.