Archives
All the articles I've archived.
Deep Learning Audio Note 4: Language, Knowledge, and Deployment in ASR
Connect the five dimensions of ASR through language model fusion, knowledge sources, and the constraints of offline, streaming, and on-device deployment.
Deep Learning Audio Note 3: Alignment in ASR
From DTW and HMMs to CTC, LAS, and RNN-T: how continuous speech maps to discrete text, and the tradeoffs between alignment constraints, output dependencies, and streaming deployment.
Deep Learning Audio Note 2: ASR and Speech Representations
Updated:From ASR applications, WER/CER, and real-world challenges to MFCCs, CNNs, Transformers, Conformers, and self-supervised speech representations.
Deep Learning Audio Note 1: Speech Tasks and Audio Representations
An introduction to speech tasks, sound, sampling rates, Fourier transforms, STFT, Mel spectrograms, and MFCCs.
Omni Model Architectures: One Framework, Two Families
A visual comparison of omni model architectures across audio encoding, Thinker, Talker, vocoder, discrete codecs, continuous latents, and ASR.
Encode Before You Admit
How Fun-ASR moved full-context audio encoding ahead of LM admission with batching, single-flight deduplication, and LRU caching.
Making a content cache mean what it says
Investigating a subtle batch-invariance bug in MOSS-TTS Local v1.5's reference-audio cache, where content-addressed keys hid execution-shape-dependent values.
Design agentic alpha discovery system
Design agentic alpha discovery system
Design Google Translate
Updated:System design for Google Translate: ML framing, architecture, NMT training, and inference pipeline
Minimax Speech 2.0
Updated:Reading notes on the Minimax Speech 2.0 paper: technical details of a speech synthesis model
Seedream 3.0 Technical Report
Updated:Reading notes on the Seedream 3.0 technical report
2025-03-23 Weekly Podcast Notes
Updated:Weekly podcast notes: Japan's health insurance system, short dramas entering Hollywood, Tesla/NVIDIA, Google's Willow quantum chip, and more
The Art of Subtraction in Setting-Based Mystery: Why I Do Not Like A Sweet Death for the Famous Detective
Updated:A review and reflection on setting-based mystery novels, centered on A Sweet Death for the Famous Detective
Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model
Updated:Reading notes on the Seedream 2.0 technical report: a Chinese-English bilingual image generation foundation model
2025-03-15 Weekly Podcast Notes
Updated:Weekly podcast notes: Silicon Valley 101 on MicroStrategy/Bitcoin, a Latent Space interview, and more
Thoughts on Exercise
Updated:Thoughts on exercise, warm-up and rehab, the benefits of exercise, and different happy hormones
Why Write a Blog
Updated:Some thoughts on why I started writing a blog: recording ideas, keeping my ability to think, and building a personal knowledge base