Hi, I'm
Yuhao Chen
SGLang Core Contributor
About Me
I build and optimize LLM and multimodal inference systems, focusing on model serving, scheduling, batching, and caching in SGLang and SGLang Omni.
Looking for Inference, RL, or AI Infrastructure Engineer roles.
Recent Posts
Deep Learning Audio Note 3: Alignment in ASR
From DTW and HMMs to CTC, LAS, and RNN-T: how continuous speech maps to discrete text, and the tradeoffs between alignment constraints, output dependencies, and streaming deployment.
Deep Learning Audio Note 2: ASR and Speech Representations
Updated:From ASR applications, WER/CER, and real-world challenges to MFCCs, CNNs, Transformers, Conformers, and self-supervised speech representations.
Deep Learning Audio Note 1: Speech Tasks and Audio Representations
An introduction to speech tasks, sound, sampling rates, Fourier transforms, STFT, Mel spectrograms, and MFCCs.
Omni Model Architectures: One Framework, Two Families
A visual comparison of omni model architectures across audio encoding, Thinker, Talker, vocoder, discrete codecs, continuous latents, and ASR.