Skip to content
YC's Blog
Go back

Deep Learning Audio Note 4: Language, Knowledge, and Deployment in ASR

This is the last post in this round of ASR notes. After several consecutive days of writing, I was getting a little tired, so this draft started with AI-generated text that I then rewrote and organized.

The previous post explored Alignment: how continuous speech maps to discrete text. This post covers Language, Knowledge, and Deployment, connecting all five dimensions:

ASR = Representation + Alignment + Language + Knowledge + Deployment

This is an analytical framework for understanding systems, rather than five mandatory modules executed in sequence.

阅读中文版。

Language: when several words sound alike, which makes sense?

Alignment connects positions in speech to output tokens. Even with a valid alignment, acoustic evidence alone may not determine the transcript.

Consider two / to / too in:

I want ___ go home

Their pronunciations can be identical or very similar, but language regularities favor to. A text-only language model estimates:

P(yᵤ | y<ᵤ)

This is the probability of the next token given preceding text. ASR must also use audio X. Language priors cannot replace acoustic evidence: otherwise, a recognizer may rewrite a real but uncommon expression into a familiar sentence.

Language modeling has several routes. A hand-written grammar can specify COMMAND → ACTION OBJECT. An n-gram LM estimates probabilities from a limited history, such as P(wₜ | wₜ₋₂, wₜ₋₁). Neural LMs, including LSTMs and Transformers, can use longer histories. More regularities have become learned from data, while explicit grammars and finite-state constraints remain useful for particular tasks.

How does language enter ASR?

ASR architectures use text history in different ways.

CTC factorizes path probabilities across time given the audio, without an explicit autoregressive text state. Its encoder can still learn language-related regularities from paired training data. An external LM can score candidates during beam search or rerank them afterward:

CTC frame scores + optional LM scores
                 ↓
          Beam search → hypotheses
                 ↓
          Optional LM rescoring

LAS / AED directly conditions its decoder on previous text and acoustic context:

P(Y | X) = ∏ᵤ P(yᵤ | y<ᵤ, X)

RNN-T organizes audio and text history into two information streams:

Audio → Encoder ──────────┐
                          ├→ Joiner → token / blank
Text history → Predictor ─┘

The predictor represents previously emitted non-blank tokens, and the joiner combines this with encoder output. RNN-T explicitly handles acoustic information, monotonic alignment, and text history. Its predictor has an LM-like role, but is not equivalent to an independently trained LM; some variants also restrict its history. An external LM can still be added.

Rescoring and fusion

One direct use of an external LM is to generate several hypotheses, then rerank them. A representative score is:

S(Y) = log P_ASR(Y | X) + λ log P_LM(Y) + β len(Y)

λ controls the LM weight; β is an optional length or insertion reward whose definition depends on the decoder. Rescoring cannot recover a correct transcript that is absent from the candidate set.

Weighted scores can also guide search itself. The integration point matters:

MethodHow the LM participates
RescoringScores N-best candidates or lattice paths after first-pass recognition
Shallow fusionCombines ASR and LM scores during decoding, influencing candidate expansion
Deep fusionCombines hidden representations from separately trained decoders and LMs, then trains the fusion components
Cold fusionIntroduces a pretrained, typically frozen LM during Seq2Seq training, with gates that learn how to use it

These names come from sequence generation research; ASR implementations can vary. The distinctions include when the LM joins the system, whether scores or representations are combined, and which parameters are trained.

More language capability inside ASR does not eliminate external LMs. Additional text, domain vocabulary, and decoding constraints can still help. Their value should be measured through recognition errors and unwanted corrections in the target setting.

Knowledge: where do these capabilities come from?

Here, Knowledge means sources of knowledge and how it is acquired, rather than factual question-answering ability. It is an analytical dimension in this article, not a standard ASR component strictly separate from Language.

Language asks what linguistic regularities a model knows. Knowledge asks how it learned them.

Preferring I want to go home over I want two go home reflects a language regularity. That preference might come from a grammar, labeled transcripts, or large text corpora. Pronunciation lexicons provide word-to-pronunciation mappings; unlabeled speech, self-supervised objectives, weakly supervised pairs, and pseudo-labels provide different learning signals.

Knowledge therefore affects several dimensions:

                Knowledge
                    ↓
        ┌───────────┼───────────┐
        ↓           ↓           ↓
Representation  Alignment   Language

From expert structure to learning from data

Traditional systems contain substantial human structure: phoneme inventories, lexicons, HMM topologies, and task grammars. In GMM-HMM systems, people define structures while data supplies acoustic distributions, transition probabilities, and statistics such as n-gram probabilities.

DNN-HMM systems use neural networks to provide acoustic state posteriors, converted into decoding scores through operations such as prior correction. Lexicons, state structures, and external LMs remain. Deep learning did not suddenly remove human-designed structure.

End-to-end methods such as CTC, RNN-T, and LAS let neural networks learn mappings from audio to output tokens, reducing dependence on explicit phoneme states and pronunciation lexicons:

Traditional system:
features + acoustic scores + HMM + lexicon + LM → text

End-to-end system:
audio → neural sequence model → text tokens

This is a structural simplification. Traditional components often participate in a joint search rather than a mechanical pipeline. End-to-end systems still require choices about tokenization, networks, alignment constraints, and objectives. CTC, RNN-T, and AED also model output dependencies differently; they do not all contain the same kind of internal LM.

Self-supervision changes knowledge acquisition

wav2vec 2.0 learns representations from unlabeled speech before ASR fine-tuning with transcribed audio. Its key contribution concerns knowledge acquisition, rather than a new alignment mechanism replacing CTC.

These concepts describe different dimensions:

ConceptWhat it primarily describes
Transformer / ConformerNetwork architecture
CTCA training objective with a blank-collapse alignment rule
RNN-TA sequence model and objective with a predictor, joiner, and marginalization over monotonic paths
Self-supervised learning (SSL)A strategy for constructing learning signals from data and learning representations

Whisper uses a Transformer encoder-decoder trained on roughly 680,000 hours of multilingual, multitask weakly supervised audio. It demonstrates the value of scale and data diversity for generalization across datasets, rather than inventing a new basic architecture. Generalization does not establish reliability for every accent, noise condition, or specialist domain.

Foundation model training signals can come from paired audio and text, unlabeled speech, web text, weak labels, pseudo-labels, synthetic data, and multilingual or multimodal data. Architecture alone cannot explain capability differences: data quality, coverage, objectives, and training scale also matter.

Deployment: under what constraints must the model run?

Representation, Alignment, and Language explain how speech becomes text. Deployment asks under what practical conditions this must happen, shaping model design and training.

Offline ASR

Offline recognition can process the complete input X = x₁:T, allowing bidirectional context, global attention, larger searches, and second-pass rescoring.

It can spend more computation on accuracy, but still faces throughput, memory, cost, and completion-time constraints. Long recordings may require segmentation because of memory limits.

Streaming ASR

Streaming recognition produces results as audio arrives. A causal model uses available audio; a model with limited right context can be written as:

hₜ = f(x₁:ₜ₊ᵣ)

Computing hₜ requires waiting for the necessary r future frames. More context may improve recognition, at the cost of waiting. Quality gains are not necessarily monotonic and depend on training and data.

RNN-T’s monotonic paths support incremental decoding, but streaming also requires a causal encoder or bounded future context. The original RNN-T paper discusses bidirectional networks too. CTC and AED designed for streaming can also support online recognition.

Deployment shapes the encoder. Standard global self-attention has quadratic attention computation in sequence length T. Common optimizations include chunked or local attention, bounded right context, causal convolutions, and stronger subsampling.

FastConformer improves efficiency through changes including subsampling, and studies limited-context attention for long recordings. Support for long recordings does not establish low-latency streaming: right context, convolutional access to future frames, caching, and chunking still need inspection.

Streaming latency is more than compute time

For a particular result being committed, a rough diagnostic model is:

L ≈ L_buffer/context + L_compute + L_queue + L_commit/endpoint

These components can overlap. End-to-end experience may also include capture, transport, and display time. An evaluation should specify whether it measures the first partial result, stable text, or final output after speech ends.

Real-Time Factor (RTF) generally divides processing time by audio duration. Low RTF indicates fast processing under the measured conditions, but does not guarantee low first-token latency, commitment delay, or tail latency under load.

Edge ASR: balancing several objectives

Phones, cars, and earbuds add memory, compute, power, battery, thermal, and privacy constraints. The goal becomes minimizing WER within resource and latency budgets.

CTC greedy decoding is simple:

Encoder → vocabulary projection → argmax → collapse

Collapse first merges consecutive repeated symbols, then removes blanks. A simple decoder can reduce overhead, but the encoder may dominate total cost. The CTC label alone does not establish suitability for a device.

RNN-T adds text-history modeling and sequential decoding computation, providing a way to combine history dependencies with monotonic incremental output. Its suitability needs comparison on the same data, hardware, and latency budget. Quantization, model size, caching, search width, and runtime implementation also affect the outcome.

Connecting all five dimensions

The framework can now be written as:

Speech → Representation H → Alignment + Language → Text Y

The last two dimensions influence the whole system; Alignment and Language act together in many models:

                 Knowledge sources / training signals
                                  ↓
Speech → Representation → [Alignment + Language] → Text
                                  ↑
                       Deployment constraints

Long-term trends include learning more from data, training components jointly, and considering data scale alongside deployment design. These approaches coexist with explicit constraints, modular systems, and external LMs.

From GMM-HMM and DNN-HMM to CTC, RNN-T, LAS, Whisper, and speech foundation models, the change is more than increasing model size. Representations, alignment, output dependencies, knowledge sources, and practical operating constraints have all evolved together.


Share this post on:

Next Post
Deep Learning Audio Note 3: Alignment in ASR