automatic-speech-recognition · forced alignment · WASM · ~241 MB
Speech-to-text tells you what was said. Forced alignment tells you when: give this model an audio clip and the transcript that belongs to it, and it times every word — the exact start and end of each one — entirely on your device. That's how subtitles, karaoke, and pronunciation apps are built.
| Model | onnx-community/mms-300m-1130-forced-aligner-ONNX · original MahmoudAshraf/mms-300m-1130-forced-aligner |
|---|---|
| Task | automatic-speech-recognition (CTC) — used here for forced alignment, not transcription |
| Architecture | Wav2Vec2ForCTC — Meta's MMS-300m backbone fine-tuned on a forced-alignment dataset; 31-symbol char alphabet ( |
| Params / dtype | 300M · q4 (4-bit quantized ONNX — the int8 export uses ConvInteger, which has no WASM kernel, so this demo honestly uses q4) |
| Backend | WebAssembly (CPU) — no GPU required |
| Download | ~241 MB, cached after first load |
| License | CC-BY-NC-4.0 (non-commercial — recorded honestly) |
| Web APIs | Web Workers, Web Audio (decode + resample), Cache Storage (via Transformers.js) |
Bundled JFK sample: inaugural address excerpt (20 January 1961), spoken by John F. Kennedy — public domain, U.S. government work. JFK audio source and attribution record.
word timestamps — click a word to seek the audio
Wav2Vec2 models score every ~20ms audio frame against a tiny alphabet. We take the transcript as ground truth, run the model once, and monotonically match each character of your transcript to the frame where it fires — that's the alignment. The readout shows the real numbers: how many frames, the frame length, and how many transcript characters were matched (fewer than 100% means silence or a transcript mismatch — honest, never faked). We also compare with the model's own greedy reading below.
what the model heard (greedy CTC — for comparison, not used for alignment)
Almost every product that shows when words happen — subtitles, karaoke, lyric videos, pronunciation feedback, podcast transcripts, search-within-audio — depends on forced alignment. Running it in the browser means audio never leaves the device: private by construction.
The model itself is a plain CTC speech model; alignment is an algorithm on top of its frames:
import { AutoProcessor, AutoModelForCTC } from "@huggingface/transformers";
const processor = await AutoProcessor.from_pretrained(MODEL);
const model = await AutoModelForCTC.from_pretrained(MODEL, { dtype: "int8" });
const { logits } = await model(await processor(audio)); // [1, frames, 31]
// 1. tokenize the KNOWN transcript into character ids (spaces → id 3)
// 2. argmax each frame → emission strip; collapse CTC blanks/repeats
// 3. monotonic-match transcript chars to emission frames
// 4. group chars into words → start/end = frame × 20 ms
We deliberately call the model directly instead of the ASR pipeline — the pipeline hides the per-frame output we need. Everything runs in a Web Worker; audio is decoded and resampled to 16 kHz mono via Web Audio first.