← All models

automatic-speech-recognition · forced alignment · WASM · ~241 MB

MMS forced alignment — when exactly did they say it?

Speech-to-text tells you what was said. Forced alignment tells you when: give this model an audio clip and the transcript that belongs to it, and it times every word — the exact start and end of each one — entirely on your device. That's how subtitles, karaoke, and pronunciation apps are built.

At a glance

Modelonnx-community/mms-300m-1130-forced-aligner-ONNX · original MahmoudAshraf/mms-300m-1130-forced-aligner
Taskautomatic-speech-recognition (CTC) — used here for forced alignment, not transcription
ArchitectureWav2Vec2ForCTC — Meta's MMS-300m backbone fine-tuned on a forced-alignment dataset; 31-symbol char alphabet ( + letters + apostrophe), one CTC distribution per ~20ms frame
Params / dtype300M · q4 (4-bit quantized ONNX — the int8 export uses ConvInteger, which has no WASM kernel, so this demo honestly uses q4)
BackendWebAssembly (CPU) — no GPU required
Download~241 MB, cached after first load
LicenseCC-BY-NC-4.0 (non-commercial — recorded honestly)
Web APIsWeb Workers, Web Audio (decode + resample), Cache Storage (via Transformers.js)
License: CC-BY-NC-4.0 — non-commercial. The MMS-300m forced-aligner weights are released under a Creative Commons Non-Commercial license. This demo is for personal and educational evaluation; commercial use of the weights requires a separate license. Model credit: Meta AI (MMS-300m backbone), MahmoudAshraf (forced-alignment fine-tune), onnx-community (ONNX conversion). See the model card.

Run it

…or drop / choose your own audio (wav, mp3, m4a, webm)

Bundled JFK sample: inaugural address excerpt (20 January 1961), spoken by John F. Kennedy — public domain, U.S. government work. JFK audio source and attribution record.

See inside — the CTC frame strip

Wav2Vec2 models score every ~20ms audio frame against a tiny alphabet. We take the transcript as ground truth, run the model once, and monotonically match each character of your transcript to the frame where it fires — that's the alignment. The readout shows the real numbers: how many frames, the frame length, and how many transcript characters were matched (fewer than 100% means silence or a transcript mismatch — honest, never faked). We also compare with the model's own greedy reading below.


      

Why it matters

Almost every product that shows when words happen — subtitles, karaoke, lyric videos, pronunciation feedback, podcast transcripts, search-within-audio — depends on forced alignment. Running it in the browser means audio never leaves the device: private by construction.

Use cases

Basics

Align the JFK clip and click through the words — hear exactly where each one lands.

Practical

Build subtitles: align your own audio + transcript and export a copyable SRT block.

Wild

Your voice, your words — record yourself, type what you said, watch the alignment follow you.

How the API works

The model itself is a plain CTC speech model; alignment is an algorithm on top of its frames:

import { AutoProcessor, AutoModelForCTC } from "@huggingface/transformers";

const processor = await AutoProcessor.from_pretrained(MODEL);
const model = await AutoModelForCTC.from_pretrained(MODEL, { dtype: "int8" });
const { logits } = await model(await processor(audio)); // [1, frames, 31]

// 1. tokenize the KNOWN transcript into character ids (spaces → id 3)
// 2. argmax each frame → emission strip; collapse CTC blanks/repeats
// 3. monotonic-match transcript chars to emission frames
// 4. group chars into words → start/end = frame × 20 ms

We deliberately call the model directly instead of the ASR pipeline — the pipeline hides the per-frame output we need. Everything runs in a Web Worker; audio is decoded and resampled to 16 kHz mono via Web Audio first.

References