Whisper is OpenAI's speech-to-text model trained on 680,000 hours of multilingual audio. It shows that training on diverse, weakly-labeled data scales better than clean, curated datasets.
The Traditional ASR Problem
Before Whisper, speech recognition systems were:
Traditional ASR:
1. Collect clean, labeled audio (expensive)
2. Train on single language
3. Works perfectly on that language
4. Fails on accents, background noise, other languages
Whisper's Insight: Weak Supervision at Scale
Whisper approach:
1. Collect 680,000 hours of YouTube video-audio
2. Use existing closed-caption system for labels (weak)
3. These are NOT hand-labeled, may have errors
4. Train huge model on this imperfect data
Result:
- Model learns from noise, accents, background
- Generalizes better than clean-data models
- Works in 99 languages!
This is scale over quality strategy.
Training Data
Multilingual audio from YouTube:
English: 117,000 hours
Spanish: 32,000 hours
German: 15,000 hours
Russian: 11,000 hours
Chinese: 12,000 hours
French: 12,000 hours
... (99 languages total)
Total: 680,000 hours = 97 years of continuous audio
Collected without manual labeling!
Architecture
Encoder:
- Input: 16 kHz audio (STFT)
- Output: 1500-dim hidden state
- 24 Transformer blocks
- 4 attention heads per layer
Decoder:
- Autoregressive transformer
- Generates tokens one-by-one
- Conditioned on encoder output
Special tokens:
- <|startoftranscript|>
- <|en|> (language token)
- <|transcribe|> or <|translate|>
Model Sizes
Model | Parameters | Speed | Accuracy
------|-----------|-------|----------
Tiny | 39M | ~10x | 85%
Base | 74M | ~8x | 89%
Small | 244M | ~6x | 91%
Med. | 769M | ~2x | 92%
Large | 1.5B | 1x | 94%
Speed = relative to real-time speech
Trade-off: smaller is faster, larger is more accurate.
Performance
Accuracy on test sets:
English (clean audio):
- Human baseline: 95%
- Whisper-Large: 96% (better than humans!)
Multilingual:
- Whisper: 80-85% accuracy across 99 languages
With background noise:
- Whisper-Large: 92% even with 10dB noise
Practical Deployment
from faster_whisper import WhisperModel
model = WhisperModel("base")
segments, info = model.transcribe("audio.mp3")
for segment in segments:
print(segment.text)
Limitations
Whisper can't:
- Handle multiple speakers well
- Do speaker identification
- Do speaker diarization
- Reliably translate between non-English pairs
