Loading comparison…
IPTV SPEECH · SIDE BY SIDE
Hear the voice.
Read the difference.
Ten clean Hindi speech clips. Play the audio, then compare the original Parakeet transcript from the cleaning manifest with a fresh Qwen3-ASR-0.6B transcription of that same clip.
01 / WORD DIFFERENCES
How far apart are they?
Parakeet is the comparison reference. These figures measure disagreement between two model outputs, not recognition accuracy.
Every clip, by edit type
CLICK A ROW TO INSPECTBar length is the clip’s WER. Colored sections show the share from each edit type; the full track equals the highest clip WER.
WER = (substitutions + deletions + insertions) ÷ Parakeet words. The pooled figure adds edits and Parakeet words across all ten clips before dividing. Word matching ignores punctuation and <hi-IN> tags; spelling, word breaks, and digits versus spoken numbers still count as differences.
02 / THE LISTENING ROOM
Choose a clip
Click a source in the list. The two transcripts refer to exactly the same audio.
CLIP 01 / 10
Loading clips…
The waveform shows level, not transcript alignment. Use the audio controls to play, pause or change speed.
See every changed word
Each tile shows Parakeet above Qwen. Alignment follows word order; there are no word timestamps.
03 / READING THE RESULTS
What “clean Hindi” means here
Each clip comes from a distinct IPTV episode and a segment marked good_audio.passed=true and downstream eligible in its cleaning manifest. The exact 48 kHz crop follows that segment’s original or enhanced audio route. We encoded the crop to MP3 for playback.
The manifest transcript contains Hindi writing. Independently, Whisper-small identified Hindi as the most likely language from the audio with probability at least 0.80. Qwen also identified the published audio as Hindi. Language scores help select examples; they are not human verification.
The Parakeet pane preserves the manifest transcript, including its language tags and possible errors. Qwen3-ASR-0.6B transcribed each published MP3 with automatic language identification and no context. These are model outputs, not corrected references. Read differences as observations, not accuracy scores. The WER shown uses Parakeet as a provisional reference and is directional.
The manifest’s episode ASR summary can say “English” even where its segment text contains <hi-IN> and Hindi speech. Clip selection uses the segment text and the audio check; the conflicting metadata is retained in the source manifest.