Line of inquiry
Inquiring lines›Why are language models fragile de…›How do learned model representatio…›this line of inquiry
Why does direct speech processing outperform transcription-based approaches?
A broader line of inquiry — a family of 21 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 21
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does the articulatory substrate explain direct speech-to-speech superiority over transcription pipelines?
- Do speech models learn the articulatory processes that produce acoustic signals?
- Do speech encoders actually learn the physics of how vocal tracts produce sound?
- Can speech embeddings carry articulatory structure that text cannot?
- How do speech encoders learn articulatory physics without phonetic labels?
- Why do current speech benchmarks fail to measure reasoning over audio?
- What information does transcription destroy that direct speech-to-speech models preserve?
- Why does transcription destroy prosodic information in speech processing?
- Why does articulatory probing predict SSL model performance better than phonetic probing?
- What information does transcription destroy that direct speech pathways preserve?
- Can feature disentanglement in gesture synthesis generalize to completely unseen voice distributions?
- Why do speech benchmarks still measure transcription instead of comprehension?
- Can skipping transcription reduce speech dialogue latency below 300 milliseconds?
- How does removing transcription change speech-to-speech generation latency?
- Can articulatory inversion serve as a window into what speech models have learned?
- Does direct speech-to-speech generation really eliminate transcription latency?
- How do different speech encoder layers capture different types of gesture information?
- How much latency improvement comes from collapsing the speech pipeline?
- What moves become possible when you represent ASR as a noisy observation model?
- Why do handcrafted acoustic features outperform neural speaker embeddings for personality?
- What paired speech data is needed to train end-to-end models?