Line of inquiry
Inquiring lines›Where does language-model reasonin…›How do language models represent m…›this line of inquiry
What articulatory information do speech signals carry that text cannot?
A broader line of inquiry — a family of 34 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 34
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does the articulatory substrate explain direct speech-to-speech superiority over transcription pipelines?
- Can speech embeddings carry articulatory structure that text cannot?
- Can feature disentanglement in gesture synthesis generalize to completely unseen voice distributions?
- Do speech models learn the articulatory processes that produce acoustic signals?
- Do speech encoders actually learn the physics of how vocal tracts produce sound?
- How does causal multimodal modeling differ from encoder-decoder architectures?
- Do discrete tokenized modalities preserve information better than continuous embeddings?
- How do speech encoders learn articulatory physics without phonetic labels?
- How do sparse mixture-of-experts models resolve modality capacity competition?
- What scaling exponent would audio or other modalities require in a truly multimodal system?
- Why do multimodal models fail on rare and underrepresented concepts?
- What emergent abilities appear only in truly unified multimodal systems?
- Can one streaming model handle turn-taking better than cascaded ASR-LLM-TTS?
- How do different speech encoder layers capture different types of gesture information?
- What information does transcription destroy that direct speech-to-speech models preserve?
- Why does articulatory probing predict SSL model performance better than phonetic probing?
- Can dense models partially address modality friction without full expert specialization?
- Why does transcription destroy prosodic information in speech processing?
- Can multimodal architectures successfully integrate vision without replicating past failures?
- What makes multimodal conditioning effective when features are decomposed to the right granularity?
- How much latency improvement comes from collapsing the speech pipeline?
- Can articulatory inversion serve as a window into what speech models have learned?
- Why do image captions create different friction than pure video data?
- Can skipping transcription reduce speech dialogue latency below 300 milliseconds?
- What makes internal embeddings useful as multimodal input for language model training?
- What information does transcription destroy that direct speech pathways preserve?
- How does removing transcription change speech-to-speech generation latency?
- Does direct speech-to-speech generation really eliminate transcription latency?
- Can multimodal LLMs be made to spontaneously adapt their language for efficiency?
- How do multimodal AI architectures compare to human brain export pathways?
- What paired speech data is needed to train end-to-end models?
- Why do cascade pipelines fail to capture global motion structure?
- How does mixture of experts enable flexible capacity sharing between modalities?
- What temporal and spatial constraints does Space-Time U-Net solve?