Line of inquiry
Inquiring lines›How do language models construct a…›How do dialogue systems achieve ge…›this line of inquiry
Do language model representations contain causally steerable task-specific features?
A broader line of inquiry — a family of 20 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 20
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can interventions on individual features reliably steer language model behavior?
- How do semantic features in representations become steerable task-specific directions?
- How do language models transmit traits through semantically unrelated data?
- Do representations in models causally influence text generation?
- Is gradient behavior in language functional or a sign of ambiguity?
- Do reading vectors from activation space causally control model behavior?
- Do all semantic steering effects follow predictable patterns based on feature alignment?
- Can steering vectors prove that representations are genuinely organized?
- Can models transmit behavioral traits through semantically unrelated synthetic data?
- Can a single SAE feature control reasoning behavior across model families?
- Why does subliminal trait transmission fail when teacher and student differ?
- What makes some concepts more steerable than others in activation space?
- Does the Assistant Axis exist in pre-trained models before instruction tuning?
- What causes gradient-based steering via natural language descriptions to work?
- What neural or architectural mechanism allows selective override of frequency effects?
- How does LatentQA differ from predefined concept steering like representation engineering?
- How does Western-dominance bias propagate through multimodal training data?
- What other behavioral properties exist as linear directions in activation space?
- Why can data filtering fail to remove transmitted behavioral traits?
- Why do handcrafted acoustic features outperform neural speaker embeddings for personality?