Line of inquiry
Inquiring lines›What enables authentic and grounde…›How should retrieval-augmented gen…›this line of inquiry
How do prompt structure and constraints affect model instruction reliability?
A broader line of inquiry — a family of 30 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 30
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does decomposed prompting formalize prompt libraries as reusable software modules?
- How do output format constraints compare to input exemplar brittleness?
- Can structured output formats reduce instruction following degradation?
- How does prompt brittleness across dimensions affect real-world applications?
- How do execution and planning tokens differ in their entropy dynamics?
- Does input length alone explain instruction density performance loss?
- Can algorithmic control flow over prompts simulate traditional programming languages?
- Why do semantically related prompts converge into attractor states in middle layers?
- Do recency-focused prompts and in-context examples work equally well for order recovery?
- Why does sandboxed execution matter more than monolithic prompting?
- Does minimal code engagement during vibe coding harm students' long-term programming comprehension?
- What makes draft-centric systems better anchors for coherence than feed-forward outputs?
- How do ordering effects compound across different prompt component scales?
- How does entropy-based patching compare to fixed token vocabularies in practice?
- How much does shared-prefix sampling reduce token redundancy empirically?
- Why does token ordering in LLMs create sequences rather than true temporal flow?
- How do early-prefix tokens control the generation of entire continuations?
- How can we reorganize repositories to make behaviors easier to locate?
- How should headers index procedural intent differently from keyword chunking?
- Why is digital context more volatile than conventional software context?
- How do language agents implement prompts as executable computational graphs?
- Can this whole-artifact principle apply to other generative tasks?
- How does token generation as flow differ from print's archival storage?
- What role does rigid output format play in function calling failure modes?
- How do logic units preserve document structure better than fixed-size chunking?
- How do execution traces represent state and dynamics in codebase modeling?
- How does prior coding experience change the way students use vibe coding tools?
- How do RAG and prompting techniques differ in supporting each granularity level?
- What failure modes does the negative-space checklist generation method actually catch?
- Can this approach handle continuously changing product inventories in production?