Theme of inquiry
How should retrieval-augmented generation systems be architected and triggered?
A question within its area, explored through 9 lines of inquiry below — each a family of specific questions the research asks.
45 specific questions
- Does self-reflection help models notice their own constraint violations?
- Why does self-reflection during training fail to improve model self-correction?
- When does self-reflection actually help reasoning models improve?
- How does metacognitive self-correction enable models to revise failed strategies?
- Does reflection training actually teach models to self-correct their mistakes?
- Why does self-critique fail without external verification signals?
- How do prior errors in context history amplify future failures over time?
59 specific questions
- Can automated tools close the gap between AI generation and verification?
- How does the generation-verification gap limit AI self-improvement capabilities?
- Can verification tools keep pace with AI artifact generation speed?
- Can AI evaluation tools solve the verification problem they help create?
- Where does the generation-verification gap appear in test-time compute?
- Why does AI generation outpace verification across the research lifecycle?
- How can agents verify research artifacts faster than they generate them?
25 specific questions
- How does process supervision relate to execution-signaled feedback approaches?
- Can self-supervised process models replace human annotations at scale?
- Can trajectory structure alone provide process supervision without human annotation?
- Does self-supervised process supervision work for domains with ambiguous correctness?
- Do self-supervised process reward models scale better than human annotation?
- Can self-supervised methods replace human annotations for process reward models?
- What other trajectory structures could reveal hidden process supervision signals?
15 specific questions
- Why does greater automation actually obscure rather than eliminate research failure modes?
- Does refining around bad results risk cascading errors in automated research?
- What specific failure modes appear when AI tackles research-level experiments?
- How does executable evaluation feedback sustain autonomous discovery at scale?
- Can automating failure absorption hide problems that governance needs to surface?
- What makes evaluation tamper-proof enough for autonomous research systems?
- Can evaluators investigate dependencies without accumulating mistakes over time?
47 specific questions
- How can we turn reasoning model failures into useful training signals?
- How does error avalanching compound failures in self-training iterations?
- What makes preventative lessons from failures more valuable than success patterns?
- Can AI outputs inspire new directions even when they seem like failures?
- Can population diversity in self-improvement prevent error avalanching failures?
- Can model training address failures that really originate in harness gaps?
- What happens when error accumulation and preference signal collapse occur together?
30 specific questions
- How does decomposed prompting formalize prompt libraries as reusable software modules?
- How do output format constraints compare to input exemplar brittleness?
- Can structured output formats reduce instruction following degradation?
- How does prompt brittleness across dimensions affect real-world applications?
- How do execution and planning tokens differ in their entropy dynamics?
- Does input length alone explain instruction density performance loss?
- Can algorithmic control flow over prompts simulate traditional programming languages?
33 specific questions
- How does externalizing reasoning into harness artifacts improve agent reliability?
- Why does externalized state beat parameter scaling for agent reliability?
- How does external context control compare to agents managing their own state internally?
- How do externalizing cognitive work and coordination infrastructure relate to agent reliability?
- How does structured environment-side state reduce multi-turn agent failure better than transcript replay?
- Can externalizing bookkeeping to a stateful harness replace internalized memory control?
- Where does agent reliability come from if not better tools?
21 specific questions
- Can delegation prevent silent corruption in long delegated workflows?
- Why does increased model capability make detection harder in delegated workflows?
- Why do frontier models corrupt more documents than weaker models during workflows?
- Can tool use or self-conditioning fix degradation in extended LLM workflows?
- Why does workflow position amplify malicious signals in multi-agent relay chains?
- Which model capabilities actually matter for sustained workflow delegation?
- How do workflow-inspecting defenses fail when contamination enters at planning time?
33 specific questions
- Can minimal adversarial triggers disrupt reasoning across multiple unrelated queries?
- Are reasoning models more vulnerable to adversarial manipulation than standard models?
- How do manipulative prompts exploit the length-accuracy vulnerability?
- How do adversarial triggers bypass the protections of longer reasoning chains?
- Can manipulative prompts reduce reasoning model accuracy without fine-tuning?
- Can reasoning models distinguish between new evidence and manipulative reframing?
- How can simple prompt injection attacks extract reasoning trace content?