SYNTHESIS NOTE
Topicsthis note

How does test-time scaling work for individual research agents?

This explores whether the scaling laws that apply to reasoning tokens also apply to search steps and information retrieval in agentic systems. Understanding this could reveal new ways to improve AI research capabilities through compute allocation.

Synthesis note

Core Insights


description: Navigation hub for deep research scaling, agentic system architectures, knowledge graph reasoning, and search-as-TTS — individual agent-level test-time compute type: topic-map created: 2026-02-24 topics: ["How do you navigate synthesis across fragmented research topics?"]

deep research and agentic systems topic map

How does test-time scaling work at the individual agent level? This sub-map covers deep research scaling (where TTS law generalizes from reasoning tokens to search steps), agentic system architectures (data efficiency, skill libraries, adaptation paradigms), and knowledge graph reasoning (externalized reasoning, graph-structured training data synthesis).

The key insight: search budget follows the same scaling curve as reasoning tokens, making deep research a TTS problem. At the agent level, data efficiency is extreme — 78 curated demonstrations outperform 10K samples for agency.

Parent map: How does test-time scaling work at the agent level?

Deep Research and Search Scaling

(From Arxiv/Deep Research)

Proactive Search Evaluation (2026-05-28 — VibeSearchBench)

The evaluation-experience gap in search: benchmarks reward what users never struggle with, and realistic evaluation needs vague intent, multi-turn dialogue, and open-ended structure.

Production Deployment

Agentic Systems

(From Arxiv/Agents — agent architectures, team optimization, data efficiency for agency, adaptation paradigms)

Knowledge Graph Reasoning and Training Data Synthesis

(From Arxiv/Knowledge Graphs — graph-structured reasoning, KG externalization, and synthetic training data from KGs)

Autonomous Science and Ideation — Batch #3 backlog (2026-06-03)

Three papers on AI doing research: two architectures for long-horizon autonomy, and one reframing of why LLM ideation underwhelms.

AI-for-AI and Research Venues — Batch #3 wave 2 (2026-06-03)

Related Areas

New — 2026-06-27

Inquiring lines that read this note 38

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can model routing outperform monolithic scaling as an efficiency strategy? Do autonomous architecture discoveries follow predictable scaling laws? Does parallel reasoning outperform sequential thinking under fixed compute budgets? What actually drives chain-of-thought reasoning improvements in language models? How should iterative research systems allocate reasoning per search step? Can inference-time compute substitute for scaling up model parameters? How do knowledge graphs enable efficient multi-hop reasoning over alternatives? How can LLM user simulators model realistic goal-driven conversation? When does architectural design matter more than raw model capacity? How can identical external performance mask different internal representations? How do adversarial and manipulative prompts attack reasoning models? How can AI agents autonomously learn and transfer skills across tasks? When do multi-agent approaches outperform single model extended thinking? How should inference compute be adaptively allocated based on prompt difficulty? Can single-axis benchmarks accurately predict agent deployment success? What drives capability and cost efficiency in agent systems? How do multi-agent systems achieve genuine cooperation and reasoning?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

deep research and agentic systems topic map