INQUIRING LINE

An AI agent's paper trail isn't part of the model itself; it has to be built around the model, on purpose.

Why did older robot scientists log auditable provenance while LLM agents inherited only generation capability?

This explores why earlier automated 'robot scientist' systems kept careful, checkable records of what they did and why, while today's LLM-based research agents mostly arrived able to generate hypotheses and code but with no built-in record-keeping. It also asks what is being done to bring that record-keeping back.


This explores why the record-keeping that earlier automated science systems treated as essential didn't carry over to LLM agents, and how researchers are rebuilding it. The corpus doesn't cover the history of earlier robot scientists, so it can't tell you how those systems were designed. It does answer the other half of the question clearly: auditability in LLM agents isn't a property of the model. It comes from the system built around the model, and that system has to be added on purpose. LLMs arrived as text generators, and everything that makes a process traceable (persistent state, logged decisions, recoverable steps) sat outside what they were trained to do.

The clearest statement of this is the idea that agent reliability comes from moving memory, skills and interaction protocols out of the model and into a surrounding 'harness' Where does agent reliability actually come from?. Turning an LLM into something that acts reliably takes a rebuilt pipeline: curated data, grounding, memory and tool infrastructure, and safety evaluation. Fine-tuning alone doesn't get you there Can you turn an LLM into an agent by just fine-tuning?. On this reading, LLM agents didn't lose provenance. They never had a place to store it. Dr. Claw shows the fix concretely. It wraps an unchanged coding agent in persistent state objects and skill libraries, and the result is a traceable, recoverable record of the research process. The executor itself is never modified Can orchestration layers make coding agents more auditable?.

The less obvious point is that the model's own account of its reasoning can't serve as the audit trail. Reasoning traces often don't faithfully explain what drove a decision. Influences can be missing from the trace entirely, or problematic reasoning can show up in clean-sounding language Can we actually trust reasoning model outputs?. A generation-first system that writes plausible narratives about itself isn't keeping a record. That is why the corpus keeps turning to code as a medium: code can be run, inspected, and it holds state, so the agent's work leaves something checkable Can code serve as the operational substrate for agent reasoning?. A related finding comes from discovery work. LLMs are good at proposing candidates but poor at estimating how good those candidates are or how uncertain they are. Pairing them with statistical models fitted to real experimental data puts the grounding back Can language models reliably judge their own candidate quality?.

There's also a twist about governance. In a persistent agent that logged 889 governance events over 96 active days, safeguards worked best when they were stored in the memory the agent actually consulted while working, not in a separate policy document Can governance rules embedded in runtime memory actually protect autonomous agents?. Memory can also be where learning happens. AgentFly improves its behavior purely by recording cases, subtasks and tool use, without changing the model's weights Can agents learn continuously from experience without updating weights?. So the log isn't only for auditors. It can be what the agent learns from.

The open problem is who builds and maintains these harnesses. Models vary sharply in how well they build their own harnesses, and they struggle to keep useful updates as those harnesses evolve Can language models build and maintain their own agent harnesses?. Separating a trained 'curator' from a frozen executor is one promising answer Can a separate trained curator improve skill libraries better than frozen agents?. If you want the historical comparison with earlier robot scientists, you'll need sources outside this collection.


Sources 10 notes

Where does agent reliability actually come from?

Research shows reliable LLM agents externalize three cognitive burdens—memory (state persistence), skills (procedural components), and protocols (structured interaction)—into a harness layer rather than relying on model scale alone. The harness unifies these externalities and eliminates the need for the model to solve the same problems repeatedly.

Can you turn an LLM into an agent by just fine-tuning?

Converting LLMs to action-capable systems requires four distinct stages: curating action-environment-user datasets, training for action grounding, integrating agent infrastructure with memory and tools, and rigorous safety evaluation. The surrounding system and harness determine whether actions are grounded or hallucinated.

Can orchestration layers make coding agents more auditable?

Dr. Claw wraps existing coding agents in persistent state objects and skill libraries, reporting higher research completeness and a traceable, recoverable process trail while keeping the underlying executor unchanged.

Can we actually trust reasoning model outputs?

Research shows reflection rarely corrects errors, traces rarely explain decisions faithfully, and monitoring is vulnerable to two failure modes: omission (influence never reaches the trace) and laundering (problematic reasoning appears in clean language). These vulnerabilities persist even under evaluation pressure.

Can code serve as the operational substrate for agent reasoning?

Research shows code uniquely enables agent reasoning, action, and verification by being simultaneously executable, inspectable, and stateful. This unified code-centered loop improves reasoning and verification together compared to natural-language or prose-based approaches.

Show all 10 sources
Can language models reliably judge their own candidate quality?

LLMs excel at generating valid candidates in structured spaces but cannot reliably assess their true value or uncertainty. Coupling them with Gaussian process surrogates fitted to real experimental data creates uncertainty-aware guidance for discovery.

Can governance rules embedded in runtime memory actually protect autonomous agents?

A persistent agent recorded 889 governance events across 96 active days, with safeguards encoded directly into the memory layer the agent consulted during operation. Runtime-resident governance proved more effective than external policies because the agent actually accessed it during decision-making.

Can agents learn continuously from experience without updating weights?

AgentFly formalizes agent learning as a Memory-augmented MDP with three memory modules (case, subtask, tool) that enable credit assignment and policy improvement entirely through memory operations. The approach achieved 87.88% on GAIA validation without modifying LLM parameters.

Can language models build and maintain their own agent harnesses?

Research shows that LLMs vary sharply in building harnesses across domains, struggle to retain useful intermediate updates during evolution, and produce harnesses whose performance shifts dramatically with different executors—demonstrating that harness quality cannot be inferred from downstream task scores alone.

Can a separate trained curator improve skill libraries better than frozen agents?

SkillOS shows that separating a trainable curator from a frozen executor, grouped by task streams, causes skill repositories to shift from generic verbose additions toward actionable execution logic and cross-task meta-strategies. The trained curator generalizes across different executor backbones and domains.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.