Apodex 1.1: Scaling Agentic Intelligence for Complex Work

Paper · arXiv 2608.23283 · Published August 24, 2026
Agent Harness

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this working capability: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. Environment Scaling expands the diversity and verifiability of executable file, search, and code environments, while Agentic Coordination Scaling trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What coordination failures limit multi-agent LLM systems as they scale? Does externalizing cognitive work and state improve agent reliability? How should systems govern persistent agent-generated code in shared infrastructure? Why do agents confidently report success despite actually failing tasks? Do harness improvements transfer across model scales or memorize shortcuts? What dimensions of recommendation quality do standard metrics miss? What drives capability and cost efficiency in agent systems? Can single-axis benchmarks accurately predict agent deployment success? When should tasks involve human-AI partnership versus full automation? Why do persona-level simulations fail to predict individual preferences accurately?