Equipping agents for the real world with Agent Skills

Paper · Source
Foundation ModelsLLM AgentsMulti-Agent ArchitecturesTool Use and Computer-Use Agents

Published Oct 16, 2025 Claude is powerful, but real work requires procedural knowledge and organizational context. Introducing Agent Skills, a new way to build specialized agents using files and folders.

Introduction. As model capabilities improve, we can now build general-purpose agents that interact with full-fledged computing environments. Claude Code, for example, can accomplish complex tasks across domains using local code execution and filesystems. But as these agents become more powerful, we need more composable, scalable, and portable ways to equip them with domain-specific expertise. This led us to create Agent Skills: organized folders of instructions, scripts, and resources that agents can discover and load dynamically to perform better at specific tasks. Skills extend Claude’s capabilities by packaging your expertise into composable resources for Claude, transforming general-purpose agents into specialized agents that fit your needs.

Discussion / Conclusion. Skills are a simple concept with a correspondingly simple format. This simplicity makes it easier for organizations, developers, and end users to build customized agents and give them new capabilities. We’re excited to see what people build with Skills. Get started today by checking out our Skills docs and cookbook.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What are the consequences of models training on synthetic data? Can AI-generated outputs constitute genuine knowledge or valid claims? Do language models learn genuine linguistic structure or just surface patterns? How does AI-generated content transformation affect public discourse quality? What makes AI persuasion effective and how can we counter it? How can AI alignment serve diverse human preferences at scale? How can language models sustain linguistic synchrony and intersubjectivity during dialogue? When does optimizing for quality undermine the value of diversity? Why does reinforcement learning suppress output diversity compared to supervised fine-tuning? Does alignment training create blind spots in detecting genuine safety threats? How do language models inherit human biases from training data? How do multi-agent systems achieve genuine cooperation and reasoning? What determines success in training models on multiple tasks? What factors beyond surface content determine how readers extract meaning differently?