SkillOpt: Train Your Agent Skill Docs Like a Neural Network
Microsoft's open-source text-space optimization framework that lets agent skill documents evolve automatically, ranking best or tied-best across all 52 evaluation combinations.


SkillOpt: Train Your Agent Skill Docs Like a Neural Network
Microsoft's open-source text-space optimization framework that lets agent skill documents evolve automatically, ranking best or tied-best across all 52 evaluation combinations.
Have you noticed that writing agent skill documents (CLAUDE.md, Codex skill files, system prompts) is essentially a manual process of trial and error? Write a version, run a few tasks, revise when results disappoint, then run again. SkillOpt, open-sourced by Microsoft, automates this process — it treats the skill document as "trainable parameters" and optimizes your agent skill docs with a loop much like training a neural network.
Across all 52 evaluation combinations spanning 7 models, 6 benchmarks, and 3 execution environments, skill documents trained by SkillOpt achieved the best or tied-best result every single time. The GitHub repo picked up 3.3k stars within a week of launch.
What Is SkillOpt
SkillOpt is Microsoft's open-source text-space optimization framework. The core idea: instead of training model weights, train only the natural-language skill document that guides agent behavior. Treat the skill document as the agent's "external weights" — since internal weights can be optimized with gradient descent, external weights deserve a systematic training method too.

Key resources:
- Official site: https://microsoft.github.io/SkillOpt/
- GitHub: https://github.com/microsoft/SkillOpt
- Paper: https://arxiv.org/abs/2605.23904
The Training Loop: Four Core Steps
SkillOpt's training loop maps directly onto deep learning's "forward propagation - backpropagation - parameter update," but executes in text space.

Step 1: Rollout (forward propagation)
The frozen target model takes the current version of the skill document and executes a batch of tasks, recording complete execution traces — including messages, tool calls, verification feedback, and final scores. What this step produces is "evidence," the equivalent of a neural network's forward-pass results.
Step 2: Reflect (backpropagation)
A separate optimizer model analyzes the batch of execution traces. The key design choice: failures and successes are reflected on separately. The failed minibatch is used to discover "which operating rules need fixing," while the successful minibatch confirms "which existing rules are working and must not be touched." This is the equivalent of computing "gradients in text space."
Step 3: Edit (parameter update)
Based on the reflection results, the optimizer model proposes structured edit operations on the skill document: add new rules (add), delete stale rules (delete), and replace rules that need correction (replace).
Step 4: Gate (validation gating)
A candidate new skill document must be run on a held-out validation set, and is accepted only when performance strictly improves. This prevents overfitting and ensures every update is a genuine improvement.
The whole loop runs multiple epochs, with multiple steps inside each epoch — exactly the rhythm of training a neural network.
Two Elegant Training Tricks
Text learning rate
When training a neural network, a learning rate that is too large causes catastrophic forgetting. SkillOpt runs into exactly the same problem in text space — a single overly large edit can wipe out previously learned effective rules.
The solution is a "text learning rate": a cap on the number of edit operations allowed per step, defaulting to lr=4, i.e. at most 4 add/delete/replace operations per step. Ablation studies confirm this design is necessary: without the learning-rate constraint, SearchQA performance drops from 87.1% to 84.6%, and SpreadsheetBench from 77.5% to 75.7%.
Rejected-Edit Buffer (negative-feedback memory)
When an edit proposal is rejected by the validation gate, it is not simply discarded — it goes into a buffer. The optimizer can see these "failed attempts" in later reflection phases, avoiding repeated proposals of similar useless edits. This effectively gives the optimizer negative-gradient information. Remove this buffer, and SpreadsheetBench plummets from 77.5% to 72.9%.

Evaluation Results: Leading Across All 52 Combinations
SkillOpt's evaluation coverage is remarkably comprehensive:
- Target models: GPT-5.5, GPT-5.4, GPT-5.4-mini, GPT-5.4-nano, GPT-5.2, Qwen3.5-4B, Qwen3.6-35B-A3B
- Benchmarks: SearchQA (question answering), SpreadsheetBench (code generation), OfficeQA (tool-augmented QA), DocVQA (document visual QA), LiveMathematicianBench (mathematical reasoning), ALFWorld (embodied agents)
- Execution environments: direct chat, OpenAI Codex, Anthropic Claude Code

A few standout numbers:
| Model + Environment | Benchmark | Gain |
|---|---|---|
| GPT-5.5 direct chat | SpreadsheetBench | +38.9 |
| GPT-5.5 direct chat | OfficeQA | +39.0 |
| GPT-5.4-nano (smallest model) | DocVQA | +49.4 |
| GPT-5.5 + Codex | SpreadsheetBench | +57.5 |
| GPT-5.5 + Claude Code | SpreadsheetBench | +58.3 |
Smaller models actually gain more — a good operations manual helps novices far more than it helps experts.
Transferability: Train Once, Deploy Everywhere
Skill documents trained by SkillOpt show strong transferability:
- Cross-model transfer: a LiveMath skill trained on GPT-5.4 transfers directly to GPT-5.4-nano for a 15.2-point gain
- Cross-environment transfer: a SpreadsheetBench skill trained in the Codex environment transfers directly to the Claude Code environment for a 31.8-point gain
- Self-optimization: even with GPT-5.4-nano serving as both target model and optimizer (optimizing itself), SpreadsheetBench still improves by 10.4 points
Deployment is minimal: in the end you only need a single best_skill.md file — no optimizer model, no memory module, no extra inference overhead whatsoever.
How to Use It
SkillOpt's workflow boils down to these steps:
- Prepare the target model and task set: pick the agent model you want to optimize (e.g. GPT-5.4) and prepare a set of tasks with validation functions
- Write the initial skill document: it can be a rough manual draft, or even an empty document
- Configure the optimizer model: choose a model to act as the optimizer (usually a stronger model than your target model)
- Set training parameters: including the text learning rate (default lr=4), number of epochs, and steps per epoch
- Start the training loop: run the Rollout -> Reflect -> Edit -> Gate cycle
- Get the final skill document: once training completes, it outputs
best_skill.md, ready to deploy directly into your agent
See the GitHub repo's documentation and examples for detailed usage.
Who Should Use It
- Agent developers: anyone maintaining agent config files such as CLAUDE.md or Codex skill files
- Prompt engineers: practitioners who need to systematically optimize long-document prompts
- AI application teams: teams that want to improve agent performance without swapping the underlying model
SkillOpt leaves us with a key insight: everything about an agent can self-learn — including the very skill document that governs its behavior.
Toolin Editorial Team
Categories
Related articles

GPT-5.6 Is Here: Three New Model Tiers + ChatGPT Work, Codex Merged into the Desktop App
OpenAI has launched the GPT-5.6 family (Sol/Terra/Luna) and the ChatGPT Work agent, and merged Codex into the ChatGPT desktop app — with full pricing and availability details.

ChatGPT's Big Voice Upgrade: GPT-Live Full-Duplex Mode Goes Live
OpenAI has released the GPT-Live full-duplex voice model family. ChatGPT can finally talk while listening, chime in naturally, and know when to stay silent, replacing the original Advanced Voice Mode.

LingBot-World 2.0: An Open-Source, Playable Interactive World Model with Effectively Unlimited Duration
Ant Group's Robbyant has open-sourced LingBot-World-Infinity: hour-level real-time generation at 720p/60fps, with a built-in agent proposing events automatically — turning watching video into entering a world.

Two Open-Source Skills: Have Codex/Claude Code Make Your Illustrations and Web Designs
The open-source guizang-material-illustration article-illustration skill and web-design-taste web design skill install into Codex with one command — goodbye to AI-flavored illustrations and cookie-cutter web pages.

Tencent Hunyuan Hy3 Official Release: A Homegrown Coding Agent Model Matching GLM-5.2
The official Tencent Hunyuan Hy3 is live: 295B MoE architecture, agent task success rate up to 90%, API price of ¥1 per million tokens, with a limited-time free trial inside WorkBuddy.

Vidu S1: The Open-Class Interactive Video Model That Lets Digital Humans Talk in Real Time
Shengshu Technology's Vidu S1 achieves real-time voice-driven video generation at 540P/25FPS, runs on consumer GPUs, and turns a single image into a conversable digital human.