Darwin Skill 2.0: Let Your AI Skills Evolve on Their Own
darwin-skill 2.0 is an open-source Skill/Prompt auto-optimization tool that distills the best of two Microsoft papers, using multi-judge independent review plus human checkpoints to lift your AI skill documents from 80 points to over 90.


Darwin Skill 2.0: Let Your AI Skills Evolve on Their Own
darwin-skill 2.0 is an open-source Skill/Prompt auto-optimization tool that distills the best of two Microsoft papers, using multi-judge independent review plus human checkpoints to lift your AI skill documents from 80 points to over 90.
If you maintain multiple AI Skills or Prompts, you know how painful manual review is: every Skill has to be read end to end, problems found, edits made, then read again — it is simply not sustainable in human hours. But left alone, a Skill slowly "drifts" and keeps getting worse.
darwin-skill 2.0 is an open-source Skill auto-optimization tool. What it does is simple: score your Skill, propose improvements, re-score after the edit, and roll back if the score did not go up. The whole flow works like biological evolution — generation after generation of mutation, selection, and elimination, and what survives is always the stronger version.
- Open source: https://github.com/alchaincyf/darwin-skill
- License: MIT
From 1.0 to 2.0: What Two Microsoft Papers Changed
The 1.0 era ran for a month, gaining an average of 13.5 points with 0 rollbacks. Sounds good, but the author realized that 0 rollbacks does not fully prove the algorithm is precise — the results come out exactly as strict as you set the scoring criteria.
The turning point was two papers Microsoft Research posted on the same day, May 22, 2025:
- SkillLens (arXiv 2605.23899): Studies how Skills should be evaluated. Core finding — letting a single AI score a Skill is only 46.4% accurate, worse than a coin flip. But after adding three key dimensions to the scoring criteria, accuracy can rise to 73.8%.
- SkillOpt (arXiv 2605.23904): Studies how Skills should be optimized. Core idea — treat the Skill document as the neural network's "externally trainable state" and optimize it through backpropagation.
2.0 directly absorbed the essence of both papers and shipped four core upgrades.
The Four Core Upgrades in 2.0
1. Scoring criteria upgraded from 8 to 9 dimensions
It directly adopts the "prescription" from the SkillLens paper that pulled accuracy from 46.4% to 73.8%:
- Failure mode encoding: You must spell out explicit branches like "if X happens, do Y; otherwise do Z" — writing only the happy path is not allowed
- Executable specificity: Softeners like "suggest," "could consider," and "it depends" are explicitly banned, and more than three occurrences lose points
- High-risk action blacklist: Every Skill must have a standalone "things you must absolutely never do" section
2. Independent multi-judge review
No more relying on a single AI judge. Each round spins up two independent judges (neither knows the other exists), and only the consensus score counts. The next round brings in two brand-new judges to avoid anchoring effects.
If scores enter a plateau (single-round gain < 1 point), the early-stop mechanism halts automatically.
3. Human in the loop checkpoints
This is Darwin's biggest difference from SkillOpt. SkillOpt is a benchmark-driven, fully automated pipeline suited to enterprise scenarios. But for individual developers, the benchmark itself is hard to define — subjective dimensions like "does this read smoothly to me" cannot be crammed into an automated loop.
Darwin 2.0 sets explicit human intervention points at every stage:
- Stage one: Judges run and score automatically; a human reviews the report and decides what to change
- Stage two: The lowest dimension is fixed automatically; a CHECKPOINT forces a pause to wait for user confirmation
- Stage three: New judges re-evaluate; if the gain is below the threshold, it forcibly stops
4. Counter-example blacklist
It adds 8 anti-patterns drawn from 40 real-world optimization runs, including: the same AI both editing and judging, using git reset --hard as the rollback mechanism, stuffing in redundancy to pad the score, scoring without running tests, and changing multiple dimensions in one round.

Real-World Results
Tested on a real 368-line Skill (huashu-gpt-image):
| Stage | Score | Change |
|---|---|---|
| Baseline | 80.8 | Consensus of two independent judges |
| Round 1 | 91.5 | +10.7, only failure modes changed |
| Round 2 | 91.65 | +0.15, early stop triggered |
Key finding: only the "failure mode" dimension was changed, yet the "workflow" dimension jumped from 7.5 to 9.0 — because failure modes demand explicit branches, and once they are written out, the flow clarifies itself. This is called a "correlated dimension cluster".
Validation at Scale
2.0 scanned the entire Skill library, nearly 30 Skills in total. Each ran two rounds of independent judges, 9-dimension scoring, and validation-gated rollback:
- steve-jobs-perspective: 64 -> 94 (+30, done in a single round)
- huashu-weread-advisor: 80+ -> 91.4
- darwin-skill itself: 86.05 -> 92.7 (self-referential evaluation)
- Multiple content Skills all landed in the 90+ range
Average gain: +15 points. Every Skill's optimization has a complete git commit chain to trace back through.

How to Use It
Drop the repo link to your Agent and have it install for you:
Install this Skill for me: https://github.com/alchaincyf/darwin-skillThen just tell your Agent "run Darwin optimization on XX skill". One baseline evaluation plus one optimization round takes about 15-30 minutes, with most of the time spent waiting for the judge Agents to return.
Who It Is For
- Developers maintaining multiple Skills / Prompts
- Teams that care about Prompt quality but have no time to review each one
- Individual creators who want steadier output from AI tools
Toolin Editorial Team
Categories
Related articles

Doubao Seed 2.1 Pro, Hands-On: Coding Enters the Top Tier, with Multimodal Surprises
A hands-on review of ByteDance's Doubao Seed 2.1 Pro: agent coding and multimodal capability have crossed the production-ready line, including rebuilding front-end interactions from screenshots, at a price nearly 80% lower than Claude Opus 4.6.

Hyper3D Rodin Gen-2.5: A Million Polygons in 4 Seconds as Thinking Comes to 3D Generation
Deemos has released Hyper3D Rodin Gen-2.5, the first to bring an LLM-like Thinking mechanism to 3D generation — million-polygon models in 4 seconds, 10-million-polygon precision, and native 12K texturing.

Hands-On with WeChat's "Xiaowei" AI Assistant: 12 Entry Points Covering Chat, Content, and Documents
WeChat's native AI assistant Xiaowei is in gray testing. Its main model is the in-house WeLM; it can search chat history, summarize official-account articles, and invoke local-life services, with a second confirmation required for sensitive operations.

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories
The Gaoling School of AI at Renmin University of China has released DeNovoSWE, the first long-horizon training set for generating complete repositories from documents, with 4818 real task instances; Qwen3-30B improved from 5.8% to 47.2% on BeyondSWE-Doc2Repo.

Hands-On with Baidu's DuMate: A Tutorial for the Homegrown Office Agent, from Installation to Automation
A step-by-step walkthrough of Baidu DuMate's request → authorize → execute → deliver pipeline — installation, interface, hands-on office scenarios, and automation, all covered in one article

Claude Tag: Making AI a True Member of Your Team
Anthropic has launched Claude Tag, embedding Claude into Slack workflows with shared context, persistent memory, and proactive intervention — a new paradigm for team collaboration.