Darwin Skill 2.0: Let Your AI Skills Evolve on Their Own

·Toolin Editorial Team

darwin-skill 2.0 is an open-source Skill/Prompt auto-optimization tool that distills the best of two Microsoft papers, using multi-judge independent review plus human checkpoints to lift your AI skill documents from 80 points to over 90.

Darwin Skill 2.0: Let Your AI Skills Evolve on Their Own

If you maintain multiple AI Skills or Prompts, you know how painful manual review is: every Skill has to be read end to end, problems found, edits made, then read again — it is simply not sustainable in human hours. But left alone, a Skill slowly "drifts" and keeps getting worse.

darwin-skill 2.0 is an open-source Skill auto-optimization tool. What it does is simple: score your Skill, propose improvements, re-score after the edit, and roll back if the score did not go up. The whole flow works like biological evolution — generation after generation of mutation, selection, and elimination, and what survives is always the stronger version.

From 1.0 to 2.0: What Two Microsoft Papers Changed

The 1.0 era ran for a month, gaining an average of 13.5 points with 0 rollbacks. Sounds good, but the author realized that 0 rollbacks does not fully prove the algorithm is precise — the results come out exactly as strict as you set the scoring criteria.

The turning point was two papers Microsoft Research posted on the same day, May 22, 2025:

  • SkillLens (arXiv 2605.23899): Studies how Skills should be evaluated. Core finding — letting a single AI score a Skill is only 46.4% accurate, worse than a coin flip. But after adding three key dimensions to the scoring criteria, accuracy can rise to 73.8%.
  • SkillOpt (arXiv 2605.23904): Studies how Skills should be optimized. Core idea — treat the Skill document as the neural network's "externally trainable state" and optimize it through backpropagation.

2.0 directly absorbed the essence of both papers and shipped four core upgrades.

The Four Core Upgrades in 2.0

1. Scoring criteria upgraded from 8 to 9 dimensions

It directly adopts the "prescription" from the SkillLens paper that pulled accuracy from 46.4% to 73.8%:

  • Failure mode encoding: You must spell out explicit branches like "if X happens, do Y; otherwise do Z" — writing only the happy path is not allowed
  • Executable specificity: Softeners like "suggest," "could consider," and "it depends" are explicitly banned, and more than three occurrences lose points
  • High-risk action blacklist: Every Skill must have a standalone "things you must absolutely never do" section

2. Independent multi-judge review

No more relying on a single AI judge. Each round spins up two independent judges (neither knows the other exists), and only the consensus score counts. The next round brings in two brand-new judges to avoid anchoring effects.

If scores enter a plateau (single-round gain < 1 point), the early-stop mechanism halts automatically.

3. Human in the loop checkpoints

This is Darwin's biggest difference from SkillOpt. SkillOpt is a benchmark-driven, fully automated pipeline suited to enterprise scenarios. But for individual developers, the benchmark itself is hard to define — subjective dimensions like "does this read smoothly to me" cannot be crammed into an automated loop.

Darwin 2.0 sets explicit human intervention points at every stage:

  • Stage one: Judges run and score automatically; a human reviews the report and decides what to change
  • Stage two: The lowest dimension is fixed automatically; a CHECKPOINT forces a pause to wait for user confirmation
  • Stage three: New judges re-evaluate; if the gain is below the threshold, it forcibly stops

4. Counter-example blacklist

It adds 8 anti-patterns drawn from 40 real-world optimization runs, including: the same AI both editing and judging, using git reset --hard as the rollback mechanism, stuffing in redundancy to pad the score, scoring without running tests, and changing multiple dimensions in one round.

Darwin 2.0 workflow

Real-World Results

Tested on a real 368-line Skill (huashu-gpt-image):

StageScoreChange
Baseline80.8Consensus of two independent judges
Round 191.5+10.7, only failure modes changed
Round 291.65+0.15, early stop triggered

Key finding: only the "failure mode" dimension was changed, yet the "workflow" dimension jumped from 7.5 to 9.0 — because failure modes demand explicit branches, and once they are written out, the flow clarifies itself. This is called a "correlated dimension cluster".

Validation at Scale

2.0 scanned the entire Skill library, nearly 30 Skills in total. Each ran two rounds of independent judges, 9-dimension scoring, and validation-gated rollback:

  • steve-jobs-perspective: 64 -> 94 (+30, done in a single round)
  • huashu-weread-advisor: 80+ -> 91.4
  • darwin-skill itself: 86.05 -> 92.7 (self-referential evaluation)
  • Multiple content Skills all landed in the 90+ range

Average gain: +15 points. Every Skill's optimization has a complete git commit chain to trace back through.

Batch optimization results

How to Use It

Drop the repo link to your Agent and have it install for you:

Install this Skill for me: https://github.com/alchaincyf/darwin-skill

Then just tell your Agent "run Darwin optimization on XX skill". One baseline evaluation plus one optimization round takes about 15-30 minutes, with most of the time spent waiting for the judge Agents to return.

Who It Is For

  • Developers maintaining multiple Skills / Prompts
  • Teams that care about Prompt quality but have no time to review each one
  • Individual creators who want steadier output from AI tools

Related articles

Doubao Seed 2.1 Pro, Hands-On: Coding Enters the Top Tier, with Multimodal Surprises
AI Products

Doubao Seed 2.1 Pro, Hands-On: Coding Enters the Top Tier, with Multimodal Surprises

A hands-on review of ByteDance's Doubao Seed 2.1 Pro: agent coding and multimodal capability have crossed the production-ready line, including rebuilding front-end interactions from screenshots, at a price nearly 80% lower than Claude Opus 4.6.

Toolin Editorial Team
Hyper3D Rodin Gen-2.5: A Million Polygons in 4 Seconds as Thinking Comes to 3D Generation
AI Products

Hyper3D Rodin Gen-2.5: A Million Polygons in 4 Seconds as Thinking Comes to 3D Generation

Deemos has released Hyper3D Rodin Gen-2.5, the first to bring an LLM-like Thinking mechanism to 3D generation — million-polygon models in 4 seconds, 10-million-polygon precision, and native 12K texturing.

Toolin Editorial Team
Hands-On with WeChat's "Xiaowei" AI Assistant: 12 Entry Points Covering Chat, Content, and Documents
AI Products

Hands-On with WeChat's "Xiaowei" AI Assistant: 12 Entry Points Covering Chat, Content, and Documents

WeChat's native AI assistant Xiaowei is in gray testing. Its main model is the in-house WeLM; it can search chat history, summarize official-account articles, and invoke local-life services, with a second confirmation required for sensitive operations.

Toolin Editorial Team
DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories
AI Products

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories

The Gaoling School of AI at Renmin University of China has released DeNovoSWE, the first long-horizon training set for generating complete repositories from documents, with 4818 real task instances; Qwen3-30B improved from 5.8% to 47.2% on BeyondSWE-Doc2Repo.

Toolin Editorial Team
Hands-On with Baidu's DuMate: A Tutorial for the Homegrown Office Agent, from Installation to Automation
AI Tutorials

Hands-On with Baidu's DuMate: A Tutorial for the Homegrown Office Agent, from Installation to Automation

A step-by-step walkthrough of Baidu DuMate's request → authorize → execute → deliver pipeline — installation, interface, hands-on office scenarios, and automation, all covered in one article

Toolin Editorial Team
Claude Tag: Making AI a True Member of Your Team
AI Products

Claude Tag: Making AI a True Member of Your Team

Anthropic has launched Claude Tag, embedding Claude into Slack workflows with shared context, persistent memory, and proactive intervention — a new paradigm for team collaboration.

Toolin Editorial Team