OpenAI in Practice: Using Skills to Speed Up Open Source Maintenance, 45% More PRs Merged

·Toolin Editorial Team

The OpenAI team overhauled Agents SDK maintenance with Codex Skills — via AGENTS.md, local skills, and GitHub Actions, merged PRs climbed from 316 to 457 in three months, a 45% increase.

OpenAI in Practice: Using Skills to Speed Up Open Source Maintenance, 45% More PRs Merged

The OpenAI team used Codex to overhaul how the Agents SDK repositories are maintained, turning repetitive engineering work into repeatable workflows through repo-local Skills, AGENTS.md files, and GitHub Actions.

The results are striking: from December 1, 2025 to February 28, 2026, the two repositories merged 457 PRs in total, versus only 316 in the previous three months (September 1 to November 30, 2025) — a 45% increase.

  • Python repository: 182 → 226 PRs
  • TypeScript repository: 134 → 231 PRs

The setup is remarkably simple, yet the effect is obvious. If you maintain open source projects, this approach is worth borrowing.

Background: High-Traffic SDK Repositories

The OpenAI Agents SDK ships in Python and TypeScript versions, providing the core components for building agent applications, and it is also a clean way to build voice agents on top of the Realtime API.

Usage scale:

  • Python package (PyPI): roughly 14.7 million downloads in the last 30 days
  • TypeScript package (npm): roughly 1.5 million downloads in the last 30 days

That level of activity means a flood of PRs, Issues, and maintenance work. Under the traditional approach, repetitive engineering tasks eat up enormous amounts of time.

Four-layer Skills system architecture

Core Setup: A Four-Layer Architecture

The whole system is very lean:

  1. Policy layer: repository policies live in AGENTS.md
  2. Skill layer: repo-local skills sit in the .agents/skills/ directory
  3. Script layer: skills can internally include scripts and reference material
  4. CI layer: when the same workflow needs to run in CI, use the Codex GitHub Action

This setup gives Codex a stable context for how the repository operates, making repetitive engineering work faster and more accurate.

The Progressive Disclosure Model of Skills

Skills use a progressive disclosure model:

  1. First you only see metadata such as name and description
  2. When selected, the full content of SKILL.md is loaded
  3. When needed, reference material is read or scripts are run

This way the agent's context isn't bloated from the start, while still carrying rich instructions, scripts, and reference material.

Progressive disclosure model of skills

Python Repository: 8 Core Skills

The Python repository is the leaner baseline version, containing 8 skills:

1. code-change-verification Whenever code or build behavior changes, run the required formatting, lint, type-check, and test pipelines.

2. docs-sync Audit the documentation against the codebase to find missing, incorrect, or outdated docs. Treat docstrings and comments in the source code as the authoritative source for generated reference documentation.

3. examples-auto-run Run the examples in automatic mode, producing logs and retry helper files.

4. final-release-review Compare the previous release tag against the current release candidate to check release readiness.

5. implementation-strategy Before touching the runtime or making API changes, pin down the compatibility boundaries and the implementation plan.

6. openai-knowledge Pull the latest OpenAI API and platform docs through the official Docs MCP workflow.

7. pr-draft-summary Prepare branch name suggestions, a PR title, and a draft description at handoff time.

8. test-coverage-improver Run coverage checks, find the biggest gaps, and propose high-impact test suggestions.

JavaScript Repository: 3 Additional Exclusive Skills

The JavaScript repository follows the same overall pattern, adding 3 repo-specific skills for its npm monorepo and release process:

1. changeset-validation Check whether changesets and version-bump levels genuinely match the package diffs.

2. integration-tests Publish the packages to a local Verdaccio (local npm package registry) registry and verify installation and runtime behavior across the supported runtimes.

3. pnpm-upgrade Coordinate updates to the pnpm toolchain and the version pinning in CI.

Key Design Principles

1. Contracts with clear responsibilities Every skill has a crisp trigger condition and a concrete output.

2. "Report first, then act" workflows docs-sync and test-coverage-improver are "report first, then act" workflows: first inspect the current diffs or coverage artifacts, prioritize them, then ask for approval before editing.

3. Narrow, purpose-built skills The JavaScript-only pnpm-upgrade skill is deliberately narrow: it coordinates pnpm version upgrades and does nothing else.

4. Workflows live in the repository Put these workflows next to the code instead of hiding them in some document or wiki.

  • Python repository: .agents/skills/
  • JavaScript repository: .agents/skills/

A Real-World Case: The PR Review Process

Take PR review as an example. The traditional process requires:

  1. Manually checking code formatting
  2. Running lint and type checks
  3. Running the test suite
  4. Checking whether docs were updated
  5. Checking test coverage
  6. Writing the PR description

Now, Codex can:

  1. Invoke the code-change-verification skill to validate automatically
  2. Invoke the docs-sync skill to check the docs
  3. Invoke the test-coverage-improver skill to analyze coverage
  4. Invoke the pr-draft-summary skill to generate the PR description

The whole process is automated; maintainers only need to review the results.

How to Get Started?

1. Apply for Codex for Open Source Eligible open source project maintainers can apply for:

  • ChatGPT Pro with Codex
  • API credits
  • Conditional access to Codex Security

Application address: https://openai.com/codex-for-open-source

2. Create AGENTS.md Create AGENTS.md at the repository root and write in the repository policies and workflow instructions.

3. Create the .agents/skills/ directory Create the .agents/skills/ directory in your repo, with one subdirectory per skill.

4. Write the skill manifests Each skill contains:

  • SKILL.md: the manifest file describing the skill's responsibilities, trigger conditions, and outputs
  • scripts/: optional scripts directory
  • references/: optional reference-material directory
  • assets/: optional assets directory

5. Configure GitHub Actions Use the Codex GitHub Action to run the skills in CI.

Toolin's Take

Who it's for?

  • Teams maintaining high-traffic open source projects
  • Projects with lots of repetitive engineering tasks
  • Projects that need a standardized PR review process

Core strengths:

  • Simple setup, obvious effect (45% more PRs merged)
  • Progressive disclosure model, so the context doesn't balloon
  • Workflows live in the repository, easy to maintain and hand down
  • The same skill set can be reused locally and in CI

Limitations:

  • Requires applying for Codex for Open Source eligibility
  • Upfront time investment to write the skill manifests
  • Possibly over-engineered for small projects

Summary: If the open source project you maintain involves a lot of repetitive engineering work, this approach is worth trying. The OpenAI team proved the effect with hard numbers: a 45% increase in PRs merged over three months, on a remarkably simple setup.

The key is to "solidify" your workflows into the repository, giving the AI a stable context to rely on.

Related articles

Baidu DuMate, A Practical Guide: From Installation to Office Automation
AI Tutorials

Baidu DuMate, A Practical Guide: From Installation to Office Automation

A full walkthrough of DuMate, the general-purpose office agent from Baidu, covering installation, skills, app connections, and automation — get up and running in 3 minutes and hand your daily office chores to AI.

Toolin Editorial Team
DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories
AI Products

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories

Renmin University's Gaoling School has released the DeNovoSWE dataset — 4818 real task instances that train Code Agents to generate complete repositories from documentation, lifting Qwen3-30B from 5.8% to 47.2% on BeyondSWE-Doc2Repo.

Toolin Editorial Team
Doubao Seed 2.1 Pro, Hands-On: Coding Enters the Top Tier, with Multimodal Surprises
AI Products

Doubao Seed 2.1 Pro, Hands-On: Coding Enters the Top Tier, with Multimodal Surprises

A hands-on review of ByteDance's Doubao Seed 2.1 Pro: agent coding and multimodal capability have crossed the production-ready line, including rebuilding front-end interactions from screenshots, at a price nearly 80% lower than Claude Opus 4.6.

Toolin Editorial Team
Hyper3D Rodin Gen-2.5: A Million Polygons in 4 Seconds as Thinking Comes to 3D Generation
AI Products

Hyper3D Rodin Gen-2.5: A Million Polygons in 4 Seconds as Thinking Comes to 3D Generation

Deemos has released Hyper3D Rodin Gen-2.5, the first to bring an LLM-like Thinking mechanism to 3D generation — million-polygon models in 4 seconds, 10-million-polygon precision, and native 12K texturing.

Toolin Editorial Team
Hands-On with WeChat's "Xiaowei" AI Assistant: 12 Entry Points Covering Chat, Content, and Documents
AI Products

Hands-On with WeChat's "Xiaowei" AI Assistant: 12 Entry Points Covering Chat, Content, and Documents

WeChat's native AI assistant Xiaowei is in gray testing. Its main model is the in-house WeLM; it can search chat history, summarize official-account articles, and invoke local-life services, with a second confirmation required for sensitive operations.

Toolin Editorial Team
DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories
AI Products

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories

The Gaoling School of AI at Renmin University of China has released DeNovoSWE, the first long-horizon training set for generating complete repositories from documents, with 4818 real task instances; Qwen3-30B improved from 5.8% to 47.2% on BeyondSWE-Doc2Repo.

Toolin Editorial Team