Hermes Agent: The Open-Source Python Project That Beat OpenAI Codex

·Toolin Editorial Team

Hermes Agent cut its startup time by 63% through three engineering optimizations and beat the Rust-written OpenAI Codex 6:5 across 11 CLI benchmarks, with GitHub stars passing 160,000.

Hermes Agent: The Open-Source Python Project That Beat OpenAI Codex

An open-source Agent framework written in pure Python has beaten OpenAI's Rust-written Codex in CLI benchmarks. That is not a stunt — it is a real 6:5 record.

Hermes Agent is currently the fastest-growing open-source Agent framework on GitHub. Launched in February 2026, it is only three months in: stars have passed 160,000, and daily active Token consumption has hit 353B, nearly double that of comparable projects.

If you are choosing an AI coding Agent, Hermes deserves serious consideration.

What Hermes Agent Is

Hermes is an open-source Agent framework from NousResearch — think of it as an open-source Claude Code or Codex CLI. It runs right in your terminal, reading your project, understanding context, planning changes, and editing code files.

Its killer feature is a closed-loop learning architecture: after each complex task, the Agent automatically distills the solution into a reusable Skill (skill). The next time a similar task shows up, it calls the existing skill directly and skips reasoning from scratch. Official numbers show that instances with more than 20 self-created skills complete similar tasks 40% faster than fresh instances.

Even better is the autonomous Curator introduced in v0.12, a background Agent that runs automatically and periodically scores, prunes, and merges your skill library. Hermes does not just learn — it also manages what it learns.

How It Beat Codex: Three Cuts That Took 63% off Startup Time

Before the optimizations, Hermes's record against Codex was 5:6. Afterward it flipped straight to 6:5. And this reversal came not from swapping models or piling on compute, but from three purely engineering optimizations.

Cut one: a Bitwarden disk cache

Previously, every startup called the Bitwarden Secrets Manager API to fetch credentials, at 380 milliseconds a pop. And the cache was purely in-process, so running twice back to back still meant re-fetching.

The fix: add an L2 disk cache. The cache file's permissions are locked down to 0600, stored at /cache/bws_cache.json, with a default TTL of 300 seconds. The access token itself never touches disk. One cut takes off 380ms.

Cut two: lazy-loading the model catalog

hermes_cli.models._PROVIDER_MODELS is a giant dictionary holding every AI provider's model information, and it used to be eagerly imported at module load, eating about 55ms.

The team used PEP 562's module-level getattr to make it lazy, paying that cost only when the model catalog is actually accessed. Another 55ms saved.

Cut three: deduplicating config reads

The top of main.py originally read config.yaml twice — one yaml.safe_load for masking secrets, and one full load_config() just to check a single boolean. They were merged into a single raw load. 17ms saved.

Add the three cuts together and startup time plunged from 701ms to 258ms, a 63% reduction.

Why Python Can Beat Rust

The result looks counterintuitive, but the logic behind it is direct: in the Agent race, architectural decisions at the framework level matter more than raw speed at the language level.

A single LLM call routinely costs hundreds of milliseconds or even seconds. The 443 milliseconds Hermes optimized away is already the limit of what the framework layer can squeeze out. What really shapes the Agent experience is architecture design, not interpreted versus compiled.

Hermes co-founder and chief scientist Teknium put his finger on it: migrate to Rust and "you can no longer edit the code, or improve and iterate in real time". Python's advantage is not being fast — it is being alive -- developer friendliness and iteration speed are the biggest performance advantages.

How to Get Started

# GitHub URL
git clone https://github.com/nousresearch/hermes-agent

# Configure your API Key per the README and you are ready to go

Hermes supports a variety of underlying models, including SkyClaw-v1.0, the Qwen series, Claude, and GPT. Pick based on your own budget and needs.

Best-Fit Scenarios

  • Daily coding assistance: An AI programmer in your terminal — reads code, edits code, runs tests
  • Complex project refactoring: Coordinated multi-file edits, long-chain reasoning
  • Skill accumulation and reuse: Automated Skills keep piling up as you use it, and efficiency keeps climbing
  • Budget-constrained developers: Open source and free, pairable with low-cost models

GitHub: https://github.com/nousresearch/hermes-agent