GPT-5.5-Cyber at Full Power: OpenAI Welds an AI Security Engineer into Codex

·Toolin Editorial Team

OpenAI expands its Daybreak security program with the full release of GPT-5.5-Cyber (CyberGym 85.6%, ahead of Mythos 5), a Codex Security plugin update, and the Patch the Planet open-source remediation effort — pushing vulnerabilities all the way from discovery to landed patches.

GPT-5.5-Cyber at Full Power: OpenAI Welds an AI Security Engineer into Codex

OpenAI has bundled its cybersecurity line into a program called Daybreak, and today it dropped three things at once: a full release of GPT-5.5-Cyber, a Codex Security plugin update, and a Patch the Planet program aimed at the open-source ecosystem. The core narrative comes down to one sentence — AI can now help defenders complete the entire loop "from vulnerability discovery to landed patch," instead of just producing more vulnerability reports.

The full version of GPT-5.5-Cyber scored 85.6% on CyberGym, ahead of Anthropic Mythos 5 at 83.8% and the standard GPT-5.5 at 81.8%. It is the highest single-model CyberGym score OpenAI has measured.

CyberGym 85.6%, ahead of Mythos 5 at 83.8% and GPT-5.5 at 81.8%.

What GPT-5.5-Cyber Is

It is a specialized model for advanced, authorized security work: more capable than the general-purpose version and far less prone to unnecessary refusals. The preview stage mainly tackled "false refusals in professional workflows"; this update goes further, letting it sustain deep analysis across large codebases.

A standard Cyber task loop runs like this:

  1. Identify security-relevant components
  2. Trace whether the vulnerable code is actually reachable
  3. Validate suspected issues in a controlled environment
  4. Develop and test patches
  5. Prepare evidence for human review

GPT-5.5-Cyber's goal is not to generate yet another vulnerability report, but to help defenders complete the entire remediation loop.

Beats GPT-5.5 across all three real-world security benchmarks.

A Clean Sweep Across Three Benchmarks

  • CyberGym (can it reproduce known vulnerabilities in a software environment): Cyber 85.6% vs standard 81.8%
  • ExploitGym (turning known vulnerabilities into working exploits): Cyber 39.5% vs standard 25.95%
  • SEC-bench Pro (long-horizon vulnerability discovery + PoC generation): Cyber 69.8% vs standard 63.1%

Codex Security: A Security Engineer at Every Developer's Side

If GPT-5.5-Cyber is the spear, Codex Security is the shield placed at every developer's side. OpenAI has integrated it directly into the Codex workflow, and the out-of-the-box capability chain looks like this:

  • Understand the team's code and threat model (auto-generating one if none exists)
  • Identify potential vulnerabilities and determine whether the affected code is reachable
  • Collect evidence and provide verification steps
  • Develop targeted patches and verify the fix

Humans stay in control of the key decisions: which findings to investigate, which changes to apply, and what information to share.

Codex Security workflow

Scan an entire codebase, part of one, or a specific change and commit.

The scale numbers speak for themselves: since the cloud preview launched in March, Codex Security has scanned more than 30 million commits across over 30,000 codebases, human reviewers have marked more than 70,000 findings as fixed, and over 500,000 findings have been automatically confirmed as fixed.

The plugin also supports triaging and validating existing findings (from scanners, security advisories, bug bounty reports, or tickets), then auto-generating patches at scale to help teams burn down vulnerability backlogs. Results can be fed into other tools via SARIF files or CodeQL queries, and pipelined with the Codex CLI for automation.

Patch the Planet: Getting Open-Source Fixes Actually Landed

Finding vulnerabilities matters, but what actually protects the world is getting fixes landed. Patch the Planet is co-launched by OpenAI and Trail of Bits, working with HackerOne and Calif to fund professional security researchers, equip them with Codex Security and advanced models, and have them collaborate directly with open-source maintainers.

The program's core is expert-level human security review, because the reality is brutal: among widely used open-source projects, 94% rely on fewer than 10 developers to maintain more than 90% of the code added in a year. AI is making vulnerability discovery faster, which in turn dumps thousands of reports on maintainers — the majority of them low-quality false positives.

Patch the Planet program

The first cohort of 30+ open-source projects includes cURL, Go, Python, Sigstore, and pyca/cryptography.

The working model: researchers complete deduplication and validation before anything reaches maintainers, delivering clean patches instead of dumping noise. Participating projects get conditional access to ChatGPT Pro and Codex Security, plus API credits for core development, automation, and release workflows.

The first five-day sprint has already surfaced hundreds of issues pending review across 19 projects, with dozens of patches merged.

An Awkward Footnote: Codex Logs Hammering the Disk

Worth noting: on the very same day, Codex itself was hit by reports of an "epic bug." Multiple developers reported that during streaming tasks and long-running sessions, Codex writes TRACE logs to the local ~/.codex/logs_2.sqlite at a frantic pace of roughly 5MB/s (peaking at 16MB/s in real measurements) — an estimated 640TB of writes per year, enough to burn through the rated write endurance of a consumer SSD.

And it happens "silently": the file size looks unremarkable because the logs are written and deleted over and over, but the actual write volume hitting the flash far exceeds what is visible. The related issue has existed since April, and the core write-rate problem remains unfixed. For developers who lean hard on Codex, keeping an eye on disk health is a prudent move.

Who It's For

  • Authorized defenders / security teams: GPT-5.5-Cyber targets authorized advanced defensive work and requires application through a trusted access mechanism
  • Development teams: use GPT-5.5 + Codex Security as the everyday starting point, putting security-engineer capability into every developer's workflow
  • Open-source maintainers: Patch the Planet provides expert-level human review + patches, helping small teams withstand the surge of vulnerability reports in the AI era
  • Enterprise security product vendors: integrate the model's capabilities into your own products through the Daybreak Cyber Partner Program (nearly 30 major security companies have joined)

💡 Tip: For most defenders, GPT-5.5 + Trusted Access for Cyber + Codex Security is the right starting point. GPT-5.5-Cyber is aimed at vetted defenders who need stronger capability and more permissive model behavior, paired with stronger verification, monitoring, and scope controls.

Daybreak program portal: openai.com/daybreak (applications for the partner program are now open)

Related articles

Qwen-Image-3.0: Alibaba Qwen's Third-Generation Image Generation Model
AI Products

Qwen-Image-3.0: Alibaba Qwen's Third-Generation Image Generation Model

Up to 4.5k tokens of input, crisp rendering of 10px-scale small text, and 12 languages — renders complex layouts like posters, exam papers, and storyboard comics in a single pass.

Toolin Editorial Team
Build an Automated Reddit Overseas Customer Acquisition Workflow with Qwen 3.8
AI Tutorials

Build an Automated Reddit Overseas Customer Acquisition Workflow with Qwen 3.8

An eight-step fully automated pipeline: Apify scraping plus Qwen 3.8 scoring/drafting/self-review/retrospective — 260 posts filtered down to 22 opportunities, at a tenth of the cost.

Toolin Editorial Team
Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion
AI Products

Qwen-Audio-3.0-TTS: The Speech Synthesis Model That Can Express Emotion

Alibaba's new-generation TTS model controls laughter, gasps, and anger with tags, delivers 48kHz film-grade audio, and tops the global Speech Arena.

Toolin Editorial Team
Qwen3.8-Max Preview: A Hands-On Early Access Guide
AI Products

Qwen3.8-Max Preview: A Hands-On Early Access Guide

Qwen's 2.4T-parameter flagship preview is live on Token Plan, Qoder, and the Qwen website; officially rated second only to Fable 5 overall.

Toolin Editorial Team
Making a Full-Scene Infographic with SenseNova U1 Pro
AI Tutorials

Making a Full-Scene Infographic with SenseNova U1 Pro

A hands-on tutorial for SenseTime's flagship multimodal model U1 Pro: turn raw data into a deliverable 8K infographic and full-match panoramic visual, automatically.

Toolin Editorial Team
Claude for Teachers: A Free AI Teaching Assistant for Every K-12 Teacher in the US
AI Products

Claude for Teachers: A Free AI Teaching Assistant for Every K-12 Teacher in the US

Anthropic launches a free AI teaching assistant for certified US K-12 teachers, wired into 50 states' standards, with lesson-plan alignment, differentiated tiering, and student data analysis.

Toolin Editorial Team