Codex GPT-6 Sol Performance Tracker

We are collecting a new GPT-6 Sol/high baseline from runs beginning September 24, 2026. Degradation detection is paused.

- • Updated daily: Daily benchmarks on a curated subset of SWE-Bench-Pro

- • Collecting baseline: Degradation detection resumes once the new model baseline is established

- • What you see is what you get: We benchmark GPT-6 Sol directly in Codex CLI, with no custom agent harness.

Summary

Daily Trend

Pass rate over time

Toggle 95% CI to view uncertainty around each point.

Enable 95% CI checkbox to show confidence intervals

Weekly Trend

Aggregated 7-day pass rate

The same uncertainty toggle applies here for 7-day windows.

Enable 95% CI checkbox to show confidence intervals

Baseline data is being collected.

Performance deltas will be available once baseline is established.

Change Overview

Performance delta by period

Other Metrics

Daily benchmark resource and execution trends

Input Tokens

Daily total input usage

Output Tokens

Daily total output usage

Average Runtime

Per-instance runtime by day

Tool Calls

Total tool invocations by day

Methodology

The goal of this tracker is to detect statistically significant degradations in Codex with GPT-6 Sol performance on SWE tasks. We are an independent third party with no affiliation to frontier model providers.

Context: As AI coding assistants become increasingly relied upon for software development, monitoring their performance over time is critical. This tracker helps detect regressions that may affect developer productivity.

We run a daily evaluation of Codex CLI on a curated, contamination-resistant subset of SWE-Bench-Pro. We use the latest available Codex release with GPT-6 Sol. Scheduled evaluations currently use high reasoning effort. Benchmarks run directly in Codex without a custom agent harness, so results reflect what actual users can expect. This allows us to detect degradation related to both model changes and harness changes.

Each daily evaluation attempts N=50 test instances, so daily variability is expected. Pass rates use completed test outcomes; infrastructure failures and cancellations are reported as attempted evaluations but excluded from the statistical denominator. Weekly and monthly results are aggregated for more reliable estimates.

We are collecting a new GPT-6 Sol/high baseline from runs beginning September 24, 2026. The previous gpt-5.6-sol baseline of 83.40% (613 passes among 735 valid trials) belongs to that earlier model and is not used for degradation verdicts during collection. Earlier high and xhigh reasoning runs remain available on the historical performance page.

Pass rates and 95% confidence intervals remain descriptive while the baseline is collected. Degradation detection, significance thresholds, and performance-change comparisons are paused until a baseline for the new model is established.