Codex GPT-6 Sol Performance Tracker
We are collecting a new GPT-6 Sol/high baseline from runs beginning September 24, 2026. Degradation detection is paused.
- • Updated daily: Daily benchmarks on a curated subset of SWE-Bench-Pro
- • Collecting baseline: Degradation detection resumes once the new model baseline is established
- • What you see is what you get: We benchmark GPT-6 Sol directly in Codex CLI, with no custom agent harness.
Summary
Daily Trend
Pass rate over time
Toggle 95% CI to view uncertainty around each point.
Enable 95% CI checkbox to show confidence intervals
Weekly Trend
Aggregated 7-day pass rate
The same uncertainty toggle applies here for 7-day windows.
Enable 95% CI checkbox to show confidence intervals
Baseline data is being collected.
Performance deltas will be available once baseline is established.
Change Overview
Performance delta by period
Other Metrics
Daily benchmark resource and execution trends
Input Tokens
Daily total input usage
Output Tokens
Daily total output usage
Average Runtime
Per-instance runtime by day
Tool Calls
Total tool invocations by day
Methodology
The goal of this tracker is to detect statistically significant degradations in Codex with GPT-6 Sol performance on SWE tasks. We are an independent third party with no affiliation to frontier model providers.
Context: As AI coding assistants become increasingly relied upon for software development, monitoring their performance over time is critical. This tracker helps detect regressions that may affect developer productivity.
We run a daily evaluation of Codex CLI on a curated, contamination-resistant subset of SWE-Bench-Pro. We use the latest available Codex release with GPT-6 Sol. Scheduled evaluations currently use high reasoning effort. Benchmarks run directly in Codex without a custom agent harness, so results reflect what actual users can expect. This allows us to detect degradation related to both model changes and harness changes.
Each daily evaluation attempts N=50 test instances, so daily variability is expected. Pass rates use completed test outcomes; infrastructure failures and cancellations are reported as attempted evaluations but excluded from the statistical denominator. Weekly and monthly results are aggregated for more reliable estimates.
We are collecting a new GPT-6 Sol/high baseline from runs beginning September 24, 2026. The previous gpt-5.6-sol baseline of 83.40% (613 passes among 735 valid trials) belongs to that earlier model and is not used for degradation verdicts during collection. Earlier high and xhigh reasoning runs remain available on the historical performance page.
Pass rates and 95% confidence intervals remain descriptive while the baseline is collected. Degradation detection, significance thresholds, and performance-change comparisons are paused until a baseline for the new model is established.