Same coding benchmarks, different price tags. Approximate $/task across OpenAI Codex (Sol, Terra, Luna), Kimi K3, Grok 4.5, Cursor Composer 2.5, and DeepSeek V4 Flash. Sorted by API cost, lowest first.
Click any column to sort. Default: API $/task, lowest first.
| Model ↕ | Effort / Harness ↕ | Coding Index ↕ | DeepSWE ↕ | TerminalBench ↕ | SWE Atlas ↕ | API / PAYG $/task ↕ | GPT Credits −20% ↕ | Pro 20x rough est. ↕ |
|---|---|---|---|---|---|---|---|---|
| Sol | Medium / Codex | 61 | 64% | 78% | 40% | $2.99 | $2.39 | $1.50–$1.87 |
| Sol | High / Codex | 64 | 65% | 83% | 45% | $4.14 | $3.31 | $2.07–$2.59 |
| Sol | XHigh / Codex | 65 | 67% | 86% | 42% | $5.24 | $4.19 | $2.62–$3.28 |
| Sol | Max / Codex | 67 | 69% | 88% | 43% | $7.08 | $5.66 | $3.54–$4.43 |
| Terra | Medium / Codex | 48 | 46% | 69% | 28% | $0.72 | $0.58 | $0.36–$0.45 |
| Terra | High / Codex | 56 | 60% | 76% | 31% | $1.27 | $1.02 | $0.64–$0.79 |
| Terra | XHigh / Codex | 57 | 58% | 81% | 32% | $1.52 | $1.22 | $0.76–$0.95 |
| Terra | Max / Codex | 62 | 67% | 84% | 36% | $2.21 | $1.77 | $1.10–$1.38 |
| Luna | Medium / Codex | 42 | 37% | 63% | 27% | $0.09 | $0.07 | $0.04–$0.06 |
| Luna | High / Codex | 51 | 53% | 72% | 29% | $0.19 | $0.15 | $0.10–$0.12 |
| Luna | XHigh / Codex | 55 | 57% | 76% | 31% | $0.25 | $0.20 | $0.12–$0.16 |
| Luna | Max / Codex | 59 | 63% | 80% | 33% | $0.31 | $0.25 | $0.15–$0.19 |
| Kimi K3 | Kimi Code | 61 | 64% | 84% | 37% | $3.18 | — | — |
| Grok 4.5 | High / Grok Build | 64 | 60% | 85% | 48% | $2.59 | — | — |
| Composer 2.5 | Standard / Cursor | 38 | 16% | 67% | 31% | $0.08 | — | — |
| Composer 2.5 | Fast / Cursor | 38 | 16% | 67% | 31% | $0.55 | — | — |
| DeepSeek V4 Flash | Max / Codex | 55 | 43% | 85% | 39% | $0.07 direct API | — | — |
As of . Prices, discounts, and benchmark scores move; treat this as a snapshot, not a live feed.
Pro 20x: rough estimate only, based on the working assumption that included Pro 20x compute provides about 1.6x to 2.0x the value of base PAYG credits. It is not an official OpenAI per-task price.
GPT Credits −20%: applies the 20% credit discount discussed in the source comparison. Non-OpenAI models are left blank.
Benchmark comparability: these rows combine model and harness configurations. Harness choice can materially affect agentic coding results.
Practical guide: How to Have an Endless Supply of Tokens in Codex.