Operator noteAleksei Balchunas

Which coding model is cheapest per task?

Same coding benchmarks, different price tags. Approximate $/task across OpenAI Codex (Sol, Terra, Luna), Kimi K3, Grok 4.5, Cursor Composer 2.5, and DeepSeek V4 Flash. Sorted by API cost, lowest first.

Click any column to sort. Default: API $/task, lowest first.

Model Effort / Harness Coding Index DeepSWE TerminalBench SWE Atlas API / PAYG $/task GPT Credits −20% Pro 20x rough est.
SolMedium / Codex6164%78%40%$2.99$2.39$1.50–$1.87
SolHigh / Codex6465%83%45%$4.14$3.31$2.07–$2.59
SolXHigh / Codex6567%86%42%$5.24$4.19$2.62–$3.28
SolMax / Codex6769%88%43%$7.08$5.66$3.54–$4.43
TerraMedium / Codex4846%69%28%$0.72$0.58$0.36–$0.45
TerraHigh / Codex5660%76%31%$1.27$1.02$0.64–$0.79
TerraXHigh / Codex5758%81%32%$1.52$1.22$0.76–$0.95
TerraMax / Codex6267%84%36%$2.21$1.77$1.10–$1.38
LunaMedium / Codex4237%63%27%$0.09$0.07$0.04–$0.06
LunaHigh / Codex5153%72%29%$0.19$0.15$0.10–$0.12
LunaXHigh / Codex5557%76%31%$0.25$0.20$0.12–$0.16
LunaMax / Codex5963%80%33%$0.31$0.25$0.15–$0.19
Kimi K3Kimi Code6164%84%37%$3.18
Grok 4.5High / Grok Build6460%85%48%$2.59
Composer 2.5Standard / Cursor3816%67%31%$0.08
Composer 2.5Fast / Cursor3816%67%31%$0.55
DeepSeek V4 FlashMax / Codex5543%85%39%$0.07 direct API
Higher benchmark scores are better. Lower cost per task is better.

As of . Prices, discounts, and benchmark scores move; treat this as a snapshot, not a live feed.

Pro 20x: rough estimate only, based on the working assumption that included Pro 20x compute provides about 1.6x to 2.0x the value of base PAYG credits. It is not an official OpenAI per-task price.

GPT Credits −20%: applies the 20% credit discount discussed in the source comparison. Non-OpenAI models are left blank.

Benchmark comparability: these rows combine model and harness configurations. Harness choice can materially affect agentic coding results.

Practical guide: How to Have an Endless Supply of Tokens in Codex.