LLM Model Comparison for OpenCode Go — July 2026

Some links on this page are affiliate links. I earn a commission at no extra cost to you.

This guide compares 13 large language models available for AI coding agents as of June 2026. Every model is ranked by cost, output speed, capabilities, and specialization — so you can pick the right model for your workflow and budget. Pricing and speed data are sourced from official provider pages, independent benchmarks, and real-world API measurements.

I use OpenCode Go with these models daily — it's my go-to coding agent whether I'm in the terminal or VS Code. OpenCode Zen is the free tier; Go gives you generous credits for $5 your first month, then $10/month after that (cancel anytime). Highly recommended if you work with AI-assisted development. Use my referral link and you'll get $5 in usage credits to start.

Speed notes: Values shown are typical throughput on managed API providers. Actual tok/s varies by provider, quantization, hardware, and concurrent load.
Capabilities legend: 🖼️ = image input, 🎬 = video input, 🎤 = audio input, 🧠 = reasoning/thinking mode, 🛠️ = tool/function calling, 📖 = structured output, 🔓 = open weights.
Model Provider Cost/1M tokens Rank Speed Capabilities Best For
Qwen3.7 Max Alibaba Cloud $2.50 in / $7.50 out 1 — Most Expensive ~117 🧠🛠️📖 Complex agentic workflows, sustained autonomous coding
GLM-5.2 DeepInfra, Fireworks, Z.ai $1.40 in / $4.40 out 2 ~120–150 🧠🛠️📖🔓 Open-source agentic coding, long-context tasks
GLM-5.1 DeepInfra, Fireworks, Z.ai $1.40 in / $4.40 out 2 ~50–75 🧠🛠️📖🔓 Self-hosted agents, cost-sensitive open-source
Kimi K2.7 Code Moonshot AI $0.95 in / $4.00 out 3 ~100–180 🖼️🎬🧠🛠️📖🔓 Code generation/refactoring, agentic coding
Kimi K2.6 Moonshot AI $0.95 in / $4.00 out 3 85–380 🖼️🎬🧠🛠️📖🔓 Multi-agent orchestration, autonomous research
Qwen3.6 Plus Alibaba Cloud $0.50 in / $3.00 out 4 ~50–85 🖼️🎬🧠🛠️📖 Repo-level coding, 3D/game dev, multimodal
Qwen3.7 Plus Alibaba Cloud $0.40 in / $1.60 out 5 ~47 🖼️🎬🧠🛠️📖 Vision+reasoning, multimodal agent workflows
MiniMax M3 MiniMax $0.30 in / $1.20 out 6 ~45–145 🖼️🎬🧠🛠️📖🔓 Code from screenshots, desktop GUI automation
MiniMax M2.7 MiniMax $0.30 in / $1.20 out 6 ~40–50 🧠🛠️📖 Self-improving agent loops, multi-agent collaboration
MiMo-V2.5-Pro Xiaomi MiMo $0.435 in / $0.87 out 7 ~44 / 1000+ (US) 🖼️🎤🎬🧠🛠️📖🔓 Demanding agentic tasks, multimodal+code
DeepSeek V4 Pro DeepSeek $0.435 in / $0.87 out 7 ~88 🧠🛠️📖🔓 Cost-sensitive frontier coding, full-codebase analysis
DeepSeek V4 Flash DeepSeek $0.14 in / $0.28 out 8 ~83–150 🧠🛠️📖🔓 High-throughput, cost-critical applications
MiMo-V2.5 Xiaomi MiMo $0.14 in / $0.28 out 8 — Least Expensive ~80–120 🖼️🎤🎬🧠🛠️📖🔓 High-volume, cost-sensitive multimodal tasks

Output Price per 1M Tokens (USD)

Model                     Output $/1M tokens
──────────────────────────────────────────────────
Qwen3.7 Max               ██████████████████████████████████████████████████ $7.50
GLM-5.2                   ███████████████████████████████████                $4.40
GLM-5.1                   ███████████████████████████████████                $4.40
Kimi K2.7 Code            ██████████████████████████████████                 $4.00
Kimi K2.6                 ██████████████████████████████████                 $4.00
Qwen3.6 Plus              ████████████████████████                           $3.00
Qwen3.7 Plus              ██████████████                                     $1.60
MiniMax M3                ██████████                                         $1.20
MiniMax M2.7              ██████████                                         $1.20
MiMo-V2.5-Pro             ██████                                             $0.87
DeepSeek V4 Pro           ██████                                             $0.87
DeepSeek V4 Flash         ██                                                 $0.28
MiMo-V2.5                 ██                                                 $0.28

█ = $0.15


Cost vs Speed Side-by-Side

Model                  Cost ($/1M out)              Speed (tok/s)
─────────────────────────────────────────────────────────────────────────────
Qwen3.7 Max            ██████████████████████████████████████████████████ $7.50 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 117
GLM-5.2                ███████████████████████████████████                $4.40 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 135
GLM-5.1                ███████████████████████████████████                $4.40 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓             62
Kimi K2.7 Code         ██████████████████████████████████                 $4.00 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 140
Kimi K2.6              ██████████████████████████████████                 $4.00 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓   85
Qwen3.6 Plus           ████████████████████████                           $3.00 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓   67
Qwen3.7 Plus           ██████████████                                     $1.60 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓       47
MiniMax M3             ██████████                                         $1.20 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓   95
MiniMax M2.7           ██████████                                         $1.20 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓      45
MiMo-V2.5-Pro          ██████                                             $0.87 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓      44
DeepSeek V4 Pro        ██████                                             $0.87 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓   88
DeepSeek V4 Flash      ██                                                 $0.28 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 116
MiMo-V2.5              ██                                                 $0.28 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 100

█ = $0.15/1M out (cost)  ▓ = 3 tok/s (speed)


Speed × Cost Grid

Models placed by cost bucket (columns) × speed bucket (rows). Empty cells are gaps in the market.

                    Cost per 1M output tokens
                  <$1.00    $1-2      $2-3      $3-4      $4-5      $5+
                ─────────────────────────────────────────────────────────
125-150+ tok/s    ·········· ·········· ·········· K27C ····· GLM52 ····· ·· 🏎️ Max speed
100-125  tok/s  DFlash  M25 ·········· ·········· ·········· ·········· ·· QM ·· ⚡ Very fast
 75-100  tok/s  DPro ····· MM3 ······· ·········· K26 ····· ·········· ·········· 🔥 Fast
 50-75   tok/s    ·········· Q37P ····· Q36P ····· ·········· GLM51 ···· ·········· ⚙️ Moderate
 <50     tok/s  M25P ····· M27 ······· ·········· ·········· ·········· ·········· 🐢 Slower

Model Key

QM — Qwen3.7 Max K27C — Kimi K2.7 Code GLM52 — GLM-5.2 K26 — Kimi K2.6 GLM51 — GLM-5.1 M3 — MiniMax M3 Q37P — Qwen3.7 Plus M27 — MiniMax M2.7 Q36P — Qwen3.6 Plus M25P — MiMo-V2.5-Pro DPro — DeepSeek V4 Pro M25 — MiMo-V2.5 DFlash — DeepSeek V4 Flash

What the Grid Tells You

Observation Models
Best value — fast AND cheap DeepSeek V4 Flash, DeepSeek V4 Pro, MiMo-V2.5, MiniMax M3
Pay for speed — fastest regardless of cost Kimi K2.7 Code, GLM-5.2, Qwen3.7 Max
💡 Hidden value — surprising speed at lower cost MiniMax M3 ($1–2 cost, 75–100 speed), DeepSeek V4 Flash (<$1, 100–125 speed)
🐢 Affordable but slow MiMo-V2.5-Pro, Qwen3.7 Plus, MiniMax M2.7
🕳️ Market gap — empty $2–3 column (entire $2–3 range empty across all speed tiers)

Speed-per-Dollar Value Rating

Sorted by throughput per dollar (tok/s per $1M output cost). Higher = more tokens for your money.

Model                        Cost  Speed  tok/s per $  Rating
──────────────────────────────────────────────────────────────
DeepSeek V4 Flash            $0.28   116         414  ★★★★★
MiMo-V2.5                    $0.28   100         357  ★★★★★
DeepSeek V4 Pro              $0.87    88         101  ★★★★
MiniMax M3                   $1.20    95          79  ★★★
MiMo-V2.5-Pro                $0.87    44          51  ★★★
MiniMax M2.7                 $1.20    45          38  ★★
Kimi K2.7 Code               $4.00   140          35  ★★
GLM-5.2                      $4.40   135          31  ★★
Qwen3.7 Plus                 $1.60    47          29  ★★
Kimi K2.6                    $4.00    85          21  ★★
Qwen3.6 Plus                 $3.00    67          22  ★★
Qwen3.7 Max                  $7.50   117          16  ★
GLM-5.1                      $4.40    62          14  ★
Rating Threshold (tok/s per $1 output)
★★★★★200+
★★★★80–200
★★★40–80
★★20–40
<20

Speed Benchmark Sources

Model Speed (tok/s) Source
Qwen3.7 Max~117Design for Online Leaderboard (June 2026)
GLM-5.2~120–150Estimated from provider benchmarks (Z.ai, DeepInfra)
GLM-5.1~50–75LLM Stats — FriendliAI
Kimi K2.7 Code~100–180Estimated — similar architecture to K2.6
Kimi K2.685–380Artificial Analysis
Qwen3.7 Plus~47Artificial Analysis
Qwen3.6 Plus~50–85LLM Stats — Together
MiniMax M345–145Telnyx Benchmark
MiniMax M2.7~40–50Estimated from provider benchmarks
MiMo-V2.5-Pro~44 / 1000+Xiaomi MiMo Blog
DeepSeek V4 Pro~88Artificial Analysis
DeepSeek V4 Flash~83–150Range from provider benchmarks (DeepInfra, Fireworks, OpenRouter)
MiMo-V2.5~80–120Estimated — lighter variant of V2.5-Pro

🥋 Try OpenCode Go

I use it daily with every model in this guide. Use my referral link and you get $5 in usage credits to start.

opencode.ai/go →

Capability Notes


Quick Decision Guide

If you need… Pick this model
Best agentic reasoning (quality first)Qwen3.7 Max
Best open-source coding agentGLM-5.2 or MiMo-V2.5-Pro
Multi-agent swarmsKimi K2.6
Cheapest frontier codingDeepSeek V4 Flash
Cheapest multimodalMiMo-V2.5
Code from screenshots / desktop automationMiniMax M3
Self-hosted, cost-sensitive agentsGLM-5.1 or DeepSeek V4 Flash
Long-context (1M+) at low costDeepSeek V4 Flash or MiniMax M3

Self-Hosting Hardware

Running open-weight models locally? You'll want a machine that can handle them. Here's what pairs well with self-hosted coding agents:

🖥️ Mini PC
Compact, power-sipping home server for lighter models like DeepSeek V4 Flash or GLM-5.1. Great for a dedicated 24/7 coding agent box.
🎮 NVIDIA GPU
If you're running larger models (DeepSeek V4 Pro, MiMo-V2.5-Pro), a dedicated GPU makes a real difference in tokens per second. Search for "RTX 3090 24GB used" for the best price-to-performance value.
💾 NVMe SSD
Model weights eat up 50–300GB+. A fast NVMe drive means loading models in seconds instead of minutes.
🧠 DDR5 RAM
Models with 1M+ context windows need serious memory to breathe. 64GB+ gives you room for full-context reasoning without swapping.
🔌 5-Port Network Switch
Cheap, silent, plug-and-play way to wire your mini PC, NAS, and other gear together on a dedicated LAN. Beats Wi-Fi for reliability when you're running a coding agent 24/7.
⌨️ USB-C Hub / Dock
Don't forget the desk setup. A solid hub makes it easy to plug your coding machine into monitors, peripherals, and networking without crawling behind the desk.

Standard disclaimer: As an Amazon Associate I earn from qualifying purchases. Prices and availability change frequently.


Pricing Sources Used

Source Models
Z.AI PricingGLM-5.2, GLM-5.1
Kimi PlatformKimi K2.7 Code, K2.6
Xiaomi MiMo PricingMiMo-V2.5-Pro, MiMo-V2.5
Alibaba Cloud Model StudioQwen3.7 Max, Qwen3.7 Plus, Qwen3.6 Plus
MiniMax Pay-as-You-GoMiniMax M3, M2.7
DeepSeek PricingDeepSeek V4 Pro, V4 Flash

Last updated: July 4, 2026. Prices and speeds change frequently — verify against provider pages before making purchasing decisions.