LLM Model Comparison for OpenCode Go — July 2026
Some links on this page are affiliate links. I earn a commission at no extra cost to you.
This guide compares 13 large language models available for AI coding agents as of June 2026. Every model is ranked by cost, output speed, capabilities, and specialization — so you can pick the right model for your workflow and budget. Pricing and speed data are sourced from official provider pages, independent benchmarks, and real-world API measurements.
I use OpenCode Go with these models daily — it's my go-to coding agent whether I'm in the terminal or VS Code. OpenCode Zen is the free tier; Go gives you generous credits for $5 your first month, then $10/month after that (cancel anytime). Highly recommended if you work with AI-assisted development. Use my referral link and you'll get $5 in usage credits to start.
Speed notes: Values shown are typical throughput on managed API providers. Actual tok/s varies by provider, quantization, hardware, and concurrent load.
Capabilities legend: 🖼️ = image input, 🎬 = video input, 🎤 = audio input, 🧠 = reasoning/thinking mode, 🛠️ = tool/function calling, 📖 = structured output, 🔓 = open weights.
| Model | Provider | Cost/1M tokens | Rank | Speed | Capabilities | Best For |
|---|---|---|---|---|---|---|
| Qwen3.7 Max | Alibaba Cloud | $2.50 in / $7.50 out | 1 — Most Expensive | ~117 | 🧠🛠️📖 | Complex agentic workflows, sustained autonomous coding |
| GLM-5.2 | DeepInfra, Fireworks, Z.ai | $1.40 in / $4.40 out | 2 | ~120–150 | 🧠🛠️📖🔓 | Open-source agentic coding, long-context tasks |
| GLM-5.1 | DeepInfra, Fireworks, Z.ai | $1.40 in / $4.40 out | 2 | ~50–75 | 🧠🛠️📖🔓 | Self-hosted agents, cost-sensitive open-source |
| Kimi K2.7 Code | Moonshot AI | $0.95 in / $4.00 out | 3 | ~100–180 | 🖼️🎬🧠🛠️📖🔓 | Code generation/refactoring, agentic coding |
| Kimi K2.6 | Moonshot AI | $0.95 in / $4.00 out | 3 | 85–380 | 🖼️🎬🧠🛠️📖🔓 | Multi-agent orchestration, autonomous research |
| Qwen3.6 Plus | Alibaba Cloud | $0.50 in / $3.00 out | 4 | ~50–85 | 🖼️🎬🧠🛠️📖 | Repo-level coding, 3D/game dev, multimodal |
| Qwen3.7 Plus | Alibaba Cloud | $0.40 in / $1.60 out | 5 | ~47 | 🖼️🎬🧠🛠️📖 | Vision+reasoning, multimodal agent workflows |
| MiniMax M3 | MiniMax | $0.30 in / $1.20 out | 6 | ~45–145 | 🖼️🎬🧠🛠️📖🔓 | Code from screenshots, desktop GUI automation |
| MiniMax M2.7 | MiniMax | $0.30 in / $1.20 out | 6 | ~40–50 | 🧠🛠️📖 | Self-improving agent loops, multi-agent collaboration |
| MiMo-V2.5-Pro | Xiaomi MiMo | $0.435 in / $0.87 out | 7 | ~44 / 1000+ (US) | 🖼️🎤🎬🧠🛠️📖🔓 | Demanding agentic tasks, multimodal+code |
| DeepSeek V4 Pro | DeepSeek | $0.435 in / $0.87 out | 7 | ~88 | 🧠🛠️📖🔓 | Cost-sensitive frontier coding, full-codebase analysis |
| DeepSeek V4 Flash | DeepSeek | $0.14 in / $0.28 out | 8 | ~83–150 | 🧠🛠️📖🔓 | High-throughput, cost-critical applications |
| MiMo-V2.5 | Xiaomi MiMo | $0.14 in / $0.28 out | 8 — Least Expensive | ~80–120 | 🖼️🎤🎬🧠🛠️📖🔓 | High-volume, cost-sensitive multimodal tasks |
Output Price per 1M Tokens (USD)
Model Output $/1M tokens ────────────────────────────────────────────────── Qwen3.7 Max ██████████████████████████████████████████████████ $7.50 GLM-5.2 ███████████████████████████████████ $4.40 GLM-5.1 ███████████████████████████████████ $4.40 Kimi K2.7 Code ██████████████████████████████████ $4.00 Kimi K2.6 ██████████████████████████████████ $4.00 Qwen3.6 Plus ████████████████████████ $3.00 Qwen3.7 Plus ██████████████ $1.60 MiniMax M3 ██████████ $1.20 MiniMax M2.7 ██████████ $1.20 MiMo-V2.5-Pro ██████ $0.87 DeepSeek V4 Pro ██████ $0.87 DeepSeek V4 Flash ██ $0.28 MiMo-V2.5 ██ $0.28
█ = $0.15
Cost vs Speed Side-by-Side
Model Cost ($/1M out) Speed (tok/s) ───────────────────────────────────────────────────────────────────────────── Qwen3.7 Max ██████████████████████████████████████████████████ $7.50 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 117 GLM-5.2 ███████████████████████████████████ $4.40 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 135 GLM-5.1 ███████████████████████████████████ $4.40 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 62 Kimi K2.7 Code ██████████████████████████████████ $4.00 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 140 Kimi K2.6 ██████████████████████████████████ $4.00 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 85 Qwen3.6 Plus ████████████████████████ $3.00 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 67 Qwen3.7 Plus ██████████████ $1.60 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 47 MiniMax M3 ██████████ $1.20 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 95 MiniMax M2.7 ██████████ $1.20 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 45 MiMo-V2.5-Pro ██████ $0.87 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 44 DeepSeek V4 Pro ██████ $0.87 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 88 DeepSeek V4 Flash ██ $0.28 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 116 MiMo-V2.5 ██ $0.28 ▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓▓ 100
█ = $0.15/1M out (cost) ▓ = 3 tok/s (speed)
Speed × Cost Grid
Models placed by cost bucket (columns) × speed bucket (rows). Empty cells are gaps in the market.
Cost per 1M output tokens
<$1.00 $1-2 $2-3 $3-4 $4-5 $5+
─────────────────────────────────────────────────────────
125-150+ tok/s ·········· ·········· ·········· K27C ····· GLM52 ····· ·· 🏎️ Max speed
100-125 tok/s DFlash M25 ·········· ·········· ·········· ·········· ·· QM ·· ⚡ Very fast
75-100 tok/s DPro ····· MM3 ······· ·········· K26 ····· ·········· ·········· 🔥 Fast
50-75 tok/s ·········· Q37P ····· Q36P ····· ·········· GLM51 ···· ·········· ⚙️ Moderate
<50 tok/s M25P ····· M27 ······· ·········· ·········· ·········· ·········· 🐢 Slower Model Key
What the Grid Tells You
| Observation | Models |
|---|---|
| ⭐ Best value — fast AND cheap | DeepSeek V4 Flash, DeepSeek V4 Pro, MiMo-V2.5, MiniMax M3 |
| ⚡ Pay for speed — fastest regardless of cost | Kimi K2.7 Code, GLM-5.2, Qwen3.7 Max |
| 💡 Hidden value — surprising speed at lower cost | MiniMax M3 ($1–2 cost, 75–100 speed), DeepSeek V4 Flash (<$1, 100–125 speed) |
| 🐢 Affordable but slow | MiMo-V2.5-Pro, Qwen3.7 Plus, MiniMax M2.7 |
| 🕳️ Market gap — empty $2–3 column | (entire $2–3 range empty across all speed tiers) |
Speed-per-Dollar Value Rating
Sorted by throughput per dollar (tok/s per $1M output cost). Higher = more tokens for your money.
Model Cost Speed tok/s per $ Rating ────────────────────────────────────────────────────────────── DeepSeek V4 Flash $0.28 116 414 ★★★★★ MiMo-V2.5 $0.28 100 357 ★★★★★ DeepSeek V4 Pro $0.87 88 101 ★★★★ MiniMax M3 $1.20 95 79 ★★★ MiMo-V2.5-Pro $0.87 44 51 ★★★ MiniMax M2.7 $1.20 45 38 ★★ Kimi K2.7 Code $4.00 140 35 ★★ GLM-5.2 $4.40 135 31 ★★ Qwen3.7 Plus $1.60 47 29 ★★ Kimi K2.6 $4.00 85 21 ★★ Qwen3.6 Plus $3.00 67 22 ★★ Qwen3.7 Max $7.50 117 16 ★ GLM-5.1 $4.40 62 14 ★
| Rating | Threshold (tok/s per $1 output) |
|---|---|
| ★★★★★ | 200+ |
| ★★★★ | 80–200 |
| ★★★ | 40–80 |
| ★★ | 20–40 |
| ★ | <20 |
Speed Benchmark Sources
| Model | Speed (tok/s) | Source |
|---|---|---|
| Qwen3.7 Max | ~117 | Design for Online Leaderboard (June 2026) |
| GLM-5.2 | ~120–150 | Estimated from provider benchmarks (Z.ai, DeepInfra) |
| GLM-5.1 | ~50–75 | LLM Stats — FriendliAI |
| Kimi K2.7 Code | ~100–180 | Estimated — similar architecture to K2.6 |
| Kimi K2.6 | 85–380 | Artificial Analysis |
| Qwen3.7 Plus | ~47 | Artificial Analysis |
| Qwen3.6 Plus | ~50–85 | LLM Stats — Together |
| MiniMax M3 | 45–145 | Telnyx Benchmark |
| MiniMax M2.7 | ~40–50 | Estimated from provider benchmarks |
| MiMo-V2.5-Pro | ~44 / 1000+ | Xiaomi MiMo Blog |
| DeepSeek V4 Pro | ~88 | Artificial Analysis |
| DeepSeek V4 Flash | ~83–150 | Range from provider benchmarks (DeepInfra, Fireworks, OpenRouter) |
| MiMo-V2.5 | ~80–120 | Estimated — lighter variant of V2.5-Pro |
🥋 Try OpenCode Go
I use it daily with every model in this guide. Use my referral link and you get $5 in usage credits to start.
opencode.ai/go →Capability Notes
- Multimodal models (image input) — Qwen3.7 Plus, Qwen3.6 Plus, Kimi K2.7 Code, Kimi K2.6, MiniMax M3, MiMo-V2.5-Pro, MiMo-V2.5
- Video input — Same set as multimodal above
- Audio input — MiMo-V2.5-Pro, MiMo-V2.5 (Xiaomi models)
- Open weights (can self-host) — GLM-5.2, GLM-5.1, Kimi K2.7 Code, Kimi K2.6, MiniMax M3, MiniMax M2.7, MiMo-V2.5-Pro, MiMo-V2.5, DeepSeek V4 Pro, DeepSeek V4 Flash
- 1M+ context window — Qwen3.7 Max, Qwen3.7 Plus, Qwen3.6 Plus, GLM-5.2, MiniMax M3, MiMo-V2.5-Pro, DeepSeek V4 Pro, DeepSeek V4 Flash, MiMo-V2.5
- Reasoning/CoT mode — Qwen3.7 Max (always-on), Qwen3.7 Plus, Qwen3.6 Plus, GLM-5.2, GLM-5.1, Kimi K2.7 Code (always-on), Kimi K2.6 (toggle), MiniMax M3, DeepSeek V4 Pro (toggle), DeepSeek V4 Flash (toggle)
Quick Decision Guide
| If you need… | Pick this model |
|---|---|
| Best agentic reasoning (quality first) | Qwen3.7 Max |
| Best open-source coding agent | GLM-5.2 or MiMo-V2.5-Pro |
| Multi-agent swarms | Kimi K2.6 |
| Cheapest frontier coding | DeepSeek V4 Flash |
| Cheapest multimodal | MiMo-V2.5 |
| Code from screenshots / desktop automation | MiniMax M3 |
| Self-hosted, cost-sensitive agents | GLM-5.1 or DeepSeek V4 Flash |
| Long-context (1M+) at low cost | DeepSeek V4 Flash or MiniMax M3 |
Self-Hosting Hardware
Running open-weight models locally? You'll want a machine that can handle them. Here's what pairs well with self-hosted coding agents:
Compact, power-sipping home server for lighter models like DeepSeek V4 Flash or GLM-5.1. Great for a dedicated 24/7 coding agent box. 🎮 NVIDIA GPU
If you're running larger models (DeepSeek V4 Pro, MiMo-V2.5-Pro), a dedicated GPU makes a real difference in tokens per second. Search for "RTX 3090 24GB used" for the best price-to-performance value. 💾 NVMe SSD
Model weights eat up 50–300GB+. A fast NVMe drive means loading models in seconds instead of minutes. 🧠 DDR5 RAM
Models with 1M+ context windows need serious memory to breathe. 64GB+ gives you room for full-context reasoning without swapping. 🔌 5-Port Network Switch
Cheap, silent, plug-and-play way to wire your mini PC, NAS, and other gear together on a dedicated LAN. Beats Wi-Fi for reliability when you're running a coding agent 24/7. ⌨️ USB-C Hub / Dock
Don't forget the desk setup. A solid hub makes it easy to plug your coding machine into monitors, peripherals, and networking without crawling behind the desk.
Standard disclaimer: As an Amazon Associate I earn from qualifying purchases. Prices and availability change frequently.
Pricing Sources Used
| Source | Models |
|---|---|
| Z.AI Pricing | GLM-5.2, GLM-5.1 |
| Kimi Platform | Kimi K2.7 Code, K2.6 |
| Xiaomi MiMo Pricing | MiMo-V2.5-Pro, MiMo-V2.5 |
| Alibaba Cloud Model Studio | Qwen3.7 Max, Qwen3.7 Plus, Qwen3.6 Plus |
| MiniMax Pay-as-You-Go | MiniMax M3, M2.7 |
| DeepSeek Pricing | DeepSeek V4 Pro, V4 Flash |
Last updated: July 4, 2026. Prices and speeds change frequently — verify against provider pages before making purchasing decisions.