Claude Opus 4.7
Anthropic
Default to Claude Opus 4.7 inside Claude Code for repo-scale edits. GPT-5.6 via Codex is the strongest first-party OpenAI alternative; pick Cursor when you need multi-model routing (including Grok 4.5) in the same workflow.
Directional ratings for the AI models people are actually using right now. Three buckets — code, image and video, writing and agents — refreshed after every release.
Anthropic
Default to Claude Opus 4.7 inside Claude Code for repo-scale edits. GPT-5.6 via Codex is the strongest first-party OpenAI alternative; pick Cursor when you need multi-model routing (including Grok 4.5) in the same workflow.
OpenAI
Sora 2 for video, Midjourney v7 for still images. Imagen 4 and the Nano Banana image line inside Gemini are the best one-surface answer when you want both from a single account.
Anthropic
Claude Opus 4.7 leads for long-form editorial and stable agent loops. GPT-5.6 wins when you need broader tool integrations and voice, especially around Microsoft and OpenAI surfaces.
Gemini 3.6 Flash (workhorse), 3.5 Flash-Lite and a 3.5 Flash Cyber security variant land, plus Nano Banana 2 Lite for images; the Gemini 3.5 Pro flagship slips again.
Google DeepMind · Gemini 3 Pro
DeepSeek promotes V4 Pro and V4 Flash out of preview with new peak-hour API pricing; V4 supersedes the R2 reasoning preview.
DeepSeek · DeepSeek V4
GPT-5.6 rolls out in Luna, Terra and Sol variants after a limited government-gated preview, alongside the GPT-Live real-time voice models.
OpenAI · GPT-5.6
xAI ships Grok 4.5, a mixture-of-experts model with a 500K context window focused on coding and agents, available in Cursor and Grok Build.
xAI · Grok 4.5
Mistral confirms a new open-weight model family in partner early access, with a broader release expected later in the summer.
Mistral · Mistral Large 3
GLM-5.2 ships as an open-weight model with a 1M-token context window, rated the strongest open-weight model available at launch.
Z.ai · GLM-5.2
Kimi K2.7 Code lands as a coding-focused successor to K2.6 with a 256K context window and a forced-thinking mode.
Moonshot · Kimi K2.7 Code
Anthropic enables background agent tasks up to 60 minutes inside Claude Code.
Anthropic · Claude Opus 4.7
OpenAI exposes shareable tool registries for the Responses API.
OpenAI · GPT-5.5
Image generation inside Gemini 3 switches to Imagen 4 with stricter prompt adherence.
Google DeepMind · Gemini 3 Pro
| Model | Provider | Code | Image / Video | Writing / Agents | Released |
|---|---|---|---|---|---|
| Claude Opus 4.7Default for coding and agent loops in Claude Code. Claude Code fast mode (/fast) runs Opus with faster output rather than a smaller model, available on Opus 4.6 and 4.7. | Anthropic | 92 Leads SWE-Bench Verified and holds the longest reliable agent loops in coding.↗ | — | 90 Top-tier long-form prose, strongest tool-use compliance across published evals.↗ | |
| Claude Sonnet 4.6Mid-tier Claude; default on claude.ai Free and Pro, 1M-token context in beta. | Anthropic | 85 Full upgrade over Sonnet 4.5 on coding and agent planning; trails Opus 4.7 on the hardest multi-file work.↗ | — | 86 Strong long-context reasoning and stable tool use at a lower price than Opus.↗ | |
| Claude Haiku 4.5Fastest, lowest-cost current Claude; the budget pick for high-volume agents. | Anthropic | 76 Fast and cheap for routine edits; not the pick for hard repo-scale agent loops.↗ | — | 74 Reliable for short-form and high-throughput tool use; less depth on long-form editorial.↗ | |
| GPT-5.6OpenAI's current default; three variants Luna, Terra, Sol (Sol most capable). GPT-Live adds simultaneous listen-and-speak voice. | OpenAI | 89 Strongest OpenAI coder to date; still trails Claude Opus 4.7 on the longest agent loops.↗ | 83 Best general-purpose image generation; long-form video remains a separate surface.↗ | 88 Reliable agent runner with the broadest tool and voice integrations.↗ | |
| GPT-5.5Superseded by GPT-5.6 (Jul 2026) as OpenAI's default; still widely deployed. | OpenAI | 87 Competitive on SWE-Bench, weaker on multi-file refactors than Claude Opus 4.7.↗ | 82 Best general-purpose image generation; video remains Sora-2 surface.↗ | 86 Reliable agent runner, strong function calling, slightly looser editorial voice.↗ | |
| Sora 2Video-first; 60s coherent clips, audio bed. | OpenAI | — | 91 Best long-form video coherence and prompt adherence in published comparisons.↗ | — | |
| Gemini 3 ProStrong multimodal context, integrated Imagen 4 and Veo 3 routing. July 2026 refresh adds the Gemini 3.6 Flash workhorse and Nano Banana 2 image model; a Gemini 3.5 Pro flagship has slipped repeatedly. | Google DeepMind | 84 Good repo-scale reasoning, behind Claude on agent loop stability.↗ | 88 Imagen 4 plus Veo 3 combination is the most consistent image-plus-video pair.↗ | 84 Long-context wins on research-style tasks; weaker tool-call discipline.↗ | |
| Copilot SparkWraps OpenAI GPT-5.x with Microsoft IDE and Office tooling. | Microsoft | 83 Best IDE-integrated experience; raw model behind Claude on hard tasks.↗ | 70 Uses GPT image generation under the hood; trails dedicated image leaders.↗ | 80 Solid for Office-tethered workflows; less flexible as a standalone agent.↗ | |
| Grok 4.5xAI's most capable model; mixture-of-experts, 500K context, roughly $2/$6 per million tokens, available in Cursor. Musk positions it as comparable to Opus 4.7 but faster (vendor claim, unverified). | xAI | 85 Big step up on coding and agents; independent SWE-Bench confirmation still pending.↗ | 74 Aurora image gen is competent; no first-party long-form video.↗ | 82 Strong with real-time data and tool use; long-form structure trails the top tier.↗ | |
| Grok 4Superseded by Grok 4.5 (Jul 2026); real-time X integration, looser safety posture. | xAI | 78 Improving fast; SWE-Bench still trails the top three.↗ | 74 Aurora image gen is competent; no first-party long-form video.↗ | 76 Strong with real-time data, weaker on stable long-form structure.↗ | |
| Llama 4 405BOpen-weights flagship; the practical pick for self-hosting. | Meta | 80 Best open-weights coder; closes the gap to GPT-5.5 on SWE-Bench Lite.↗ | 68 Companion Emu 3 model lags Imagen 4 and Midjourney 7.↗ | 78 Long-form is verbose but stable; tool-use is recent and improving.↗ | |
| DeepSeek V4V4 Pro is a 1.6T-parameter mixture-of-experts model (about 49B active), V4 Flash is a lighter sibling. Official launch mid-July 2026 with new peak-hour API pricing; supersedes the R2 preview. | DeepSeek | 85 Excellent on competitive-programming and SWE-Bench Lite; strong value coder.↗ | — | 81 Strong reasoning, terser prose; cost-effective for batch agents.↗ | |
| DeepSeek R2Superseded by DeepSeek V4 (Jul 2026); the R2 preview is no longer the current public tier. | DeepSeek | 85 Excellent on competitive-programming and SWE-Bench Lite; weaker on long agent loops.↗ | — | 80 Strong reasoning, terser prose; cost-effective for batch agents.↗ | |
| Mistral Large 3Strong European hosting and data-residency story. A new open-weight frontier family entered partner early access in July 2026 (name and benchmarks not yet disclosed). | Mistral | 76 Improved on Large 2; still behind the top three on agent loops.↗ | — | 79 Tight, neutral prose; tool use is reliable but not best-in-class.↗ | |
| GLM-5.2Open-weight, roughly 753B parameters (about 40B active) with a 1M-token context window. Widely rated the most capable open-weight model at release; strong on Chinese-language tasks. | Z.ai | 79 Best open-weight coder alongside Llama 4; long-context handling improved sharply over GLM-5.↗ | — | 76 Solid agent runner; English long-form still trails the closed top tier.↗ | |
| Kimi K2.7 CodeOpen-sourced coding-focused successor to K2.6, with a 256K context window and a forced-thinking mode; long-context specialist lineage. | Moonshot | 80 Long-horizon coding and repo-wide reading; competitive on agentic benchmarks after the K2.7 upgrade.↗ | — | 78 Excellent at synthesising large corpora; looser prose voice.↗ | |
| Midjourney v7Image only; no chat or tool use. | Midjourney | — | 90 Best aesthetic image fidelity; weaker on strict prompt adherence.↗ | — | |
| Runway Gen-4Video editing primitives plus generation. | Runway | — | 86 Strong on directed edits and motion control; behind Sora 2 on raw coherence.↗ | — |
No models match those filters.
OpenAI's first-party coding agent for CLI, IDE, cloud, and GitHub workflows.
Strongest agentic SWE-Bench results in the matrix; fast mode (/fast) speeds Opus output.
Best multi-model IDE; lets you swap providers per task.
Tight VS Code and PR-review integration.
Best terminal-native option; strongest cost control.
VS Code extension with explicit plan/act split.
Cline fork with multi-mode workflow.
Directional ratings curated from public benchmarks, model cards, and hands-on use. Scores are 0–100 within a bucket, not across buckets. Each rating links to the primary evidence. Refreshed after any model release. Not authoritative; AI-assisted editorial.