Skip to main content
Tech AI Pulse

AI Pulse

Directional ratings for the AI models people are actually using right now. Three buckets — code, image and video, writing and agents — refreshed after every release.

Updated 25 July 2026 18 models tracked Methodology
Coding Best right now

Claude Opus 4.7

Anthropic

Default to Claude Opus 4.7 inside Claude Code for repo-scale edits. GPT-5.6 via Codex is the strongest first-party OpenAI alternative; pick Cursor when you need multi-model routing (including Grok 4.5) in the same workflow.

Runner-up GPT-5.6
Image / Video Best right now

Sora 2

OpenAI

Sora 2 for video, Midjourney v7 for still images. Imagen 4 and the Nano Banana image line inside Gemini are the best one-surface answer when you want both from a single account.

Runner-up Midjourney v7
Writing / Agents Best right now

Claude Opus 4.7

Anthropic

Claude Opus 4.7 leads for long-form editorial and stable agent loops. GPT-5.6 wins when you need broader tool integrations and voice, especially around Microsoft and OpenAI surfaces.

Runner-up GPT-5.6

Latest releases

17 entries · newest first
  1. New model

    Google ships the Gemini 3.6 Flash family

    Gemini 3.6 Flash (workhorse), 3.5 Flash-Lite and a 3.5 Flash Cyber security variant land, plus Nano Banana 2 Lite for images; the Gemini 3.5 Pro flagship slips again.

    Google DeepMind · Gemini 3 Pro

  2. New model

    DeepSeek V4 graduates to general availability

    DeepSeek promotes V4 Pro and V4 Flash out of preview with new peak-hour API pricing; V4 supersedes the R2 reasoning preview.

    DeepSeek · DeepSeek V4

  3. New model

    OpenAI publicly launches GPT-5.6

    GPT-5.6 rolls out in Luna, Terra and Sol variants after a limited government-gated preview, alongside the GPT-Live real-time voice models.

    OpenAI · GPT-5.6

  4. New model

    xAI releases Grok 4.5

    xAI ships Grok 4.5, a mixture-of-experts model with a 500K context window focused on coding and agents, available in Cursor and Grok Build.

    xAI · Grok 4.5

  5. Capability

    Mistral opens early access to a new open-weight frontier family

    Mistral confirms a new open-weight model family in partner early access, with a broader release expected later in the summer.

    Mistral · Mistral Large 3

  6. New model

    Z.ai releases GLM-5.2 open weights

    GLM-5.2 ships as an open-weight model with a 1M-token context window, rated the strongest open-weight model available at launch.

    Z.ai · GLM-5.2

  7. New model

    Moonshot open-sources Kimi K2.7 Code

    Kimi K2.7 Code lands as a coding-focused successor to K2.6 with a 256K context window and a forced-thinking mode.

    Moonshot · Kimi K2.7 Code

  8. Capability

    Claude Code adds long-running background agents

    Anthropic enables background agent tasks up to 60 minutes inside Claude Code.

    Anthropic · Claude Opus 4.7

  9. Capability

    GPT-5.5 picks up custom tool registries

    OpenAI exposes shareable tool registries for the Responses API.

    OpenAI · GPT-5.5

  10. Version bump

    Gemini 3 Pro routes to Imagen 4 by default

    Image generation inside Gemini 3 switches to Imagen 4 with stricter prompt adherence.

    Google DeepMind · Gemini 3 Pro

Comparison matrix

Sort and filter; scores are within a bucket, not across.
Model Provider Code Image / Video Writing / Agents Released
Claude Opus 4.7Default for coding and agent loops in Claude Code. Claude Code fast mode (/fast) runs Opus with faster output rather than a smaller model, available on Opus 4.6 and 4.7. Anthropic 92 Leads SWE-Bench Verified and holds the longest reliable agent loops in coding. 90 Top-tier long-form prose, strongest tool-use compliance across published evals.
Claude Sonnet 4.6Mid-tier Claude; default on claude.ai Free and Pro, 1M-token context in beta. Anthropic 85 Full upgrade over Sonnet 4.5 on coding and agent planning; trails Opus 4.7 on the hardest multi-file work. 86 Strong long-context reasoning and stable tool use at a lower price than Opus.
Claude Haiku 4.5Fastest, lowest-cost current Claude; the budget pick for high-volume agents. Anthropic 76 Fast and cheap for routine edits; not the pick for hard repo-scale agent loops. 74 Reliable for short-form and high-throughput tool use; less depth on long-form editorial.
GPT-5.6OpenAI's current default; three variants Luna, Terra, Sol (Sol most capable). GPT-Live adds simultaneous listen-and-speak voice. OpenAI 89 Strongest OpenAI coder to date; still trails Claude Opus 4.7 on the longest agent loops. 83 Best general-purpose image generation; long-form video remains a separate surface. 88 Reliable agent runner with the broadest tool and voice integrations.
GPT-5.5Superseded by GPT-5.6 (Jul 2026) as OpenAI's default; still widely deployed. OpenAI 87 Competitive on SWE-Bench, weaker on multi-file refactors than Claude Opus 4.7. 82 Best general-purpose image generation; video remains Sora-2 surface. 86 Reliable agent runner, strong function calling, slightly looser editorial voice.
Sora 2Video-first; 60s coherent clips, audio bed. OpenAI 91 Best long-form video coherence and prompt adherence in published comparisons.
Gemini 3 ProStrong multimodal context, integrated Imagen 4 and Veo 3 routing. July 2026 refresh adds the Gemini 3.6 Flash workhorse and Nano Banana 2 image model; a Gemini 3.5 Pro flagship has slipped repeatedly. Google DeepMind 84 Good repo-scale reasoning, behind Claude on agent loop stability. 88 Imagen 4 plus Veo 3 combination is the most consistent image-plus-video pair. 84 Long-context wins on research-style tasks; weaker tool-call discipline.
Copilot SparkWraps OpenAI GPT-5.x with Microsoft IDE and Office tooling. Microsoft 83 Best IDE-integrated experience; raw model behind Claude on hard tasks. 70 Uses GPT image generation under the hood; trails dedicated image leaders. 80 Solid for Office-tethered workflows; less flexible as a standalone agent.
Grok 4.5xAI's most capable model; mixture-of-experts, 500K context, roughly $2/$6 per million tokens, available in Cursor. Musk positions it as comparable to Opus 4.7 but faster (vendor claim, unverified). xAI 85 Big step up on coding and agents; independent SWE-Bench confirmation still pending. 74 Aurora image gen is competent; no first-party long-form video. 82 Strong with real-time data and tool use; long-form structure trails the top tier.
Grok 4Superseded by Grok 4.5 (Jul 2026); real-time X integration, looser safety posture. xAI 78 Improving fast; SWE-Bench still trails the top three. 74 Aurora image gen is competent; no first-party long-form video. 76 Strong with real-time data, weaker on stable long-form structure.
Llama 4 405BOpen-weights flagship; the practical pick for self-hosting. Meta 80 Best open-weights coder; closes the gap to GPT-5.5 on SWE-Bench Lite. 68 Companion Emu 3 model lags Imagen 4 and Midjourney 7. 78 Long-form is verbose but stable; tool-use is recent and improving.
DeepSeek V4V4 Pro is a 1.6T-parameter mixture-of-experts model (about 49B active), V4 Flash is a lighter sibling. Official launch mid-July 2026 with new peak-hour API pricing; supersedes the R2 preview. DeepSeek 85 Excellent on competitive-programming and SWE-Bench Lite; strong value coder. 81 Strong reasoning, terser prose; cost-effective for batch agents.
DeepSeek R2Superseded by DeepSeek V4 (Jul 2026); the R2 preview is no longer the current public tier. DeepSeek 85 Excellent on competitive-programming and SWE-Bench Lite; weaker on long agent loops. 80 Strong reasoning, terser prose; cost-effective for batch agents.
Mistral Large 3Strong European hosting and data-residency story. A new open-weight frontier family entered partner early access in July 2026 (name and benchmarks not yet disclosed). Mistral 76 Improved on Large 2; still behind the top three on agent loops. 79 Tight, neutral prose; tool use is reliable but not best-in-class.
GLM-5.2Open-weight, roughly 753B parameters (about 40B active) with a 1M-token context window. Widely rated the most capable open-weight model at release; strong on Chinese-language tasks. Z.ai 79 Best open-weight coder alongside Llama 4; long-context handling improved sharply over GLM-5. 76 Solid agent runner; English long-form still trails the closed top tier.
Kimi K2.7 CodeOpen-sourced coding-focused successor to K2.6, with a 256K context window and a forced-thinking mode; long-context specialist lineage. Moonshot 80 Long-horizon coding and repo-wide reading; competitive on agentic benchmarks after the K2.7 upgrade. 78 Excellent at synthesising large corpora; looser prose voice.
Midjourney v7Image only; no chat or tool use. Midjourney 90 Best aesthetic image fidelity; weaker on strict prompt adherence.
Runway Gen-4Video editing primitives plus generation. Runway 86 Strong on directed edits and motion control; behind Sora 2 on raw coherence.

Coding agents

IDE wrappers and CLI agents. SWE-Bench Verified, higher is better.
  1. 1

    Codex

    OpenAI · default model: GPT-5.6 / GPT-5-Codex

    OpenAI's first-party coding agent for CLI, IDE, cloud, and GitHub workflows.

    75 SWE-B
  2. 2

    Claude Code

    Anthropic · default model: Claude Opus 4.7

    Strongest agentic SWE-Bench results in the matrix; fast mode (/fast) speeds Opus output.

    71 SWE-B
  3. 3

    Cursor

    Anysphere · default model: GPT-5.6 / Claude Opus 4.7 / Grok 4.5

    Best multi-model IDE; lets you swap providers per task.

    64 SWE-B
  4. 4

    GitHub Copilot

    GitHub · default model: GPT-5.6

    Tight VS Code and PR-review integration.

    58 SWE-B
  5. 5

    Aider

    Open source · default model: Pluggable

    Best terminal-native option; strongest cost control.

    55 SWE-B
  6. 6

    Cline

    Open source · default model: Pluggable

    VS Code extension with explicit plan/act split.

    52 SWE-B
  7. 7

    Roo Code

    Open source · default model: Pluggable

    Cline fork with multi-mode workflow.

    50 SWE-B

Methodology

Directional ratings curated from public benchmarks, model cards, and hands-on use. Scores are 0–100 within a bucket, not across buckets. Each rating links to the primary evidence. Refreshed after any model release. Not authoritative; AI-assisted editorial.

  • Scores are within a bucket. A 90 in code is not comparable to a 90 in image and video.
  • Each rating links to the primary evidence. Where no link exists, treat the score as editorial judgment.
  • This section is AI-assisted editorial. It is decision support, not an authoritative leaderboard.

Read the full methodology →