[JJ’s LLM Insight #2] The AI Coding Agent Market – Have you ever seen a market shift this fast? I certainly haven’t.

Written in

by

 

In Part 1, we explored the structural differences between frontier and open-weight models, and why open-weight matters for data sovereignty. This time, we’re diving into where those models actually do the work: the AI coding agent market.


The Coding Agent Market Right Now

The first half of 2026 brought a seismic shift to the AI coding agent landscape.

According to JetBrains’ 2026 Developer Ecosystem Survey (May-July, 15,000+ respondents), Claude Code is now the most widely adopted AI coding tool, used by 39% of professional developers at work worldwide. Just six months earlier, in January, it was at 18%.

Meanwhile, GitHub Copilot – the longtime leader – dropped from 29% to 21% in the same period, surrendering the top spot.

ToolJan 2026May-Jul 2026Change
Claude Code18%39%+21pp
GitHub Copilot29%21%-8pp
OpenAI Codex3%16%+13pp
Cursor18%12%-6pp
OpenCode–7%–
JetBrains AI–~9%–

(Source: JetBrains AI Pulse (Jan 2026) and Developer Ecosystem Survey 2026 (May-Jul). Respondents could select multiple tools, so totals exceed 100%.)

I don’t think I’ve ever witnessed this pace of change before. Claude Code jumped from 3% in mid-2025 to 39% in just 18 months – a 36 percentage point leap. In the US, that number hits 47%. And 31% of Claude Code users call it their “primary AI coding tool” – that’s an 80% conversion rate from user to primary tool.

Here’s the thing worth noting. Microsoft still owns the massive world that is GitHub. Out of GitHub’s 225 million users, 50 million use Copilot, and over 90% of Fortune 500 companies are on GitHub. But “distribution” and “choice” are different things. The numbers prove it. Personally, I don’t use Copilot anymore either.


What Happened in Six Months?

I felt it happening, but honestly, I was still surprised. A market leader changing in six months is rare even in SaaS history. And this happened against Copilot, which has Microsoft’s massive distribution channel behind it.

Let me share a few thoughts on what’s going on here.

First, we hit the tipping point from “autocomplete” to “agent.”

Copilot started as autocomplete. You type, it suggests the next bit of code. But Claude Code was designed as an agent from day one. Say “fix this bug,” and it finds the files, modifies the code, runs the tests, and if they fail, fixes and retries – all autonomously in a loop.

What developers wanted wasn’t “faster typing” – it was “delegation of work.” According to JetBrains, 90% of professional developers now use AI coding agents at least weekly, and 68% use them daily. This isn’t experimental behavior anymore. The shift in how developers work has already happened.

Second, satisfaction beat distribution.

Claude Code recorded the highest satisfaction scores from the start (91% CSAT, NPS 54). And 80% of users converted it to their “primary tool.” That conversion rate is the key.

Copilot dominates in distribution (50 million users), but for many developers, it’s just “the tool that comes pre-installed.” The tool they actually choose to use is different. When JetBrains asked developers with 10+ years of experience which tool they’d “want to use daily,” 46% picked Claude Code. Copilot got 9%.

Third, the Cursor paradox is fascinating.

Cursor’s revenue grew 4x from $1B (Nov 2025) => $4B (Jun 2026), but market share actually dropped from 18% => 12%. How does revenue go up while share goes down?

The answer: “existing users are using it more.” Cursor shifted to usage-based pricing, and power users on the $20 plan are spending $60-$200/month. But the rate of attracting new developers fell behind Claude Code. According to Ramp data, Cursor’s share of AI coding tool spending dropped from 41% (Jun 2025) to 26% (May 2026), and Anthropic captured half of that loss.

Fourth, what this change means

The competitive axis in this market has shifted from “code completion UX” to “agent capability.” To put it bluntly: expectations moved from “a tool that helps me type” to “a tool that does work when I tell it to.”

But here’s what we need to address. Claude Code being #1 means full lock-in to Anthropic. When Copilot was beating Cursor, it meant lock-in to GitHub/Microsoft. Even as the market leader changes, the fundamental problem of “being locked into a single vendor” hasn’t changed.

That’s the core issue I want to address in this article.


Now let’s look at what these tools actually do – and what they can’t do.


What Each Tool Actually Does – And Doesn’t Do

GitHub Copilot – Stepping Back from the Frontline

Copilot is no longer simple autocomplete. It’s become a full development platform with agent mode, coding agents, CLI, and code review. As of September 2026, it supports GPT-6 Astra, Claude Fable 5.1, and Gemini 3.8 Flash.

Key Numbers:

  • 50 million users (Jul 2026, Microsoft earnings)
  • 4.7 million paid subscribers (Jan 2026, 75% YoY growth)
  • Estimated ARR: $900M-$1.1B
  • 46% of newly pushed code on GitHub is AI-generated (GitHub’s own claim)

Pricing (shifted to usage-based “GitHub AI Credits” on June 1, 2026):

PlanPriceIncluded AI Credits
Free$0Limited (2,000 completions/month)
Pro$10/mo$15 ($10 base + $5 flex)
Pro+$39/mo$70
Max$100/mo$200
Business$19/user/mo$19 (org-wide pooling)
Enterprise$39/user/mo$39 (org-wide pooling)

(Source: GitHub official pricing page, Sep 2026)

Strengths:

  • Seamless GitHub ecosystem integration
  • IP indemnification (crucial for enterprise)
  • Lowest barrier to entry (from $10/mo)
  • Enterprise admin permission system (GA Sep 2026)
  • Multi-model support (GPT-6 Astra, Claude, Gemini, etc.)

Limitations:

  • Agent mode lags behind Claude Code – JetBrains shows Copilot usage at half of Claude Code’s level
  • Unpredictable costs after June 2026 usage-based pricing switch
  • Promotional credits ended Sep 1 – Business lost 37%, Enterprise lost 44% of included credits

It helps. But it doesn’t execute on its own. This isn’t a version problem – it’s a fundamental design philosophy difference.

Cursor – The Fastest-Growing SaaS in History, Now in SpaceX’s Hands

Cursor forked VS Code and put AI at the center of the editor. Its Composer agent modifies multiple files simultaneously, executes terminal commands, and autonomously loops through fixing errors. In 2026, they added the ability to run up to 8 parallel agents simultaneously.

Key Numbers:

  • ARR growth: $100M (Jan 2025) => $1B (Nov 2025) => $2B (Feb 2026) => $3B (Apr 2026) => $4B (Jun 2026)
  • Acquired by SpaceX for $60B (Aug 14, 2026) – the largest startup acquisition in history
  • Previous valuation: $29.3B (Nov 2025 Series D)
  • Used by 64%+ of Fortune 500
  • Workplace adoption: 12% (May-Jul 2026, down from 18% in Jan)

Pricing:

PlanPriceFeatures
HobbyFreeLimited Agent requests
Pro$20/moExtended Agent limits, Cloud Agents
Pro+$60/mo3x Pro usage
Ultra$200/mo20x Pro usage
Teams$40-$120/user/moSSO, central management

(Source: Cursor official pricing page, Sep 2026)

Strengths:

  • Best-in-class multi-file editing
  • Multi-model support (Claude, Gemini, Grok, proprietary models)
  • Codebase indexing – understands the entire repo
  • Cloud Agents – run agents in the cloud
  • “Projects” – coordinator that orchestrates multiple sub-agents (Sep 2026)

Limitations:

  • VS Code fork means no JetBrains, Vim, etc.
  • OpenAI cutting off Cursor’s model access on Nov 12, 2026 (post-SpaceX acquisition)
  • $20 Pro plan, but typical agent users spend $60-$100/month
  • Power users exceed $200/month
  • Market share declining (18% => 12%)

(Source: Second Talent, Forbes, Bloomberg, CNBC, TechCrunch)

A “$20 plan” no longer means “spending $20.”

Claude Code – The New #1, But Fully Locked to Anthropic

Claude Code is an autonomous coding agent that runs in the terminal. It works from the CLI rather than a browser or IDE, reading and writing files, executing shell commands, running tests, managing git, and looping autonomously until the job is done.

Key Numbers:

  • Workplace adoption: 39% (May-Jul 2026), 47% in the US
  • Run-rate revenue: $2.5B+ (Feb 2026)
  • Anthropic total annualized revenue: $65B (end of Jul 2026)
  • User satisfaction: 91% CSAT, NPS 54 – highest among all surveyed
  • 4% of all GitHub commits made with Claude Code (Feb 2026, 2x vs. previous month)

Pricing:

PlanPriceUsage
Free$0Claude Code not included
Pro$20/mo ($17 annual)Base (5+ Free per 5-hour session)
Max 5x$100/mo5x Pro
Max 20x$200/mo20x Pro
Team Standard$25/user/mo1.25x Pro
Team Premium$125/user/mo6.25x Pro

(Source: Anthropic official pricing page, Sep 2026)

Strengths:

  • Full codebase analysis capability
  • MCP integration for external tools
  • Sub-agent support – splits complex tasks
  • Headless mode – auto-runs in CI/CD pipelines
  • Highest user satisfaction (91% CSAT)
  • May 2026: Claude Code limits doubled after SpaceX compute deal

Limitations:

  • Terminal-based – high barrier if you’re not comfortable with CLI
  • No inline autocomplete – no suggestions as you type
  • Only Anthropic models – can’t switch to other models
  • No offline/local model support – can’t use in air-gapped environments
  • Performance degrades in long sessions (context rot)
  • Claude Code and regular chat share the same usage pool – heavy coding reduces chat limits
  • Enterprise average cost: $150-$250/month per active developer

(Source: JetBrains, Anthropic, Bloomberg, Financial Times)

It’s true that Claude Code solves the hardest problems. But that capability is completely locked to Anthropic’s cloud.

OpenAI Codex – The ChatGPT Ecosystem’s Coding Agent

OpenAI Codex showed the sharpest growth, jumping from 3% to 16% market share. It started as a CLI in April 2025 and has since become a comprehensive coding agent supporting terminal, desktop app, IDE extension, cloud service, and mobile app.

At DevDay September 2026, OpenAI announced it “can run locally, remotely from your phone, or entirely in the cloud.” They also added voice commands, multi-task management via the /agents view, and session resume.

Key Numbers:

  • Workplace adoption: 16% (May-Jul 2026, up from 3% in Jan, +13pp)
  • GitHub stars: 120K+, 474 contributors
  • Current models: GPT-6 Astra (GA-Super hottest trend now), GPT-5.6 Sol/Terra/Luna
  • Apache-2.0 license (commercially usable)

Pricing (included with ChatGPT subscription):

PlanPriceCodex Usage (5-hour window)
Free$0Limited
Plus$20/moBase (10-100 Sol)
Pro 5x$100/mo5x Plus
Pro 20x$200/mo20x Plus
Business$20/user/moSame as Plus

(Source: OpenAI official pricing page, Sep 2026)

Strengths:

  • ChatGPT subscribers can use at no extra cost
  • Multiple interfaces: CLI, desktop app, VS Code extension, iOS app, cloud
  • Codex Cloud – reusable cloud dev environments, shared team configs
  • Codex Security Cloud – vulnerability scanning and auto-fix suggestions
  • Auto PR review on GitHub/GitLab (@codex review)
  • Voice commands to start and steer tasks

Limitations:

  • OpenAI models only – can’t switch to other models
  • 5-hour rolling window limit – hit the cap mid-session and you wait
  • April 2026: Plus plan usage reduced (per OpenAI release notes)
  • Business plan has same usage as Plus – may be insufficient for enterprise
  • No offline/air-gap support (except Ollama local mode)

(Source: OpenAI DevDay 2026, OpenAI official docs, AIToolsReview)

Codex’s biggest advantage is that ChatGPT subscribers can use it immediately. But that also means lock-in to the OpenAI ecosystem.

OpenCode – The Open Source Alternative

While all the tools above are commercial products, OpenCode takes a completely different approach. It’s an open-source coding agent released under MIT license.

Key Numbers:

  • GitHub stars: 205K+ (surpassing Claude Code’s 122K)
  • Monthly active developers: 16 million
  • Workplace adoption: 7% (May-Jul 2026, JetBrains survey)

Strengths:

  • Model-agnostic: 75+ LLM providers supported – Anthropic, OpenAI, Google, Bedrock, and local models via Ollama
  • Air-gap support: Run completely locally with Ollama
  • Plan/Build modes: Tab to switch – Plan is read-only planning, Build executes
  • AGENTS.md: Document project-specific context and conventions
  • MCP integration: Connect external tools
  • GitHub Actions integration: Auto-run in CI/CD pipelines
  • Free: The agent itself is free, you just pay model costs directly

Limitations:

  • January 2026: Anthropic blocked Claude Code credential tunneling, so no direct Anthropic API access
  • Less polished experience compared to commercial tools
  • Relies on community support

(Source: GitHub, JetBrains Developer Ecosystem Survey, AWS blog)

OpenCode’s real significance is proving that “coding agents don’t have to be locked to a specific vendor.” But leveraging that flexibility requires significant setup and operational expertise.

Open-Weight Models with Coding Specializations

The open-weight models we covered in Part 1 are also releasing coding-specialized versions. Moonshot AI announced Kimi K2.7 Code, and MiniMax released the M2.7 coding agent. These models share a common thread: delivering frontier-level coding performance in open-weight form.

But a great model alone doesn’t solve coding agent problems. The model is the engine. How you harness that engine is what matters.


The Problems Everyone Shares, Nobody Has Solved

Each of these tools has different strengths. But some fundamental problems remain unsolved by all of them.

1. The Limits of Context Management – “A bigger window doesn’t fix the problem”

A million-token context window didn’t make the problem disappear. Failures just got quieter.

Anthropic themselves acknowledge this: “Model performance degrades as context grows. Attention spreads across more tokens, and old irrelevant content starts interfering with current tasks.” They call this “context rot.”

Every tool’s response to this problem is compaction – summarizing old conversations to free up space. But compaction has a fatal paradox:

  1. Compaction is inherently lossy. What the summarizer deems unimportant vanishes.
  2. Worse, compaction runs when the agent is at its weakest. It’s deciding “what to throw away” while already degraded by context rot.

If the variable name needed at turn 200 got discarded during turn 50’s compaction? The agent either re-derives it or just makes it up.

2. Vendor Lock-in – “Context is the stickiest layer”

Each tool stores session state, compacted context, and work history on its own servers. There’s no standardized export format for this data.

Why does this matter? Switch tools and you start from scratch. Months of accumulated project context, the agent’s learned understanding of your codebase, approved patterns – all of it stays on that vendor’s servers.

“Context is the stickiest layer. The more you use it, the higher your exit cost climbs at the same rate.”

Claude Code only supports Anthropic models. Cursor’s OpenAI access will be cut off after the SpaceX acquisition. Copilot is starting to support other models, but core infrastructure remains tied to GitHub/Microsoft.

3. Cost Opacity – What a “$20 plan” really means

The AI coding agent market is shifting from per-seat fixed pricing to token-based variable pricing. This seems user-friendly but actually makes cost prediction extremely difficult.

  • Cursor Pro ($20/mo): Typical agent users spend $60-$100/mo, power users exceed $200
  • Claude Code: Enterprise average $13/day per active developer, $150-$250/mo (Anthropic official docs)
  • GitHub Copilot: Promotions ended Sep 1, 2026 – Business lost 37% of included credits, Enterprise lost 44%

There’s also the concept of “Verification Tax” – the additional cost of verifying agent-generated code. Review time, test execution, fix-up work. None of this shows up on the tool’s price tag.

4. Security and Data Sovereignty

Security issues hit harder with coding agents. These tools don’t just send prompts – they analyze your entire codebase and modify files.

  • All major tools require cloud connectivity
  • Session logs and context stored on vendor servers
  • No true air-gap (closed network) support (except OpenCode + Ollama)
  • Limited self-hosting options

Claude Code explicitly states “no offline/local model support.” Cursor’s Cloud Agents run on Cursor infrastructure. For environments where code cannot leave – finance, defense, public sector – none of these tools are even an option.

5. The Reliability Paradox – The gap between “code generation” and “production deployment”

Recent research shows that while AI coding tools deliver meaningful productivity gains in coding activity itself, those gains drop sharply in the transition from writing code to deploying reliable software.

Agents “cannot verify their own work” – this is a structural limitation. Tests passing is evidence, not proof that requested behavior is correct. Agents can infer the wrong package manager, run incomplete test commands, misunderstand errors, weaken tests, and modify unrelated config files.

“The key research question is shifting from ‘how much code can agents generate’ to ‘how much production-verified value can engineering systems deliver per dollar, per reviewer hour, per unit of operational risk.’”


So What Do We Actually Need?

Summarizing the problems we’ve covered:

  1. Context Management: Bigger windows aren’t the answer. Intelligent context engineering is needed.
  2. Vendor Lock-in: Sessions, context, and history need to be portable.
  3. Model Flexibility: Ability to choose different models for different tasks.
  4. Cost Transparency: Predictable cost structures.
  5. Data Sovereignty: Self-hosting and air-gap options required.
  6. Reliability: Pipelines that go from generation to verified deployment.

Kimchi Harness Platform – CAST AI’s Approach to These Problems

CAST AI’s Kimchi is an AI coding agent harness platform designed to tackle the limitations listed above head-on.

Multi-Model Orchestration

You’re not locked to one model. Choose different models based on the nature of the work – fast and cheap models for exploration, strong reasoning models for complex architecture decisions, coding-specialized models for code generation. It’s OpenCode’s proven “model agnosticism” implemented as a production-grade harness.

Sub-Agent Architecture to Prevent Context Pollution

Explore, Plan, Build, and Review are separated into different agents. Each agent works with only the context needed for its role. This structurally avoids the “needed turn 50’s info at turn 200 but it got compacted away” problem.

Teleport – From Local to Cloud, Seamlessly

Upload a session started locally to a cloud sandbox and continue working. Close your laptop and the agent keeps running in the cloud. Session state, context, and history travel with it. (This feature is covered in detail in Part 4.)

Air-Gap Installation Support

Install and operate Kimchi even in environments where internet connectivity isn’t allowed – finance, defense, public sector. Combined with open-weight models, your data never leaves your organization’s infrastructure. An option that Claude Code, Cursor, and Copilot none of them provide. (This feature is covered in detail in Part 5.)

Native Integration with Open-Weight Models

Run the open-weight models we covered in Part 1 – Kimi, MiniMax, GLM, DeepSeek, Nemotron – directly on the Kimchi harness. A practical alternative when frontier model API costs are prohibitive or data sovereignty is the priority.

For more details, check out the official documentation at https://docs.kimchi.dev/.


In the Next Part

This part covered the current state of the AI coding agent market and the unsolved problems that major tools share. Next, we’ll dig deeper into “the multi-model dilemma” – why going all-in on a single model is risky, and what it actually means to route models by task.


Sources and References

Market Share and Adoption:

  • JetBrains AI Pulse Survey (Jan 2026)
  • JetBrains Developer Ecosystem Survey 2026 (May-Jul 2026, 15,000+ respondents)

Company Data:

  • Microsoft FY26 Q4 Earnings Call (Jul 29, 2026)
  • Anthropic official announcements, Bloomberg, Financial Times
  • Second Talent AI Coding Statistics
  • Forbes, CNBC, TechCrunch, Reuters

Pricing Information:

  • GitHub Copilot official pricing page (Sep 2026)
  • Anthropic Claude official pricing page (Sep 2026)
  • Cursor official pricing page (Sep 2026)

(Prices and features change rapidly – please verify the latest information before making adoption decisions.)

Tags

Categories

댓글 남기기

JJ's technical blog

A blog about the messy, fascinating intersection of Kubernetes and AI – where infrastructure decisions increasingly shape whether AI initiatives actually succeed. Drawing on hands-on experience with enterprise customers across Korea and APAC, this space covers everything from cluster optimization to the evolving demands AI workloads place on cloud-native systems.

Seoul,
South Korea

JJ's technical blog에서 더 알아보기

지금 구독하여 계속 읽고 전체 아카이브에 액세스하세요.

계속 읽기