AI Coding Isn't a Silver Bullet — Don't Rush to Call It a Productivity Win
A few months ago, Sundar Pichai said 75% of new code at Google is now AI-generated, up from 50% just six months earlier. At OpenAI, engineers reportedly get ~99% of their output tokens from Codex now. Typing your own code has become almost symbolic.
Numbers like that made me think AI had basically caught up to — maybe even surpassed — most engineers.
We've been running AI Coding hard in production for about six months now. And I had the same reaction at first. Then I watched what actually happens when you drop these agents into a real, aging, high-stakes enterprise system. The story gets a lot more interesting.
Here's the short version: AI didn't reveal that coding is solved. It revealed everything that was never really about coding in the first place.
The Honeymoon Was Real
I'm not here to talk AI down. The honeymoon phase happened, and it was legitimate.
Work that used to take weeks — requirements analysis, research, design discussions, actual coding — started getting done in a day or two. We had an agent produce 20-30K lines of C++ for a real feature. Some self-contained problems went from "a sprint" to "a couple of hours." And it wasn't just fast — it was giving us solution designs that looked genuinely professional before a human ever touched them.
It's easy to look at that and think: okay, humans propose, AI implements, we're basically done here.
We were wrong. Or rather — we were only seeing half the picture.
The Two-Day Feature That Took Two More Weeks
We asked an agent (running on a frontier GPT model) to extend some of DolphinDB's Join operators to support cross-time-partition computation — needed for 24/7 markets like crypto, where data doesn't cleanly split by calendar day the way it does for traditional trading hours.
In under two working days, the agent analyzed the existing codebase, designed a solution, implemented it, wrote tests, and generated documentation. This is work that would normally take a human team weeks. Initial tests looked clean.
Then we widened the test coverage. And the cracks showed.
Not "wrong code" cracks — the logic and design were basically sound. The problem was that the agent didn't know about the dozens of scattered constraints living inside hundreds of thousands of lines of legacy code: partitioning quirks, unsupported data types, capabilities that exist in SQL but not in the functional API, old limitations nobody wrote down anywhere. We ended up finding dozens of issues across several categories.
I want to be clear about what conclusion I didn't draw from this: "AI still isn't good enough." What I actually concluded was closer to — AI is already good enough that we now have to fix the rest of the system around it.
Three Things Speed Exposed
Problem 1: We sped up coding, not delivery.
Coding was never the whole pipeline — review, testing, release were always there too. When coding took most of the time, speeding it up mattered a lot. Now that AI can compress days into hours, those other stages become the bottleneck. If AI cuts dev time from 5 days to 1, but review still takes 2 and testing still takes 3, you haven't gained 5x — you might have made things worse, because now there are more PRs stacking up against the same review capacity. This is Amdahl's Law playing out in software engineering, and it means the real ROI conversation has to move from "coding efficiency" to "delivery efficiency" — CI/CD, review, and release all need to level up too.
Problem 2: The bottleneck isn't hallucination. It's institutional knowledge.
It's tempting to blame every enterprise AI failure on "hallucination." That's not really what we saw. The agent's research, structure, and reasoning were often perfectly sound — just incomplete, because the real constraints of a mature system were never written down anywhere. They live in old interface decisions, defensive code from a past incident, and tribal knowledge. A human engineer at least knows where the landmines probably are. An agent is a brilliant new hire with zero onboarding — great on a clean, isolated module, and blind to the debt buried in a decade-old system. This isn't a "bigger context window" problem. It's a discipline in its own right — I'd call it Environment Engineering, not just Context Engineering.
Problem 3: Cheap code generation makes reinventing the wheel dangerously easy.
What a mature company actually has isn't "code" — it's production-hardened databases, engines, SDKs, and components built over years. A good engineer defaults to reusing them. An agent that doesn't know they exist will just quietly rebuild them from scratch.
Three Conclusions We're Actually Building On
After deeply using AI Coding, our takeaway isn't "AI is overhyped" — quite the opposite. We've become increasingly convinced that AI's capabilities are already strong enough to substantively change how software is produced. But it's precisely because AI is so capable that it exposes a deeper truth: code has only ever been one part of software production, and what enterprises truly need to improve goes far beyond just writing code.
From what we've seen, at least three conclusions are worth rethinking.
From Coding Efficiency to Delivery Efficiency
The easiest things to quantify in AI Coding are code volume, token consumption, PR counts, and development time. But none of these metrics naturally translate into the business outcomes a company actually cares about.
Imagine AI lets one engineer open ten PRs a day, while code review, testing, and release can still handle only two. The company doesn't gain 10x delivery capacity — it gains eight PRs piling up in a queue.
So the ROI of AI Coding ultimately has to come back to end-to-end software delivery efficiency. That means enterprises must upgrade automated testing, code review, CI/CD pipelines, release strategy, and quality verification in lockstep, so that the code output AI generates actually converts into deliverable software.
In short: AI Coding solves "producing code." What enterprises need to solve is "digesting code."
Building an Engineering Environment Agents Can Understand
When an agent faces a complex enterprise system, the problem usually isn't that it can't do the work at all — it's that it can't grasp the implicit constraints scattered across code, documentation, and people's experience.
Enterprises need to re-examine their own engineering environment: making architecture clearer, components easier to discover, knowledge easier to search, historical experience reusable, engineering standards explicit, tools callable, and permission boundaries transparent to agents. These things used to exist to make human collaboration easier — now they also determine whether an agent can actually get work done. So we believe what enterprises need to build is no longer just traditional Context Engineering, but a more complete Agent Engineering Environment.
This is also a core consideration behind how we built DolphinX.
DolphinX isn't simply about giving an enterprise "one more coding agent." It's about providing the environment an agent needs to actually work inside an enterprise: knowledge that can be accumulated and retrieved, repetitive work that can be packaged into Skills, historical interactions that can form Memory, existing enterprise capabilities that can be invoked through mechanisms like MCP, and user identity, permissions, and execution processes that can be brought under unified governance.
This way, when an agent encounters an enterprise system, it doesn't need to start understanding everything from zero each time, and it doesn't need to duplicate the enterprise's existing knowledge and capabilities into some external AI system.
In the past, engineers adapted to the enterprise's software environment. In the AI era, we also need the software environment itself to be understandable by agents.
Infrastructure Is Leverage, Not Something AI Replaces
AI generating SQL doesn't replace the database. Petabyte-scale storage, real-time stream processing, distributed compute — these still need real, hardened infrastructure, not a clever prompt. What AI is genuinely great at is understanding intent and orchestrating the capabilities that already exist.
Take DolphinDB as an example. It has been production-validated over many years across finance, energy, industrial, and other sectors, accumulating a mature data model, a library of computation functions, a stream-processing framework, and distributed execution capability. These aren't things a few lines of code can replace — they're infrastructure an enterprise has gradually refined through real business use. An agent doesn't need to, and shouldn't, "reimplement" a compute engine — it just needs to know these capabilities exist, understand how to invoke them, and put its energy into understanding user intent and composing existing capabilities.
Closing Thoughts: When Code Is No Longer the Moat
If implementation is getting cheap, where's the value going?
My honest take: it's not disappearing; it's moving up. As "writing the code" gets commoditized, judgment, architecture, experience, and accountability get more valuable, not less. The winners in this next phase won't be the teams whose AI writes code the fastest. They'll be the teams with the best engineering environment, the strongest infrastructure, and the deepest understanding of their own business.