AI did not eliminate software maintenance
It exported it to every department
If you liked this piece, you should check out:
Specialty Tokens: is your AI transformation partner. I run this firm to assess, train and implement AI and agents across your teams successfully.
Community: Throw in questions and meet other readers in the chat. Learning and requesting new content together.
The person leading data at a large, publicly listed media company told me they were doing a lot of internal AI education.
People were experimenting. Teams were building. In his words, “We get a lot of vibe coding.”
He wanted to go all in.
Then I asked a less exciting question.
Who maintains what they build?
How does a prototype make its way back into the real system? Who owns it when the person who made it gets busy, changes roles, or leaves?
The conversation changed. He was not resistant. He was overwhelmed.
At first I thought this was the familiar gap between an AI pilot and production. I no longer think that goes far enough.
The company did not have too few AI experiments. It was about to own more software than it knew existed.
AI removed the queue
For most of the history of enterprise software, building was expensive.
A team wanted a tool. It wrote a specification, found a budget, waited for IT or engineering, and eventually received something. The queue was slow and frustrating, but it also rationed how much software a company could create.
AI removed much of that queue.
A marketer can build a lead-research agent. Finance can make a workflow that drafts close commentary. Operations can connect a form, a model, and the CRM before lunch. A founder can replace five handoffs with one agent and a Slack message.
This is real progress. It is also a change in who produces software inside a company.
AI did not turn everyone into an engineer. It made engineering consequences available to everyone.
The moment another person depends on the lead-research agent, it is no longer a personal experiment. It is software. It has inputs, permissions, failure modes, users, and an implied promise that it will work again tomorrow.
Nobody has to call it software for the obligation to appear.
The bottleneck moves downstream
We can already see this inside software engineering, where the measurements are unusually legible.
A 2026 NBER study of more than 100,000 GitHub developers found that autonomous coding agents were associated with 180% more commits. The increase fell to 50% at the project level and 30% at the release level.
That is not evidence that the agents failed. It is evidence that they succeeded at producing more than the surrounding system could absorb. Code still had to be reviewed, integrated, tested, released, and attached to something users wanted.
OpenAI described the same shift from inside an unusually aggressive experiment. Its team built a product with roughly a million lines of agent-generated code in about one-tenth the time it estimated manual development would have taken. Then code generation stopped being the main constraint.
Human QA became the bottleneck.
The team’s response was not to tell the model to try harder. It made the environment more legible. Logs, metrics, tests, architectural rules, review feedback, and recurring cleanup became part of the system the agent could use. Human corrections were turned into documentation, tooling, or enforced rules.
That is the part most companies are missing.
They are giving more people the ability to create, but they have not built a place for what those people learn to go.
A company learns only when the correction survives
Imagine a sales agent drafts a follow-up using an outdated pricing deck.
The account executive notices, fixes the email, and sends the correct version. The customer receives accurate information. The immediate problem is solved.
But unless that correction changes the source, the instruction, or a test, the company has learned nothing. The next person runs the agent and gets the same mistake.
Now imagine the correction updates the pricing source, adds a check, and reaches every future run. The next account executive does not need to discover the same failure.
That is organizational learning in a form you can inspect.
The point of using intelligence at work is not only to create more output. It is to create new intelligence the company can use next time.
So I would measure AI productivity differently.
Not by the number of prompts, pilots, agents, or hours people say they saved. Measure how often a correction becomes reusable.
The model makes the first version. The company earns the second.
Every department is becoming a software team
On the morning of my conversation with the data lead, I spoke to a successful founder whose company was already using AI for commercial work: branding, websites, ads, emails, follow-ups, and similarity search for clients.
Those use cases are attractive because the result is visible. You can point to the page, the ad, or the lead list.
Operations are different.
An ad is an output. An operational workflow is a dependency.
The moment the company depends on it, someone has to know which system is the source of truth, what the agent may change, where approval happens, how failure becomes visible, and when the workflow should be retired.
That work used to sit mostly with software and operations teams. Now it is spreading into sales, finance, marketing, support, legal, and HR.
This does not mean every agent needs a committee. Central approval can become its own failure mode. Slow gates encourage people to build around them.
It means every company needs a faster path from personal invention to shared memory.
Airtable offers a useful counterexample to the idea that AI adoption has to be centralized. During a 60-day internal program, 435 of roughly 700 employees built production agents. The company says more than 200 workflows became agent-driven.
But “let everyone build” was not the whole system.
Employees spent the first two days mapping workflows before they touched the tools. Agents went into a shared registry so people could find and reuse what already existed. Outcomes were measured weekly. One person owned the program, and each function had a captain carrying lessons across teams.
The freedom produced hundreds of agents. The structure gave those agents a chance to become company capability instead of personal debris.
Put every pilot through the Monday test
The first builder proves that the model can do the job.
The second person proves there is a workflow.
The first correction that improves every future run proves the company can learn.
Before calling a pilot successful, ask five questions:
Can somebody other than the builder find it and run it?
Is one owner visible where the workflow lives?
Can that owner see what it did and where it failed?
Does a correction change the next run for everyone?
Is there a clear way to pause, replace, or delete it?
If the answer is no, the pilot may still be useful. It is not yet organizational capability.
It is future archaeology.
The next AI race is a memory race
The next enterprise AI problem will not be an inability to build.
It will be an inability to know what has been built, who owns it, what it is allowed to do, what the company learned from it, and whether it should still be running.
The model race matters. Better models create new possibilities. But frontier capability diffuses quickly. Most companies will be able to buy access to similar intelligence.
The advantage is what happens after access.
One company fixes an agent’s mistake in a chat window and forgets it. Another turns that mistake into a better source, instruction, test, or permission rule. One repeats the same correction across twenty disconnected pilots. The other makes the correction once.
That is the living ecosystem I want when I talk about an AI-native company. Not a company with the most pilots. A company in which useful corrections have somewhere durable to go.
Your AI pilot worked. Good.
Now see whether it survives Monday.
Maintenance is not the boring part after the innovation. It is where intelligence becomes institutional.


