JOURNAL

Everyone can build an agent now. Nobody has signed off on them

The software factory was the proof, not the destination. The lawyer, the bookkeeper and the marketing lead are building their own agents, and in most places nobody said yes.

23 September 2026·13 min read·internal ai tools · ai agents · ai consultant · claude · agent governance · ai workflow automation · claude news

TL;DR: AI agents have left the engineering department. Anthropic classified 1.2 million Claude Cowork sessions across more than 600,000 organisations, and software development accounts for 8.7% of them. Business operations accounts for 33.4%. The agents being built inside your company right now are probably being built by someone who is not a developer, in software that is not a developer tool. Building is no longer the bottleneck. Deciding who may build is.

For four years each layer has been stacked on the last one. First the good instruction. Then the loop that could run several steps. Then the graph that could steer many loops. Then the software factory, where agents write, test and move code along while humans approve at the gates.

Every layer made production cheaper. None of them made trust cheaper.

Now the pattern is leaving software. That is the actual news, and it is a harder shift to handle than the software factory was.

Anthropic, 2026 Agentic Coding Trends Report (PDF, primary source)

The lawyer who built their own tools

There is a passage in Anthropic's agentic coding report that is easy to read straight past. Their legal team cut marketing review turnaround from two or three days down to 24 hours. Not by hiring. By building workflows that absorb the repetitive part: contract redlining and content review.

Then comes the sentence that matters. Using Claude Code, a lawyer with no coding experience built self-service tools that triage issues before they reach the legal queue. Claude Code is still a developer tool, and that is the point: when the report was written the lawyer had to go there. The same capability now sits inside Word and Excel.

Consider what had to be true for that to happen. The lawyer did not file a ticket. Did not wait for a sprint. Was not ranked against the product roadmap. There was a problem, there was a tool, and the thing got built.

Animation of a work queue where items disappear from the middle instead of being taken from the front, so the queue empties without passing the person who used to assign the work.
The queue is not being worked down from the front. It empties from the inside, because the person with the problem solved it themselves.

For many years automation inside a company went through one channel. Someone with a problem described it to someone who could build. That handover was slow. It was also, with nobody planning it, a filter. Every automation passed a human whose job was to ask whether it should exist at all.

The handover is gone. The filter went with it, and nobody decided that.

8.7%. That is how much of Claude Cowork is coding

In May, Anthropic classified 1.2 million anonymised Cowork sessions from more than 600,000 organisations into 20 categories of work. The split looks like this:

  • Business process and operations: 33.4%. Reports, onboarding checklists, reconciling spreadsheets. Anthropic notes these are tasks that sit in finance, HR and administration.
  • Content creation and copywriting: 16.4%. Drafts, decks, posts, proposals.
  • Software development: 8.7%.
  • Then DevOps and infrastructure at 7%, research at 6.4%, data analysis at 5.8%, document processing at 4.1%, and sales and revenue operations at 4%.
Infographic of Claude Cowork sessions by category of work, with business operations at 33.4 percent, content and copywriting at 16.4 percent and software development at 8.7 percent.
Anthropic's own count across 1.2 million Cowork sessions, 11 to 31 May 2026.

That number needs an honest footnote. The 8.7% does not mean coding has become a small part of AI use. Developers write code in Claude Code, not in Cowork, so they sit largely outside this sample. The figure says something more precise and more useful: when a company is handed a general agent tool that is not aimed at developers, office work is what shows up. Half of it is admin and text.

It matches the direction Anthropic predicts itself. Trend 7 in the report is called "Non-technical use cases expand across organizations", and the prediction is that sales, marketing, legal and operations gain the ability to automate workflows with little or no engineering involvement.

There is also a figure for how much new work gets created, and that figure belongs to engineering. In the software development section Anthropic puts about 27% of AI-assisted work in that category: work that would not have been done otherwise. Not faster work. Work that could not justify its cost before. Whether the same holds outside code, the figure does not say.

At the large end: Anthropic's report puts Zapier at 89% AI adoption across the whole company with more than 800 agents deployed internally, and TELUS at over 13,000 custom AI solutions. Both are the companies' own numbers, reported by Anthropic.

Ask yourself how many of those 800 went past an approval.

The spreadsheet waited. The agent does not

Shadow IT is an old problem. A colleague built a spreadsheet, it became business critical, nobody documented it, and one day the colleague resigned. Companies have lived with that for decades. It is a question of access and documentation, and mostly it is just irritating.

Here is my position. Shadow agents are not that problem in new packaging. The difference is that the spreadsheet waited.

A spreadsheet does nothing until someone opens it. It has no schedule. It sends no email. It does not call an API at three in the morning. The worst an undocumented spreadsheet can do is give a wrong answer to the person who opened it, and that person is sitting right there and can see the number looks odd.

Animation with two objects side by side. A spreadsheet sits completely still until a hand opens it. Beside it an agent starts on its own against a clock and sends something onward with nobody touching it.
One object waits for a human. The other one has a schedule.

An agent acts. It runs to a schedule, it holds credentials, it writes into systems, and it does all of that whether or not anyone is watching.

That moves the question from IT to liability. A spreadsheet that is wrong is an error in a document. An agent that is wrong is something your company did. To a customer, a supplier or a regulator, with your name on it.

A ban moves the problem. It does not solve it

The obvious reaction is to forbid it. Nobody builds agents without approval. I understand the instinct. The ban does not work, and the reason is not culture.

The tool to build them is inside the software your colleagues already have open. Claude for Financial Services shipped on 5 May 2026 with ten ready-made agent templates, among them one that reconciles general ledger accounts, one that runs the month-end close and a KYC screener. It shipped add-ins for Excel, PowerPoint and Word, with Outlook still to come. A week later, on 12 May, Claude for Legal shipped with twelve practice-area plugins and more than twenty connectors into Westlaw, iManage, Ironclad and the rest. Artificial Lawyer reports that those plugins install in one click.

So the question of whether employees should be allowed to build agents has already been answered by the way the product is distributed. A ban is a sign on a door in a building with no walls.

This is still my judgement and not a rule anyone has passed: the only control that works is making the sanctioned route faster than the unofficial one. People go around the process when the answer takes three weeks. They rarely go around it for fun.

Read that with this in mind. I sell this setup. I make money building the sanctioned route for companies, so I have a clear interest in recommending it. I also run it on myself, including when nobody is paying me to, and it is drawn further down this page.

Who may build internal AI tools where you work?

This lands differently depending on where you sit. The subject is the same. The consequence is not.

Owner or leader in an SMB

You probably have no IT department to route a ban through, and right now that is an advantage. You can make this call in an afternoon where a larger company needs three departments to agree. The decision is not whether agents are a good idea. It is which two systems an agent may never write to without a human seeing it first. Banking and payroll are where I usually land.

Operations and transformation lead

This arrives on your desk as an operations problem, long after the agents were built. The cheapest thing you can do this week is find out which agents are already running and what they can reach. Not to shut them down. To know. In my experience a list of ten agents takes a morning to write now, and it gets harder every week you leave it.

Sales and RevOps lead

Your agents touch contact data, which makes them a data protection matter and not only an efficiency one. An agent writing into your CRM is changing a register of identifiable people. An agent sending on its own is communicating with them. The first needs you to be able to show who changed what. The second needs a lawful basis that holds, and human-in-the-loop AI before anything sends, at least until you have watched that agent be wrong once.

Product and engineering lead

You are the only person in the building who already owns the machinery. Version control, tests, review, and a place where work exists without counting yet. Your job shifts from being the one who builds to being the one who lends the gates to everyone else. That is an unpopular reading, because it looks like more work with no more headcount. The alternative is five departments each inventing their own version of something you have run for a decade.

What a factory looks like when it is not building software

There is one question that sorts agents faster than any maturity model. Put it to every agent that is running or about to run:

If this agent is wrong on Monday, who finds out, and when?

There are really only three answers.

  1. A human sees the output before it counts. Ready. The mistake gets caught while it is still free.
  2. Someone notices at next month's reconciliation. Not ready. You have built an error with a month of delay baked into it.
  3. You hear about it from a customer. Not ready, and it should be switched off today.

Answers 2 and 3 are not arguments for dropping the agent. They are arguments for moving it behind a gate until the answer becomes 1.

Animation where the same wrong result travels along three different routes, and the length of each route shows how far the error gets before anyone notices: stopped immediately, caught at the monthly reconciliation, or out with the customer.
It is not the error that decides the damage. It is how far it gets before anyone notices.

Here is mine. This is the Brinvik content factory, and it does not build software. It makes marketing, which is exactly the category sitting at 16.4% in Anthropic's count.

It runs Tuesday and Thursday nights in the cloud. Seven gates, each with a cap on rounds, one to three, because a gate that blocks the build is a broken gate. It builds everything and publishes nothing. I read every draft before anything goes live.

Animation stepping through the Brinvik content factory in five views: two sources converge on Gate 0, the seven gates run on with a loop back from Gate 3 to Gate 1, Gate 7 branches down to run_complete.sql, and past the night-to-morning divide sit the queue and my own read.
My own factory in five views. The bar along the bottom shows where you are in the diagram.

If you want all seven gates at once, the whole diagram is here at full size.

The interesting part is not the number of gates. It is that it has the four things that made the software factory work, none of which were invented for marketing:

  • A record of what happened. Seven tables per run. What was seen, what was written, what was queued, what was published, what ran, what broke, and what the run handed to the morning.
  • Checks that run themselves. The gate scripts are the mechanical half, and they look for reasons to stop the run rather than reasons to wave it through. A gate that never stops anything has not approved anything. It just has not looked.
  • A human gate. Me, in the morning.
  • A state where the work exists without counting. A draft.

That last one is the one people skip, and it is the cheapest of the four. A draft is the most underrated safety mechanism there is. It costs nothing to produce, and it makes every mistake free right up to the moment somebody presses the button.

There is no git revert on a sent email. There is a draft.

Building the agent is rarely the hard part. The hard part is what comes after: getting it behind a gate without the gate making it useless. Most people stall exactly there, because they know neither where the approval should sit nor who should be looking at what. That is the conversation I get called in for. I am an AI consultant and a standing AI implementation consultant for a handful of Danish companies, and the work is rarely building more agents. It is setting up the internal AI tools already running so they can keep running without anyone holding their breath.

The order matters more than the tooling. AI workflow automation that starts by counting what already exists gets there faster than the version that starts by buying more. If you want to see what that looks like where you work, read about internal AI tools or AI advisory on a retainer. If you would rather see what people actually build first, I went through 43 workflows from Claude and judged each one against a real company.

Right now you can still write every agent in your company on a single page, because there are few enough of them. That will not stay true.

If you want a structured version of that list, my AI readiness audit is free on the site. It takes five to ten minutes, and you do not have to talk to me to use it.

The list is not worth making because agents are dangerous. It is worth making while it is short.

Sources

Primary source

Third party

FAQ

Frequently asked questions

An AI agent is software that carries a task through from start to finish instead of answering a question. It has access to your systems, it can run to a schedule, and it can write data and send messages. The difference between an agent and a chatbot is that the agent acts.

You decide, but in practice the answer today is anyone with Claude or a similar tool open. The tool sits inside Word and Excel. If you have not written a rule down, the unwritten rule is that anyone may.

You can write the ban, but you cannot enforce it, because the tool is already inside the software they work in. What works is making the sanctioned route faster than the unofficial one. Nobody goes around the process for fun.

In Claude Cowork, software development is 8.7% of sessions, measured by Anthropic across 1.2 million sessions in May 2026. Business operations is 33.4% and content work is 16.4%. The figure is low because developers code in Claude Code and are therefore not in that sample.

Ask one question: if the agent is wrong on Monday, who finds out and when? If a human sees the output before it counts, it is ready. If the error only surfaces at the next month-end, or you hear about it from a customer, it is not.

The first step costs a morning and no money: write down which agents are running, who built them and what they can reach. Only after that is it worth discussing what a setup with approval and logging should cost.

They become one the moment they touch personal data. The legal requirement is the lawful basis: Article 6 has to hold before the agent sends anything out. The rest is my practice and not statute. If you cannot trace a change back to whoever made it, you cannot put it right either.

Because the software factory is where the pattern was proven. Code can be tested automatically and rolled back, so it was the cheapest place to let agents work. What is interesting in 2026 is that the pattern is moving into functions that have neither of those two things.

Get new essays by email.

Roughly twice a month. Same voice. No list rental, no retargeting.

Sign up for the Brinvik journal. Unsubscribe anytime. See our privacy policy.

Protected by Cloudflare Turnstile. No challenge, no CAPTCHA. Brinvik never shares your address.