Why we left automatic execution switched off.
We built a system that can change ad accounts on its own. It works. It is off, on purpose, and the reason says more about the job than about the software.
Here is something you rarely hear from a company that builds software. We built a feature, tested it, and turned it off.
The platform behind our agency can carry out changes on ad accounts by itself. Pause a wasteful placement. Add a negative keyword. Move budget toward a campaign that is holding its return. The code exists. It has been tested. It is switched off, and it has been off the whole time it has been running on client accounts.
That can look like timidity in an industry that talks about AI agents running accounts and campaigns that run themselves. I think it is the opposite. It is the one decision in this business where being bold is cheap and being right is expensive.
Finding a problem and fixing it are different skills
Think about a good doctor reading a scan. Spotting the shadow is pattern recognition, and machines have become very good at it. Deciding what to do about the shadow needs context the scan does not contain: the patient’s history, what they can tolerate, what they want. Nobody objects to software that flags the shadow. Most people would object to software that schedules the surgery.
Ad accounts are the same. Our daily checks are good at noticing things: a search term with spend and no sales, a campaign whose conversions fell while the store kept taking orders, a creative that is tiring. What the checks cannot see is the context around them. The product that is about to go out of stock. The margin that changed last week. The promotion the brand is launching on Thursday that makes today’s dip irrelevant. The decision needs the business, and the business is not in the data.
The asymmetry that decides it
When a person reviews a proposal before it ships, the cost of a wrong recommendation is a few seconds of their attention. When a system ships the change itself, the cost of a wrong recommendation is real money, spent or lost before anyone looks. Worse, the damage compounds. Bidding algorithms learn from what happens next, so a bad change does not just cost a day. It teaches the platform something false.
So the question is never whether the system is usually right. It is whether it is right often enough, on the specific kind of change, that removing the person is worth the tail risk. That bar is higher than most demos suggest, and it is different for each kind of change. Adding a negative keyword to a term that has spent heavily with no sales is close to safe. Cutting a budget because conversions dropped is not, because half the time the conversions did not drop. The tracking did.
What we do instead
The platform reads each account every day: the ad platforms, the analytics, the store, our own first-party tracking, and the customer list, side by side. It ranks what it finds by the money at stake, not by how dramatic the percentage looks. Anything that should change is written up as a proposal that says what it would change, why, and how to undo it.
Then a senior strategist decides. Changes that affect a client go through an approval the client controls, in the same portal where they see their numbers. Every action is logged. The watching is automated. The judgment is not.
This is slower than autopilot by minutes. It is faster than the traditional agency by weeks, because the problems are found the day they appear instead of at the monthly review. That trade seems right to me for money that belongs to someone else.
The case for automation, taken seriously
It would be easy to write this as a story about caution beating recklessness. That is not the argument. A great deal of advertising is already automated, and most of it should be.
Smart Bidding sets a bid for every auction. Meta’s delivery system decides who sees which ad, many times a second. Nobody sensible wants a person approving those decisions, because the platforms see signals no person can and the decisions are bounded by settings a person chose: a budget, a target, an audience, a set of creatives.
The line I care about is different. It is between automation that works inside the limits a person set and automation that changes the limits. A bidding algorithm spending a budget is the first kind. A system that raises the budget, pauses the campaign, or rewrites the structure is the second. The first is a tool. The second is a decision maker, and a decision maker has to be accountable to someone.
What a good proposal contains
Because a person approves each change, the quality of the proposal decides the quality of the decision. A proposal that says “pause this” invites a rubber stamp. A useful one reads more like a short brief.
- What would change, stated exactly: which campaign, ad group, keyword, or budget, and the value before and after.
- Why: the finding that triggered it, with the dollars at stake rather than a percentage on its own.
- The evidence across ledgers: whether the store, analytics, and first-party tracking agree with what the platform shows.
- What could make it wrong: a promotion, a stock change, a tracking break, a data delay.
- How to undo it, and how long it would take to know whether it worked.
Written that way, a proposal can be approved in a minute by someone who understands the business, and rejected just as fast when it misses something. Every rejection is also useful. It is a record of what the system could not see, and that record is exactly what decides when a kind of change is safe to automate.
What earning it will look like
- A long record, on real accounts, of proposals of one specific kind and how many were approved as written.
- A measured cost for the misses, including how quickly they would have been noticed and reversed.
- A narrow scope to begin with, such as negative keywords for terms with heavy spend and no sales, rather than budgets or structure.
- The client’s explicit agreement, account by account, with a way to switch it back off.
Why this is also a measurement argument
The strongest reason to keep a person in the loop is not philosophical. It is that the data an automated change would act on is wrong more often than people expect.
Consider the most common emergency in paid media: conversions fall off a cliff overnight. An automated system reading the ad platform alone sees a campaign that has stopped working and cuts its budget. A person reading the ledgers side by side checks the store first. If the store kept taking orders at the usual rate, the campaign did not stop working. The tag did. Cutting the budget would have turned a measurement problem into a revenue problem, and taught the bidding system that the campaign is worse than it is.
That is why our checks read the store, analytics, first-party tracking, and the platforms together before anything is proposed, and why a proposal carries the cross-check with it. An automated change is only as good as the number it trusted. Until every number it might trust has been checked against the others, the check has to be a person.
The cost of being slower
The fair objection is speed. Some problems cost money every hour they run, and waiting for a person seems wasteful. In practice the delay is rarely the approval. It is noticing. A traditional agency finds a runaway search term at the weekly review or the monthly report. Daily checks find it the next morning, ranked by the money at stake, with the proposed fix already written. The approval that follows takes minutes.
The remaining gap between a next-morning fix and an instant one is real, but small, and it buys something that matters more than those hours: every change on a client’s account has a person who can explain it.
When we will switch it on
Not on a date. On evidence. We are improving how reliable the analysis is step by step, and each step has to pass its own checks before we lean on it. When we can show, on real accounts, that a specific kind of change is recommended correctly almost every time, and that the rare miss is cheap and reversible, that kind of change becomes a candidate for automatic execution. Probably narrow ones first, with the client’s agreement, and never for everything at once.
Until then, the capability stays built and off. I would rather tell a client that a person looked at every change than tell them the software was confident.
A question worth asking anyone
If your agency or your tools make changes automatically, ask two things. Which kinds of change happen without a person looking, and what evidence says that is safe? And when one of those changes was wrong, how quickly did anyone notice?
Neither question is hostile. Automation is valuable, and some of it is plainly safe. But the confidence of the software is not the evidence. The record is. If nobody can show you the record, the person approving the changes is you, whether or not anyone asked.
Written by Sam Nouri, founder, adsrunner. If this resonated and you want to apply it to your own account, you can book a strategy call or run a free audit.
How we research, source figures, and handle corrections: editorial policy.