A promising demonstration shows that an AI media buyer can explain an account. A useful pilot shows whether it can make decisions your team would trust under ordinary operating conditions. Shadow mode creates room to evaluate that judgment while the existing operator remains responsible for campaign changes.
The important boundary is execution. The AI may read approved information and produce recommendations, but its outputs do not automatically become campaign actions. Define that boundary using an action-level permissions matrix, and verify the vendor can enforce it.
Choose the question the pilot must answer
Start with one sentence: “We want to learn whether this operator can identify useful next actions for this account without creating an excessive review burden.” That is a different question from “Will it increase revenue?”
A shadow pilot can evaluate account understanding, evidence selection, recommendation clarity, scope discipline, and reviewer effort. It cannot directly establish the return from recommendations that were never executed. If a human selectively adopts only the best suggestions, the observed results also reflect that selection.
Write down the distinction before the pilot starts. It prevents a good analysis demonstration from being retold as an independently measured improvement in advertising performance.
Give both operators the same starting context
Prepare an account brief with the objective, conversion definition, timezone, currency, product economics, exclusions, recent changes, and known measurement gaps. Include the campaign responsibilities of other teams. The human should not have an invisible advantage because only the human knows that a promotion ends tomorrow.
At the same time, do not leak future outcomes into an earlier decision. Save the report snapshot that was available when each recommendation was made. If conversions later arrive, preserve both the original snapshot and the revised report.
This matters because recent CPA and ROAS can change as conversions mature. Google's conversion-lag documentation describes that reporting issue. In your pilot, late data is part of the environment the operator must handle, not something to erase from the record after the fact.
Use a parallel decision journal
Ask the AI and the human to record their intended next action before reviewing one another's recommendation. Perfect independence may be impractical in a small team, but recording the sequence makes the limitation visible.
| Journal field | What to record |
|---|---|
| Decision time | When the recommendation was produced |
| Available evidence | Report dates, account scope, and known missing data |
| Proposed action | Exact object, proposed change, or explicit decision to wait |
| Rationale | The observation and business assumption behind the action |
| Alternative explanation | A plausible reason the apparent pattern could mislead |
| Expected result | What should change if the hypothesis is useful |
| Review point | When there will be enough evidence to revisit it |
“Do nothing” can be a strong recommendation when the account is stable or the data is incomplete. A pilot that rewards the number of changes creates pressure to intervene even when intervention is unnecessary.
Evaluate disagreements instead of counting matches
Agreement with a human buyer is not a complete quality measure. The human may be wrong, the AI may identify an overlooked issue, or both may be making reasonable choices under uncertainty.
Group disagreements by their cause. Did one operator use the wrong account? Did they disagree about attribution? Did one know the product was out of stock? Did they have different beliefs about how much evidence was enough? Did the AI propose a useful action but fail to explain the spending consequence?
This produces a better improvement list than a single agreement percentage. A permission error needs a different response from an unclear explanation. Missing business context needs a different response from a weak performance hypothesis.
For an illustrative review, imagine ten recommendations: four useful as written, three useful after clarification, two premature because of incomplete data, and one outside scope. Those counts describe review outcomes, not a success benchmark. The out-of-scope action may matter more than several correct summaries.
Measure the work created for reviewers
Track the minutes required to understand, verify, correct, and approve each recommendation. Include the time spent finding the underlying report. An assistant that writes a convincing paragraph but forces a buyer to reconstruct every calculation may move work rather than remove it.
Also count repeated recommendations, contradictory suggestions, and proposals made after a teammate already completed the task. Record whether the system acknowledges the changed account state. These are practical indicators of whether the operator can fit into an existing team.
Use an ordinary shared journal if that is sufficient. The pilot does not need a new reporting platform to answer these questions. What matters is consistent definitions and a record that can be revisited.
Decide what would justify the next stage
Set acceptance criteria before seeing results. Examples include: no account-scope errors in the pilot sample; a clearly defined review owner for every executable recommendation; evidence windows attached to budget proposals; and reviewer time low enough to justify the workflow.
Choose thresholds according to the consequences of a mistake. Passing a small sample does not prove an error rate is zero. Document how many opportunities were observed and what kinds of situations were absent. A quiet week without promotions or tracking problems may not exercise the hardest cases.
The next stage can be a limited execution pilot for one action family, rather than an immediate switch to broad autonomy. Keep other actions in recommendation mode while the team develops confidence in the narrower workflow.
Use an experiment for the performance question
If the remaining question is whether the operator improves outcomes, plan a test with a suitable control, a predefined outcome, a spending boundary, and a measurement window. The Google Ads experiments overview describes platform-supported experiments; availability and setup depend on the campaign type.
Do not assume that splitting a budget between two differently configured campaigns creates a valid experiment. Audience overlap, campaign history, allocation, and concurrent promotions can all affect interpretation. A before-and-after comparison is useful operational context, but it does not isolate the effect of the AI.
Close shadow mode with a short decision memo: what the operator handled well, which failure modes appeared, what changed during the pilot, and what remains untested. Use a weekly operator review to maintain that discipline after launch, and a pilot scorecard when comparing vendors. The output should be a defensible next step, not a general declaration that AI works.
