Ad performance anomaly detection without false alarms

Ad performance anomaly detection avoids false alarms when each campaign is judged against its own recent history, not a fixed rule. Flag a movement only when there is enough volume, the change is large enough to matter and it is unusual for that campaign, measured with the median and MAD. Remove weekday patterns first, merge related metrics into one finding and rank by share of spend.

Key takeaways

  • Fixed thresholds fail both ways: they fire on volatile campaigns’ normal swings and miss real moves on steady ones.
  • The same +18% CPA can be a robust z of 12.8 on one campaign and 1.3 on another.
  • Use the median and MAD, not the mean and standard deviation: one bad day can hide a real incident.
  • Gate on volume, remove the weekday pattern and flag only changes that are material and unusual.
  • Report one finding per campaign, ranked by share of spend; a ROAS gain outranks a CPA rise.

Why do fixed threshold alerts fail?

A fixed threshold fails because it applies one number to campaigns that behave very differently. “Alert me if CPA rises 20%” is too tight for a campaign whose CPA swings daily, so it fires on noise, and too loose for one that barely moves, so it stays quiet while something real breaks.

Both failures are costly: noise teaches people to ignore alerts, and an ignored alert is how a real problem runs for a week. The better question is “is this unusual for this campaign?”, which takes each campaign’s own baseline and a measure of how much it normally varies.

SituationFixed thresholdBaseline-aware detection
What counts as normalOne number for every campaignEach campaign’s own recent history
Volatile campaignFires on ordinary swingsWide normal range, stays quiet
Steady campaignMisses moves under the thresholdNarrow normal range, flags early
Few conversionsFires on luckWaits for a minimum volume
WeekendsVolume “drops” every SaturdayWeekday pattern removed first
CPA, conversion rate and conversions move togetherThree alertsOne finding, with its driver
CPA up, ROAS upAlerts on CPAReports the ROAS gain

Is an 18% CPA increase an incident or just noise?

It depends on how much that campaign’s CPA normally moves. Example: two campaigns, each with a median daily CPA of $40 over 14 baseline days, both had a CPA of $47.20 over the last 7 days, an 18% increase. Their baseline daily CPAs:

Campaign A (steady):   39, 41, 40, 38, 42, 40, 41, 39, 40, 42, 38, 41, 40, 39
Campaign B (volatile): 45, 30, 52, 39, 26, 49, 41, 33, 60, 38, 28, 55, 30, 50

A never left $38–$42; B ranged from $26 to $60. With the robust z-score explained further down:

Campaign A: median $40, MAD $1
z = (47.20 − 40) ÷ (1.4826 × 1 ÷ √7) ≈ 12.8 → incident

Campaign B: median $40, MAD $10
z = (47.20 − 40) ÷ (1.4826 × 10 ÷ √7) ≈ 1.3 → noise

Same percentage, opposite conclusions. A fixed 20% rule gets both wrong: it stays silent on A’s 7-day CPA, and on daily CPA (“above $48”) it would have fired on 5 of B’s 14 ordinary baseline days.

How do you detect ad performance anomalies, step by step?

Run every campaign through the same checks and keep only what passes all of them. This is how Borealis works for Google Ads monitoring, and you can reproduce each step in a spreadsheet.

1. Compare each campaign with its own recent history

Use a current window and the baseline right before it, for example the last 7 days against the 28 days before. Both are whole weeks, so they hold the same mix of weekdays. Compute rates such as CPA and ROAS from window totals (for CPA, total spend ÷ total conversions) and compare volumes as daily averages.

Planned peaks break this comparison, because your own history stops being a fair reference; for promotion weeks, see Black Friday ad monitoring.

2. Don’t judge a rate on too little volume

A rate built on a handful of conversions swings on luck. Example: a campaign converted 9 times from 360 clicks in the baseline (2.5%) and once from 100 clicks this week (1.0%). That looks like a 60% conversion-rate drop, but two more conversions this week would make it 3.0%, a rise. Set a minimum, for example 15 baseline conversions for CPA or conversion rate and 150 clicks for CTR or CPC.

3. Require a change that is both material and unusual

Material means big enough to act on, for example 12% or more. Unusual means large compared with the campaign’s normal variability, for example a robust z of 2.5 or more. Campaign B’s +18% is material but not unusual. Campaign A at $41.60 would be the reverse: z ≈ 2.9 on a 4% move, statistically real and commercially trivial.

Why use the median and MAD instead of the mean and standard deviation?

Because one bad day drags the mean and inflates the standard deviation, while the median and MAD barely notice it. MAD (median absolute deviation) is the median distance between each day and the median. Multiplied by 1.4826, it sits on the same scale as a standard deviation for normally distributed data.

robust z = (x − median) ÷ (1.4826 × MAD)

for a 7-day average rather than a single day:
robust z = (x − median) ÷ (1.4826 × MAD ÷ √7)

The √7 (about 2.65) is there because a 7-day average moves less than a single day. Example: take Campaign A and suppose that on day 10 a tracking glitch left most conversions unrecorded, so CPA read $112 instead of $42.

StatisticClean baselineWith the $112 day
Mean$40.00$45.00
Standard deviation$1.30$19.32
Median$40.00$40.00
MAD$1.00$1.00

With the mean and standard deviation, this week’s $47.20 is only 4.9% above “normal”, and z = (47.20 − 45.00) ÷ (19.32 ÷ √7) ≈ 0.30. The incident disappears, and it stays hidden for as long as the glitch sits in the baseline. With the median and MAD nothing moved, and z is still about 12.8.

How do you handle weekends and weekday patterns?

Remove the weekday pattern before you measure how much a volume metric normally varies. Conversions and clicks often follow a weekly rhythm: day by day, every Saturday looks like a drop, and across mixed weekdays the rhythm makes a campaign look more volatile than it is, so a real drop passes as noise.

Example, daily conversions with a two-week baseline:

DayWeek 1Week 2This week
Monday615951
Tuesday565848
Wednesday555346
Thursday495143
Friday454337
Saturday313327
Sunday383530
Baseline:  667 ÷ 14 = 47.64 conversions a day
This week: 282 ÷ 7 = 40.29 a day (−15.4%)

Raw days: MAD 7.5
z = (40.29 − 47.64) ÷ (1.4826 × 7.5 ÷ √7) ≈ −1.75 → looks like noise

After subtracting each weekday’s median: MAD 1.0
z = (40.29 − 47.64) ÷ (1.4826 × 1.0 ÷ √7) ≈ −13.1 → a real drop

Each weekday’s median (Monday 60, Tuesday 57 and so on) removes the rhythm and leaves the day-to-day noise, about one conversion here. For volumes, center on the baseline daily average rather than the median of raw days: over whole weeks it weighs every weekday equally.

How do you rank anomalies so the list stays short?

Merge related metrics, weight severity by spend and let value beat cost.

One finding per campaign

When CPA, conversion rate and conversions move together, they are one event. Example: conversion rate falls 23% while CPC and clicks hold. Conversions fall 23% and CPA rises 29.9%, because CPA = CPC ÷ conversion rate and 1 ÷ 0.77 = 1.299. Report one finding, “CPA up 29.9%, driven by conversion rate”, not three alerts. To go from driver to cause, see why your CPA went up.

Severity is materiality × direction × significance

Share of spend says how much a movement matters, direction whether it is a problem or an opportunity, and significance how sure you can be. Example: two campaigns both show CPA +40% in a $10,000 week. One spends $100 (1% of the account) at a $25 CPA; the other $4,500 (45%) at a $45 CPA. At the same spend, a 40% higher CPA buys 28.6% fewer conversions (1 − 1 ÷ 1.4): 4 conversions become 2.9 on the small campaign, and 100 become 71.4 on the large one. Same percentage, 25 times the damage.

A simple tiering: critical for a large, very unusual bad move on a campaign with a real share of spend (for example 10% or more); attention for other bad moves; opportunity for good ones; informational below a few percent of spend.

Let value win over cost

If ROAS improved, a higher CPA is the price of that return, not an incident. Example, per day: spend holds at $400, conversions fall from 10 to 8 and conversion value rises from $1,600 to $1,800. CPA goes from $40 to $50 (+25%), yet ROAS goes from 4.0 to 4.5 (+12.5%), because each conversion is now worth $225 instead of $160. Report the ROAS gain and suppress the CPA alert.

How do you build this in Google Sheets or Excel?

Use one tab per campaign and one row per day: the 28-day baseline in rows 2–29 and the current 7 days in rows 30–36. Columns: A date, B weekday (=WEEKDAY(A2)), C spend, D conversions, E daily CPA (=C2/D2). If many days have zero conversions, the campaign is too small for daily CPA; judge conversions instead.

CellWhat it holdsFormula
J2Baseline median=MEDIAN(E2:E29)
J3MAD=MEDIAN(ABS(E2:E29-J2))
J4Current CPA=SUM(C30:C36)/SUM(D30:D36)
J5Change=J4/J2-1
J6Robust z=(J4-J2)/(1.4826*J3/SQRT(7))
J7Volume gate=SUM(D2:D29)>=15
J8Flag=AND(J7,ABS(J5)>=0.12,ABS(J6)>=2.5)

In Google Sheets, wrap the MAD formula in ARRAYFORMULA(); Excel 365 takes it as written, and older Excel needs Ctrl+Shift+Enter. If most days are identical, MAD can be zero; use a floor such as MAX(J3,0.02*J2) instead.

For conversions or clicks, add a residual column F in rows 2–29: =D2-MEDIAN(FILTER($D$2:$D$29,$B$2:$B$29=B2)). This needs FILTER (Google Sheets, Excel 365 or 2021). Take the MAD of column F the same way, then compute =(AVERAGE(D30:D36)-AVERAGE(D2:D29))/(1.4826*MAD/SQRT(7)). Then sort the flagged campaigns by share of spend.

Across dozens of accounts, the sheet becomes a job of its own. Borealis runs these checks every morning on Google Ads, names the driver and campaign behind each change and sends one email per project; other platforms are coming soon.

Frequently asked questions

What z-score threshold should I use for PPC anomaly detection?

There is no universal value: the threshold trades missed incidents against false alarms. A robust z of 2.5 is a reasonable start for flagging, with a higher bar, such as 3.5, for anything urgent. After a month, review the flags: if most were noise, raise it; if real problems slipped through, lower it.

How much history does a new campaign need before anomaly detection works?

At least two full weeks, and four is better. With less, you cannot estimate how much the campaign normally varies, so every movement looks alarming or meaningless. Use whole weeks so the weekday mix matches, and until then watch budget, delivery and tracking directly.

Can Google Ads automated rules detect anomalies?

Only in the fixed-threshold sense. Automated rules act when a metric meets a condition you set, such as CPA above a value, so one number covers campaigns with very different volatility. Google Ads scripts can compute medians and deviations, but you write and maintain that code yourself.

Should I detect anomalies at campaign level or account level?

At campaign level, then roll up. Account totals hide offsetting moves: one campaign’s CPA can double while another improves and the account looks flat. When the account moves, recompute its metric as if each campaign had kept its baseline rate at its current volume; the biggest gap shows where to look first.

How do I avoid anomaly alerts after a planned budget increase?

Expect the movement rather than silence it. Check that what moved matches the plan: more budget should raise spend and conversions, possibly with a somewhat higher CPA. If size and direction match, dismiss it; if CPA rises far more than planned, or conversions don’t follow spend, treat it as an incident.