← work
Analytics QA · Automation · [CHECK] · GA4, Looker Studio, Notion, Automated browser QA

Automated Nightly Tag-Health Monitoring

<30 min
weekly tag QA, down from 2–4 hours
6+
web properties checked nightly
[CHECK]
issues caught (count / period)
problem:

Rather than a single client engagement, this is a capability Datacraft built and now runs across its own portfolio of client web properties — because the same tag-reliability problem shows up for nearly every client running any meaningful amount of marketing tracking.

  • Manual verification doesn't scale. Confirming that every tracked event was still firing correctly meant opening GA4 DebugView and Tag Assistant and manually firing test events, property by property — a process that ate 2 to 4 hours a week even when nothing was wrong.
  • It only caught what someone remembered to test. A manual process is only as thorough as the person running it that day; an untested event path could silently break for weeks before anyone noticed a dip in a report and traced it back to a tagging failure.
  • Releases break tags without warning. Marketing sites and platforms change constantly — a redesign, a new page template, a checkout update — and any of those can silently break a tag with no error message pointing back to the cause.
what i built:

We built an automated nightly QA workflow that runs its own browser session and fires every tracked event across all six-plus monitored properties, independent of anyone remembering to test that day. Rather than spot-checking a handful of paths, it exercises the full set of tracked events every single night.

Any event that comes back with zero recorded hits gets flagged automatically — that's usually the signature of a release that broke something, whether it's a redesign that moved a button, a checkout change that altered a dataLayer variable, or a third-party script that stopped loading. Those flags feed directly into health-check dashboards built in GA4 and Looker Studio, so the status of every property's tracking is visible at a glance rather than buried in raw event data.

To keep the process from just generating noise, we log weekly triage notes in Notion — a running record of what was flagged, what turned out to be a real issue versus a false positive, and what was fixed — so patterns across properties become visible over time instead of every flag being treated as a one-off.

Nightly schedule trigger
  -> Automated browser fires every tracked event
  -> 6+ client properties
  -> Zero hits recorded?
       |
       +-- Yes -> Flag as an issue --+-> GA4 / Looker Studio health dashboard
       |                             +-> Notion triage log
       +-- No  -> Healthy, no action

# Every tracked event gets fired automatically every night; only the ones
# that come back silent ever reach a human.
outcome:
  • Weekly tag-QA time dropped from 2–4 hours of manual, property-by-property checking to under 30 minutes of triaging only real, flagged issues.
  • Tag failures are now caught automatically and consistently, rather than depending on someone remembering to test the right thing on the right day.
  • The health-check dashboards give an at-a-glance view of tracking status across the entire property portfolio, rather than requiring a manual pass to know whether everything is working.

[CHECK] Still to confirm: a specific example or count of issues this workflow has caught over a given period, if a concrete number is preferred alongside the time-savings figure.

Tag failures are usually silent until someone notices a reporting gap weeks later. If your team is still checking tags by hand, automating that check is one of the highest-leverage things you can do for data trust.

[CHECK] [screenshot] the server container