Building an AI-powered marketing team (Part 2)
Is your team's AI actually safe — and is it working?
This is Part 2 of a 3-part series on building an AI-powered marketing team. Last week I covered the individual trap of AI and how to build your first shared system. This week: how to tell if that system is actually paying off, and whether it's safe to trust.
80% of marketers feel pressure to adopt AI right now, and 89% of that pressure comes from the C-suite or the board, per Supermetrics’ 2026 Marketing Data Report.
Here’s the other half of that same report: 52% of marketers say external teams define their data strategy and measurement, not marketing, and nearly four in ten still can’t prove ROI across channels.
Put those two together and you get a familiar shape. Marketing is told to move fast, without owning the controls that make it safe or the measurement that proves it’s working. That’s not two problems. It’s the same problem twice: responsibility without ownership.
Once a workflow moves off one person’s laptop and into shared infrastructure — which is what Part 1 was about, a mistake stops being one person’s mistake. It’s the whole team’s exposure.
Why “is it safe” just got a deadline
The EU AI Act reaches full enforcement in August 2026 — five weeks from this issue — and it mandates documented controls for high-risk AI systems. Marketing content generation, personalisation, and targeting can fall inside that scope depending on how they’re used.
Many of us aren’t based in the EU, the practical effect is the same: “we’ll figure out governance later” is no longer a safe default, because “later” now has a date attached to it.
But the deadline is really just forcing a question that shared AI workflows already raised in Part 1: once a workflow has one owner and touches shared data, who decides what needs a human check before it ships? Without an answer, you get one of two failure modes — either everything gets manually reviewed (which kills the speed AI was supposed to buy you), or nothing does (which is how a hallucinated stat or an off-brand claim ends up in a client-facing deck).
The three-tier guardrail test
Not every AI output carries the same risk, so it shouldn’t get the same review. Sort outputs into three tiers before you decide what needs a human in the loop:
Tier 1: Internal and exploratory. Brainstorming, first-draft outlines, research summaries for your own use. Nobody outside the team sees this. Light or no review needed, and the cost of an error is simply a wasted five minutes.
Tier 2: Team-facing but not public. Internal briefs, sales enablement, drafts that other teammates will build on. These need a peer review, not a formal sign-off — someone else on the team should read it before it becomes the basis for someone else’s work.
Tier 3: External, high-stakes, or regulated. Anything that makes a claim, cites a number, mentions pricing, touches personal data, or goes external (ads, client decks, published content). This tier requires a named human sign-off before it ships, and a record of who reviewed it and when.
Most teams don’t need a review board, just this triage step done once, written down, and applied consistently — the same discipline Part 1 recommended for documenting a prompt: treat “what needs review” as a process, not a judgment call made fresh every time.
Why “is it working?” is not that easy to answer
I wish I could just answer “the team likes it” and “we’re faster now” when finance asks what a workflow is worth, but feeling faster isn’t the same as proving value.
It’s easy to know a workflow saves time, but it’s much harder to know whether it’s actually moving a number the business cares about, and the data backs this up: despite the pressure to adopt AI, nearly four in ten marketers still can’t prove ROI across channels, and over half report pressure to cut costs while maintaining results.
Teams are being asked to do more with AI and prove it’s worth the investment, often without ever having set up a way to measure it. The usual failure here isn’t a bad tool. It’s that “time saved” became the only metric anyone tracked, and time saved doesn’t survive a budget conversation.
A measurement framework you can actually run
Pick one business metric it should move. Not “productivity”, but something finance already tracks: conversion rate, cost per lead, time-to-first-draft on a deliverable that has a deadline, win rate on proposals. If you can’t name the metric, you’re not ready to claim ROI yet.
Baseline it before you scale the workflow. You can’t prove lift without knowing the starting point. If the workflow’s already team-wide, use the last quarter before rollout, or a comparable team/segment that hasn’t adopted it yet.
Instrument the actual touchpoint, not the team’s general sense of speed. Track the specific step the AI workflow touches — first-draft turnaround, review cycles, time-to-send — not a vague team-wide productivity feeling.
Set a kill-or-scale review date. Same cadence as the quarterly workflow review from Part 1, but with a decision attached: if the metric hasn’t moved by the review date, the workflow gets fixed, downgraded back to a personal tool, or retired. Infrastructure that nobody’s allowed to kill isn’t infrastructure, just a sunk cost with a Slack channel.
Where this leaves you
Safe and working aren’t two workstreams. They’re one question asked twice. A workflow with guardrails but no proof is safe and unaccountable, nobody can tell if it’s worth keeping. A workflow with ROI but no review tier is measured and safe, right up until an unreviewed claim ships under your name. Once it’s shared infrastructure, you don’t get to pick one.
Part 3 covers adoption: getting the team to actually use what you’ve built, once it’s proven and safe. See you then!
The AI;DR
Elsewhere in the AIverse
OpenAI launches Presence, its first real enterprise agent platform. Presence deploys voice and chat agents scoped to one job at a time — billing, claims, IT support — with OpenAI’s own engineers doing the implementation and a Codex-powered loop proposing updates as gaps surface in production. This is a services play as much as a product one — OpenAI’s field team, not a self-serve dashboard.
Black Forest Labs puts one model behind images, video, and robots. FLUX 3 generates up to 20-second video clips with synced audio, and the same backbone is already running on Audi’s productiohgjb mn line handling soft-material assembly. FLUX 3 Video is in early access now.
Washington and Beijing both start reaching for the export-control lever. Beijing is reportedly weighing new controls that would block foreign users from downloading Chinese open-weight models onto local servers, while the Trump administration is debating a reciprocal move and weighing sanctions over alleged IP theft via model distillation. If your stack includes any Chinese open-weight model (Kimi, DeepSeek, Qwen), access and hosting terms could shift with little warning.


