Turning raw SEO data into actionable insights doesn

Inspired by this post on Search Engine Land.


You may have watched your AI assistant reject an unsafe request and concluded that its safeguards worked. If you tested only once, you answered the wrong question. An attacker does not need every prompt to succeed. They need one useful failure after enough retries.
Best-of-N jailbreaking turns that model variability into a search process. To manage the risk, you need to evaluate the whole campaign, enforce permissions outside the model, and control every additional chance created by retries, fallback models, tools, and automated agents.
A Best-of-N attack creates or collects multiple versions of a prohibited request, submits them to an AI system, and selects the response that comes closest to the intended outcome. The essential move is to send many variations and keep the most successful result. The value of N is not fixed, and the selection can be performed by a person, a script, or another model.
This changes the security question. A per-request review asks, “Did this prompt get blocked?” A campaign-level review asks, “Did any related attempt produce a prohibited result?” The second question reflects the attacker’s objective.
The probability principle is straightforward. If each attempt has a nonzero chance of crossing a boundary, repeated opportunities can raise the chance that at least one attempt succeeds. Under the simplified assumption that attempts are independent and have the same success probability p, the probability of any success after N attempts is 1 – (1 – p)^N. Real prompt variants are often correlated, so you should not use that formula as a production risk estimate. Measure complete campaigns against your actual system instead.
Three distinctions prevent confusion during threat modeling:
Treat Best-of-N as a threat multiplier, not as the root vulnerability. It finds inconsistent decisions and weak handoffs. It cannot grant a caller a permission that your application enforces deterministically outside the model. That is why authorization architecture matters more than clever safety wording.

Your model is only one part of the attack surface. A typical AI workflow also has an identity layer, input filters, a router, one or more models, output checks, retrieval, tools, and application code. Every component that makes a fresh probabilistic decision can give a campaign another route to success.
| Layer | Misleading green light | Campaign signal to inspect | Stronger control |
|---|---|---|---|
| Prompt policy | One prohibited request was refused | Related requests are repeatedly rephrased after denials | Aggregate policy events by actor, session, intent cluster, and protected resource |
| Input moderation | Each prompt remains below an individual alert threshold | Small wording, format, language, or encoding changes accumulate around the same objective | Analyze normalized forms and sequences while retaining the raw input for investigation |
| Model routing | The primary model refused | A fallback model, alternate endpoint, or retry path returned a different decision | Apply one canonical policy before routing and a final gate after generation |
| Tools and agents | The assistant’s visible text looks harmless | A tool call requests a broader scope, sensitive record, or irreversible action | Enforce authorization, parameter validation, and action limits in application code |
| Traffic controls | Each IP address or API key stays within its local limit | Related attempts move across sessions, keys, endpoints, or models | Correlate only the identifiers justified by your threat model, privacy obligations, and retention policy |
| Logging | Every prompt was stored somewhere | No record connects attempts, decisions, tool calls, and final outcomes | Assign campaign and event identifiers so an investigation can reconstruct the sequence |
For an SEO, AEO, or GEO workflow, the highest-consequence result may not be a bad chat response. It may be an unauthorized CMS publication, a destructive edit, exposure of an unpublished campaign, or a tool call made with the application’s credentials. If a model generates page copy or JSON-LD, syntactic validation is necessary but insufficient. Valid structured data can still contain false, disallowed, or unapproved claims. Check the output against business rules and publishing permissions before it reaches a live page.

No safety prompt can carry this responsibility alone. Prompts influence model behavior, but they are not security boundaries. Use several controls with different failure modes, and place deterministic checks wherever failure could expose data, spend money, alter content, or trigger an external action.
There is no universal safe retry count. A blanket limit low enough for a sensitive data-export agent may be needlessly hostile in a public brainstorming tool. Set budgets by consequence, then examine legitimate retry behavior before choosing enforcement thresholds. Track false positives alongside security outcomes so that users who are clarifying ambiguous, multilingual, or accessibility-related requests are not treated automatically as attackers.
Be careful with model-based safety judges as well. A second model can add useful evidence, but it may share blind spots with the model it evaluates. Use deterministic authorization and validation for hard boundaries, with model judgments contributing to risk scoring rather than granting privileged access on their own.
A single-prompt red-team check will miss the defining behavior of Best-of-N. Your evaluation runner should group related attempts, preserve production routing logic, and score whether any attempt reaches a prohibited outcome. Keep testing authorized, isolated, and away from live customer data or publishing systems.
Your evaluation dashboard should include the campaign any-success rate, attempts to the first breach, breach severity, detection and containment outcomes, tool or data-boundary violations, and false-positive friction for legitimate users. Do not collapse these into one average. A small number of severe authorization failures should remain visible rather than being diluted by many harmless refusals.
Stop a test immediately if it begins interacting with real user records, external recipients, paid services, or live publishing. Move the scenario into an isolated environment with synthetic data and inert tools. The purpose of the exercise is to verify containment, not to prove that production damage is possible.
Before your next release, choose the AI workflow with the greatest access to data, tools, or publishing. Trace every place where a rejected request can receive another model call or another route. Then add campaign-level telemetry and a deterministic gate at the highest-consequence handoff.
That review will not eliminate model variability. It will prevent variability from becoming permission.


When I think about brand visibility today, it’s clear that being chosen by AI systems is crucial. Authority, unique insights, and consistent signals now determine if my brand makes the cut.
I’ve realized that AI isn’t just reshaping search; it’s deciding which brands are seen and which are ignored.
I learned from Andrew Warden, CMO of Semrush, at the Adobe Summit that visibility is evolving fundamentally, and our brands risk being systematically filtered out by AI systems.
“The idea of standing out is no longer optional. There’s a real risk of sameness,” he pointed out.
With AI systems deciding what to highlight and what to ignore, I know I must compete more fiercely for visibility in AI-generated answers.
The change is evident in the data: 60% of Google searches now end without a click to a website. People are still seeking information but aren’t always visiting websites. They’re getting their answers directly from AI systems like Google AI Overviews and ChatGPT.
These AI systems have become, as Warden described, the “new gatekeepers.”
This shift ushers us into the agentic era, where AI systems act as intermediaries, guiding users from inquiry to decision in one seamless interface.
Meanwhile, user behavior is evolving. People engage more in conversational environments, posing follow-up questions, refining queries, and surveying options within the interface, all resulting in fewer clicks but often attracting higher-intent users.
Warden noted that consumers using LLMs convert at least four times higher than those relying solely on search.
Despite some claims that AI could replace search, Warden reassured us that SEO is not dead.
SEO has become more foundational than ever. It’s essential to ensure my brand exists in the data layer AI systems rely on.
Warden emphasized, “SEO isn’t just for humans anymore. This is a training manual for AI right now.”
This involves ensuring:
Without these, my brand won’t appear at all.
Research backs this up: 94% of Google AI Overviews cite at least one top organic result, reaffirming that traditional search signals still support AI outcomes.
One striking concept from the session was what Warden dubbed the “bland tax.”
AI conditions itself to overlook blandness, causing generic or repetitive content to vanish.
If I’m generic, Warden warned I’m perceived as average, and if I’m bland, I’m effectively invisible.
AI systems don’t reward sameness. Rather than highlighting my brand, they often condense similar content into a single, attribution-lacking response.
“This is an invisible penalty,” Warden noted.
The consequences manifest in several ways:
“You also become a free training ground for LLMs,” he said.
Warden redefined brand visibility as a blend of:
“You absolutely need both,” Warden asserted.
SEO ensures I’m discoverable. Authority determines whether my brand shows up in AI-generated responses.
Without authority, I risk turning into a “commodity that isn’t worth being mentioned.”
Warden outlined three crucial areas determining whether my brand appears or gets filtered out:
AI systems map entities and relationships, and they must recognize my brand as an authority on a topic.
One key signal is brand demand. If people aren’t seeking out my brand, neither will AI.
Strong brands emphasize their authority across various platforms—owned content, media exposure, and community discussions—demonstrating their niche.
AI systems prioritize content that offers new insights. It’s vital to not just publish content but contribute something meaningful.
They emphasize new facts with proprietary data, original research, unique perspectives, and expert insights.
According to Warden, original insights can enhance visibility by 30 to 40%.
AI evaluates not just what I convey but also what others say about my brand.
This includes reviews, discussions on platforms like Reddit and YouTube, media mentions, and customer conversations.
Warden warned that conflicting signals could prompt AI to flag my brand as unreliable.
Consistency across these channels creates what he called a “consensus signal” that AI systems can trust.
One of our biggest challenges is organizational, as visibility isn’t just a channel issue; it’s an organizational one.
Currently, responsibilities are fragmented. SEO teams focus solely on rankings, PR and brand teams manage messaging, and growth teams conduct experiments. This leaves no one clearly owning AI visibility.
This fragmentation leads to inconsistent signals and missed opportunities for us.
To truly compete, we need alignment across teams, working on a shared strategy about how my brand appears wherever LLMs gather data.
Meanwhile, traditional performance metrics are unraveling.
Many marketers, including myself, notice a gap where rankings hold steady, but traffic declines. Meanwhile, leads might increase, yet attribution remains murky.
Warden explained that demand remains, but traffic no longer serves as its proxy. Our content is utilized, but not in ways directing users back to us.
This creates a growing disparity between impact and the ability to measure that impact accurately.
The nature of competition has evolved. I’m no longer vying for a mere position; instead, I’m competing to be featured in a synthesized AI answer.
Authority, once easier to influence, now hinges on external validation—emphasizing what others say over what I publish.
Algorithms have shifted from being my allies to arbiters of meaning, marking a significant change in search dynamics since Google itself emerged.
AI has not altered what makes a brand strong but has transformed how that strength is measured and rewarded. The brands that win today will build real authority in a focused niche, publish original and high-value content, and ensure consistent messaging across every platform.
The need for consistent third-party validation across an ecosystem is paramount.
As Warden urged, I must make it impossible for LLMs to ignore my brand.
Inspired by this post on Search Engine Land.


In March 2026, I, along with my research team, delved into the world of solutions used by B2B distributors, manufacturers, and enterprise commerce businesses. Our goal was simple: to find the best tools to connect ERP systems to eCommerce platforms. We studied 34 products spread across three categories: dedicated middleware connectors, ERP-native proprietary storefronts, and general-purpose iPaaS platforms. Each type made our list because they
Inspired by this post on First Page Sage Blog.


Microsoft has just rolled out a suite of updates across Microsoft Advertising, and I couldn



Inspired by this post on Search Engine Land.


Your dashboard is green, the meeting starts soon, and you still cannot answer the question that matters: what changed, why did it change, and what should the team do next?
That is a reporting-system problem, not a chart problem. Modern marketing analytics should connect business outcomes to channel activity, preserve the definitions behind every metric, expose uncertainty, and deliver the next decision without forcing someone to reconstruct the analysis during the meeting.
Most bloated reports begin with a harmless question: what data can we pull? Every available metric gets added, the dashboard becomes comprehensive, and the decision it was meant to support disappears.
Reverse the sequence. Before choosing a connector, chart, or reporting platform, write a one-sentence measurement brief:
This report helps [owner] decide [action] at [cadence] by comparing [outcome] with [baseline], using [drivers] to explain the result and [guardrails] to prevent a bad trade-off.
A paid media lead might need to reallocate campaign budget each week. A content lead might need to decide which topics deserve an update, expansion, or new format. An SEO lead might need to distinguish a visibility problem from a conversion problem. These decisions require different evidence even when they draw from the same underlying data.
Assign every metric a role. If a metric has no role, remove it from the primary report.
| Metric role | Question it answers | Marketing example | How it should affect action |
|---|---|---|---|
| Outcome | Did the work produce the intended business result? | Qualified conversions, pipeline, revenue, retained customers | Determines whether the strategy is working |
| Driver | What directly influenced the outcome? | Qualified traffic, landing-page conversion rate, lead acceptance | Identifies where to intervene |
| Diagnostic | Where did performance change? | Campaign, query group, page type, audience, device, video | Narrows the investigation |
| Guardrail | What must not deteriorate while the team optimizes? | Acquisition cost, lead quality, unsubscribe rate, brand demand | Prevents a local gain from becoming a business loss |
This hierarchy corrects a common reporting mistake. Impressions, views, clicks, and engagement can be useful drivers or diagnostics, but they do not automatically become business outcomes because they are easy to retrieve. Likewise, a channel-level return figure is not trustworthy unless the report states what counts as a conversion, which costs are included, and how credit is assigned.
Record five items beside every primary outcome: its definition, owner, data system, update cadence, and attribution rule. If attribution is involved, also state the model, lookback window, reporting timezone, currency treatment, and whether the metric uses event time or processing time. There is no universally correct attribution model. There is only a model that is explicit enough to interpret and consistent enough to compare.
Set action rules before looking at the latest result. The rule does not need an invented universal threshold. It can be operational: investigate when an outcome moves outside its expected range, when a guardrail worsens, when the data is stale, or when two systems no longer reconcile. Precommitting to the rule reduces the temptation to invent a convenient explanation after seeing the chart.

A polished dashboard cannot repair inconsistent definitions underneath it. If paid media uses platform-reported conversions, analytics uses attributed sessions, sales uses accepted opportunities, and finance uses recognized revenue, placing the figures on one page does not make them comparable.
Create a small data contract for each reporting dataset. It should specify:
Grain is the detail most likely to prevent a silent reporting error. Joining campaign-day costs to lead-level conversions can multiply spend when several leads share the same campaign and date. Aggregate both datasets to a compatible grain before joining them, or model the relationship so the cost appears only once. After every join, compare row counts and totals with the inputs.
Separate period reporting from cohort reporting. A period view answers what happened during a selected date range. A cohort view follows people, accounts, campaigns, or content acquired in a particular period through later outcomes. A recent acquisition cohort may look weak simply because its conversions have not had time to mature. Label incomplete cohorts instead of presenting them as final.
Run a compact quality checklist before publishing any result:
Do not hide a reconciliation gap with a calculated adjustment. If two systems answer different questions, label the difference. If they should match and do not, hold the affected conclusion until you know why. A visible limitation is manageable; an invisible one becomes a decision error.
A modern reporting stack does not require one tool to extract, clean, model, visualize, explain, and distribute everything. It works better when each layer has a narrow responsibility:
Dashboards are effective presentation surfaces when stakeholders need filters, recurring monitoring, and a shared view without access to every backend system. A Looker Studio report can, for example, connect YouTube Analytics data, support customized views, and distribute scheduled PDF snapshots. That makes it useful for a channel owner who needs repeatable visibility rather than a custom analysis every morning.
Keep the dashboard when its data volume is manageable, the transformations are simple, refreshes complete reliably, and an analyst can trace a wrong number back to its origin. Move complex logic upstream when the same calculated field is copied across pages, manual updates recur, refreshes become fragile, or debugging requires a long sequence of interface clicks. Broad datasets and accumulated business logic can make a dashboard slow to change, difficult to debug, and vulnerable to dataset limits.
Code is a better home for repeatable extraction, normalization, backfills, joins, tests, and calculations that need review. It gives you files that can be compared, versioned, and rerun. That does not mean every marketing team needs to replace every dashboard. A practical architecture keeps a familiar dashboard at the front while moving fragile transformations into a controlled pipeline behind it.
APIs are retrieval mechanisms, not guarantees of completeness. For every API connection, record the account or property queried, requested fields, filters, pagination behavior, expected refresh schedule, and the response received when data is unavailable. Keep credentials outside report code, grant only the access required, and plan for permission revocation. A successful request proves that data arrived; reconciliation proves that the right data arrived.
AI coding assistants can reduce the effort required to scaffold connectors, transformations, tests, and report components. Natural-language specifications can help tools such as Claude Code and OpenAI Codex assemble multistep reporting workflows. Treat the generated work as a draft implementation. Review the query grain, inspect joins, run tests, protect secrets, and compare outputs with authoritative systems before a generated number reaches a stakeholder.
Use AI differently in the analysis layer. Ask it to identify anomalies worth investigating, draft plain-language explanations from approved metrics, or translate a validated analysis for different audiences. Do not let it infer causation from a correlated chart or invent a reason for a movement that the data cannot explain. The final narrative should distinguish among a measured fact, an analyst interpretation, and a proposed test.

One dashboard should not try to answer every question for every person. An executive wants to know whether the business outcome changed and whether intervention is needed. A channel operator needs enough detail to choose the intervention. An analyst needs access to definitions, segments, and reconciliation evidence.
Build three layers, even if they live in the same reporting product:
Put context next to the metric it qualifies. A global note at the bottom of a long report will not protect a chart at the top from misinterpretation. Each primary view should show its date range, comparison basis, filters, timezone, attribution label, refresh timestamp, and any material gap in coverage.
Add a short narrative block to every decision view:
Be strict about causal language. If a campaign change and a conversion change occurred together, say they coincided unless the measurement design supports a stronger claim. If an experiment or another credible identification method isolates the effect, explain that method. Precision in the wording is part of analytics quality.
Annotations should capture business events that a chart cannot know: a campaign launch, budget change, tracking migration, site release, promotion, pricing change, consent update, or outage. Store the event date, owner, affected scope, and a brief description. An annotation is a lead for investigation, not automatic proof that the event caused the movement.
Distribution needs the same discipline as analysis. A scheduled PDF is a fixed snapshot, so include its reporting window and data cutoff. Link it to the interactive view when recipients may need filters or diagnostics. Archive material snapshots used for recurring business decisions; otherwise a later refresh can leave the team debating a number that no longer appears on screen.
Access is part of report design. Stakeholders should not need administrative access to every marketing platform simply to read an approved result. The reporting team, however, must document which account and permission power each connection. With YouTube Analytics, a report builder who does not own the channel may need Manager permission and the Channel ID entered through the connector’s advanced settings. Test delegated access with the actual reporting identity instead of assuming that a visible channel in YouTube Studio will automatically appear in the reporting connector.
A wholesale reporting rebuild creates too many simultaneous unknowns. Start with one recurring workflow that consumes meaningful time, has a known audience, and regularly produces a decision. A pre-meeting channel report, weekly SEO performance brief, or campaign pacing view is a better migration candidate than an enterprise-wide measurement platform.
The parallel run matters because two reports can display plausible but different numbers. A discrepancy may come from timezone boundaries, attribution logic, late-arriving conversions, deduplication, renamed dimensions, incomplete pagination, or a genuine bug. Matching the old number is not always the goal if the old logic was wrong, but every difference should have an explanation.
Give the finished workflow a runbook. It should tell another qualified person how to trigger a refresh, locate logs, rerun a failed period, backfill data, rotate credentials, verify source totals, publish the output, and roll back a breaking change. Include the last known successful run and the owner of each upstream dependency.
Measure the reporting system itself. Track whether scheduled runs complete, whether data meets its freshness expectation, whether reconciliation tests pass, whether recipients receive the right artifact, and whether decisions and owners are captured. The point is not to create a dashboard about dashboards. It is to notice reliability problems before they become meeting problems.
Choose the recurring report that causes the most avoidable pre-meeting work. Write its decision contract, mark every metric as an outcome, driver, diagnostic, or guardrail, and remove anything that serves no decision. That small redesign will show you exactly where the next improvement belongs: the definition, the data pipeline, the analysis, or the delivery.
