I’m thrilled to share how Yahoo Scout is revolutionizing the way we experience AI-powered searches. By anchoring responses in Yahoo’s esteemed content ecosystem, it ensures that the information we receive is not only consistent but also reliable.
By prioritizing sourcing, consistency, and enduring distribution, Yahoo Scout flips traditional AI search paradigms on their heads. This approach not only enhances user trust but also sets a new standard for how search engines can function within a trusted network.
I’ve been contemplating how even when content ranks well on search engines, it can still falter when it comes to AI retrieval. These AI systems assess pages very differently, based not just on their rank, but also on how information is extracted, embedded, and structured.
There’s an intriguing disconnect between traditional ranking and being successfully parsed by AI. A webpage can comply with excellent SEO guidelines and still miss the mark with AI-generated responses and citations.
In many situations, content quality isn’t the issue. It’s about whether the information can be reliably extracted after being segmented and embedded by AI systems.
This challenge is becoming increasingly common as search engines view pages as complete entities, but AI systems dive into the raw HTML to extract meaning from fragments rather than entire pages.
Crucial insights can get lost if they’re not appropriately structured or if they rely too heavily on visual rendering or inference.
This leads to a divergence between what’s visible in search and what’s accessible via AI, where content might exist in an index but lacks substantial meaning for AI retrieval.
The visibility gap is something I’ve been grappling with: Understanding the difference between ranking versus retrieval is key.
As search winds its processes around rankings, AI systems engage with fragments operated within a different representation of similar information. It’s here the visibility gap takes shape.
A page might rank high, but if its embedded content is incomplete or poorly organized, then the AI retrieval process becomes unreliable.
Treat retrieval as an entirely unique visibility factor. It doesn’t override SEO, but increasingly defines whether content can be effectively surfaced, summarized, or cited when AI filters come into play.
Another structural issue arises when content never even becomes accessible to AI. Many AI crawlers only parse raw HTML without executing JavaScript or client-side rendering. This creates blind spots, especially for JavaScript-heavy sites where the core content may appear in Google’s index but remains invisible to AI.
Testing if your content appears in initial HTML is quite straightforward. Simply inspect the HTML response at fetch time rather than the version rendered in a browser.
Running requests with AI user agents like “GPTBot” reveals if your site returns blank HTML even if it appears fully populated to users, highlighting its absence in initial responses.
Tools like Screaming Frog can validate this at scale. Disabling JavaScript rendering can reveal what AI systems see—if your essential content only displays with JavaScript, it can be indexed by Google’s search but not by AI retrieval systems.
Keep in mind that even with content returned, excessive code and scripts can hinder extraction by AI systems. Cleaner HTML results in more reliable embeddings, enhancing AI visibility.
To tackle this, deliver fully rendered HTML when AI systems fetch your content. Pre-rendering can often fix these retrieval issues, ensuring content is present in initial responses.
Delivery can be managed effectively at the edge layer, providing AI crawlers with complete pages instantly. Human users receive a dynamic version while AI sees what it needs to extract meaning.
If pre-rendering isn’t viable, focus on ensuring primary content is accessible in a clean initial HTML response, even without script execution.
Columns laden with excessive markup can interfere with proper extraction, diminishing the content’s value.
The next structural failure to consider is when content is optimized for keywords rather than the entities AI seeks. Traditional SEO applies keyword relevance, but AI retrieves based on entity relationships.
Without clear definition, entity signals can weaken, causing pages to underperform in retrieval even if they rank well for queries.
AI evaluates sections independently once extracted, making the consistency of header tags essential to maintaining coherence.
Ensuring sections have a single, defined purpose allows for better embedding when isolated from larger context.
Finally, conflicting signals or metadata can dilute the semantics retrieved by AI, creating noise and ambiguity.
SEO doesn’t have to mean choosing between ranking and retrieval anymore. Both must be prioritized to succeed in today’s landscape.
Your map-pack position has not moved, yet calls and website visits are down. Before you blame demand, seasonality, or your sales team, inspect the result customers actually saw. An AI-generated local answer may have shortened the list, substituted different businesses, or removed the call and website controls that once turned visibility into action.
Your local search program now has to answer four separate questions: Was your business available to the search system? Was it included in the result? Could the searcher act from that result? Did the interaction become a lead or customer? A ranking report answers only part of the second question. Here is how to measure and improve the rest of the funnel.
A stable rank can conceal a smaller conversion opportunity
The traditional local pack gave businesses a familiar bargain: earn a prominent position and receive a visible route to a phone call, website visit, or direction request. AI local results change both sides of that bargain. They can show fewer businesses, choose a different set of businesses, and present a generated explanation without the action buttons attached to a conventional listing.
The reduction is not merely theoretical. Sterling Sky’s 2026 market analysis found that AI local packs surfaced only 32% as many unique businesses as traditional map packs. The total number of visible businesses fell in 88% of the 322 markets examined. That does not establish an identical loss for every industry or location, but it shows why a business can retain its conventional rank while losing exposure in the interface customers increasingly encounter.
Advertising adds another layer. Sponsored listings, Local Services Ads, and expanded Google Ads units can occupy space around or inside local results. In some layouts, organic listings lose their direct call or website controls even when the businesses themselves remain visible. Your listing can therefore register an impression without offering the same conversion opportunity that an impression used to represent.
This is the practical meaning of zero-click local search. It does not always mean that the searcher received no value or that your business received no exposure. It means the result may satisfy part of the decision journey inside Google while giving you less traffic, less interaction data, and fewer immediate actions.
Key takeaways
A traditional map-pack rank measures one result type, not your visibility across AI answers, paid local units, and other discovery surfaces.
Track inclusion and actionability separately. Being named in an AI answer is not equivalent to receiving a call button or website link.
Treat a decline in actions per impression as a funnel diagnosis problem before treating it as a ranking problem.
Audit business identity, primary category, services, and real-world positioning before investing in another round of authority building.
Use paid local search to fill a verified conversion gap, then judge it by qualified outcomes rather than the visibility it buys.
Build a scorecard around the local search funnel
Start by retiring the idea that one visibility number can describe local performance. A useful scorecard separates availability, inclusion, actionability, and outcomes. This distinction prevents you from applying the wrong fix to the wrong failure.
What you observe
What it may mean
What to inspect next
Traditional rank is stable, but calls and website visits fall
The visible surface or its action controls changed
Capture the actual results, including AI packs, ads, and the presence of call, website, booking, and direction controls
Your business appears in the traditional pack but not the AI local answer
You may have an eligibility, classification, or corroboration gap
Compare your business name, primary category, services, local pages, structured data, and third-party descriptions
Your business is mentioned by AI, but no direct action follows
You have exposure without an immediate conversion path
Check whether the result links to your site or profile, then strengthen owned conversion paths and evaluate paid coverage
Impressions remain steady while the action rate declines
The denominator may include less actionable exposure
Review calls, clicks, bookings, and direction requests independently instead of treating impressions as visits
Both impressions and actions move sharply
Demand, seasonality, tracking issues, campaigns, or interface changes may be interacting
Annotate known platform issues and paid activity before assigning the movement to SEO
Build the scorecard from a fixed set of commercially important service-and-location queries. For each query, record which surface appears, which businesses are included, how each business is described, and which action controls are available. Keep the location, device context, and query wording consistent when comparing observations. A national rank scan cannot represent what a customer sees from a particular service area.
Add an AI inclusion measure alongside your conventional rank: the share of sampled AI local answers in which the business appears. Label it as a sampled visibility metric, not an official Google ranking. Also record the context of the mention. A recommendation for your core service is materially different from a passing mention or an appearance for a service you do not provide.
For engagement, calculate a diagnostic action rate by dividing recorded profile actions by impressions, while preserving calls, website clicks, bookings, and direction requests as separate lines. This rate is not a perfect conversion metric. AI-generated mentions can count as impressions even when they do not produce the familiar listing actions, and current reporting does not cleanly separate every organic, paid, and AI exposure. Its value is diagnostic: it tells you when the relationship between exposure and action has changed.
Do not stop at Google Business Profile data. Connect tagged website visits, call records, booking completions, form submissions, and qualified leads wherever your systems permit. A call count tells you whether the interface generated activity. A qualified-lead count tells you whether that activity was commercially useful. Preserve both because a campaign can raise calls while lowering lead quality.
Annotate the scorecard when advertising changes, tracking fails, an API issue is known, or seasonal demand moves. U.S. action trends have been less stable than trends in markets exposed to fewer search-interface experiments, which supports investigating result-format changes without proving they caused every decline. An annotation keeps a coincidental movement from becoming an expensive SEO diagnosis.
Fix AI eligibility before chasing another ranking gain
Traditional local SEO asks how strongly a business competes on proximity, relevance, prominence, reviews, citations, and engagement. AI-mediated local search adds an earlier gate: whether the system considers the business an appropriate candidate for the specific request.
This is the difference between ranking and eligibility. A ranking problem means the system understands what you are and prefers another eligible business. An eligibility problem means the system may not place you in the candidate set at all. More links or reviews will not reliably solve a classification mismatch.
Run the eligibility audit in this order:
Write the real-world promise in one sentence. State what the location actually does, for whom, and where. Use this as the control statement against which every profile, page, and citation is checked.
Verify the business name. It should represent the name used in the real world, not a string expanded with services or locations for ranking purposes. A manipulated name may create inconsistency instead of clarity.
Reassess the primary category. Choose the category that best describes the location’s main operation. Do not use an aspirational category simply because it matches a valuable query.
Reconcile services with operations. The profile service list, local landing page, navigation, visual assets, and customer-facing language should agree about what the location provides. Remove stale services and add real services that are missing.
Check location boundaries. Make the address, service area, hours, and availability claims consistent wherever they appear. Do not imply a staffed location or service footprint that does not exist.
Inspect the machine-readable version. LocalBusiness JSON-LD should mirror the visible page and the verified business facts. Use the most specific accurate business type available, and keep core properties such as name, URL, telephone, address, opening hours, and service information aligned with the customer-facing content.
Retest the query set. Separate queries where you are absent from queries where you appear but rank poorly. The first group remains an eligibility investigation; the second can move into competitive ranking work.
Structured data is a consistency mechanism, not a way to manufacture eligibility. Marking up a service that the location does not visibly offer creates another contradiction. The same principle applies to categories and landing pages: describe the operation precisely before trying to make it look broader.
Your Google Business Profile is still central, but it is no longer the whole representation of your business. AI systems encounter business facts and reputation signals across maps, directories, review platforms, community discussions, social channels, and your own site. If those descriptions disagree, the system has to decide which version is trustworthy.
Create one governed record for each location. It should hold the approved name, address or service area, phone number, URL, hours, primary category, secondary categories, active services, accessibility details, and a short factual description. Give local operators a defined way to report temporary hours, moves, closures, and service changes. Central control protects identity; local input keeps the record true.
Then audit the places that can independently corroborate that record:
Major map and review ecosystems: correct identity and operational facts, resolve duplicate listings, and update stale categories or hours.
Industry and local directories: prioritize sources that customers in the market genuinely use rather than creating large volumes of low-value listings.
Community references: earn accurate mentions through real associations, events, partnerships, sponsorships, customer recommendations, and local coverage. Do not manufacture forum conversations or undisclosed endorsements.
Owned location pages: include the services, service boundaries, hours, contact route, local proof, and useful answers that belong to that specific location. Avoid pages that differ only by a place name.
Reviews and responses: monitor whether customer language reflects the services and experience you actually want associated with the location. Respond to factual problems and operational changes rather than inserting target phrases into every reply.
Photos and video: publish current, high-quality visuals that show the premises, team, equipment, products, or service process when those elements are relevant and safe to display. Visuals should provide evidence, not decorative stock imagery.
Useful local content answers questions that arise before and after the initial business search: service limitations, preparation, availability, local conditions, the decision process, and what happens next. Assign the content to someone who understands the location’s work. A central team can supply structure and quality controls, but it should not invent local facts on the location’s behalf.
Recover the next customer action on every surface
Eligibility gets you considered. Corroboration makes you easier to trust. Neither guarantees that the result will contain a usable conversion control. You still need a plan for the next action when Google changes the interface.
Start with the result itself. For every priority query, note whether the searcher can call, visit the site, request directions, book, or continue into another Google experience. If the business is visible but the intended action is missing, classify that as an actionability gap. Do not send the SEO team looking for a ranking fix when the interface is the constraint.
Organize content around those fuller decisions. Explain which needs the location handles, which it does not, where service is available, what information a customer should have ready, and which contact route fits the request. Use direct language that can be understood in a conversational answer. Do not bury a crucial eligibility or booking condition in promotional copy.
Paid local search becomes a tactical option when a high-value organic result repeatedly lacks the call or website control you need. Test Local Services Ads or another appropriate paid format against the specific gap you observed. Set a controlled budget, separate paid calls from organic calls where measurement permits, and evaluate qualified leads, booked work, and acquisition cost. Buying back a prominent button is useful only when the resulting customers justify the spend.
Do not assume every location needs permanent paid coverage. A location that already receives actionable organic visibility may gain little from paying for duplicate exposure, while a location pushed below ads or stripped of direct controls may have a clearer case. The decision belongs in the scorecard: interface gap, paid coverage, qualified outcome, and cost.
For a multi-location organization, review performance at the location level before rolling out a network-wide response. AI inclusion, ad pressure, community signals, demand, and conversion economics can differ by market. Use central standards for data, schema, measurement, and brand identity, then let each location supply the facts, media, relationships, and service detail that make its local evidence genuine.
Begin with one priority query and trace it from result format to qualified outcome. Record whether the location was eligible, included, actionable, and commercially successful. Once that chain is visible, you can fix the actual break instead of defending a rank that no longer guarantees the engagement you need.
I’ve discovered how custom GPTs can revolutionize how we handle SEO, transforming repetitive tasks into efficient workflows. By leveraging AI, we can speed up our processes, from planning and analysis to reporting and technical work.
If you don’t have access to paid ChatGPT, don’t worry. You can still utilize these prompts by saving them as standalone references in your notes. Remember, they’re just starting points, so modify them to fit your team’s requirements.
Working with AI requires trial and error. My advice is to start with small tasks to practice writing prompts. Iterate on them and take notes on what produces good outputs.
AI can sometimes be verbose, so it’s helpful to set strict formatting guidelines and clear context. Upload resources and articles to guide AI results, and always define the role and audience upfront.
Let’s dive into seven prompts that I’ve found incredibly useful for developing custom GPTs dedicated to planning, analysis, and ongoing SEO tasks:
1. Project plan GPT
By analyzing previous project plans, I can create a GPT that assists in drafting this year’s focus areas.
How to set it up
Input project plans from previous years.
Specify a format for consistency.
Determine the number of items or sections to include.
Include specific details unique to your team.
Optionally, integrate team feedback and retrospectives.
Example prompt
Based on last year’s project plan, outline this year’s focus. List three critical items for each quarter, ensuring at least one covers link building.
Include a one-sentence summary for each recommended item and at least two KPIs to measure success.
[Insert last year’s plan.]
Now critique the plan. Offer three reasons against focusing on these items, providing sources for your notes.
By connecting performance dashboards or custom GA reports to ChatGPT, it can handle initial issue identification. This allows me to focus on investigating critical trends.
How to set it up
Hook up reporting tools or upload data directly.
Direct AI on specific aspects to investigate.
Set frequency for data review, such as daily or weekly.
Provide examples of pages or categories to analyze.
Example prompt
Here’s the weekly site report. Analyze this week’s performance against last week’s data, summarizing sessions, conversions, and engagement.
Highlight three successes and three areas needing improvement, color-coded by significance.
[Insert report doc.]
3. Competitor analysis GPT
I’ve found it invaluable to scrutinize what works on competitor sites. This often involves tools like Semrush or Ahrefs.
How to set it up
Integrate Ahrefs, Semrush, or upload relevant reports.
Select competitors and identify top-performing pages.
List key metrics for evaluation.
Create unique prompts for various levels of analysis.
Now, more than ever, custom GPTs are making a significant impact alongside existing SEO tools and workflows. They’re not about replacing the tools we use, but about making initial tasks smoother so that we can focus on insightful and strategic actions. By integrating them into our everyday processes, from planning to technical checks, we can really enhance our productivity.
I’ve come to realize that SEO now serves as both a brand and performance channel. The traditional traffic model has been disrupted by AI Overviews and zero-click SERPs, making brand strength crucial for SEO ROI.
For years, SEO was straightforward: rank higher, get more traffic, then boost the sales pipeline. However, this simple equation is rapidly evolving, much to the frustration of marketing leaders.
With AI Overviews and users getting answers directly from LLMs, the idea of “rank and receive traffic and leads” is less effective now. Even top keyword positions don’t guarantee the clicks they once did.
This shift has sparked challenging discussions in boardrooms. Executives often question, “If traffic is down, how can we measure SEO success?”
It’s obvious now: the traffic model has changed, yet the demand for ROI remains. We must treat SEO as a brand-dependent performance channel, not just a traffic provider.
Why traffic and pipeline are no longer in lockstep
Linear attribution has never fully reflected the dynamic nature of organic search. While ChatGPT isn’t replacing Google, it’s augmenting it.
Users now verify information across platforms due to skepticism of search and LLM results. Where research once happened solely within Google’s ecosystem, it has become more scattered.
Today’s organic search is akin to a pinball machine, with buyers bouncing across channels unpredictably. This introduces complexity that traditional attribution software struggles to follow.
Such complexity has broken the linearity executives crave. Traffic and pipeline charts, once aligned, now often diverge.
Across B2B SaaS portfolios, a common pattern emerges: organic sessions may be flat or declining, yet rankings for high-intent terms stay stable, and the pipeline from organic search grows.
This mismatch doesn’t indicate SEO failure. Rather, it shows that traffic is no longer a reliable business impact measure.
The traffic lost to zero-click searches often consists of informational, low-intent content. What remains is higher-intent traffic, closer to conversion.
We’re seeing the “atomization” of search demand. Short-head, broad keywords are declining, while specific, long-tail queries with higher intent are rising.
Many leaders mistakenly react to dropping sessions by pushing for quantity, aiming to regain the lost numbers through top-of-funnel content. This often inflates vanity metrics without delivering qualified leads.
SEO ROI is now the downstream outcome of brand traction
For years, SEO was viewed as a pure performance channel. We believed optimizing some keywords would suffice.
In reality, SEO has always depended on brand strength. The rise of AI-driven engines highlights this, expecting reputations, not just keywords.
If your brand lacks authority, technical optimizations alone won’t elevate your status. Brand strength determines organic performance limits. Search engines seek web-wide consensus, and weak associations hinder results.
Brand strength for LLMs means owning topical authority, aligning with customer queries, being validated by trusted sources, and having clear positioning.
SEO captures pre-existing demand validated by your brand, not creating it from nothing.
The new defensibility metrics for SEO
As traffic no longer headlines KPIs, new defensibility metrics are necessary. Successful teams focus on revenue and reputation impact, not just volume.
Metrics proving business impact include stable top-10 rankings for commercial keywords, increased Ahrefs traffic value, stable solution page traffic, growing homepage traffic, and developing LLM referral traffic.
When pipeline per organic visitor rises, even with falling sessions, the dialogue shifts from “SEO is broken” to recognizing SEO’s evolution.
Modern SEO is moving from acquisition to influence
Successful SEO isn’t about recovering traffic but influencing buyer decisions and enhancing organic visibility. In an AI-first context, zero-click doesn’t imply zero-value.
SEO remains key in building market readiness, positioning brands as authorities even before buyers enter the funnel.
When I first stumbled upon the concept of query fan-out, I realized how misunderstood it often is in the world of AEO and SEO. It’s fascinating how AI searches can take a single prompt and transform it into numerous sub-queries, expanding the scope of search in unimaginable ways.
Understanding this process opened my eyes to the hidden potential these sub-queries hold. By leveraging the data generated from them, I discovered new strategies to enhance SEO effectiveness, making my digital marketing efforts more robust.
You have a decision to prepare for, but not yet a reliable switch to flip. Google has discussed letting publishers opt out of AI Overviews and AI Mode, yet it has not disclosed a clear, feature-specific implementation. Adding a guessed crawler rule or sitewide directive now could affect more than the AI feature you meant to control.
Do the policy work first. Decide which content you would exclude, what outcome would justify exclusion, how you would detect collateral damage, and what would trigger a rollback. Then, if Google releases a documented control, you can test it as an operating decision instead of reacting with a blanket yes or no.
The opt-out question is ahead of the actual control
Google has been exploring ways for websites to opt out of AI-generated search features. What publishers still need is the operational detail: whether a control would apply to AI Overviews, AI Mode, or both; whether it could be used on individual URLs or only an entire site; how quickly a change would take effect; and whether it would alter eligibility for traditional search.
Until those questions have documented answers, nobody can responsibly give you an exact implementation recipe. A directive intended for an AI training crawler is not automatically a control for an AI-generated search result. A general search restriction is not automatically limited to AI. The names may sound related, but the scope and business consequences are different.
The disagreement makes sense because “block AI” is not a business objective. One publisher may prioritize broad discovery. Another may place more value on controlling the reuse of expensive original work. A third may want visibility in AI results but only when those appearances send qualified readers or reinforce the brand. You cannot resolve those positions with a technical toggle alone.
Keep three decisions separate in every internal discussion:
AI training access: whether a named crawler may collect content for a training-related purpose.
Traditional search access: whether Google can crawl, index, and present a page in established search results.
AI search presentation: whether content can contribute to or appear in AI Overviews and AI Mode.
Build the policy around content classes, not one domain-wide answer
A sitewide decision is simple to announce and difficult to evaluate. Your domain probably contains pages with different economics and different jobs: original reporting, evergreen reference material, product or service pages, subscriber content, documentation, archives, and pages built primarily to acquire search visitors. A future control may or may not support URL-level rules, but your policy should be ready for that possibility.
Create an inventory by template or content class. You do not need to classify every URL manually. Start with the groups that account for most of your search traffic, revenue, subscriptions, leads, or editorial investment.
Name the page class. Use a stable label such as original news, analysis, evergreen guide, product page, documentation, archive, or subscriber-only content.
State its primary job. Choose one: attract new readers, convert demand, retain subscribers, establish authority, support customers, or generate direct revenue.
Record its dependency on Google discovery. Use your own impressions, clicks, landing sessions, conversions, and revenue rather than an editorial assumption.
Identify the use you want to control. Say “AI Overviews and AI Mode” if that is the target. Do not write only “AI,” because that leaves training, search presentation, and other uses mixed together.
Assign a provisional status: allow, exclude when a verified control exists, or include in the first test.
Name the owner who can approve implementation and the owner who can order a rollback.
The three provisional statuses keep uncertainty visible without forcing a premature technical change:
Allow: discovery is the dominant objective, so the current state remains in place unless measured harm changes the decision.
Exclude when possible: the content conflicts with a declared reuse or rights policy, but implementation waits for a documented control whose scope is understood.
Test: the trade-off is uncertain, so the content becomes a candidate for a limited, reversible experiment.
Add the reason beside every status. “Editorial leadership requested it” is an approval trail, not a decision rule. A usable reason sounds like this: “These pages depend on search acquisition, so exclusion will be retained only if targeted AI use declines without pushing qualified organic visits or conversions below our predeclared guardrails.”
If Google ultimately offers only a domain-wide setting, your classification work still matters. It shows which page groups carry the benefit and which carry the cost. That gives leadership a defensible basis for accepting or rejecting the broader control.
Decide what success and failure look like before changing anything
A publisher test fails when the team changes a setting first and chooses the interpretation later. Traffic can move for many reasons. If your success criteria remain unwritten, almost any result can be used to defend the decision someone already preferred.
Build a measurement sheet with four layers:
Business outcome: qualified leads, purchases, subscriptions, advertising value, or another result tied to the selected page class.
Search referral outcome: impressions, clicks, click-through rate, landing sessions, and the queries sending those visits.
AI feature observation: whether the chosen URLs or brand appear for a fixed set of queries in AI Overviews or AI Mode.
Technical guardrails: continued crawling, indexation, and appearance in the traditional search surfaces you intended to preserve.
Do not assume your normal analytics can isolate every AI feature appearance. If they cannot, create a manual observation set. Select queries before the test, record the page and feature being checked, keep the location, account state, and device conditions as consistent as practical, and save dated evidence. The purpose is not to estimate all AI visibility from a small sample. It is to check whether the behavior of known query-URL pairs changed after the control.
Use queries where the page had previously appeared in the targeted feature whenever possible. If an AI Overview does not appear for a query on a later check, that single absence does not prove the exclusion worked; the feature itself may not have appeared. Verification needs to distinguish “the feature was present without our content” from “the feature was not present at all.”
Write the retention rule in advance. A practical template is:
We will retain exclusion for [content class] only if the targeted use declines in our logged sample, organic search outcomes remain above our chosen floor, the primary business metric stays within its guardrail, and traditional search eligibility shows no unintended change.
Publisher decision template
Choose the floors from your own historical volatility and business tolerance. There is no credible universal percentage that tells every publisher when loss of reach is worth greater content control. A subscription publisher, a lead-generation site, and an advertising-funded newsroom can assign very different values to the same traffic movement.
Test a documented control with the smallest reversible scope
When Google publishes an actual control, verify what it governs before deploying it. The label is not enough. Read for its target feature, supported scope, interaction with traditional search, activation behavior, verification method, and rollback procedure. If the documentation does not answer one of those questions, record it as an unresolved risk rather than filling the gap with an assumption.
Then run the test in this order:
Choose a narrow cohort. Prefer one content class or template over the entire site when the documented control permits it.
Select a comparison cohort. Match pages as closely as practical on purpose, query demand, historical performance, update pattern, and publication timing.
Capture a baseline. Include a period that reflects your normal publishing or business cycle, and note promotions, seasonal events, migrations, algorithm changes, or major editorial updates that could distort it.
Freeze avoidable confounders. Do not simultaneously rewrite titles, change internal links, redesign templates, or move URLs unless those changes are part of the test.
Apply one documented control. Log the exact setting, scope, time, implementer, approver, and expected outcome.
Verify the target behavior. Check the tracked query-URL pairs and confirm that any observed change concerns AI Overviews or AI Mode rather than a broader loss of search access.
Compare business results and guardrails. Use the predeclared rule, not a newly chosen metric that happens to support the preferred conclusion.
Roll back if the blast radius is larger than intended. Preserve the implementation log so the team can separate recovery from later unrelated changes.
If the control is sitewide only, you lose the cleanest form of an internal comparison. Do not pretend a before-and-after chart proves causation. Keep a dated change log, use the same tracked query set, document concurrent events, and require stronger evidence before making the setting permanent.
Operational cost belongs in the result as well. A page-level control that must be maintained across several publishing systems creates a different burden from a stable sitewide setting. Record implementation time, quality-assurance failures, ownership gaps, and rollback effort. A policy that cannot be maintained reliably is not an effective control, even when its strategic intent is sound.
Key takeaways
Google has discussed publisher opt-outs for AI Overviews and AI Mode, but a clear feature-specific implementation has not been established here.
Blocking an AI training crawler is not the same as opting out of an AI-generated search feature.
Classify content by business purpose and Google dependency before choosing allow, exclude, or test.
Predeclare the target behavior, primary business metric, search guardrails, technical checks, and rollback condition.
When a documented control arrives, begin with the smallest reversible cohort its scope permits.
Your useful next step is a one-page control brief, not a speculative configuration change. Assign an owner, classify the page groups that matter, capture their baseline, and list the documentation questions Google must answer. When a real control becomes available, you will be ready to evaluate it with evidence instead of making a domain-wide bet under deadline pressure.
If your pages rank in Google but disappear when a buyer asks ChatGPT, Gemini, or Perplexity what to choose, you do not have a conventional ranking problem. You have a chain-of-trust problem. The assistant must be able to reach your information, understand what it means, reconcile it with information elsewhere, and decide that it is relevant and credible enough to use.
That changes where you should start. Publishing more content or adding AI-related keywords will not repair a blocked crawler, a confused business identity, or conflicting location data. Audit the full path to an AI answer, then fix the earliest point at which your visibility breaks.
AI visibility is a connected system, not a single ranking
Traditional rank tracking asks where a page appears for a query. AI search visibility covers several different outcomes: whether an assistant mentions your brand, uses your content, links to your site, states your facts accurately, or recommends you as a suitable choice. A brand can succeed at one outcome and fail at another.
A practical audit separates the system into these stages:
Access: Can retrieval systems and permitted bots reach the important public pages without being blocked by robots rules, authentication, a firewall, or a challenge page?
Interpretation: Does each page make the subject, claim, location, product, and relationship between entities explicit?
Corroboration: Do your website, business profiles, reviews, and other public records agree on the facts that matter?
Selection: Does your information answer the user’s actual task well enough to be cited or recommended?
The order matters. Better copy cannot compensate for a page that cannot be retrieved. Perfect crawl access cannot resolve two different addresses for the same location. Consistent facts do not guarantee selection when the page never answers the question behind the prompt.
What you observe
Likely bottleneck
First check
Important public pages are absent from retrieval or crawler logs
Access
Robots rules, authentication, CDN controls, and firewall challenges
Assistants state an old address, name, or service detail
Interpretation or corroboration
The canonical page and every prominent public profile carrying that fact
Your pages are cited for facts, but your brand is not recommended
Confidence or task fit
Reputation signals, comparative evidence, and whether the offer fits the prompt
Google visibility is strong while assistant visibility is weak
Selection
A separate prompt-level baseline for each assistant
Do not label every absence a crawl problem. If an assistant accurately summarizes a page but does not mention your brand, it obtained the information through some path. Your next work belongs farther down the chain, usually in attribution, corroboration, or selection.
Prove access before you rewrite the content
Start with the pages closest to discovery, evaluation, and conversion. These are usually your main service or product pages, location pages, comparison resources, original research, documentation, pricing explanations, and pages that answer recurring pre-sale questions. The goal is not to make every URL equally prominent. It is to ensure that your most useful public information is technically reachable.
Fetch each priority URL without a login. Confirm that the response contains the intended page, not a consent wall, security challenge, empty shell, or error message.
Read robots.txt as a set of instructions. Look for broad disallow rules, overlapping bot-specific directives, and stale rules left by a migration or staging environment.
Inspect controls outside robots.txt. A CDN, web application firewall, rate limit, or bot-management product can reject a request even when the robots file allows it.
Follow redirects to the final page. The destination should remain public, load the substantive content, and identify the stable canonical version of the URL.
Review server and security logs. Look for successful requests, repeated rejections, redirects, and challenge responses associated with the crawlers you intend to permit.
Retest after changing a rule. A configuration edit is not proof that the final URL is reachable through the full delivery stack.
Refining robots.txt and maintaining a useful llms.txt file can improve the conditions under which AI bots discover your content. The files serve different jobs. Robots.txt communicates crawl permissions. An llms.txt file can act as a concise map to important, canonical resources.
If you publish llms.txt, keep it selective. Point to pages that explain who you are, what you offer, and where your strongest reference material lives. Remove redirected, duplicated, expired, and thin URLs. Update the file when important destinations change. A stale directory creates another version of your site for machines to reconcile.
Treat llms.txt as a signpost, not an access-control system or a visibility guarantee. It does not override robots.txt, authentication, firewall rules, or a broken page. It also does not replace ordinary internal links and crawlable site architecture. Do not expose private, administrative, customer, or staging URLs merely to make a crawler test pass.
Your access audit passes when a priority public URL can be retrieved without credentials, returns the intended substantive content, survives the redirect path, identifies a stable canonical destination, and is not rejected by a rule or security control you meant to allow.
Make your identity, evidence, and suitability easy to resolve
Build pages around complete, extractable answers
An extractable page does not need robotic prose. It needs explicit relationships. A reader and a retrieval system should both be able to identify what the page answers, which entity the answer concerns, where the claim applies, and what supports it.
Use a descriptive heading that matches a real question or decision rather than a vague slogan.
Name the company, product, service, or location before relying on pronouns such as it, this, or we.
Give the direct answer first, then add conditions, exceptions, evidence, and next steps.
Keep supporting evidence close to the claim it supports. Do not make a reader hunt through unrelated pages to understand the basis of an important statement.
Distinguish facts from positioning. Availability, location, compatibility, and eligibility should not be buried inside promotional language.
Use internal links with descriptive anchor text so the relationship between an overview, supporting evidence, and a detailed resource is apparent.
Keep structured data, including JSON-LD, aligned with the visible page. Markup should clarify information that users can verify on the page, not introduce a separate set of claims.
Page structure is especially important when a fact has a limited scope. If a service is available only in a particular region, a feature applies only to one plan, or a result depends on stated conditions, carry that qualifier into the answer itself. A technically accurate sentence can still create a wrong AI answer when its limiting context is several paragraphs away.
Give every team one record of core business facts
Create an internal fact sheet for the details that assistants and customers must not get wrong. Include the official brand and location names, canonical URLs, contact details, addresses, operating hours, service areas, categories, and current descriptions of the main products or services. Assign an owner to each field so an operational change has somewhere to go before conflicting versions spread.
Audit those facts across your own site and the external platforms likely to carry them, including Google Maps, Yelp, and Facebook. Check each location separately. A correct corporate address does not repair an incorrect branch profile, and a correct branch page does not erase stale hours elsewhere.
Consistency does not require identical marketing copy on every platform. It requires agreement on verifiable facts. Preserve platform-appropriate descriptions, but remove conflicts in identity, location, availability, and contact information. When you find a discrepancy, correct the system that owns the bad record rather than merely publishing another page with the right answer.
Treat reputation as a confidence signal, not decoration
AI recommendations are markedly selective in the local context measured by SOCi’s 2026 Local Visibility Index. Across nearly 350,000 locations belonging to 2,751 multi-location brands, ChatGPT recommended 1.2% of locations, Gemini recommended 11%, and Perplexity recommended 7.4%. Brands appeared in Google’s local three-pack 35.9% of the time. The resulting gap ranged from about three to 30 times within that dataset.
Those percentages describe a particular multi-location sample, not a universal multiplier for every query, industry, or business. They still expose a costly assumption: strong local Google performance is not a dependable proxy for AI recommendations.
Profile accuracy also differed by assistant in the same dataset. Gemini returned accurate business information in 100% of the measured cases, while ChatGPT and Perplexity reached 68%. That variation is a reason to inspect individual answers and platforms, not to calculate one blended visibility score that hides factual errors.
Ratings appeared to work more like a confidence filter than a simple ranking boost. Locations recommended by ChatGPT averaged 4.3 stars, with slightly lower averages for Gemini and Perplexity. Do not turn 4.3 into a supposed eligibility threshold; it is an observed average, not a published cutoff. Use it as a prompt to examine the underlying customer experience, recurring complaints, unresolved listing errors, and whether your public reputation supports the recommendation you want an assistant to make.
Measure mentions, citations, accuracy, and recommendations separately
A conventional position report cannot show whether an assistant named your brand, recommended it, cited it, or repeated an incorrect fact. Build a prompt-level measurement set around the tasks your audience actually performs.
Discovery prompts: The user is identifying possible approaches, providers, products, or locations.
Comparison prompts: The user is weighing alternatives against explicit requirements.
Suitability prompts: The user wants to know what fits a particular situation, industry, location, or constraint.
Factual prompts: The user needs an address, capability, policy, compatibility detail, operating hour, or other verifiable fact.
Branded prompts: The user already knows your name and expects an accurate explanation.
Non-branded prompts: The user describes the need without giving the assistant your brand as a hint.
For every test, record the exact prompt, platform, model or product surface when identifiable, location context, account state, test date, complete answer, cited URLs, brand mentions, recommendation status, and factual errors. Preserve the response itself. AI answers can vary, and a result you did not save cannot be audited later.
Keep the core metrics separate:
Visibility rate: the share of eligible responses that mention your brand.
Recommendation rate: the share that present your brand as a suitable option, not merely as background.
Citation rate: the share that link to or explicitly identify your owned content.
Factual accuracy: whether the material facts stated about your brand are correct and current.
Cross-platform consistency: whether different assistants produce materially compatible descriptions of the same entity.
A single answer is an observation, not a trend. Retest the same prompt set under documented conditions and look for direction across repeated runs. Change a small, named group of inputs, log the change, and then use the same prompts again. Otherwise, you will not know whether an apparent improvement came from your work, answer variability, or a different testing context.
Keep Google and AI results side by side, but never substitute one for the other. Fewer than half of the brands leading local Google visibility also led their sectors in AI outcomes. In retail, only 45% of the top 20 local-search brands also reached the leading group for AI recommendations. That is dataset-specific evidence for maintaining separate dashboards and separate diagnoses.
Use the following sequence to turn the audit into work:
Baseline the prompts connected to your highest-value customer decisions.
Resolve access failures on the pages that should answer those prompts.
Correct conflicting identity, location, product, and availability facts.
Rewrite weak pages so the direct answer, scope, evidence, and entity relationships are explicit.
Repair inaccurate external profiles and address the operational causes of recurring negative sentiment.
Retest the same prompt set and classify each remaining failure as an access, interpretation, corroboration, or selection problem.
Key takeaways
Google rankings are useful context, but they do not predict whether an AI assistant will cite or recommend you.
Fix the earliest broken stage: access, interpretation, corroboration, or selection.
Robots.txt and llms.txt can support discovery, but neither repairs firewall blocks, private pages, weak answers, or conflicting facts.
Your site, Google Maps, Yelp, Facebook, and other prominent profiles should agree on verifiable business details.
Structured data should reinforce visible content, not create claims that users cannot verify on the page.
Measure mentions, recommendations, citations, and factual accuracy separately for each assistant.
Review averages from a multi-location dataset are diagnostic context, not universal eligibility thresholds.
Start with one high-value query cluster rather than a site-wide rewrite. Confirm that its best pages are reachable, align the facts across your public presence, strengthen the direct answers and supporting evidence, and capture a baseline in the assistants your audience uses. That gives you a controlled unit of work and a result you can actually diagnose.
If your average positions look steady while organic growth feels weaker, you may be measuring a journey that no longer happens in the same number of steps. A person can express a fuller need in one query, receive a synthesized answer, and skip follow-up searches that once gave you several chances to earn a click.
That changes visibility in two ways. Search sessions are becoming more compressed, and AI recommendations are less stable than conventional rankings. Your response should be an intent-based system that measures repeated presence, gives machines unambiguous evidence, and still helps a person make the decision in front of them.
Search demand can persist while the journey loses steps
Datos/SparkToro behavioral data from millions of users found that desktop Google searches per U.S. user fell by nearly 20% year over year. The decline in the EU and U.K. was much smaller, at roughly 2% to 3%. This is a per-user change, not proof that Google suddenly lost its audience.
The surrounding numbers make that distinction important. Traditional search remained about 10% of U.S. desktop activity through 2025. Dedicated AI tools accounted for only 0.77%, while Google AI Mode represented about 0.06% of U.S. desktop events by December. AI adoption is growing, but those shares are too small to support a simple story in which everyone abandoned Google for a chatbot.
These figures do not prove that AI caused every missing search. They are consistent with a more practical mechanism: AI answers and instant results can resolve part of a need before a person performs a second, third, or fourth query. Search remains central, but each session may generate fewer opportunities for publishers.
Query shape is changing at the same time. Six-to-nine-word searches are increasing rapidly in the U.S. Very long queries of 15 words or more remain uncommon and volatile, but they show that people are experimenting with more complete descriptions of what they need. You should therefore plan around the decision contained in a query, not just the keyword string that introduces it.
Choose one commercially meaningful decision. Examples include selecting a product for a constrained use case, deciding whether a service fits a particular situation, or comparing two approaches.
List the modifiers that change the answer. Audience, budget, compatibility, location, urgency, skill level, risk tolerance, and intended use can turn superficially similar prompts into different decisions.
Write down the facts required to answer each version. Include suitability, exclusions, specifications, limitations, evidence, availability, and the next action.
Map every important fact to a crawlable location. A claim should have a clear home on a page, not exist only in an image, sales call, private document, or advertising campaign.
Consolidate wording variants, but split genuinely different intents. If ten phrasings lead to the same criteria and answer, one strong resource can serve them. If the criteria change, create a distinct section or page rather than forcing every audience into generic copy.
This exercise gives you an intent map rather than another keyword list. It also exposes a common visibility gap: the page may mention the right topic while failing to provide the specific facts a search engine or AI system needs to answer the actual decision.
Measure AI visibility as repeated presence, not a fixed rank
An AI recommendation is generated for a particular request and context. It is not a stored, universally ordered result. Across nearly 3,000 executions of 12 identical prompts by more than 600 volunteers, an identical recommendation list appeared fewer than once in 100 responses. Getting the same list in the same order was rarer still, at fewer than once in 1,000.
A single screenshot therefore cannot tell you that your brand ranks third in AI search. It tells you that your brand appeared third in one response. Running the same prompt once more and reporting the better result is no more defensible; it replaces one anecdote with another.
The more useful signal is visibility percentage: how often your brand appears across a defined set of valid responses. Presence proved more stable than exact order, even when the lists themselves changed. Smaller niche categories tended to produce more consistent answers than large markets, so you should not compare percentages across unrelated categories as though they shared the same competitive conditions.
Define the prompt universe before collecting results. Select the audience, decision, market, language, and meaningful constraints. Do not add favorable prompts after seeing the outcome.
Create wording variants that preserve intent. Natural prompts can differ substantially in phrasing while expressing the same underlying need. Keep these in one family.
Separate prompts when the purpose changes. A general product recommendation and a recommendation for gaming, accessibility, enterprise security, or noise cancellation are different intent families if their selection criteria differ.
Repeat tests under documented conditions. Record the product or model, interface, date, locale, login or personalization state when known, exact prompt, and complete response.
Classify the outcome before calculating a rate. A passing mention, a direct recommendation, a citation, and an accurate description are not interchangeable forms of visibility.
Aggregate by intent family. Calculate repeated presence within each decision context before combining anything into an overall number.
There is not yet a validated universal minimum number of runs, and API output may not reproduce what a person sees in a consumer interface. Treat a small sample as directional. Keep the protocol consistent, retain the underlying responses, and widen the sample before making an expensive content or positioning decision.
You can still record list order for diagnosis. A persistent pattern may lead you to inspect what distinguishes frequently preferred brands. But exact position should not become the executive KPI, agency guarantee, or performance bonus when the output is inherently variable.
Make every important claim retrievable, specific, and verifiable
The next visibility problem is eligibility: can a system identify your entity, retrieve the relevant facts, and determine whether your offer fits the user’s constraints? A page can be persuasive to a person while remaining ambiguous to a machine because the product name changes between sections, limitations are missing, specifications live in images, or structured data conflicts with visible copy.
Moving from discovery to transaction inside one AI conversation is still a forecast rather than established behavior at scale. It is nevertheless sensible to make product and service information machine-readable now. The same cleanup also helps conventional search, feeds, internal search, accessibility, and human comparison.
Use this content pattern for each important decision page:
Entity: State the exact product, service, organization, person, or location being described. Use the same canonical naming across headings, copy, metadata, and structured data.
Direct answer: Address the central decision early. Say who or what the option is for, rather than making the reader assemble an answer from feature copy.
Qualifiers: State compatibility requirements, exclusions, prerequisites, geographic limits, and material tradeoffs. Missing limits invite incorrect assumptions.
Comparable facts: Present specifications, capabilities, availability, and policies in labeled text or tables where a comparison genuinely helps.
Evidence: Add original measurements, first-party data, expert explanation, examples, or a documented method. Include enough context for someone to judge what the evidence does and does not establish.
Freshness: Show when time-sensitive facts were reviewed, and correct outdated pages instead of allowing contradictory versions to coexist.
Structured data: Apply the most specific relevant schema types and properties, using the same facts shown to the reader. Markup labels evidence; it does not replace evidence or make an unsupported claim true.
Generic summaries are easy to reproduce and hard to distinguish. Proprietary data and distinctive first-party content give other sites and AI systems information they cannot obtain from another lightly rewritten overview. The useful part is not merely owning data. You need to publish the method, scope, date, definitions, and limitations that make the result interpretable.
Specificity also protects brand accuracy. When your trial policy, service boundary, compatibility, or availability is unclear, a generative system may fill the gap with a category-level pattern that applies to competitors but not to you. Put the correction on the canonical page, align related pages and schema, and make the wording explicit enough to quote without reconstruction.
Do not create a separate thin page for every prompt variation. Build around meaning. A strong resource can answer several phrasings when the intended decision is the same, while modular sections can address the qualifiers that materially change the answer.
Treat video as visual, audio, text, and metadata
Video can supply evidence that prose struggles to carry: a product in use, a software workflow, a physical dimension, an expert’s explanation, or the exact state of an interface. AI systems can process visual frames, speech, on-screen text, and relationships between them. Some handle these streams together; others depend on separate recognition and transcription components. Either way, clarity determines how much useful information survives.
Optimize all four layers rather than uploading a polished file and relying on its title:
Visual layer: Publish crisp 1080p video where practical. OCR can struggle with footage below 360p, and enhancement cannot reliably restore text that was never captured clearly. Use high contrast, bold readable type, and close enough framing for labels and interface states to be legible.
Temporal layer: Keep a key object, label, or action on screen long enough to appear in sampled frames. Rapid cuts may look energetic to a person while causing an automated system to miss the one frame that establishes the fact.
Audio layer: Use clear speech, identify speakers, reduce competing noise, and align narration with the action on screen. Deliberate pauses can separate important statements and reduce ambiguity.
Text layer: Provide human-verified captions and a transcript. A transcript gives text-dependent systems access to the substance and reduces errors introduced by automatic speech recognition.
Metadata layer: Use accurate titles and descriptions, then add applicable VideoObject markup. Properties such as hasPart, transcript, and interactionStatistic should describe real, visible content and verified data.
Review the finished video without sound, then review only the audio and transcript. If either version loses the core claim, the layers are not reinforcing one another. Fix the asset itself before adding schema; metadata cannot rescue an unreadable demonstration, an incorrect caption, or a missing limitation.
Use a scorecard that separates exposure, accuracy, and value
Traffic remains useful, but it no longer describes the whole journey. An answer can mention your brand without linking to it, cite you without recommending you, recommend you inaccurately, or send a visitor who converts. Those are different outcomes and should occupy different rows in your reporting.
Key takeaways
Fewer searches per person do not mean Google has become irrelevant; they mean each journey may contain fewer opportunities.
An AI list position is an observation from one response, not a durable rank.
Measure repeated brand presence across defined intent families and documented conditions.
Separate mentions, recommendations, citations, accuracy, and business outcomes.
Improve visibility eligibility with explicit facts, distinctive evidence, consistent structured data, and machine-readable media.
A practical scorecard can use the following definitions. Set the inclusion rules before testing, and keep the denominator visible beside every percentage.
Metric
How to calculate it
What it helps you decide
AI visibility rate
Valid responses that mention your brand divided by all valid responses in the defined prompt set
Whether you enter the answer set for that intent
Recommendation rate
Valid responses that present your brand as a suitable option divided by all valid responses
Whether appearances are incidental or decision-relevant
First-party citation rate
Responses that cite a page you control divided by valid responses on citation-capable surfaces
Whether your own evidence is being used, rather than only third-party descriptions
Accuracy rate
Reviewed appearances with all predefined material claims correct divided by appearances reviewed
Whether greater exposure is reinforcing the right brand facts
Intent coverage
Intent families in which the brand appears divided by all intent families tested
Which audiences or use cases have evidence gaps
Human search performance
Impressions, clicks, landing-page behavior, and conversions reported by page and intent group
Whether conventional discovery and on-site usefulness are improving
Business outcome
Qualified actions, leads, sales, or other agreed outcomes from attributable journeys
Whether visibility work is connected to value rather than exposure alone
Store the prompt and complete response behind every AI observation. Also retain the model or product, interface, collection date, locale, and personalization state when known. Compare like with like. If a platform changes, preserve the old series and label a new baseline instead of hiding the discontinuity inside a blended average.
Do not force no-click visibility into a revenue number you cannot defend. Report correlation as correlation, keep attributable conversions separate, and use brand visibility trends to decide where to investigate. The purpose of the scorecard is to improve decisions, not manufacture certainty from a probabilistic system.
On your next reporting cycle, start with one high-value customer decision. Build its prompt family, collect a documented baseline, identify the most obvious evidence or accuracy gap, and correct that gap on the canonical page. Then rerun the same protocol. That gives you a repeatable visibility practice while the interfaces, models, and search journeys continue to change.
You do not need the agency with the longest service list. You need one that understands the constraint most likely to derail your growth: a difficult website, a regulated approval process, local-market competition, a narrow buyer group, or a team with little time to implement recommendations.
That changes how you should build a shortlist. Instead of beginning with agency rankings, start with your operating reality, define the evidence each candidate must provide, and make every contender answer the same questions. The result is a decision you can defend after the sales presentation is over.
Choose for the constraint that can break the engagement
“Specialized SEO” is not one service. A telecom company may need JavaScript troubleshooting, mobile-first technical work, Core Web Vitals improvements, lead generation, and a reliable compliance workflow. A pharmaceutical business may have medical, legal, and regulatory review requirements that determine what can be published. A contractor usually depends more heavily on geographically specific demand, calls, map visibility, and service-area pages. A small business may have a sound strategy but no spare team to execute it.
An agency’s industry label is therefore only a filter. A relevant client logo shows that the agency entered the market before; it does not show what the team diagnosed, changed, or measured. Even a firm featured among small-business SEO agencies still has to prove that its delivery model fits your staff, margins, geography, and sales process.
Write a short constraint brief before contacting candidates. Include:
The business event SEO should influence, such as a qualified inquiry, booked consultation, application, purchase, or sales opportunity.
The buyer and the problem that brings that person to search.
The geographic market you can actually serve.
The technical environment the agency will inherit, including the CMS, JavaScript dependencies, analytics setup, and development resources.
The people who can approve content, technical work, and regulated claims.
The capacity available for writing, subject-matter review, design, development, and sales follow-up.
The search surfaces that matter to you, including conventional results, local results, answer engines, and generative AI systems.
This brief prevents a common procurement error: buying a strategy that assumes resources you do not have. If every recommendation will wait for an unavailable developer or subject-matter expert, the agency’s theoretical sophistication will not rescue the engagement.
Set or adjust the criteria before you know which agency scores well. Otherwise, an impressive presenter can quietly redefine what “best” means during the meeting. A pharmaceutical buyer might elevate governance and compliance evidence. A contractor might place more emphasis on local execution and lead attribution. A resource-constrained business might value prioritization and implementation support more than awards.
Criterion
Telecom starting weight
Evidence to request
Technical SEO competency
20%
An anonymized audit excerpt, the affected templates, the proposed fix, implementation responsibility, and the validation method.
Industry experience and track record
15%
A relevant engagement with a similar buyer, business model, search problem, and operational constraint.
Team composition
15%
The named strategist, technical specialist, writer or editor, analyst, and day-to-day account lead who would do the work.
Leadership experience
12%
Who makes strategic decisions, when senior specialists participate, and how an escalation reaches them.
Geographic presence and reviews
10%
Evidence that the team understands the target market, plus review patterns rather than a single testimonial.
Client satisfaction and results
10%
Baseline, measurement window, intervention, business outcome, and a clear explanation of what the agency can substantiate.
Innovation and future-readiness
10%
A practical AEO or GEO workflow covering query selection, source-page improvement, entity clarity, citations, monitoring, and limitations.
Media recognition and industry awards
8%
Recognition relevant to the work you are buying, separated from paid placements and general promotional visibility.
Do not award points for a capability merely because it appears on a slide. Define what earns full, partial, or no credit. For example, “technical SEO” should not receive full credit for a generic site-audit screenshot. The candidate should be able to explain a real diagnosis, the implementation path, the dependency that made it difficult, and the evidence used to verify the result.
Future-readiness deserves the same discipline. AEO and GEO are not synonyms for publishing more AI-generated copy. Ask how the agency identifies questions worth answering, strengthens the underlying page, clarifies entities and claims, uses structured data where appropriate, and observes whether the brand appears accurately in answer systems. No agency controls whether a frontier model cites or recommends a page, so guaranteed inclusion should reduce confidence rather than increase it.
Make every proof point survive a follow-up question
A polished case study can conceal the information you need most. Traffic may have grown while qualified inquiries remained flat. A ranking increase may concern a low-value query. A chart may begin after a migration problem was already corrected. A client may also have supplied writers, developers, and public-relations support that you will not have.
Use the same evidence ladder for every claim:
Relevance: Was the client similar in buyer, geography, sales motion, platform, and operating constraint?
Baseline: What was happening before the work, and which measurement defined the problem?
Intervention: What did the agency actually change, as distinct from work performed by the client or another vendor?
Mechanism: Why was that change expected to affect discovery, evaluation, or conversion?
Verification: Which analytics, search, local, CRM, or sales records supported the claimed outcome?
Transferability: Which conditions made the result possible, and which of those conditions are absent in your business?
If a candidate cannot answer the baseline and intervention questions, you cannot tell whether its work caused the result. If it cannot answer the transferability question, you cannot tell whether the example applies to you.
For telecom, request technical and compliance evidence
A credible telecom SEO team should be able to discuss rendering, crawl paths, mobile templates, Core Web Vitals, product architecture, lead journeys, and the review of regulated or sensitive claims. Ask for an anonymized technical finding and follow it from diagnosis through implementation and validation. You are testing whether the agency can move from an audit to a shipped fix, not whether it owns an auditing tool.
For pharmaceuticals, inspect the publishing controls
When comparing pharmaceutical SEO agencies, ask who separates search recommendations from medical or legal approval, how claim-supporting material is recorded, how reviewers receive context, and what happens when an approved statement changes. A content calendar is not enough. The agency needs a workflow that preserves accuracy and approval status from briefing through publication and later revision.
For contractors, trace visibility to serviceable demand
A contractor SEO agency should explain how it handles Google Business Profile ownership, service-area relevance, location and service-page architecture, duplicate or thin pages, reviews, calls, forms, and lead quality. Ask it to distinguish increased visibility from increased demand inside the area you can serve. Traffic from the wrong location is not a business win.
For a small business, test prioritization under constraint
A small-business engagement often fails at the handoff between recommendation and implementation. Give each candidate the same hypothetical constraint: limited writing capacity, limited development help, or a narrow service area. Ask what it would do first, what it would defer, what it needs from you, and what would invalidate its initial plan. The quality of those trade-offs tells you more than the length of the proposed deliverable list.
Also ask who will write and review specialist content. A general copywriter can organize information, but your business still needs a defined subject-matter review path. The agency should identify where expert input enters the workflow, how factual changes are resolved, and who owns the final approval.
Protect access, accountability, and exit rights before signing
An SEO proposal mixes three different things: work the agency controls, work your team controls, and outcomes neither party can guarantee. Separate them in the agreement. The agency can control whether it delivers an audit, brief, page, schema recommendation, implementation, or report. It cannot guarantee a particular ranking, AI citation, lead volume, or revenue result.
Resolve these operating terms before work begins:
Account ownership: analytics, Search Console, Google Business Profile, tag management, advertising, CMS, call tracking, and reporting accounts should be created or retained in your business’s name where the platforms allow it.
Access level: give each person the permissions needed for the work, document who has administrative access, and include a revocation process for the end of the engagement.
Implementation responsibility: state whether the agency, your team, or another vendor edits templates, publishes pages, adds structured data, redirects URLs, and validates releases.
Approvals: name the person responsible for brand, factual, medical, legal, security, and technical sign-off where those controls apply.
Measurement definitions: define a qualified lead, branded versus non-branded demand, the reporting data set, attribution limitations, and how CRM outcomes will be reconciled with web analytics.
Change records: require a useful record of material content, technical, schema, and tracking changes so later performance shifts can be investigated.
AI use: document where generative tools may be used, what human review follows, and whether confidential business or customer information may enter an external model.
Exit package: specify the files, briefs, content, credentials, dashboards, change records, and unresolved recommendations you receive when the relationship ends.
Account and data ownership are not administrative trivia. If a vendor controls a critical profile, tracking number, dashboard, or analytics property, changing agencies can interrupt reporting or customer contact. Resolve ownership in writing and have appropriate legal or security reviewers examine any term that creates material exposure for your business.
Use the sales call to test how the working relationship will behave under pressure. Ask:
Which part of our constraint brief changes your usual process?
What would you investigate before recommending new content?
Show us a recommendation that required development, compliance, or subject-matter approval. How did it reach production?
Who performs each part of our work, and which responsibilities would be subcontracted?
Which result in your proposal is a deliverable, which is a forecast, and which is outside your control?
How would you connect search visibility to qualified opportunities in our sales process?
What would cause you to change the strategy?
What will we still own and be able to use if the engagement ends?
Listen for boundaries as well as confidence. A trustworthy answer names assumptions, dependencies, and uncertainty. Be cautious when a candidate guarantees rankings or AI citations, avoids naming the delivery team, presents traffic as the only business measure, recommends large content volume before understanding the market, or makes essential data available only through a proprietary dashboard you lose on exit.
Key takeaways for your shortlist
Choose around the constraint that can block results, not around the broadest service menu.
Define and weight the scorecard before meeting agencies so presentation quality cannot rewrite your criteria.
Require every result claim to identify the baseline, intervention, verification method, and conditions needed to repeat it.
Match the proof to the market: technical and compliance depth for telecom, controlled review for pharmaceuticals, serviceable local demand for contractors, and realistic prioritization for small businesses.
Treat AEO and GEO as measurable discovery work, not as a promise that an AI system will cite or recommend you.
Keep business accounts, data, implementation records, and reusable deliverables under terms that survive the agency relationship.
Before you book another sales call, finish the constraint brief and scorecard. Send both to every contender and require evidence in the same format. That small piece of procurement discipline will make the pitches comparable and expose the gaps while you can still walk away.