I recently dove deep into the fascinating world of ChatGPT Ads with insights from Adthena. It turns out, the advertising space on ChatGPT is a treasure trove of competitive information that many search teams are missing out on.
Your competitors are running stealth campaigns via ChatGPT, and the frustrating part is that it’s not immediately visible what they’re bidding on or what creative strategies they’re adopting. Unlike Google Ads, there’s no native way—yet—to get a behind-the-scenes look at this in ChatGPT.
When OpenAI launched advertising within AI-generated responses, brands jumped on board quickly. With the Ads Manager and lowered spending thresholds, this new ad channel grew rapidly. And with plans to expand to U.K. markets soon, there’s a quickly closing window for early adopters to gain a significant advantage.
From the start, we’ve been closely monitoring these developments, and what we’ve found is eye-opening.
What Does the Current ChatGPT Ads Landscape Look Like?
Our analysis spans nearly a million queries across 20 industries in five markets, telling a clear story of the current landscape.
It’s Primarily a U.S. Channel—Other Markets are Catching Up
In the U.S., ads are run on about 4.5% of queries. In contrast, during the same period, the U.K. had none. The U.S. dominates, accounting for 90% of ChatGPT ad placements in our dataset, with Canada and New Zealand also active and Australia at 1.6%.
For U.K. teams, it means while the channel isn’t live yet, U.S. competitors are already fine-tuning prompts and creative strategies, placing them at a strategic advantage when the U.K. market opens.
The Majority of Responses Contain Just One Ad
On average, ChatGPT presents only 1.06 ad items per response in the U.S., implying a single sponsored slot per query. This level of exclusivity changes the game completely compared to multi-slot Google Ads.
Industry Restrictions Still Apply
Certain sectors, like Legal and Pharma, show no ad activity due to what seems to be OpenAI’s deliberate restrictions, although this could change, providing proactive teams an edge.
Unexpected Hot Categories
Logistics, Home & Garden, and Beauty & Cosmetics are leading in ad frequency, indicating high potential for growth in these sectors.
Retail Leads in Ad Spend
Retail & Fashion accounts for a vast share of U.S. ad items, indicating robust advertiser demand, far surpassing the national average. This suggests the significant investments made by retail brands in this space.
Current Challenges in Competitive Intelligence
Without tools like Auction Insights, understanding your competitive landscape on ChatGPT is practically impossible. You’re spending budget where you can barely track competitor activity. It’s a gap that Adthena aims to close.
Achieving Full Market Visibility with Adthena
Adthena’s ChatGPT Ads Intelligence offers broader insights by monitoring a plethora of prompts daily, providing a competitive overview previously unavailable.
You can now see who bids on your prompts, track share of voice, and spot open prompts ripe for targeting before competitors do.
In a new and rapidly evolving channel, being an early mover is an opportunity that shouldn’t be missed. Try ChatGPT Ads Intelligence free for 21 days and unlock the full potential of your advertising strategy.
Beyond Just ChatGPT: Expanding Your Search Horizons
As users move towards AI-driven searches for high-intent queries, such as product recommendations, it’s essential for search practitioners to adapt. Simply put, the game is changing.
If you’re attentive to ChatGPT Ads now, you’ll be hard to budge later. Our data shows a window of opportunity open now, similar to the early days of Google Ads. Capitalize on this before it closes.
Start your free 21-day trial of Adthena’s ChatGPT Ads Intelligence today to discover what’s unfolding in the ChatGPT ad space.
An industry-focused SEO agency should offer more than a portfolio containing familiar company names. Its real value lies in understanding how a sector’s customers search, which evidence earns their trust, and what technical or geographic constraints shape the path to conversion.
Three 2026 agency reports covering solar, agriculture, and local SEO reveal a useful selection framework. They also show why a ranking should begin due diligence rather than settle the decision.
Key takeaways
Relevant client experience, review quality, and leadership expertise recur across all three agency evaluations.
Specialization should be tested at the level of search behavior, content, technical requirements, geography, and commercial outcomes.
Local SEO is a distinct operating capability, not a substitute for knowledge of a client’s industry.
Scorecard weights reveal what a ranking values, but buyers still need to examine the evidence behind each score.
The best agency is the one whose delivery model fits the organization’s actual bottleneck, whether that is authority, local visibility, branding, or technical execution.
What specialization should change in practice
The three reports share a basic premise: experience close to the client’s market matters. The solar evaluation gave notable clients 28% of its score and also considered home-services experience when an agency had less direct solar work. The agriculture evaluation assigned 25% to notable clients and emphasized leadership experience in agriculture-specific strategy. The local SEO report made demonstrated local experience its largest factor, at 25%.
Those criteria point to different kinds of relevance. Vertical expertise concerns the market itself: its audiences, terminology, buying process, content opportunities, and standards of credibility. Local expertise concerns how a business competes across places, including location pages, structured information, and visibility in map-oriented results. An agency may possess one capability without the other.
The solar report illustrates how varied agencies within one vertical can be. It described First Page Sage as using thought-leadership content, geographically targeted landing pages, and white papers for mid-market and enterprise providers. Siana Marketing was presented as combining SEO and generative engine optimization, or GEO, with knowledge of solar sales cycles. Anchour was positioned around branding for smaller companies, while XEN Solar was associated with technical SEO and HubSpot optimization. These profiles are source-reported positioning, not independently verified performance, but they demonstrate that an industry label can encompass substantially different delivery models.
What the agency scorecards measure – and omit
Report
Agency pool reviewed
Most heavily weighted evidence
Distinctive considerations
Solar SEO
31 agencies
Notable clients, 28%; leadership experience, 22%; average reviews, 22%
Year founded, 16%; company size, 12%
Agriculture SEO
81 companies
Average reviews, 25%; notable clients, 25%; leadership experience, 20%
Services, founder involvement, and media references, each 10%
Local SEO
48 firms
Local SEO experience, 25%; average reviews, 20%
Technical expertise and local-pack effectiveness, each 15%; leadership, employee tenure, and media references
The overlap is meaningful. All three reports considered client reviews and leadership experience, while the two vertical studies placed substantial weight on recognizable or relevant clients. Taken together, the reports treat market evidence, reputation, and senior expertise as complementary signals rather than interchangeable ones.
The differences are just as instructive. The solar methodology rewarded longevity and company size. The agriculture methodology considered whether the founder remained active and how often the company appeared in media. The local evaluation gave explicit weight to technical SEO, local-pack results, and median employee tenure. A buyer that values stable account teams may find tenure more informative than media visibility; a multi-location operator may care more about local-pack evidence than an agency’s founding date.
Methodological transparency also needs scrutiny. The agriculture article says it used seven factors, but the supplied methodology names six: reviews, clients, leadership, services, founder involvement, and media references. Their stated weights total 100%, yet the mismatch between the announced and enumerated factor count is a reminder to inspect the underlying rubric rather than rely only on the final order.
How to test an agency’s claimed industry expertise
Interrogate the case evidence
A logo establishes that some relationship existed; it does not explain the scope, duration, baseline, or result. Buyers can ask what the agency was responsible for, which search problems it addressed, and how outcomes were measured. Reviews deserve similar examination. The agriculture report said it consulted G2, Clutch, and Google Reviews, while the local report described a composite drawn from Google, Clutch, and other verified platforms. The solar report referred more generally to publicly available reviews and gave additional weight to solar-client feedback.
That makes review composition more important than a headline average. Relevant questions include whether comments describe SEO work, whether they come from comparable organizations, and whether they discuss communication and execution as well as satisfaction.
Distinguish leadership credentials from delivery capacity
Leadership experience appeared in every methodology, receiving 22% in solar, 20% in agriculture, and 10% in local SEO. Senior expertise can shape strategy and quality standards, but buyers also need to learn who will actually conduct research, create content, implement technical changes, and report results. The local report’s inclusion of employee tenure offers one possible signal of delivery continuity; the solar report instead used company size as an indicator of capacity and client support.
Request a diagnosis specific to the business
A credible proposal should connect tactics to an identified constraint. An authority problem may call for expert-led content. A location-discovery problem may require technically sound location architecture and local visibility work. A weak market position may require branding before publishing at scale, while an implementation backlog may favor a technically oriented partner. This diagnosis is more revealing than whether an agency repeats the vocabulary of the sector.
Match the engagement model to the actual search problem
The sources suggest that industry specialization is not a single service category. In the agriculture report, First Page Sage was described as offering SEO, GEO, advertising, and web development, with thought leadership at the center of its positioning. The report said the company was founded in 2009 and began adapting to generative AI in 2023, while also crediting it with early GEO research. Those are claims made by the source and should be assessed alongside work samples and client evidence.
The appearance of GEO in both the agriculture and solar coverage indicates that some sector-focused firms are extending their positioning beyond conventional search results. That does not remove the need for foundational SEO. A buyer can ask the agency to separate established deliverables – such as site architecture, content, and location optimization – from newer visibility initiatives, then explain how each will be measured.
Organizational fit matters as well. The solar report associated one agency with enterprise thought leadership, another with small-company branding, and another with agile technical support. A specialist can therefore be relevant to the industry but wrong for the client’s scale, internal resources, technology stack, or immediate commercial objective.
Turn selection criteria into an accountable engagement
Before contracting, the organization should translate its selection rationale into a clear operating agreement. The scope can identify the audiences and markets being pursued, the technical and content responsibilities of each party, the approval process, and the business actions that count as meaningful conversions. Reporting should distinguish completed work and search visibility from qualified commercial outcomes.
The same evidence used to select the agency can become a review standard. If leadership involvement influenced the decision, its expected role should be explicit. If local-pack effectiveness was decisive, the relevant locations and queries should be agreed upon. If industry content expertise won the work, editorial quality and access to subject-matter experts should be built into the process.
As search interfaces and agency offerings continue to evolve, the strongest partnerships will be those that define specialization through observable decisions and accountable work, rather than through category labels alone.
You are not hiring an enterprise SEO agency because your team needs more keyword ideas. You are hiring because something has become difficult to coordinate: technical changes stall, content quality varies across business units, reporting does not connect visibility to revenue, or your brand is missing from AI-generated answers.
The agency landscape becomes easier to navigate when you stop looking for a universal winner. Start with the constraint you need removed, then make each contender prove that its delivery model can work inside your organization.
Read the landscape by operating model, not ranking
A June 1, 2026 evaluation weighted leadership experience at 30%, notable clients at 25%, third-party review averages at 25%, years in business at 12%, and company size at 8%. That lens favors established vendors with recognizable accounts. It does not establish pricing, contract flexibility, technical depth, international coverage, or the quality of the people assigned to your account.
There is another limitation worth keeping visible: First Page Sage produced the ranking and placed itself first. Treat the order as a discovery aid, not an independent verdict. The more useful information is how the firms differ.
A demand team that wants organic and paid search managed as connected acquisition channels
Use the size bands as capacity signals, not quality scores. A larger company may offer more specialists and coverage, but your account can still receive a small delivery team. A smaller firm may provide better access to senior people, but it may have less room to absorb a sudden international rollout. Ask who will actually do the work.
Define the bottleneck before you build the shortlist
An enterprise SEO brief that asks for more traffic invites generic proposals. Replace it with an operating problem. Your brief should name the business outcome, the part of the search system that is failing, and the internal constraint the agency must work around.
If authority is the problem: ask how the agency will extract expertise from executives, product leaders, sales teams, or clinicians without turning every page into a slow approval project. Thought-leadership capability matters more than raw publishing volume.
If technical scale is the problem: describe the platforms, templates, faceted navigation, migrations, international sites, and release process in scope. Look for an agency that can translate crawl and indexation findings into requirements your engineers can ship.
If fragmented channels are the problem: decide which relationship must improve: SEO and paid search, search and social, brand and demand generation, or content and video. Favor the operating model built around that connection.
If AI visibility is the problem: define what you mean by success. It could include accurate brand representation, stronger coverage of customer questions, clearer entity relationships, or visibility in relevant AI answers. Do not accept a promise of guaranteed inclusion.
If market expansion is the problem: require evidence from the actual region, language, search environment, and approval structure involved. A generic global capability claim is not a substitute for local operating knowledge.
This step may remove impressive names from consideration. That is useful. A well-known full-service agency can still be the wrong choice for a technical migration, while a focused specialist can be wrong for a multinational program requiring continuous coverage across several disciplines.
Make every contender prove enterprise readiness
Client logos show that a commercial relationship existed. They do not tell you what the agency owned, whether the work resembled your problem, or whether the people responsible are still there. Ask for evidence that exposes the delivery system behind the pitch.
A named account team: request each person’s role, expected involvement, location, and relevant experience. Clarify which people are committed to delivery and which appear only during sales.
A sample diagnostic: give contenders a bounded scenario from your environment and ask how they would investigate it. You are testing prioritization and reasoning, not collecting free consulting.
Redacted working artifacts: ask to see a technical requirement, content brief, editorial workflow, measurement specification, or executive report. Polished case-study slides reveal less than the documents teams use every week.
A route from recommendation to release: have the agency explain who converts an SEO finding into an engineering ticket, who validates the implementation, and what happens when another team blocks it.
Content governance: ask how subject-matter experts, legal reviewers, brand teams, editors, and local markets participate. The answer should cover ownership and approvals, not merely writing.
Measurement ownership: require a clear distinction between activity, search visibility, qualified visits, conversions, pipeline, and revenue. Confirm who supplies each data set and how disagreements will be resolved.
AI-search methods: ask which work is distinct from established SEO and which work overlaps with technical accessibility, entity clarity, authoritative content, structured data, and off-site reputation. A credible answer should acknowledge uncertainty and avoid guaranteed placements.
Capacity under pressure: present a plausible launch, migration, or reputation issue and ask how staffing and escalation would change. The answer will tell you more than the agency’s total headcount.
References should also be problem-specific. Speak with a client whose organization resembles yours in complexity and ask what slowed the engagement, how senior access changed after the sale, and which promised capability required the most client-side support.
Use a decision scorecard that procurement cannot flatten
Procurement comparisons often make unlike services look interchangeable. Prevent that by marking each criterion as pass, concern, or fail and recording the evidence beside it. Do not average away a failure in an area that can stop the engagement.
Decision area
Question to settle
Evidence to retain
Strategic fit
Does the proposed program address the bottleneck in your brief?
Problem statement, priorities, exclusions, and expected business outcome
Technical execution
Can recommendations survive your CMS, engineering, security, and release constraints?
Sample requirements, validation process, and ownership map
Content operations
Can the agency obtain expertise and move work through your approvals?
Workflow, role definitions, briefs, and quality controls
SEO, AEO, and GEO scope
Are conventional search and AI discovery connected without vague claims?
Defined activities, measurement limits, and reporting examples
Measurement
Can the agency connect its work to outcomes your leadership recognizes?
Metric definitions, data dependencies, attribution assumptions, and reporting cadence
Team quality
Are the proposed specialists the people who will serve the account?
Named staffing plan, responsibilities, availability, and escalation path
Commercial clarity
Can you tell what is included and what triggers more cost?
Deliverables, dependencies, change process, renewal terms, and exit provisions
Treat access to the delivery team, measurement ownership, and implementation responsibility as gates. A strong brand name or attractive review average should not compensate for ambiguity in those areas. Record concerns during the pitch process; memory becomes generous once polished proposals arrive.
Key takeaways
Choose an operating model that fits your bottleneck, not the agency with the highest overall rank.
Use company size as a capacity clue, then verify the people and time assigned to your account.
Replace client-logo proof with relevant artifacts, named team members, and problem-specific references.
Define AI-search success before buying GEO or AEO services, and reject guaranteed-inclusion claims.
Make technical execution, measurement ownership, and delivery-team access non-negotiable gates.
Your next move is to write a brief around the constraint that is costing your organization the most. Send the same scenario and evidence requests to every contender. The right agency will make the work, ownership, and tradeoffs clearer before the contract is signed.
If several 2026 rankings have left you with several different best agencies, the rankings aren’t necessarily contradictory. Each one reflects a different candidate pool, industry context and definition of fit. Your job is not to accept the published order. It is to decide whether the order still holds for your business.
Use industry rankings to discover credible candidates, then re-rank those candidates around your search demand, operating constraints and commercial risk. The process below gives you a defensible way to do that without turning agency selection into a contest between sales presentations.
Rankings help you discover candidates, not declare a universal winner
An agency’s position is conditional. It depends on which firms entered the evaluation, which criteria were used, how those criteria were weighted and when the underlying information was checked. A first-place agency for a luxury fashion brand does not automatically become the best choice for a biotech platform, regional med spa or international logistics provider.
Can it protect scientific accuracy while making complex concepts discoverable to distinct audiences?
Those pool sizes show the breadth of consideration, but they are not confidence scores. Fashion does not have a more reliable winner merely because its candidate field was larger than the med-spa field. The pool may be larger because the market contains more plausible candidates, because the inclusion criteria differ or because relevant experience is defined differently.
The luxury-brand condition also narrows what the fashion evidence means. It can be highly relevant when premium positioning, controlled language and brand presentation are central to the assignment. It may be less diagnostic for a discount marketplace, an apparel manufacturer selling through distributors or a retailer whose primary problem is managing a large and frequently changing catalog.
Freshness deserves the same care. The biotech field looks ahead to 2026 but carries a May 14, 2025 update date. That does not prove any position is wrong. It does mean you should verify the agency’s current team, client mix, conflicts, service scope and technical capabilities before treating its rank as current.
Do not average positions across different industry rankings or treat them as if they came from one league table. Start with the vertical closest to your business model. If your company spans verticals, identify the harder search problem and use that as the primary filter. A biotech logistics provider, for example, may need scientific governance and supply-chain demand generation; neither label alone establishes fit.
Real industry specialization changes how the agency works
Industry logos are weak evidence on their own. Specialization becomes meaningful when it changes discovery, keyword and entity research, website architecture, content approval, measurement and reporting. Ask candidates to show how their process changes for your vertical rather than merely showing that they recognize its vocabulary.
Logistics and supply chain: test the commercial architecture
A logistics website may need to organize demand by service, geography, shipment or operational problem, customer industry and buying role. Those dimensions can overlap. Publishing a page for every possible combination creates duplication; collapsing everything into broad service pages can hide the specific expertise a buyer is trying to find.
Give the agency a representative service line and ask it to sketch the path from search query to qualified inquiry. The answer should cover page hierarchy, supporting content, internal links, proof, conversion language and how irrelevant leads will be screened out. If the response jumps immediately to a calendar of generic thought-leadership topics, the commercial model has been skipped.
Also listen for the way the team handles operational terminology. It should be able to preserve the language practitioners use while explaining the offering clearly enough for procurement, finance or leadership. Replacing precise terminology with high-volume but poorly matched phrases can increase visibility while reducing lead quality.
Fashion: test catalog mechanics and brand restraint
Fashion SEO sits at the intersection of brand presentation, product discovery, merchandising and technical catalog management. Category pages, product pages, editorial content and seasonal collections can compete with one another if their roles are not clearly defined. Changes in inventory can also leave valuable internal links pointing toward thin, unavailable or retired destinations.
Ask the agency to choose a representative category and explain what it would optimize, what it would preserve and why. For a transactional site, the response should address indexation, canonical choices, filters, internal linking, product availability and structured data alongside copy. It should also identify where search-led wording would damage the brand rather than assuming every available keyword belongs on the page.
Luxury-brand experience is most useful when your own positioning requires similar restraint. If your growth model depends on frequent promotions, marketplace visibility or a broad value-oriented catalog, ask for evidence from that operating model rather than accepting prestige logos as a substitute.
Med spas: test local intent and content governance
Med-spa discovery is often both local and treatment-specific. A candidate therefore needs to connect location information, service detail, practitioner or facility trust signals, reviews and conversion paths without manufacturing interchangeable city pages. A page that merely swaps place names is not a local strategy.
Ask the team to walk through a treatment page from query selection to publication. Who verifies medical or treatment-related statements? How are candidacy, limitations and expected outcomes described without drifting into unsupported promises? How do local pages differ when locations offer different services or have different staff? The agency does not need to make clinical decisions, but it does need a workflow that routes health-related claims to an appropriate reviewer.
Measurement should reach beyond local rankings. Define what happens after a visitor arrives: a call, consultation request, booking or another meaningful action. Then establish how your team will feed appointment quality and service-line value back into SEO decisions. Otherwise, attractive traffic reports can conceal low-value inquiries.
Biotech: test scientific review and entity consistency
Biotech content has to preserve scientific precision while serving readers with different levels of technical knowledge. Researchers, prospective partners, buyers and investors may look for different answers even when they use overlapping terminology. Treating them as one audience usually produces pages that are dense but directionless.
Ask who translates the search opportunity into a technical brief, who reviews scientific statements and how corrections propagate across the site. The agency should be able to keep platform names, indications, mechanisms, development stages and organizational relationships consistent across navigation, page copy, metadata and structured data where structured data is appropriate.
Then ask how old claims are retired. Updating one prominent page is not enough when an outdated statement remains in an executive biography, resource page, downloadable asset or schema implementation. A credible workflow includes an inventory of dependent content and a named approval path.
If an agency gives essentially the same answer for every vertical after swapping industry nouns, its specialization is surface-deep. The strongest signal is not familiarity with jargon. It is an operating model built around the consequences of getting the content, architecture or measurement wrong.
Build your own decision matrix before requesting proposals
Agencies cannot respond comparably when each one receives a different version of the assignment. Prepare a short brief before outreach. Include your priority offerings, markets, audiences, primary conversion, current platform, internal implementation resources, approval constraints and known measurement gaps. State whether you need strategy only, production, technical implementation or an accountable combination.
Send the same brief to every candidate and evaluate each response as pass, concern or fail against the same matrix. Do not turn the labels into a mechanical total. A failure involving data ownership, claim approval or an undisclosed conflict can outweigh several softer passes.
Criterion
A pass looks like
A warning looks like
Business-model fit
The team maps SEO activity to your actual offering, buyer, market and conversion path.
The strategy would work only if your business behaved like a different client shown in the pitch.
Evidence quality
Examples identify the starting problem, work performed, relevant outcome and agency’s actual scope.
Charts lack context, screenshots have no meaningful baseline, or credit is claimed for work performed by others.
Industry workflow
Research and approval steps reflect your terminology, risk level, internal experts and publishing constraints.
Specialization is supported mainly by client logos and generic claims about understanding the audience.
Technical depth
The agency connects crawling, indexation, rendering, templates, internal links and structured data to specific site problems.
A standard audit is presented as the strategy, with no explanation of who will implement or validate changes.
Content operations
Briefing, subject-matter review, editing, approval, updating and retirement all have clear owners.
The proposal promises content volume without explaining accuracy control, differentiation or maintenance.
AI discovery
The team explains how entity clarity, answerable content, supporting evidence, crawlability, internal relationships and schema fit together, while acknowledging measurement limits.
It guarantees placement in AI answers or treats GEO and AEO as labels for producing more generic copy.
Measurement
Primary conversions, diagnostic metrics, lead quality and reporting decisions are defined before work begins.
Success is reduced to traffic, impressions, keyword counts or a proprietary score that cannot be reconciled with business outcomes.
Commercial safety
Account access, content and data ownership, subcontracting, change control, cancellation and transition duties are explicit.
The agency controls essential assets, avoids documenting handoff obligations or leaves implementation costs outside an apparently complete fee.
AI visibility deserves particular scrutiny in a 2026 selection. An agency should distinguish conventional search performance from appearances in answer engines or model-generated responses. It should also explain which observations are reproducible, which depend on prompts or platforms and which cannot be attributed cleanly. A polished AI dashboard is not useful if nobody can explain what its metrics mean or what decision will change when they move.
Ask how structured data fits the plan, but do not accept schema volume as a goal. Markup should represent the visible content and the entities the page actually describes. It cannot repair vague positioning, unsupported claims, inaccessible pages or contradictory facts elsewhere on the site.
Set hard stops before presentations begin. Ranking guarantees, refusal to provide access to your own accounts, undisclosed subcontracting, publication without required review and ambiguous ownership of your domain, analytics or content all deserve resolution before a contract is signed. If a candidate will not resolve them in writing, remove it from the shortlist.
Test the agency’s thinking with a real working session
Presentation fluency can hide weak diagnosis. Give every finalist the same bounded working exercise using a real part of your site. You are not asking for a free strategy. You are testing how the team frames a problem, handles missing information and converts analysis into an implementable decision.
Select a representative service, category, treatment or platform page tied to a meaningful conversion.
Provide the same business context, technical constraints and available performance information to each finalist.
Ask the team to identify search intent, relevant entities, architectural issues, content gaps, proof requirements and conversion friction.
Require prioritization. Each recommendation should identify its expected role, dependencies, implementation owner and validation method.
Ask what the team still does not know and how it would obtain the missing information after kickoff.
A strong response distinguishes evidence from assumption. It may decline to estimate an outcome until analytics, indexation, competition or lead-quality data has been checked. That is disciplined diagnosis, not evasiveness. A weak response manufactures certainty, reaches for a familiar tactic before establishing the problem or produces a long backlog with no decision logic.
Pay attention to who attends. If senior specialists lead the sale, ask who will conduct discovery, write briefs, review technical recommendations and join reporting meetings after signature. Request the names or roles of the delivery team and clarify how substitutions are handled. Industry experience held only by an executive who disappears after the pitch will not improve day-to-day work.
Reference calls are most useful when you ask about operating behavior rather than general satisfaction. Ask what the agency actually owned, what delayed the work, how disagreements were resolved, whether senior involvement changed after the sale, how reporting affected decisions and what happened when a recommendation or published claim was wrong. A reference chosen by the agency will naturally be favorable, but precise process questions can still reveal the conditions behind the success.
Read the final contract against the proposal and your matrix. Confirm deliverable definitions, implementation responsibilities, approval timing, account access, data and content ownership, use of third parties, change control, cancellation and transition support. A vague exit clause can turn an ordinary mismatch into an expensive migration, so resolve the handoff before work starts rather than after the relationship has deteriorated.
Finally, name an internal owner. Even a capable specialist agency cannot approve scientific claims, supply merchandising decisions, verify service availability or judge lead quality without your team. The contract should make that dependency visible instead of allowing delays to become a recurring dispute about who was waiting for whom.
Key takeaways
An industry ranking is a candidate-discovery tool, not a transferable verdict about the best agency for every company in that vertical.
Verify freshness, current team composition, relevant client work, conflicts and service scope before relying on any 2026 position.
Real specialization changes architecture, content review, technical execution and measurement; industry logos alone do not establish it.
Compare candidates against the same written brief and use hard stops for ownership, approvals, access, conflicts and unsupported guarantees.
Use a real working session to test prioritization, assumptions and implementation thinking before you sign.
Your next move is concrete: open the ranking closest to your operating model, create a small candidate set, run freshness and conflict checks, and send every remaining agency the same brief. Choose the team that makes dependencies, uncertainty and commercial risk visible while showing how it will solve your particular search problem. That is a stronger basis for a 2026 decision than the number beside an agency’s name.
You can have a healthy SEO dashboard and still be nearly invisible when a buyer asks an AI assistant what to choose. The difficult part isn’t collecting another visibility score. It’s knowing whether a change reflects stronger retrieval, a different mix of prompts, or noise in the answers you sampled.
A useful measurement system starts with a repeatable prompt panel, distinguishes mentions from citations, checks whether your brand is represented accurately, and connects that evidence to business outcomes. Here is how to build one without turning a handful of AI responses into false precision.
Measure what happens inside the answer, not just after the click
Traditional search measurement follows a familiar sequence: query, ranking, impression, click, session, conversion. Generative search compresses much of that journey into an answer. A user can discover your brand, compare it with alternatives, absorb a claim about it, and make a decision without visiting your site.
That makes traffic an incomplete visibility measure. Some studies cited in current GEO coverage put traditional-result clicks at only 8% when AI-generated summaries are present. Treat that figure as a warning about measurement gaps, not as a universal click-through benchmark for your site. The practical point is that an off-site answer can influence demand even when analytics records no session.
Measure AI search visibility across four layers. Presence tells you whether the brand appears. Use tells you whether an owned page is retrieved or cited. Representation tells you whether the answer describes the brand accurately and in the right context. Impact tells you whether that exposure is associated with qualified visits, branded demand, leads, sales, or another business outcome.
These layers prevent a common reporting error. A brand mention is not automatically an owned-content citation. A citation is not proof that the answer framed the brand correctly. Visibility is not proof of commercial influence. Each is useful, but each answers a different question.
Key takeaways
Use a stable set of prompts so one reporting period can be compared with another.
Keep mentions, citations, observable retrieval, entity accuracy, sentiment, and conversions as separate measures.
Report results by platform, topic, intent, and prompt cohort before calculating an overall score.
Save the underlying answer and its citations. A percentage without evidence cannot be audited.
Use visibility metrics to choose an action, then judge that action by the specific metric it was intended to change.
Build a prompt panel you can rerun without moving the goalposts
Your prompt panel is the measurement instrument. If the prompts change whenever a campaign changes, the resulting trend line cannot tell you whether visibility improved or the test simply became easier.
Start with topics and decisions that matter
List the topics your brand should credibly be associated with, then map the questions a real buyer asks while learning, solving, comparing, choosing, and validating. This creates a panel that covers informational discovery as well as decision-stage visibility.
Learn: What is the category, process, or concept?
Solve: How should someone handle a defined problem or constraint?
Compare: What are the meaningful differences between available approaches?
Choose: Which options fit a particular use case, audience, budget, or requirement?
Validate: Is a named brand suitable, credible, compatible, or known for the relevant capability?
Include branded and unbranded prompts, but don’t blend their results. An unbranded prompt tests discovery and competitive consideration. A branded prompt tests entity recognition, factual accuracy, and reputation. A dashboard that combines them can look strong simply because the model answers direct questions about a brand that the user already named.
Apply audience, industry, location, or product qualifiers only when they change the decision. Keep them in dedicated cohorts. Otherwise, an increasingly narrow prompt may manufacture visibility that does not exist for the broader market question.
Create a prompt registry before collecting answers
Give every prompt a permanent record. At minimum, store its ID, exact wording, topic, intent, audience qualifier, branded or unbranded status, platform and mode, relevant competitor set, target page, and the brand facts you expect an accurate answer to preserve.
Freeze the wording used for your baseline. If you improve a prompt later, create a new version instead of overwriting the old one. Keep retired prompts in the registry so historical rates retain their original denominator. This is less convenient than editing a shared list in place, but it prevents an invisible change in the test from masquerading as an improvement in performance.
Use a consistent collection protocol
Run the exact registered prompt in the intended platform and mode, such as an answer with web search enabled rather than a model-only response.
Record the platform, mode, timestamp, prompt version, full response, visible citations, cited URLs, and any named competitors.
Score the answer with a written rubric. Preserve the raw response so another reviewer can check the decision.
Repeat the panel on a fixed cadence. If resources permit, run prompts more than once so a single response is not mistaken for a stable pattern.
Log failed captures, blocked responses, and unavailable features separately. Do not score a technical failure as brand absence.
Keep platform results separate. Google AI Overviews, ChatGPT search, and other answer systems are different surfaces with different retrieval and citation behavior. You can create a portfolio view later, but first calculate each platform’s rate against its own eligible observations.
If you do publish an aggregate, state its weighting. An unweighted average gives every prompt-platform pair the same influence. A business-weighted score gives priority cohorts more influence. Neither is inherently correct; an unexplained blend is the problem.
Use a metric stack instead of one opaque visibility score
Eligible answers containing a qualifying brand mention or traceable use of owned content, divided by all eligible answers in the cohort.
Does the brand enter the answer at all?
AI citation frequency
Eligible answers containing a visible citation connected to the brand, divided by all eligible answers. Report any-brand citation and owned-domain citation separately.
Is the answer visibly supported by material associated with the brand, and does it cite the brand’s own site?
Share of model voice
The brand’s unique inclusions divided by unique inclusions for the entire predefined competitor set. Count a brand once per answer so repetition does not inflate share.
How much of the observable category conversation does the brand occupy?
Entity recognition accuracy
Brand-discussing answers that preserve the required facts divided by all answers that discuss the brand.
Does the system understand who the brand is, what it offers, and how its entities relate?
Sentiment and framing
Counts of favorable, neutral, critical, or mixed descriptions, paired with issue codes and the exact claim being evaluated.
How is the brand characterized before the user reaches its site?
Prompt coverage
Priority prompt cells with at least one qualifying inclusion divided by all eligible priority prompt cells.
Across how much of the intended buyer journey is the brand visible?
Observable retrieval success
Runs in which a relevant owned page is visibly retrieved or cited, divided by runs where that page is an eligible answer source.
Can the system access and use the content you expected it to use?
Conversion influence
Qualified visits, conversions, lead quality, revenue, branded demand, or other outcomes associated with AI referrals and visibility changes.
Is AI visibility connected to business value?
The denominator matters as much as the numerator. Show both on every metric card. A 50% inclusion rate based on two eligible answers carries very different weight from the same rate across a broad, repeated panel.
Keep citation frequency and retrieval success distinct. A brand can be mentioned because a third-party page was retrieved. An owned page can be cited without the brand becoming a recommended option. A model may also name the brand without exposing any source. Consumer-facing outputs rarely reveal every internal retrieval step, so call the measure observable retrieval rather than claiming access to hidden model behavior.
Share of model voice also needs a locked competitor set. Adding weak competitors lowers everyone’s apparent share; removing a dominant competitor raises it. Version the set just as you version prompts, and show absolute inclusion alongside share. If absolute visibility holds steady while share falls, competitors may be gaining rather than your brand disappearing.
For entity accuracy, write the answer key before scoring responses. Include only facts the brand can substantiate, such as its official name, category, product relationships, supported markets, or current positioning. Record each error type separately. A single accuracy percentage will not tell your content team whether the problem is an outdated name, a category mismatch, a confused product relationship, or a claim that is too broad.
Sentiment needs the same discipline. A neutral answer that omits the brand’s relevant capability is different from a critical answer containing a factual error. Save the exact sentence, its context, the issue code, and the affected prompt. Automated labels can help sort a large collection, but consequential or ambiguous cases still need human review.
Read metric combinations as a diagnostic system
No metric tells you what to change by itself. The useful signal comes from combinations. Start with the smallest cohort where the problem appears, then diagnose the layer most likely to be responsible.
Low inclusion plus low observable retrieval
Begin with access and extractability. Check whether the intended page can be crawled, whether the primary answer is available in parseable text, whether important information is current, and whether structured data accurately describes the visible content and entity relationships. Crawlability, schema use, freshness, and parsing quality all belong in a retrieval-success investigation.
Do not add schema merely to produce more markup. Structured data can clarify supported facts; it cannot make a thin, contradictory, or inaccessible page authoritative. Validate the markup, align it with what users can see, and retest the affected prompt cohort after the page can be revisited.
Inclusion without owned citations
The system recognizes the category connection, but your site is not supplying the visible evidence. Inspect which domains are cited instead and what those pages make easy to extract. Then improve the relevant owned page with a direct answer, clear definitions, explicit comparison dimensions, supported claims, and enough surrounding context for a passage to stand on its own.
Do not treat matching wording as proof that the model used your page. Unless the interface exposes a citation or retrieval record, hidden sourcing remains unknown. Score what you can observe and use citation gains as the validation target for this change.
Strong visibility with weak entity accuracy
This is a representation problem, not an awareness problem. Compare the wrong claim with the corresponding signals on your site, structured data, product pages, and corroborating profiles. Standardize names and relationships, remove obsolete descriptions, and make the canonical explanation explicit. Retest the prompts that produced the error rather than waiting for the global score to move.
Informational coverage without decision-stage visibility
The brand may be recognized as an educator but absent from the consideration set. Examine compare, choose, and validate prompts. If the cited pages answer selection questions that your pages avoid, create or improve content around fit, limitations, use cases, evaluation criteria, and meaningful alternatives. The goal is not to declare yourself the best. It is to supply the facts an answer system needs to explain when the offering is or is not a fit.
Visibility gains without measurable business impact
First check intent. More citations on broad educational prompts may be valuable without creating immediate demand. Next check whether the cited or visited page offers a sensible next step for that query. Then inspect referral classification, landing-page engagement, conversion quality, direct traffic, and branded search movement.
Do not force a revenue claim from a coincident trend. Off-site AI interactions are often not connected to an identifiable user journey. Call the result influence unless you have instrumentation that supports stronger attribution.
Change one measurement layer at a time
Turn each diagnosis into a recorded experiment. State the affected cohort, observed gap, proposed change, page or entity being changed, metric expected to move, business guardrail, and next review point. If you rewrite the prompts, replace the target pages, and change the scoring rubric together, you will not know which change produced the new result.
Keep a control cohort of unchanged prompts when practical. It gives you context when visibility moves across the platform rather than only on the pages you changed.
Report evidence, decisions, and business influence in one workflow
A dashboard should shorten the distance between an observed gap and the person who can address it. Clutch, for example, places Conductor-powered visibility analysis inside its AI Visibility Dashboard. The useful principle is workflow integration: a report creates more value when operators can move from the trend to the affected prompt, answer, citation, topic, and page.
Give each audience the view it needs
Leadership view: priority-topic inclusion, share of model voice, entity accuracy, major reputation issues, qualified AI traffic, and conversion influence.
Evidence view: exact prompt, full response, visible links, scoring decision, timestamp, reviewer, and prompt version.
Every summary card should show the current value, comparison baseline, numerator, denominator, included cohort, and last collection date. Avoid a global visibility score that cannot be traced to those components. It may look tidy, but it cannot tell a content, technical SEO, brand, or analytics team what to do next.
Keep the collection cadence and the decision cadence separate
Collect on a consistent schedule that your team can sustain. Review urgent factual errors when they appear, but make strategic decisions only after you have enough comparable observations to distinguish a pattern from one answer. Annotate changes to prompts, pages, structured data, competitor sets, platform modes, and scoring rules directly on the timeline.
When a platform introduces a materially different mode or answer experience, create a new cohort. Do not splice it into the old series as if the measurement environment stayed constant.
Triangulate AI visibility with analytics and search data
No single product captures the complete path. Combine controlled prompt testing with analytics, server or referral evidence where available, Search Console, traditional SEO tools, technical audits, and business data. This mixed approach reflects the reality that GEO measurement currently requires multiple tools and methods.
In GA4, isolate known AI-platform referrals and compare their landing pages, engagement, conversion rate, conversion value, and lead quality with relevant baselines. Keep the referral rules documented because platforms and referrer behavior can change. Review direct and branded-search demand alongside those sessions, but present the relationship as supporting evidence rather than proof that every change came from AI exposure.
Search Console still helps you see traditional query demand, page performance, and technical conditions around the topics in your prompt panel. It will not expose every AI interaction, but it can reveal whether a page has a broader indexing, relevance, or demand problem that also limits its usefulness to generative systems.
Evaluate tools by the decisions they support
Before buying an AI visibility platform, ask whether it supports the exact environments you need to measure and whether you can audit its results. A useful evaluation checklist includes:
Named platforms and modes rather than a generic claim of model coverage.
Exact prompt storage, prompt versioning, cohort management, and repeatable scheduling.
Preservation or export of full responses, citations, cited URLs, timestamps, and scoring evidence.
Transparent definitions and denominators for inclusion, citations, share of voice, sentiment, and coverage.
A configurable competitor set and the ability to retain historical versions of that set.
Segmentation by topic, intent, platform, geography where relevant, brand, competitor, and target page.
Human review, issue coding, annotations, ownership, and an audit trail for score changes.
Connections to analytics and business outcomes rather than visibility reporting alone.
Do not compare vendor scores as though they were interchangeable. One may count every mention, another only cited mentions, and another may use a proprietary weighted index. Compare the underlying prompts, observations, scoring rules, and denominators before comparing the headline numbers.
Start with one commercially important topic. Freeze its prompts, capture a baseline, and identify the largest localized gap: presence, citation, retrieval, accuracy, competitive share, or impact. Assign one change to that gap and name the metric that should respond. When the dashboard can tell your team what to inspect next, AI search visibility stops being a vanity score and becomes an operating system for better decisions.
If you are deciding whether ChatGPT advertising deserves budget, do not start by asking whether it resembles paid search. Start with the moment the ad enters: the user has already described a need, added constraints, and moved partway toward a decision.
That is enough activity to reveal recurring creative conventions. It is not enough to establish a universal cost per acquisition, return on ad spend, or incrementality benchmark. The observations come from a vendor-tracked index during a trial, span materially different verticals, and do not provide one standardized performance baseline for every advertiser.
Use the data to answer questions such as how much copy the format can carry, which information tends to appear first, and how closely creative reflects the conversation. Do not use it to forecast your return before you have campaign-level evidence from your own offer, audience, and destination.
Before assigning meaningful budget, make sure your pilot can answer a defined question:
Can you identify a narrow group of commercial topics where the user is likely to be comparing options or preparing to act?
Do you have a specific, verifiable benefit that can be understood without several lines of explanation?
Does the destination continue the exact promise made in the ad?
Can you separate ChatGPT placements from your other paid traffic when evaluating outcomes?
Have you defined what would justify expanding, revising, or stopping the test before spend begins?
Rollout status is time-sensitive, so confirm actual inventory and account eligibility before committing budget or launch dates. A projected geographic expansion is not the same thing as inventory you can buy.
Write an answer fragment, not a compressed search ad
A traditional search ad often has several components competing for attention: multiple headlines, descriptions, sitelinks, extensions, and other assets. The early ChatGPT format is more restrained. That makes every word carry more of the decision.
The strongest working model is an answer fragment. It should make sense beside the assistant’s response, acknowledge the user’s decision criteria, and introduce a next step without pretending to be the neutral answer.
Lead with the decision-driving benefit. Do not spend the available space on a generic slogan.
Headline opening
Most begin with the brand name
Test a Brand: Benefit construction when recognition and accountability matter.
Body
About 19 words, commonly split into two sentences
Use the first sentence for proof and the second for a low-friction action.
Relevance
Stronger creative mirrors the user’s context
Reflect the category, constraint, or desired outcome instead of repeating a loose keyword.
Offer detail
Dollar signs, rates, and concrete figures were associated with stronger conversion performance
Prioritize a specificity test, but treat the pattern as a hypothesis to validate in your own campaign.
Build each variation from three prompt components
When a user asks for accounting software for a small team, for example, accounting software is only the category. Small team is the constraint. The unstated decision criterion might be fast setup, predictable cost, or limited administrative work. Creative that reflects only the category will feel generic even if it contains the right keyword.
Extract the category: what kind of product, service, or action does the user want?
Extract the constraint: what price, use case, location, feature, risk, or timing narrows the choice?
Choose one decision criterion your offer can substantiate.
Write the headline as Brand: Verified Benefit.
Use the body for one proof point and one proportionate call to action.
Remove any claim that the landing page cannot immediately confirm.
A useful template is: Brand: [specific outcome]. [Proof tied to the user’s constraint]. [Simple next action]. The brackets are not an invitation to stuff several benefits into one placement. Choose one reason to continue.
Specificity needs controls. If you advertise a price, rate, discount, delivery window, or availability claim, it must be current, approved, and visible at the destination. A concrete figure can improve clarity, but an outdated figure creates both conversion friction and potential compliance exposure. When the value changes frequently, build a review process before testing it in ad copy.
Test in an order that explains the result
Changing the headline, proof, call to action, and landing page at the same time may produce a winner, but it will not tell you why it won. Start with the variables most closely tied to conversational relevance:
Specific offer versus general benefit.
Query-matched benefit versus broad category language.
Quantified proof versus qualitative proof.
Low-commitment call to action versus immediate purchase or signup language.
General landing page versus a page that continues the same constraint and benefit.
Hold the other elements steady during each comparison. The point is not merely to improve the ad. It is to learn which part of the conversation your audience needs resolved before moving forward.
Measure prompt coverage and response duplication before calling it reach
Clicks and conversions still matter, but they do not tell you whether your brand is present across the conversations that matter. Conversational inventory needs an observation layer organized around topics, prompts, and individual responses.
That becomes especially important because one brand has been observed appearing twice within the same ChatGPT response. This double-parked behavior creates more placements, but it does not automatically create more unique reach. Counting each placement as a separate conversation would overstate coverage.
For every observed placement, record the topic, prompt or prompt class, response identifier, timestamp, position, advertiser, headline, body, and destination. Add post-click outcomes when your analytics can connect them. That record supports several more useful measurements:
Observed prompt coverage: the portion of your monitored commercial prompts in which your brand appeared.
Observed response presence: responses containing your brand divided by eligible responses you actually monitored.
Duplication rate: brand-present responses containing more than one placement for the same brand.
Competitor overlap: responses where your brand and a named competitor appeared together.
Creative-context match: whether the ad reflects the category, constraint, and decision criterion in the prompt.
Post-click continuity: whether the destination preserves the offer and language that earned the click.
Business outcome: qualified lead, sale, signup, or another result defined before the pilot.
Call these observed rates, not platform-wide impression share. A monitoring sample cannot tell you the total number of eligible conversations unless the platform provides that denominator. This naming discipline prevents a directional visibility metric from turning into a false market-share claim.
Review duplication separately from performance. Two appearances might reinforce recall, or they might add no incremental value. The placement pattern alone cannot settle that question. Compare duplicated and single-placement responses only when you have enough campaign data to evaluate their downstream outcomes.
Your landing-page review should be just as specific. Check whether the advertised benefit appears without searching, whether the price or rate matches, whether the next action is obvious, and whether the page answers the constraint expressed in the originating conversation. A relevant ad that lands on a general homepage throws away the context that made the placement useful.
Coordinate ChatGPT ads with AEO and GEO without merging the KPIs
Paid presence and organic AI visibility can occur in the same conversational environment, but they are not the same achievement. A sponsored placement buys labeled exposure. An organic citation, recommendation, or brand mention depends on how the system constructs its answer. Early placement observations do not establish that buying ads improves organic answer inclusion.
Keep the two lanes separate in reporting. If you combine them into one AI visibility number, you will not know whether a change came from media spend, content improvements, brand demand, or answer-engine behavior.
Use one shared topic map. Organize paid monitoring and organic visibility work around the same commercial questions, constraints, entities, and decision criteria.
Give paid media its own outcomes. Track observed presence, duplication, clicks, qualified actions, and campaign economics.
Give AEO and GEO their own outcomes. Track whether the brand is mentioned, cited, represented accurately, and connected to the intended category across monitored answers.
Align the factual layer. Prices, rates, features, availability, and offer terms should agree across ad copy, visible page content, and applicable structured data.
Investigate cross-channel clues. A commercial prompt with competitor ads but weak organic answers may expose a content opportunity. Strong organic visibility with no paid presence may identify a conversation worth testing, but neither observation guarantees demand or return.
JSON-LD can clarify entities, products, offers, and other machine-readable facts when it accurately represents visible content. It does not purchase inventory, guarantee inclusion in an AI response, or repair a weak offer. Use structured data to reduce ambiguity, then use advertising to test whether a clear commercial promise earns action.
This coordinated model also gives you a cleaner competitive view. You can distinguish a competitor that is buying exposure from one that is repeatedly earning non-sponsored visibility. The response is different: one may call for a media test, while the other may require better content, stronger entity signals, clearer proof, or a more competitive offer.
Key takeaways for your first ChatGPT ad pilot
Treat early placement data as evidence about format and creative conventions, not as a guaranteed ROI benchmark.
Write for a user who has already supplied context: lead with the brand, one verified benefit, one proof point, and one next action.
Use the observed 30-character headline and 19-word body patterns as editing discipline, not as assumed platform limits.
Test concrete figures before vague claims when your offer supports them, but keep every price, rate, and term synchronized with the destination.
Measure prompts and unique responses as well as placements, because two appearances in one response do not equal two reached conversations.
Coordinate paid, AEO, GEO, landing-page content, and structured data around one topic map while reporting paid and organic outcomes separately.
Your next move is a narrow pilot, not a platform-wide commitment. Choose a small set of high-intent topics, document the user’s constraints, create controlled variations, and establish an organic visibility baseline before ads run. You will then be able to decide from your own evidence whether conversational advertising adds qualified demand, merely adds placements, or reveals a larger content opportunity.
You need a test that shows where visibility breaks: whether an AI system retrieves your brand, mentions it, cites it, explains it correctly, places it on a shortlist, or recommends it. The framework below turns those separate outcomes into a prompt panel, a repeatable scorecard, and an experiment you can act on.
Start with the decision, not a visibility score
AI visibility is not a single event. Your brand can be cited without being recommended, mentioned without receiving a citation, or described accurately but placed behind competitors. Treating all three situations as visible conceals the problem you need to fix.
Separate each answer into five measurement states:
Retrieval: the AI answer appears and has an opportunity to include your brand.
Inclusion: your brand, product, or page is mentioned.
Attribution: an owned URL or a third-party page about your brand is cited.
Positioning: the answer gives your brand a particular order, category, use case, or authority level.
Recommendation: the answer actively includes your brand in the decision set for the intended user.
Denominators matter, especially on search surfaces that do not generate an AI answer for every query. A missing AI Overview is not the same result as an AI Overview that appears but omits your brand. Track both:
AI answer trigger rate = attempts that produced an AI answer divided by all attempts.
Among-answer mention rate = rendered AI answers mentioning the brand divided by all rendered AI answers.
End-to-end mention rate = attempts mentioning the brand divided by all attempts, including attempts without an AI answer.
Do not compress these outcomes into one proprietary visibility score. A composite can rise because branded prompts improved while the unbranded prompts that create new demand deteriorated. Show the component rates and their numerators so a change remains interpretable.
Build a prompt panel that can be rerun
A useful prompt panel is a measurement instrument, not a loose keyword list. Every prompt needs a defined intent, an eligible engine or surface, and a reason for being in the panel.
Branded identity prompts test whether the system knows what the brand is, who it serves, and how it differs.
Category prompts remove the brand name and test discovery for the problem or product class.
Comparison prompts test alternatives, versus questions, and the attributes used to separate competitors.
Decision prompts add a buyer constraint, such as audience, use case, risk, or required capability, and test whether the brand is recommended.
Validation prompts test reputation, limitations, suitability, or factual claims that a buyer may check before acting.
Keep a stable core panel for trend reporting and a separate exploratory panel for new questions. If you rewrite, remove, or add core prompts, create a new panel version. Do not splice the results into the previous trend line as though the test stayed constant.
Run each target engine as its own surface. A first-month fictional-brand test covering 825 prompts and 15,835 answers found materially different behavior across ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, and Gemini. Google AI Mode was comparatively stable for branded questions, Perplexity surfaced new material quickly, ChatGPT recognition strengthened during the month, and Gemini produced substantial citation gaps. Because the brand was artificial and the observation window was short, those results are evidence that engines differ, not a permanent ranking of the engines.
A permanent prompt ID, prompt family, and panel version.
The exact prompt text without silent edits.
The engine and specific surface, such as Google AI Mode or Google AI Overviews.
The date, run number, locale, and any account or session conditions you can keep consistent.
The complete answer, ordered brand mentions, cited URLs, and first cited URL.
Whether your brand was recommended, how it was framed, and whether the description was accurate.
Use a fresh conversation for each conversational-engine run so earlier messages do not become an uncontrolled input. Run repetitions in the same measurement window, then rerun the complete batch on a fixed cadence. Weekly measurement can suit an active intervention; a monthly cadence may be enough for an established baseline. Consistency matters more than choosing an arbitrary universal interval.
Evaluate tracking tools against this test design. Familiar SEO integration can still leave you with narrow LLM coverage and no optimization workflow. Before committing to a platform, confirm that it covers your target surfaces, retains raw answers and cited URLs, distinguishes mentions from citations, preserves prompt versions, records repeated runs, and exports answer-level rows. A polished summary dashboard cannot compensate for missing evidence.
Score each answer without losing its context
Create one row per answer, not one row per prompt. Aggregating three runs before storing them destroys the variation you are trying to measure.
Inclusion: record brand absent or present. Calculate mention rate separately for branded, category, comparison, decision, and validation prompts.
Attribution: distinguish an owned-domain citation from a citation to an independent page about the brand. Then record whether the owned page was the first or main cited source. A third-party citation can improve brand exposure without giving your site attribution.
Order and recommendation: record the brand’s position among listed options and whether the language explicitly recommends it. Do not treat a neutral appearance in a list as a recommendation.
Explanation depth: apply a small internal rubric consistently. Score 0 for absent, 1 for a name-only list appearance, 2 for a short explanation containing one defined claim, and 3 for a substantive explanation covering the audience, use case, or reason to choose. This is an operational rubric, not an industry benchmark.
Framing and accuracy: label the tone as positive, neutral, cautionary, or negative. Record authority labels such as leader, challenger, or niche option only when the answer actually uses that framing. Mark factual descriptions as correct, incomplete, or incorrect in a separate field.
Stability: with three runs, report whether the brand appeared in none, one, two, or all three. Keep that distribution visible beside the average rate.
Accuracy is the non-negotiable guardrail. A confidently worded but false recommendation is not a visibility win. Keep inaccurate claims in the visibility totals so you do not hide the problem, but flag them separately and prioritize correction over reach.
Report each metric with its numerator and denominator. A percentage without the number of eligible answers conceals small samples, missing AI-answer triggers, and changes to the prompt mix. Break results down by engine, prompt family, branded versus unbranded intent, and run consistency before looking at an overall total.
Turn signal patterns into controlled content changes
Diagnose the gap before editing
The scorecard should point to a failure mode. It should not merely tell you that visibility is low.
Observed pattern
Likely reading
Next test
Strong branded mentions, weak category mentions
The entity is recognized, but its association with the wider problem or category is weak.
Test a page that connects the brand clearly to the category, audience, and use cases.
Frequent mentions, few owned citations
The brand is known, but the main site is not being selected as evidence.
Consolidate definitive facts on an owned page and inspect which independent URLs are being cited instead.
Citations without recommendations
Your material is useful as evidence, but the brand’s decision position is unclear.
Test explicit audience fit, differentiators, selection criteria, and honest limitations.
Name-only appearances
The system has too little usable information for a deeper explanation.
Test one comprehensive page that answers what the brand is, who uses it, and how to choose it.
Top placement in only one run
The apparent lead may be output volatility rather than a stable gain.
Repeat the batch and report the run distribution instead of publishing the best screenshot.
Visibility on one engine only
The gain is surface-specific.
Inspect that engine’s citations and distribution path; do not describe the result as universal AI visibility.
Positive but inaccurate descriptions
Repeated claims are shaping the narrative without adequate verification.
Correct the canonical brand information and monitor the exact false claim across owned and independent pages.
For your site, give each supporting page a distinct job tied to a real prompt or decision. Measure which URL is cited. Remove or correct pages that merely repeat claims, especially when repetition could amplify an error.
Test one explanation at a time
Most AI visibility work is a structured before-and-after test, not a true randomized A/B test. Retrieval systems change, answers vary, and you do not control when every engine discovers a revision. You can still make the evidence more useful:
Write a falsifiable hypothesis. For example, clarifying audience and category on the canonical brand page should increase explanation depth on branded identity prompts.
Capture a triplicate baseline batch. If the three runs conflict sharply, repeat the baseline before changing the site.
Make the smallest coherent intervention. Update the entity page, publish a comparison resource, or improve a specific claim set, but do not combine a redesign, a large publishing sprint, and a distribution campaign if you want to know what helped.
Record the changed URLs, publication date, affected claims, internal links, and prompt families expected to move.
Use discovery as the gate instead of assuming every engine follows the same calendar. Begin interpreting the post-change period only after the new or revised material appears in citations or is otherwise demonstrably available to the surface being tested.
Rerun the same panel, engine mix, session setup, and scoring rules. Keep newly discovered prompts in the exploratory panel until the current test ends.
Compare the target prompts with unaffected prompt families and competitor patterns. If every brand moves in the same direction, engine drift is a stronger explanation than your page change.
Repeat the result in another scheduled window. Call a one-engine or one-run gain directional, not conclusive.
Describe a before-and-after movement as associated with the intervention unless you have stronger controls. That language is not timidity; it is an accurate reflection of a system whose retrieval, citations, and generated wording can all change outside your test.
Keep AI response metrics beside traditional SEO and business outcomes. Citations do not guarantee visits, and visits do not prove that the answer influenced a decision. Some ChatGPT journeys continue on Google as users verify what they were told, so direct AI referrals may miss part of the path. Compare AI visibility with organic landing-page activity, branded demand, qualified visits, and conversions, but do not assign causation merely because two lines moved together.
Key takeaways
Choose the decision you need to make before choosing a visibility metric.
Separate AI-answer triggers, mentions, owned citations, independent citations, mention order, recommendations, explanation depth, framing, accuracy, and stability.
Keep branded, category, comparison, decision, and validation prompts in separate cohorts.
Measure each engine and surface independently, and run the exact prompt three times as a practical volatility check.
Store one row per answer with the raw response and cited URLs. Do not rely on a composite score or a selected screenshot.
Diagnose the missing stage, change one coherent content element, wait for discovery, and rerun the versioned panel.
Track AI visibility beside rankings, traffic, and conversions without treating any one of them as a substitute for the others.
Your first useful measurement system can be a spreadsheet: a stable core prompt panel, three runs per prompt, one row per answer, and one intervention tied to one failure mode. Automate it after the process can explain why a number moved. That is the point at which AI visibility becomes an operating metric instead of a collection of interesting screenshots.
Heidi Sturrock, a seasoned paid search consultant, shared her insights with me in a recent episode of PPC Live The Podcast. With over two decades of industry experience, Heidi discussed a memorable campaign blunder that surprisingly turned into a strategic win, as well as her experiences with AI Max across numerous accounts.
Heidi’s career story includes a significant misstep she made while running a competitor campaign using broad match without negative keywords. Launched on a Friday, this led to a weekend surge of calls from irate customers of a competitor, a situation both alarming and chaotic for the client’s call center.
Unexpectedly, her client saw potential in this turmoil. Instead of dwelling on the mistake, they chose to transform these calls into sales opportunities by offering a discounted first month to switchers. By dividing the campaign for better focus, they turned a problem into a pathway for growth.
From this experience, I learned two valuable lessons: never initiate major campaigns on a Friday and ensure all stakeholders are involved in client meetings. Having both the business owner and the sales leader aware of the situation allowed for quick, effective problem-solving.
When facing a mistake, I’ve realized the importance of halting the issue swiftly, taking responsibility, and presenting a clear plan for resolution. Clients value honesty, and this approach can reinforce trust even in difficult times.
Common failures in account management often involve misaligned attribution windows and undue focus on secondary KPIs. It’s crucial to align metrics with the primary goals, ensuring that higher CPCs are understood within the broader context of achieving ROAS targets.
Regarding AI tools, Heidi’s exploration of AI Max across various accounts delivered mixed results. Success often hinged on the availability of comprehensive historical data and well-defined goals. Her advice is to experiment gradually and prepare upcoming guidelines on her blog.
For those in the industry, embracing technological changes, especially in AI, is essential. Mastering these tools can propel us ahead as marketers.
Stay connected with Heidi on LinkedIn or visit HeidiSturrock.com for her expert guides, including tips on crafting effective ad copy. Also, catch her live at SMX Advanced in Boston this June, where she’ll participate in an engaging expert panel discussion.
Your organic traffic moved during March 2026, and the tempting response is to rewrite every page that lost clicks. Resist that impulse. Google’s first core update of 2026 arrived close to separate spam and Discover changes, so a simple month-over-month chart cannot tell you what happened.
Your first job is attribution: isolate the affected search surface, query set, page group, and shared weakness. Then change only what the evidence supports. This protects strong pages from panic edits and gives you a credible way to judge whether the work helps.
Key takeaways
Google said the March 2026 core update could take up to two weeks to roll out. Treat movement inside that window as provisional rather than a final verdict.
Do not attribute every March change to the core update. A March spam update, a February Discover update, your own site releases, tracking problems, and changing demand can produce different patterns.
Diagnose at the level of search surface, query cluster, page group, and template. A sitewide traffic total hides the pattern you need to fix.
Audit whether losing pages satisfy the searcher’s task more clearly and completely than competing results. Cosmetic rewrites and extra keywords are not a recovery strategy.
There is no universal or immediate repair. Improvements can appear gradually, including after later core updates, so preserve evidence and measure each coherent batch of changes.
Treat March as an attribution problem, not a verdict
A core update is a broad reassessment of how Google’s systems surface useful results across many sites and searches. Google characterized this release as a regular update focused on relevant and satisfying content. A ranking loss does not, by itself, prove that a page violated a rule, received a manual penalty, or needs to be deleted.
The surrounding timing matters. The core update followed a March 2026 spam update and a February 2026 Discover update. Those events are not interchangeable. A change confined to Discover should not automatically become a core-update content project. A Web Search decline should not be blamed on Discover. A sitewide drop across every acquisition channel may point to measurement, demand, or a site release rather than Google rankings.
Build a timeline before opening your content editor. Mark the core, spam, and Discover milestones; Google’s confirmed rollout completion; and every meaningful change your team shipped nearby. Include migrations, URL changes, template releases, internal-link changes, tracking updates, large content batches, and availability or pricing changes that could affect demand. The purpose is not to choose a convenient explanation. It is to keep plausible causes separate long enough to test them.
Use comparable reporting periods on either side of the event. Match the length and weekday mix, and note seasonal or campaign-driven demand. If a comparison period overlaps the rollout, label the result provisional. For historical analysis, anchor the post-update period after Google’s confirmed completion marker rather than assuming the announcement date was the moment every ranking changed.
Build a page-and-query evidence map
Start with Google Search Console and your analytics platform, but do not begin with total organic sessions. First separate Web Search from Discover and other channels. Within Web Search, compare impressions, clicks, click-through rate, and average position by query and page. Within Discover, examine the available page-level reporting on its own terms rather than forcing it into a Web Search query analysis.
Group affected pages by the reason they exist: topic, search intent, content format, audience, template, authoring process, or business line. The useful unit is rarely one isolated URL. If a collection of similar pages declined together while the rest of the site held steady, the shared pattern is more informative than the site’s average.
Save an untouched baseline export before editing anything. Preserve page, query, device, country, impressions, clicks, position, and conversion data where available.
Separate losses in visibility from losses in response. Falling impressions or positions indicate a search-visibility problem. Stable impressions with fewer clicks point toward result presentation, changed result features, or user choice. Stable search clicks with weaker conversions point downstream to the landing experience, offer, tracking, or audience fit.
Rank page groups by material impact, then look for repeated behavior. A cluster losing across many related queries deserves attention before a single volatile term.
Record winners as well as losers. Unchanged and improving pages show which formats, topics, and approaches Google continued to surface on your own domain.
Inspect the current results for the affected queries. Compare the task served, answer depth, format, specificity, freshness needs, and intended audience. Do not copy the winners; identify what searchers can accomplish there that they cannot accomplish on your page.
Observed pattern
Working interpretation
Next check
Web Search impressions fall across one topic cluster
The cluster may have lost relevance or competitiveness for those searches
Compare query intent, result types, answer depth, and the pages that replaced it
Discover declines while Web Search remains stable
The evidence does not support a sitewide core-update diagnosis
Analyze Discover separately and account for the February 2026 Discover update
One template declines across unrelated topics
A shared presentation, technical, or content-production pattern may be involved
Compare affected and unaffected templates, including rendering, indexing, internal links, and visible page structure
Impressions remain stable but clicks decline
Visibility may not be the primary problem
Review titles, descriptions, competing result features, and whether the displayed promise matches the query
Search clicks remain stable but conversions decline
The ranking update is not sufficient to explain the business loss
Check tracking, page behavior, offer changes, availability, and conversion flow
All channels fall at the same time
A Google core update is unlikely to be the only cause
Check analytics integrity, site releases, outages, demand, and commercial changes
These interpretations are starting hypotheses, not automatic diagnoses. Require the pattern to appear in the underlying page and query data before assigning work to it.
Fix satisfaction gaps rather than chasing signals
Google’s standing direction remains to create helpful content for people. That advice becomes useful only when you turn it into page-level questions. “Make it better” is not an action. “Move the procedure ahead of the company background because the dominant queries ask how to complete the task” is an action.
Test the page against the searcher’s actual job
Write the main task in one sentence before reviewing the page. Is the person trying to learn, compare, troubleshoot, verify, calculate, choose, or complete a process? Then locate the first point where the page materially serves that task. If the answer is buried beneath a generic introduction, brand narrative, or loosely related background, fix the order before adding more words.
Check whether the title, opening, headings, body, examples, and call to action serve the same intent. A page often weakens when it promises one job in the search result, explains another in the body, and pushes a third in the call to action. Alignment matters more than repeating the target phrase.
Find the missing decision support
A page can be factually correct and still leave the reader unable to act. Look for absent prerequisites, constraints, tradeoffs, failure modes, definitions, examples, or next steps. Add only what closes a real decision gap. A longer page that delays the answer is not inherently more satisfying than a concise one.
Ask a hard comparative question: what can someone decide or do after reading the results now ranking above you that they could not decide or do after reading your page? The answer should become a concrete edit. If you cannot identify a meaningful difference, do not manufacture one by expanding every section.
Verify accuracy, ownership, and maintenance
Check every consequential claim, named feature, date, process, and recommendation. Remove unsupported certainty. Replace stale instructions. Make authorship and editorial responsibility clear where the reader needs them to judge the advice. Cite the originating authority when a claim depends on a standard, policy, specification, or official announcement.
Do not simulate freshness by changing a date while leaving old guidance intact. A meaningful update should have a reason you can record: a corrected fact, a changed process, a better explanation, a newly addressed intent, or clearer decision support.
Keep schema aligned with the visible page
JSON-LD can clarify the entities, properties, and relationships already represented on a page. It cannot turn thin, mismatched, or unsupported content into a satisfying result. Treat structured data as a consistency layer, not a core-update recovery switch.
After a substantive edit, verify that the markup still matches the visible content. Remove properties the page no longer supports, keep entity names and relationships consistent, and avoid adding types merely because they appear SEO-friendly. The content, metadata, internal links, and schema should describe the same thing without contradiction.
Make controlled changes and measure recovery honestly
Prioritize shared weaknesses that affect a meaningful group of pages. An isolated decline with no repeatable pattern is a poor reason for a sitewide rewrite. A clear intent mismatch across an entire template or topic cluster is a stronger candidate because the diagnosis and expected effect can be stated in advance.
Preserve the baseline data and a recoverable copy of every page before making material changes.
Resolve measurement, indexing, rendering, redirect, or deployment problems before judging content quality. Content edits cannot repair missing data or a broken delivery path.
Choose a coherent page group with one identifiable weakness. Define the intended change and the metric that should respond.
Make the smallest batch large enough to test the shared diagnosis. Avoid mixing unrelated URL, template, copy, schema, and commercial changes when they can be separated.
Annotate what changed, where, why, and when. Keep unaffected pages steady where practical so later comparisons retain context.
Re-evaluate the same page and query groups after Google has processed the changes. Judge visibility and qualified outcomes together rather than celebrating a traffic increase that does not serve the audience or business.
Choose the treatment page by page. Refresh a URL when its purpose remains valid but its answer is stale, incomplete, unclear, or poorly ordered. Consolidate pages when several weak URLs divide the same intent and none earns a distinct role. Leave a strong page alone when the evidence is inconclusive. Retire a page only when it no longer serves a user or business purpose; preserve the evidence first, and map a relevant redirect before removing a URL when a genuine replacement exists.
Google has not supplied a special one-step repair for this update. Recovery may be gradual and may become visible around subsequent core updates. That does not mean you should wait passively, but it does mean you should reject guaranteed recovery dates and avoid claiming that one edit caused a later movement without supporting evidence.
Your next action is straightforward: annotate the core, spam, and Discover context; preserve a clean baseline; map the largest losses by surface, query intent, page group, and template; and approve edits only where you can name the satisfaction gap. That turns a volatile month into a controlled recovery program instead of a trail of untraceable changes.
I’ve found an incredible new way to streamline content creation, competitive analysis, reporting, and monitoring with the latest Profound Agents feature. We can now effortlessly integrate prompt volume data directly into any Profound Agent, bringing together all our workflows into a single platform. This innovation is perfect for marketers looking to enhance efficiency.