You are not choosing an AEO agency because you need another content supplier. You are choosing one because your brand is missing, misrepresented, or overlooked when prospects ask answer engines questions connected to a purchase.
The difficulty is that an agency can promise visibility, but it cannot control what an external AI platform generates or cites. A sound selection process therefore focuses on what you can inspect: the agency’s diagnosis, evidence standards, implementation method, measurement protocol, and ownership terms.
Write the selection brief before you look at agencies
AEO can mean content production, technical SEO, structured data, entity management, digital PR, prompt monitoring, or some mixture of them. The market already spans agency-led strategy, creative content, AI-driven analysis, and DIY-oriented approaches. Those options become comparable only after you define the problem they must solve.
Start by choosing the primary outcome. Most AEO briefs contain one or more of these problems:
- Presence: Your brand does not appear in answers to relevant non-branded questions.
- Accuracy: Answers mention your brand but get important facts, capabilities, availability, or positioning wrong.
- Preference: Your brand appears, but competitors receive the recommendation, supporting explanation, or citation.
- Conversion: You earn mentions or referral visits, but the cited pages do not help qualified visitors take the next step.
These are not interchangeable. A mention-tracking campaign will not fix unsupported product claims. Schema work will not repair weak third-party authority. More content will not solve a conversion problem on an already cited page. Ask every candidate to state which problem it believes you have, what evidence supports that diagnosis, and what it would deliberately leave out of scope.
Your brief should also identify:
- The answer platforms and interfaces that matter to your audience, named explicitly rather than grouped under AI.
- The markets, languages, locations, and audience segments in scope.
- The product lines, services, topics, and entities the engagement covers.
- The questions that matter across discovery, comparison, validation, and purchase.
- The claims that require legal, compliance, product, medical, or subject-matter review.
- The systems the agency may need to touch, including your CMS, analytics, tag manager, schema implementation, product data, and reporting tools.
- The business event you ultimately care about, such as a qualified inquiry, signup, demo request, purchase, or assisted conversion.
Use this brief template: Improve [presence, accuracy, preference, or conversion] for [audience] asking [question groups] on [named platforms and interfaces], within [market and language], while protecting [brand, compliance, security, or editorial constraints].
Give each shortlisted agency the same brief. If one candidate is allowed to redefine the objective while another must answer your original request, their proposals will not be comparable.
Attach a baseline where you can. Include your approved brand facts, current priority pages, analytics definitions, known technical constraints, and a representative query set. For observed answers, record the exact question, platform, interface, date, location or language context, account state when relevant, generated answer, cited URLs, and whether the brand description was correct. AI outputs can vary, so a screenshot without its run conditions is weak evidence.
Inspect the method from question to business outcome

A serious AEO method connects audience questions to evidence, content, technical implementation, external authority, and measurement. If a proposal jumps from keyword research directly to publishing pages, ask what happened to the other layers.
Question demand and entity facts
A search keyword export is useful input, but it is not a complete model of answer demand. People ask full questions, add constraints, compare alternatives, challenge claims, and continue a conversation. The agency should show how it groups those behaviors without pretending it can enumerate every possible prompt.
Ask for a sample question map containing:
- The audience and decision stage behind each question group.
- The answer the user needs, not merely the phrase they typed.
- The entities, attributes, comparisons, and evidence required for a useful response.
- The pages or external assets that currently support the answer.
- The gap: missing evidence, ambiguous language, conflicting facts, poor retrieval, weak authority, or an unsuitable destination page.
- The assumptions used to choose platforms, markets, and query variants.
Look for an entity-fact process as well. Your company name, products, executives, locations, prices, policies, credentials, and other important attributes may appear across many owned and third-party properties. The agency should identify a canonical fact owner, the approved wording, where each fact is published, and how changes propagate. Otherwise, content teams can create the same inconsistency they were hired to fix.
Keep part of the evaluation set separate from the questions used to shape the work. Testing only the prompts the agency optimized against encourages dashboard overfitting. A separate evaluation set will not eliminate output variability, but it gives you a cleaner check on whether the work generalizes.
Content and technical implementation
AEO content should make useful claims easy to understand without stripping away the conditions that make them true. That requires more than short answers. It requires clear definitions, explicit relationships, comparison criteria, supporting evidence, qualified claims, suitable authorship, and a page structure that keeps the answer connected to its context.
Ask the agency to walk through a real content brief. It should show the target question, intended reader, factual inputs, missing evidence, subject-matter reviewer, answer structure, internal links, citation needs, conversion path, and update owner. If the brief is mostly a word count and a list of keywords, the operating model is still conventional content production with an AEO label.
Technical work should be equally concrete. The proposal should explain how crawlers reach the relevant content, how client-side rendering or access controls affect retrieval, how duplicate or conflicting URLs are handled, and how structured data maps to visible page content.
JSON-LD can express entities and relationships in a machine-readable form, but valid markup does not prove the underlying claim and does not guarantee inclusion in an answer. Ask for a content-to-schema crosswalk showing which visible fact supports each property, where the data comes from, who maintains it, how it is validated, and what happens when the page changes. The deployment plan should include staging, approval, monitoring, and rollback rather than direct, unreviewed changes to production.
Authority beyond your own website
Your website is only one place where an answer system may encounter your brand. A complete plan should consider the wider set of public materials that describe the business, while distinguishing assets you control from mentions you must earn.
Ask the agency to separate:
- Owned corrections: Resolving inconsistent facts across your site, profiles, documentation, feeds, and public company information.
- Earned authority: Creating evidence and expert contributions that can merit independent coverage, citations, or relevant links.
- Community participation: Answering real questions under the rules and norms of the relevant platform.
- Manipulative activity: Synthetic reviews, disguised promotion, fabricated expertise, or mass-produced third-party placements.
Do not accept the last category as an unavoidable shortcut. It creates platform, reputation, and potentially legal exposure while giving you assets that may disappear as soon as the vendor relationship ends. Ask who performs off-site work, whether subcontractors are involved, how placements are disclosed, and which tactics the agency refuses to use.
Measurement that separates observation from attribution
An AI visibility score is not self-explanatory. You need its denominator, query set, run conditions, treatment of citations, treatment of answer variation, and rules for adding or removing prompts. Without those definitions, a rising score may reflect a changed dashboard rather than changed market visibility.
Require a metric dictionary before implementation. It should separate:
- Implementation signals: Content coverage, supported entity facts, access issues, schema validity, editorial completion, and distribution work.
- Observed answer signals: Brand presence, factual accuracy, cited URLs, competitor inclusion, recommendation context, and answer consistency across the defined evaluation protocol.
- Business signals: Referral sessions where identifiable, engagement on cited landing pages, assisted conversions, qualified leads, purchases, and downstream value where your analytics can support the connection.
The reporting system should retain raw observations and a change log. If an answer changes after a page update, that is an association worth investigating. It is not automatically proof that the update caused the change. A trustworthy agency will mark that distinction instead of converting every favorable movement into a success claim.
Demand evidence you can audit

Polished decks show communication skill. They do not, by themselves, show that the agency can diagnose your problem or execute safely. Ask for work artifacts that expose how decisions were made.
| Agency claim | Evidence to request | Warning sign |
|---|---|---|
| We improve AI visibility | A redacted baseline and result captured under a defined protocol, plus the intervention, observation conditions, and limitations | A favorable screenshot with no query denominator, run conditions, or losing examples |
| We produce AEO content | A content brief, before-and-after page, factual evidence requirements, reviewer workflow, and edit rationale | Publishing volume presented as the outcome, with no evidence or governance process |
| We implement structured data | A page-to-schema mapping, validation output, data ownership model, deployment process, monitoring plan, and rollback path | A list of schema types with no explanation of whether the pages support the properties |
| We measure answer performance | The metric dictionary, prompt-set governance, raw observation export, change log, and treatment of variable outputs | A proprietary score whose components or historical inputs cannot be exported |
| We know your industry | Work showing how the team handled your industry’s claims, evidence, review, buying process, and constraints | A client-logo slide with no explanation of the work performed |
| We can execute the strategy | Names and roles of the delivery team, sample handoffs, approval responsibilities, and dependencies on your staff | Senior specialists lead the sale but the delivery team remains unnamed |
For each case example, ask what the agency delivered, what the client delivered, what changed, what failed, and how the outcome was measured. Improvements can come from a site migration, brand campaign, product launch, public relations event, demand shift, or internal content work happening alongside the engagement. The agency does not need to prove laboratory-style causality, but it should disclose important concurrent changes.
Reference calls are most useful when you ask operational questions:
- Which promised deliverables were actually usable without rework?
- How much access to internal experts and editors did the engagement require?
- What did the agency try that did not work, and how did it respond?
- Could the client export the raw data and continue the process independently?
- What became difficult during renewal or offboarding?
Listen for specificity rather than universal praise. A reference who describes tradeoffs, dependencies, and a failed idea may tell you more than one who offers only a positive verdict.
Use a paid diagnostic as the final audition
When the expected engagement is substantial, use a bounded paid diagnostic before committing to a broad retainer. Payment lets you request real work without disguising free strategy as procurement. A narrow scope limits your commitment while revealing how the agency reasons, communicates, handles uncertainty, and works with your team.
Choose a real business area, not a toy exercise. Give the candidate access only to the information required for that area and ask for:
- A baseline built from the agreed question set and observation protocol.
- An inventory of supported, missing, ambiguous, and conflicting entity facts.
- A diagnosis that separates content, technical, authority, measurement, and conversion problems.
- An opportunity map ranked by expected value, confidence, effort, dependencies, and risk.
- A sample content or schema intervention detailed enough for your team to review.
- A measurement plan connecting implementation, observed answers, and business outcomes.
- A backlog that names the owner, required input, approval path, and completion evidence for each item.
- A list of assumptions, unknowns, and conditions that could change the recommendation.
Do not judge the diagnostic by the size of its opportunity forecast. Judge whether it finds a real constraint, distinguishes evidence from inference, prioritizes work your organization can execute, and makes its data reviewable.
Set pass-or-fail gates before scoring presentation quality. A candidate should fail the process if it guarantees placement in external answers, refuses to explain its metrics, will not transfer usable data, proposes unsafe access, hides the delivery team, or relies on tactics your brand cannot defend publicly. A strong creative idea should not cancel out a basic ownership or integrity problem.
Turn the operating model into contract language
Vague contract language turns a clear pitch into an unmanageable engagement. Optimize content is an activity, not a deliverable. Replace it with named outputs, acceptance criteria, owners, and evidence of completion.
Make the agreement explicit about:
- The platforms, interfaces, markets, languages, entities, and content areas in scope.
- The agreed deliverables, review process, revision boundaries, and acceptance criteria.
- Which implementation work the agency performs and which work remains with your internal teams.
- How the question set, measurement method, and reporting definitions may change.
- Your ownership of briefs, content, schema, research outputs, dashboards, prompt sets, raw exports, and configuration files.
- Your right to retrieve historical data in a usable format when the engagement ends.
- The named delivery roles, subcontractor rules, and process for replacing key personnel.
- How confidential information may be entered into AI tools, whether providers retain it, and which security or privacy approvals apply.
- The access model for your CMS, analytics, search tools, repositories, and production systems.
- Change approval, backups, rollback responsibilities, incident handling, and offboarding.
- The activities excluded from scope, including development, public relations, design, analytics engineering, legal review, or subject-matter validation where relevant.
Use least-privilege access. A diagnostic rarely requires broad production permissions. Prefer read-only access, scoped accounts, staging environments, backups, and an approved deployment path. At offboarding, revoke accounts and credentials, transfer source files and historical exports, and confirm that scheduled automations no longer act on your systems.
External answer placement should never be the guaranteed deliverable because the agency does not control the platform. It can commit to work it controls: audits, briefs, implementations, reviews, monitoring, reporting, experiments, and documented response times. If data rights, privacy, indemnity, regulated claims, or intellectual-property terms create material exposure, have the appropriate legal or compliance owner review them before signature.
Key takeaways
- Define whether you need presence, accuracy, preference, or conversion improvement before requesting proposals.
- Require a method that connects questions, entity facts, content, technical implementation, external authority, and business measurement.
- Evaluate artifacts and raw observations, not screenshots, client logos, publishing volume, or an unexplained visibility score.
- Use a bounded paid diagnostic to test the agency’s reasoning and operating fit on a real part of your business.
- Make guarantees, data portability, asset ownership, delivery-team transparency, and safe access pass-or-fail conditions.
- Contract for named outputs and acceptance evidence rather than broad optimization activity.
Your next move is simple: put the brief, evidence requests, diagnostic output, and pass-or-fail gates into one request and send the same version to every shortlisted agency. Choose the team that makes its work inspectable, its uncertainty visible, and its assets transferable. That gives you something more durable than a forecast: an AEO program you can govern after the sales meeting ends.

Leave a Reply