You want AI systems to recognize and cite your expertise, but you don’t want a generated answer to replace the page, dataset, or original work that paid for it. A blanket allow-or-block decision cannot resolve that conflict.
The workable approach is to decide separately what should be discoverable, available for live answers, eligible for model training, or kept behind real access controls. Connect those decisions to business value and rights status before anyone edits a crawler directive.
Stop treating crawl access as one permission
Traditional search indexing, result previews, live retrieval for an AI answer, and model training are different uses. A platform may offer separate controls for some of them, combine others, or provide no control that matches the choice you actually want to make.
Google-Extended shows why the distinction matters. It can prevent content from being used for Gemini training without preventing live website information from contributing to AI-generated answers. Content already indexed by Google may also remain eligible to appear in AI Overviews. Blocking training, therefore, is not the same as blocking answer generation.
The European Commission’s antitrust investigation puts this lack of choice at the center of the dispute: publishers argue that they cannot meaningfully reject generative use without jeopardizing search visibility. The investigation does not settle what is lawful for your content, but it does expose the strategic mistake of treating search inclusion as consent to every downstream use.
For every important group of URLs, answer four separate questions:
- Should an ordinary search crawler be allowed to index this content?
- Should a search result be allowed to display a preview or snippet?
- Do you want an AI system to retrieve this page when constructing a live answer?
- Do you want the content used to train or improve a model?
Do not assume that one directive answers all four questions. Write down the desired outcome first, and then identify whether each platform provides a documented control for it.
A robots.txt rule is also not a security boundary. It communicates a preference to crawlers that honor it; it does not make public material confidential or prevent every form of copying. If disclosure of a dataset, licensed report, client deliverable, or proprietary method would cause serious commercial or legal harm, protect it with authentication or another genuine access control. If ownership or licensing terms are unclear, have intellectual-property counsel review them before changing access or reuse terms.
Build a rights-to-visibility matrix before changing directives

Make decisions at the URL-family level rather than applying one sitewide rule. A public glossary, a product page, an original investigation, and a licensed database do not carry the same discovery value or substitution risk.
| Decision factor | What to record | How it should affect your posture |
|---|---|---|
| Business role | Discovery, authority building, conversion, support, or paid deliverable | Discovery content usually benefits from broader access; a paid deliverable needs a stronger boundary |
| Rights status | Owned, licensed, contributor-supplied, user-supplied, or uncertain | Uncertain or restricted rights require review before you authorize new uses |
| Substitution risk | Whether a generated answer could satisfy the need without a visit | High-risk pages may need a useful public summary with the full asset kept under access control |
| Visibility dependency | Search impressions, qualified visits, leads, sales, or assisted conversions | Do not restrict a high-dependency URL group without a baseline and rollback plan |
| Distinctive value | Original data, reporting, methodology, tools, templates, or expert analysis | The harder the asset is to replace, the more deliberate its public surface should be |
| Available controls | Crawler, directive, affected product, documented behavior, and owner | Implement only controls that match the intended use closely enough to justify the tradeoff |
Turn that matrix into an implementable policy:
- Group URLs by template and business function. Start with categories such as public reference content, commercial pages, original editorial work, licensed material, and authenticated assets.
- Assign a default posture to each group: open for discovery, public but bounded, restricted, or licensed for specific uses.
- Record which team owns the decision. SEO can explain visibility consequences, but it should not silently decide rights questions for editorial, product, or legal teams.
- Inventory the current
robots.txtrules, page-level directives, authentication boundaries, and contractual restrictions before changing anything. - For each crawler instruction, record the exact crawler and product behavior it is meant to affect. Do not infer behavior from the directive’s name.
- Apply the first change to a non-critical URL family. Preserve the previous configuration, capture the baseline, and define the condition that would trigger a rollback.
The same caution applies to noai, nopreview, and similar emerging conventions. A label does not tell you which systems honor it, whether it affects training or live retrieval, or whether it changes ordinary search eligibility. Platform-specific documentation has to answer those questions.
Make the public layer easy to cite and hard to confuse
Protecting high-value material does not require making your whole brand invisible. A stronger architecture separates a public reference layer from the asset that contains the complete commercial value.
Build a useful public reference layer
The public page must contain enough substance to deserve selection. A vague teaser gives an answer engine little reason to cite you, while publishing the entire asset may let the generated response replace you.
- Put the core answer in fully rendered HTML. Googlebot can process JavaScript well, but other AI crawlers may not render a JavaScript-dependent page reliably.
- Use descriptive headings and answer one recognizable question directly under the relevant heading. Follow the short answer with scope, exceptions, evidence, and the next action.
- Name your organization, authors, products, and subject entities consistently. Make authorship, expertise, editorial responsibility, and update history visible rather than leaving authority to be inferred.
- Add structured data that agrees with the visible content. Appropriate schema, complete metadata, and meaningful image alt text can help machines connect the page to the correct entities, but markup does not grant a license or compel an AI system to cite you.
- Show provenance for consequential claims. Identify who produced original data, explain the method at a useful level, state important limitations, and distinguish an observed fact from your interpretation.
- Give the reader a reason to continue beyond the extracted answer: an interactive tool, complete dataset, implementation workflow, downloadable resource, consultation path, or transaction that the summary cannot reproduce.
Generic explanations are especially vulnerable to substitution because the answer contains little that belongs distinctly to your entity. The public layer should carry something attributable: a clear framework, original evidence, a named expert’s analysis, a transparent method, or a maintained record of change.
Keep the irreplaceable asset behind a real boundary
- Keep full proprietary datasets, premium templates, licensed archives, and account-specific outputs behind authentication when public exposure is not an acceptable cost of discovery.
- Publish a useful summary only if you are comfortable with that summary being publicly accessible and potentially reused.
- State ownership and permitted uses in clear terms, and provide a licensing or permissions contact for organizations that want broader access.
- Do not publish confidential material and rely on a bot instruction to protect it. Remove it from public delivery or require authorized access.
This creates a deliberate exchange: machines can understand what you know and why your entity is relevant, while the complete experience or asset still requires a relationship with you.
Measure whether visibility creates value or merely extraction

Organic sessions alone no longer describe search performance. Many AI interactions end without a click, so referral traffic cannot capture every useful mention or every instance in which your material satisfies the user elsewhere.
Some publishers have reported traffic declines of 20% to 50% on informational queries. That range is not a forecast for your site. It is a warning that rankings can remain visible while the economic value of the result changes.
Capture a baseline before changing access controls, then monitor five layers:
- Answer visibility: Use a fixed set of important prompts and record whether your brand, product, expert, or content appears. Keep the prompt wording stable enough to compare observations.
- Attribution quality: Record whether the answer names you, links to the correct page, represents the claim accurately, and distinguishes you from similarly named entities.
- Discovery: Track ordinary search impressions, clicks, AI referrals that can be identified, landing pages, and changes by URL family.
- Business value: Measure qualified conversions, assisted conversions, sales conversations, subscriptions, branded search, and other downstream outcomes that matter to the page’s assigned role.
- Exposure: Review server logs for crawler activity and document cases where protected or distinctive material appears elsewhere without the attribution or use you expected.
Interpret combinations of signals instead of chasing a single metric:
- If AI mentions rise and qualified conversions also rise, the public layer is probably supporting discovery even when direct clicks are limited.
- If mentions rise but links and downstream value do not, inspect whether the answer reproduces too much of the page, the citation is missing, or the page lacks a compelling next step. Blocking should not be your automatic first response.
- If visibility falls after a directive change, compare crawler logs, indexing, and the affected URL family against the recorded intent. Roll back when the lost discovery is more valuable than the use you prevented.
- If an AI answer misstates your position, improve the page’s explicit definitions, entity relationships, evidence, and limitations. Preserve examples of the error so you can determine whether the problem changed.
- If licensed, confidential, or access-controlled material is reproduced, preserve the output, URL, date, relevant access logs, and configuration. Escalate to the platform and qualified counsel rather than trying to settle the rights question through SEO settings alone.
Keep a change log with the affected URL family, intended behavior, implementation owner, prior configuration, observed result, and rollback condition. Without that record, a later traffic change will tempt the team to assign causation to whichever AI event is most visible.
Key takeaways
- Search indexing, snippets, live AI retrieval, and model training are separate uses, even when a platform does not provide separate controls for all of them.
- Google-Extended can address Gemini training without necessarily removing indexed content from AI Overviews or preventing live use in generated answers.
- Make rights decisions by URL family and business role, not with one sitewide allow-or-block rule.
- Schema and clear HTML improve machine understanding; they do not create access control, waive rights, or guarantee attribution.
- Use authentication for assets that must remain protected. Crawler preferences are not a substitute for a security boundary.
- Judge AI visibility by attribution, accuracy, qualified outcomes, and exposure as well as traffic.
Your next move is to choose one important URL family and complete the rights-to-visibility matrix before touching its directives. Capture the current configuration and performance, decide which uses you actually want, and change only the control that can credibly serve that decision. The durable strategy is neither maximum exposure nor total disappearance. It is a deliberately designed public surface with a defensible boundary around the value you cannot afford to give away.
References
- Search Engine Land — Google vs. publishers: What the EU probe means for SEO, AI answers and content rights
- Search Engine Land — From SEO to GEO: How marketing leaders stay visible in AI-driven search

Leave a Reply