Schema.org adoption can now be discussed with more evidence than anecdote. A reported monthly dataset shows how broadly individual Schema.org types and properties appear across domains observed through Google’s public web crawling infrastructure.
The figures are best treated as directional adoption signals, not exact market-share measurements or proof that a term improves search performance. Because the supplied material contains one report, the dataset details below are attributed to that report and are not independently corroborated here.
Key takeaways
- The reported statistics count unique domains using a Schema.org term, rather than every page or markup instance.
- Results appear in broad ranges such as 10K-100K domains instead of as exact counts.
- The source says the files are updated monthly and available in JSON, CSV and summary JSON formats.
- Adoption data can support prioritization and benchmarking, but it does not establish implementation quality, eligibility for search features or business impact.
What the adoption metric actually measures
According to the supplied CrushPress.AI report, Schema.org term frequencies are evaluated within Google’s public web crawling infrastructure and aggregated at the domain level. If one domain uses the same term on 100 pages, that still contributes one domain to the reported range for that term.
This unit of measurement answers a particular question: how widely has a term spread among observed websites? It does not answer how many pages contain the term, how frequently it appears within a site or how much content the markup describes.
The report says each record identifies whether the term is a type, such as Person or Event, or a property, such as price or telephone. It also includes the term’s official URI and a domain-count bucket. Those fields make it possible to distinguish the vocabulary item being measured from the range used to express its adoption.
Why ranges are more useful than they first appear

The source reports that Schema.org publishes ranges such as 10K-100K domains rather than precise totals. It says this approach reduces the effect of daily fluctuations and helps preserve website privacy. Monthly updates provide recurring snapshots without suggesting a level of precision the underlying observation process may not support.
That design changes the appropriate analysis. A bucket can reveal whether a term is niche, moderately adopted or broadly established, but it cannot support an exact adoption rate. Two terms in the same range also cannot be reliably ranked from the bucket alone, and movement within a range will remain invisible until a boundary is crossed.
Month-to-month comparisons therefore require restraint. Remaining in one bucket does not prove that usage was static, while entering a new bucket indicates a threshold crossing rather than disclosing the precise size or timing of the change.
A practical way to use the dataset

Start with relevance, not popularity
A term should first match the entity, attribute or relationship a site genuinely needs to describe. A large adoption bucket can show that implementation is common across domains, but popularity cannot make an irrelevant term appropriate.
Use adoption as supporting evidence
When several relevant terms compete for development time, the reported ranges can add an external signal to the decision. Teams can pair that signal with content coverage, technical effort, maintenance ownership and the specific purpose of the markup. The source suggests that visible adoption may also help make the case for implementation to development stakeholders.
Preserve the reporting context
Any internal dashboard or recommendation should record the term, whether it is a type or property, its official URI, the observed bucket and the monthly dataset snapshot used. The source says raw files are available through the Google Public Stats dataset on GitHub in JSON and CSV, with a summary JSON format containing aggregated bucket distributions.
The conclusions the figures cannot support
Domain adoption is not a quality score. The reported metric does not state whether markup is valid, complete, current or faithful to the visible content. It also does not show whether a search system used the markup, whether a search feature appeared or whether traffic and conversions changed.
The crawling context matters as well. The source ties the frequencies to Google’s public web crawling infrastructure, so the figures describe domains observed within that system rather than an independently established census of every website. Broad buckets further limit fine-grained comparisons.
The most defensible role for this dataset is as a recurring map of vocabulary diffusion. Used alongside implementation audits and site-specific objectives, future monthly snapshots can make structured-data planning more evidence-aware without turning adoption into a substitute for relevance or quality.

Leave a Reply