The state of AI-search readiness, 2026
We ran the AICite audit on 25 of the internet's most recognized sites. Only 1 got an A. 2 got an F. The AI companies fared worst.
The full ranking
Sorted highest to lowest by AICite score. Every grade is deterministic against the same 13-signal rubric. Click any hostname for its full AI-search readiness breakdown.
| # | Site | Grade | Score | Notes |
|---|---|---|---|---|
| 1 | cloudflare.com | A | 91 | Full llms.txt, robots policy for every AI crawler, complete Organization + WebSite JSON-LD. |
| 2 | mistral.ai | B | 83 | Best-scoring AI company. Ships llms.txt, robust structured data. |
| 3 | hubspot.com | B | 80 | Marketing platform practicing what it preaches on schema. |
| 4 | stripe.com | B | 77 | Solid across the board; missing llms-full.txt keeps it from an A. |
| 5 | netlify.com | B | 77 | Robust structured data, complete robots.txt. |
| 6 | shopify.com | B | 77 | Strong OpenGraph, Organization schema, canonical. |
| 7 | linear.app | C | 69 | Great SEO fundamentals, no AI-specific files. |
| 8 | github.com | C | 68 | Missing llms.txt. Ironic given how much of the AI-training web lives on it. |
| 9 | notion.so | C | 66 | Half-ships JSON-LD; missing Article schema. |
| 10 | slack.com | C | 66 | Solid meta, missing AI-crawler policy. |
| 11 | figma.com | C | 62 | Missing AI files; JSON-LD present but light. |
| 12 | techcrunch.com | C | 62 | Weirdly better than The Verge or Wired. |
| 13 | apple.com | D | 57 | Marquee brand, sparse structured data. No llms.txt, no explicit robots policy for AI crawlers. |
| 14 | wired.com | D | 57 | News org with weak structured data. Missing llms.txt. |
| 15 | nytimes.com | D | 55 | Missing llms.txt, no AI-crawler robots policy. |
| 16 | meta.com | D | 55 | Corporate site is thin on structured data. |
| 17 | discord.com | D | 55 | Missing AI files, half-complete schema. |
| 18 | theverge.com | D | 51 | Media site, minimal AI-search prep. |
| 19 | vercel.com | D | 49 | Deploys a lot of AI apps; hasn't shipped its own llms.txt. |
| 20 | reddit.com | D | 46 | One of the internet's most-scraped sites, no AI policy files. |
| 21 | anthropic.com | D | 45 | Ships Claude. Doesn't ship llms.txt. |
| 22 | microsoft.com | D | 45 | Corporate homepage; sub-domains do better. |
| 23 | huggingface.co | D | 41 | The AI model hub — no llms.txt, no AI-crawler robots rules. |
| 24 | ycombinator.com | F | 32 | Founder-heavy audience but no AI-search prep at all. |
| 25 | wikipedia.org | F | 28 | The most-cited source on the internet fails 10 of 13 checks. Zero structured data on the homepage. |
5 things this tells us
1. AI companies aren't AI-search ready
The single sharpest finding. Of the 5 AI-focused companies we could grade (Anthropic, HuggingFace, Mistral, Meta, Microsoft), only Mistral broke into the B range. Anthropic — who co-authored the llms.txt spec — doesn't ship one on its own homepage. If the AI labs aren't tuning their sites for AI search, they may be assuming their own retrieval systems will figure it out. Most won't.
2. Cloudflare is quietly the model citizen
The only A in the dataset. Ships /llms.txt,
explicit robots policy for every major AI crawler, complete
Organization JSON-LD with rich sameAs array.
Cloudflare's business incentive around AI crawler transparency
is obvious — but that doesn't diminish the fact that they've
actually done the work.
3. News media has near-zero AI-search prep
NYTimes, The Verge, Wired, TechCrunch — a majority of the
media set landed in C or D territory. Given how heavily AI
answer engines reformulate news content, this is a strategic
blind spot. A publisher who ships llms.txt plus
a well-formed Organization block is likely to
get name-checked in AI answers where their competitors don't.
4. Wikipedia's F is real and consequential
Wikipedia gets grade F/28. Zero JSON-LD on the homepage,
no llms.txt, no explicit AI-crawler robots
policy. Because Wikipedia is served under such heavy caching
and infrastructure focus, the marketing/discovery layer has
been deprioritised. This matters because Wikipedia is
arguably the single most-cited source in every major LLM's
training data — its own homepage failing 10 of 13 signals
illustrates that our rubric is measuring "is this site
trying to be discoverable" not "does it get
retrieved anyway".
5. The bar to top-quartile is low
Only 6 sites made it into the B range or above. That means shipping the three files this site's SEO trilogy covers — llms.txt, robots.txt AI policy, Organization JSON-LD — would land any site in the top 25% of this leaderboard. That is a real advantage for any site willing to spend 30 minutes on it.
See where your site ranks
Methodology
Every site scored against the same 13 signals AICite uses for every audit. Weights sum to 100:
/llms.txtdiscovery file (weight 20)- JSON-LD present (8), Organization/WebSite (8), Article/FAQ/Product/HowTo (8)
/robots.txtwith explicit AI-crawler policy (8)/llms-full.txtfull-content dump (8)- Sitemap referenced from robots.txt (6)
- Canonical link (6), meta description (6)
- OpenGraph title + type (5)
<title>(4), HTTPS (4),<html lang>(3)
Each site was fetched live at audit time from Cloudflare's
edge network. No caching, no retries. Sites that returned
403 or throttled responses to Cloudflare Worker
outbound IPs are noted below and excluded from the ranking.
Notable exclusions
Four candidate sites blocked our audit fetches with
403 Forbidden or 202 Accepted
(implied throttling): openai.com,
perplexity.ai, amazon.com,
and producthunt.com. They may serve the same
content to consumer browsers but return errors to
cloud-datacenter IPs — a common anti-bot posture. The
practical implication: any AI crawler that fetches from a
cloud IP range (which is most of them) sees the same 403
response, meaning these sites are actively hostile
to AI-search retrieval regardless of their homepage schema.
The leaderboard is a snapshot. Grades change when sites update. We'll refresh this dataset periodically; the underlying rubric — the free AICite audit — never does.