AI readiness study: what we found on 652 small-business websites
We checked 652 small-business websites across five industries for public access, structured information and selected agent-discovery files. The results show specific gaps businesses can investigate, with clear limits on what a homepage scan can prove.
Reviewed October 1, 2026 · 9 min read
On this page7 sections
What information does an automated visitor find on a typical small-business homepage? We examined that question using a defined sample and a deterministic scanner. The study measures detected website signals, not search rankings, AI recommendations or completed transactions.
The final dataset contains 652 unique websites across restaurants, health, home trades, personal care and independent retail. Of those, 253 sites (38.8%) had no parseable JSON-LD in the scanned homepage response, and 27 (4.1%) met the scanner’s structured-pricing check.
None passed the five JSON discovery-file probes used in this scan. That is a finding about specific paths and response checks. It does not mean the businesses have no APIs, booking integrations or agent-assisted purchase options elsewhere.
The tables below preserve the August 2026 results. For this October 1 editorial review, we rechecked the stored aggregates and clarified their interpretation. No new websites were scanned for this revision, and we have replaced the earlier claim of agent invisibility with descriptions of what was actually measured.
Key findings from the 652-site sample#
- 253 sites (38.8%) had no parseable JSON-LD in the scanned homepage response. This includes responses the scanner could not successfully retrieve.
- 27 sites (4.1%) met the structured-pricing check, while 148 (22.7%) contained a price-like expression in extracted homepage text.
- 381 sites (58.4%) had a detected action link or form, and 208 (31.9%) had a detected structured action signal. Neither check verified completion.
- No site passed any of the five JSON discovery-file probes: two agent.json locations, an A2A card location, an MCP card location and an OpenAPI location.
- 58 sites (8.9%) had no successful homepage response recorded. Causes included denied requests, missing pages, server errors and fetch failures.
- 33 sites (5.1%) had a robots.txt rule interpreted as blocking at least one checked AI token. Separately, 215 (33.0%) passed the scanner’s basic response check at /llms.txt.
The mean scanner score was 49.1 out of 100, with a median of 51. This is a score under Nexez’s version 2 rubric, not an industry-standard measure or the probability that an agent will recommend or transact with a business.
What the original 30.7% headline measured#
The original article used a composite measure that combined an access result with missing markup and discovery-file signals. We retain the calculation here so the earlier figure can be understood and reproduced.
A site met the composite when its homepage status was outside 200 through 399, or when it had neither parseable JSON-LD nor a positive result from any of six checks: /agent.json, /.well-known/agent.json, /.well-known/agent-card.json, /.well-known/mcp.json, /openapi.json and /llms.txt. Exactly 200 of 652 sites met it, or 30.7%.
We call this the access/markup gap in the table below. It is not proof that an agent cannot understand a business. Agents can read ordinary HTML, search indexes and other sources, and a homepage scan does not cover all of them.
Excluding the /llms.txt response flag raises the composite to 253 sites (38.8%), a difference of 53 sites, or 8.1 percentage points. Another count, 56 sites with that flag but no parseable JSON-LD, includes three unsuccessful homepage responses already counted in the composite. Those counts should not be treated as interchangeable. The llms.txt guide explains the proposal and its limits.
The per-industry breakdown#
| Industry | Sites | Mean score | No parseable JSON-LD | Structured pricing | Access/markup gap |
|---|---|---|---|---|---|
| Restaurants and cafes | 156 | 48 | 40% | 3% | 33% |
| Health (clinics, dentists, doctors) | 145 | 50 | 32% | 5% | 28% |
| Home trades (plumbers, electricians, HVAC) | 83 | 52 | 36% | 13% | 31% |
| Personal care (salons, beauty, massage) | 129 | 49 | 38% | 2% | 33% |
| Independent retail | 139 | 48 | 47% | 1% | 29% |
Home trades had the highest detected structured-pricing share at 13%, compared with 1% to 5% in the other groups. Eighteen percent of trade sites also had a detected offer-schema signal. The study did not establish whether a particular website builder or business practice caused that difference.
Independent retail had the largest share without parseable homepage JSON-LD, at 47%. Health had the smallest access/markup gap, at 28%, while restaurants and personal care were both 33%. Percentages in the table are rounded to whole numbers; the industry groups have different sample sizes.
Action signals are not completed transactions#
The scanner found action links or forms on more than half the sites. That gives us evidence of a possible next step, but does not tell us whether a particular agent can use it successfully.
The five JSON discovery probes returned no positive results across the sample. These checks looked for expected fields at fixed locations. An MCP server on another host, an API documented elsewhere or a third-party booking flow could be missed. A successful discovery-file check would also not prove that an action works.
To evaluate completion, test an actual supported journey: finding a suitable offer, checking availability, obtaining approval and receiving confirmation. The service booking guide covers those routes. This study did not attempt bookings or purchases.
Check the signals on your own site
The public Nexez scanner reports access and machine-readable information for your site. Use the findings as a diagnostic starting point, then inspect relevant product pages and test checkout or booking separately. The public scanner may evolve beyond the version used for this study.
Scan your site freeRobots rules and the /llms.txt response flag#
The robots.txt results describe directives observed by the scanner. They do not demonstrate how every operator behaves or whether a particular request was blocked at the network edge.
The most frequently disallowed tokens were GPTBot at 4.9%, ClaudeBot at 4.6% and Google-Extended at 4.1%. Each of the other checked tokens was below 1%. Google-Extended is a robots control token rather than a separate HTTP crawler identity. The crawler policy guide explains why training, search and user-requested access should be considered separately.
The 58 unsuccessful homepage results included 25 HTTP 403 responses, 24 fetches with no HTTP status recorded, four 404s, two 429s, and one each of 500, 503 and 526. These are not all firewall blocks. A failed request from our scanner does not establish that every other agent receives the same result.
The 33.0% /llms.txt figure needs particular care. Version 2 accepted a successful response with at least 20 trimmed characters; it did not validate that the body was a genuine llms.txt document. A generic HTML fallback could pass. Although 159 sites (24.4%) had both this flag and parseable JSON-LD, the study cannot establish llms.txt adoption, builder defaults or the operator’s intent from those flags alone.
Methodology#
The sample came from OpenStreetMap through the public Overpass API during August 10-11, 2026. Stored scans in the reported dataset are dated August 11 in UTC. Twelve US metros were selected in advance: Columbus OH, Raleigh NC, Tucson AZ, Spokane WA, Grand Rapids MI, Chattanooga TN, Boise ID, Worcester MA, Baton Rouge LA, Reno NV, Des Moines IA and Richmond VA. The frame is reproducible from the documented selection procedure, but OSM itself changes over time.
- Eligibility: an OSM record needed a website or contact:website tag. The selection excluded known platform-hosted pages and records with chain-identifying brand tags. Those filters are imperfect proxies for independently owned business websites.
- Selection: domains were ordered within each metro-industry cell using a SHA-256 hash seeded with the cohort label, with a cap of 14 per cell. Domains were deduplicated rather than manually selected.
- Scanning: scanner version 2 fetched the homepage and probed fixed discovery paths and robots.txt. Its scoring logic used response data and deterministic heuristics, without an LLM judging the page. Parseable JSON-LD meant that a script body parsed as JSON, not that every schema property was correct.
- Access preferences: the harness checked robots.txt for its own identified user agent and skipped five disallowed targets. Scans ran in small batches. The homepage checks did not reproduce every commercial agent’s browser, network or credentials.
- Accounting: 722 targets were sampled. There were 655 completed scan records, 62 targets classified as errors after retries, and five robots exclusions. Deduplicating completed results by the hashed final domain left 652 sites. Completed records could still contain an unsuccessful homepage response.
Limits of the sample and scanner
This is not a probability sample of all US small businesses. OSM coverage, website tagging and eligible counts vary by location and industry. Businesses relying on platform pages are excluded. We checked homepages and fixed paths, without a full browser crawl or transaction test. Failed targets were excluded from the final denominator, and a basic response check can produce false positives. These limits prevent us from turning the percentages into a national invisibility or booking-failure rate.
What businesses can use from the findings#
Begin with concrete questions: does your public page load for the clients you want to serve, does it explain your business accurately, and do prices and booking terms appear where customers need them? Missing homepage JSON-LD is a prompt to inspect the implementation, not evidence that the business cannot be found.
Add accurate structured data where it describes visible content and supports a relevant consumer. Keep it synchronized with current offers. Choose additional feeds or agent interfaces based on the customers and platforms you need to support, as explained in the agentic commerce guide.
Finally, test the action that matters: a quote request, booking or purchase. Record the result in the system that owns it. Future comparisons should report changes in sampling, scanner logic and measurement limits so a higher score is not mistaken for proven growth in recommendations or sales.
Turn clear business information into a useful next step
Nexez publishes structured business listings and offers with supported checkout and scheduling paths. See how those listings can complement your existing website and customer journey.
See how it worksFrequently asked questions
How many websites did the study scan?
The study sampled 722 targets across 12 US metros and five industries. It recorded 655 completed scans, 62 errors and five robots exclusions. Deduplicating the completed records by final-domain hash left the 652 sites used in the reported percentages.
Does the study prove that 31% of sites are invisible to AI?
No. The 30.7% figure combines unsuccessful homepage responses with missing JSON-LD and selected discovery-file signals. It is a scanner composite, not a test of whether an assistant can find, understand or transact with a business. The revised article explains the calculation and its limits.
Where did the sample of websites come from?
The sampling frame was OpenStreetMap records with business website tags, queried across 12 selected US metros. Known platform pages and brand-tagged chains were excluded, and domains were selected through a deterministic hash order. The result is a defined sample, not a representative survey of every US small business.
Was an AI or LLM used to score the websites?
No. Version 2 used deterministic parsing and heuristic checks on retrieved responses. The same captured inputs and evaluation conditions produce the same result, but rescanning a live website can differ because pages, network conditions and time-dependent signals change.
Does the 33% llms.txt result measure confirmed adoption?
No. The original check accepted a successful response with at least 20 trimmed characters at /llms.txt. It did not verify that the body was a genuine llms.txt document, so generic fallback pages could pass. The result is a response flag and should not be treated as a confirmed adoption rate.
Will the study be repeated?
The documented procedure supports a follow-up, and the original plan called for a six-month comparison. A repeat should preserve or disclose changes to the sample frame and scanner version. This October editorial revision rechecked the stored August aggregates; it did not run a new cohort.
Keep reading
Which AI crawlers should your business allow?
Search crawlers, training crawlers and user-triggered fetchers serve different purposes.
7 min readAgent readinessJSON-LD for AI agents: practical schema examples
JSON-LD gives software an explicit description of the facts on your page.
6 min readAgent readinessWhat is llms.txt, and when is it useful?
An llms.txt file gives an agent a short guide to your website.
8 min read