File viewer

live_checks_12082026_0255.md

/app/data/llm/analysis/bupa/live_checks_12082026_0255.md

The probe data from a single datacenter IP shows zero blocking across all 21 requests — all seven user agents, including AI crawlers, received 200 or 404 responses with no 403s or errors. This is the headline: from this vantage point, the site is fully open to the tested bots.

What stands out

  • 0% blocked across all AI crawlers. GPTBot, ClaudeBot, PerplexityBot, Google-Extended, OAI-SearchBot, Bytespider — every request was allowed. The blocked_pct is 0.0 for every agent on every path. This is not a server-side blocking problem, at least from this probe IP.
  • Two of the three probed paths return 404. /contact/ and /services/ both yield 404 for every agent, including the browser. The only 200 is /robots.txt. If these are intended to be core pages (contact, services), they are missing or the probes used incorrect URLs. The 404s are consistent and not agent-specific.
  • Response times are healthy. Overall average 207.2ms, p95 274ms. No request exceeded 445ms. The browser was slightly slower (avg 275.7ms) than the bots, but the difference is marginal and not a concern.
  • No differential treatment. The differential_count is 0 — no instance where a bot was treated differently from the browser. This reinforces the finding of uniform access.

What’s concerning

The 404s on /contact/ and /services/ are the only red flag. If those pages are meant to exist, they are broken and represent lost SEO value. If they are not the correct URLs, the probe simply targeted the wrong endpoints — but the pattern (two plausible paths both 404) suggests a higher likelihood that the pages are genuinely missing or have moved without redirects. This needs immediate verification against the site’s actual URL structure.

What’s positive

No blocking of AI crawlers. If the site is aiming to be cited in AI-generated answers, the technical door is open — at least from this probe. The 200 on /robots.txt means bots can fetch directives, so any restrictions would be in the file itself, not at the HTTP level.

Connections between data points

The combination of zero blocking and consistent 404s on two high-value paths points to a potential root cause: AI visibility problems are not a server-access issue. If the site is not appearing in AI search results, the bottleneck is likely content quality, crawlability of the correct pages, or the fact that key pages simply don’t exist. The healthy response times rule out timeout-based crawl abandonment.

Hypothesis to verify

The probe IP is a single datacenter address. If the site uses Cloudflare, the production bot IPs may be whitelisted while this IP is treated differently. The 0% block rate here does not prove that Googlebot or Bingbot would get the same treatment. To confirm, we would need server logs or probes from IP ranges the actual crawlers use. This is a suggestive but not definitive finding.

Key signals

  • 0% blocking across all bots — all 21 requests allowed, no 403s, no errors. AI crawlers are not being blocked at the HTTP level from this probe.
  • Two of three probed paths return 404 — /contact/ and /services/ are missing or incorrectly targeted. If these are intended live pages, this is a critical content gap.
  • Uniform treatment across all agents — browser and bots see identical status codes, no differential blocking. Consistent access.
  • Response times all under 500ms — no performance penalty that would discourage crawling.

Potential featured finding: If AI citation counts are low, the root cause is not HTTP-level blocking. The data points toward content/crawlability issues, specifically the 404s on pages that likely matter. Fixing those missing pages — or redirecting them to correct URLs — would unlock the most immediate value, assuming the URLs are correct targets.