live_checks_12082026_0142.md
/app/data/llm/analysis/house730/live_checks_12082026_0142.md
Analytical Commentary
Overview
This probe data is clean but paints a stark picture: the site is returning 403s to almost everything from this datacenter IP. 18 of 21 probes came back as 403, which is an 85.7% block rate. There are no DNS or connection errors — every probe resolved and got a response — so the infrastructure is healthy, but access control is extremely tight.
The Central Finding: Near-Universal 403 Blocking
Every user agent except PerplexityBot received a 403 on all three paths. That includes the browser user agent, which is the most significant data point here.
A browser receiving a 403 from a datacenter IP is not unusual — it's consistent with a Cloudflare or similar WAF blocking non-residential IP ranges at the edge. This is a critical interpretive constraint: we cannot conclude from this data alone that any specific crawler is blocked in production. The 403s are likely IP-based, not UA-based. If real Googlebot hits from a Google IP range, this same WAF rule would likely let it through. We need server logs or IP-verified crawler tests to confirm whether legitimate search engine crawlers are affected.
The PerplexityBot Exception
PerplexityBot is the only UA that got through and received non-403 responses across all three paths:
robots.txtreturned a 200/services/and/contact/both returned 404
The 200 on robots.txt means PerplexityBot can read the directives. The 404s on the other paths mean those pages legitimately don't exist, or the paths don't resolve — PerplexityBot wasn't blocked, it just hit dead ends.
Why is PerplexityBot allowed when everything else, including the browser UA, is blocked? Two hypotheses to verify:
- PerplexityBot is explicitly whitelisted at the WAF or origin level, perhaps as part of an intentional AI crawler strategy — though that would be unusual given that all other AI bots are blocked.
- PerplexityBot's requests originate from an IP range that happens to fall outside the datacenter-blocking rule. If Perplexity's infrastructure uses residential or cloud IPs that the WAF doesn't flag, this is a coincidental bypass, not an intentional whitelist.
The response time data supports hypothesis 2: PerplexityBot averaged 215.7ms across its three probes, while the blocked bots averaged 62–87ms. The 403 responses are fast because they're rejected at the edge. PerplexityBot's slower responses suggest it's actually reaching the origin server — consistent with its traffic not triggering the edge block. The browser UA also hit the origin once (558ms on robots.txt), likely because it got a 403 that still took time to generate, but this is a single outlier in an otherwise blocked pattern.
AI Crawler Impact: All Major Bots Blocked
The six AI crawler bots tested — GPTBot, Google-Extended, OAI-SearchBot, ClaudeBot, Bytespider, and PerplexityBot — break down into two clear groups:
- Blocked (5 of 6): GPTBot, Google-Extended, OAI-SearchBot, ClaudeBot, Bytespider — all 100% blocked, every path, every probe. 15 probes, 15 403s.
- Allowed (1 of 6): PerplexityBot — all 3 probes returned 200 or 404, never 403.
If these blocks are IP-based and not UA-based (as the browser 403 suggests), then the actual production impact on AI crawlers is unknown. If these bots hit from non-datacenter IPs, they may be crawling the site freely. If they share IP ranges with the datacenter probe source, the site is invisible to every major AI crawler except Perplexity.
This is a high-priority question to resolve: check server logs for actual GPTBot hits and their response codes.
The robots.txt 403 Problem
The robots.txt path returning 403 to every UA except PerplexityBot is the single most concerning technical signal in this data — with the browser UA 403 as the important caveat.
If a legitimate crawler receives a 403 on robots.txt, it typically stops and does not crawl the site at all. This file is meant to be publicly readable. Blocking it at the edge is a misconfiguration risk if the WAF rule is too broad and catches legitimate crawler IPs.
However, the browser UA also got a 403 on robots.txt, which reinforces the IP-based blocking interpretation. If Googlebot hits from a proper Google IP range, it likely gets through. We cannot confirm the blocking risk without log data showing real crawler responses.
Response Time Patterns
The overall response time average is 111ms, but this number is misleading because it mixes two distinct populations:
- Blocked responses (403s): Fast, averaging 62–87ms across all blocked UAs except the browser outlier. This is consistent with edge-level rejection — the request never reaches the origin.
- Allowed responses: Significantly slower. PerplexityBot averaged 215.7ms, and the single browser probe on robots.txt took 558ms. These are origin-level response times.
The p95 of 178ms across all probes is dominated by the fast 403 responses. There is no latency concern worth flagging — the edge infrastructure is responsive.
Path Availability
All three paths respond, but only robots.txt returned a 200 (to PerplexityBot). /services/ and /contact/ both return 404 to PerplexityBot. This means:
/services/either doesn't exist or PerplexityBot isn't hitting the right URL variant (trailing slash vs non-trailing, case sensitivity)./contact/same — the path doesn't resolve.- Both 404s are consistent across paths, so it's not a one-off routing issue.
If the site has a /services/ or /contact/ page under a different URL pattern, that's a non-issue. If those pages exist and this is a canonicalisation problem, it needs attention. Either way, it's low severity compared to the blocking questions.
What This Data Cannot Tell Us
- Whether Googlebot, Bingbot, or any other search crawler is actually blocked in production. A 403 from a datacenter IP does not equate to a production crawler block.
- Whether the 403s are from Cloudflare, a different WAF, or origin-level rules.
- Whether PerplexityBot's access is intentional or coincidental.
- What content exists behind the blocked responses — we have no successful page renders beyond the robots.txt text.
What Additional Data Would Answer These Questions
- Server access logs showing actual crawler hits, response codes, and originating IP ranges.
- IP-verified crawler tests (e.g., hitting the site from known Googlebot IPs or using a residential proxy).
- Cloudflare WAF rules audit to see what's triggering the datacenter IP blocks.
- Search Console data on crawl availability and indexing.
Key signals
- 85.7% block rate across all probes, with 18 of 21 requests returning 403 from this datacenter IP. No connection errors — blocking is the single dominant signal.
- Browser UA blocked on all 3 paths, confirming the 403s are IP-based, not UA-targeted. This means the AI crawler blocks observed may not apply in production if those bots use non-datacenter IPs.
- PerplexityBot is the only unblocked UA, returning 200 on robots.txt and 404 on the other two paths. Response times (avg 215.7ms vs 62–87ms for blocked bots) suggest it reaches the origin, consistent with an IP-based bypass rather than a UA whitelist.
- robots.txt returns 403 to all UAs except PerplexityBot. If legitimate crawler IPs are also caught by this rule, the site is uncrawlable. Log verification is needed to confirm whether production crawlers are affected.
- All major AI crawlers blocked (GPTBot, Google-Extended, OAI-SearchBot, ClaudeBot, Bytespider) with PerplexityBot as the sole exception. If these blocks are IP-specific and not production-real, the site may still be visible to AI crawlers. If the IP blocking is broad, the site is invisible to every major AI indexer except Perplexity.
Potential featured finding: The IP-based 403 across all UAs (including browser) on all paths, including robots.txt, is the root signal. If the WAF rule responsible for this also catches legitimate crawler IPs, it would explain a complete absence of crawling, indexing, and AI visibility in a single root cause. Verification of production crawler access is the highest-priority next step.