File viewer

live_checks_12082026_0118.md

/app/data/llm/analysis/house730/live_checks_12082026_0118.md

The probe data from a single datacenter IP paints a clear picture: house730.com is aggressively blocking almost everything from this location. 18 of 21 requests returned 403, including the browser. PerplexityBot is the sole exception — allowed through on all three paths.

The block is not driven by user-agent. The browser's 403 on /contact/, /robots.txt, and /services/ confirms this is an IP-level restriction, almost certainly imposed by a CDN/WAF (likely Cloudflare) that flagged the source IP as hostile or high-risk. All other tested agents — Bytespider, ClaudeBot, GPTBot, Google-Extended, OAI-SearchBot — are swept up in the same blanket block. Their response times are consistently fast (55–119ms, with one 318ms outlier for Google-Extended on /contact/), consistent with edge blocking that never reaches origin.

PerplexityBot stands alone. It received a 200 on /robots.txt (135ms) and 404s on /services/ and /contact/ (190ms, 201ms respectively). The 200 proves the bot can fetch the robots.txt file; the 404s indicate those specific paths don't exist, but the bot was permitted to check. This asymmetry suggests Perplexity's crawler infrastructure is either explicitly whitelisted in the WAF configuration, or it operates from an IP range that isn't classified as malicious.

From this data alone, I wouldn't conclude that search engine bots (Googlebot, Bingbot) are blocked in production. Those crawlers typically originate from IP ranges that Cloudflare automatically whitelists, and this datacenter IP is not representative of a search engine's official IP pool. However, the data does flag a real risk: if any legitimate crawler uses an IP that falls into the same blocked pool, it will be silently dropped. That's a hypothesis worth testing with log data — see whether major search bots are hitting 403s or 200s in practice.

The 404 responses for PerplexityBot on /services/ and /contact/ are worth noting. If those URLs are supposed to exist, it's a content issue. If they're intentional dead ends, fine — but they're included in the probe set, so the consultant should verify that these paths are meant to be crawlable. The /robots.txt returning 200 is a positive signal: the bot can read directives, which matters for any whitelisted AI crawler.

The response time differential between blocked and allowed requests is clear and reinforces the edge-blocking pattern: blocked requests resolve in <120ms on average; PerplexityBot's allowed requests take 135–201ms, suggesting they hit origin and generated a response. The browser's 391ms on /robots.txt (blocked) is an outlier possibly caused by a slow 403 page generation — not a concern.

Overall, the most material finding is the combined total block of five major AI crawlers (including Google-Extended, which feeds Gemini, and GPTBot/ClaudeBot for their respective products) alongside a full pass for PerplexityBot. If the client cares about visibility in AI-generated answers, this is a direct, high-impact finding. The next step is to reconcile with live crawl data and server logs to see whether the same blocking pattern holds across the full production IP pool.

Key signals
- From this datacenter IP, 18 of 21 requests (85.7%) are blocked, including a standard browser — an IP-level blanket restriction, not UA targeting.
- PerplexityBot is the only agent allowed through (3/3, 100%), while Bytespider, ClaudeBot, GPTBot, Google-Extended, and OAI-SearchBot are all 100% blocked.
- Blocked requests are resolved at the edge (fast 403s under 120ms median); allowed PerplexityBot requests reach origin (135–201ms, returning actual statuses 200/404).
- PerplexityBot's 200 on /robots.txt confirms the file is fetchable; the 404s on /services/ and /contact/ indicate either dead paths or content issues that would prevent indexing even for an allowed bot.
- Potential featured finding: The blanket block suppresses 5 of 6 AI crawlers tested, consistent with low or zero visibility in AI search results outside Perplexity. This data does not prove production search engine blocking but strongly indicates an IP-level WAF rule that, if not carefully scoped, could silently block legitimate crawlers.