File viewer

live_checks_12082026_0916.md

/app/data/llm/analysis/bupa/live_checks_12082026_0916.md

Here is the analytical commentary based on the blind HTTP probe data for bupa.com.hk.

The most significant finding isn't a blocking issue—it's a content availability problem. 14 of the 21 probes returned a 404, completely overshadowing the perfect 0% block rate. The site isn't preventing AI crawlers from accessing content; it's failing to serve content at all on two of the three probed paths.

There is no evidence of differential treatment between bots and browser. All 21 probes across 7 user agents were allowed. The digest_browser_vs_bot_delta is empty (count: 0), confirming identical HTTP-level access. The hypothesis that Cloudflare bot-blocking is impacting AI crawler visibility on this site is not supported by this data. Every AI bot tested—GPTBot, ClaudeBot, Google-Extended, OAI-SearchBot, PerplexityBot, Bytespider—returned the exact same status codes as the browser, path for path.

The 404 errors on /contact/ and /services/ are the dominant concern. This is not a transient error; the pattern is absolute. All 7 user agents received a 404 on both /contact/ (browser elapsed: 412ms) and /services/ (browser elapsed: 448ms). These URLs either never existed at these paths, have recently moved without proper redirects, or are being served by a misconfigured platform. If these pages are intended to be live content, their unavailability means they are invisible to every crawler tested, rendering any discussion of AI bot permissions moot.

The /robots.txt file is accessible and healthy, returning a 200 to every agent. Response times for this file were among the fastest overall (average ~94ms), which is normal for a static file. This confirms the server is operational and capable of serving 200s, sharpening the contrast with the 404 paths.

Response times show a clear pattern where 404s cost more than 200s. Bots took longer to receive a 404 on missing pages than a 200 on the existing robots.txt. For example, OAI-SearchBot averaged 79ms for robots.txt (200) but spiked to 369ms average across the two 404 paths. The browser shows an even starker gap: 115ms for the 200, versus 412–448ms for the 404s. This is consistent with a CMS or application layer performing a full page lookup before determining the content is missing, rather than serving a cached 404 quickly. It could point to an absence of page caching or a slow backend that impacts even non-existent paths.

A hypothesis to verify is that the /contact/ and /services/ URLs were selected for probing based on an assumed generic URL structure, and that the actual live content exists under different, locale-specific or CMS-driven paths (e.g., /en/contact/ or /health-services/). If these are not the correct URLs, the 100% 404 rate tells us nothing about content visibility—only that URL discovery is a challenge.

The key limitation of this data is its scope. Only 3 paths were tested from a single datacenter IP. A 0% block rate here does not prove that production search engine crawlers (Googlebot, Bingbot) are never blocked, as they originate from a different set of IPs that Cloudflare will typically whitelist. We can only state that the tested User-Agent strings are not filtered at the HTTP level.

Key signals - 14 of 21 probes returned 404, accounting for 100% of probes on /contact/ and /services/. Content is absent, not blocked. - 0.0% block rate across 7 user agents, including all AI crawlers. No differential treatment is detected. - Browser vs. bot differential count is 0. The access behaviour is completely uniform. - 404 response times for the browser (412ms, 448ms) are significantly higher than the 200 for robots.txt (115ms), suggesting an uncached application-layer lookup that fails slowly. - Potential featured finding: The 100% 404 rate on tested content pages is an availability failure that renders AI crawling permissions irrelevant for these paths until corrected or the correct URLs are identified.