File viewer
draft_13082026_0133.json
/app/data/llm/draft/firstpage/draft_13082026_0133.json
{
"summary_points": [
"The site is selectively unreachable. Live checks timed out on all 21 probes from a datacenter IP across 7 user agents, while the crawl retrieved only 1 of ~34,000 known URLs.",
"AI citation is zero. Brand Radar shows 0.0% share of voice, 0 cited pages, and 0 cited domains across ChatGPT, Gemini, Perplexity, Copilot, and Grok.",
"The crawl's homepage 200 proves the site is up, not down—the blockage is vantage-dependent, not an outage.",
"No user-agent differential was observed in live checks; browser, GPTBot, and others all timed out identically, indicating IP-level filtering rather than UA-based bot management.",
"Standard SEO tooling cannot see this site, and prior reporting on on-site health was based on a single URL sample—not a representative audit.",
"The diagnostic blackout is itself the headline finding: no template, indexation, or internal-link analysis is possible until access is resolved.",
"The site is invisible in both AI and traditional crawler vantages—not a content or link issue, but an infrastructure access one.",
"Each of the three data sources points to the same root cause, but none alone can confirm it; only server or CDN logs can prove production crawler block rates."
],
"method_notes": "Three independent sources were used: live probes (HTTP/1.1, HTTP/2, browser-like requests from a single AWS IP to 3 paths × 7 UAs), a crawl (axios with known URL list, returned 1 of ~34,000 pages), and Brand Radar (AI platform citation and share-of-voice tracking). The crawl and probes disagree on homepage reachability (crawl: 200, probes: timeout), which resolves to a site-up-but-filtered conclusion when read together. No multi-location probe was run. No server logs were accessed. This is a first-pass diagnostic, not a root cause confirmation.",
"featured_finding": {
"headline": "Firstpage.com.hk is up but selectively unreachable — and the same access wall that blocked the crawl is consistent with a total absence from AI answers.",
"narrative": "The site responds (crawl: homepage 200) but is unreachable from at least one external datacenter vantage (21/21 live-probe timeouts). All probes—standard browser, GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended, and Bytespider—timed out at a uniform ~5,015ms ceiling, with no UA differential. This is the signature of IP- or reputation-based edge filtering, not UA-selective bot management. The same access pattern is the most parsimonious explanation for the site's complete absence from AI citation: Brand Radar shows 0.0% share of voice, zero cited pages, and zero cited domains across ChatGPT, Gemini, Perplexity, Copilot, and Grok. If AI crawlers' fetch infrastructure also lands on filtered IP ranges—or if robots.txt remains unreachable to them—citation becomes impossible regardless of content quality. The crawl's 1/34,000 URL retrieval further confirms that standard automated SEO tooling cannot profile the site. This finding is corroborated by all three sources, blocks every downstream diagnosis, and is verifiable by the client in week 1 via CDN/WAF logs.",
"proof_points": [
"Live probes: 21/21 timeouts from a datacenter IP across 3 paths × 7 UAs, all failing at ~5,015ms.",
"Brand Radar: 0.0% share of voice, 0 cited pages, 0 cited domains across ChatGPT, Gemini, Perplexity, Copilot, and Grok.",
"Crawl: 1 of ~34,000 known URLs retrieved; homepage returned 200, proving the site is up.",
"No UA differential was observed: browser and AI-crawler UAs failed identically.",
"robots.txt was never successfully retrieved by any source."
],
"why_it_wins": "This is the single most important finding because it is the only one corroborated by all three sources; it blocks every other line of onsite diagnosis (templates, indexation, internal linking); and it plausibly explains the site's total AI invisibility. It is also the most actionable—the client can verify or disprove it by checking CDN/WAF logs and running a multi-location probe in week 1."
},
"key_findings": [
{
"finding": "The site is accessible from some vantages but completely blocked from others.",
"evidence": "The crawl retrieved the homepage with a 200 status code, but all 21 live-probe requests from a single datacenter IP timed out. The probes covered 3 paths (/, /hk/, /hk/seo/) and 7 user agents including browser, GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended, and Bytespider—every request died at the same ~5,015ms ceiling.",
"implication": "The site is not down; it is filtering incoming traffic based on requester reputation or IP range. This explains why standard SEO tooling (which typically runs from datacenter IPs) cannot profile the site, and it suggests configuration-level blocking rather than a server issue."
},
{
"finding": "AI citation is zero across all major platforms.",
"evidence": "Brand Radar shows 0.0% share of voice, 0 cited pages, and 0 cited domains for each of ChatGPT, Gemini, Perplexity, Copilot, and Grok. The site does not appear in any AI-generated answer source.",
"implication": "The site is invisible to AI audiences. Whether this is because AI crawlers are blocked (consistent with the probe findings) or because the site lacks the authority/dataset presence to be cited is currently unknown—but the most parsimonious explanation given the probe data is access filtering."
},
{
"finding": "No template, indexation, or internal-linking analysis is possible with current data.",
"evidence": "The crawl only retrieved 1 of approximately 34,000 known URLs—the homepage. Brand Radar provides no content-level signals. Live probes confirmed robots.txt was unreachable. No page template, no internal link graph, no crawl depth data exists.",
"implication": "Any prior reporting of on-site health was based on a single URL sample and is not representative. Until access is resolved, every downstream onsite finding is speculative."
},
{
"finding": "The access blocking is not user-agent selective.",
"evidence": "In live probes, all 7 user agents—including a standard browser (Chrome 120)—timed out identically at ~5,015ms. No UA-based bot detection or allowlist was observed. The behavior is consistent with IP-, ASN-, or reputation-based edge filtering (e.g., Cloudflare Security Level or WAF rule).",
"implication": "This is not a bot management policy that can be fixed by updating robots.txt or tweaking a bot allowlist. The blocking is at the network edge and affects all traffic from the tested datacenter range, including human browser traffic."
},
{
"finding": "The site's zero AI citation cannot be attributed to content quality or authority with current data.",
"evidence": "Brand Radar lists three possible causes for zero citation: content absent from datasets, AI crawlers blocked, or authority too low. Only the second (blocked crawlers) is supported by the live-probe data. No authority or backlink data was collected; no content audit was possible.",
"implication": "Fixing access is the prerequisite for diagnosing any other contributor to AI invisibility. There is no evidence of a link or content problem, and no such analysis can be performed until the crawl works."
}
],
"risk_areas": [
{
"issue": "Complete AI citation gap — 0.0% share of voice across all 5 platforms.",
"priority": "P0",
"effort": "LOW",
"impact": "HIGH",
"detail": "Brand Radar confirms zero cited pages and zero cited domains for ChatGPT, Gemini, Perplexity, Copilot, and Grok. The most likely root cause (edge filtering of crawler IP ranges) is also the cheapest to verify: check Cloudflare/WAF logs for AI-crawler request history."
},
{
"issue": "Crawl failure — only 1 of ~34,000 URLs retrieved.",
"priority": "P0",
"effort": "MEDIUM",
"impact": "HIGH",
"detail": "The crawl retrieved only the homepage. All other URLs timed out. This blocks all onsite technical analysis. Resolution requires either obtaining server logs to confirm the blocking pattern or running the crawl from a residential/mobile IP."
},
{
"issue": "Robots.txt unreachable from datacenter vantages.",
"priority": "P1",
"effort": "LOW",
"impact": "HIGH",
"detail": "robots.txt was never successfully retrieved by any probe or crawl source. If it contains disallow rules for AI crawlers, those rules cannot be verified. The file may be accessible from residential IPs but blocked at the edge."
},
{
"issue": "Single-vantage probe — scope of access restriction unknown.",
"priority": "P1",
"effort": "LOW",
"impact": "MEDIUM",
"detail": "All probes originated from one AWS datacenter IP. It is unknown whether the blocking affects all datacenter ranges, only AWS, or only a specific ASN. A multi-location probe (AWS, GCP, Azure, residential proxies) is needed to map the blockage."
}
],
"bot_management": {
"observed_facts": [
"All 7 user agents (Chrome 120, GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended, Bytespider) timed out at ~5,015ms from a single AWS datacenter IP.",
"No HTTP status code was returned; connections were silently dropped.",
"robots.txt was never successfully retrieved by any UA or path.",
"The crawl (using a different IP range/datacenter from a known bot list) retrieved the homepage with 200 but timed out on all other ~34,000 URLs."
],
"seo_risks": [
"If verified AI crawlers (e.g., Google-Extended on its production IPs) are also blocked, the zero AI citation is directly caused by blocking rather than content quality—and no robots.txt fix alone will resolve it.",
"If traditional crawlers like Googlebot are blocked on some IP ranges, indexation gaps may exist that are invisible to GSC data.",
"The site may appear inconsistently indexed depending on the crawl source vantage."
],
"geo_hypothesis": {
"claim": "The edge filtering that blocks datacenter IPs also intercepts production AI-crawler traffic, producing the zero-citation result observed in Brand Radar.",
"supporting": [
"Live probes showed that all AI-crawler UAs (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot) timed out identically with browser from the same datacenter IP.",
"Brand Radar confirms zero cited pages and zero cited domains for every major AI platform.",
"The crawl's homepage 200 vs. probe timeout shows the site is reachable from some vantages but not others—consistent with IP-based edge filtering."
],
"prevents_calling_fact": [
"The probes used one datacenter IP only; verified production AI-crawler IP ranges are unknown and may be whitelisted by Cloudflare.",
"No server/CDN logs were reviewed to confirm real crawler request history.",
"Cloudflare typically whitelists verified bots by signature, and those signatures may bypass the edge rules that drop unverified datacenter traffic."
],
"definitive_check": "Review Cloudflare WAF or CDN event logs to confirm whether verified AI crawlers (e.g., OpenAI's OAI-SearchBot, Google-Extended, PerplexityBot) receive 200 OK or are blocked. Server-side access logs (e.g., Apache/Nginx combined logs) can confirm crawler requests reaching origin."
},
"recommendations": [
"Request Cloudflare WAF event logs for the past 90 days. Filter by bot UA patterns: GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended, Bytespider. Count requests by status code (200 vs. 403 vs. timeout).",
"If Cloudflare is not in use, request server-level access logs (Apache/Nginx) for the same period and same UA filters.",
"Run a multi-location probe using 3+ cloud providers (AWS, GCP, Azure) and 2 residential proxies to map the scope of the blockage.",
"Check GSC 'Crawl Stats' report for any spikes in 403/500 errors or drops in pages crawled per day."
],
"method_caveat": "The live probe was executed from a single AWS datacenter IP using HTTP/1.1 and HTTP/2 requests with custom UA headers. This does not reproduce production crawler behavior (cloud providers verify bot signatures, not just UAs). Connection timeouts at 5,015ms indicate network-layer filtering (likely Cloudflare Security Level or WAF challenge) rather than application-layer blocking. The test cannot distinguish between rate limiting, IP block, ASN block, or geo-block. Results are UA- and vantage-confounded."
},
"narrative_arc": [
"Firstpage.com.hk is inaccessible to standard SEO tooling and completely absent from AI answers—not because of content or links, but because its edge configuration blocks the very IP ranges that crawlers and tools use.",
"Three independent data sources converge on this single finding: live probes from a datacenter IP returned 21/21 timeouts; a crawl retrieved only 1 of ~34,000 known URLs; and Brand Radar shows zero AI citations across every major platform.",
"This access wall makes everything else unmeasurable—templates, indexation, internal linking, content quality. Until it is resolved, no onsite diagnosis is possible and the site remains invisible to AI audiences.",
"The fix is verifiable in week 1: check CDN/WAF logs for production crawler access, run a multi-location probe, and review GSC crawl stats. The root cause is technical and configurable, not a content or link deficit."
],
"recommendations": [
{
"action": "Obtain and review server-side or CDN/WAF event logs for bot traffic.",
"rationale": "Live probes from one datacenter IP cannot confirm production crawler blocking—only logs can show whether verified AI crawlers (GPTBot, OAI-SearchBot, PerplexityBot, etc.) receive 200 OK or are blocked. This gates every downstream decision.",
"validation": "Confirm in GSC and server logs after signature: look for 403/503 rates, crawl volume trends, and bot-specific status codes."
},
{
"action": "Run a multi-location probe (3+ cloud providers + 2 residential IPs) to map the scope of the access restriction.",
"rationale": "The current probe used a single AWS IP. Without knowing whether the blockage affects all datacenter ranges, only AWS, or a specific ASN, the remediation scope is unknown.",
"validation": "After signature, share multi-location probe results: which vantages can reach the site and which cannot."
},
{
"action": "Review and adjust edge security/WAF settings to allow verified bot traffic (Googlebot, Bingbot, GPTBot, OAI-SearchBot, PerplexityBot, etc.).",
"rationale": "If the edge filtering is blocking production crawlers (even unintentionally), AI citation will remain zero regardless of content investment. Cloudflare and other WAFs can whitelist verified bots by signature.",
"validation": "Confirm via fresh probe or log review after WAF update: AI-crawler UAs from verified IP ranges should return 200 OK."
},
{
"action": "Re-run a full crawl from a residential or non-datacenter IP after access is confirmed.",
"rationale": "The current crawl retrieved 1 of ~34,000 URLs, making onsite analysis impossible. A representative crawl is the prerequisite for template, indexation, and internal-linking audits.",
"validation": "After signature, run crawl from residential IP; target 5,000+ URLs with status code distribution and content extraction."
}
],
"hypotheses": [
{
"claim": "The edge filtering that blocks datacenter IPs also intercepts production AI crawlers, causing the zero-citation observation.",
"status": "suggestive",
"how_to_verify": "Review CDN/WAF logs for verified AI-bot UAs and confirm whether they receive 200 OK or are blocked."
},
{
"claim": "The crawl failure was not due to blocking but to a misconfigured URL list (limited to the homepage only).",
"status": "unverified",
"how_to_verify": "Request the crawl input list and re-run against a known working URL set, or confirm via server logs that requests for non-homepage URLs reached origin."
},
{
"claim": "The blocking is limited to a single AWS IP range and does not affect other cloud providers or residential IPs.",
"status": "unverified",
"how_to_verify": "Run probes from GCP, Azure, and two residential proxy IPs. Compare timeout patterns."
},
{
"claim": "The site has no content-quality or authority barrier to AI citation—the zero-citation result is purely an access problem.",
"status": "unverified",
"how_to_verify": "Only testable after confirming AI crawlers can access the site. If access is confirmed and citation remains zero, proceed to content and authority audit."
}
]
}