synthesis_13082026_0133.md
/app/data/llm/analysis/firstpage/synthesis_13082026_0133.md
Synthesis — firstpage.com.hk
Where the sources corroborate
- Live checks + crawl: Two independent measurement systems failed to retrieve the site from datacenter infrastructure. The crawl fetched 1 of ~34,000 known URLs; the probes got 21/21 timeouts across 3 paths and 7 user agents. Different failure signatures (partial vs. total), but both are consistent with edge-level filtering of non-residential IP ranges.
- Crawl + live checks, read together: The crawl got a 200 on the homepage while the probes couldn't retrieve even the homepage. Read naively these contradict; read together they prove the site is up but vantage-dependent — it answers some requesters and silently drops others. That is the signature of IP/reputation-based edge filtering, not a server outage and not UA-based bot management (the probes showed zero UA differential; browser and GPTBot failed identically).
- Brand Radar + live checks: Brand Radar shows zero cited pages and zero cited domains across ChatGPT, Gemini, Perplexity, Copilot, and Grok. The probes show AI-crawler UAs (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended, Bytespider) timing out — alongside everything else. The AI invisibility co-occurs with an access pattern that drops datacenter traffic. Whether production AI crawlers are actually affected is unproven (Cloudflare whitelists verified bots; the probe was one datacenter IP), but the observed zero-citation pattern is exactly what that condition would produce.
Where the sources contradict or complicate
- Crawl (homepage 200) vs. live checks (100% timeout): Not a single condition — different IPs, different vantage. The correct reading is selective reachability, not outage. Do not merge these into "the site blocks crawlers."
- Crawl's "bot-blocking" hypothesis vs. live checks' "no UA differential": If the crawl failed due to blocking, it was IP-based, not UA-based. The alternative explanation — a misconfigured sample list containing only the homepage — remains open and is cheap to rule out.
- Brand Radar
digest_sov_history= flat 1.0 vs. zeros everywhere else: Data artefact. Effective share of voice is zero; the history metric should be discarded.
Narratives that only emerge in combination
- Up but invisible. The site responds (crawl: 200, indexable homepage) yet is unreachable from at least one external datacenter vantage (live checks) and completely absent from every AI answer engine (Brand Radar). No single source shows this; only the triangulation does.
- The diagnostic blackout is itself the finding. Standard SEO tooling cannot see this site, and AI platforms demonstrably do not cite it. Every downstream question — templates, indexation, internal linking, content quality — is currently unmeasurable. Any prior reporting of on-site "health" was reporting on a 1-URL sample.
- Zero AI footprint, cause undetermined but technical in shape. Brand Radar's three candidate explanations (content absent from datasets, AI crawlers blocked, authority too low) are all distribution/technical hypotheses — none of the three sources surfaces any evidence of a content or links problem. There is no link-centric story in this data.
Data gaps visible only in combination
- Server/CDN logs (the decisive dataset): The only way to confirm whether verified Googlebot, Bingbot, and AI crawlers get through in production. Probes cannot answer this; logs can.
- robots.txt content: Never successfully retrieved by any source. AI-bot disallow rules are unverifiable.
- Crawl input list: Was the 1-URL crawl a configuration error or real blocking? Unknown.
- Multi-location probes: The scope of the drop (one IP range? all datacenters? geoblock?) is unmeasured.
- GSC access: Production crawl stats and indexation are unobserved.
- No authority/backlink data exists in this engagement — Brand Radar's "low authority" hypothesis is currently untestable. Supporting context only; not the bottleneck on present evidence.
Ranked priorities
- Resolve edge accessibility and obtain logs — gates both onsite diagnosis and the GEO question. (Onsite + GEO, corroborated by all three sources.)
- Confirm production AI-crawler access via logs — determines whether the zero AI share of voice is an access failure or a dataset/authority absence. (GEO.)
- Re-run a representative crawl — unlocks template, indexation, and architecture analysis that is currently impossible. (Onsite.)
- Authority baseline — only after access is fixed; no evidence currently implicates links.
Featured finding
Firstpage.com.hk is up but selectively unreachable — and the same access wall that blinded the crawl is consistent with the site's total absence from AI answers.
- Live checks: 21/21 probes timed out from a datacenter IP — 3 paths, 7 user agents including a standard browser — with robots.txt itself unreachable and all responses dying at a uniform ~5,015ms ceiling. A silent drop at the network edge, not UA-selective bot management.
- Crawl: 1 of ~34,000 known URLs retrieved. The homepage's 200 proves the site responds; the near-total coverage failure proves automated access from standard tooling effectively fails.
- Brand Radar: 0.0% share of voice, zero cited pages, zero cited domains across ChatGPT, Gemini, Perplexity, Copilot, and Grok — the exact downstream signature expected if AI platforms cannot retrieve Firstpage content.
Mechanism (hypothesis, unconfirmed): if the edge filtering that drops datacenter IPs also intercepts AI crawlers' fetch infrastructure — or leaves robots.txt unreachable to them — citation becomes impossible regardless of content quality, producing the uniform zeros observed. The crawl's 200 vs. the probes' total timeout shows access is vantage-dependent, which is what IP/reputation-based filtering looks like. Whether verified production crawlers are affected cannot be determined from probes alone; only server or CDN logs can answer that.
Why this matters more than anything else in the data: it is the only finding corroborated by all three sources; it blocks every other diagnosis (no template, indexation, or internal-linking analysis is possible until the crawl works); and it plausibly gates the entire AI-visibility channel, which currently sits at zero. It is also verifiable by the client's own team in week 1 — pull the CDN/WAF firewall event logs, run a multi-location probe, and check GSC crawl stats. Nothing else in this dataset can be fixed, or even measured, until access is resolved.