File viewer

synthesis_12082026_0142.md

/app/data/llm/analysis/house730/synthesis_12082026_0142.md

Featured finding

A datacenter-IP blocking rule at the edge is simultaneously blinding house730's measurement stack and potentially locking out every major AI crawler — and the one bot that gets through is the one non-Google platform where citations actually register.

Three sources triangulate this independently:

  • Live checks: 18 of 21 probes return 403 from a datacenter IP — including the browser UA and robots.txt itself. Fast 403s (62–87ms) mean rejection at the edge, not the origin. The sole exception is PerplexityBot, whose slower responses (avg 215.7ms, 200 on robots.txt) show it reaching the origin.
  • Ahrefs: digest_authority reports 0 backlinks, 0 refdomains; digest_internal_linking returns 0; digest_branded_vs_generic classifies 0 of 1,000 keywords as branded — while "house730" is the #1 keyword by traffic. Other Ahrefs endpoints return rich historical data. This is exactly the fingerprint of AhrefsBot (a datacenter crawler) being blocked from recrawling while stale indexes persist.
  • Brand Radar + Ahrefs AI citations: Perplexity — the only crawler confirmed to reach the origin — is also the only non-Google AI platform with meaningful citation presence (500 citations in Ahrefs). GPTBot, OAI-SearchBot, ClaudeBot, Bytespider, and Google-Extended all took 403s from the probe IP.

Mechanism: an IP-based WAF rule rejects datacenter traffic at the edge → third-party crawlers can't recrawl → Ahrefs summary endpoints compute impossible zeros → and any AI crawler whose retrieval infrastructure shares those IP ranges never reaches the content.

Critical caveat: this does not prove Googlebot, Bingbot, or any production AI crawler is blocked. Google's pipeline clearly works (151K monthly organic visits, AI Overviews dominance). The open question — answerable only in server logs and GSC — is whether the rule that catches AhrefsBot and datacenter probes also catches production AI crawler IPs.

Why this matters more than anything else: every other finding depends on it. You cannot trust the backlink data, the internal-linking data, or the branded-split data until access is fixed — and the GEO strategy (below) cannot expand beyond Google until you know whether GPTBot and OAI-SearchBot can reach the site. It is a configuration fix, verifiable in week 1 with server logs and GSC crawl stats — not a 12-month program.


Narrative 2: GEO visibility is real but structurally fragile — one platform, one content type

  • Brand Radar: 80% of AI impression volume (14,190 of 17,710) comes from Google AI Overviews alone, across just 40 questions. Ahrefs corroborates: Google's AI surfaces account for the bulk of 3,024 citations (1,609 AI-overview keyword citations across 1,040 pages).
  • Both sources agree citations are listing-driven, not authority-driven: Brand Radar's top-cited page is /rent/t1b8/ (9 responses); editorial content is nearly invisible (one /property-tips/article/1963/ page, 2 responses). Own-pages share of cited content is only 19% (160 of 844 pages) despite house730 leading all domains (188 responses vs Centanet's 89).
  • Competitors already out-cite house730 on the most contested topic — short-term rentals (28Hse page with 7 responses, Spacious blog with 5, Weave Living 4). That is a concrete content-gap target.
  • Single-platform dependency + single content type = a channel one algorithm shift away from erosion. The fix (Cantonese district-pricing explainers, short-term rental guides) is also the diversification play.

Narrative 3: The traffic decline is a channel shift, not a rankings or authority failure

  • Ahrefs: rankings are dominant — 546 of 1,000 keywords in positions 1–3, only 34 in 11–20; "居屋" (pos 2) and "租屋" (pos 3) both carry AI Overview appearances. Yet traffic fell 30.7% while tracked search volume fell only 15.9%.
  • Brand Radar: every top-volume AI question is a Google AI Overview on exactly these HK property-pricing queries.
  • Together this is consistent with AI Overviews absorbing clicks on queries house730 still wins — the visibility moved into the AI answer, where house730 is cited but not clicked. This reframes the engagement: the client does not need "more rankings" or "more links"; they need to own and monetise the AI answer layer.

Contradictions to manage

  • Crawl vs. everything else: the crawl shows a Shopify architecture (/products/, /collections/, /cdn/, /checkouts/); Ahrefs and Brand Radar show a property portal (/rent/, estate profiles, /property-tips/). Either the crawl hit the wrong seed or a Shopify-hosted section. It cannot diagnose the money pages.
  • Crawl 200s vs. live-check 403s: different source IPs, different conditions — consistent with IP-based rules that let the crawl through while rejecting the datacenter probe. Do not merge into one "site is blocked" story.
  • Brand Radar SOV 1.0 vs. 194 cited domains: the SOV metric is unusable (tracked_brands misconfigured); the cited-domain data (188 vs. Centanet 89, 28Hse 49) is the real competitive picture. Verify the tracked set before the proposal.
  • Spam attack vs. decline timeline: the coordinated spam wave (identical first-seen timestamps 2026-01-31, "TELEGRAM @happygrannypies" anchors) postdates the traffic decline that began in 2024 — it cannot be the cause. If a prior vendor blamed links, the data says otherwise. Secondary effect worth noting: the same crawler-blocking that zeros Ahrefs also blinds the client's ability to detect this attack.
  • Perplexity: 500 citations (Ahrefs) vs. 320 impressions / 9 questions (Brand Radar): citations without modelled traffic — consistent with confirmed access (live checks) but a small HK user base.

Data gaps visible only in combination

  • No valid onsite data exists for the pages that drive the business. The crawl sampled 18 Shopify-flavoured pages; the pages AI cites and users land on (/rent/, listing/estate URLs) were never crawled. Combined with digest_internal_linking returning 0, internal architecture is a complete blind spot. A clean re-crawl (exclude /cdn/, seed portal paths) is mandatory before any template claim.
  • No production crawler truth. Live checks are one datacenter IP; Ahrefs zeros are circumstantial. Only server logs + GSC resolve whether real AI crawlers are blocked.
  • /services/ and /contact/ 404 even for PerplexityBot — low severity, likely wrong probe paths, but worth confirming.

Ranked priorities (all sources considered)

  1. WAF/bot-access verification and fix (live_checks + Ahrefs + Brand Radar) — featured finding; unlocks data trust and AI-crawler access; verifiable in week 1.
  2. GEO diversification (Brand Radar + Ahrefs) — defend the 14,190-impression Google AI Overviews channel and close the short-term-rental and informational content gaps where 28Hse/Spacious/Weave already win; Cantonese pricing-intent only (991/1,000 keywords zh-HK).
  3. Clean onsite crawl of the actual portal (crawl + Ahrefs + Brand Radar) — currently zero onsite visibility into cited/traffic pages.
  4. Backlink cleanup only, never acquisition (Ahrefs) — monitor/disavow the Jan-2026 spam wave; fix the 13 broken am730 links (DR 75, %0D%20 template error). Rankings prove links are not the bottleneck.

Next step requiring signature: server log access and GSC verification — the two data sources that convert the featured finding from "strongly corroborated" to "confirmed," in week 1.