File viewer

synthesis_12082026_0118.md

/app/data/llm/analysis/house730/synthesis_12082026_0118.md

Synthesis: house730.com

Where sources corroborate each other

  • Ahrefs + Brand Radar: the traffic decline is an AI-answer story, not a rankings story. Ahrefs shows organic traffic −30.6% over 36 months while rankings hold (546/1,000 sampled keywords in positions 1–3, 420 in 4–10), search volume is down only 15.5%, and keyword count is flat. Brand Radar shows 80% of AI impression volume (14,190/17,710) sits in Google AI Overviews on 100% transactional Cantonese queries, and Ahrefs confirms 3,024 AI citations with "租屋" (11,000 vol) surfacing as an ai_overview result. Rankings intact + traffic down + heavy citation presence is consistent with AI answers absorbing clicks. This is co-occurrence, not proven causation — GSC CTR-by-query data is needed to confirm.

  • Live checks + Brand Radar + Ahrefs citations: crawler access maps to citation presence, platform by platform. PerplexityBot is the only AI crawler confirmed through the WAF (200 on /robots.txt), and Perplexity holds 500 citations in Ahrefs — outsized for Perplexity's HK market share. OAI-SearchBot, GPTBot, ClaudeBot, and Google-Extended are all 403'd from the probe IP, and ChatGPT appears nowhere in either dataset's platform mix while Gemini is marginal (1,260 vs 14,190 impressions). Google AI Overviews dominates — which itself proves verified Googlebot reaches origin fine.

  • Ahrefs + Brand Radar: the money pages earn citations with zero links. /rent/t1b8/ is simultaneously a top-traffic page with 0 referring domains (Ahrefs) and the #1 own-cited page with 9 responses (Brand Radar). AI citation eligibility on this site is driven by on-page relevance, not authority. This directly undercuts any links-first strategy.

  • Crawl + Ahrefs broken backlinks: configuration hygiene, not content quality, is the onsite theme. The crawl's only scaled signal is 179 /cdn/ asset URLs returning indexable 200s with no evidence of noindex or robots exclusion; Ahrefs' 13 broken backlinks contain %0D%0A line-break characters — a systematic URL formatting bug. Both are plumbing failures, not editorial ones.

Contradictions — and how they resolve

  • Crawl shows 200s; live checks show 403s. Different IPs, different UAs. The crawl's requests passed; the probe's single datacenter IP was WAF-flagged (even the browser got 403). Do not merge these into "the site blocks crawlers."
  • 3,024 AI citations vs "AI crawlers blocked." Both true at once. Citations flow through platforms with verified access (Googlebot → AI Overviews; Perplexity explicitly allowed). The bots blocked in the probe correspond to the platforms where citations are absent or thin. The apparent contradiction is the finding.
  • Brand Radar SOV 1.0 flat vs "#1 of 194 domains." The 1.0 is a single-brand tracker artefact. The real competitive signal is 19% cited-page share with a 2.1x lead over centanet.
  • 875K spam backlinks vs stable top rankings. The January 2026 spam wave (avg DR 6.7, "TELEGRAM @happygrannypies" anchors) has not visibly moved rankings — Google may already be discounting it. Links are not the current bottleneck; DR is inflated but harmless-looking so far.

Strongest cross-source narratives

1. Organic is now the only channel, and it's being squeezed at the click, not the position. Paid collapsed 92.8% (−231K monthly visits, Ahrefs), so every organic loss hits revenue directly. The loss mechanism visible in the data is AI-answer interception of high-intent queries (Ahrefs + Brand Radar), not ranking or demand collapse. No link campaign addresses this.

2. The citation moat is real but shallow, and it's a land grab. #1 share, 2.1x lead, but breadth-over-depth: 169 pages, 188 responses, 1.1 avg, top page only 9/217 (Brand Radar). Competitors are carving out specific clusters — 28hse, spacious, runhotel, weave-living all cited on short-term rental queries where house730 has no depth. Meanwhile the crawl cannot verify the on-page health of the very pages being cited, because 91% of the sample is /cdn/ assets.

3. (Supporting, deprioritised) Authority is misallocated and poisoned at the edges. Money pages have 0 referring domains while estate profiles hoard links (122 RDs, 74 visits); internal linking digest is empty; 13 broken links from DR-75 HK media carry a fixable URL bug. Frame strictly as cleanup/redistribution — the rankings data proves acquisition is not the lever.

Data gaps visible only in combination

  • Brand Radar names the cited pages (/rent/t1b8/, listing/estate pages); the crawl sampled almost none of them. We cannot confirm title/H1/indexation health of the pages driving both traffic and citations. Recrawl excluding /cdn/.
  • No server logs or WAF config access — the production condition of AI crawlers (the pivotal verification) is unknown.
  • Internal linking digest empty — equity flow from estate pages to category pages is unmodelled, right where Ahrefs shows the misallocation.
  • Unknown whether /cdn/ assets send X-Robots-Tag noindex — the 162 "indexable" count can't be converted to actual index pollution without header/robots checks.
  • No GSC CTR data to confirm cannibalisation; no multi-brand SOV or absolute mention trend to measure whether the 2.1x lead is growing or eroding.
  • /services/ and /contact/ 404'd for PerplexityBot — intended or broken? Uncorroborated elsewhere.

Ranking (all sources together)

  1. GEO/AI visibility: verify and fix production AI-crawler access; defend the Google AI Overviews lead; close the short-term-rental citation gap.
  2. Onsite: recrawl without /cdn/; apply noindex/robots controls to asset paths; fix the %0D%0A URL bug (also recovers broken backlinks); audit the cited-page templates.
  3. Internal linking: pull the digest; route estate-page equity to /buy/ and /rent/ category pages.
  4. Backlink hygiene (supporting only): disavow audit on the spam wave; reclaim the 13 broken links. Never the headline.

Featured finding

The platform-by-platform split in AI crawler access matches the platform-by-platform split in citations — and the access-confirmed platforms are exactly where house730 is winning.

Cross-references live checks + Brand Radar + Ahrefs. From the probe IP, five major AI crawlers (GPTBot, ClaudeBot, OAI-SearchBot, Google-Extended, Bytespider) were 403'd at the edge (sub-120ms, never reaching origin) while PerplexityBot fetched /robots.txt with a 200. This proves datacenter-IP blocking, not production blocking — verified crawlers are typically whitelisted. But the citation data shows the same split: Perplexity (confirmed access) holds 500 citations (Ahrefs) despite negligible HK market share; ChatGPT appears in neither dataset; Gemini is marginal at 1,260 impressions vs Google AI Overviews' 14,190 — and AI Overviews runs on Googlebot, whose production access is proven by that dominance.

The mechanism: a platform can only cite pages its crawler can fetch. The two access-confirmed crawlers correspond to the two platforms with meaningful citation presence; the blocked bots correspond to the gaps. This is co-occurrence — alternative explanations exist (Brand Radar may not track ChatGPT; Gemini adoption in HK is low) — which is precisely why server logs and WAF config are the pivotal next step.

Why this matters more than anything else in the data: house730 is already the #1 cited domain with a 2.1x lead and proven citation conversion wherever access exists (Brand Radar). Organic traffic is down 30.6% with AI answers absorbing clicks (Ahrefs), making owned citations the one asset that monetises the AI shift. If any production AI crawler is being 403'd, every other GEO, content, and internal-linking initiative is irrelevant to that platform — and the fix is a WAF rule change, verifiable in week 1, that opens platforms competitors haven't locked up yet.