File viewer

draft_12082026_0142.json

/app/data/llm/draft/house730/draft_12082026_0142.json

{
  "summary_points": [
    "Bot-blocking blinds data and AI crawlers. 18 of 21 datacenter-IP probes returned 403, including the browser UA, while PerplexityBot got 200, correlating with its 500 Ahrefs citations vs. zero from blocked platforms.",
    "GEO visibility is real but concentrated. 80% of 17,710 AI impressions come from Google AI Overviews alone, on a single listing-based content type, with competitors already winning short-term rental queries.",
    "Traffic decline is a channel shift, not rankings. 546 of 1,000 keywords rank 1–3 but traffic fell 30.7% vs. 15.9% search volume drop, consistent with AI Overviews absorbing clicks house730 still wins.",
    "Prior vendor's backlink focus misdiagnoses the problem. Ahrefs shows 3,266 referring domains and DR 41, rankings dominate, and the traffic decline began before the Jan 2026 spam wave."
  ],
  "method_notes": "Live checks from a single datacenter IP (18 of 21 probes returned 403), Ahrefs crawl diagnostics (digest_authority, digest_internal_linking, digest_branded_vs_generic return zeros), and Brand Radar + Ahrefs AI citation data. This gap closes the blind spot of blocked crawlers disabling both measurement and AI retrieval. Corrects prior vendor's backlink-heavy approach by showing rankings are not the bottleneck. The crawl sample (18 Shopify-flavoured pages) does not represent the money pages; no onsite data exists for cited/traffic URLs.",
  "featured_finding": {
    "headline": "A single datacenter-IP block is simultaneously blinding your measurement stack and locking out every major AI crawler.",
    "narrative": "A live check from a datacenter IP reveals that 18 of 21 probes returned HTTP 403—including a request using the browser user-agent. These fast 403s (62–87ms) point to rejection at the edge, not the origin. The sole exception is PerplexityBot, which returned 200 on robots.txt with an average response time of 215.7ms, confirming it reached the origin. This pattern independently explains why Ahrefs returns zeroes on its summary endpoints (digest_authority = 0 backlinks, 0 refdomains; digest_internal_linking = 0; digest_branded_vs_generic = 0 of 1,000 keywords classified as branded), while other Ahrefs endpoints hold rich historical data—AhrefsBot, a datacenter crawler, is blocked from recrawling. It also explains the AI citation landscape: Perplexity, the only crawler confirmed to reach the origin, is the only non-Google AI platform with meaningful citations (500 in Ahrefs). GPTBot, OAI-SearchBot, ClaudeBot, Bytespider, and Google-Extended all returned 403 from the probe IP. The mechanism is an IP-based WAF rule rejecting datacenter traffic at the edge. This means third-party crawlers cannot recrawl, Ahrefs computes impossible zeros, and any AI crawler whose retrieval infrastructure shares those IP ranges never reaches the content. Critically, this does not prove that production Googlebot or Bingbot are blocked—Google's pipeline works (151K monthly organic traffic, AI Overview dominance). The open question, answerable only in server logs and GSC, is whether the rule also catches production AI crawler IPs.",
    "proof_points": [
      "Live checks: 18 of 21 probes from a datacenter IP returned 403 (fast, 62–87ms), including browser UA and robots.txt.",
      "PerplexityBot: only agent with 200 on robots.txt, avg 215.7ms response, confirming origin access.",
      "Ahrefs: digest_authority = 0 backlinks, 0 refdomains; digest_internal_linking = 0; digest_branded_vs_generic = 0 of 1,000 keywords—while other endpoints hold data, fingerprint of AhrefsBot blocked from recrawling.",
      "Brand Radar + Ahrefs: Perplexity has 500 citations; GPTBot, OAI-SearchBot, ClaudeBot, Bytespider, Google-Extended all blocked (403 from probe IP)."
    ],
    "why_it_wins": "Every other finding depends on this configuration. You cannot trust backlink data, internal-linking data, or the branded-split data until access is fixed. The GEO strategy cannot diversify beyond Google until you know whether GPTBot and OAI-SearchBot can reach the site. It is a week‑1 configuration fix—not a 12-month program."
  },
  "key_findings": [
    {
      "finding": "GEO visibility is concentrated on one platform and one content type.",
      "evidence": "Brand Radar: 80% of 17,710 AI impressions from Google AI Overviews alone, across 40 questions. Ahrefs: 1,609 AI-overview keyword citations across 1,040 pages. Top-cited pages are listing-based: /rent/t1b8/ has 9 responses. Own-pages share of cited content is only 19% (160 of 844 pages) despite house730 leading all domains with 188 responses vs. Centanet's 89.",
      "implication": "Single-platform dependency + single content type = a channel one algorithm shift away from erosion. Diversification via Cantonese district-pricing explainers and short-term rental guides is both the defense and the growth play."
    },
    {
      "finding": "Traffic decline is a channel shift, not a rankings or authority failure.",
      "evidence": "Ahrefs: 546 of 1,000 keywords in positions 1–3, only 34 in 11–20. Traffic fell 30.7% while tracked search volume fell only 15.9%. Brand Radar: every top-volume AI question is a Google AI Overview on HK property-pricing queries.",
      "implication": "You don't need more rankings or more links. The visibility moved into the answer layer; you need to own and monetise the AI answer surface."
    },
    {
      "finding": "Prior vendor's backlink focus misdiagnoses the problem.",
      "evidence": "Ahrefs: DR 41, 3,266 referring domains, 546 keywords in positions 1–3. The coordinated spam wave (identical first-seen timestamps 2026-01-31, 'TELEGRAM @happygrannypies' anchors) postdates the traffic decline that began in 2024, so it cannot be the cause. 13 broken am730 links (DR 75, '%0D%20' template error) are a fixable technical issue, not a link-building need.",
      "implication": "The answer is on-site and technical, not off-site and link-based. Backlink cleanup is warranted; acquisition is not."
    }
  ],
  "risk_areas": [
    {
      "issue": "Onsite blind spot: no crawl data exists for pages that drive the business.",
      "priority": "P0",
      "effort": "LOW",
      "impact": "HIGH",
      "detail": "The crawl sampled 18 Shopify-flavoured pages; the pages AI cites and users land on (e.g., /rent/, listing/estate URLs) were never crawled. Combined with Ahrefs digest_internal_linking returning 0, internal architecture is a complete blind spot. A clean re-crawl excluding /cdn/ and seeding portal paths is mandatory before any template claim."
    },
    {
      "issue": "Bot-blocking zeroes Ahrefs summary data, masking the real competitive picture.",
      "priority": "P0",
      "effort": "MEDIUM",
      "impact": "HIGH",
      "detail": "Ahrefs digest_authority, digest_internal_linking, and digest_branded_vs_generic all return zeros. This is likely caused by AhrefsBot being blocked from recrawling, not by actual lack of content or links. Fixing the WAF rule will restore the data quality needed for accurate diagnostics."
    },
    {
      "issue": "Competitors already win contested AI topics: short-term rentals.",
      "priority": "P1",
      "effort": "MEDIUM",
      "impact": "MEDIUM",
      "detail": "On short-term rental queries, 28Hse has 7 responses, Spacious blog 5, Weave Living 4—all ahead of house730. This is a concrete content gap that is both an opportunity and a risk if left unaddressed.",
    },
    {
      "issue": "13 broken am730 backlinks from a high-DR domain.",
      "priority": "P2",
      "effort": "LOW",
      "impact": "LOW",
      "detail": "am730 (DR 75) links to house730 with a broken URL template embedding '%0D%20' from a GSC import. Fixable with a simple redirect. Not a driver of the traffic decline but an easy win to recover authority flow.",
    }
  ],
  "bot_management": {
    "observed_facts": [
      "18 of 21 probes from a datacenter IP returned HTTP 403, including requests using the browser user-agent and robots.txt.",
      "403s were fast (62–87ms), indicating rejection at the edge, not the origin server.",
      "PerplexityBot was the only agent to receive a 200 on robots.txt, with an average response time of 215.7ms, confirming origin access.",
      "GPTBot, OAI-SearchBot, ClaudeBot, Bytespider, and Google-Extended all returned 403 from the probe IP."
    ],
    "seo_risks": [
      "AhrefsBot is blocked from recrawling, causing summary metrics (digest_authority, digest_internal_linking) to compute impossible zeros—this falsifies the competitive analysis and may mislead SEO strategy.",
      "Confirmed blockage of PerplexityBot competitors (GPTBot, OAI-SearchBot, ClaudeBot) suggests the AI citation gap is self-inflicted.",
      "If the same IP rule affects production AI crawler IPs, AI visibility on non-Google platforms is structurally prevented."
    ],
    "geo_hypothesis": {
      "claim": "The datacenter-IP blocking rule is a root cause of the AI citation gap on non-Google platforms.",
      "supporting": [
        "PerplexityBot, the only confirmed-to-be-unblocked crawler, is the only non-Google AI platform with citations (500 in Ahrefs).",
        "GPTBot, OAI-SearchBot, ClaudeBot all returned 403 from the probe IP.",
        "80% of AI impressions (14,190 of 17,710) come from Google AI Overviews, consistent with only Google's crawler being unaffected."
      ],
      "prevents_calling_fact": [
        "The live check is a single datacenter IP; it does not prove production AI crawler IPs are blocked.",
        "Google's pipeline clearly works (151K monthly organic traffic, AI Overviews dominance), so Googlebot is not blocked—this is consistent with the hypothesis that only datacenter IPs are affected.",
        "Correlation does not equal causation: traffic decline and bot blocking coexist, but causation is unproven."
      ],
      "definitive_check": "Server logs showing whether production IPs for GPTBot, OAI-SearchBot, ClaudeBot, etc. receive 200 or 403. GSC crawl stats for AI crawlers. Cloudflare (or equivalent) WAF logs."
    },
    "recommendations": [
      "Audit the WAF/edge rule that blocks datacenter IP ranges. Whitelist all verified AI crawler IP ranges (GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, etc.).",
      "Verify the fix by re-running the live check from the same datacenter IP and confirm 200 for each AI crawler.",
      "Cross-check in GSC crawl stats for AI crawl activity after the change."
    ],
    "method_caveat": "The live check is UA-based from a single datacenter IP and is confounded by edge routing, Cloudflare/other WAF, and possible IP-based exceptions. It does not diagnose production crawler behaviour. Only server logs resolve this."
  },
  "narrative_arc": [
    "Your website blocks the very crawlers that your measurement and AI visibility depend on.",
    "Ahrefs sees zeros; AI platforms other than Google cannot reach you; yet Perplexity—the one crawler that gets through—is your only non-Google citation source.",
    "Meanwhile, your rankings are dominant, your traffic decline is a channel shift into AI Overviews, and your prior vendor's backlink focus missed the real bottleneck.",
    "The fix is a configuration change, verifiable in week 1, that unlocks accurate data and a path to diversify AI visibility beyond Google."
  ],
  "recommendations": [
    {
      "action": "Conduct immediate audit of WAF/edge rules blocking datacenter IP ranges.",
      "rationale": "The live check confirmed 18 of 21 probes return 403 from a datacenter IP. This blinds Ahrefs and blocks all tested AI crawlers except PerplexityBot. Restoring access is prerequisite for any other SEO or GEO work.",
      "validation": "After fix, re-run live checks from the same datacenter IP and confirm 200 for all AI crawler user-agents. Confirm in GSC crawl stats."
    },
    {
      "action": "Diversify AI content beyond listings: produce Cantonese district-pricing explainers and a short-term rental guide.",
      "rationale": "80% of AI impressions are Google AI Overviews on listing content. Competitors already out-cite house730 on short-term rentals. This gap is both a risk and a concrete growth opportunity.",
      "validation": "After publication, track citations per platform in Brand Radar and Ahrefs; monitor share of voice on targeted question sets."
    },
    {
      "action": "Perform a clean crawl of the actual portal pages (exclude /cdn/, seed /rent/, estate profiles, /property-tips/).",
      "rationale": "The current crawl sampled only Shopify-flavoured pages and provides zero data on the pages AI cites and users visit. Internal architecture, templates, and on-page signals are a complete blind spot.",
      "validation": "After crawl, cross-check crawled pages vs. Brand Radar's cited-page list and Ahrefs' top-traffic pages to ensure coverage."
    }
  ],
  "hypotheses": [
    {
      "claim": "The datacenter-IP blocking rule is a root cause of the AI citation gap on non-Google platforms.",
      "status": "suggestive",
      "how_to_verify": "Server logs confirming production AI crawler IPs are also blocked. GSC crawl stats for AI crawlers."
    },
    {
      "claim": "Traffic decline is caused by AI Overviews absorbing clicks that house730 still wins in rankings.",
      "status": "suggestive",
      "how_to_verify": "Analyse GSC click-through-rate by query type (AI Overview vs. non-AI Overview). Correlate with Brand Radar impression data."
    }
  ]
}