synthesis_12082026_0223.md
/app/data/llm/analysis/bupa/synthesis_12082026_0223.md
Synthesis: bupa.com.hk
Where the sources corroborate each other
- Live checks + crawl: Both confirm unobstructed access. The crawl reached 5,191 URLs without systemic 403s; live checks got 200s for every AI bot on every path. No contradiction — the site is open to machines, full stop.
- Crawl + Brand Radar: The crawl shows a substantial real content corpus exists (/tc/ = 1,216 URLs, /en/ = 988, /zh/ = 344). Brand Radar shows only 26–27 of ~3,627 indexable pages are ever cited by AI. Together: content existence and crawl access are not the bottleneck. Something between "page exists and is fetchable" and "page gets cited" is failing.
- Brand Radar + crawl on the language mismatch: The highest-volume AI queries are Traditional Chinese ("香港自愿医保哪家好?", 2,100 vol), and Bupa's largest content directory is /tc/. Bupa has Chinese-language content at scale, yet Chinese-language commercial queries are answered from bowtie.com.hk, 10life, hkvhis, moneyhk101. The gap is not content volume.
Where sources complicate each other
- Brand Radar's own conclusion vs. the crawl: Brand Radar framed the citation gap as "a content structure and citation authority problem, not a technical SEO problem." The crawl contradicts the second half of that sentence: ~4,507 empty H1s and ~4,503 missing meta descriptions across crawled URLs indicate the content structure problem is a template-level technical problem. The crawl doesn't just complicate the Brand Radar diagnosis — it supplies its mechanism.
- Crawl caveat that tempers the on-page finding: The H1/meta counts span all crawled HTML, including /files/, /PDF/, and /_Incapsula_Resource/ assets where empty H1s are expected. The analyst's conclusion ("virtually every indexable page lacks a meta description") holds directionally, but the exact failure rate on editorial pages in /tc/, /en/, /zh/ is unquantified. Verification step, not a refutation.
- Brand Radar SOV is a measurement artefact: 100% share of voice with a flat 365-day trendline means only Bupa is tracked. Do not quote it as a competitive signal. The real GEO metrics are own-citation share (6.7%) and Bowtie 48 responses vs Bupa 29.
Narratives that only emerge with all sources together
1. Open door, empty room. Live checks eliminate blocking as an explanation (0% block rate, 42–84 ms for AI bots). The crawl eliminates content absence (/tc/ library exists and is crawlable). Brand Radar documents the outcome (6.7% own-citation share; homepage = 34% of own citations; zero product, claims, or network pages cited). The only explanation consistent with all three: when AI crawlers fetch Bupa's deep pages, the templates give them nothing to extract — no H1 hierarchy, no descriptive metadata — so models cite the one page with brand-level clarity (the homepage) or, far more often, a competitor or aggregator whose pages are structured for extraction. Bowtie's 47 cited deep pages vs Bupa's homepage-and-contact-page citation profile is exactly the pattern this predicts.
2. The competitor gap is structural, not authority-based. No source provided shows backlinks or domain authority as the differentiator. The visible difference between Bupa and Bowtie is cited-page depth and count (47 vs 27 pages; product/education content vs homepage/about/customer-care). Combined with the crawl's template finding, this reframes the GEO problem as fixable on-site — and directly counters any "you need more authority" pitch.
3. Crawl budget dilution compounds everything. 23% of crawled URLs are redirects; /files/ + /_Incapsula_Resource/ + /PDF/ are ~30% of the crawl. Every bot — Googlebot or AI — spends fetch budget on junk, slowing discovery and refresh of the citable /tc/ and /en/ content. Secondary to the template problem, but it means even fixed templates will take longer to pay off without hygiene work.
Data gaps visible only in combination
- robots.txt content unknown. Live checks prove HTTP-level openness, but compliant bots (GPTBot, ClaudeBot) obey robots.txt. Copilot and Google AI Overviews citations don't settle this — they draw on Bing/Google indices, not those bots. One fetch closes this gap.
- No cross-walk between cited pages and crawl data. Are the 26 cited pages the few with intact H1s, and the uncited product/network/claims pages the broken ones? If yes, that's direct mechanism proof. Both datasets exist; they just haven't been joined. Week-1 verifiable.
- Crawl covers 5,191 of ~34K URLs, seeded mid-architecture. Depth data is an artefact; true architecture, orphan status, and the state of the full /tc/ library are unknown. Needs a homepage/sitemap-seeded crawl.
- No GSC/indexation data. 3,627 "indexable" pages ≠ indexed pages.
- Bowtie not crawled. The claim "Bowtie's pages are better structured" is inference from citation patterns, not direct evidence.
Ranked priorities
- Template signal failure as the mechanism of the AI citation gap (featured below) — GEO + onsite, jointly.
- robots.txt audit — gates the entire GEO narrative; trivial effort.
- Cited-page ↔ crawl cross-walk — converts "consistent with" into proof; the client's team can do it themselves in week 1.
- Crawl budget hygiene — redirect cleanup, exclude /files/, /_Incapsula_Resource/, /PDF/ from crawl paths.
- Full-corpus crawl + Brand Radar competitor configuration — fixes the measurement gaps (depth artefact, meaningless SOV).
- Backlinks — absent from all three sources as a factor; no evidence they are the bottleneck. Do not lead with them.
Featured finding
Bupa's AI invisibility is a self-inflicted, on-site problem — and all three data sources converge on it.
- Live checks prove every major AI crawler (GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended, Bytespider) reaches bupa.com.hk with a 0% block rate and sub-90 ms response times. Access is not the problem.
- Brand Radar shows that despite open access, Bupa owns only 26 of 389 cited pages (6.7%) in its own AI query set; its most-cited page is the homepage (10 of 29 responses), and no product detail, claims, or network hospital page is cited at all. Meanwhile Bowtie earns 48 responses from 47 deep pages, and the 2,100-volume query "香港自愿医保哪家好?" is answered from comparison sites and competitors.
- The crawl shows why: empty H1s on ~4,507 URLs and missing meta descriptions on ~4,503 — near-universal template-level absence of the heading structure and descriptive signals AI models extract when choosing what to cite.
The mechanism is eliminative and verifiable: blocking is ruled out (live checks), content absence is ruled out (1,216 /tc/ URLs in the crawl), so the citation gap is consistent with the one failure all three sources leave standing — pages that present no extractable structure to any machine reader. The confirming test is simple: check whether Bupa's 26 cited pages are the rare ones with intact H1s, and whether Bowtie's 47 cited pages have them. Either the client's team or we can do this in week 1.
This matters more than anything else in the data because it sits on the highest-value demand (12,420 monthly AI impressions, 70.4% via Google AI Overviews, led by a 2,100-volume commercial query), it is losing that demand to a named direct competitor, the fix is a template deployment rather than a multi-quarter authority campaign, and every alternative explanation has already been eliminated by the data in hand. Nothing else in the sources — redirects, crawl budget, metadata CTR — comes close to this combination of value at stake, clarity of cause, and speed of fix.
Next step requiring signature: GSC access plus server logs (to confirm real AI-bot fetch behaviour at production scale) and a sitemap-seeded full crawl — the three items needed to convert this synthesis into a scoped template-fix engagement.