crawl_12082026_0223.md
/app/data/llm/analysis/bupa/crawl_12082026_0223.md
This crawl provides a partial view (5,191 URLs from an initial 20-seed list) of a much larger site (~34K known URLs). The findings point to two systemic problems that will compound one another: widespread on-page neglect and a heavy overhead of non-content URLs consuming crawl budget.
Indexability and status codes
- 1,141 permanent redirects (301) plus 76 temporary ones (302) mean over 23% of crawled URLs simply forward elsewhere. A redirect-to-content ratio this high likely reflects legacy URL structures, domain migrations, or tracking parameters that were never cleaned up. While redirects themselves aren't a problem, at this volume they waste crawl budget and can slow discovery of real content.
- 19 404s and 2 500s are low in absolute terms and not a priority unless they represent important pages that were recently removed.
- 3,627 indexable URLs out of 5,191 gives a 69.9% indexability rate. The “lost” 30% is almost entirely redirects. That’s blunt evidence that the crawl is spending significant time on paths that lead nowhere.
On-page fundamentals (scope: all crawled HTML resources)
- Meta descriptions: 4,503 URLs have no meta description. With only 3,953 200-status pages in the crawl, this number slightly exceeds the total of live HTML pages, which means it likely includes some redirects or non-HTML resources. Nonetheless, the implication is clear: virtually every indexable page that does exist lacks a meta description. That will depress click-through from search results at scale.
- H1 tags: 4,507 URLs have an empty H1. Again covering essentially all pages. This is not a “some pages lack headings” situation — it is an absence of any meaningful heading structure across the entire crawled portion of the site.
- Duplicate titles: 210 pages share a title with at least one other page. Among a corpus dominated by missing H1 and meta, duplicate titles are a secondary concern, but they confirm a lack of editorial attention to page-level signals.
Taken together, these three numbers describe a site where the majority of pages present no clear topic, no persuasive snippet, and no structural heading hierarchy to search engines. This is the single most actionable finding in the data.
URL composition and likely bloat
The directory breakdown reveals where the crawl budget is being spent:
- /files/ (1,215 URLs) and /_Incapsula_Resource/ (317 URLs) together account for nearly 30% of all crawled URLs. Both are almost certainly not editorial content. /files/ likely contains uploaded assets (PDFs, images, documents); /_Incapsula_Resource/ is an artifact of the Incapsula/Imperva CDN and is serving internal JavaScript or CSS. Neither belongs in search engine indices.
- /PDF/ and /pdf/ add another 139 URLs. If these are patient forms, brochures, or policy documents, they probably lack titles, H1s, and search-optimised descriptions — exactly matching the broader metadata gap.
- The three language directories /tc/ (1,216), /en/ (988), /zh/ (344) contain the bulk of genuine content. Traditional Chinese pages heavily outnumber English and Simplified Chinese, suggesting a content effort (likely blog or health library) primarily in TC. The /-/ directory (890 URLs) may be a CMS media handler (common in Sitecore), further inflating asset crawls.
Crawl depth caveat and site architecture
The depth bucket shows 5,038 URLs at depth 3+ from the seed list, and only 1 at depth 0, 1 at depth 1, and 151 at depth 2. This distribution is almost certainly an artefact of the crawl configuration: the 20-seed list did not include the homepage or main navigation anchor points, so the crawl started deep and all further internal links appear at depth 3+. From this data alone, no conclusion can be drawn about the real architectural depth of important pages. A follow-up crawl from the homepage (or a sitemap) is required to assess whether key product and content pages are too many clicks from the root.
However, what the depth data does confirm is that the seed URLs (wherever they sat in the architecture) are linked to thousands of other pages via internal links — which means the crawl was able to discover a significant web of content. The risk is that without a representative crawl, the true hierarchy and orphan status of important pages remain unknown.
Connections and implications
The combination of massive basic on-page gaps and a crawl budget diluted by redirects, CDN resources, and legacy file directories explains multiple symptoms at once:
- Even if Google can discover and index all 3,627 indexable pages, those pages will present no