File viewer

crawl_12082026_0118.md

/app/data/llm/analysis/house730/crawl_12082026_0118.md

This crawl is heavily distorted by a sampling artefact. 179 of the 197 URLs (90.9%) live under /cdn/ — almost certainly static assets (images, maybe CSS/JS) rather than HTML pages. Everything that follows about metadata and content quality is downstream of this.

What the data actually shows

  • URL composition keeps the sample from representing the real site. Only 13 URLs are from content-adjacent paths: 6 /products/, 5 /pages/, 2 /collections/. This sample cannot be used to assess title/meta quality for products or editorial content.
  • The 82.2% indexable rate (162 URLs) is misleading. If the /cdn/ directory returns 200 for image requests without noindex tags, those 179 URLs are all technically indexable. That’s not a win — it’s a likely configuration problem where static assets can be crawled and indexed.
  • Empty H1 (181) and missing meta descriptions (185) are artefactual, not diagnostic. CDN resource URLs won’t have H1s or meta tags. The 6 actual product pages may or may not have these issues; we don’t know from this export.
  • Duplicate titles showing 0 is noise. With fewer than a dozen likely content pages, the sample isn’t large enough for duplicate detection to be meaningful.
  • Depth distribution is consistent with CDN asset organisation. A single depth-0 homepage, then 65 depth-1 URLs and 95 depth-2 URLs — typical when /cdn/images/ paths sit at depth 1 or 2. The real content pages (e.g., products) could be deeper but we have only a handful.
  • Two status-0 URLs and one 302 are minor. The 302 hits /checkouts/ (1 URL) — likely a normal flow redirect. The status-0 errors need retesting but are too few to read into.

Connection: CDN URLs and indexable bloat

The co-occurrence of 179 /cdn/ URLs and 162 indexable URLs points to a potential root cause: the CDN endpoint serves assets without X-Robots-Tag: noindex or appropriate robots.txt blocking, and the crawler followed links to them. This doesn’t prove Google is indexing them, but it means nothing is actively preventing indexing. This can lead to crawled budget waste and index pollution. Hypothesis to verify: check if the CDN subdomain or path is excluded in robots.txt or returns proper directives.

What stands out most

  • The audit sample is functionally a CDN inventory, not a site audit.
  • We can draw zero conclusions about on-page SEO for products, collections, or editorial content. The only safe inference is that the CDN setup may be exposing a large number of non-HTML, indexable URLs to crawlers.

Limitations that require additional data

  • No visibility into whether /cdn/ URLs are images, scripts, or something else. A handful of live checks or a content-type header sample will clarify.
  • No canonical tags or robots directives in the export to confirm how indexable these assets are in practice.
  • 20 URLs out of ~34K means we’re missing the distribution of true content directories. Without a full crawl or site: query data, we can’t assess site structure or depth for real pages.

Key signals: - 179 out of 197 crawled URLs fall under /cdn/, making the sample unusable for content quality assessment. - 162 indexable URLs, mostly driven by CDN resources, suggests no indexing prevention is in place on those paths. - Only 13 content-path URLs (6 products, 5 pages, 2 collections) were captured — no meaningful title/meta evaluation possible. - Two status-0 errors and one 302 to checkout are negligible; CDN indexable bloat is the only issue with scale signals. - Potential featured finding: The CDN endpoint likely lacks noindex directives, artificially inflating indexable counts and risking crawled budget waste; this is the single most actionable observation before any other SEO work can be scoped.