File viewer

crawl_12082026_0142.md

/app/data/llm/analysis/house730/crawl_12082026_0142.md

The crawl data is dominated by Shopify asset URLs—179 out of 197 URLs live under /cdn/. These are image and static resource files, not content pages. Because Screaming Frog applies on-page checks to every crawled URL, the headline figures for empty <h1> (181) and missing meta descriptions (185) are almost entirely generated by those assets, not by the site’s actual product, collection, or information pages.

What the numbers actually show for real pages

The crawl contains only 18 non-asset URLs: 6 /products/, 5 /pages/, 2 /collections/, and 1 each for /, /checkouts/, /search/, /cart/, and /customer_authentication/. Among these, duplicate titles count is zero, which suggests all 18 page titles are unique—a positive signal, though the sample is tiny. At least 16 of those 18 pages have an <h1> (197 total – 181 empty = 16 with <h1>), and at least 12 have a meta description (197 – 185). A few real pages may be missing those elements, but without isolating the HTML pages, you cannot identify which ones or how severe the gap is. The 2 status‑0 connection errors and single 302 are trivial noise.

Indexability is not a crisis

162 of the 197 URLs are marked indexable (82.2%). That figure includes images that aren’t blocked by noindex or robots.txt—this is typical and not inherently a problem. The 35 non‑indexable URLs likely stem from disallowed or noindexed asset paths. There is no indication that critical content pages are blocked. The single /search/ page could be an indexation risk (can generate infinite thin pages if not noindexed), but with only one URL in the sample, that’s untested.

The crawl sample is functionally useless for page‑level SEO diagnosis

Because 90% of the crawl is asset noise, all aggregated on‑page metrics are misleading. The sample also captures only 18 out of ~34K real site URLs—far too few to assess title/H1/meta patterns at scale. The depth distribution (65 at depth 1, 95 at depth 2, 36 at depth 3+) likely reflects asset embedded‑resource crawling rather than the information architecture of content pages.

Immediate implication

A clean crawl—excluding /cdn/, and optionally /cart/, /checkouts/, /search/, /customer_authentication/—is required before any reliable statement can be made about page‑level SEO health. The zero duplicate titles among the 18 real pages is encouraging but needs to be validated against a representative HTML crawl.


Key signals

  • 90.8% of crawled URLs are /cdn/ assets (179/197). This corrupts every on‑page aggregate metric—empty H1 (181), missing meta (185)—making the data unfit for content diagnosis.
  • Only 18 actual content pages appear in the sample from ~34K known site URLs. No claims about title duplication, H1 coverage, or meta completeness can be generalised from this.
  • Duplicate titles are zero across those 18 pages. A positive micro‑signal that title uniqueness may be well‑handled, pending a clean crawl.
  • Indexability is not a concern: 162 indexable URLs largely reflect image files not blocked by directives, not valid SEO pages. Nothing indicates a noindex or robots.txt error on core pages.
  • Potential featured finding: The crawl is structurally misconfigured for this Shopify site. The single action of excluding asset paths (especially /cdn/) would transform the output from misleading noise into actionable page‑level insight. That is the root cause that, if fixed, unlocks the most value from future crawls.