Back to Blog

Can a crawler actually read your site?

TL;DR: Key Takeaways

Before GEO workshops, ask an unkind question: is the offer in the HTML, or only in client state after JavaScript?

  • Many marketing sites render services in the browser.
  • Robots.txt, sitemaps, canonicals and soft 404s remain the base.
  • SOURCE/01 by GlasBox (Search) lists blocked agents, render-dependent service pages and index conflicts.

Before GEO workshops, ask an unkind question: is the offer in the HTML, or only in client state after JavaScript?

Many marketing sites render services in the browser. Humans see cards and tabs. A plain fetcher sees an empty shell. The same applies to consent walls that empty the first HTML byte, and to `noindex` on the very pages that define the offer.

Robots.txt, sitemaps, canonicals and soft 404s remain the base. What is new is reading the bot classes vendors document for retrieval. Blocking "all AI bots" as a reflex can hit training corpora and also harm systems that fetch live. The decision should be written down, not left as a forgotten banner toggle.

SOURCE/01 by GlasBox (Search) lists blocked agents, render-dependent service pages and index conflicts. Without that list every content plan is speculation.

---

Next step

GlasBox measures classic SEO and AI visibility separately. No ranking or mention guarantee. The free short check stays thin on purpose.

Request an introductory call

Frequently Asked Questions

What question comes before a GEO workshop?
Is the offer in the HTML, or only in client state after JavaScript?
Should you block all AI bots by default?
Usually not. It can hit training corpora and also harm systems that fetch live. The decision should be written down.
What does SOURCE/01 by GlasBox (Search) list?
Blocked agents, render-dependent service pages and index conflicts. Without that list every content plan is speculation.