Crawl architecture

Robots- und Sitemap-Kreuzprüfung

Robots.txt und Sitemap-XML gemeinsam lesen und Blockierungen, Redirects, noindex, defekte Antworten sowie Canonical-Konflikte finden.

Der öffentliche Crawl ist auf 20 URLs begrenzt.

Tool richtig interpretieren

Methode, Daten und Grenzen

Was macht das Tool?

Cross-check robots.txt, sitemap declarations, sitemap XML, and sampled page indexability.

Wie funktioniert es?

The checker reads robots.txt, follows a bounded sitemap index, then inspects only the configured sample limit.

Welche Daten werden analysiert?

Robots directives, sitemap XML locations, HTTP status, redirects, meta robots, and canonical tags.

Wie ist das Ergebnis zu lesen?

Sitemap URLs should be reachable, indexable, unblocked, and self-canonical.

Einschränkungen

  • — The public scan is deliberately bounded.
  • — JavaScript-rendered directives are not evaluated.
  • — Large crawls belong in the queue-backed audit workflow.

Datenschutz und Speicherung

  • — Only the origin you submit is contacted, from our server, within the configured crawl limit.
  • — We store the tool key, the origin, the score, and how many URLs were sampled — not the crawled documents.
  • — Credential-like query parameters are redacted before the run is recorded.
  • — Your IP address is stored only as an irreversible hash for abuse protection.
  • — Results are not published, are not indexed, and never appear in our XML sitemap.