« Web Scraping

Canonical tags & meta robots directives

robots.txt controls whether a crawler fetches a page at all. These two mechanisms work at the page level instead, after a page has already been fetched: a canonical tag says which URL a piece of content should be attributed to, and a robots directive says what to do with this specific page once it's been read.

Canonical tags

Meta robots & X-Robots-Tag