Robots.txt, noindex and canonicals do different jobs

You will be able to choose the right control for keeping a page out of results or consolidating duplicates.

A small accounting firm in Singapore launched a new website after months of work. The designer had built it on a test address first, and to keep Google away from the half-finished version, ticked a box that said "discourage search engines". The site went live. Three months later, the firm noticed that enquiries from Google had stopped almost completely. The box was still ticked.

Stories like this are common, and they come from mixing up three controls that sound similar and do quite different jobs: robots.txt, noindex and canonical tags. This lesson explains each one and when to use it.

Robots.txt controls crawling

Robots.txt is a plain text file at the root of your site, at your domain followed by /robots.txt. It gives instructions to crawlers about which parts of the site they may fetch. A line such as "Disallow: /admin/" asks crawlers not to visit pages under that folder. You can open your own robots.txt in a browser right now by typing that address.

The thing to understand is that robots.txt controls crawling, not indexing. Google's documentation says that a page blocked by robots.txt can still be indexed and appear in results if other pages link to it. In that case Google knows the address exists but cannot read the page, so the result may show the address with little or no description.

So robots.txt is the right tool when you want to stop crawlers spending time on parts of a site that have no value in search, such as internal search results or endless filtered versions of a product list. It is the wrong tool for hiding a page from Google's results.

Noindex keeps a page out of results

To keep a page out of Google's results, use a noindex rule. It is a small instruction, placed either in a meta tag in the page's code or in the page's HTTP response header, that tells Google not to include the page in its index. Many website builders show it as a setting such as "hide this page from search engines" on each page.

There is a catch that trips people up. Google can only see a noindex rule if it can crawl the page. If you block the page in robots.txt and also add noindex, Googlebot never fetches the page, never sees the noindex, and the page could still appear in results through links. Google's documentation is explicit: for noindex to work, the page must not be blocked by robots.txt.

Good uses for noindex include thank-you pages after a form, internal pages for staff, thin tag or archive pages your builder creates automatically, and test pages you need to keep live. For Priya, the confirmation page parents see after booking a trial class is a sensible noindex candidate.

Canonicals pick one version among duplicates

Lesson 1.1 explained that Google groups duplicate or near-duplicate pages and chooses one as the canonical, the version it may show. You can tell Google which version you prefer with a canonical tag: a line in the page's code that points to the preferred address.

Duplicates are more common than people think. The same page can be reachable with and without "www", with a tracking code added to the address, or through two different menu paths. An online shop might show the same product at several addresses, one per colour. A canonical tag on each version pointing to one main address tells Google which you prefer.

Google treats the canonical tag as a strong hint, not a command. If other signals disagree, for example if your sitemap lists a different version or your internal links all point elsewhere, Google may choose another version. Keep your signals consistent: link to the canonical version, list it in your sitemap and point the canonical tags at it.

A canonical tag is the wrong tool for keeping a page out of results. It tells Google which version to show, so the content still ends up in the index under the preferred address.

The staging accident

Many sites are built on a staging version first, a private copy where changes can be tested. Designers often block it from search engines with a blanket robots.txt rule that disallows the whole site, or a noindex on every page. Both are sensible on a staging site.

The trouble starts when the site is copied to the live address with those settings still in place. A blanket "Disallow: /" in robots.txt asks every crawler to stay away from everything. A site-wide noindex removes every page from results over time. The accounting firm in the opening had the second. Neither is visible to a person browsing the site, which is why they can go unnoticed for months.

After any site launch, redesign or platform move, check robots.txt and the noindex setting on a few pages within the first day.

Choosing the right control

Ask what you want to happen. If you want a page kept out of results, use noindex and leave it crawlable. If you want Google to show one version of several near-identical pages, use a canonical tag on each. If you want crawlers to skip a section that has no search value and you do not mind if its addresses are known, use robots.txt.

Start with your own site. Open your robots.txt, then view three important pages and check whether any carry a noindex rule or a canonical tag pointing somewhere unexpected. Your website builder's SEO settings or the URL Inspection tool in Search Console will show you both. The activity below asks you to note anything that blocks a page you want found.

Read your site's robots.txt, check three key pages for noindex and canonical tags, and note anything that blocks a page you want found.

Course

Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).