Sitemap and robots.txt

Serve an XML sitemap with hreflang alternates for the languages each page is translated into, and a robots.txt that points to it.

Both routes are off by default. Turn them on in config.php:

site/config/config.php
return [
    'johannschopplich.helpers' => [
        'sitemap' => [
            'enabled' => true
        ],
        'robots' => [
            'enabled' => true
        ]
    ]
];

The sitemap is served at /sitemap.xml and robots.txt at /robots.txt.

What It Lists

The sitemap lists every listed and unlisted page – the home and error pages included – but no drafts. Each entry carries the page URL, its modification date, and a priority:

<url>
  <loc>https://example.com/exhibitions</loc>
  <lastmod>2026-10-09</lastmod>
  <priority>0.8</priority>
  <changefreq>weekly</changefreq>
</url>

priority and changefreq resolve like meta values: from the page model's metadata(), the meta.defaults option, or a page field of that name. priority defaults to 0.5, is clamped between 0 and 1, and skips the site field. changefreq falls back to a changefreq field in the site content and is left out without any value. As blueprint fields:

site/blueprints/pages/default.yml
fields:
  priority:
    label: Sitemap Priority
    type: range
    min: 0
    max: 1
    step: 0.1
    default: 0.5
  changefreq:
    label: Change Frequency
    type: select
    options:
      - always
      - hourly
      - daily
      - weekly
      - monthly
      - yearly
      - never

Exclude Pages

Three ways keep a page out of the sitemap. A page excluded by one of them keeps its children in the sitemap unless they are excluded too.

By Template

exclude.templates takes template names, matched against the page's intended template:

site/config/config.php
return [
    'johannschopplich.helpers' => [
        'sitemap' => [
            'enabled' => true,
            'exclude' => [
                'templates' => ['error', 'search']
            ]
        ]
    ]
];

By Page ID

exclude.pages takes page IDs as regular expressions, matched against the whole ID regardless of case. blog excludes the blog page but not its posts; blog(/.*)? excludes both:

site/config/config.php
return [
    'johannschopplich.helpers' => [
        'sitemap' => [
            'enabled' => true,
            'exclude' => [
                'pages' => ['legal/privacy', 'blog(/.*)?']
            ]
        ]
    ]
];

An ID is the page's folder path without sorting numbers, so it carries no language prefix and no translated slug. Instead of an array, exclude.pages takes a closure that returns one:

site/config/config.php
return [
    'johannschopplich.helpers' => [
        'sitemap' => [
            'enabled' => true,
            'exclude' => [
                'pages' => fn () => site()->index()->filterBy('noindex', 'true')->pluck('id')
            ]
        ]
    ]
];

In the Blueprint

site/blueprints/pages/private.yml
title: Private Page
options:
  sitemap: false

Multi-Language Sites

Each page has one entry at its default-language URL. Below it, the entry lists an hreflang alternate for every language the page has a content file in, and an x-default alternate for the default-language URL:

<url>
  <loc>https://example.com/about</loc>
  <lastmod>2026-10-09</lastmod>
  <priority>0.5</priority>
  <xhtml:link rel="alternate" hreflang="de-de" href="https://example.com/de/ueber-uns" />
  <xhtml:link rel="alternate" hreflang="en-us" href="https://example.com/about" />
  <xhtml:link rel="alternate" hreflang="x-default" href="https://example.com/about" />
</url>

A page without content in any language, such as a virtual page created without content, gets an alternate for every language.

The hreflang value comes from the language's locale in lowercase, so de_DE.utf8 becomes de-de.

Caching

With Kirby's pages cache enabled, the sitemap is cached with it. Edits saved through Kirby, in the Panel or from PHP, refresh it. A change to config.php, a blueprint, a page model, or a content file on disk shows once the cache is emptied.

robots.txt

The route answers with a fixed body that allows every crawler and names the sitemap:

User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml

It names the sitemap even when the sitemap route is off. A robots.txt file in the web root is served by the web server before Kirby sees the request, so delete it for the route to answer.