Next.js robots.txt: App Router setup and launch checks
Add robots.txt to a Next.js App Router site, declare the sitemap, and diagnose missing or cached output without confusing crawl rules with access control.
Choose one robots.txt implementation
In the Next.js App Router, you can use a static app/robots.txt file or generate the response with app/robots.ts. If your app lives under src/app, use that directory instead. Pick one owner for the route so a static asset and a generated handler do not compete.
Use a static file when the production rules rarely change. Use a typed metadata route when the configuration belongs in code. The examples follow the Next.js 15 App Router API; a Pages Router project can serve a static file from public/robots.txt.
Static production example
Replace the domain and private route pattern with your own. The slash in /dashboard/ matches paths beneath that prefix; check whether /dashboard without the slash also exists and what it returns. Public product and article routes should remain available if you want search engines to crawl them.
# app/robots.txt (or src/app/robots.txt)
User-agent: *
Allow: /
Disallow: /dashboard/
Sitemap: https://example.com/sitemap.xmlTyped App Router example
Use the canonical production origin for the sitemap address. Do not derive it from an untrusted request host. If environment variables provide the origin, validate their deployment values before relying on them in crawler-facing output.
// app/robots.ts (or src/app/robots.ts)
import type { MetadataRoute } from 'next';
export default function robots(): MetadataRoute.Robots {
return {
rules: {
userAgent: '*',
allow: '/',
disallow: ['/dashboard', '/account'],
},
sitemap: 'https://example.com/sitemap.xml',
};
}Keep crawling, indexing, and authentication separate
robots.txt controls crawling for compliant bots. It does not protect a private route and does not reliably remove a known URL from search. Protect private data with authentication and authorization. For an accessible page that should stay out of search, use an appropriate noindex instruction and allow a crawler to read it.
Google does not support the crawl-delay directive. Avoid adding it as a fix for a sitemap fetch failure or assuming it changes Googlebot's behavior.
Check the deployed response
| Observation | Next check |
|---|---|
| 404 at /robots.txt | File placement, deployment output, rewrites, and route ownership |
| 200 but body is HTML | A middleware or SPA fallback is serving the wrong response |
| Unexpected Disallow: / | Production configuration inherited a staging rule |
| Rules refer to an old domain | Hard-coded origin, environment variable, or cached deployment output |
| Correct local response, stale production response | CDN cache and the exact deployment currently serving the domain |
curl -L -D robots-headers.txt https://example.com/robots.txt -o robots-body.txt
curl -L -D sitemap-headers.txt https://example.com/sitemap.xml -o sitemap-body.xmlTroubleshoot middleware and caching
If middleware applies login or locale redirects to every path, make sure crawler files receive their intended public responses. Use the project's existing matcher strategy rather than copying an unrelated matcher that excludes the wrong routes.
Compare the body and headers from production with the code you changed. Avoid assuming a local edit is live. A metadata route can be cached; choose the framework's supported caching controls when your rules truly need request-time behavior. Recheck after deployment and any cache purge.
Ready for a pre-launch audit?
Run the public analyzer and get a prioritised report for the URL you are about to share.
Run the SEO audit