Key takeaways
- Private paths (admin, accounts, drafts, customer data) should be disallowed, and still protected by login.
- A blanket block of the whole site also blocks the agents that recommend you.
- robots.txt is a request to well-behaved crawlers. It is not a lock.
- An /llms.txt brief gives answer agents the public facts in one file, so they crawl fewer pages.
“How do I stop AI from scraping my site?” usually means two fears at once. One is real: a crawler copying customer data, drafts, or anything behind a login. The other is expensive in a different way: blocking the public pages, then wondering why assistants recommend someone else.
Treat them as different pages. Close the first set. Leave the second set readable, and make it short.
Separate private pages from public ones
Private means admin, accounts, internal APIs, staging, and anything that names a customer. Public means the pages you already want a stranger to read: services, prices, hours, location, and how to get in touch.
An assistant fetching the public pages is how a customer gets an answer. An assistant fetching /admin or a draft is the problem people mean by scraping. The robots.txt file is where you say which is which.
Close the private paths
Disallow those paths for the agents you name, and for the * fallback, so a specific bot is not waved through a door you closed for everyone else. Then put a login on anything that is actually sensitive. Some crawlers ignore robots.txt. A disallow line does not replace authentication.
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /api/
Disallow: /account/
Sitemap: https://yourdomain.com/sitemap.xmlBest Practice robots.txt for the AI Age has the fuller template, including separate blocks for Googlebot and GPTBot, and how to check the file in a browser before you publish it.
Leave the public pages open
Disallow: / tells every polite crawler to skip the whole site. Search stops. Answer agents stop. The business is still online for humans who already know the URL, and absent from the reply when someone asks an assistant who to hire.
Allow the marketing pages, the service pages, and the contact page. If you need to opt out of a vendor’s training use, use that vendor’s own setting as well as robots.txt. Your public pages can stay available for answers while a training program is declined.
Give agents a file so they crawl less
A crawler that cannot find a short source of truth walks the site. It fetches menus, repeats, and old posts, looking for the price and the hours. That is the scrape that feels endless, and it still often ends in a guess.
Publish /llms.txt, a plain-text AI Website Profile: who you are, what you offer, what it costs, and how to reach you. A well-behaved agent can read that file and stop. The public pages stay open. The crawl gets smaller because the facts are no longer scattered through the design.
Fetch https://yourdomain.com/robots.txt and confirm the private paths are the ones disallowed. Then confirm /llms.txt loads. Those two files are the control and the brief.
