[{"data":1,"prerenderedAt":118},["ShallowReactive",2],{"seo-verification":3,"blog-crawlers-ia-search-training-agent-en":6},{"google":4,"bing":5},"EycwPY2XMyTkVzas3n1ygeNJFGAH513qrMjfDljzsMQ","",{"id":7,"slug":8,"title":9,"excerpt":10,"readTime":11,"views":12,"isPinned":13,"publishedAt":14,"category":15,"categories":21,"featuredImage":23,"bgImage":24,"posterImage":25,"relatedSolution":23,"intro":26,"sections":27,"ctaTitle":71,"ctaBody":72,"ctaButton":73,"ctaUrl":74,"relatedPosts":75},201,"crawlers-ia-search-training-agent","AI Crawlers: taking back control over your clients' sites","Cloudflare now splits AI crawlers into three families. How to configure robots.txt and Cloudflare rules to protect your clients' sites.",7,0,false,"2026-08-01T00:00:00+00:00",{"id":16,"name":17,"slug":18,"color":19,"icon":20},8,"Security & Monitoring","securite-monitoring","bg-rose-500\u002F10 text-rose-400","security",[22],{"id":16,"name":17,"slug":18,"color":19,"icon":20},null,"\u002Fblog\u002Fcovers\u002Fbg.svg","\u002Fblog\u002Fcovers\u002Fcrawlers-ia-search-training-agent-poster.svg","Since July 2026, Cloudflare AI Audit classifies AI crawlers into three distinct categories: search, training and autonomous agents. Blocking all AI bots indiscriminately can penalise indexation in AI search results, while allowing everything exposes content to being captured for model training.",[28,32,42,45,64,68],{"type":29,"title":30,"body":31},"h2","The three AI crawler families to distinguish","Cloudflare AI Audit, updated on 1 July 2026, classifies AI bots into three categories with very different intentions: search crawlers (GPTBot, PerplexityBot, Amazonbot) feed AI answers and generative search engines; training crawlers (CCBot, Common Crawl, ClaudeBot in scraping mode) collect content to build or fine-tune models; autonomous agents drive browsers without a human session to perform tasks on behalf of a third party.",{"type":33,"title":34,"items":35},"ul","Why differentiating these three families matters",[36,37,38,39,40,41],"**Visibility vs. exposure** — allowing search crawlers increases chances of being cited in AI answers; blocking training crawlers protects high-value content.","**Granular control** — robots.txt per User-agent allows a per-family policy without cutting classic SEO indexation.","**Agent attack surface** — an autonomous agent can scrape a form or trigger actions; blocking them at WAF level is independent of the rest.","**Editorial intent** — some content (paid guides, tutorials) is reserved for human readers and should not feed a training corpus without agreement.","**Server load** — training crawlers are often aggressive and do not honour Crawl-delay; identifying them allows throttling or network-level blocking.","**Auditability** — Cloudflare AI Audit traces traffic per category, making the decision visible and reversible.",{"type":29,"title":43,"body":44},"robots.txt: the first line of defence","The robots.txt file is the most readable declaration of intent. Well-configured crawlers (GPTBot, PerplexityBot, Amazonbot) honour it. Less scrupulous training crawlers (CCBot) respect it variably. Autonomous agents generally ignore it. This is why robots.txt must be complemented by active rules at the network level.",{"type":46,"title":47,"steps":48},"steps","Audit and configuration step by step",[49,52,55,58,61],{"title":50,"body":51},"Audit traffic with Cloudflare AI Audit","In the Cloudflare dashboard, open Security > Bots > AI Audit. The interface displays volume per category (search, training, agents) and per user-agent. Identify the bots present and their actual weight before configuring anything.",{"title":53,"body":54},"Declare policy per family in robots.txt","Add targeted blocks: User-agent: GPTBot \u002F Allow: \u002F for search crawlers to keep; User-agent: CCBot \u002F Disallow: \u002F for training crawlers to block; User-agent: * to maintain the default SEO rule.",{"title":56,"body":57},"Create a Cloudflare WAF rule for autonomous agents","In Security > WAF > Custom Rules, create a rule targeting cf.bot_management.js_check_failed combined with a low bot score. Choose Block or Challenge depending on your client's tolerance.",{"title":59,"body":60},"Block unruly training crawlers at the WAF","For CCBot and bots with unreliable user-agents, add a WAF rule on http.user_agent contains \"CCBot\" with action Block.",{"title":62,"body":63},"Verify classic SEO indexation is preserved","After deployment, check in Cloudflare Analytics that Googlebot, Bingbot and usual SEO crawlers are not impacted.",{"type":65,"title":66,"body":67},"tip","Do not block everything: search crawlers have value","GPTBot (OpenAI), PerplexityBot and Amazonbot feed the answers of generative search engines. A site blocked for these crawlers does not appear in AI summaries from ChatGPT or Perplexity. The recommended strategy: allow search crawlers, block training crawlers, filter autonomous agents at the WAF.",{"type":29,"title":69,"body":70},"Client sites without Cloudflare: what to do","For sites hosted without Cloudflare as proxy, robots.txt remains the only declarative lever. Complement it with User-agent directives for each known bot. On a VPS with root access, fail2ban can read access logs and ban by IP the crawlers that ignore robots.txt.","Protect your clients' sites","ServOrbit helps agencies manage hosting and security for their clients centrally. Discover how to simplify supervision and protection of your portfolio.","Solutions for agencies","\u002Fsolutions\u002Fagences",[76,91,105],{"id":77,"slug":78,"title":79,"excerpt":80,"readTime":81,"views":12,"isPinned":13,"publishedAt":82,"category":83,"categories":88,"featuredImage":23,"bgImage":24,"posterImage":90,"relatedSolution":23},165,"pourquoi-votre-site-est-plus-rapide-protege","Powered by Cloudflare: Why Your ServOrbit Site Starts Out Better Protected","ServOrbit uses Cloudflare for DNS, CDN, HTTPS and basic network protection, with no implied promise of a partnership.",6,"2026-07-05T00:00:00+00:00",{"id":84,"name":85,"slug":86,"color":87,"icon":86},9,"News","actualites","bg-sky-500\u002F10 text-sky-400",[89],{"id":84,"name":85,"slug":86,"color":87,"icon":86},"\u002Fblog\u002Fcovers\u002Fpourquoi-votre-site-est-plus-rapide-protege-poster.svg",{"id":92,"slug":93,"title":94,"excerpt":95,"readTime":16,"views":96,"isPinned":13,"publishedAt":97,"category":98,"categories":99,"featuredImage":23,"bgImage":24,"posterImage":101,"relatedSolution":102},108,"securiser-vps-crowdsec","Securing your VPS with CrowdSec","Deploy CrowdSec on your VPS to block attacks thanks to behavioral detection and a shared community blocklist.",596,"2026-03-04T00:00:00+00:00",{"id":16,"name":17,"slug":18,"color":19,"icon":20},[100],{"id":16,"name":17,"slug":18,"color":19,"icon":20},"\u002Fblog\u002Fcovers\u002Fsecuriser-vps-crowdsec-poster.svg",{"categorySlug":103,"appSlug":104},"securite","crowdsec",{"id":106,"slug":107,"title":108,"excerpt":109,"readTime":81,"views":110,"isPinned":13,"publishedAt":111,"category":112,"categories":113,"featuredImage":23,"bgImage":24,"posterImage":115,"relatedSolution":116},105,"superviser-vps-uptime-kuma","Monitoring your VPS with Uptime Kuma","Deploy Uptime Kuma on your VPS to monitor your sites and services self-hosted, with alerts and a public status page.",1520,"2026-03-07T00:00:00+00:00",{"id":16,"name":17,"slug":18,"color":19,"icon":20},[114],{"id":16,"name":17,"slug":18,"color":19,"icon":20},"\u002Fblog\u002Fcovers\u002Fsuperviser-vps-uptime-kuma-poster.svg",{"categorySlug":103,"appSlug":117},"uptime-kuma",1785628423418]