# persian360.com # # Text and data mining rights are expressly reserved under Article 4(3) of # Directive (EU) 2019/790. The reservation is stated in full at # https://persian360.com/legal/terms#tdm and in machine-readable form at # https://persian360.com/ai.txt # # Ordinary search indexing is welcome. Collecting this site to train, fine-tune, # evaluate or ground a machine-learning model is not, and is refused below by # name for every crawler that publishes one. # ── Search engines: welcome ────────────────────────────────────────────────── User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: DuckDuckBot Allow: / User-agent: Applebot Allow: / # ── Model training and bulk collection: refused ────────────────────────────── # Google's separate signal for Gemini/Vertex training. Blocking it does not # affect Googlebot or search ranking. User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Disallow: / User-agent: ChatGPT-User Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-Web Disallow: / User-agent: anthropic-ai Disallow: / User-agent: CCBot Disallow: / User-agent: PerplexityBot Disallow: / User-agent: Perplexity-User Disallow: / User-agent: Bytespider Disallow: / User-agent: meta-externalagent Disallow: / User-agent: meta-externalfetcher Disallow: / User-agent: FacebookBot Disallow: / User-agent: Amazonbot Disallow: / User-agent: cohere-ai Disallow: / User-agent: Diffbot Disallow: / User-agent: omgili Disallow: / User-agent: Timpibot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: AI2Bot Disallow: / # ── Everyone else ──────────────────────────────────────────────────────────── # Content Signals, stated so the machine-readable position and the legal one # cannot drift apart. Each line is the terms at /legal/terms#tdm, compressed: # # search=yes ordinary search indexing is welcome, and always has been # ai-input=no no grounding, retrieval-augmented generation or answer # synthesis. The terms forbid using any part of the service to # "ground a machine-learning model" and to "build a derived # dataset, index or corpus" — which is what an AI retrieval # crawler does. Blocking those agents by name above is the # enforcement; this is the declaration. # ai-train=no no training, fine-tuning, evaluation or benchmarking # # HISTORY, so this is not silently re-enabled: until 2026-08-13 Cloudflare # prepended a managed block to this file. It declared `use=reference` — express # permission for the AI grounding the terms deny — and added a second # `User-agent: *` group, so the served file both contradicted itself and was # ambiguous about which group applied. It was turned off in Cloudflare → AI # Crawl Control, and this file is now the single statement. If a managed block # ever reappears above, that setting has been switched back on. # # /catalog? stays blocked: the catalogue filters client-side and no query form # of it is canonical, so a ?-URL is a tracking parameter or a mistake. # # /account and /legal/v1.0/ were Disallowed here until 2026-08-13. They are not # any more, deliberately: both now carry a real `noindex`, and a crawler has to # be allowed to FETCH a page before it can read that header. Disallowing them # would leave the URLs indexable from inbound links with none of the # instruction attached. Block or noindex — never both on the same URL. User-agent: * Content-Signal: search=yes, ai-input=no, ai-train=no Disallow: /catalog? Allow: / Sitemap: https://persian360.com/sitemap.xml