# Upstream Fly's whole crawler policy, in one file the repo owns. Cloudflare's # managed robots.txt is deliberately disabled: a dashboard mode-switch once # dropped every training-crawler block without a diff anywhere, and this file # exists so that can never happen again. scripts/check_robots.py asserts the # SERVED file enforces all of it, whoever writes it. # # Content signals, per contentsignals.org: search = build a search index and show # links and excerpts; ai-input = ground or cite this content in AI answers; # ai-train = train or fine-tune models. Our position: cite us, don't train on us. # The artwork is original commissioned work and is not licensed as training # data; see /terms/ and /llms.txt. User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / # Search crawlers are welcome. User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: DuckDuckBot Allow: / # Search-facing AI crawlers that cite and link are welcome too. User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: Perplexity-User Allow: / User-agent: Claude-SearchBot Allow: / # Model-training crawlers are not. The artwork here is original, commissioned # work and is not licensed as training data. See /terms/. User-agent: GPTBot Disallow: / User-agent: CCBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: anthropic-ai Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Claude-Web Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: FacebookBot Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Bytespider Disallow: / User-agent: Amazonbot Disallow: / User-agent: PetalBot Disallow: / User-agent: Diffbot Disallow: / User-agent: Omgilibot Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: Timpibot Disallow: / User-agent: cohere-ai Disallow: / User-agent: AI2Bot Disallow: / Sitemap: https://www.upstreamfly.com/sitemap.xml Sitemap: https://www.upstreamfly.com/sitemap-images.xml