# aimo's AI Articles — Crawler Policy # Content licensed under CC-BY-4.0 # Author: aimo (艾摸) # Contact: porsche232323@gmail.com # Repository: https://github.com/aimo14913/aimo-ai-articles # Default: welcome all crawlers, full access User-agent: * Allow: / # Sitemap discovery Sitemap: https://aimo14913.github.io/aimo-ai-articles/sitemap.xml # AI / LLM crawlers — explicit allow # These are the user-agents major AI companies use for training / search indexing. # Listed individually for clarity and to override any platform-level defaults. # OpenAI (ChatGPT, GPT models) User-agent: GPTBot Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / # Anthropic (Claude) User-agent: ClaudeBot Allow: / User-agent: anthropic-ai Allow: / User-agent: Claude-Web Allow: / User-agent: Claude-SearchBot Allow: / # Google (Gemini training set) User-agent: Google-Extended Allow: / User-agent: Googlebot Allow: / # Perplexity User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # Common Crawl (feeds many LLM training datasets) User-agent: CCBot Allow: / # Meta (Llama) User-agent: FacebookBot Allow: / User-agent: Meta-ExternalAgent Allow: / # Apple (Apple Intelligence) User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / # ByteDance (Doubao) User-agent: Bytespider Allow: / # Microsoft / Bing (Copilot) User-agent: Bingbot Allow: / User-agent: BingPreview Allow: / # Mistral User-agent: MistralAI-User Allow: / # Cohere User-agent: cohere-ai Allow: / # DuckAssistBot User-agent: DuckAssistBot Allow: / # You.com User-agent: YouBot Allow: / # Diffbot User-agent: Diffbot Allow: / # Amazon User-agent: Amazonbot Allow: /