# ====================================================================== # robots.txt — Weather Shield (SEO + AI Search optimized, safe allowances) # Updated: 2025-10-02 # Goals: # 1) Keep SEO crawlers and social previews open # 2) OPEN AI SEARCH REFERRALS (user + searchbot UAs) # 3) Explicitly disallow AI MODEL TRAINING crawlers (policy-based) # 4) Conserve resources by disallowing non-essential crawlers # ====================================================================== # --------------------------- # XML Sitemaps # --------------------------- Sitemap: https://weathershield.com/sitemap_index.xml # --------------------------- # Global baseline (site is open) # --------------------------- User-agent: * Disallow: # (Optional) Site-specific sensitive paths — add if applicable # Disallow: /cart/ # Disallow: /checkout/ # Disallow: /account/ # Disallow: /search Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php # --------------------------- # Core Search Engines (OPEN via baseline) # Google: Googlebot (+Image/Video/News/InspectionTool), GoogleOther, GoogleOther-Image, GoogleOther-Video # Bing/MSN: bingbot, msnbot, BingPreview, AdIdxBot # DuckDuckGo: DuckDuckBot # Apple: Applebot # --------------------------- # --------------------------- # Social / Link Preview Bots (OPEN via baseline) # facebookexternalhit, facebot, FacebookBot # LinkedInBot # Twitterbot # Pinterestbot # Discordbot # TelegramBot # Meta external agents (meta-externalagent/meta-externalfetcher/meta-webindexer) — left OPEN # --------------------------- # ====================================================================== # AI SEARCH REFERRALS — OPEN (no disallow lines) # PerplexityBot, Perplexity-User # ChatGPT-User # OAI-SearchBot # Claude-SearchBot, Claude-User # MistralAI-User # ====================================================================== # ====================================================================== # AI MODEL TRAINING & BROAD HARVESTERS — DISALLOW (policy-based) # Flip to allow ONLY if leadership approves corpus inclusion # ====================================================================== User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: CCBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: cohere-ai Disallow: / User-agent: cohere-training-data-crawler Disallow: / User-agent: AI2Bot Disallow: / User-agent: Ai2Bot-Dolma Disallow: / User-agent: DeepSeekBot Disallow: / User-agent: Diffbot Disallow: / User-agent: QuillBot Disallow: / User-agent: quillbot.com Disallow: / User-agent: Bytespider Disallow: / # ====================================================================== # EXTRANEOUS / SMALLER / NON-ESSENTIAL BOTS — DISALLOW to conserve resources # (Open only if you explicitly rely on them) # ====================================================================== User-agent: AddSearchBot Disallow: / User-agent: Andibot Disallow: / User-agent: Awario Disallow: / User-agent: bedrockbot Disallow: / User-agent: bigsur.ai Disallow: / User-agent: Brightbot 1.0 Disallow: / User-agent: Crawlspace Disallow: / User-agent: Datenbank Crawler Disallow: / User-agent: Devin Disallow: / User-agent: DuckAssistBot Disallow: / User-agent: Echobot Bot Disallow: / User-agent: EchoboxBot Disallow: / User-agent: Factset_spyderbot Disallow: / User-agent: FriendlyCrawler Disallow: / User-agent: Gemini-Deep-Research Disallow: / User-agent: Google-CloudVertexBot Disallow: / User-agent: Google-Firebase Disallow: / User-agent: GoogleAgent-Mariner Disallow: / # (INTENTIONALLY OPEN) GoogleOther / GoogleOther-Image / GoogleOther-Video — no disallow blocks here User-agent: iaskspider/2.0 Disallow: / User-agent: ICC-Crawler Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: img2dataset Disallow: / User-agent: ISSCyberRiskCrawler Disallow: / User-agent: Kangaroo Bot Disallow: / User-agent: LinerBot Disallow: / # (INTENTIONALLY OPEN) meta-externalagent / Meta-ExternalAgent / meta-externalfetcher / Meta-ExternalFetcher / meta-webindexer — no disallow blocks User-agent: MyCentralAIScraperBot Disallow: / User-agent: netEstate Imprint Crawler Disallow: / User-agent: NovaAct Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / # Avoid broad "OpenAI" catch-all to prevent accidental blocking of allowed referral UAs # (Do not add "User-agent: OpenAI" here) User-agent: Operator Disallow: / User-agent: PanguBot Disallow: / User-agent: Panscient Disallow: / User-agent: panscient.com Disallow: / User-agent: PetalBot # Flip to allow if Huawei/Petal Search visibility matters Disallow: / User-agent: PhindBot Disallow: / User-agent: Poseidon Research Crawler Disallow: / User-agent: QualifiedBot Disallow: / User-agent: SBIntuitionsBot Disallow: / User-agent: Scrapy Disallow: / User-agent: SemrushBot-OCOB Disallow: / User-agent: SemrushBot-SWA Disallow: / User-agent: ShapBot Disallow: / User-agent: Sidetrade indexer bot Disallow: / User-agent: TerraCotta Disallow: / User-agent: Thinkbot Disallow: / User-agent: TikTokSpider Disallow: / User-agent: Timpibot Disallow: / User-agent: VelenPublicWebCrawler Disallow: / User-agent: WARDBot Disallow: / User-agent: Webzio-Extended Disallow: / User-agent: wpbot Disallow: / User-agent: YaK Disallow: / User-agent: YandexAdditional Disallow: / User-agent: YandexAdditionalBot Disallow: / User-agent: YouBot Disallow: / # ====================================================================== # Experimental AI opt-out signals (non-standard, advisory) # These DO NOT block crawling; they signal training opt-out to participating bots. # May also inadvertently shadow-block crawl requests - use with caution. # ======================================================================