# ============================================================================== # Roadways Robots Policy # Allow legitimate search engines to crawl and index public timetable records. # Block automated AI scrapers, LLM training bots, and unauthorized data miners. # ============================================================================== # Search Engine Crawlers (Allowed) User-agent: Googlebot User-agent: Bingbot User-agent: YandexBot User-agent: DuckDuckBot User-agent: Baiduspider User-agent: Slurp Allow: / Disallow: /admin Disallow: /api/ # OpenAI Crawlers, Search & ChatGPT Scrapers (Blocked) User-agent: GPTBot User-agent: ChatGPT-User User-agent: OAI-SearchBot User-agent: OpenAI Disallow: / # Google AI Crawlers (Blocked while keeping Google Search allowed) User-agent: Google-Extended User-agent: GoogleOther Disallow: / # Apple Intelligence Training Crawler (Blocked while keeping Apple Search allowed) User-agent: Applebot-Extended Disallow: / # Anthropic Claude Crawlers (Blocked) User-agent: ClaudeBot User-agent: Claude-Web User-agent: anthropic-ai Disallow: / # Common Crawl & Mass Training Scrapers (Blocked) User-agent: CCBot Disallow: / # ByteDance / TikTok AI Scraper (Blocked) User-agent: Bytespider Disallow: / # Meta / Facebook AI Training (Blocked) User-agent: FacebookBot User-agent: Meta-ExternalAgent User-agent: Meta-ExternalFetcher Disallow: / # Perplexity AI Scraper (Blocked) User-agent: PerplexityBot Disallow: / # Cohere AI Training (Blocked) User-agent: cohere-ai Disallow: / # Other Commercial AI & Web Scrapers (Blocked) User-agent: MistralAI-User User-agent: AI2Bot User-agent: Diffbot User-agent: Omgilibot User-agent: Omgili User-agent: ImagesiftBot User-agent: Amazonbot User-agent: Scrapy User-agent: Seekport User-agent: YouBot User-agent: Timpibot User-agent: VelenPublicWebCrawler Disallow: / # Default Policy for all other compliant search crawlers User-agent: * Allow: / Disallow: /admin Disallow: /api/ # Master Sitemaps Sitemap: https://bustime.haryanabusinfo.in/sitemap.xml