RobotsDisallowed is a public catalog that tracks websites and organizations explicitly blocking AI and web-scraping crawlers in their robots.txt or related mechanisms. It focuses on documenting the growing trend of content owners asserting control over how their data is used for model training and automated harvesting. The project aggregates domains, notes the targeted bots or user agents, and surfaces patterns for researchers, policymakers, and tool builders. It serves both as a transparency effort and as a resource for people designing allow/deny strategies for automated access. The dataset invites community contributions to keep the picture current as new bots emerge and policies shift. It also highlights the intersection of web standards, ethics, and AI governance by showing how site owners operationalize consent and restriction at scale.

Features

  • Curated list of domains that disallow AI or scraping bots
  • Identification of targeted user agents and blocking patterns
  • Community-updated dataset reflecting policy changes
  • Reference for researchers and builders of crawl-aware tools
  • Snapshot of evolving norms around data usage and consent
  • Lightweight format for analysis and reuse

Project Samples

Project Activity

See All Activity >

Categories

Libraries

Follow RobotsDisallowed

RobotsDisallowed Web Site

Other Useful Business Software
Build Agents and Models on One Platform Icon
Build Agents and Models on One Platform

Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
Try It Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of RobotsDisallowed!

Additional Project Details

Registered

2025-10-28