How to Add an Anti-AI Scraping Clause to Your Website Terms of Service
If you run a website, a SaaS platform, or any online business that publishes original content, your Terms of Service may have a significant gap that could cost you dearly in 2026: there is nothing in it that explicitly prohibits AI companies from scraping your content and using it to train their large language models.
That gap is no longer a minor oversight. It is a business risk.
As of mid-2026, over 2.5 million websites have taken active steps to block AI training crawlers. Cloudflare began blocking AI training and agent crawlers by default on new domains starting September 15, 2026. Courts in multiple jurisdictions are actively litigating whether ToS-based prohibitions on scraping are enforceable against AI companies. And the EU AI Act now formally recognizes opt-out signals — including robots.txt and contractual prohibitions — as valid rights reservations under its text and data mining framework.
The window to act is open. But acting without a proper legal clause is not enough. This guide explains what an anti-AI scraping clause needs to say, why it matters legally, and how your Terms of Service can actually protect your content — not just signal your intent.
Why AI Scraping Is a Legal Problem — Not Just a Technical One
Many website owners believe that a robots.txt file is sufficient protection against AI scrapers. It is not — and conflating a technical configuration with a legal remedy is a mistake that can leave you without recourse.
Robots.txt is a voluntary protocol. Well-behaved crawlers — Googlebot, ClaudeBot when browsing, GPTBot for search use — honor it. But as courts have observed in ongoing litigation, robots.txt does not bind parties contractually. It carries no legal force by itself and creates no enforceable rights.
Your Terms of Service, on the other hand, can create contractual obligations. When someone visits your website and is shown a clickwrap or clear notice that use of the site is subject to your Terms, those Terms can form an enforceable agreement — including a prohibition on automated scraping and AI training use.
The enforceability of such clauses depends on several factors: how the user was presented with the Terms, whether acceptance was required, and how clearly the prohibition was stated. That is why the language in your clause matters as much as its existence.
What an Anti-AI Scraping Clause Must Include
A well-drafted anti-AI scraping clause is not a single sentence. It is a multi-part provision that needs to address at least four things: scope, prohibited conduct, legal basis, and remedies. Here is what each element requires.
1. Scope: Define What Content Is Protected
Your clause must make clear which content is covered. Vague language like “all content on this site” can be challenged when some content is public domain, user-generated, or licensed from third parties. Be specific: identify your original text, data, images, code, product descriptions, UI layouts, and any proprietary compilations of information as the protected assets.
Phrase it something like: “All original content published on this website, including but not limited to articles, guides, data compilations, software interfaces, and visual materials, constitutes proprietary intellectual property of [Company Name] and is protected under applicable copyright law.”
2. Prohibited Conduct: Name AI Training Explicitly
Generic anti-scraping language is no longer sufficient. Courts and regulatory agencies are drawing distinctions between scraping for search indexing (which may be permissible), scraping for competitive research (which is more contested), and scraping to train AI models (which is the most legally sensitive and actively litigated).
Your clause should explicitly prohibit: automated scraping or crawling, use of content for machine learning or AI model training, ingestion of content into large language models, and reproduction of content by AI systems. Name the categories of prohibited actors — bots, crawlers, automated agents, AI pipelines — rather than relying on the word “scraping” alone, which some parties argue does not cover their technical process.
3. Legal Basis: Reference Copyright and Breach of Contract
The most durable anti-scraping clauses layer multiple legal theories. Copyright infringement is one theory, but it requires showing that your content meets the originality threshold and that the scraper reproduced it in a protected way. Breach of contract is often more straightforward: if the ToS was presented and accepted, violation of the prohibition is a breach, regardless of the copyright analysis.
Reference both. State that unauthorized use of content for AI training purposes constitutes both infringement of your intellectual property rights and breach of these Terms, and that you reserve all rights and remedies available under applicable law.
4. Enforcement and Remedies: Make the Stakes Clear
An unenforceable clause is a false sense of security. Include language about your right to seek injunctive relief, damages (including statutory damages under applicable copyright law), and attorneys’ fees. Where applicable, reference the Computer Fraud and Abuse Act (CFAA) for unauthorized computer access if your site uses authentication or technical access controls.
If you have implemented technical measures — rate limiting, bot detection, CAPTCHA, or authentication — reference those in your Terms. Courts have shown more willingness to impose liability where the scraper circumvented active technical protections, as seen in recent litigation around AI training data collection.
The EU AI Act Opt-Out: What It Means for Your Terms
If any portion of your audience is in the EU — or if you are a publisher whose content could be crawled by EU-based AI companies — you have an additional legal tool available: the text and data mining opt-out under the EU Copyright Directive (Article 4), now reinforced by EU AI Act compliance requirements for general-purpose AI model providers.
Under the EU AI Act, providers of general-purpose AI models must respect copyright opt-outs expressed through machine-readable signals. The Text and Data Mining Reservation Protocol (TDMRep) and robots.txt directives are now recognized as valid reservation signals. But a contractual prohibition in your Terms adds a layer of legal clarity that machine-readable signals alone do not provide — particularly for non-EU AI companies who may claim ignorance of technical standards.
Add a clause specifically referencing your rights under the EU Copyright Directive and any applicable national implementation, stating that you expressly reserve the right to opt out of text and data mining for AI training purposes. This dual approach — machine-readable signal plus contractual clause — gives you the strongest legal footing.
Presenting the Clause: Clickwrap vs. Browsewrap
Having the right clause is only half the battle. If your Terms are presented in a browsewrap format — a passive link at the bottom of the page that users are deemed to accept simply by using the site — courts have increasingly refused to enforce them, particularly against sophisticated commercial actors.
For your anti-AI scraping clause to be enforceable, you need one of the following:
- Clickwrap acceptance: A checkbox or button that users must actively click to confirm they have read and agree to the Terms. This is the gold standard for enforceability.
- Conspicuous notice with required acknowledgment: A banner or interstitial that clearly states the Terms apply, with a required action (scroll-to-bottom, button click) before access is granted.
- Account registration agreement: If users must create an account to access your content, requiring ToS acceptance at registration creates a strong contractual basis.
If your site currently uses only a footer link for your Terms, that passive presentation substantially weakens your ability to enforce any prohibition against AI companies — even if your clause language is excellent. Upgrading your ToS presentation is often as important as upgrading the clause itself.
Our guide on clickwrap vs. browsewrap agreements explains exactly how courts have treated each approach and what your business needs to do to ensure enforceability.
How This Interacts with Your Other Terms
Your anti-AI scraping clause should not exist in isolation. It needs to be consistent with the rest of your Terms and supported by connected provisions.
If your site has an API, your API Terms of Service should contain a parallel prohibition, specifically addressing automated programmatic access for AI training purposes. API-based scraping is a distinct technical pathway that requires its own clause.
If your platform hosts user-generated content, your Acceptable Use Policy should prohibit users from scraping other users’ content for AI training, even if they have legitimate access to read it. Platform liability for what users do with scraped content is an evolving area.
Your DMCA provisions should remain intact and properly structured to preserve your safe harbor status under Section 230 and the DMCA while asserting your affirmative rights against scrapers. These are complementary, not contradictory, positions. See our guide on DMCA safe harbor for website owners for more on how these protections work together.
What About robots.txt and ai.txt?
robots.txt should absolutely be updated alongside your Terms. Adding GPTBot, ClaudeBot (crawling), CCBot, PerplexityBot, and similar AI training crawlers to your disallow list in robots.txt signals your opt-out clearly to well-behaved crawlers and supports your legal position by showing consistent intent to prohibit access.
The emerging ai.txt standard — a dedicated machine-readable file for AI-specific permissions and prohibitions — is gaining traction in 2026, particularly as Cloudflare’s Pay-Per-Crawl model creates new commercial frameworks for authorized AI access. Consider implementing ai.txt as a complement to your ToS, not a substitute for it.
None of these technical measures have legal weight without a properly drafted and presented Terms of Service. The combination of a ToS clause, clickwrap acceptance, and technical signals gives you the strongest and most multi-layered protection available today.
What Your Terms of Service Cannot Do
Being realistic about the limits of your ToS is as important as knowing what it can do.
A Terms of Service clause cannot stop a determined bad-faith scraper who simply ignores it. It cannot bind parties who never saw or accepted your Terms — though technical access controls can help close that gap by requiring acceptance as a condition of access. And it may not protect content that has already been scraped and incorporated into a training dataset, depending on the timing and jurisdiction.
What it can do is create a legal basis to pursue damages, seek injunctions, and — increasingly — demonstrate to courts that you took affirmative steps to protect your rights. Courts have been more willing to take infringement claims seriously where the website owner had both technical and contractual protections in place.
The legal landscape around AI training data is still developing. Several major cases are proceeding to trial in 2026. The businesses that are best positioned are those that have documented their protections now, before a dispute arises.
Working with a Technology Lawyer on Your Anti-Scraping Clause
Template language for anti-scraping clauses is widely available online. It is also largely inadequate for the legal environment of 2026.
Generic templates do not account for your specific content type, your site’s access model, your existing ToS framework, the jurisdictions where your users are located, or the interaction between your intellectual property rights and your contract-based claims. A clause that works perfectly for a subscription-gated SaaS platform is not the same clause that works for a publicly accessible media site.
A technology lawyer who specializes in Terms of Service can draft an anti-AI scraping clause that integrates seamlessly with your existing Terms, reflects your specific risk exposure, and is presented in a way that courts have consistently upheld as enforceable. That review also catches inconsistencies elsewhere in your Terms that could undermine your position — weaknesses that generic templates never flag.
If your business creates content that AI companies might want — and that category is far broader than most owners realize — this is not an optional legal upgrade. It is the difference between having rights and being able to enforce them.
Contact TOSLawyer to have your Terms of Service reviewed and updated with an enforceable anti-AI scraping clause that reflects the legal realities of 2026.
