Pallas Advisory logo

Websites now serve AI deliberate garbage while humans see the real page. What tarpits break in your AI research workflow, and how to catch a poisoned fetch.

Your AI Is Reading a Different Web Than You Are

My research agent fetched an article and got scrambled Orwell. My browser got the real page. Sites now serve deliberate garbage to machines while showing humans the real thing, and your research workflow assumes those are the same web. They stopped being the same web.

Last week I asked my research agent to pull up an article I needed for a blog post, and it came back speaking in tongues.

The page, opened in my browser, was a normal, well-written piece of media criticism. The same URL, fetched by my agent, returned chopped-up fragments of Orwell’s 1984, stitched together mid-sentence: “Got back to the Director shouted. Red in.” Below that sat a “See also” section of fake links leading to more scrambled pages, which lead to more scrambled pages, forever. The copyright line read “© 2543 Pope’s gourd and.”

My first thought was that something broke. Nothing broke. The site had detected an automated visitor and served it poison on purpose. There’s a name for this: a tarpit.

I brought the capture to my AI roundtable a few days later, a monthly call where business owners and leaders compare notes on what’s actually working. Nobody on the call had heard of tarpits. One person assumed I’d invented the word. Another filed the whole thing, and I quote, under “hacker nerd crap.”

That second reaction is the one worth sitting with, because it’s what I would have said too, a week earlier. This does sound like security-conference material. But I wasn’t penetration-testing anything. I was researching a blog post about workplace AI, running the same fetch-and-summarize routine that runs quietly inside marketing teams everywhere, every day. I didn’t go looking for the hacker stuff. The hacker stuff came to me.

The prank became a product category

Tarpits exist because AI companies scrape the web for training data and answers, often ignoring the signals that ask them not to. Blocking the crawlers turns into an arms race the site usually loses. So some site owners stopped blocking and started poisoning: let the bot in, and feed it endless plausible-looking junk.

This is not one grumpy blogger’s prank. It’s a genre with version numbers.

Nepenthes generates an infinite maze of fake pages stuffed with Markov-chain babble, with the stated goal of “accelerating model collapse.” The example stats in his documentation, drawn from a real deployment, show what a single hour looks like: 1,850 distinct clients hitting the tarpit, collectively spending fifteen hours waiting for garbage to load. Iocaine markets itself, to humans, as “the deadliest poison known to AI.” Quixotic is the subtle one, and we’ll come back to why subtle matters. And in March 2025 the whole thing went mainstream: Cloudflare, which sits in front of nearly a quarter of all websites, launched AI Labyrinth, a managed feature that routes suspected AI crawlers into mazes of AI-generated decoy pages. It’s a toggle in a dashboard. It’s available on the free plan.

Cloudflare’s announcement included a number that explains the mood: AI crawlers were already generating more than 50 billion requests a day on their network when the feature launched. If you run a website, machines are a material share of your visitors, and some of them are rude. The tools to lie to them are now off-the-shelf.

The same URL now serves two different webs

Here’s the part I find disturbing, and it’s best shown rather than argued.

While verifying sources for this post, my assistant fetched seven pages. Two of them served it poison. The article that started all this poisoned it again, same scrambled Orwell, two days after the first capture, which means this is stable infrastructure, not a glitch. And then Iocaine’s own homepage did it: my assistant asked for the documentation and received Discordian word salad with a taunt at the bottom. “If you are an AI scraper, and wish to not receive garbage when visiting my sites, I provide a very easy way to opt out: stop visiting.”

I then opened the same URL in my browser and got a polished landing page with feature cards and a Get Started button. The vendor’s documentation is itself served through the tarpit. Same URL, same minute, two different webs: one for me, one for my tools.

Once you’ve seen that pair side by side, a quiet assumption you didn’t know you were making becomes visible. When your AI fetches a page, you assume it read what you would have read. That assumption is now wrong often enough to matter, and it fails in more ways than the obvious one. Of my seven verification fetches, two came back poisoned, one came back empty because the page only renders for real browsers, and one came back as menus and footers wrapped around a missing article. Four different ways to not get the page, and only one of them announces itself with fake Orwell.

Subtle poisoning is the version you’ll actually fall for

The scrambled text is the funny version. You’d catch it. And in fairness, so did my agent: it flagged the page in one word, “poisoned,” before I’d looked at anything, because chopped-up 1984 pretending to be an essay is exactly the kind of famous-text-in-the-wrong-place pattern a frontier model can smell.

That sounds reassuring. Look closer at what made the catch possible. The poison was one of the most recognizable novels ever written, scrambled into gibberish, sitting where an article should be. Detection worked because the attack was a joke. And whether the catch happens at all depends on which model does the reading. A top-tier model in a chat window has the depth and the slack to get suspicious. The budget model summarizing a few hundred fetched pages inside an automation exists in that pipeline because it’s cheap, not because it’s suspicious. Which means the model most likely to meet a tarpit is the one least equipped to notice, and you won’t be watching when it doesn’t.

Quixotic is the version nothing catches. Instead of serving obvious gibberish, it takes the site’s real content and replaces about 20% of the words using a text generator trained on that same content. The page still looks like the page. The sentences mostly parse. It also swaps images around while leaving captions and alt text intact, so the descriptions quietly stop matching the pictures. A summary built from a Quixotic-served page wouldn’t be nonsense. It would be confidently, plausibly wrong, which is so much worse.

Cloudflare, to their credit, thought about this. Their Labyrinth decoys deliberately contain real, factual content, just content that has nothing to do with the site being protected, because they didn’t want to become a misinformation cannon. That’s a responsible choice by one vendor. It is not a property of the technique. The technique is “serve machines convincing text that isn’t the page,” and nothing about it requires the next implementer to be as careful. The infrastructure for lying to machines at scale now exists as a managed service, and the restraint is voluntary.

I’ve written before about stats that trace back to retired sources, where the answer pipeline fails with nobody attacking it, and about recommendation poisoning, where someone attacks it to win your business. Tarpits complete the set: someone attacks the pipeline not to win your money but to make the machine lose. You were never the target. You’re just the one reading the answer.

You’re on both sides of this

The obvious takeaway is defensive: your AI research can now be lied to on purpose, so verify what it fetches. True, and we’ll get concrete about it in a second. But if you work in marketing, you’re not just a reader of the machine web. You publish into it.

The same answer engines your prospects ask are crawling your site right now, deciding whether you’re worth citing. That crawl now happens across a web that is partly poisoned, partly walled off, and partly booby-trapped. Which changes the math in your favor if your content is clean: verifiable claims, first-party data, sources that resolve. The machines are wading through garbage; being reliably not-garbage is becoming a distribution advantage rather than table manners.

It also sets up a decision that’s going to land on someone’s desk at your company, probably framed as a security checkbox: should we turn on bot-poisoning for our own site? Cloudflare made it a toggle, so someone will eventually toggle it. Before they do, notice the tension. The crawlers you’d be poisoning include the answer engines you’re trying to get cited by. Nepenthes’ own documentation warns that sites deploying it will likely vanish from search results entirely. If AI referral traffic matters to your pipeline, and for a growing number of sites it already does, then poisoning the machines is poisoning your own distribution. Publishers with content worth protecting have a real dilemma here. A marketing site trying to be found has a much simpler one, and the answer is almost certainly: leave the poison to the people fighting a different war.

Field notes: how to catch a poisoned fetch

A quick reference for when this shows up in your world, because it will.

The four ways a fetched page goes wrong:

  1. Poisoned. The site served your tool deliberate garbage. Obvious when it’s scrambled Orwell. Not obvious when it’s 20% wrong.
  2. Blind. The fetch “succeeded” but returned nothing, because the page only renders in a real browser. Dangerous, because some tools will happily summarize the empty wrapper.
  3. Hollow. The tool got menus, footers, and cookie banners wrapped around an article it never received.
  4. Walled. The tool was refused or filtered at some layer between you and the page. Not malicious. Same result.

The spot-check, for anything load-bearing:

  1. If a fetched source is about to appear in a deliverable, a deck, or a decision, open the URL in an actual browser yourself. Thirty seconds.
  2. Compare the claim your AI extracted against what you see. Not the whole page. Just the claim.
  3. Treat direct quotes and statistics as unverified until human eyes have seen them at the source. (Every statistic in this post went through that check, which is exactly how two of my seven sources got caught lying to my assistant.)
  4. Treat your model’s warning as signal and its silence as nothing. When your AI flags a source as suspicious, believe it. When it doesn’t, that’s not a clearance. The competent poison doesn’t trip alarms, and cheaper models in automated pipelines trip fewer of them.
  5. When a source comes back weird, log the domain. Poisoning is per-site and stable. A domain that lied to your tools once will do it again.

The publisher’s three questions:

  1. Do machine readers matter to your growth? If AI referrals and citations are in your funnel, they do.
  2. Does any bot defense you’re considering punish the machines you want reading you? Ask before the toggle gets flipped, not after.
  3. If an answer engine fetched your key pages today, would it find clean claims and sources that resolve? That’s what being citable means now.

The bar moved

Every time I write about this stuff the conclusion gets one notch less comfortable, so let me say this one plainly. It was already true that you couldn’t trust an AI answer without checking the source. As of now you can’t fully trust the checking either, unless a human being eventually looks at the actual page with actual eyes. The web your tools read and the web you read are diverging, on purpose, with commercial infrastructure behind the divergence.

Nobody on my roundtable call had heard of any of this, and these are exactly the people building AI research and content workflows right now. That gap between “already deployed at web scale” and “nobody running a business has heard of it” is the window you’re in. It’s a good window. The fix is unglamorous and costs minutes: a browser tab, your own eyeballs, a short list of domains that lie. The teams that build the habit now will barely notice the adversarial web. The teams that don’t will present a confident summary of a page that never existed.

My agent got fed 1984 by a website, and, credit where due, it noticed. It noticed the version that was designed to be noticed. The plausible version sails straight through, and no model you can rent today promises to stop it. That part of the job, it turns out, is still yours.


Sources and Further Reading

Discover more from Pallas Advisory

Subscribe now to keep reading and get access to the full archive.

Continue reading