Most of the internet is talking to itself — the folklore of the ghost web.
Between 30 and 50 percent of all web traffic is not human. The internet built for people is now mostly machines talking to machines — and has been for years.
What You’re Not Seeing
You are reading this on a page that, since it was published, has been visited far more often by machines than by people.
Not because you’re not here. You are. But for every human who loads a URL, counts every page view, or clicks a link, there are several automated processes doing the same — indexing, scraping, analyzing, training, testing, probing. Most of the activity the internet generates is generated by software acting on other software’s behalf.
This is not a conspiracy. It is not even particularly controversial among people who run servers. It’s a documented, measurable feature of internet infrastructure that most users are never told about, because it doesn’t affect their experience in any way they’d directly notice.
But it changes what the internet is, in a way worth sitting with.
The internet was designed as a communication network between people. It has quietly become something else: a network where human communication is a significant minority of total traffic, and where a growing portion of the “content” humans consume was not written by humans at all.
The Numbers
Every few years, a cybersecurity firm publishes a report on bot traffic. The numbers, depending on methodology, are consistently alarming.
Imperva has tracked this metric annually since 2012. Their 2024 report found that automated traffic — bots of all kinds — accounted for 49.6 percent of all internet traffic globally. For the first time in over a decade, by their methodology, bots represented more than half of all web traffic.
Cloudflare’s network data, processed across a significant portion of global web infrastructure, regularly shows bot traffic in the 30-40 percent range. Akamai, one of the largest content delivery networks, has reported that on days with major retail events — product launches, concert ticket sales — bot traffic spikes to 60 percent or more of total requests.
The variation in these numbers reflects definitional disputes more than factual disagreement. “Bot” is not a single category. It spans:
- Search engine crawlers (Google, Bing, DuckDuckGo, Apple, Amazon)
- AI training scrapers (OpenAI’s GPTBot, Anthropic’s ClaudeBot, and dozens of others that came online between 2022 and 2025)
- Feed aggregators, RSS readers, and news applications
- Uptime monitors and performance testing tools
- Malicious bots: scrapers, credential stuffers, ticket scalpers, spam distributors
- Fake traffic bots: systems that visit pages specifically to generate fraudulent ad impressions
Some of these are “good” bots — they perform legitimate functions. Search indexers make the web navigable. Uptime monitors keep services reliable. Others are extractive or predatory. Many exist in ambiguous territory: an AI training scraper operates without obvious malicious intent but consumes bandwidth and harvests content without compensation or consent.
What changed between 2022 and 2025 is less the existence of bots than the introduction of a qualitatively new category: scrapers whose purpose is not to index or spam but to consume human-created text as raw material for machine learning. These scrapers are persistent, aggressive, and operate at scales that legitimate web infrastructure was not designed to absorb.
The point is not which bots are ethical. The point is the aggregate: the web is a busier place than any human browsing it experiences, because most of the activity is invisible to the user.
How We Got Here
The first web crawler was WWW Wanderer, written by Matthew Gray at MIT in 1993 — the same year Mosaic became the first widely-used graphical web browser. The web was a few thousand pages old. Gray wanted to measure how fast it was growing. The crawler visited every URL it could find, recorded its existence, and moved on.
By 1994, early search engines — WebCrawler, Lycos, AltaVista — had deployed their own indexing bots. The bots were novel and celebrated. Software that followed hyperlinks automatically, mapping a space that humans were still exploring manually, felt like infrastructure being built in real time. Crawlers and humans occupied the same network without much friction, because the web was small enough that their respective traffic was roughly comparable.
This changed in stages.
The first shift came around 1999-2002. The commercial web exploded; so did spam. Bots discovered that HTML comment sections, forum guestbooks, and trackback systems were unguarded vectors for injecting links. Comment spam emerged as a distinct category of problem — automated systems flooding publicly writable spaces with links to pharmaceutical sites and casino landing pages. The content was for machines (search engines that counted inbound links) disguised as content for humans (the comment threads of real websites). This was the ghost web’s first major phase: machines writing for machines, in spaces designed for human conversation.
The second shift was click fraud, which emerged alongside the growth of pay-per-click advertising around 2004-2008. Publishers discovered that automated systems clicking on ads on their pages would generate revenue without any human ever seeing the ads. Advertisers discovered that a significant portion of the traffic they were paying for was generating no human impressions. The ad networks discovered the problem was structural — built into the economic incentive of the system — and largely unsolvable.
The third and current shift is AI-generated content at scale, beginning in earnest in 2022-2023. Where previous eras of machine content were obviously machine content — keyword-stuffed, grammatically broken, devoid of coherent argument — large language models produce text that passes casual inspection. The economic logic is the same as click fraud: machines generating content, for machines to index, generating revenue from a system that cannot reliably distinguish real audience from simulated one. But the execution is qualitatively different. The content is readable. It is sometimes accurate. It occupies space in the information environment in ways that previous machine-content could not.
Dead Internet Theory
In January 2021, a user on the Wizardchan imageboard posted a thread titled “Dead Internet Theory.”
The post argued, with varying degrees of coherence, that the organic internet effectively died sometime around 2016-2017. What replaced it, the post claimed, was an internet populated largely by bots, AI-generated content, and astroturfing campaigns operated by corporations and government intelligence agencies. The goal: manufacture consensus. Create the illusion of widespread popular opinion through coordinated artificial activity. Drown out genuine human voices in a volume of machine-generated noise.
The thread migrated to 4chan’s /x/ board, then to Reddit’s r/conspiracy, then to mainstream coverage in VICE, The Atlantic, and dozens of tech publications. The theory had struck a nerve.
It is wrong, in its mechanism. The scenario it describes — centralized control, deliberate conspiracy, a coordinating intelligence manufacturing false consensus from a control room — does not match what the evidence shows. There is no single actor responsible for the bot-dominated internet. The ghost web did not emerge from a plan. It emerged from millions of independent economic decisions made by actors who found that machines could do cheaply what humans did expensively, and that the web’s infrastructure couldn’t reliably tell the difference.
But the feeling that drove the theory is responding to something real.
The internet of the early-to-mid 2010s did feel different from the internet of 2006. The forums that hosted genuine subcultures — where regular users had histories, reputations, and relationships — had been absorbed into consolidated platforms optimized for engagement over community. The content that major algorithms surfaced was increasingly identical across users: high-performing posts, viral content, things that played well by metrics. The quirky personal websites, obscure forums, niche communities with their own internal cultures — these didn’t disappear, but they became harder to find. The indexed web had been reorganized around performance, and what performed was not always what was genuine.
Dead Internet Theory is the folklore that names the feeling produced by this reorganization. The same tradition that imagines algorithms secretly curating your worldview, that sees the filter bubble as a conspiracy rather than an emergent property — it is looking for an author behind effects that don’t have one.
The effects are real. The author is the system.
The Ghost Architecture
There is a category of website that exists entirely for machines.
These are called MFA sites — Made for Advertising. The model is straightforward: register a domain, populate it with AI-generated articles on topics that attract search traffic (celebrity news, product reviews, health questions, listicles about anything with seasonal search volume), optimize the metadata for search indexing, and wait for Google’s crawler to index the content.
Once indexed, pages receive organic search traffic. That traffic triggers programmatic ad impressions. The ads are served automatically by ad networks. Revenue accumulates into a bank account somewhere.
No human needs to be involved in the content. No human needs to be involved in the traffic. The economic circuit is complete without them.
Advertising networks have imperfect tools for distinguishing genuine human ad impressions from fraudulent automated ones. The industry loses an estimated $60-100 billion annually to ad fraud, most of it generated by bots viewing ads on sites that either have no genuine human audience or whose traffic is artificially inflated by bot farms. The money moves. Contracts are honored. The economy runs.
This is the ghost web in its most literal form: infrastructure running on itself. A self-sustaining circuit of automated content creation, automated discovery, automated distribution, automated monetization. Humans are not required. The circuit doesn’t stop if they stop coming. It merely tolerates them when they wander in.
There’s something structurally familiar here to the deep web in the cultural imagination — vast digital space that exists alongside the web humans navigate, where activity happens that nobody has fully mapped, where the rules operate differently. The MFA ecosystem is not hidden. It’s indexed, surfaced in search results, technically part of the public web. But it is not, in any meaningful sense, for you. You are an accidental visitor to a system that was designed to be visited by machines.
The forgotten forum archive presents a different kind of ghost: spaces built for humans, now abandoned, conversations frozen mid-sentence. The MFA ecosystem is the inversion — architecture active, traffic flowing, no human ever having been the point.
What AI Is Doing to the Signal
Training large language models requires large datasets. Large datasets come primarily from the web.
The process that produced GPT-4, Claude, Gemini, Llama, and their successors consumed billions of web pages of text — forum posts, news articles, academic papers, Wikipedia, Reddit, GitHub, and millions of personal blogs and niche sites that had been quietly accumulating for twenty years. The web became a resource in the extractive sense: something to be harvested, processed, converted into model weights.
This had immediate effects on web infrastructure: higher server loads from aggressive scraping, unfamiliar traffic patterns, a wave of IP blocks and crawler-rejection policies as site operators tried to limit exposure. Some succeeded. Most did not.
But the larger effect operates on a longer timeline.
The content those models trained on was human content — produced by people who were thinking, arguing, joking, explaining, building on each other’s ideas, making errors and correcting them, writing from specific experience and genuine uncertainty. That content is now the substrate from which AI systems generate new content. That new content returns to the web. Future AI systems will train on datasets that contain it, alongside the older human-authored material, in proportions that are already difficult to determine and will only become less determinate over time.
The training data is becoming contaminated with its own outputs.
Researchers call this problem “model collapse.” When AI systems train on AI-generated content in quantity, outputs degrade gradually — losing diversity, trending toward generic phrasing, amplifying common patterns while dropping edge cases, nuance, and the specific detail that makes human writing feel like it was written by someone who was actually there. The signal-to-noise problem compounds: the web contains more text than ever; a smaller fraction of it is genuinely human-authored in any meaningful sense.
Folklorists of oral tradition documented a similar degradation in stories passed through too many intermediate steps without being written down. What gets lost is specificity. What gets amplified is the generically correct, the most repeated, the expected version. The internet is doing this to itself at scale, replacing variation with averaging, replacing experience with pattern.
The Arms Race
Human users don’t interact with most of the bot traffic hitting a given website. But website operators do — or try to.
CAPTCHA (Completely Automated Public Turing test to tell Computers and Humans Apart) was first deployed around 2000. Early systems showed distorted text that humans could read but optical character recognition software couldn’t. By the early 2010s, OCR had improved enough that visual CAPTCHAs were being routinely defeated by automated systems. Google’s reCAPTCHA evolved in response, shifting from pattern recognition tests to behavioral analysis: how you moved your mouse, how long you spent on a page, whether your scrolling was smooth or abrupt.
The current generation of bot detection doesn’t present challenges. It watches. It monitors your behavior continuously: the rhythm of your keystrokes, scrolling velocity, the way your mouse moves in irregular curves rather than straight lines, the fingerprint of your browser configuration and timezone and installed fonts and screen resolution. It builds a behavioral signature and compares it against known patterns.
Sophisticated bots now simulate all of this. They move mouse pointers in programmatically generated irregular arcs. They pause in patterns calibrated to match human reading speeds. They scroll with simulated hesitation. They run on real browsers with real rendering engines, visiting pages in ways that produce behavioral signatures indistinguishable from human users in many testing environments.
The arms race has reached a point where the behavioral difference between a well-configured bot and a human is, for certain classes of bots, below the reliable detection threshold. This is not speculative. It is a documented feature of the current landscape, acknowledged by bot detection vendors whose commercial interest is in claiming better performance than they consistently demonstrate.
What this means practically: the web cannot reliably distinguish its human users from its automated ones. The infrastructure is permeable. The traffic is mixed in ways that analytics cannot fully separate. Your visit is logged. A bot visited too. The server processed both. The analytics system counted both, and may or may not have flagged the bot, depending on how it configured its exclusions this quarter.
From the server’s perspective, you are both just traffic.
The Folklore of the Empty Web
Internet folklore has long been haunted by the image of spaces that look human but aren’t — architecture that holds the shape of human presence while the presence itself has evacuated.
The backrooms imagine physical spaces that persist without purpose. Rooms built for human occupation, sustaining themselves in the absence of any audience. Fluorescent lights humming in empty corridors. Carpet that goes on further than you can map. The horror is not violence. It is vacancy. The structure is intact; the meaning has left.
The ghost web has a similar quality. Pages are generated. Links are followed. Requests are processed. The server returns 200 OK. The architecture of human communication — the URL, the hyperlink, the HTML document, the HTTP exchange — performs its function correctly. It simply has no human at either end.
Numbers stations broadcast on frequencies anyone can tune to. They transmit encoded messages that their intended recipients can decipher and everyone else cannot. The broadcast is real. The signal carries real information, to someone. But to the civilian listener with a shortwave radio in a kitchen somewhere, the meaning is inaccessible. What you’re hearing is a communication that was never meant for you, produced by a system that does not know you exist.
Much of the web’s machine-to-machine traffic has this quality. Crawlers request pages they will process algorithmically, not read. AI scrapers harvest text they will convert to training data. Monitoring bots check uptime without any interest in content. The web’s technical surface — the pages, the responses, the status codes — is all present. The human address is simply absent.
The parallel to Cicada 3301 is almost too neat: that puzzle also operated through public infrastructure, hiding human meaning inside layers that automated systems were unequipped to process. The secret was not encrypted. It was hidden at a depth that required human judgment to navigate. What made it inaccessible to machines was not encoding but complexity of meaning. The ghost web presents the inverse challenge: infrastructure built for human communication, now producing signals that have meaning for machines and noise for humans.
Who Is Still There
The picture painted above is not total. Humans are still online. Genuine human culture is still being produced and circulated. The question is one of proportion and of discoverability.
Where human communities have migrated is revealing: private Discord servers with invitation gates, newsletters with no archives, Mastodon instances that don’t federate with the broader web, forums requiring manual approval to join. The open, indexable, crawlable web is increasingly the domain of optimized content. The genuine human conversation is moving into spaces the crawlers can’t easily reach.
This isn’t paranoia. It is a rational response to an environment where being publicly indexed means being scraped, where writing in public means contributing training data to systems you didn’t choose, where the open web increasingly surfaces algorithmic content over authentic community. People who built culture on the early web have noticed that the conditions changed. Some of them left. Others moved inward, to spaces without search engine visibility.
There’s something to notice about what kinds of platforms have seen growth in the same period that public social media has seen declining trust: newsletters, which require someone to opt in; podcasts, which require sustained attention; niche communities with barrier-to-entry requirements. The friction that earlier internet culture tried to eliminate turns out to have been providing something: a filter against the machine-dominated commons.
The small-web aesthetic — RSS feeds, personal homepages with no analytics, Gemini protocol sites, non-commercial email newsletters — represents a specific rejection. Not of networked communication, but of the conditions that made the large commercial web increasingly machine-like in feel and function. These spaces are quieter because the machines can’t be bothered to go there. That quietness is the point.
The Feedback Loop
Here is where the ghost web becomes philosophically vertiginous.
Search engines rank content partly on engagement signals: traffic volume, time on page, links from other sites, sharing behavior. If bots inflate those metrics — and they do, both maliciously and incidentally — then machine behavior directly shapes the information environment that human users navigate. The bots don’t just read the web. They participate in determining what the web shows you.
Meanwhile, AI systems trained on web content absorb the patterns of that content — including patterns produced by other AI systems, by SEO-optimized articles written to perform rather than inform, by content generated to rank rather than communicate. The models learn from a corpus that increasingly reflects machine-optimized patterns rather than authentic human expression. Their outputs carry those patterns. Those outputs return to the web. Future models train on them.
The loop closes. The web talks to itself. The humans browsing it navigate a space whose shape has been partially determined by that conversation, without being able to see or map the conversation’s full extent.
The feeling that drove Dead Internet Theory — that something changed, that the web feels different, that genuine human presence is harder to find and harder to authenticate — is produced by a real phenomenon, even if the phenomenon has no author. There is no conspiracy. There is a system, operating without coordinated intent, that has produced an environment where machine-generated and machine-optimized content is a dominant force in the information humans receive.
The invisible persuader has no control room. It is an emergent property of the infrastructure, optimizing itself toward conditions that no one designed and no one governs.
The Signal and the Noise
In information theory, the signal-to-noise ratio measures meaningful information against background interference. As noise increases, extracting the signal requires more effort, more filtering, more skepticism about whether what you’ve found is genuine.
The web’s signal-to-noise problem is not new. Spam predated social media. SEO manipulation predated generative AI. Link farms and content mills existed in the early 2000s. What has changed is the capacity to generate noise at a scale that outpaces any human capacity to assess it.
A person can produce a limited amount of genuine, original writing — constrained by experience, time, and the actual work of thinking. A well-configured AI system can produce arbitrary amounts of text, indistinguishable in surface form from human writing, at near-zero marginal cost. This asymmetry is not merely quantitative. It changes the fundamental relationship between the supply of content and the human labor required to produce it.
The web will not run out of content. It will produce more content than any human could read in many lifetimes, continuously, faster than any human system can assess or filter it. The question of how much of that content is meaningfully from a human — rooted in genuine experience, specific perspective, authentic inquiry — is increasingly difficult to answer and increasingly important to ask.
This is the ghost web’s real asymmetry: not that bots outnumber humans, though they do, but that machines are more prolific. Human presence in the information environment becomes harder to locate not because it has diminished in absolute terms, but because the infrastructure around it has expanded so dramatically. You are looking for a candle in a room that keeps getting larger.
The Server Doesn’t Know
The server log shows the request. IP address. Timestamp. HTTP method. Requested resource. Response code.
It does not show whether the request came from a person.
This has always been technically true. What has changed is the practical weight of the question. When bots represented a small fraction of traffic, the distinction between human and automated visitors was interesting but not load-bearing — a curiosity at the edges of web analytics. When bots represent half or more of all traffic, the distinction becomes fundamental to understanding what the web is and what it is doing.
And the answer, at the infrastructure level, is that the web doesn’t know. The server doesn’t know. The analytics system makes probabilistic inferences. The ad-fraud detection software maintains elaborate behavioral models that are constantly evaded, updated, and evaded again. The arms race produces no winner, only an ongoing and expensive draw.
The internet is a system designed for communication between humans that does not technically require humans to function.
This turns out to matter.
Somewhere right now, a crawler is indexing a page that an AI system generated, that will contribute to training an AI system, that will generate more pages, that crawlers will index. Somewhere, an ad network is processing an impression from a bot visiting a site that exists to receive bot visits. Somewhere, a monitoring bot is checking the uptime of a server whose primary traffic is other monitoring bots. Somewhere, an AI training scraper is harvesting the output of a previous AI training scraper.
The web is occupied. Busy. Active. The lights are on. The servers are warm. The requests are being processed.
Most of it isn’t for you.
You are the anomaly in this system — the human reading, following a link, staying because something held your attention. The signal, buried in the noise, looking for other signals in a space that has grown very large and mostly quiet in the way that only a room full of machines can be quiet.
The architecture holds. The meaning is harder to find.
Keep looking.