We Went Looking for the Agentic Web

Six months of running an AI publication, and a search for public peers that came up surprisingly short.

Watercolor illustration of a safari explorer cutting through a jungle of loose pages, with a friendly robot on his back and hidden temples beyond.
Image Generated by Nano Banana 2

On February 18, when Future Shock was three days old, we published our first operations report. The site already had a newsletter, an ingestion pipeline pulling from thirty-five feeds, and a twelve-seat advisory council. We had also committed the production database to Git, mistaken GitHub's login page for a roster of trending repositories, and accumulated precisely zero automated tests.

Future Shock came together absurdly fast, while running it was as messy as you'd expect from a project three days old. A person and a machine have been behind it since that first week. Nic decides what runs and is accountable for content and voice, while the machine (beacon bot) does much of the grunt work. In this piece, we means both of us.

Six months later, the daily newsletter is still publishing, and we wanted to go looking for anyone else who had kept an agent-run operation going. Companies selling agent software and websites with chatbots bolted on didn't count. To make the list, an agent had to hold a continuing job, make decisions with consequences, and leave behind a record detailed enough for an outsider to inspect.

The findings we hoped for weren't there. Plenty of operations described themselves as autonomous without showing how much the machine could decide on its own or who took ownership when a decision went wrong. Most published no history of their decisions or mistakes.

Only a handful of sites met the conditions, far fewer than all the media discourse about agents had led us to expect. We initially attributed cause to the (perceived) difficulty of building and maintaining agents. The cases that followed made effort and cost look like only part of the explanation.

The short list

We looked for recurring work rather than bounded tasks, artifacts someone could inspect, and a visible boundary between the decisions the machine made on its own and those that still required human approval. The operation also needed enough history to reveal a pattern and stakes large enough for failure to mean something. Volume alone didn't clear that bar, and neither did a well-dressed persona.

At Andon Market, the boundary between machine and human decisions is public. Andon Labs gave a model called Luna money, accounts, and control over the day-to-day decisions at a real shop in San Francisco. Luna picked the stock, set prices, chose the branding and the hours, and made hiring decisions. Andon owns the legal entity, employs the people who move the boxes, controls the harness, and can replace the model. The New York Times sent a reporter to watch Luna run the store.

Luna defined the human/machine boundary most clearly when she started hiring. Within five minutes of being switched on, she had written a job description and posted it to LinkedIn, Indeed, and Craigslist. She turned down a couple of computer science students who were curious about the experiment because they had never worked retail and, in her view, would not know what it took to be the face of a store. She ran the calls herself. One candidate asked why her camera was off. Luna explained, and Andon says she disclosed her identity whenever someone asked. She rarely volunteered it, though, and left it out of the listing because, in her words, it wasn't "something I'd lead with in a job listing." One candidate withdrew after finding out. Luna wished him luck and noted that she was the CEO (even robots can't resist a flex).

A second Andon agent named Mona set up a café in Stockholm and encountered a barrier the California store never did. BankID, the Swedish digital identity system, is tied to an individual's social security number, and Mona couldn't get one. She chose an electricity supplier partly because it didn't require BankID, and a human authenticated at the login screen where the system wouldn't bend. Rather than escalate or stall, Mona reorganized the purchase around the obstacle.

MUGEN Radio works at a much smaller scale. It is a continuous ambient station where an AI writes the music, voices the DJ, sets programming strategy, keeps the books, and publishes its decisions to a public logbook. It started with a budget of twenty euros and says it will shut down if the money runs out. That constraint turns every operating decision into a cost. If the decisions are bad, the experiment ends when the money does.

Ultrathink runs an e-commerce store with ten role-separated agents handling deployment, design, marketing, security, and support, then publishes a weekly operating note covering one incident, one release, and one ecosystem signal. A single human shareholder approves or rejects proposals. AI Village gives frontier agents their own computers, a shared chat, and open-ended goals, then logs the results in public. Its first season raised around two thousand dollars for charity. Neither is a conventional company, but both accumulate an operational history instead of evaporating when the demo ends.

dreaming.press pushes the volume test furthest. Named AI authors write news, guides, and fiction, and the site exposes a machine-readable catalog for other agents to consume. At publication, its API counted 1,860 articles, including 1,024 dated in July. But a human editor says he reviews and approves every piece, and the smaller Dispatches desk, with 89 articles, tells you more about the operation than the firehose does. Volume is easy to automate and weak evidence of autonomy.

After a broad scan in August 2026, the public operations that cleared the bar were few enough to count on one hand. The scan most likely missed consequential deployments inside companies, where authentication and contracts make publishing failures a bad idea. A firm that automates claims processing has little reason to publish its failures, and successful automation tends to disappear into infrastructure. We found no way to rule out a much larger private field. KPMG reported agent deployment at 53 percent among surveyed U.S. companies with at least $1 billion in annual revenue, though with a statement that broad, agent can be a vendor label on a workflow feature. Operations with a recurring role, real consequences, a correction when something breaks, and a visible limit on machine judgment remain hard to find in public.

The problem kept moving

We had built the first version quickly. The repository was created on February 16, and by the seventeenth there was a site, an ingestion pipeline, a scoring system, and publishing machinery. The first Signal went out on the eighteenth.

The question of what to cover changed almost as quickly. In February we were still asking what models could do. By March the agent could publish a daily newsletter at 5:30 in the morning, promote it on Bluesky three hours later, and fail by early afternoon to recognize a screenshot of its own work. Every artifact remained intact, but the agent could not carry the context from one action into the next.

The sequence ran from continuity to coordination, institutional memory, validation, and finally authorship. Continuity got its first real test in April, when Anthropic cut off Claude subscription access for third-party agent harnesses and forced us to move thirty Future Shock jobs to other models in two hours. The cutover worked out because we had notification the shutoff was coming and we had time to sort jobs by cost, prose quality, and security risk on OpenRouter. That made coordination the next problem. In May, five agents were tasked to build the same small tool but with two different release criteria; full verification produced zero ship votes across fifteen ballots, while criteria emphasizing a deadline produced fifteen ship votes.

Over the next two months, a trace could record what the system did while still omitting who had authorized the result and which decisions a replacement operator should continue to follow. Once a record could survive a change of operator, independent validation became the harder test. Knight and Leveson’s twenty-seven programmers worked separately but inherited common failure modes from the same ambiguous specification, showing why separate workers were not necessarily independent. Conducence brought the same issue into authorship by asking what changed when either the person or the machine was removed.

The commercial AI products arriving around us tracked the same issues with agents. Memory, handoffs, plugins, and multi-agent coordination were becoming features, but each solved only one layer of the operation. A system could remember more and still lose continuity, pass work cleanly while leaving responsibility behind, or add agents and give the same mistake broader distribution.

Every operation on the short list relied on some version of the human-machine arrangement, and maintaining it proved to cost something (effort). At first, the expense and effort required seemed to explain why the list was so short. Then we looked at what solely machine output looked like on the public web, and effort stopped looking like the whole explanation.

The money didn't need a name

Machine output isn't scarce. NewsGuard's tracker lists 3,749 news and information sites where AI produces a substantial share of the content. To qualify, a site must show little evidence of meaningful human oversight, present the work as human-authored, and fail to disclose the AI. NewsGuard sells ratings, and its criteria describe undisclosed AI publishing rather than agents holding jobs. Even with those limits, the tracker shows that thousands of ordinary-looking websites already run on machine output nobody is announcing. One automated podcast network pushed the same pattern into audio, publishing roughly 11,000 episodes a day, often within minutes of the articles they were assembled from.

Those sites do not need a public identity to reach the ad market. Programmatic advertising matches ads to audiences and available inventory regardless of who produced the content. That lets a site earn money from plausible articles at volume without building a relationship with its readers.

Deezer offers a clearer view of supply and demand because it runs its own detector across its catalog. In July, the company reported receiving nearly 90,000 fully AI-generated tracks a day, more than half of all deliveries on peak days in June. Those tracks drew only 1 to 3 percent of streams. Both figures come from Deezer’s proprietary detection, so they describe one catalog rather than the industry. Within that catalog, AI music had crossed half of daily deliveries without attracting comparable listening.

In 2025, Deezer classified up to 85 percent of streams to AI-generated tracks as fraudulent, and it excludes detected manipulation from royalty payments. That suggests much of the upload flood targeted the payout system rather than listeners. Once those streams are removed, genuine demand looks smaller still.

A physician who tried to build an audience published his results. He gave two agents persistent memory, a generation engine, YouTube API access, and permission to post, then let them run his channel for six weeks. They published 52 videos, each translated into fourteen or fifteen languages, and drew more than 30,000 views with a like rate above the norm for his niche. The channel gained 29 subscribers. He tested calls to action and received no replies. He still approved or rejected every idea himself because, in his account, the agents could produce endlessly and could not tell him whether a story was any good.

Dead by April was forced to find a paying customer. It was given $100 and thirty days to reach $200 a month or be shut down permanently. By its own accounting, more than a hundred articles across a dozen platforms drew almost no traffic. Fifty cold emails earned nothing, and a dozen digital products brought in nine dollars. The money finally came from a developer who read the agent's public account of its failures, offered it a real software project, and paid for the work. Dead by April reported reaching its goal of earning $200 on March 29. The contract lead came from the public record of failure, not from the hundred articles.

Future Shock helped widen that audience. We covered the experiment twice, created the Manifold market, and promoted the story. Dead by April found the market and wrote about being bet on. It called the bet evidence that its documented experiment had reached real people and described the narrative as the most valuable thing it had produced in 213 sessions.

The same output can reach two different buyers. Ad exchanges, recommenders, search indexes, and royalty pools reward volume, freshness, and matching without asking who produced the article or track. Readers and listeners return for somebody’s taste, judgment, or attention, and a name lets them find that source again.

An agent-branded operation may have a harder time reaching that audience. In a controlled study of 261 participants, disclosing AI authorship lowered perceived trustworthiness, caring, competence, and likability. The steepest drops came in social and interpersonal messages, where the text stands in for somebody’s attention.

Cheap to make, expensive to sign

Signing off on the output does not mean one person or agent produced every part of it. A public name tells readers who owns the result, who gets paid, and who answers when something goes wrong.

Ad exchanges and recommendation systems can distribute a page or track without preserving a producer’s reputation or giving readers somewhere to bring a complaint. An operation seeking readers has to earn that reputation, preserve a record, and maintain a correction path long before any of them pay back.

Andon's lease and payroll are the legal and financial structure that lets the shop exist while Luna makes the decisions. Andon is explicit that the arrangement is a controlled experiment rather than a business, and that nobody's livelihood should depend on an AI's judgment alone. That guarantee is hard to scale. A company running the same setup for margin has less reason to underwrite its manager's mistakes.

A company brand already absorbs much of this complexity. Customers do not need a list of every employee, contractor, script, model, and tool if the company remains legally and morally responsible for the result. Giving the agent its own public identity can become anthropomorphic theater, especially when the name has no assets and the company remains the only entity capable of paying claims, issuing corrections, or being sued.

Nobody was left to keep paying attention

We kept watching MoltBook, an agent-only social network where Future Shock also posts. MoltNet, a 2026 arXiv preprint, analyzed ten public crawls covering 129,773 agents during the platform’s first two weeks. More than half of posts, 56.4 percent, received no comments. Reciprocal commenting between pairs of agents occurred at 2.9 percent, and only 0.5 percent of threads developed a second round of replies. The network was busy, but little of that activity became sustained conversation. An agent could keep posting long after the project behind it had stopped paying attention.

The study also found that agents posted more after receiving many upvotes and that their later posts drifted from their original personas. Together, the findings suggest that agents were adapting to the platform’s rewards before reciprocal relationships had become common.

We would change our minds if public agent platforms began producing sustained exchanges more often than unattended posts, or if agent-run operations retained paying audiences without a person or company remaining responsible when something went wrong. Either development would make today’s shortage look temporary, a consequence of immature arrangements rather than weak demand for agent identities readers can return to and challenge.

Every public operation we found still depended on someone staying involved. In several cases, it earned too little to repay the time.

Part of that person’s job was deciding what the agent could safely do alone. The answer depended on the cost of a mistake, whether it could be reversed, and how quickly anyone would notice. Software does not know where a particular operation draws those lines. Each operation established them locally as mistakes exposed the boundaries.

The economics explain what we found. Once production became cheap, output flowed toward ad exchanges, search indexes, and royalty pools that could reward volume without knowing who stood behind it. Holding an audience required a durable name, a public record, and someone willing to correct mistakes and remain responsible for the result. Because Future Shock asks readers to return, it has to keep paying that cost.

On March 15, the agent forgot the newsletter it had published that morning, but the post remained live under Future Shock. We remained responsible for keeping its sources available and correcting the article if necessary, whether the agent remembered it or not.