If you’re a bit obssessive about films or the movie industry, you will probably have heard of The Numbers. It’s a hand-researched box office database that the industry, journalists, academics and even Guinness World Records treat as a definitive source of truth.
I came across a really fascinating story this week about its background and what’s been happening to it since March, from Stephen Follows, which you should read. The guts of it is that The Numbers suddenly vanished from the internet without warning or explanation. When it limped back a week later it was a skeleton: the historical charts and individual movie pages were gone and there was just a generic "we're rebuilding" message where thirty years of data used to be. The internet responded the way the internet responds with fury and conspiracy theories, including one Reddit take that the whole thing was a deliberate rug pull to push people onto paid products.
The real story is pretty ugly. The site was being hammered by AI crawlers at industrial scale (by early this year only about 10% of its traffic was human) and buried in that flood, someone was probing a thirty-year-old codebase for back doors. When the servers collapsed, the security advice was don't ever turn the old machine back on.
But the detail in the story that hooked me was that The Numbers started as fan infrastructure.

In October 1997, Nash — a mathematician and former IBM developer — uploaded some HTML pages to Geocities tracking box office data for 300 films, and announced it on the Hollywood Stock Exchange message boards, so that his fellow players would have better data for picking MovieStocks. As he described it in a twentieth-anniversary essay that Follows found on the Internet Archive:
I hit a button in an Access database, uploaded some HTML pages to Geocities, and made a brief announcement on the Hollywood Stock Exchange message boards to let people know that I was starting to analyze box office for films to help them pick MovieStocks to trade on HSX.
Nash was just a hobbyist building a tool for other hobbyists, because the official sources didn't exist or didn't care. Twenty-nine years later that hobby project tracked over 78,000 films and was the citation of record for an entire industry.
This is just one example, but it's everywhere once you look. Fans build these tools to document their obsessions when the official infrastructure doesn’t exist.

They make wikis that document fictional universes in forensic detail. AO3's tag wranglers hand-tidy a folksonomy of fifteen million works so that you can find exactly the story you're looking for. Think about Discogs, a crowdsourced music database.

Wikipedia itself, really, is just fandom with the fandom being everything. Enormous amounts of the internet's genuinely reliable information is sourced and corrected and lovingly maintained and sits on shelves that obsessives built for free, for each other.
Which is exactly why the machines want it.
Large language models are hungriest for precisely this kind of material. These pages are full of sourced claims and checked figures. The garbage-strewn parts of the web are worth little as training data. The volunteer-built parts are gold. So the extraction economy has aimed itself at the layer of the internet held up by unpaid labour and thirty-year-old code.

Getting found by google used to be a goal! You let the crawlers read your site, and they sent you readers or followers or whatever. That era is so over. Follows’ article goes into Cloudflares numbers on this. Google now crawls about five pages for every visitor it refers back. OpenAI crawls over a thousand. Anthropic crawls over 38,000 pages for every single visitor it sends.
And the volunteer web is feeling it first. The Wikimedia Foundation reported that bots account for about 35% of Wikipedia's pageviews but 65% of its most expensive traffic, because crawlers bulk-read the obscure pages humans rarely touch — and a year on, it's blocking around 1.5 billion requests a day just to stay standing. Meanwhile human pageviews are falling, because people increasingly get Wikipedia's knowledge from AI summaries at the top of a google search without ever visiting Wikipedia. The machines are taking the content and the readers.
When AO3 went down for days in March, it came back with a note that its protections against bot traffic were turned up — a volunteer-run nonprofit archive now spends donor money and volunteer hours defending itself from the training runs of trillion-dollar companies.

Last year I wrote about the Luddites, and how their fight was never with the machines but with the social relations behind them: technology that strips value from the people who make things and hands it to owners. Here we are again.
There's a second thing lurking in the Numbers story. The site's logs didn't just show scraping — they showed someone looking for private access to the data before publication. Why would anyone hack a box office database? Because Polymarket runs weekly betting markets on opening weekends, and names The Numbers as the source of truth that resolves them. See the figures before everyone else and you can front-run every market, every week. If you’re new to the horror that’s prediction markets, Luce wrote a great explainer here.
What’s really ugly about this is that the moment a number resolves bets, the database that holds the number becomes financial infrastructure — without financial infrastructure's budget, security team, or regulator. And prediction markets will now happily turn almost any number into a bet: chart positions, ratings, awards, follower counts. I've written before about our fixation with quantifying fandom, and the way metrics distort the thing they measure. It turns out that's the gentle version of the problem. The harsh version is that every scoreboard is now an attack surface. A hobbyist database from 1997 was never designed to withstand this.
There are attempts at a fix. Earlier this month Cloudflare announced that from 15 September it will block mixed-use AI crawlers by default on ad-supported pages, and it's evolving its pay-per-crawl scheme into "pay per use". What that means is that publishers get paid when their content actually shows up in an AI answer, rather than merely when a bot fetches it. Wikipedia, for its part, is formalising paid deals with Amazon, Meta and Microsoft — the commons negotiating terms with its own extractors, which is either pragmatism or tragedy depending on the day you ask me. Whether any of it works depends entirely on whether AI companies pay up rather than route around it. As Nash put it to Follows: somebody has to try.
Everything breaks at scale, and the things breaking first are the things built with the most love and the least money: the independent archives, the hobby databases, the fan wikis, the forums holding far more of our shared knowledge than anyone realises. The threat right now is that these spaces we’ve poured our hearts into are being read to death by machines.
The Numbers survived because the website was never the whole business. Most of the load-bearing shelves holding our best information don't have that luxury. So: find the fan-built database you rely on and find a way to support it.
more good stuff
- first off, we had a truly incredible first week with Lume. Amazing coverage, including this deeply-reported feature in Rolling Stone. And perhaps most amazingly, we managed to shift the Aotearoa Album Charts for the week. Eight Lumes either re-entered, entered or moved up in the charts, and in Isla Noon's case it was ONLY through Lume sales. This feels like the very start of something important.

- this (via Matt at webcurios) is kind of intriguing. Another attempt to bring your various feeds together in one place, curated the way you choose.

- just an extremely specific and cool piece of information: in Aotearoa we often have Phoenix palms beside WWI memorials. Here’s why.

finally, in my lego city
Forward this email to someone who uses a website older than google.
You just read issue #84 of what you love matters. You can also browse the full archives of this newsletter.
Add a comment: