Crashers, explained

A ten-minute film

The back end, in detail

The first film explains Crashers to someone who does not write software. This one is for someone who does — or who wants to. It goes through every program in the pipeline: how each website is actually read, what the cheap filters throw away and why, exactly what is asked of Claude and what is never taken on trust, and how one file gets from a laptop to a phone.

9 min 59 s · narrated, no music · download the MP4 · full transcript below

The shape of it, in one paragraph

Thirty-four programs and twenty-one modules — about 8,000 lines of Python — run once a morning on one Windows laptop. They read sixteen public event sites, throw away everything that plainly fails the rule, hand what is left to Claude to read, and write a single 318 KB JSON file. Git pushes that file; Cloudflare hands it out. There is no server, no database and no bill. The back end never talks to the front end — it writes a file, and the file is the interface.

The eight stages

The rail along the foot of every scene in the film is these eight boxes, so you can always see which one is being explained.

Crawl, combine, what is new, judge, merge, publish, decorate, ship Crawl 16 sources 3,649 rows Combine price, date, place, dedupe, keywords New what has never been classified Judge Claude reads 370 candidates Merge append only, never overwrite Publish events.json 525 rows Decorate times, blurbs, prune, pics, names Ship git push → Cloudflare edge ONE LAPTOP, 6:00 AM THE INTERNET
One script, scripts/refresh.py, runs all eight in order and records the result of every step in a log.

Stage one — sixteen sites, six ways in

There is no feed and nobody hands the project data. Each site has a small adapter whose only job is to turn whatever that site does into one shape: an Event record with a title, a start, a place, a price flag and a description. Everything downstream sees only that shape, which is why adding a source is one file and not a rewrite.

What differs is the route in, and it differs a lot:

SourceHow the data is actually obtainedThe catch
City of Sydney Every page ships its whole record in a __NEXT_DATA__ script tag. No HTML parsing at all; their sitemap lists every event page. None worth the name. This is the easy one.
Eventbrite Search pages carry window.__SERVER_DATA__: ~1,850 free Sydney events, 20 a page. Detail pages are fetched only for survivors. The listing summary almost never mentions food, so the second request is unavoidable. Seven location slugs are needed — “sydney” does not reach Penrith. australia--windsor resolves to Queensland and is deliberately not used.
Meetup The search page embeds ten schema.org Event objects with the full description — the richest text of any source. One request per day, thirty days. Paging does not exist: page, offset, skip all return page one. Timestamps are UTC, so an 8pm event lands on the wrong day unless converted.
Humanitix A POST to their own search API, sliced by bounding box (one per catchment) × date preset. Page numbers are ignored and every slice caps at 32 rows; customDateRange is accepted and silently returns nothing.
Seven councils
Ryde, Woollahra, Hornsby, Ku-ring-gai, Parramatta, Fairfield, Campbelltown
All run OpenCities, which renders identical markup, so one parser and a config table covers all seven. Paging is an ASP.NET postback — every link is href="#", and ?page=2 returns page one. The way through is the page-number <select> posted as a form, with __VIEWSTATE re-read each time because it is single-use.
The rest
North Sydney, Canada Bay, Inner West, Mosman, Penrith
Hand-written CSS selectors per council, in one file driven by a config table rather than one adapter each. Drupal is the CMS, not an events module, so no two themes agree. Mosman writes times with a dot — “10.00am” — which a naive parser reads as 12am.

Four councils that are deliberately not in there

Canterbury-Bankstown emits utility classes with no stable hook to anchor on. Bayside's date element is empty in the HTML and filled in later by JavaScript. Cumberland has grid classes and nothing event-specific. Georges River is one flat page with no container per event. All four need a real browser, which is a different job — so the well-behaved sources ship and the awkward ones wait, with the reason written down rather than forgotten.

Permission, and pace

Before a source is added, a script reads its robots.txt and the answer is recorded in the project's README. If a site names a Claude agent, the project leaves it alone — TryBooking and Ticketebo both do, and both stay out, even though the pages are public and could be taken on a technicality. Humanitix used to name one and no longer does, which is why it is in.

Every request then goes through one function that spaces requests per host: 0.5 s for the big ticketing platforms, 1.5 s for Mosman, 2.5 s for Inner West Council — who quite fairly returned a 429 when crawled faster. A 429 does not merely pause the crawler; it doubles that host's delay for the rest of the run. Every request also carries a user agent naming the project and an email address.

Roughly 3,500 fetched pages are cached on disk as JSON. A second run reads from disk and asks nobody for anything, which is why a full crawl takes about an hour and a rebuild takes seconds.


Stage two — the free filters

Everything crawled lands in one funnel, and every filter in it costs nothing. These are the real numbers from the run of 21 September 2026:

Everything crawled
3,649
Not stated paid
2,748
In the next 28 days
2,225
In a catchment
2,144
After dedupe
2,081
Worth an LLM call
370

The price flag has three values, not two

is_free can be True (the source states it is free), False (the source states a price) or None (the listing never says). Only False is a refusal. Dropping None alongside it looks harmless and was in fact a bug in two separate files: Meetup states no price anywhere in its markup, so every Meetup event was being fetched and then silently discarded, with nothing in the output to say why. Whether an unpriced listing is free is exactly the classifier's job.

Merging duplicates

The same gallery opening can appear on City of Sydney, Eventbrite and Humanitix at once. Two events merge only if they start at the same minute and their titles agree by one of three measures:

The third exists because a sequence matcher reads a shared phrase as different text once the words move. “Building Tech Safety for Women at Pyrmont Community Centre” against “Building tech safety for women” scores 0.67 on sequence and 1.00 on words — one event, two listings. The matcher is conservative on purpose: a false merge hides a real event, which is worse than one duplicate slipping through.

The keyword gate

About fifty regular expressions — refreshment, canap[eé]s?, morning tea, sausage sizzle, grazing (table|board), complimentary — sort every listing into three tiers. It is tuned for recall, not precision: junk getting through is fine, because Claude sorts it out, but a real free feed dropped here can never be recovered by anything downstream.

Evidence and prior are kept strictly apart in the code, because a signal is evidence and a shape is a guess — telling a reader “openings usually have wine” is a different promise from “this event has wine”.


Stage three — paying only for what is new

The classifier is the only step that could ever cost money, so a much cheaper question runs first: which of these have never been classified? The feed is compared by URL against every verdict ever reached. On this run: 370 candidates, 601 already on record, zero new — and the expensive step was skipped entirely.


Stage four — what is actually asked of Claude

One prompt, two judgements, and the prompt is explicit that they must never contaminate each other.

1 · Stated

Does the listing explicitly say attendees receive food or drink at no extra cost? The prompt names what does not count — cash bar, food trucks on site, available for purchase, BYO, licensed venue, a cafe on site — and it is equally explicit about reading the fine print. Venues advertise FREE DRINK in the title and qualify it in an asterisk at the bottom: “free drink voucher redeemable with your first drink purchase at the bar”. That is a discount, the attendee pays before receiving anything, and it is the single case the classifier is tested against hardest.

2 · Likelihood

An honest estimate that an attendee actually gets fed, even when the text says nothing: gallery openings 70–85, corporate and university networking 65–80, community festivals and markets 10–25, guided tours and workshops 5–15. The prompt asks for calibration rather than generosity and says why: a reader sees that percentage next to a warning that the description mentions nothing, so an inflated number wastes somebody's evening.

No quote, no claim

If Claude reports that a listing states free food, it must quote the exact sentence — and that quote is not taken on trust. A function checks the words really do appear in the listing they are attributed to. If they don't, the claim is not downgraded, it is deleted: has_food goes back to false, the quote is wiped, and the event drops to a labelled estimate.

Across 601 verdicts the check has never once caught a fabrication. Both apparent catches were the check itself being too strict, and the second version of it cost six genuine verdicts before that was noticed. A check that is too tight here silently deletes real events, which is the worse failure and much harder to see — so the comparison is made on words, ignoring punctuation, emphasis markers and emoji, which are the formatting of a quote rather than its content.

Two backends, one judgement

The prompt, the verdict shape and the quote check live in one file, and both routes import them rather than restating them — so the backends cannot drift apart. One route is the Anthropic API, which costs money per event. The other shells out to Claude Code itself on a subscription that is already paid for, which costs nothing, and that is the one that runs.

The subprocess is given nothing: --restricted, no tools, slash commands disabled, no setting sources, and a scratch working directory. It cannot read the repository, run anything, or pick up its own project instructions — it is a text-in, JSON-out call that happens to be a process. Listings go twelve at a time, numbered, and the reply is matched back by number rather than by position, so a model that drops or reorders one corrupts nothing: the missing row comes back empty and is skipped.

For the paid route there is still a ceiling. The run estimates its cost first, and over a dollar it publishes what it already understands and stops. That is not really about a dollar — it is so that a source changing shape and dumping two thousand rows into the feed costs an explanation rather than money.


Stage five — the answer key

Verdicts are never written over the top of the file that holds them. A URL not yet present is appended; a URL already there is left exactly as it was, whoever classified it.

The reason is that 344 of those 601 rows were reached by reading each listing by hand, in session, and they are the answer key the automated classifier is scored against — currently at about 90% agreement. A re-crawl cannot rebuild them. Overwrite them and you do not just lose work: you remove the only thing the model is measured by, and every score afterwards is the classifier grading its own homework.


Stage six — building the payload

publish.py turns verdicts into the file the app downloads. It runs the duplicate matcher a second time over the classified rows — free, and it means a fix to the matcher reaches the app without re-reading a single listing — then writes out only the fields the app reads, under short names: source_url → url, likelihood_reason → why, access_tier → access. The payload ships to every phone that opens the site, so a field name is bytes.

It also enforces rule one. Where the source flagged an event free but Claude read the text and found “$15.00 fee on arrival”, “$5 entrance” or “must be a member”, the event is dropped. Seventy went that way on this run: 595 classified rows became 525 published.

Stage seven — five decorators, and why the order matters

  1. add_times — 279 clocks recovered

    City of Sydney and Inner West print their hours in prose; Eventbrite needs a detail fetch. Events whose source publishes no time get none, and the app says so rather than guessing.

  2. add_blurbs — 244 written, 265 extracted

    One line saying what the event actually is. Where no line has been written, a line is picked out of the host's own description — nothing is paraphrased, so it cannot claim anything the listing does not.

  3. prune_past — 525 → 335

    Drops events the calendar has passed, by date only. Whether a 6pm event is over depends on what time it is when the page is open, so that call belongs in the browser.

  4. add_images — 316 pictures

    Reads each page's own og:image and hotlinks it, so the host serves the bytes and nothing is republished. This is why the order matters: it is the only decorator that costs a request per event, so it runs after pruning and reads the 335 listings the app will show rather than the 525 publish built. Its cache keeps the negatives too — a URL mapped to null means “read it, there is no picture” — so a quiet morning costs nothing at all.

  5. name_events — 329 short names

    A host's title is written to sell a listing, not to be scanned in a rail. “Soda Tuesday | Musical Bingo | Free Drink on Arrival, $1 HotDogs” wants to read as “Musical Bingo”, and no rule gets there — so it is read, by the same free Claude route, under one hard constraint: every word in the short name must already appear in the host's own title. It is deleting, not writing, and that is checked mechanically rather than trusted. Anything that invents a word falls back to the full title.

Stage eight — shipping it

If events.json changed, the laptop commits and pushes. Cloudflare sees the push and copies the site onto its own machines around the world. There is no Worker script and no build step — the config file is nine lines long and says serve this directory.

All of it is triggered by Windows Task Scheduler at 6am. Not GitHub Actions, and that was a decision: the repository is private, so those minutes are metered, and a polite crawl is an hour of wall clock — a nightly hour would eat the free allowance and then start costing money. The setting that matters is StartWhenAvailable: this laptop is usually off overnight, so a plain 6am trigger would simply never fire. Instead the run happens the moment the machine is next switched on.


Transcript

The full narration, scene by scene, with the time each one starts.

0:00

The back end

This is the back end of Crashers, in detail. Thirty four small programs, twenty one modules, about eight thousand lines of Python, and one command that runs it all.

0:10

There is no server

The thing that surprises people first. The back end is not a server. Nothing is listening on a port, there is nothing to log into, and there is no bill. It is a folder of programs on one laptop that wake up, read the public web, think about what they found, and write a single file.

0:26

The conductor

The conductor is a script called refresh. It runs every other step as a separate process and watches the exit code. That matters: if a council website is down, that step fails, the failure goes in the log, and the run carries on. One source being down is never a reason to publish nothing.

0:43

One adapter per site

Stage one. Reading the websites. There is a small adapter per site, because no two sites give up their data the same way. Some publish a structured record inside the page. Some need their search endpoint poked in a particular order. Some hide behind a form that only answers a post. The adapter's only job is to turn whatever that site does into one shape.

1:04

The easy one

The City of Sydney is the easy one. Their site is built in Next dot J S, and every page ships a script tag holding the entire record the page was rendered from. So the adapter never parses H T M L at all. It reads a proper record, with a free event boolean and real coordinates. Their sitemap lists every event page too.

1:23

Eventbrite, in two stages

Eventbrite is the biggest source. Two thirds of everything the app shows. Their search pages give about eighteen hundred free Sydney events cheaply, but the summary almost never mentions food. The food is in the full description, and that is a second request per event. So the crawl is two stage. Read the cheap list, drop the online events and anything outside Sydney, then fetch detail pages for the survivors only. A few hundred requests instead of eighteen hundred.

1:49

Meetup will not be paged

Meetup will not be paged. Page, offset, skip, all of them return page one, because the real list loads over GraphQL as you scroll. But the search page embeds ten schema dot org event objects with the full description. So the adapter asks for one day at a time. Thirty days, thirty requests. Meetup is also the only source that timestamps in U T C. Left unconverted, an eight p m event lands on the wrong day, and the next step deletes it as finished.

2:17

Humanitix, sliced two ways

Humanitix has the same problem and two more. Their page number is ignored, every slice returns the same thirty two rows, and their custom date range silently returns nothing. So the query is sliced two ways at once. One bounding box per catchment, times their named date presets. A slice under thirty two rows is complete rather than truncated.

2:37

Seven councils, one parser

Then the councils. Most New South Wales councils don't run their own event system. They run a platform called OpenCities, which renders identical markup on every one of them, so seven councils share one adapter. The interesting part is paging. Every page link there is an A S P dot net postback, and asking for page two returns page one, which is a convincing way to collect the same ten events fifteen times. The way through is the page number dropdown, submitted as a form post, with the view state re-read every time.

3:06

Being a good guest

This is all done politely, and deliberately so. Before a source is added, a script reads its robots file and the answer goes in the README. If a site names a Claude agent, we leave it alone. TryBooking and Ticketebo both do, and both stay out. Every request then goes through one function that spaces requests per host. Half a second for the big platforms, two and a half for Inner West Council, who quite fairly returned a four two nine when we crawled them faster. And a four two nine doesn't just pause us, it doubles that host's delay for the rest of the run.

3:37

The funnel

Stage two combines the lot. Sixteen files, three thousand six hundred and forty nine rows, into one funnel. Anything the source states costs money, gone. Two thousand seven hundred and forty eight. Outside the next twenty eight days, gone. Two thousand two hundred and twenty five. Outside the nine catchments, gone. Two thousand one hundred and forty four. Sixty three duplicates merge, leaving two thousand and eighty one. Every filter there is free, and none of them reads English.

4:06

Three values, not two

One subtlety there. The price flag has three values, not two. Free, paid, and nothing stated. Only the middle one is a refusal, and dropping the third with it would silently throw away every Meetup event.

4:20

Merging duplicates

The duplicate matcher deserves a minute. The same gallery opening can appear on City of Sydney, Eventbrite and Humanitix at once. Two events merge only if they start at the same minute and their titles agree, by one of three measures. Near identical, on its own. Similar character by character, plus venues within three hundred metres. Or the same words in a different order, plus that venue check. The third exists because a title that is another title plus a venue name scores badly on sequence, perfectly on words.

4:49

The cheap word search

Then the keyword pass, and it is deliberately dumb. About fifty regular expressions. Refreshments, canapés, morning tea, sausage sizzle, complimentary. Tuned for recall, not precision. Junk getting through is fine, because Claude sorts it out, but a real free feed dropped here can never be recovered. Every listing lands in one of three tiers. Evidence, prior, or skip. Three hundred and fourteen, fifty six, seventeen hundred and eleven.

5:16

Only pay for what is new

Stage three is the only step that could ever cost money, so before it runs, one script asks a much cheaper question. Which of these have we never seen? On the last run, three hundred and seventy candidates, six hundred and one already known, and zero new. The expensive step was skipped entirely.

5:34

Two judgements, kept apart

When there is something new, each listing goes to Claude with a prompt that asks for two judgements and insists they never contaminate each other. The first is stated. Does the listing explicitly say attendees get food or drink at no extra cost? The prompt is blunt about what doesn't count. Cash bar, food trucks, available for purchase, B Y O. And blunt about the fine print, because venues advertise free drink in the title, then qualify it at the bottom. Redeemable with your first purchase.

6:02

The honest guess

The second judgement is likelihood. An honest estimate that you actually get fed, even when the text says nothing. Gallery openings, seventy to eighty five. Corporate networking, sixty five to eighty. Guided tours and workshops, five to fifteen. Be calibrated, not generous, and the prompt says why. A reader sees that percentage next to a warning that the description mentions nothing.

6:25

No quote, no claim

Then the rule the whole app rests on. If Claude says the listing states free food, it has to quote the exact sentence, and that quote is not taken on trust. A function checks the words really do appear in the listing they are attributed to. If they don't, the claim isn't downgraded, it is deleted, and the event drops to an estimate. Across six hundred verdicts it has never once caught a fabrication. Both apparent catches were the check being too strict, and one cost six genuine verdicts.

6:51

Two backends, one judgement

There are two ways to run that classifier, and only one judgement. The prompt, the verdict shape and the quote check live in one file, and both routes import them rather than restating them. One route is the Anthropic A P I, which costs money per event. The other shells out to Claude Code itself, on a subscription already paid for, which costs nothing. That is the one that runs. The subprocess gets no tools, no file access and no settings. It is a text in, JSON out call that happens to be a process.

7:20

The answer key

Then the most conservative rule in the codebase. Verdicts are never written over the top of the file that holds them. New addresses are appended, and one already there is left exactly as it was. Three hundred and forty four of those rows were reached by reading each listing by hand. They are the answer key the classifier is scored against, and it agrees with them ninety per cent of the time.

7:41

Building the payload

Stage four builds the file the app downloads. It takes every verdict, runs the duplicate matcher a second time, which is free and lands a fix without re-reading a listing, then writes only the fields the app reads, under short names. Source url becomes url. Likelihood reason becomes why. That payload ships to every phone that opens the site, so a field name is bytes. It also enforces rule one. Where Claude read the text and said this is not free, the event is dropped. Seventy went that way.

8:10

Five decorators, in order

Then five small programs decorate that file in place, and the order is load bearing. Times are recovered, two hundred and seventy nine of them pulled out of prose the source never structured. Blurbs are stamped on. Past events are pruned, and five hundred and twenty five becomes three hundred and thirty five. Only then are pictures fetched, because that is the one decorator costing a request per event, and pruning first means reading three hundred listings rather than five hundred.

8:35

Deleting, not writing

The last decorator is the nicest. A host's title is written to sell a listing, not to be scanned in a list. Soda Tuesday, bar, Musical Bingo, bar, free drink on arrival, one dollar hot dogs. You want Musical Bingo. No rule gets there, so it is read, by that same free Claude route, under one hard constraint. Every word in the short name must already appear in the host's own title. It is deleting, not writing, and that is checked.

9:02

Shipping it

Stage five. If the events file changed, the laptop commits it and pushes. Cloudflare sees the push and copies the site onto its machines around the world. No worker script, no build step. The config file is nine lines long and says, serve this directory. The back end never talks to the front end. It writes a file, and the file is the interface.

9:23

Six in the morning

All of it is triggered by Windows Task Scheduler at six each morning. Not GitHub Actions, because the repository is private and those minutes are metered. The setting that matters is start when available. This laptop is usually off overnight, so a plain six a m trigger would simply never fire.

9:40

The shape of it

Eight thousand lines, sixteen sources, nine filters, and a language model used in exactly one place, where judgement is genuinely needed. Everywhere else, plain boring code you can read. No database, no server, no accounts, no bill. Websites in, one file out, and a phone that tells you where dinner is.

Where the numbers come from

Every figure here was read off the project itself on 21 September 2026 — the refresh log at data/logs/refresh-20260921-113507.log, the answer key at data/classified.jsonl, and the payload app/local/events.json as it stood that morning. Nothing is illustrative. The numbers move every day; these are one real day's.