A ten-minute film
The back end, in detail
The first film explains Crashers to someone who does not write software. This one is for someone who does — or who wants to. It goes through every program in the pipeline: how each website is actually read, what the cheap filters throw away and why, exactly what is asked of Claude and what is never taken on trust, and how one file gets from a laptop to a phone.
9 min 59 s · narrated, no music · download the MP4 · full transcript below
The shape of it, in one paragraph
Thirty-four programs and twenty-one modules — about 8,000 lines of Python — run once a morning on one Windows laptop. They read sixteen public event sites, throw away everything that plainly fails the rule, hand what is left to Claude to read, and write a single 318 KB JSON file. Git pushes that file; Cloudflare hands it out. There is no server, no database and no bill. The back end never talks to the front end — it writes a file, and the file is the interface.
The eight stages
The rail along the foot of every scene in the film is these eight boxes, so you can always see which one is being explained.
scripts/refresh.py, runs all eight in order and
records the result of every step in a log.
Stage one — sixteen sites, six ways in
There is no feed and nobody hands the project data. Each site has a small
adapter whose only job is to turn whatever that site does into one shape: an
Event record with a title, a start, a place, a price flag and a
description. Everything downstream sees only that shape, which is why adding a
source is one file and not a rewrite.
What differs is the route in, and it differs a lot:
| Source | How the data is actually obtained | The catch |
|---|---|---|
| City of Sydney | Every page ships its whole record in a __NEXT_DATA__ script tag. No HTML parsing at all; their sitemap lists every event page. |
None worth the name. This is the easy one. |
| Eventbrite | Search pages carry window.__SERVER_DATA__: ~1,850 free Sydney events, 20 a page. Detail pages are fetched only for survivors. |
The listing summary almost never mentions food, so the second request is unavoidable. Seven location slugs are needed — “sydney” does not reach Penrith. australia--windsor resolves to Queensland and is deliberately not used. |
| Meetup | The search page embeds ten schema.org Event objects with the full description — the richest text of any source. One request per day, thirty days. |
Paging does not exist: page, offset, skip all return page one. Timestamps are UTC, so an 8pm event lands on the wrong day unless converted. |
| Humanitix | A POST to their own search API, sliced by bounding box (one per catchment) × date preset. | Page numbers are ignored and every slice caps at 32 rows; customDateRange is accepted and silently returns nothing. |
| Seven councils Ryde, Woollahra, Hornsby, Ku-ring-gai, Parramatta, Fairfield, Campbelltown |
All run OpenCities, which renders identical markup, so one parser and a config table covers all seven. | Paging is an ASP.NET postback — every link is href="#", and ?page=2 returns page one. The way through is the page-number <select> posted as a form, with __VIEWSTATE re-read each time because it is single-use. |
| The rest North Sydney, Canada Bay, Inner West, Mosman, Penrith |
Hand-written CSS selectors per council, in one file driven by a config table rather than one adapter each. | Drupal is the CMS, not an events module, so no two themes agree. Mosman writes times with a dot — “10.00am” — which a naive parser reads as 12am. |
Four councils that are deliberately not in there
Canterbury-Bankstown emits utility classes with no stable hook to anchor on. Bayside's date element is empty in the HTML and filled in later by JavaScript. Cumberland has grid classes and nothing event-specific. Georges River is one flat page with no container per event. All four need a real browser, which is a different job — so the well-behaved sources ship and the awkward ones wait, with the reason written down rather than forgotten.
Permission, and pace
Before a source is added, a script reads its robots.txt and the
answer is recorded in the project's README. If a site names a Claude
agent, the project leaves it alone — TryBooking and Ticketebo both do,
and both stay out, even though the pages are public and could be taken on a
technicality. Humanitix used to name one and no longer does, which is why it
is in.
Every request then goes through one function that spaces requests per host:
0.5 s for the big ticketing platforms, 1.5 s for Mosman,
2.5 s for Inner West Council — who quite fairly returned a
429 when crawled faster. A 429 does not merely pause
the crawler; it doubles that host's delay for the rest of the run.
Every request also carries a user agent naming the project and an email address.
Roughly 3,500 fetched pages are cached on disk as JSON. A second run reads from disk and asks nobody for anything, which is why a full crawl takes about an hour and a rebuild takes seconds.
Stage two — the free filters
Everything crawled lands in one funnel, and every filter in it costs nothing. These are the real numbers from the run of 21 September 2026:
- Everything crawled
- 3,649
- Not stated paid
- 2,748
- In the next 28 days
- 2,225
- In a catchment
- 2,144
- After dedupe
- 2,081
- Worth an LLM call
- 370
The price flag has three values, not two
is_free can be True (the source states it is free),
False (the source states a price) or None (the listing
never says). Only False is a refusal. Dropping
None alongside it looks harmless and was in fact a bug in two
separate files: Meetup states no price anywhere in its markup, so every Meetup
event was being fetched and then silently discarded, with nothing in the output
to say why. Whether an unpriced listing is free is exactly the classifier's job.
Merging duplicates
The same gallery opening can appear on City of Sydney, Eventbrite and Humanitix at once. Two events merge only if they start at the same minute and their titles agree by one of three measures:
- Near-identical (≥ 0.90) on its own.
- Similar character-by-character (≥ 0.72) plus venues within 300 m.
- The same words in a different order (≥ 0.85 token overlap) plus the same venue check.
The third exists because a sequence matcher reads a shared phrase as different text once the words move. “Building Tech Safety for Women at Pyrmont Community Centre” against “Building tech safety for women” scores 0.67 on sequence and 1.00 on words — one event, two listings. The matcher is conservative on purpose: a false merge hides a real event, which is worse than one duplicate slipping through.
The keyword gate
About fifty regular expressions — refreshment,
canap[eé]s?, morning tea, sausage sizzle,
grazing (table|board), complimentary — sort every
listing into three tiers. It is tuned for recall, not precision:
junk getting through is fine, because Claude sorts it out, but a real free feed
dropped here can never be recovered by anything downstream.
- evidence — 314. The text claims food or drink.
- prior — 56. The format usually feeds you (a gallery opening, a book launch) but the text is silent.
- skip — 1,711. No reason to spend a token.
Evidence and prior are kept strictly apart in the code, because a signal is evidence and a shape is a guess — telling a reader “openings usually have wine” is a different promise from “this event has wine”.
Stage three — paying only for what is new
The classifier is the only step that could ever cost money, so a much cheaper question runs first: which of these have never been classified? The feed is compared by URL against every verdict ever reached. On this run: 370 candidates, 601 already on record, zero new — and the expensive step was skipped entirely.
Stage four — what is actually asked of Claude
One prompt, two judgements, and the prompt is explicit that they must never contaminate each other.
1 · Stated
Does the listing explicitly say attendees receive food or drink at no extra cost? The prompt names what does not count — cash bar, food trucks on site, available for purchase, BYO, licensed venue, a cafe on site — and it is equally explicit about reading the fine print. Venues advertise FREE DRINK in the title and qualify it in an asterisk at the bottom: “free drink voucher redeemable with your first drink purchase at the bar”. That is a discount, the attendee pays before receiving anything, and it is the single case the classifier is tested against hardest.
2 · Likelihood
An honest estimate that an attendee actually gets fed, even when the text says nothing: gallery openings 70–85, corporate and university networking 65–80, community festivals and markets 10–25, guided tours and workshops 5–15. The prompt asks for calibration rather than generosity and says why: a reader sees that percentage next to a warning that the description mentions nothing, so an inflated number wastes somebody's evening.
No quote, no claim
If Claude reports that a listing states free food, it must quote the exact
sentence — and that quote is not taken on trust. A function checks the words
really do appear in the listing they are attributed to. If they don't, the
claim is not downgraded, it is deleted:
has_food goes back to false, the quote is wiped, and the event
drops to a labelled estimate.
Across 601 verdicts the check has never once caught a fabrication. Both apparent catches were the check itself being too strict, and the second version of it cost six genuine verdicts before that was noticed. A check that is too tight here silently deletes real events, which is the worse failure and much harder to see — so the comparison is made on words, ignoring punctuation, emphasis markers and emoji, which are the formatting of a quote rather than its content.
Two backends, one judgement
The prompt, the verdict shape and the quote check live in one file, and both routes import them rather than restating them — so the backends cannot drift apart. One route is the Anthropic API, which costs money per event. The other shells out to Claude Code itself on a subscription that is already paid for, which costs nothing, and that is the one that runs.
The subprocess is given nothing: --restricted, no tools, slash
commands disabled, no setting sources, and a scratch working directory. It
cannot read the repository, run anything, or pick up its own project
instructions — it is a text-in, JSON-out call that happens to be a process.
Listings go twelve at a time, numbered, and the reply is matched back
by number rather than by position, so a model that drops or reorders
one corrupts nothing: the missing row comes back empty and is skipped.
For the paid route there is still a ceiling. The run estimates its cost first, and over a dollar it publishes what it already understands and stops. That is not really about a dollar — it is so that a source changing shape and dumping two thousand rows into the feed costs an explanation rather than money.
Stage five — the answer key
Verdicts are never written over the top of the file that holds them. A URL not yet present is appended; a URL already there is left exactly as it was, whoever classified it.
The reason is that 344 of those 601 rows were reached by reading each listing by hand, in session, and they are the answer key the automated classifier is scored against — currently at about 90% agreement. A re-crawl cannot rebuild them. Overwrite them and you do not just lose work: you remove the only thing the model is measured by, and every score afterwards is the classifier grading its own homework.
Stage six — building the payload
publish.py turns verdicts into the file the app downloads. It runs
the duplicate matcher a second time over the classified rows — free, and it
means a fix to the matcher reaches the app without re-reading a single listing —
then writes out only the fields the app reads, under short names:
source_url → url,
likelihood_reason → why,
access_tier → access. The payload ships to
every phone that opens the site, so a field name is bytes.
It also enforces rule one. Where the source flagged an event free but Claude read the text and found “$15.00 fee on arrival”, “$5 entrance” or “must be a member”, the event is dropped. Seventy went that way on this run: 595 classified rows became 525 published.
Stage seven — five decorators, and why the order matters
-
add_times — 279 clocks recovered
City of Sydney and Inner West print their hours in prose; Eventbrite needs a detail fetch. Events whose source publishes no time get none, and the app says so rather than guessing.
-
add_blurbs — 244 written, 265 extracted
One line saying what the event actually is. Where no line has been written, a line is picked out of the host's own description — nothing is paraphrased, so it cannot claim anything the listing does not.
-
prune_past — 525 → 335
Drops events the calendar has passed, by date only. Whether a 6pm event is over depends on what time it is when the page is open, so that call belongs in the browser.
-
add_images — 316 pictures
Reads each page's own
og:imageand hotlinks it, so the host serves the bytes and nothing is republished. This is why the order matters: it is the only decorator that costs a request per event, so it runs after pruning and reads the 335 listings the app will show rather than the 525 publish built. Its cache keeps the negatives too — a URL mapped to null means “read it, there is no picture” — so a quiet morning costs nothing at all. -
name_events — 329 short names
A host's title is written to sell a listing, not to be scanned in a rail. “Soda Tuesday | Musical Bingo | Free Drink on Arrival, $1 HotDogs” wants to read as “Musical Bingo”, and no rule gets there — so it is read, by the same free Claude route, under one hard constraint: every word in the short name must already appear in the host's own title. It is deleting, not writing, and that is checked mechanically rather than trusted. Anything that invents a word falls back to the full title.
Stage eight — shipping it
If events.json changed, the laptop commits and pushes. Cloudflare
sees the push and copies the site onto its own machines around the world. There
is no Worker script and no build step — the config file is nine lines long and
says serve this directory.
All of it is triggered by Windows Task Scheduler at 6am. Not GitHub Actions, and
that was a decision: the repository is private, so those minutes are metered, and
a polite crawl is an hour of wall clock — a nightly hour would eat the free
allowance and then start costing money. The setting that matters is
StartWhenAvailable: this laptop is usually off overnight, so a plain
6am trigger would simply never fire. Instead the run happens the moment the
machine is next switched on.
Transcript
The full narration, scene by scene, with the time each one starts.
The back end
This is the back end of Crashers, in detail. Thirty four small programs, twenty one modules, about eight thousand lines of Python, and one command that runs it all.
There is no server
The thing that surprises people first. The back end is not a server. Nothing is listening on a port, there is nothing to log into, and there is no bill. It is a folder of programs on one laptop that wake up, read the public web, think about what they found, and write a single file.
The conductor
The conductor is a script called refresh. It runs every other step as a separate process and watches the exit code. That matters: if a council website is down, that step fails, the failure goes in the log, and the run carries on. One source being down is never a reason to publish nothing.
One adapter per site
Stage one. Reading the websites. There is a small adapter per site, because no two sites give up their data the same way. Some publish a structured record inside the page. Some need their search endpoint poked in a particular order. Some hide behind a form that only answers a post. The adapter's only job is to turn whatever that site does into one shape.
The easy one
The City of Sydney is the easy one. Their site is built in Next dot J S, and every page ships a script tag holding the entire record the page was rendered from. So the adapter never parses H T M L at all. It reads a proper record, with a free event boolean and real coordinates. Their sitemap lists every event page too.
Eventbrite, in two stages
Eventbrite is the biggest source. Two thirds of everything the app shows. Their search pages give about eighteen hundred free Sydney events cheaply, but the summary almost never mentions food. The food is in the full description, and that is a second request per event. So the crawl is two stage. Read the cheap list, drop the online events and anything outside Sydney, then fetch detail pages for the survivors only. A few hundred requests instead of eighteen hundred.
Meetup will not be paged
Meetup will not be paged. Page, offset, skip, all of them return page one, because the real list loads over GraphQL as you scroll. But the search page embeds ten schema dot org event objects with the full description. So the adapter asks for one day at a time. Thirty days, thirty requests. Meetup is also the only source that timestamps in U T C. Left unconverted, an eight p m event lands on the wrong day, and the next step deletes it as finished.
Humanitix, sliced two ways
Humanitix has the same problem and two more. Their page number is ignored, every slice returns the same thirty two rows, and their custom date range silently returns nothing. So the query is sliced two ways at once. One bounding box per catchment, times their named date presets. A slice under thirty two rows is complete rather than truncated.
Seven councils, one parser
Then the councils. Most New South Wales councils don't run their own event system. They run a platform called OpenCities, which renders identical markup on every one of them, so seven councils share one adapter. The interesting part is paging. Every page link there is an A S P dot net postback, and asking for page two returns page one, which is a convincing way to collect the same ten events fifteen times. The way through is the page number dropdown, submitted as a form post, with the view state re-read every time.
Being a good guest
This is all done politely, and deliberately so. Before a source is added, a script reads its robots file and the answer goes in the README. If a site names a Claude agent, we leave it alone. TryBooking and Ticketebo both do, and both stay out. Every request then goes through one function that spaces requests per host. Half a second for the big platforms, two and a half for Inner West Council, who quite fairly returned a four two nine when we crawled them faster. And a four two nine doesn't just pause us, it doubles that host's delay for the rest of the run.
The funnel
Stage two combines the lot. Sixteen files, three thousand six hundred and forty nine rows, into one funnel. Anything the source states costs money, gone. Two thousand seven hundred and forty eight. Outside the next twenty eight days, gone. Two thousand two hundred and twenty five. Outside the nine catchments, gone. Two thousand one hundred and forty four. Sixty three duplicates merge, leaving two thousand and eighty one. Every filter there is free, and none of them reads English.
Three values, not two
One subtlety there. The price flag has three values, not two. Free, paid, and nothing stated. Only the middle one is a refusal, and dropping the third with it would silently throw away every Meetup event.
Merging duplicates
The duplicate matcher deserves a minute. The same gallery opening can appear on City of Sydney, Eventbrite and Humanitix at once. Two events merge only if they start at the same minute and their titles agree, by one of three measures. Near identical, on its own. Similar character by character, plus venues within three hundred metres. Or the same words in a different order, plus that venue check. The third exists because a title that is another title plus a venue name scores badly on sequence, perfectly on words.
The cheap word search
Then the keyword pass, and it is deliberately dumb. About fifty regular expressions. Refreshments, canapés, morning tea, sausage sizzle, complimentary. Tuned for recall, not precision. Junk getting through is fine, because Claude sorts it out, but a real free feed dropped here can never be recovered. Every listing lands in one of three tiers. Evidence, prior, or skip. Three hundred and fourteen, fifty six, seventeen hundred and eleven.
Only pay for what is new
Stage three is the only step that could ever cost money, so before it runs, one script asks a much cheaper question. Which of these have we never seen? On the last run, three hundred and seventy candidates, six hundred and one already known, and zero new. The expensive step was skipped entirely.
Two judgements, kept apart
When there is something new, each listing goes to Claude with a prompt that asks for two judgements and insists they never contaminate each other. The first is stated. Does the listing explicitly say attendees get food or drink at no extra cost? The prompt is blunt about what doesn't count. Cash bar, food trucks, available for purchase, B Y O. And blunt about the fine print, because venues advertise free drink in the title, then qualify it at the bottom. Redeemable with your first purchase.
The honest guess
The second judgement is likelihood. An honest estimate that you actually get fed, even when the text says nothing. Gallery openings, seventy to eighty five. Corporate networking, sixty five to eighty. Guided tours and workshops, five to fifteen. Be calibrated, not generous, and the prompt says why. A reader sees that percentage next to a warning that the description mentions nothing.
No quote, no claim
Then the rule the whole app rests on. If Claude says the listing states free food, it has to quote the exact sentence, and that quote is not taken on trust. A function checks the words really do appear in the listing they are attributed to. If they don't, the claim isn't downgraded, it is deleted, and the event drops to an estimate. Across six hundred verdicts it has never once caught a fabrication. Both apparent catches were the check being too strict, and one cost six genuine verdicts.
Two backends, one judgement
There are two ways to run that classifier, and only one judgement. The prompt, the verdict shape and the quote check live in one file, and both routes import them rather than restating them. One route is the Anthropic A P I, which costs money per event. The other shells out to Claude Code itself, on a subscription already paid for, which costs nothing. That is the one that runs. The subprocess gets no tools, no file access and no settings. It is a text in, JSON out call that happens to be a process.
The answer key
Then the most conservative rule in the codebase. Verdicts are never written over the top of the file that holds them. New addresses are appended, and one already there is left exactly as it was. Three hundred and forty four of those rows were reached by reading each listing by hand. They are the answer key the classifier is scored against, and it agrees with them ninety per cent of the time.
Building the payload
Stage four builds the file the app downloads. It takes every verdict, runs the duplicate matcher a second time, which is free and lands a fix without re-reading a listing, then writes only the fields the app reads, under short names. Source url becomes url. Likelihood reason becomes why. That payload ships to every phone that opens the site, so a field name is bytes. It also enforces rule one. Where Claude read the text and said this is not free, the event is dropped. Seventy went that way.
Five decorators, in order
Then five small programs decorate that file in place, and the order is load bearing. Times are recovered, two hundred and seventy nine of them pulled out of prose the source never structured. Blurbs are stamped on. Past events are pruned, and five hundred and twenty five becomes three hundred and thirty five. Only then are pictures fetched, because that is the one decorator costing a request per event, and pruning first means reading three hundred listings rather than five hundred.
Deleting, not writing
The last decorator is the nicest. A host's title is written to sell a listing, not to be scanned in a list. Soda Tuesday, bar, Musical Bingo, bar, free drink on arrival, one dollar hot dogs. You want Musical Bingo. No rule gets there, so it is read, by that same free Claude route, under one hard constraint. Every word in the short name must already appear in the host's own title. It is deleting, not writing, and that is checked.
Shipping it
Stage five. If the events file changed, the laptop commits it and pushes. Cloudflare sees the push and copies the site onto its machines around the world. No worker script, no build step. The config file is nine lines long and says, serve this directory. The back end never talks to the front end. It writes a file, and the file is the interface.
Six in the morning
All of it is triggered by Windows Task Scheduler at six each morning. Not GitHub Actions, because the repository is private and those minutes are metered. The setting that matters is start when available. This laptop is usually off overnight, so a plain six a m trigger would simply never fire.
The shape of it
Eight thousand lines, sixteen sources, nine filters, and a language model used in exactly one place, where judgement is genuinely needed. Everywhere else, plain boring code you can read. No database, no server, no accounts, no bill. Websites in, one file out, and a phone that tells you where dinner is.
Where the numbers come from
Every figure here was read off the project itself on
21 September 2026 — the refresh log at
data/logs/refresh-20260921-113507.log, the answer key at
data/classified.jsonl, and the payload
app/local/events.json as it stood that morning. Nothing is
illustrative. The numbers move every day; these are one real day's.