Crashers, explained

The written version

Every step, in order

The film covers the same ground in five minutes. This page is the one to read if you want the file names, the exact numbers, and the bits that are still rough. Everything here describes the project as it stood on 21 September 2026.

What the app is trying to answer

Crashers lists a free public event only if all four of these are true. Almost every piece of machinery below exists to decide the second one, because it is the only one a computer cannot read straight off a page.

RuleWhat it means
$0 to attendNo ticket price and no membership requirement. A free RSVP is fine.
Food or drink, freeProvided by the host at no extra cost. A cash bar is not it. Food trucks on site are not it. BYO is not it.
Open to the publicInvite-only and members-only events do not qualify.
In greater SydneyA union of nine catchments, not one circle — see below.

Every listed event also carries an access tier, which is the thing no other events site bothers to tell you: can you simply turn up, or do you have to book, and if you book might you still miss out?

Turn up, nothing to book
90
Free RSVP, places not capped
175
Must book, places capped
58
The listing doesn't say
12

Stage 1 — Reading the websites

There is no data feed and no partnership. For each website there is a small Python program — an adapter — that opens the site's public events pages and pulls out the title, date, venue, price and description. These are the same pages anyone can open in a browser. Thirteen sources reach the current payload:

SourceListings returnedWhy it's in
Eventbrite AU920The biggest single vein of free Sydney events. Listing pages carry no description, so the crawl geo-filters first and only fetches detail pages for survivors.
City of Sydney What's On699Publishes a sitemap and structured data including freeEvent and coordinates. Community centres are the richest source of morning teas and barbecues.
City of Parramatta390One of nine councils crawled through shared council adapters.
Meetup329The tech and startup scene, which reliably feeds a room.
Inner West Council205Marrickville, Newtown, Balmain, Leichhardt, Ashfield — outside the City of Sydney's boundary.
Fairfield City187Council listings.
Mosman Council154The best-structured council source. States its cost explicitly and runs three months ahead — 107 of its 154 are free.
Campbelltown City123Council listings.
City of Ryde102Council listings.
North Sydney95Council listings.
Hornsby Shire87Council listings.
Humanitix81The $0-ticket tail. A much higher free rate than Eventbrite, because the not-for-profit end of the market books here.
Penrith City6Small, but Penrith-specific — Eventbrite's own "sydney" search returns zero Penrith events.

Ku-ring-gai, Woollahra and Canada Bay are crawled as well; none of their events qualified on the day these numbers were taken. Across all fifteen crawl files the last full run returned 3,583 listings.

Asking permission first

Every website publishes a file at /robots.txt saying which automated visitors it welcomes. Crashers reads that file before a source is ever added, and the decision is re-checked rather than read back out of a table, because robots.txt changes in both directions.

Where a site names our bot and says no, we leave the site alone — even though the pages are public and taking them would be technically possible. Four sites currently sit in that bucket: Willoughby Council, TryBooking, Art Guide AU and Ticketebo. Humanitix used to be one of them and no longer names us, which is why it is now live: the rule did not change, the site did.

What being polite looks like in code

All fetching goes through one module, crashers/polite.py, so the manners cannot be skipped by a new adapter:

  • A minimum gap between requests per host — 0.5 s for the big ticketing platforms, up to 2.5 s for Inner West Council, which returned a 429 Too Many Requests at 0.4 s and was entirely right to.
  • Any 429 or server error doubles that host's delay for the rest of the run, then backs off exponentially before retrying.
  • The bot identifies itself with a name and a contact address, so a site owner who wants it to stop has somewhere to write.
  • Only a short evidence snippet is stored, never a republished description, and every event links back to the host.

Nothing is ever fetched twice

Every page fetched is written to data/raw/ on the laptop's disk — 1,843 Eventbrite pages, 1,109 City of Sydney pages, 259 Inner West, 175 Mosman, 81 Humanitix, 10 Penrith. Re-running the classifier, or fixing a bug in how a date is read, costs those websites nothing at all, because the pipeline re-reads the copy on disk.


Stage 2 — Throwing away the obvious no's

3,583 listings is far too many to think hard about, and thinking hard is the only expensive step. So four cheap, exact filters run first, followed by a plain keyword search. This is the funnel from the real log:

Funnel from 3,583 crawled listings down to 356 candidates 3,583 Everything the crawl returned 15 crawl files, 13 sources in the payload 2,682 Price not stated as paid − 901 carried a ticket price 2,159 Happening in the next 28 days − 523 too far off 2,078 Inside one of the nine catchments − 81 out of area 2,017 After duplicates are merged − 61 listed on two different sites 356 Worth a closer look 305 whose text claims food or drink, 51 whose format suggests it. 1,661 skipped.
Bar width is proportional to the count. The last step is the keyword gate, which is why the drop is so steep.

The nine catchments

Coverage is a union of catchments rather than one radius, because Penrith is about 50 km from the CBD and no single circle describes both without dragging in everything between them:

Sydney CBD 10 km · Parramatta 12 km · Blacktown 11 km · Penrith 15 km · Liverpool 13 km · Hawkesbury 13 km · Blue Mountains 22 km · Campbelltown 13 km · Upper North Shore 20 km.

The keyword gate

The last filter (crashers/prefilter.py) is a list of regular expressions looking for phrases like morning tea, canapés, light refreshments, sausage sizzle, complimentary, wine and cheese, book launch — plus a list of negatives like cash bar and food available to purchase.

It is deliberately tuned for recall, not precision. Letting junk through is fine, because the next stage sorts it out. Dropping a real free-feed event here is not, because nothing downstream can recover it.


Stage 3 — Claude reads what's left

This is the only stage that makes a judgement. Each of the 356 survivors is handed to Claude with the host's own listing text, and comes back with a structured verdict. The critical design decision is that it produces two answers that are never allowed to contaminate each other.

The classifier produces a stated verdict with a verbatim quote, or an estimate with a percentage The listing title, description, price, venue, date Claude claude-haiku-4-5 structured verdict Host stated The listing explicitly says so — and Claude must quote the exact sentence. No quote, no claim. 191 of the 335 published events Estimate The listing never mentions food or drink, so a percentage and a one-line reason are given instead. 144 of the 335 — always labelled, always filterable
Mixing these two would poison the confirmed feed, which is the app's whole value. Keeping them apart is what lets a reader say "only sure things" or "I'll take my chances tonight".

No quote, no claim

This is the project's firmest rule and it is enforced twice — once in crashers/classify.py, and again in crashers/blurb.py where the written line is what people actually read. A positive food or drink verdict must carry a verbatim evidence string lifted out of the host's own text. The app then shows that sentence to the reader.

“A networking reception with drinks and catering will follow”

The effect is that Crashers is reporting what a host said rather than making a promise of its own. Where a host says only “light refreshments provided”, the tag says only Light refreshments — inventing a menu for them would be making things up.

What it costs to run, and how it's checked

The classifier has two interchangeable back ends behind one shared prompt. By default, with no API key set, it runs through claude -p on the Claude Code subscription, which is what the nightly task uses and costs nothing extra. With a key it uses the Anthropic API with Haiku 4.5 instead — the dry run for the 241 genuinely new events on this particular morning estimated $0.47.

Accuracy is measured against a hand-labelled answer key: 344 listings read and judged by a person, committed to the repository as data/classified.jsonl. The model currently scores 90.4% against it. That validator has one trap worth naming, because it caught the project out: the nightly run writes its own verdicts into the same file, so the scorer has to filter down to the hand-labelled rows or it ends up marking the model's homework and drifting upward forever.


Stage 4 — It all becomes one file

Everything so far produces verdicts scattered across working files. publish.py turns them into the single file the app actually downloads, and three small programs then decorate it.

publish.py writes events.json, then add_times, add_blurbs and prune_past decorate it PROGRAM WHAT IT DOES TO THE FILE publish.py reads every verdict add_times.py stamps the clock on add_blurbs.py the one-line description prune_past.py drops old dates events .json 251 KB 595 classified → 525 published − 70 that turned out to cost money after all 279 times read from pages already on disk; 140 state none 244 written lines, 265 extracted from the host's text 525 → 335 − 190 already been and gone
Order matters: publish.py writes only the fields the app reads, and the three decorators own the rest. Run with --keep so that times and written lines survive a rebuild instead of having to earn themselves again.

One detail worth pulling out. The one-line description under each title is the single place in the whole pipeline where words are put in a host's mouth, and it still obeys the rule: the free food or drink is named only from the verbatim quote, never from the model's sense of what such an event usually serves. Everything else must come from the listing — the model compresses, it does not fill gaps.


Where the data actually lives

There is no database

No Supabase. No Postgres. No server of our own, no login, no user accounts, no API. The entire state of the app is four files in a folder, versioned with Git:

  • index.html — the app itself, including all its logic
  • events.json — the 335 events, 251 KB
  • crashers.css + tokens.css + app.css — the look
  • _headers — the caching rules

For 335 events in a quarter of a megabyte, a database would be more machinery to go wrong, not less. It would need hosting, backups, credentials, a connection pool and a migration story, and in return it would answer a question — "what's on?" — that a static file already answers for free, from a machine down the road, with no request to a server at all.

The working data upstream of the app also lives in plain files: data/raw/ holds every page ever fetched, and the intermediate stages are JSONL — one JSON object per line. Two of those files are committed to the repository on purpose, because they cannot be regenerated by re-crawling: data/classified.jsonl, the hand-labelled answer key, and data/blurbs.json, the written descriptions, so a fresh clone gets them for nothing.


Stages 5 & 6 — GitHub and Cloudflare

From the laptop through GitHub and Cloudflare Workers to a phone The laptop 6:00 am, refresh.py Did it change? git diff --quiet -- app/local GitHub the safe copy, and the starting gun Workers Build npx wrangler deploy assets only, no script Cloudflare's edge the four files, copied to machines worldwide PRIVATE, ON ONE MACHINE PUBLIC INFRASTRUCTURE no change → no push → no deploy, and nothing is wasted

Git is a history book: every change to the project, in order, permanently. GitHub is where that book is kept online. At the end of each morning's run, refresh.py checks whether events.json actually changed, and if it did, commits and pushes it. GitHub then does two jobs — it is the off-site copy, and its receiving a push is the signal that starts the deploy.

An assets-only Worker, not a server

Crashers deploys as a Cloudflare Worker serving static assets — deliberately not Pages, and that distinction cost two failed builds. The whole configuration is four lines in wrangler.jsonc:

{ "name": "crashers",
  "compatibility_date": "2026-09-08",
  "assets": { "directory": "./app/local" } }

There is no main entry, because there is no Worker script. Nothing executes when you open the app. Cloudflare has the four files already and hands out the same four to everyone — which is why it is instant, and why it survives any amount of traffic on a free plan.

Response headers live in app/local/_headers rather than in host configuration, so the rules travel with the files: events.json may be held for five minutes, index.html and / must never be held at all — a reader running last week's app against this morning's data is the one combination that breaks.

Cloudflare only rebuilds when something in app/local/ changed. A quarter of this project's commits touch the crawler or the notes and would otherwise publish a byte-identical site.

Why the hosting moved, and what it taught

The site was on Netlify until the free allowance ran out. Deploys then stopped silently: the live site sat eight commits behind for about seventeen hours, still answering 200 OK, still looking perfectly healthy, and serving a version with events the calendar had already passed. Nothing warned anyone. That failure looks exactly like success from the outside, which is the reason it is worth writing down.


Stage 7 — Your phone

You open the page; it downloads events.json once. From that moment everything happens inside the phone's browser: sorting by what's nearest, the map and its markers, the filters, the saved list, and the rule that hides an event whose stated finishing time has passed.

If you tap “Where am I?”, the browser gives the page your coordinates, the page re-sorts the list, and that is the end of it. There is nowhere to send a location to — there is no server. Until you tap it, distances are measured from the centre of the catchment an event sits in and labelled “from CBD” or “from Penrith”, so a number is never passed off as distance from you.

Every event links straight out to whoever is hosting it. Crashers is a discovery layer and never a ticketing one: it takes no bookings and no payments.


And again tomorrow

A Windows Task Scheduler entry runs scripts/refresh.py --commit at 6:00 every morning. The setting that matters is StartWhenAvailable: the laptop is usually off overnight, so a plain 6 am trigger would simply never fire; with it, a missed run happens as soon as the machine is next on.

It runs as the logged-in user rather than as SYSTEM, because that is where the API key lives if one is set. A source that fails is reported and skipped — one council being down is not a reason to publish nothing — and the classification step has a hard cost ceiling, so a source that changes shape and dumps two thousand rows into the feed cannot quietly spend money.

What it costs

Hosting
$0
Storage
$0
Classification
$0
Servers running
none

Cloudflare's free tier serves the site; GitHub stores it; the reading is done on a Claude Code subscription that is already paid for. The Anthropic API route exists as an alternative and would cost roughly fifty cents on a busy morning, but it is not the one the nightly task takes.


The rough edges

Anything that claims to be accurate has to include these.