The written version
Every step, in order
The film covers the same ground in five minutes. This page is the one to read if you want the file names, the exact numbers, and the bits that are still rough. Everything here describes the project as it stood on 21 September 2026.
What the app is trying to answer
Crashers lists a free public event only if all four of these are true. Almost every piece of machinery below exists to decide the second one, because it is the only one a computer cannot read straight off a page.
| Rule | What it means |
|---|---|
| $0 to attend | No ticket price and no membership requirement. A free RSVP is fine. |
| Food or drink, free | Provided by the host at no extra cost. A cash bar is not it. Food trucks on site are not it. BYO is not it. |
| Open to the public | Invite-only and members-only events do not qualify. |
| In greater Sydney | A union of nine catchments, not one circle — see below. |
Every listed event also carries an access tier, which is the thing no other events site bothers to tell you: can you simply turn up, or do you have to book, and if you book might you still miss out?
- Turn up, nothing to book
- 90
- Free RSVP, places not capped
- 175
- Must book, places capped
- 58
- The listing doesn't say
- 12
Stage 1 — Reading the websites
There is no data feed and no partnership. For each website there is a small Python program — an adapter — that opens the site's public events pages and pulls out the title, date, venue, price and description. These are the same pages anyone can open in a browser. Thirteen sources reach the current payload:
| Source | Listings returned | Why it's in |
|---|---|---|
| Eventbrite AU | 920 | The biggest single vein of free Sydney events. Listing pages carry no description, so the crawl geo-filters first and only fetches detail pages for survivors. |
| City of Sydney What's On | 699 | Publishes a sitemap and structured data including freeEvent and coordinates. Community centres are the richest source of morning teas and barbecues. |
| City of Parramatta | 390 | One of nine councils crawled through shared council adapters. |
| Meetup | 329 | The tech and startup scene, which reliably feeds a room. |
| Inner West Council | 205 | Marrickville, Newtown, Balmain, Leichhardt, Ashfield — outside the City of Sydney's boundary. |
| Fairfield City | 187 | Council listings. |
| Mosman Council | 154 | The best-structured council source. States its cost explicitly and runs three months ahead — 107 of its 154 are free. |
| Campbelltown City | 123 | Council listings. |
| City of Ryde | 102 | Council listings. |
| North Sydney | 95 | Council listings. |
| Hornsby Shire | 87 | Council listings. |
| Humanitix | 81 | The $0-ticket tail. A much higher free rate than Eventbrite, because the not-for-profit end of the market books here. |
| Penrith City | 6 | Small, but Penrith-specific — Eventbrite's own "sydney" search returns zero Penrith events. |
Ku-ring-gai, Woollahra and Canada Bay are crawled as well; none of their events qualified on the day these numbers were taken. Across all fifteen crawl files the last full run returned 3,583 listings.
Asking permission first
Every website publishes a file at /robots.txt saying which
automated visitors it welcomes. Crashers reads that file before a source is
ever added, and the decision is re-checked rather than read back out of a
table, because robots.txt changes in both directions.
Where a site names our bot and says no, we leave the site alone — even though the pages are public and taking them would be technically possible. Four sites currently sit in that bucket: Willoughby Council, TryBooking, Art Guide AU and Ticketebo. Humanitix used to be one of them and no longer names us, which is why it is now live: the rule did not change, the site did.
What being polite looks like in code
All fetching goes through one module, crashers/polite.py, so
the manners cannot be skipped by a new adapter:
- A minimum gap between requests per host — 0.5 s for the big
ticketing platforms, up to 2.5 s for Inner West Council,
which returned a
429 Too Many Requestsat 0.4 s and was entirely right to. - Any
429or server error doubles that host's delay for the rest of the run, then backs off exponentially before retrying. - The bot identifies itself with a name and a contact address, so a site owner who wants it to stop has somewhere to write.
- Only a short evidence snippet is stored, never a republished description, and every event links back to the host.
Nothing is ever fetched twice
Every page fetched is written to data/raw/ on the laptop's disk —
1,843 Eventbrite pages, 1,109 City of Sydney pages, 259 Inner West, 175
Mosman, 81 Humanitix, 10 Penrith. Re-running the classifier, or fixing a bug
in how a date is read, costs those websites nothing at all, because the
pipeline re-reads the copy on disk.
Stage 2 — Throwing away the obvious no's
3,583 listings is far too many to think hard about, and thinking hard is the only expensive step. So four cheap, exact filters run first, followed by a plain keyword search. This is the funnel from the real log:
The nine catchments
Coverage is a union of catchments rather than one radius, because Penrith is about 50 km from the CBD and no single circle describes both without dragging in everything between them:
Sydney CBD 10 km · Parramatta 12 km · Blacktown 11 km · Penrith 15 km · Liverpool 13 km · Hawkesbury 13 km · Blue Mountains 22 km · Campbelltown 13 km · Upper North Shore 20 km.
The keyword gate
The last filter (crashers/prefilter.py) is a list of regular
expressions looking for phrases like morning tea, canapés,
light refreshments, sausage sizzle, complimentary,
wine and cheese, book launch — plus a list of negatives like
cash bar and food available to purchase.
It is deliberately tuned for recall, not precision. Letting junk through is fine, because the next stage sorts it out. Dropping a real free-feed event here is not, because nothing downstream can recover it.
Stage 3 — Claude reads what's left
This is the only stage that makes a judgement. Each of the 356 survivors is handed to Claude with the host's own listing text, and comes back with a structured verdict. The critical design decision is that it produces two answers that are never allowed to contaminate each other.
No quote, no claim
This is the project's firmest rule and it is enforced twice — once in
crashers/classify.py, and again in crashers/blurb.py
where the written line is what people actually read. A positive food or drink
verdict must carry a verbatim evidence string lifted out of the
host's own text. The app then shows that sentence to the reader.
“A networking reception with drinks and catering will follow”
The effect is that Crashers is reporting what a host said rather than making a promise of its own. Where a host says only “light refreshments provided”, the tag says only Light refreshments — inventing a menu for them would be making things up.
What it costs to run, and how it's checked
The classifier has two interchangeable back ends behind one shared prompt.
By default, with no API key set, it runs through claude -p on the
Claude Code subscription, which is what the nightly task uses and costs
nothing extra. With a key it uses the Anthropic API with Haiku 4.5 instead —
the dry run for the 241 genuinely new events on this particular morning
estimated $0.47.
Accuracy is measured against a hand-labelled answer key: 344 listings read and
judged by a person, committed to the repository as
data/classified.jsonl. The model currently scores
90.4% against it. That validator has one trap worth naming,
because it caught the project out: the nightly run writes its own verdicts
into the same file, so the scorer has to filter down to the hand-labelled rows
or it ends up marking the model's homework and drifting upward forever.
Stage 4 — It all becomes one file
Everything so far produces verdicts scattered across working files.
publish.py turns them into the single file the app actually
downloads, and three small programs then decorate it.
publish.py writes only the fields the app reads,
and the three decorators own the rest. Run with --keep so that
times and written lines survive a rebuild instead of having to earn
themselves again.
One detail worth pulling out. The one-line description under each title is the single place in the whole pipeline where words are put in a host's mouth, and it still obeys the rule: the free food or drink is named only from the verbatim quote, never from the model's sense of what such an event usually serves. Everything else must come from the listing — the model compresses, it does not fill gaps.
Where the data actually lives
There is no database
No Supabase. No Postgres. No server of our own, no login, no user accounts, no API. The entire state of the app is four files in a folder, versioned with Git:
index.html— the app itself, including all its logicevents.json— the 335 events, 251 KBcrashers.css+tokens.css+app.css— the look_headers— the caching rules
For 335 events in a quarter of a megabyte, a database would be more machinery to go wrong, not less. It would need hosting, backups, credentials, a connection pool and a migration story, and in return it would answer a question — "what's on?" — that a static file already answers for free, from a machine down the road, with no request to a server at all.
The working data upstream of the app also lives in plain files:
data/raw/ holds every page ever fetched, and the intermediate
stages are JSONL — one JSON object per line. Two of those files are committed
to the repository on purpose, because they cannot be regenerated by
re-crawling: data/classified.jsonl, the hand-labelled answer key,
and data/blurbs.json, the written descriptions, so a fresh clone
gets them for nothing.
Stages 5 & 6 — GitHub and Cloudflare
Git is a history book: every change to the project, in order, permanently.
GitHub is where that book is kept online. At the end of each morning's run,
refresh.py checks whether events.json actually
changed, and if it did, commits and pushes it. GitHub then does two jobs — it
is the off-site copy, and its receiving a push is the signal that starts the
deploy.
An assets-only Worker, not a server
Crashers deploys as a Cloudflare Worker serving static assets
— deliberately not Pages, and that distinction cost two failed builds. The
whole configuration is four lines in wrangler.jsonc:
{ "name": "crashers",
"compatibility_date": "2026-09-08",
"assets": { "directory": "./app/local" } }
There is no main entry, because there is no Worker
script. Nothing executes when you open the app. Cloudflare has the
four files already and hands out the same four to everyone — which is why it
is instant, and why it survives any amount of traffic on a free plan.
Response headers live in app/local/_headers rather than in host
configuration, so the rules travel with the files: events.json
may be held for five minutes, index.html and / must
never be held at all — a reader running last week's app against this morning's
data is the one combination that breaks.
Cloudflare only rebuilds when something in app/local/ changed. A
quarter of this project's commits touch the crawler or the notes and would
otherwise publish a byte-identical site.
Why the hosting moved, and what it taught
The site was on Netlify until the free allowance ran out. Deploys then
stopped silently: the live site sat eight commits behind for about
seventeen hours, still answering 200 OK, still looking
perfectly healthy, and serving a version with events the calendar had
already passed. Nothing warned anyone. That failure looks exactly like
success from the outside, which is the reason it is worth writing down.
Stage 7 — Your phone
You open the page; it downloads events.json once. From that
moment everything happens inside the phone's browser: sorting by what's
nearest, the map and its markers, the filters, the saved list, and the rule
that hides an event whose stated finishing time has passed.
If you tap “Where am I?”, the browser gives the page your coordinates, the page re-sorts the list, and that is the end of it. There is nowhere to send a location to — there is no server. Until you tap it, distances are measured from the centre of the catchment an event sits in and labelled “from CBD” or “from Penrith”, so a number is never passed off as distance from you.
Every event links straight out to whoever is hosting it. Crashers is a discovery layer and never a ticketing one: it takes no bookings and no payments.
And again tomorrow
A Windows Task Scheduler entry runs scripts/refresh.py --commit
at 6:00 every morning. The setting that matters is
StartWhenAvailable: the laptop is usually off overnight, so a
plain 6 am trigger would simply never fire; with it, a missed run happens as
soon as the machine is next on.
It runs as the logged-in user rather than as SYSTEM, because that is where the API key lives if one is set. A source that fails is reported and skipped — one council being down is not a reason to publish nothing — and the classification step has a hard cost ceiling, so a source that changes shape and dumps two thousand rows into the feed cannot quietly spend money.
What it costs
- Hosting
- $0
- Storage
- $0
- Classification
- $0
- Servers running
- none
Cloudflare's free tier serves the site; GitHub stores it; the reading is done on a Claude Code subscription that is already paid for. The Anthropic API route exists as an alternative and would cost roughly fifty cents on a busy morning, but it is not the one the nightly task takes.
The rough edges
Anything that claims to be accurate has to include these.
-
The app labels distances against two catchments, not nine.
The pipeline accepts events from nine, but
index.htmlonly mirrors the CBD and Penrith. That is why a Liverpool event can currently read “37 km from Penrith”. Known, written down in the project's own README, not yet fixed. - 140 of the 525 events published that morning state no time at all. The app says “No time listed” rather than inventing one, but it means a third of the feed can't answer “is it on right now?”.
- The estimates are estimates. 144 of the 335 events are in the feed because the format suggests food, not because anyone said so. They are labelled and filterable, but a 72% is still a 72%.
- The classifier is right about 90% of the time against the hand-labelled set. One in ten is wrong somewhere.
- It depends on one laptop being switched on. The daily refresh is a scheduled task on a personal machine, not a cloud job. If the laptop stays off, the feed goes stale — though the site itself keeps serving, because the site is only files.
- Websites change shape without warning. Every adapter is a guess about the structure of someone else's HTML, and each one will break eventually.