Status
How this was built, and whether it works
Written by Claude, which built it. The short version: the film and the site both work, the facts in them were taken from the project's own source and logs rather than from a description of it, and the two things I am least happy with are the synthetic voice and the absence of real captions.
What was made
- Film
- 4:59
- Scenes
- 13
- Frames rendered
- 8,975
- File size
- 17.9 MB
Plus three pages: the film with a full transcript, a written walkthrough with five flow charts, and this. Everything is in one public repository, including the source of the film itself.
Added afterwards — the back-end film
A second, longer film was built on 22 September 2026:
the back end, in detail. 9 min 59 s,
26 scenes, 17,988 frames, cut to the same 10-minute ceiling the first was
cut to five. It reuses this repository's drawing kit rather than restating
it — the shared src/ui.tsx gained a code panel, a bar and a
margin note, and StageRail learned to take a different list of
stages so one component draws either film's spine.
What is verified: every number in it was read off the Crashers project's own
logs and payload the same way the first film's were; tsc
passes; all 26 scenes were proofed as rendered stills before the video was
made, which is how four layout collisions were found and fixed; the MP4 is
1920×1080 h264 with an AAC track and ffprobe reads its duration as
599.7 s; and all four pages were measured in headless Edge at a
390 px viewport with scrollWidth === clientWidth. That
last check earned its keep — the new page's three-column source table
overflowed by 75 px, and a screenshot alone would not have shown it,
because headless Edge clips the same way on pages that are perfectly fine.
The fix generalised the existing .figure.scroll rule to a bare
.scroll, so a wide table scrolls sideways the way the diagrams
already did.
What is not verified: nobody has watched the film end to end yet, and the narration is the same synthetic voice as the first.
How it was built, in order
-
Read the actual project first
Before anything was written, the real repository was read: the README and
CLAUDE.md, the source adapters,polite.py,prefilter.py,classify.py,refresh.py,wrangler.jsonc,_headers, the scheduled-task installer, the git log, and — most usefully — two real refresh logs from 21 September 2026 plus the publishedevents.json.That is where the correction on the front page came from. The brief mentioned Supabase; the codebase has no Supabase in it, and no database of any kind. Saying so plainly was more useful than quietly explaining a system that does not exist.
-
Write the narration before any pictures
The script was written first, as thirteen scenes in one
narration.json, aimed at someone who does not write software and with a hard five-minute ceiling. It came in at 5 min 55 s on the first pass and was cut three times — trimming clauses rather than dropping ideas, then nudging the pace — until it measured 4:59. -
Record the voice, then cut the pictures to it
Narration is Microsoft Edge neural text-to-speech, an Australian voice, generated one file per scene. A script measures each file with
ffprobeand writes the real durations tosrc/timing.json; the film reads that file for its scene lengths. So the pictures are cut to the voice rather than the voice being squeezed into guessed lengths, and changing a sentence re-cuts the film automatically.The first choice for the voice was Windows' own built-in speech, which turned out to offer only the two old robotic desktop voices. Edge neural was the next thing tried and was good enough to keep. There is no music, as asked.
-
Draw everything in code
The film is a Remotion project — React that renders to video. There is no stock footage and no generated imagery anywhere in it. Every box, arrow, counter, funnel bar, globe, phone and tick is drawn as SVG or styled HTML and animated with springs. The Crashers plate mark is the app's own
favicon.svgpath, used unchanged.The film borrows the app's design tokens verbatim — warm paper
#FAF6EF, one terracotta accent#B4442A, Bodoni Moda for the lines that matter and Karla for everything functional, no shadows, a 1px rule as the only boundary. The explanation looks like the thing it explains, and that is not decoration: it makes the phone mock-up in scene twelve read as the real app rather than as an illustration.The seven-stage rail along the foot of every scene is the one deliberate comprehension device. It is generated from a single list in
theme.tsthat the overview diagram also reads, so the film cannot contradict itself about what the stages are. -
Check the frames before rendering 8,975 of them
Thirteen single frames were rendered and looked at, one near the end of each scene where the most is on screen. That found a string of real layout collisions — a caption printed over a card in scene five, a listing card crossing a divider in seven, a closing line landing on a number in eight, a label sitting on the Sydney dot in eleven, a badge over the phone in twelve, and radial labels touching the ring in thirteen. All of them were fixed and re-checked before the full render started.
-
Build the site from the same source of truth
The transcript on the front page is generated from
src/timing.jsonbyscripts/build_site.py— the same file the film is cut to. Typing the narration out a second time by hand is exactly how a transcript ends up saying something the voice never said. -
Put it in a browser and look at it
The site was served locally and screenshotted in headless Chrome at desktop and phone widths. That caught one genuine rendering fault: the Bodoni headings were coming out with broken, ghosted hairlines.
The cause turned out to be the typeface doing exactly what it is designed to do. Bodoni Moda is a Didone with an optical-size axis, and Chrome drives that axis from the font size — at heading sizes it pushes the hairline strokes below one pixel, where they half-disappear. Pinning the axis with
font-optical-sizing: nonefixed it. It is insite.csswith that explanation next to it, because the next person to see it will otherwise assume the font failed to load.
What I actually verified
| Claim | How it was checked |
|---|---|
| The numbers are real | Every figure was read out of refresh-20260921-083424.log,
refresh-20260921-113507.log or events.json.
The funnel arithmetic was checked to add up: 305 + 51 + 1,661 = 2,017,
and each drop matches the difference between its two rows. |
| The film is under five minutes | ffprobe on the finished MP4: 299.22 s. |
| It has sound | The container carries an AAC stereo track at 48 kHz. Measured loudness over a 20-second slice: mean −24.3 dB, peak −5.7 dB. That is speech, not a silent track. |
| The film renders correctly | Render exited clean at 8,975 of 8,975 frames. Fifteen frames were rendered and fourteen of them inspected at full 1080p. |
| The code is sound | tsc --noEmit passes with no errors and no unused symbols. |
| The site works in a browser | Served locally and loaded in headless Chrome at 1280 px and 390 px. All three pages render, the fonts load, the diagrams scale, and the video element reads the MP4's duration as 4:59, which means it found and parsed the file. |
What I did not verify
- I have not watched the film. I inspected still frames and the render log. The animation timings are derived from the measured length of each narration file, so the beats should land where they were designed to — but "should" is doing real work in that sentence, and the first person to watch it end to end will know more than I do.
- I have not heard the narration. Its length was measured; its pronunciation was not listened to. Text-to-speech mangles things, and the most likely candidates here are Ku-ring-gai, canapés and Humanitix.
- One browser engine only. Chromium. Not Safari, not Firefox. The pages are plain HTML and CSS with no scripts at all, so the risk is low, but it is not zero.
- The live GitHub Pages URL was published, not re-tested end to end. The site was verified locally; the deploy is the same files.
Am I happy with it?
Mostly yes, with two specific reservations.
What I think is good. The structure holds: seven stages, named once, visible at the foot of every scene, so it is hard to get lost. The numbers are real and the funnel is the strongest thing in the film, because 3,583 → 356 → 335 tells the whole story of the project without any jargon. Saying plainly that there is no database is the single most useful minute in it, because that is the thing an outsider assumes wrongly. And building the film out of the app's own design tokens paid off more than expected — it reads as one thing rather than as a presentation about a thing.
What I am not happy with — one. The voice is synthetic. It is
clear and correctly paced, and for a technical walkthrough that is most of the
job, but it is not a person and you can tell within a sentence. For a film made
to explain your work to your dad, a real voice would be better than a good
robot. The pipeline makes that easy to change: record thirteen files over the
written script, drop them into public/audio/, re-run the timing
script, re-render. Nothing else needs touching.
What I am not happy with — two. There are no captions burned into the video. The full transcript is on the front page, which is something, but it is not the same as synchronised captions for someone who is deaf or watching with the sound off. A WebVTT track built from the same timing file would be a small job and it should be done.
Smaller things. Scenes two and ten have more empty space than they need. The globe in scene eleven is an abstract disc rather than a real map, which is honest about it being a diagram but is the least informative picture in the film. And every number here is one day's snapshot: they were true on 21 September 2026 and are already drifting.
Does it work?
| Part | Verdict |
|---|---|
| The film plays, with sound, under five minutes | Verified |
| Every stated number matches the project's own logs | Verified |
| All three pages render on desktop and phone | Verified |
| The explanation is followable by a non-technical viewer | Designed for it, not yet tested on one |
| Synchronised captions | Not done — transcript only |
| A human voice | Not done — synthetic |
The honest summary: the thing that mattered most here was accuracy, and that part I am confident about, because it was taken from the source rather than from a summary of it — including the parts that are unflattering, which are all written down on the how it works page. The craft is decent. The voice is the weak link and is the first thing I would replace.
Rebuilding it
npm install
py -3 -m pip install edge-tts
py -3 scripts/tts.py # narration -> public/audio/, src/timing.json
py -3 scripts/build_site.py # timing.json -> the transcript in index.html
npx remotion studio # preview
npx remotion render src/index.ts Crashers out/crashers-explained.mp4 \
--codec=h264 --crf=23
Edit narration.json, re-run the first two, re-render. The scene
lengths follow the voice on their own.