showrunner.nz · 2026
An AI assistant with its guardrails in code, not in the prompt
showrunner.nz finds, compares and prices any car NZ dealers sell. People ask in their own words and a language model reads it. What that model can see and spend is decided by the server, not its prompt.
- 10,278 cars from 129 NZ dealers
- 0 invented figures across 77 eval turns
- US$0.004 median cost per message
- Live at showrunner.nz

What it does
Used-car stock in New Zealand is spread across hundreds of yard websites. Showrunner puts it in one index, with a page per model, rankings, EV running costs and a trade-in estimate. The assistant is the front door: it turns a sentence into filters over real cars.
The interface
Every car, price and chart on screen is drawn by code from the index, never written by the model.


Architecture
Prices take the slow path, with a person reading the diff. A typed message takes the fast one, with the model boxed in on both sides.
Prices and pages
- Dealer stock
- Ingest service, admin only
- Postgres index
- Reviewed price snapshot
- Static model pages
The assistant
- A typed message
- Route guards
- Model, up to 3 calls
- Five read-only tools
- Sentence gate
- Blocks rendered by code
- One package owns every query
- Routes, pages and scripts all go through one inventory package, so the search rules can’t drift apart between callers.
- Prices are reviewed before they publish
- Typical prices are computed once and committed as a snapshot. A person reads the diff before it goes live, and a figure without enough listings behind it is left out.
- Ingest has no public door
- The service that writes to the index sits behind an admin token with no public domain. The public site only reads.
- Migrations only go forward
- Hand-written, forward-only migrations, tested on a database branch that gets thrown away.
Guardrails in code
A prompt is a request. These limits are enforced by the server, whatever the model decides.
- Off-topic fails closed
- The first call has to say whether a message is in scope. Out of scope, or no verdict, and no tool runs.
- Dealer text never reaches the model
- The model sees year, kilometres, price and town. Never a dealer’s name, link or description, so injected text has nowhere to hide.
- Every number is checked on the way out
- A sentence with a number that isn’t in the question or the tool results is replaced before it reaches the screen.
- Spend is claimed before it’s spent
- Each call reserves its worst-case cost against a daily ceiling, then settles at its real usage. A message gets three calls at most.
- Limits the browser can’t fake
- Thread length is counted on the server. Messages are capped at 600 characters, behind per-IP rate limits and bot protection.
- It works without the model
- Over a limit, timed out or refused, plain code reads the message and the person still gets cars.

Evals that block a change
Any change to a prompt or tool description runs the eval suite in the same commit. One invented figure fails the change.
- 98.7%
- right tool and arguments, across 77 turns
- 0
- invented figures
- US$0.0043
- median cost per message
- 8 of 8
- trade-in interviews read correctly
What testing cut
- Semantic search came out
- It couldn’t tell a vague but real request, like “a tidy first car for my daughter”, from something no yard sells. Plain filters and text search did better.
- Synonym lists came out
- Nobody can keep up a list of which cars are sold under two names. The model already knows, so it searches both.
- Stack
- Next.js, TypeScript, Postgres on Neon, Drizzle, Cloudflare Workers, Vercel
- Data
- NZ dealer listings, Rightcar safety and economy, MBIE fuel prices, UK MOT results
- Timeframe
- June 2026 to now
Want the details?
I’m happy to walk through the architecture and the evals.