Before writing a line of production code for Ballpark, I ran the concept through a six-agent research gauntlet and an adversarial red team. The verdict on my pitch: NO-GO. The version that shipped four days later, on web and iOS, is the version that survived.
A free daily estimation game: one Fermi question a day (how heavy is the Eiffel Tower?) answered with a range. Tighter catches score more. Streaks, a share grid, the familiar daily-puzzle loop.
Instead of building it, I commissioned six research agents (narrative, customer discovery, market, competitive, feature, and design), each producing a written brief and, together, a working prototype. Then an adversarial red team was pointed at the whole case with one job: kill it.
The red team's findings, verbatim from the brief:
A kill verdict isn't a landfill. The gauntlet explicitly preserved the prototype's legibility system, the question-authoring spec (every answer sourced, with a truth window), and an honest per-question content cost model. The red team's criticisms became the shipped spec:
Misses now score by distance (catches max 100, misses max 15), so honest wide
ranges beat reckless tight ones in expectation. The spec labels it
RED-TEAM-FIXED.
The results scale is a fixed log scale centered on the truth, because the red team showed an adaptive scale made a 2× miss and a 200× miss look alike.
The streak-driven daily loop was demoted; the shipped product is built around a challenge-a-friend share loop, where every result is a link someone can try to beat.
The verdict memo pre-registered what would justify further investment: retention thresholds and real interview evidence, not launch-day traffic.
Because the kill decision is the job. The expensive failure mode in 0→1 work isn't building the wrong thing badly. It is building the wrong thing well. The gauntlet cost days and killed a month of misdirected work; what shipped is smaller, more honest, and live on the web and the App Store.