A strong, honest Kraków decision sandbox with real feeds, a working traffic and disaster engine and a live deploy, held back by zero tests, single-district coverage and uncalibrated heuristic outputs.
A real consequence engine (BPR travel times, Dijkstra routing, incremental traffic assignment) running on the actual OSM road graph, plus a rare OBSERVED/PREDICTED/SIMULATED honesty layer, earns 8. A's 7 is defensible since the underlying what-if concept and heuristic formulas are not new, but the computed insight and honesty layer support the median.
Directly municipal: ZTP transit feeds, OSM streets, traffic and access analysis, and fire or flood modeled as road-closing events that shift metrics. C's 9 is generous because it is a decision-exploration sandbox rather than an operational city tool, so the median 8 stands.
The live demo gives named users preview-confirm consequence reports, search, undo/redo and version compare, but outputs are uncalibrated heuristics the team itself admits can be inflated, coverage is one district and there is no scenario export. A's 6 is harsh given the working deployed demo; the evidence supports the median 7.
The 3D city is the native UI with day/night lighting keyed to local time, DEM terrain, an OSM minimap with camera cone, color-coded HUD metrics and a data-origin legend. Polish-only labels and some generic-looking placeholder metrics keep it at 7.5 rather than C's 8.
A real end-to-end pipeline across 14,726 lines: Overpass and GTFS ingest, a hand-written protobuf decoder polling GTFS-RT, CORS proxying, IndexedDB versioning, 30 commits all within event hours and a live deploy. C's 7 dings the client-only architecture, but the measured pipeline and deploy support the median 8 despite zero tests.
claude:glm-5.3-flash71.080% agreedots-studio/dots-3-note-preview:free77.0100% agreeinclusionai/ling-3.0-flash-sante:free79.070% agreeclaude:glm-5.3–judgeThe council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.