Our team entered JustMate in the Sport & Healthcare task, didn't reach the final, and got no evaluation comments — like most teams. So we built a council that scores every finalist the same way, from what is public, and let any team send its own project for free. This page is what nine versions of panel design measured, what broke along the way, and where the council disagreed with the jury.
Three fixed models read an evidence pack built from the repo: measured facts such as code size, tests and commit history, the team's deck, committed decks and docs, the demo page text and descriptions of its screenshots — at most about 75k tokens. Each model scores every official criterion of the task from 0 to 10; the median counts, the task's official weights turn it into a score out of 100, and a fourth model writes the review. The council never runs the code, never sees the pitch, and for finalists is never told how the project placed.
The full design — the prompts, the per-task rubrics and reviewing guides built from the official rules, the queue that gives every project the identical panel — is in the repo.
Before the council, every project with public code had a blind review by one strong model. The council was measured against those blind reviews — mean absolute gap per project, and how often the top-ranked project was the same:
In every task with several reviewed projects, the council's top project was also the blind review's top project. It was never the jury's winner: in the five tasks whose winner was reviewed, the winners scored 57.5–69 (average 63.7, solidly strong) while the council's top project scored 66–83. The other five winners were never reviewed — four had no public repo found by the search, and one was deleted or made private after the event.
That gap is the finding, not a defect: the jury also had mentors' notes, the paper round and the live pitch; the council read only what was public. A high council score means strong public work, not that the jury was wrong.
The form was open for the event week, 4–11 October 2026: 62 reviews in total — 21 finalists and 41 projects teams sent in themselves. When the queue drained, the form closed and the site became this static archive; the repo keeps the council, the evidence builder, the prompts and every review.
If your project was reviewed, the number was written by the panel above, the same one every other project faced. Thanks to every team that sent its work in.
The council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.