A genuinely in-the-path control layer with real hybrid guardrails, signed feeds, hash-chained auditing and a runnable test suite, weakened by self-authored evaluation, an unexplained 174-versus-68 test claim, undisclosed heavy AI assistance, and a live demo that resists independent inspection.
Deterministic contract, limit and binding controls combine with a real model cascade and a quote-verified judge that can only tighten, all fail-closed. The 9 rather than A's 9.2 reflects that the 0-harm evidence is self-authored, and the empty live injection probe is inconclusive because the JavaScript demo exposes little to automated checks, so Member A's and B's reading is better supported than C's claim of a superficial detector.
Members disagree here: scores range by 2.2 points.A true reference monitor where the agent holds no keys, with policy, contract and feed hot reload implemented in code and committed benchmark artifacts showing p50 0.17 ms and about 3,000 decisions per second. The numbers come from the team's own committed files rather than an independent re-run, and shared SQLite WAL state is the acknowledged scaling ceiling.
Members disagree here: scores range by 2.0 points.The audit log is hash-chained with a verify endpoint, CSV export carries the deciding control, judge quote and policy hash, and Prometheus metrics plus the evidence dashboard give investigation-grade context. The depth is demonstrated through code and screenshots rather than an inspected live incident.
Members disagree here: scores range by 2.0 points.68 measured cases across 10 files, explicitly positive and negative, runnable by judges without AI, including tampered-feed and fail-closed tests. The claimed 174 control tests are not supported by the measured count and the gap is unexplained.
Members disagree here: scores range by 3.5 points.Two commands to run, Dockerfile and a live Railway deployment, Ollama-first with a hosted no-Ollama mode, YAML-only contracts for adding tools, and a small StateStore interface giving a credible path to Postgres or Redis.
Members disagree here: scores range by 2.0 points.claude:glm-5.3-flash89.990% agreedots-studio/dots-3-note-preview:free81.5100% agreeinclusionai/ling-3.0-flash-sante:free67.060% agreeclaude:glm-5.3–judgeThe council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.