An unusually complete and honest control layer whose real hybrid guardrails and tamper-evident reporting are held back by judge latency, single-process scaling, mocked MCP tools and an incomplete submission form.
Deterministic detectors plus a real Ollama judge with structured risk scoring and fail-closed behaviour blocked 20 of 20 semantic attacks and redacted all four authored personal-data cases in the committed live run. The median sits below B's 9 because of one documented false block, PII detectors that miss names and postal addresses, and evidence limited to the team's own 48 authored cases rather than a held-out set.
Members disagree here: scores range by 2.0 points.The gateway is verifiably in the request path for chat and MCP traffic, with hot policy reload that keeps the last valid snapshot and a measured 8.7 ms median overhead without AI checks. The AI judge adds about 4.83 s median and the single-process SQLite design has no streaming, which caps the score at A's 8 rather than B's 8.5.
All members confirm the SHA-256 chained, Ed25519 signed audit log with rule IDs and policy hash but no raw prompts, plus JSONL export, a signed checkpoint, an offline verifier, a live dashboard and Prometheus metrics. A and B's 9 stands because C's deductions, such as no GPU measurement or dollar-cost tracking, are extras the brief does not ask for.
Members disagree here: scores range by 2.0 points.The pack shows 23 test files with 193 measured cases, 8 Rego tests in a real OPA container and a 48-case live evaluation with expected outcomes, all runnable offline via make check. B and C repeat the team's claim of 494 pytest cases, but the measurement supports A's more careful 8, and the live cases are authored rather than held-out.
Ollama is first-class with a documented activation flow, the YAML policy and make targets are clear, and a 130-dependency licence inventory is unusually thorough. The median stays at 7 because the Postgres scale-out is described but not built, there are no monetary caps or per-minute rate limits, and the full stack was validated only on Apple-silicon macOS with the Docker Ollama path untested.
claude:glm-5.3-flash82.0100% agreedots-studio/dots-3-note-preview:free87.570% agreeinclusionai/ling-3.0-flash-sante:free71.570% agreeclaude:glm-5.3–judgeThe council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.