A genuinely engineered control layer with real gateway and proxy enforcement, hot-reloaded policy and stack-booting tests, undermined by a simulated semantic half, a mockup dashboard, README overclaims and a missing required deck.
The deterministic half is real and deep: auth, model allowlist, checksum-validated PII redaction, a signature feed with self-test vectors, budgets and fail-closed behaviour under hot reload. The semantic half is a keyword-and-weight simulation and the authority-over-detection guarantees exist only in design docs, so the hybrid requirement is half met and the median of 5 is fair.
Members disagree here: scores range by 2.0 points.Two real in-path enforcement points exist in code, including a 3-layer proxy that terminates TLS, applies policy and re-encrypts with a virtual-to-real key swap, backed by committed latency benchmark results. Member C's claim that the performance numbers are placeholders in a deck is not supported, since no deck exists and two members independently saw measured results, so the higher median of 7 stands.
Members disagree here: scores range by 2.5 points.Audit JSONL with trace IDs and a generator producing separate management and security tables from real log data are genuine. The dashboard deliverable is a static mockup not wired to the gateway, there is no live metrics endpoint and no tamper-evident log chaining, and the required deck is missing, so 4 is deserved.
Members disagree here: scores range by 2.0 points.The 46 measured cases boot the real gateway, guard service and mock LLM as processes and include genuine negative cases such as prompt injection, budget exhaustion, guard outage and rejected policy edits. The README overstates this as 76 tests plus 120 browser checks, which the measured count does not support.
Members disagree here: scores range by 2.0 points.One-command runs, pinned dependencies, three strictness profiles and prefix routing that can point at local models are real. Budgets are in-memory per process with check-then-charge, there are no Docker or Kubernetes artifacts and the two prototypes were never unified; Member C's claim that the build has not started is contradicted by roughly 11k lines of working prototype code that all members reviewed.
Members disagree here: scores range by 2.5 points.claude:glm-5.3-flash69.070% agreedots-studio/dots-3-note-preview:free51.390% agreeinclusionai/ling-3.0-flash-sante:free49.560% agreeclaude:glm-5.3–judgeThe council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.