A genuinely deep and well-tested agent control layer whose evidence matches its claims, losing points for the local-model support the task expects but does not get, an unevidenced policy hot reload and a pack with no demo checks.
The hybrid is real: deterministic guards for secret stripping, keyed-hash PII redaction, IBAN checksums and invisible characters, value tracing that keeps untrusted labels across agent handoffs, and a genuine semantic model path in the Jev detector with AI fallback and a review queue. C's 7 is the outlier and names no defect the others missed, since C's own text confirms the controls are real, so the median stays high with a point off because signature coverage skips the code-execution, deserialization and model-repo classes the task names.
Members disagree here: scores range by 2.5 points.The SDK genuinely sits in the path, wrapping model and tool calls in-process before they run, on a clean split of webhook ingest, a WebSocket control service, a worker and a Postgres-backed dashboard. It holds at 8 because hot reload amounts to a live control link with no config watcher or editable rules view in evidence, and no latency numbers beyond per-check timing prints are shown.
Members disagree here: scores range by 3.0 points.Every model call, guard decision and approval is stored as a queryable event with rule and reason, incidents trace the path from entry point to damage with replay proof, and the dashboard runs on real data. On the disputed audit-log deliverable the evidence supports B and C, who independently find the logs exportable through the event store and search, over A's absence claim, but the median holds at 7 because that export path is indirect and the pack contains no demo checks to verify reporting at runtime.
Members disagree here: scores range by 2.5 points.The measured suite is huge and real: 4,180 cases in 826 files, about one test file per source file, runnable with pnpm test, including negative cases for injection signatures and blocked payments. It lands just below the top because the claimed full coverage is unverified, the claimed count of 4,500 runs slightly ahead of the measured 4,180, and no individual negative case was read in the pack.
Deployment is a single Docker Compose with health checks and migrations, an npm-published SDK, JSON policy files and pluggable guards with x402 payment caps. The median settles at 7 because the task's local-model expectation is missed: v1 is OpenAI-only, and the base-URL hook B credits is indirect and, as B itself concedes, may be limited by the OpenAI Responses API dependency.
Members disagree here: scores range by 2.0 points.claude:glm-5.3-flash81.090% agreedots-studio/dots-3-note-preview:free91.870% agreeinclusionai/ling-3.0-flash-sante:free68.060% agreeclaude:glm-5.3–judgeThe council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.