TresMachines · ljaniec/Mathguard
A genuinely deep and honestly disclosed AI control layer whose enforcement, persistence, reporting and tests demonstrably exist in code, held back from the top band by unproven held-out detection accuracy, a synthetic signature feed, a single-node serialized deployment scope and a demo-light submission that is heavy for judges to run.
Deterministic controls are real code (PII, secret and encoded or Unicode patterns, hash-pinned artifacts, tool allowlists, reserve-before-dispatch budgets) and the semantic layer is a genuine local Ollama call, combined by a compiled Lean kernel that proves a semantic allow can never relax a hard denial. The score stays below the top band because the classifier missed an indirect-injection variant, its only passing corpus was also used for tuning, and the signature feed is self-authored synthetic content.
The gateway is verifiably in the request path: loopback HTTP ingress with bounds and host checks, a serialized engine handling roles and sessions, a long-lived compiled Lean worker on stdin/stdout with strict response validation, then the local provider, plus implemented fail-closed hot reload with monotonic epochs and separately tracked control-path p95 latency. Throughput is capped by a single engine lock with at most eight connections, a measured benign chat took about 4.3 s, and there is no separate benchmark harness.
Reporting is real rather than mocked: a hash-chained SQLite journal with intent-before-mutation, trace IDs, reason codes and per-stage latency backs the events and sanitized JSONL audit export endpoints, and the dashboard polls live authenticated endpoints for decisions, versions, quarantine and latency percentiles. Minor deductions because the screenshots are frozen offline renders with controls disabled, export is read-only with no hosted link, and sanitization means investigators must consult the private journal for content.
A measured 203 cases across 9 files run via make test against the actual compiled Lean worker with no model download, covering hostile, encoded, split, Polish-language and corrupted-journal negative cases, with 189 runtime plus 14 static checks passing and zero failures. Held below 9 because the semantic corpus is the development set rather than held out, fixtures verify enforcement rather than live detector accuracy, and totals disagree across artifacts (deck 187 vs receipt 189).
Three documented make commands with a judge quickstart and doctor script, a documented JSON policy with strict, balanced and permissive profiles, and verified Ollama support with digest preflight make the project genuinely runnable. But judges must install the Lean toolchain with pinned Mathlib and pull model weights, there is no one-command container path, and the design is explicitly single-node and single-lock with no distributed quotas, no MCP transport or OAuth, and no paid-API adapter even though budget rules model cost.
claude:glm-5.3-flash82.580% agreedots-studio/dots-3-note-preview:free79.390% agreeinclusionai/ling-3.0-flash-sante:free70.070% agreeclaude:glm-5.3–judgeThe council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.