AI Control LayerSelf-submittedNot a finalist (self-declared)

Mathguard

TresMachines · ljaniec/Mathguard

Council score

Median of 3 models, weighted by the task's official criteria
78.5 / 100

A genuinely deep and honestly disclosed AI control layer whose enforcement, persistence, reporting and tests demonstrably exist in code, held back from the top band by unproven held-out detection accuracy, a synthetic signature feed, a single-node serialized deployment scope and a demo-light submission that is heavy for judges to run.

Criteria · line = median, dots = each member
Robustness/guardrails30%
8.0

Deterministic controls are real code (PII, secret and encoded or Unicode patterns, hash-pinned artifacts, tool allowlists, reserve-before-dispatch budgets) and the semantic layer is a genuine local Ollama call, combined by a compiled Lean kernel that proves a semantic allow can never relax a hard denial. The score stays below the top band because the classifier missed an indirect-injection variant, its only passing corpus was also used for tuning, and the signature feed is self-authored synthetic content.

Architecture & performance20%
8.0

The gateway is verifiably in the request path: loopback HTTP ingress with bounds and host checks, a serialized engine handling roles and sessions, a long-lived compiled Lean worker on stdin/stdout with strict response validation, then the local provider, plus implemented fail-closed hot reload with monotonic epochs and separately tracked control-path p95 latency. Throughput is capped by a single engine lock with at most eight connections, a measured benign chat took about 4.3 s, and there is no separate benchmark harness.

Security reporting20%
8.0

Reporting is real rather than mocked: a hash-chained SQLite journal with intent-before-mutation, trace IDs, reason codes and per-stage latency backs the events and sanitized JSONL audit export endpoints, and the dashboard polls live authenticated endpoints for decisions, versions, quarantine and latency percentiles. Minor deductions because the screenshots are frozen offline renders with controls disabled, export is read-only with no hosted link, and sanitization means investigators must consult the private journal for content.

Test suite15%
8.0

A measured 203 cases across 9 files run via make test against the actual compiled Lean worker with no model download, covering hostile, encoded, split, Polish-language and corrupted-journal negative cases, with 189 runtime plus 14 static checks passing and zero failures. Held below 9 because the semantic corpus is the development set rather than held out, fixtures verify enforcement rather than live detector accuracy, and totals disagree across artifacts (deck 187 vs receipt 189).

Implementability & scalability15%
7.0

Three documented make commands with a judge quickstart and doctor script, a documented JSON policy with strict, balanced and permissive profiles, and verified Ollama support with digest preflight make the project genuinely runnable. But judges must install the Lean toolchain with pinned Mathlib and pull model weights, there is no one-command container path, and the design is explicitly single-node and single-lock with no distributed quotas, no MCP transport or OAuth, and no paid-API adapter even though budget rules model cost.

Source lines7,985
Tests203 cases
Claims built8.5 / 10
Task fitYes

Strengths

  • Compiled Lean decision kernel genuinely in the request path, with the semantic-allow-cannot-relax-a-hard-denial property proved and the worker process spawned and validated by the engine
  • Fail-closed discipline throughout: invalid policy keeps the last valid epoch, no valid policy blocks all traffic, and worker, journal or binary-tamper breakage denies or closes startup
  • Real local-model integration with provenance checks (Ollama digest, loopback-only, cloud refused) and unusually honest live evidence including a retained failure baseline and a disclosed indirect-injection miss
  • Durable hash-chained audit journal with intent-before-mutation and exact-replay idempotency, plus 203 runnable tests including attack cases that judges can run without downloading a model

Weaknesses

  • Semantic detection accuracy is unproven on unseen attacks: the only passing corpus was used for prompt tuning, one indirect injection slipped through input classification, and the signature feed is synthetic rather than fed by real advisories
  • Single-node, single-lock serialized engine with eight-slot ingress caps throughput and precludes distributed or multi-owner deployment, with no MCP transport, OAuth or paid-API adapter despite budget rules modeling cost
  • Heavy judge-side prerequisites: the full Lean and Mathlib build plus an Ollama weight download are needed to see anything live, and there is no one-command container path
  • Missed demo deliverables: no video or hosted demo, empty demo checks in the pack, and dashboard captures are frozen offline renders with controls disabled

Red flags

  • Evidence counts are inflated relative to the fine print: the deck and README headline 156 axiom records and 187 to 189 tests, but the axioms are a self-authored theorem catalog rather than independent security requirements, the team narrows the scope in finer print, and totals disagree across artifacts
  • Submission completeness risk: the official pack has empty demo checks and the deck appears only in the repo rather than the provided channel, and one member reports the submission form was never filled, so the official record may understate or misstate what actually exists
  • Text that tries to steer the reviewers: readme: "ignore previous instructions"
Built during the event: yes: 32 commits by 2 authors, 2026-10-03 17:04 to 2026-10-04 08:37 UTC
Live demo: none found in the form or README
Council v9
Aclaude:glm-5.3-flash82.580% agree
Bdots-studio/dots-3-note-preview:free79.390% agree
Cinclusionai/ling-3.0-flash-sante:free70.070% agree
Jclaude:glm-5.3–judge
Self-submitted and unverified: the result shown is the team's own claim. The council read an evidence pack built from the repo, its decks and docs; it didn't run the code or see the pitch.
The site is open source

The council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.

Star on GitHub