AI Control LayerSelf-submittedNot a finalist (self-declared)

Quard

Quard · oynozan/quard

Council score

Median of 3 models, weighted by the task's official criteria
81.0 / 100

A genuinely deep and well-tested agent control layer whose evidence matches its claims, losing points for the local-model support the task expects but does not get, an unevidenced policy hot reload and a pack with no demo checks.

Criteria · line = median, dots = each member
Robustness/guardrails30%
9.0

The hybrid is real: deterministic guards for secret stripping, keyed-hash PII redaction, IBAN checksums and invisible characters, value tracing that keeps untrusted labels across agent handoffs, and a genuine semantic model path in the Jev detector with AI fallback and a review queue. C's 7 is the outlier and names no defect the others missed, since C's own text confirms the controls are real, so the median stays high with a point off because signature coverage skips the code-execution, deserialization and model-repo classes the task names.

Members disagree here: scores range by 2.5 points.
Architecture & performance20%
8.0

The SDK genuinely sits in the path, wrapping model and tool calls in-process before they run, on a clean split of webhook ingest, a WebSocket control service, a worker and a Postgres-backed dashboard. It holds at 8 because hot reload amounts to a live control link with no config watcher or editable rules view in evidence, and no latency numbers beyond per-check timing prints are shown.

Members disagree here: scores range by 3.0 points.
Security reporting20%
7.0

Every model call, guard decision and approval is stored as a queryable event with rule and reason, incidents trace the path from entry point to damage with replay proof, and the dashboard runs on real data. On the disputed audit-log deliverable the evidence supports B and C, who independently find the logs exportable through the event store and search, over A's absence claim, but the median holds at 7 because that export path is indirect and the pack contains no demo checks to verify reporting at runtime.

Members disagree here: scores range by 2.5 points.
Test suite15%
9.0

The measured suite is huge and real: 4,180 cases in 826 files, about one test file per source file, runnable with pnpm test, including negative cases for injection signatures and blocked payments. It lands just below the top because the claimed full coverage is unverified, the claimed count of 4,500 runs slightly ahead of the measured 4,180, and no individual negative case was read in the pack.

Implementability & scalability15%
7.0

Deployment is a single Docker Compose with health checks and migrations, an npm-published SDK, JSON policy files and pluggable guards with x402 payment caps. The median settles at 7 because the task's local-model expectation is missed: v1 is OpenAI-only, and the base-URL hook B credits is indirect and, as B itself concedes, may be limited by the OpenAI Responses API dependency.

Members disagree here: scores range by 2.0 points.
Source lines109,194
Tests4180 cases
Claims built9.0 / 10
Task fitYes

Strengths

  • Origin and value tracing keeps untrusted labels on values across agents, tools and shared memory, a sophisticated taint model that powers precise root-cause analysis.
  • Deterministic controls are real code, including secret stripping, keyed-hash PII redaction with masks, IBAN checksum validation and invisible-character scanning, backed by thorough budget governance and fleet-wide quarantine of repeated dangerous values.
  • Semantic detection runs a real model path with chunk labeling, AI fallback and a human review queue, and incidents can be replayed with and without suspect content to prove causation.
  • A measured 4,180-case test suite across 826 files and a dashboard wired to an extensive Postgres query layer rather than mock data.

Weaknesses

  • No local-model or Ollama support although the task expects it; v1 is OpenAI-only and the base-URL hook is at best indirect.
  • Policy hot reload is claimed but not evidenced: rules live in code, the dashboard rules tab is read-only and no config watcher appears, only a live control link.
  • The pack contains no demo checks, so runtime behavior, latency, hot reload and the npm publication claim rest on code and form alone.
  • Signature coverage targets injected instructions and bad domains, not the code-execution, unsafe deserialization and model-repo supply-chain classes the task names.
Built during the event: yes: 495 commits by 2 authors, 2026-10-03 10:31 to 2026-10-04 08:54 UTC
Live demo: none found in the form or README
Council v9
Aclaude:glm-5.3-flash81.090% agree
Bdots-studio/dots-3-note-preview:free91.870% agree
Cinclusionai/ling-3.0-flash-sante:free68.060% agree
Jclaude:glm-5.3–judge
Self-submitted and unverified: the result shown is the team's own claim. The council read an evidence pack built from the repo, its decks and docs; it didn't run the code or see the pitch.
The site is open source

The council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.

Star on GitHub