AI Control LayerSelf-submittedNot a finalist (self-declared)

Portcullis

Portcullis · rozumalex/controllayer

Council score

Median of 3 models, weighted by the task's official criteria
81.3 / 100

A genuinely implemented, in-path control layer with unusually deep hybrid guards, hot-reloading policies and strong automated verification, held back by an optional semantic guard, an export endpoint missing from the source, and headline metrics that are self-measured rather than independently verified.

Criteria · line = median, dots = each member
Robustness/guardrails30%
8.5

The guard pipeline is real code in the request path: deterministic patterns, PII and secret redaction, a genuine LLM semantic classifier, and behavioral guards for budget, loop, lockout and data-flow egress, all exercised by tests and a 1,057-case attack corpus. It stays below 9 because the semantic half is disabled without an OpenAI key and the headline block rate is self-measured.

Members disagree here: scores range by 2.0 points.
Architecture & performance20%
8.0

The gateway is genuinely in the path for both chat and MCP traffic, reads the role policy from the database on every request so hot reload works without restart, and scopes data per organization. No aggregate latency or throughput benchmark appears anywhere to back the low-overhead claim.

Security reporting20%
7.5

Every verdict fans out to Postgres and logs with trace IDs, and the dashboard reads those traces live. The CSV/JSON export claimed in the form has no endpoint in the provided source, so that requirement is only partially evidenced; here the evidence supports the two members who read the samples and found no endpoint over the member who accepted the claim, which is why the median sits below A's 9.

Members disagree here: scores range by 2.0 points.
Test suite15%
8.5

415 pytest cases across 43 files plus a separate mutation harness that expands 58 seeds into 1,057 cases with explicit block and allow expectations and false-alarm tracking. All three members agree this is a strength and the spread is small.

Implementability & scalability15%
8.0

One-command Docker bring-up that runs without a paid key, OpenAI and Ollama behind one interface, YAML policies with strictness and budgets, and OIDC and SCIM identity with per-organization isolation. No horizontal-scaling or production-hardening story is shown.

Source lines32,183
Tests415 cases
Claims built8.5 / 10
Task fitYes

Strengths

  • A genuine in-path gateway for both agent-to-model (OpenAI-compatible, with streaming) and agent-to-MCP traffic, with role-scoped tool, model and clearance access.
  • Deep hybrid guard pipeline: deterministic patterns plus a real semantic model call plus session-aware behavioral guards such as data-flow egress, loop and lockout.
  • Policy hot reload by construction: the role policy is read from the database on every request, so admin edits apply on the next call.
  • Large runnable verification: 415 automated tests plus a 1,057-case mutation-based attack corpus with benign controls and false-alarm counting.

Weaknesses

  • The semantic guard requires an OpenAI API key or a configured endpoint, so out of the box the hybrid pipeline degrades to deterministic-only.
  • The audit-log export (CSV/JSON with filters) is claimed in the form but the export endpoint is missing from the provided source, so that task requirement is only partially met.
  • No aggregate overhead or throughput numbers exist, and the live demo check was skipped, so performance and the headline block and false-alarm rates are self-measured and unverified.
  • The repo contains zero screenshots, so the dashboard has no visual evidence beyond members reading the frontend code, and the architecture diagram lives only in the deck.

Red flags

  • The deck's 98.9 percent block rate is self-measured on the team's own corpus, and the member who read the attack harness reports the run excludes the database-dependent clearance and budget guards, so the headline number overstates coverage of the full pipeline.
Built during the event: yes: 77 commits by 2 authors, 2026-10-03 10:52 to 2026-10-04 08:54 UTC
Live demo: https://controllayer.net (skipped (not a public https URL))
Council v9
Aclaude:glm-5.3-flash88.380% agree
Bdots-studio/dots-3-note-preview:free81.3100% agree
Cinclusionai/ling-3.0-flash-sante:free71.560% agree
Jclaude:glm-5.3–judge
Self-submitted and unverified: the result shown is the team's own claim. The council read an evidence pack built from the repo, its decks and docs; it didn't run the code or see the pitch.
The site is open source

The council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.

Star on GitHub