AI Control LayerSelf-submittedNot a finalist (self-declared)

good gateway, bad gateway

wiz.dev · blatalia/wiz-dev-hackyeah-2026

Council score

Median of 3 models, weighted by the task's official criteria
65.5 / 100

A genuinely built and well-tested control layer with real policy hot reload and a strong audit dashboard, held back by unauthenticated user identity, thin injection coverage and the missing local-model support the task expects.

Criteria · line = median, dots = each member
Robustness/guardrails30%
6.0

Deterministic controls are real: PII anonymization, SQL-injection scanning, per-user tool access and budgets, and the LLM input and output judge exists in code, supporting A and C over B's claim that it is deck-only. Injection and secret detection are admitted future work, only two injection test strings exist with zero recorded hits, and user identity is client-asserted.

Architecture & performance20%
7.0

The FastAPI gateway is genuinely in the request path, with config hot reload via a background poller and per-stage latency, token and cost telemetry. No performance benchmarks are provided, only the OpenAI SDK is supported, and client-asserted identity weakens the enforcement chain.

Security reporting20%
7.0

Events are richly structured with decision, reason codes, classification scores, security scan findings, latency, cost and config version, shown in a live dashboard with filters and stats, which justifies a strong but not top score. Export is limited to copying a single event's JSON, as A and C observed against B's broader claim, and a seed SQL file for gateway events means dashboard data may be partly seeded.

Members disagree here: scores range by 2.0 points.
Test suite15%
7.0

65 measured cases across 8 files cover guardrails, tool limits, token usage, user attribution and PII, plus a judge-facing validation suite with a documented run script. Negative-case depth is not fully verifiable from the samples and injection coverage is thin.

Implementability & scalability15%
6.0

Dockerized modular services with per-user config overrides and clean default-and-override semantics make it reproducible. The stack is bound to AWS DynamoDB, config is split between YAML and DynamoDB, and there is no Ollama or local-model support, which the task details expect.

Members disagree here: scores range by 2.0 points.
Source lines6,131
Tests65 cases
Claims built7.0 / 10
Task fitYes

Strengths

  • Central policy is wired end to end: the dashboard patches a config store that the gateway hot-reloads without restart, including per-user tool toggles.
  • A real hybrid control set: deterministic PII and SQL-injection handling plus an LLM input and output judge with an enforced JSON output format.
  • A strong audit event schema with decision, reason codes, classification scores, scan findings, latency by stage, cost and config version, rendered in a polished live dashboard.
  • 65 automated tests across 8 files plus a separate validation suite with a documented run script for judges.

Weaknesses

  • Missed task requirement: no Ollama or local-model support anywhere in the code, OpenAI SDK only.
  • Identity is client-asserted: /chat trusts a user_email from the request body with no authentication, so tool permissions and audit attribution can be bypassed, and demo credentials are hardcoded in the README.
  • Prompt-injection and secret or token detection are only future extension, with two static test strings, no recorded hits and no externally fed attack signatures.
  • Audit export is limited to copying a single event's JSON with no bulk format, and there is no monetary budget cutoff despite cost tracking existing.

Red flags

  • The demo link is a Google Drive redirect (HTTP 302) with no runnable service and zero screenshots, so the live system could not be verified by the jury.
  • A seed SQL file for gateway events combined with an unverifiable gateway-to-event-store write path means dashboard activity may be partly demo-seeded data.
Built during the event: partly: 1 commit before the event; 117 commits by 5 authors, 2026-10-03 08:48 to 2026-10-04 08:56 UTC
Live demo: https://drive.google.com/drive/u/0/mobile/folders/1qWCfIIJXCKa4hqww507PJw390qKYKjDH?usp=sharing (HTTP 302)
Council v9
Aclaude:glm-5.3-flash65.890% agree
Bdots-studio/dots-3-note-preview:free70.060% agree
Cinclusionai/ling-3.0-flash-sante:free60.580% agree
Jclaude:glm-5.3–judge
Self-submitted and unverified: the result shown is the team's own claim. The council read an evidence pack built from the repo, its decks and docs; it didn't run the code or see the pitch.
The site is open source

The council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.

Star on GitHub