AI Control LayerHackYeah 2026 finalist

Clearance

Visdomers · damianlech/HackYeah26

Council score

Median of 3 models, weighted by the task's official criteria
53.5 / 100

A genuinely engineered control layer with real gateway and proxy enforcement, hot-reloaded policy and stack-booting tests, undermined by a simulated semantic half, a mockup dashboard, README overclaims and a missing required deck.

Criteria · line = median, dots = each member
Robustness/guardrails30%
5.0

The deterministic half is real and deep: auth, model allowlist, checksum-validated PII redaction, a signature feed with self-test vectors, budgets and fail-closed behaviour under hot reload. The semantic half is a keyword-and-weight simulation and the authority-over-detection guarantees exist only in design docs, so the hybrid requirement is half met and the median of 5 is fair.

Members disagree here: scores range by 2.0 points.
Architecture & performance20%
7.0

Two real in-path enforcement points exist in code, including a 3-layer proxy that terminates TLS, applies policy and re-encrypts with a virtual-to-real key swap, backed by committed latency benchmark results. Member C's claim that the performance numbers are placeholders in a deck is not supported, since no deck exists and two members independently saw measured results, so the higher median of 7 stands.

Members disagree here: scores range by 2.5 points.
Security reporting20%
4.0

Audit JSONL with trace IDs and a generator producing separate management and security tables from real log data are genuine. The dashboard deliverable is a static mockup not wired to the gateway, there is no live metrics endpoint and no tamper-evident log chaining, and the required deck is missing, so 4 is deserved.

Members disagree here: scores range by 2.0 points.
Test suite15%
6.0

The 46 measured cases boot the real gateway, guard service and mock LLM as processes and include genuine negative cases such as prompt injection, budget exhaustion, guard outage and rejected policy edits. The README overstates this as 76 tests plus 120 browser checks, which the measured count does not support.

Members disagree here: scores range by 2.0 points.
Implementability & scalability15%
5.0

One-command runs, pinned dependencies, three strictness profiles and prefix routing that can point at local models are real. Budgets are in-memory per process with check-then-charge, there are no Docker or Kubernetes artifacts and the two prototypes were never unified; Member C's claim that the build has not started is contradicted by roughly 11k lines of working prototype code that all members reviewed.

Members disagree here: scores range by 2.5 points.
Source lines11,382
Tests46 cases
Claims built5.5 / 10
Task fitYes

Strengths

  • Real enforcement code in the traffic path: a working OpenAI-compatible gateway plus a 3-layer TLS-intercepting proxy with certificate verification and virtual-to-real key swap.
  • Policy hot reload with last-known-good rejection of broken edits, plus a signature feed carrying per-rule positive and negative self-test vectors.
  • PII redaction with checksum validators for PESEL, IBAN and Luhn rather than plain regex matching.
  • Tests boot the actual stack as processes and cover real negative cases, with committed latency benchmarks comparing deterministic and guard-path enforcement.

Weaknesses

  • The required 10-slide deck was not provided and the submission form is essentially blank, leaving the repo to carry the review alone.
  • The semantic controls are keyword-and-weight simulators rather than model calls, so the hybrid deterministic plus semantic requirement is only half implemented.
  • The dashboard is a static mockup with seeded data, not a UI wired to the running gateway, and no live metrics endpoint exists.
  • Scalability is docs-only: in-memory budgets, no Docker or Kubernetes artifacts, and two prototypes that were never unified into one product.

Red flags

  • The README claims 76 automated tests plus 120 browser checks while 46 test cases were measured, and claims Visdom integration with no Visdom code in the repo.
  • Keyword-matching simulators are labelled as AI-based guardrails.
  • Mockup dashboard pages are presented as screenshots of the product.
Built during the event: yes: 37 commits by 2 authors, 2026-10-03 11:09 to 2026-10-04 08:15 UTC
Live demo: none found in the form or README
Council v9
Aclaude:glm-5.3-flash69.070% agree
Bdots-studio/dots-3-note-preview:free51.390% agree
Cinclusionai/ling-3.0-flash-sante:free49.560% agree
Jclaude:glm-5.3–judge
A HackYeah 2026 finalist, queued automatically for a council review; the council wasn't told how it placed. The council read an evidence pack built from the repo, its decks and docs; it didn't run the code or see the pitch.
The site is open source

The council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.

Star on GitHub