DefenceSelf-submittedNot a finalist (self-declared)

OpenWAF

OpenWAF-Team · AleksanderGPL/OpenWAF

Council score

Median of 3 models, weighted by the task's official criteria
73.0 / 100

A genuinely implemented WAF with real anomaly detection and a well-guarded AI investigation agent that meets the Defence brief in code, but the absent degraded-scenario demo and complete lack of visual deliverables leave its usability and design unproven.

Criteria · line = median, dots = each member
Idea & Innovation30%
7.0

The WAF-plus-LLM-triage combination is known, but the execution fills a real gap concretely: five named detectors with thresholds, warmup baselines and cooldowns, and an investigation agent restricted to read-only tools that must cite request IDs and mark assessments inconclusive when evidence is empty. C's 'incremental dashboard' reading understates how specific the failure modes are.

Members disagree here: scores range by 2.0 points.
Relation to Category20%
9.0

This is unambiguously security and resilience work: CRS prevention with per-service policies, earlier detection on a 15-second cycle including an upstream-error detector for failing services, and consequence reduction through investigations and evidence snapshots that survive log pruning. The evidence supports A and B over C's 'generic monitoring tool' reading.

Members disagree here: scores range by 3.0 points.
Practical Applicability / Usability20%
7.0

It deploys as a single Go binary with embedded frontend and SQLite, keeps recording detections with no AI key (chat returns 503), and tolerates GeoIP download failure, with acknowledge/resolve/dismiss workflows and SSE progress. C's impracticality claim is overstated by the evidence, though tuning multi-service policies, paranoia levels and anomaly settings is genuinely non-trivial.

Members disagree here: scores range by 2.0 points.
Design (visual/UI)20%
6.0

The code promises a coherent Nuxt 4 and Tailwind interface with ECharts, EN/PL i18n and severity badges, but zero screenshots, demo captures or deck exist, so visual quality is judged from framework choices alone.

Completeness & Implementation Value10%
8.0

16,277 lines by 3 authors with all 37 commits inside the event window, 27 test files with 66 cases including a gated live AI test, real Coraza rule compilation, SQLite-backed detection and investigation plumbing with crash recovery, and a 6,000-log seed generator. A working prototype, not a mock.

Source lines16,277
Tests66 cases
Claims built8.0 / 10
Task fitYes

Strengths

  • Security-critical logic is genuinely implemented: Coraza 3.8.1 with OWASP CRS 4.25.0 compiled from embedded rules, hot-reloading per-service policies, and custom rule validation with regex and CIDR checks.
  • The detection-to-investigation pipeline is real and resilience-aware: 15-second evaluation over aggregated minute buckets, queue and claim execution with five-minute leases and crash recovery, durable SSE event IDs, and evidence snapshots that survive log pruning.
  • Strong LLM guardrails: read-only telemetry tools, bounded tool budgets, mandatory request-ID citation, no blocking actions exposed to the model, and inconclusive verdicts when evidence is empty.
  • Excellent hygiene for a 24-hour build: argon2id auth with rate limiting, trusted-proxy client-IP handling, graceful shutdown, degraded operation without an AI key, and GeoIP that keeps working when its download fails.

Weaknesses

  • The task-required demo of a degraded scenario (services down, incomplete information) is missing: zero demo checks were recorded, and the seed generator produces attacks and benign traffic but not upstream errors or a service-down state, so the scenario exists only in code and docs.
  • No deck, no screenshots and no captured demo pages, so the UI, design quality and any end-to-end incident flow cannot be verified from the pack.
  • The submission form was left blank, leaving the README as the only self-description.
  • AI assessment quality is unverifiable: it needs an external API key and no executed investigation was shown, so triage value rests on tests the jury did not see run.
Built during the event: yes: 37 commits by 3 authors, 2026-10-03 10:35 to 2026-10-04 07:25 UTC
Live demo: none found in the form or README
Council v9
Aclaude:glm-5.3-flash72.090% agree
Bdots-studio/dots-3-note-preview:free74.0100% agree
Cinclusionai/ling-3.0-flash-sante:free57.050% agree
Jclaude:glm-5.3–judge
Self-submitted and unverified: the result shown is the team's own claim. The council read an evidence pack built from the repo, its decks and docs; it didn't run the code or see the pitch.
The site is open source

The council, the evidence pack, the prompts and the queue are all on GitHub. If a review helped you, a star helps other teams find it.

Star on GitHub