Operations
Cost on the request path
| Layer | What happens per request | Rough cost |
|---|---|---|
| IP reputation | A binary search over merged, sorted ranges | Microseconds; independent of the number of feeds |
| WAF, no body | Phase 1 rules over the request line, headers and cookies | Proportional to the ruleset |
| WAF, with body | Phases 1 and 2, the body materialised and handed to the engine | Proportional to the ruleset and the body size |
| WAF, response | Phases 3 and 4, the response buffered | Proportional to the response size |
A config's rules are compiled when they change — the config, a ruleset it references, its CRS options — and the result is shared by every request, so none of the costs above includes parsing or compiling.
The WAF is instrumented through Otoroshi's metrics under
cloud_apim.plugins.waf.evaluation.request and .response, so its real cost in your deployment is
measurable rather than guessed.
The single biggest lever is body inspection. If latency matters more than body coverage on a route,
set inspect_input_body: false and keep phase 1 protection — it still catches everything that lives
in the URI, headers and cookies, which is most scanning traffic.
How it fails
The extension is built so that an external dependency failing degrades protection visibly rather than breaking traffic.
| Failure | What happens |
|---|---|
| A feed provider is down | The previous snapshot keeps serving; the error appears on the feed page and in /_status |
| A feed starts serving garbage | The refresh is rejected rather than accepted; the last good snapshot stays |
| A feed refresh brings bad-but-valid content | One generation is kept — Roll back restores it |
| The CrowdSec API is unreachable | Held decisions keep being enforced; the error appears on the bouncer page |
| A CrowdSec key is revoked | Same — the mirror is never emptied by a failure |
| A geolocation database cannot be downloaded | The last good file keeps serving; the error appears on its page, and the download is retried after five minutes |
| A DNS blocklist is slow, down, or refuses the resolver | Its answers count as not listed; the zone is reported unhealthy in /_status and in the logs |
| The extension is disabled with routes still referencing it | Their requests are refused with a 503, or go through uninspected with waf.fail-open — see below |
| A route's WAF plugin names a config that does not exist | Same |
| A WAF config's rules do not compile | Same, and waf config '…' does not compile …: rule 2: … is logged each time the config changes. The API saves rules without compiling them, so this can happen |
The pattern throughout: fail open, and make the degradation visible. A WAF that silently stops protecting is worse than one that visibly stops working.
The exception is the WAF itself. When the rules a route asks for cannot run at all, the choice is
between refusing its traffic and letting it through uninspected, and that is a policy rather than
a technical default, so it is a setting: waf.fail-open (CLOUD_APIM_EXTENSIONS_WAF_FAIL_OPEN).
It is off, so such a request gets an empty 503. Either way the logs say which config and why, once
per change rather than once per request. Two cases never refuse anything: a config switched off
(enabled: false) is a decision rather than a failure, and a config in monitoring mode
(block: false) would not have refused the request had it run.
What to watch
Poll /extensions/cloud-apim/extensions/waf/reputation/_status and alert on:
| Condition | Means |
|---|---|
feeds[].snapshot.error != null | That feed is stale on this node |
feeds[].snapshot.fetched_at older than ~3× its refresh interval | Refreshes are not completing |
feeds[].snapshot.rejected suddenly non-zero | The provider probably changed format |
crowdsec[].store.last_error != null | The Local API is unreachable or the key was rejected |
crowdsec[].store.initialized == false | The initial sync has never completed |
crowdsec[].pending_push growing | Alerts are queueing — the push credentials are probably wrong |
geo[].snapshot.error != null | That geolocation database is stale on this node, or was never loaded |
rbl.zones.<zone>.healthy == false | That blocklist is not answering usefully: every lookup counts as not listed |
From the analytics side, the useful alert is a burst: a sharp rise in CloudApimWafTrailEvent
with a non-null block, grouped by source address, over a short window.
Clustered deployments
Reputation feeds are per node by design. Every node keeps its own feed index and its own CrowdSec cursor: there is no shared state and nothing to coordinate, at the cost of each node fetching each feed independently — so outbound requests scale with the number of nodes. Nodes stagger their first refresh by a random 5–25 seconds so a rolling restart does not hit every provider simultaneously. Snapshot rollback is per node and in memory; it survives a refresh, not a restart.
Everything that has to be shared is not. Bans, the ledger, fail2ban counters, challenges, tuning candidates and learning windows all go through the shared state, and on a leader/worker cluster that requires a dedicated redis — a worker never reaches your storage backend at all. This is the single most common way to end up with a page that looks reassuringly empty.
→ What a complete deployment needs
Memory
| Structure | Sizing |
|---|---|
| Feed index | Two long arrays per feed for IPv4, plus BigInt arrays for IPv6. A million merged ranges is on the order of tens of megabytes |
| Compiled rulesets | Cached by ruleset hash, not per route. integration.max-cache-items caps it, default 1000 |
| CrowdSec mirror | One entry per held decision |
max_entries on a feed is the guard that matters: it caps what a provider can make you allocate if
its list suddenly grows.
Upgrades
The extension id has been stable since the first release, so datastore keys, the API group and route paths do not move across upgrades. Entities gain new fields with defaults, so older entities stay valid. Rolling upgrades are safe — each node rebuilds its own in-memory indexes on start.