Skip to main content

Operations

Cost on the request path​

LayerWhat happens per requestRough cost
IP reputationA binary search over merged, sorted rangesMicroseconds; independent of the number of feeds
WAF, no bodyPhase 1 rules over the request line, headers and cookiesProportional to the ruleset
WAF, with bodyPhases 1 and 2, the body materialised and handed to the engineProportional to the ruleset and the body size
WAF, responsePhases 3 and 4, the response bufferedProportional to the response size

A config's rules are compiled when they change — the config, a ruleset it references, its CRS options — and the result is shared by every request, so none of the costs above includes parsing or compiling.

The WAF is instrumented through Otoroshi's metrics under cloud_apim.plugins.waf.evaluation.request and .response, so its real cost in your deployment is measurable rather than guessed.

The single biggest lever is body inspection. If latency matters more than body coverage on a route, set inspect_input_body: false and keep phase 1 protection — it still catches everything that lives in the URI, headers and cookies, which is most scanning traffic.

How it fails​

The extension is built so that an external dependency failing degrades protection visibly rather than breaking traffic.

FailureWhat happens
A feed provider is downThe previous snapshot keeps serving; the error appears on the feed page and in /_status
A feed starts serving garbageThe refresh is rejected rather than accepted; the last good snapshot stays
A feed refresh brings bad-but-valid contentOne generation is kept — Roll back restores it
The CrowdSec API is unreachableHeld decisions keep being enforced; the error appears on the bouncer page
A CrowdSec key is revokedSame — the mirror is never emptied by a failure
A geolocation database cannot be downloadedThe last good file keeps serving; the error appears on its page, and the download is retried after five minutes
A DNS blocklist is slow, down, or refuses the resolverIts answers count as not listed; the zone is reported unhealthy in /_status and in the logs
The extension is disabled with routes still referencing itTheir requests are refused with a 503, or go through uninspected with waf.fail-open — see below
A route's WAF plugin names a config that does not existSame
A WAF config's rules do not compileSame, and waf config '…' does not compile …: rule 2: … is logged each time the config changes. The API saves rules without compiling them, so this can happen

The pattern throughout: fail open, and make the degradation visible. A WAF that silently stops protecting is worse than one that visibly stops working.

The exception is the WAF itself. When the rules a route asks for cannot run at all, the choice is between refusing its traffic and letting it through uninspected, and that is a policy rather than a technical default, so it is a setting: waf.fail-open (CLOUD_APIM_EXTENSIONS_WAF_FAIL_OPEN). It is off, so such a request gets an empty 503. Either way the logs say which config and why, once per change rather than once per request. Two cases never refuse anything: a config switched off (enabled: false) is a decision rather than a failure, and a config in monitoring mode (block: false) would not have refused the request had it run.

What to watch​

Poll /extensions/cloud-apim/extensions/waf/reputation/_status and alert on:

ConditionMeans
feeds[].snapshot.error != nullThat feed is stale on this node
feeds[].snapshot.fetched_at older than ~3× its refresh intervalRefreshes are not completing
feeds[].snapshot.rejected suddenly non-zeroThe provider probably changed format
crowdsec[].store.last_error != nullThe Local API is unreachable or the key was rejected
crowdsec[].store.initialized == falseThe initial sync has never completed
crowdsec[].pending_push growingAlerts are queueing — the push credentials are probably wrong
geo[].snapshot.error != nullThat geolocation database is stale on this node, or was never loaded
rbl.zones.<zone>.healthy == falseThat blocklist is not answering usefully: every lookup counts as not listed

From the analytics side, the useful alert is a burst: a sharp rise in CloudApimWafTrailEvent with a non-null block, grouped by source address, over a short window.

Clustered deployments​

Reputation feeds are per node by design. Every node keeps its own feed index and its own CrowdSec cursor: there is no shared state and nothing to coordinate, at the cost of each node fetching each feed independently — so outbound requests scale with the number of nodes. Nodes stagger their first refresh by a random 5–25 seconds so a rolling restart does not hit every provider simultaneously. Snapshot rollback is per node and in memory; it survives a refresh, not a restart.

Everything that has to be shared is not. Bans, the ledger, fail2ban counters, challenges, tuning candidates and learning windows all go through the shared state, and on a leader/worker cluster that requires a dedicated redis — a worker never reaches your storage backend at all. This is the single most common way to end up with a page that looks reassuringly empty.

→ What a complete deployment needs

Memory​

StructureSizing
Feed indexTwo long arrays per feed for IPv4, plus BigInt arrays for IPv6. A million merged ranges is on the order of tens of megabytes
Compiled rulesetsCached by ruleset hash, not per route. integration.max-cache-items caps it, default 1000
CrowdSec mirrorOne entry per held decision

max_entries on a feed is the guard that matters: it caps what a provider can make you allocate if its list suddenly grows.

Upgrades​

The extension id has been stable since the first release, so datastore keys, the API group and route paths do not move across upgrades. Entities gain new fields with defaults, so older entities stay valid. Rolling upgrades are safe — each node rebuilds its own in-memory indexes on start.