Skip to main content

Bans and memory

Bans​

A ban attaches to an identity — an address, an apikey, a user — and is enforced by the threat gate before any inspection runs.

Threat Protection → Bans & incidents.

Two properties drive the whole design:

Reads never touch the datastore. The gate answers from a node-local map. A ban exists to be cheaper than the inspection it replaces, so a lookup that cost a network round trip would defeat the purpose. The map holds every live ban, so the cost is the same whether you hold ten or ten thousand.

Writes are instant on the node that decided. A node that bans a caller enforces it on that caller's very next request. Other nodes pick it up on their next refresh — every 10 seconds by default.

That bounded staleness is the honest trade: Otoroshi's storage abstraction has no pub/sub, so a few seconds of propagation lag buys zero I/O per request.

Making bans actually shared​

security {
// guarantees distribution regardless of the otoroshi storage backend
redis-uri = "redis://localhost:6379"
}
Without this, "distributed" depends on your storage

Bans fall back to the Otoroshi storage backend, which is only shared between nodes if that backend is. With file or inmemory storage — including the quickstart — each node has its own private ban list. The Bans & incidents page says which of the two you are running, and warns when it is the fallback.

The dedicated connection goes through the same stateful client manager Otoroshi's distributed rate limiter uses.

On what evidence​

Open a ban on the console and it shows what the caller actually did: the timeline of the incident they were part of, copied into the ban at the moment it was issued.

A copy rather than a lookup, because the two have different lifetimes. An incident is evicted after the correlation window — thirty minutes by default — and a ban routinely outlives it. Without the copy, a ban issued an hour ago has nothing behind it by the time anyone asks why, which is precisely when the question gets asked.

Bans issued by the scored response engine carry the signals as well: which detector contributed what, and how confident it was. A ban with neither — a hand-typed one against a caller nobody has an incident for — says so rather than showing an empty panel.

A ban opened, showing who issued it and what the caller did

Banning from an incident row copies that incident's timeline into the ban, so it still answers "why" an hour later when the incident itself has been evicted.

Managing them​

GET /_bansEvery live ban with its evidence
POST /_ban{"ref": "ip:1.2.3.4", "duration_seconds": 3600, "reason": "..."}
POST /_extend{"ref": "ip:1.2.3.4", "duration_seconds": 86400}
POST /_unban{"ref": "ip:1.2.3.4"}, or {"all": true}

All under /extensions/cloud-apim/extensions/waf/security. The console does the same from the row.

Extending measures from the current end, not from now. Adding an hour to a ban with an hour left leaves two. And there is nothing to extend once a ban has lapsed: the call answers that it lapsed rather than quietly minting a fresh one nobody asked for.

Every one of these actions is recorded as a CloudApimWafSecurityAudit with the operator's email on it. They are the decisions in this product that most need a name against them — taken under pressure, changing what the gateway does to real callers, and rarely asked about by the person who took them.

The allowlist​

Some callers must never be banned. The partner whose integration suite trips every rate limit, the monitoring probe, the office address.

Threat Protection → Bans & incidents → Allowlist, or one click from the row of the caller in front of you.

curl -X POST .../security/_allow \
-d '{"ref":"apikey:partner-ci","reason":"nightly integration suite","unban":true}'
curl -X POST .../security/_disallow -d '{"ref":"apikey:partner-ci"}'

The entry is permanent unless you pass duration_seconds. A bounded one is useful — leave them alone while we tune — but it is not the default, because the case the primitive exists for is the permanent one.

The allowlist panel with one permanent entry

The reason is required, not optional. It is the only thing that tells whoever finds this entry in six months whether it can go.

Two things it does, and one it does not​

It refuses the ban rather than removing it afterwards. The check lives in the ban store, which is the single point every ban goes through, from every module — fail2ban, the ledger, the response engine, the honeypot, the admin API. "This caller is never banned" is only worth promising if it holds on every path, including the ones added later.

It lifts what is already held. Allowlisting through the console also unbans the caller and forgets their ledger total. Either half on its own looks like the feature is broken: allowlisting alone leaves them banned until the current ban lapses, and unbanning alone leaves the ledger free to re-ban them inside the window.

It does not switch off inspection. The WAF still inspects, the response engine still scores, and a tier that denies still denies. The allowlist governs bans, nothing else.

To exempt a caller from the fabric entirely

Use a threat policy's exemptions instead. Those are address ranges, configured ahead of time, and they skip the gate outright. The allowlist is the operational counterpart: any kind of identity — apikey, user, fingerprint, address — added in one click during an incident, applying to every route at once.

Entries live in the shared state and reach the other nodes on their next refresh, exactly like bans. And exactly like bans, a refresh that fails keeps the stale copy: dropping the allowlist would make the fabric more aggressive, which is not the direction to fail in.

The ledger​

The ledger is the gateway's memory of a caller across requests, and it is the piece a WAF structurally cannot have on its own.

A single request is judged on its own content. A caller who trips one rule a minute for an hour looks innocent every single time. The ledger accumulates those judgements per identity over a sliding window, and once the total crosses a threshold it promotes the caller to a ban.

security.ledger {
enabled = true
window-seconds = 3600 // the window slides: misbehaving keeps you in it
ban-threshold = 100
ban-duration-seconds = 3600
}

It is fed asynchronously, from the analytics event stream, and read never on the request path — the promotion writes a ban, and ban lookups are free. A per-request round trip to read a running total would undo the reason for having a cheap layer at all.

When a ban is issued the total is consumed, so the next contribution does not re-ban instantly. The same happens when the ban is refused because the caller is allowlisted — otherwise every subsequent contribution would ask again, and be refused again, forever.

What charges it​

SourceDefaultWhy
Enforced decisions from the threat response pluginonIt knows exactly what it enforced and why
WAF blocks, via the event relayoffsecurity.ledger.feed-waf-blocks
WAF matches let through in monitoring modeoffsecurity.ledger.feed-waf-monitored
Why the WAF feeds are off by default

On a route that runs the threat response plugin, the plugin already charges the ledger for what it enforced. Charging again from the event would count the same detection twice.

Turn feed-waf-blocks on when you run the WAF without the response plugin and still want repeat offenders to accumulate towards a ban.

Turn feed-waf-monitored on only if you mean it: a WAF you are still tuning would otherwise be quietly getting callers banned.

Inspecting it​

curl -X POST .../security/_ledger -d '{}' # top offenders
curl -X POST .../security/_ledger -d '{"ref":"ip:1.2.3.4"}' # one caller
curl -X POST .../security/_ledger -d '{"ref":"ip:1.2.3.4","forget":true}'