Guardrail webhooks: design and tradeoffs

A guardrail policy rule delegates the allow/deny/mask decision for a matched request (or buffered response) to an operator-run HTTP endpoint. This closes the “semantic guardrails” gap — PII classifiers, prompt-injection detectors, custom models — without building classifiers into the gateway. This page explains the design choices and what they cost you.


The engine stays pure; the proxy makes the call

internal/policy performs no I/O. When a guardrail rule matches, Engine.Evaluate returns a pure Decision{Action: "guardrail", Guardrail: name} — nothing has been called, mutated, or timed out yet. The proxy layer resolves the name to a client and makes the synchronous HTTP callout, because the proxy already owns outbound clients, deadlines, and audit. Same layering as the rest of mcpgw: the engine decides, the proxy executes.

The payoff is that the engine remains reusable everywhere decisions are needed without side effects. mcpgw policy test and mcpgw policy replay share the live engine exactly; a guardrail decision in those tools just means “the rule matched and named this endpoint,” which is why policy test --guardrail-stub can simulate the outcome without any network. The registry of endpoints hot-swaps atomically on SIGHUP via the same pointer-swap pattern as the engine, so a reload never leaves a rule pointing at a half-built client.


Fail-closed by default

fail_open defaults to false. A guardrail that cannot be reached is a guardrail that cannot do its job, so an outage denies matched requests (403, -32001 policy_denied, audit error_kind: guardrail_unreachable) rather than silently disabling a configured security control.

fail_open: true exists for guardrails that are advisory rather than load-bearing — an observability classifier, say, where a false deny costs more than a missed scan. The bypass is never silent: the audit line and span carry guardrail_verdict: error_failopen and the gateway logs a warning on every failed call. If you cannot say out loud which posture a guardrail should have, it should be fail-closed.


Why no retries

One attempt, bounded by the configured timeout. Guardrail callouts are synchronous on the request hot path: retrying a sick endpoint doubles (or worse) the latency every matched caller pays, and the second attempt of a non-idempotent scan usually buys nothing the first did not. The fail posture is the outage answer, not retries.

The audit webhook sink does retry — deliberately different. It is asynchronous and off the hot path, so spending wall-clock there costs no caller anything. Two HTTP clients, two latency contracts.


Why there is no SSRF guard on this path

mcpgw runs a post-DNS dial guard that refuses connections to loopback, RFC1918, link-local, and CGNAT addresses. It is wired into exactly one egress: the OAuth CIMD fetcher, whose destination URL arrives inside a client’s JWT claim. That is attacker-influenced input, so the gateway must not be steerable toward internal services or the cloud metadata endpoint.

Guardrail webhooks are the opposite case, and the guard is deliberately absent. The URL comes from the operator’s config file, and private is the normal destination — the quickstart endpoint is 127.0.0.1, and an in-cluster classifier resolves into RFC1918 space. Enabling the guard here would reject the documented deployment while defending against nothing an attacker controls: anyone who can edit guardrails[].url can already edit upstreams, policy, and auth.

What does apply on this path is the HTTPS-unless-loopback rule, which defends the case that is real — a network attacker forging pass verdicts to silently disable the guardrail.


Refusing, not truncating, oversize envelopes

The gateway accepts requests up to 16 MiB, and the webhook receives the whole envelope. Handing 16 MiB to a synchronous call with a 1s budget is not a decision the operator gets to make well: the transfer alone will usually blow the deadline, so the request fails closed after the bandwidth has been spent. max_body_bytes therefore caps the envelope any endpoint can be sent, defaulting to 1 MiB — orders of magnitude above a real tool call, and an explicit bound on the exposure of the one feature that ships bodies off-box.

The interesting choice is what to do at the boundary. Truncating to fit is the tempting option and the wrong one: the guardrail would render a verdict on an envelope it did not fully see, and pass would then mean “the first megabyte looked fine.” That is a security control quietly reporting success on unexamined input. So an oversize envelope is refused outright and treated as a callout failure, taking the same fail posture as an unreachable endpoint.

It gets its own error_kind (guardrail_body_too_large) because the operator’s response differs: the endpoint is healthy and was never contacted, and the fix is raising the cap or narrowing the rule’s matchers, not debugging a webhook. For the same reason the refusal is invisible to the circuit breaker — it describes the caller’s traffic, not the endpoint’s health, and letting a run of large payloads open the breaker would take a working guardrail offline for every other request.


The breaker: no retries, but no repeated waiting either

“No retries” bounds what a single request spends on a sick endpoint. It says nothing about the thousandth request, which without a breaker would independently discover the same outage and pay the same timeout. That is the failure mode the circuit breaker closes: after breaker.failure_threshold consecutive failures (default 5), the endpoint is treated as established-sick and matched requests stop calling it for breaker.open_duration (default 5s), applying the fail posture immediately instead.

The property that makes this safe is that the breaker is decision-preserving. A short-circuited request receives exactly what an attempted-and-failed request would have: a deny under the default posture, an allow with guardrail_verdict: error_failopen under fail_open: true. No enforcement is traded away — only the waiting, and the load a failing endpoint receives while it is trying to restart. This is why the breaker needs no new audit vocabulary: error_kind: guardrail_unreachable was already the honest description, and a guardrail_latency_ms near zero is the tell that the answer came from the breaker rather than the wire.

Recovery is deliberately stingy. When the open window elapses the breaker admits exactly one probe; concurrent requests keep short-circuiting until it resolves. A recovering endpoint therefore sees a single request rather than the full backlog arriving at once — the thundering-herd shape that turns a recovering service back into a failing one. Success closes the breaker; failure re-opens it for another window.

Two things deliberately do not count. Client cancellations (a caller disconnecting mid-request) end the callout with an error that says nothing about endpoint health, so they are excluded — otherwise a burst of impatient clients could trip a perfectly healthy guardrail. And breaker state lives in the registry, so a SIGHUP reload starts every endpoint closed: an operator fixing a broken endpoint and reloading gets an immediate retest, not a wait.

The same reasoning as fail_open applies to failure_threshold: 0, which disables the breaker: it exists for operators who would rather every request test the endpoint itself, and like fail-open it should be a decision you can say out loud.


The latency budget is per-matched-request

Every request a guardrail rule matches pays the webhook round trip inline, before the upstream is ever contacted. When the endpoint is slow, matched traffic is slow by the same amount; when the endpoint hangs, matched requests pay up to the full timeout until the breaker opens (see above) — after which they pay nothing and get the same answer. That is why the timeout defaults to 1s and is capped at 10s, and why the matcher is the real tuning knob:

when:
  tool_prefix: "fs_"     # only write-shaped calls pay for the PII scan

A tool_name: "*" guardrail rule puts your webhook on the critical path of every call through the gateway. Keep matchers narrow, and watch guardrail_latency_ms in the audit log (and mcp.guardrail.latency_ms on spans) — it is the wall time of the webhook call alone, so webhook cost is never conflated with upstream latency.


redact vs guardrail: deterministic vs semantic

The two mutation actions answer different questions:

  • redact is deterministic and in-process: “replace anything that matches this RE2 pattern.” Zero added latency, zero external dependencies, but it can only find what a regex can describe.
  • guardrail is semantic and out-of-process: “ask this service what to do with the envelope.” It can run a model, consult a DLP system, apply policy that changes daily — at the price of network latency, an availability dependency, and shipping the full body off-box.

If a pattern can express the rule, use redact. Reach for guardrail when the decision needs context a pattern cannot hold. They divide the ruleset rather than stack on one request — first-match-wins means a call is governed by one action or the other — so deterministic traffic stays in-process and only the calls that need a classifier pay for it.


Why SSE interception is deferred

Guardrail callouts fire on client_to_server requests and on server_to_client buffered JSON responses. They do not fire per SSE frame in v1.

Per-frame synchronous HTTP is a latency footgun: a streaming response fans out into N frames, and N synchronous webhook calls — each potentially paying the full timeout against a sick endpoint — turns a stream into a stall. The honest v1 answer is that SSE keeps the in-process mechanisms (deny, redact, rate_limit per frame) and guardrails govern everything else. policy lint warns (guardrail-sse-direction) when a direction: server_to_client guardrail rule exists, so nobody believes SSE is covered when it is not. If per-frame callouts ship later, they will need a different shape — batching, caching, or an async verdict channel — not just the same call in a loop.


Direction is a field, not a path

agentgateway splits guardrail traffic by URL path — the gateway calls /request for one direction and /response for the other. mcpgw instead sends both directions to one endpoint with a direction field in the JSON payload (client_to_server / server_to_client), alongside method, tool_name, principal, and the full envelope as body.

One endpoint means one URL, one TLS configuration, one set of credentials — and the direction is data the webhook can branch on rather than routing it must implement. It also keeps the rule grammar mcpgw-native: when.direction already exists for SSE-frame rules, so guardrail rules read exactly like the rest of the policy block.


The webhook’s reason never leaves the box

Rejections stay generic client-side (-32001 policy_denied, nothing more) — the webhook’s reason string is operator-facing only. It is not audited, not spanned, and not returned to the caller; it goes to slog.Debug and nowhere else. This keeps the audit log metadata-only even though the webhook itself receives full bodies: what the gateway stores about a guardrail decision is name, verdict (pass|mask|reject|error_failopen), and latency — never the classifier’s explanation, which may itself quote sensitive content.

One disclosure follows from storing the marker on the request’s single audit line: if callouts fire in both directions on one request, the response-direction marker overwrites the request-direction one (last write wins — OTel SetAttributes overwrites by key, and the span behaves the same way). The request-direction verdict is then not visible in audit or on the span.