How to enable a guardrail webhook
Problem: you want a policy decision to come from something mcpgw does not have in-process — a PII classifier, a prompt-injection detector, or your own service — instead of a static allow/deny/redact rule.
Solution: name an outbound webhook under the top-level guardrails: section, then point an action: guardrail rule at it. When the rule matches, mcpgw POSTs the full JSON-RPC envelope to the webhook and maps its verdict: pass → allow, mask → forward the mutated body, reject → deny.
SECURITY: the webhook receives the full JSON-RPC body (tool arguments included). This is the first mcpgw feature that ships request bodies off-box — the audit log stays metadata-only. Treat the endpoint as a trusted data recipient: run it inside your trust boundary, over HTTPS, and mind what it logs.
Recipe
guardrails:
- name: pii-scanner # [a-z0-9_-]+, unique
url: http://127.0.0.1:8585/check # https:// required unless loopback
timeout: 1s # default 1s; max 10s
fail_open: false # default: endpoint down → deny
policy:
rules:
- id: guard-pii-on-writes
action: guardrail
guardrail: pii-scanner # must reference a defined endpoint
when:
tool_prefix: "fs_"
Reload with SIGHUP — the endpoint registry swaps atomically alongside the policy engine. A reload whose guardrails: section is invalid is rejected whole; the old registry stays live.
A reference webhook server
The contract: POST with Content-Type: application/json, one attempt bounded by the configured timeout. Anything other than HTTP 200 with a well-formed verdict — non-200, timeout, malformed JSON, unknown action — is an error handled by the endpoint’s fail posture.
The request mcpgw sends:
{ "direction": "client_to_server", "method": "tools/call", "tool_name": "fs_write", "principal": "oauth:sub:alice", "body": { "jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": {} } }
The verdicts the webhook may return (HTTP 200 only):
{ "action": "pass" }
{ "action": "reject", "reason": "pii detected" }
{ "action": "mask", "body": { "jsonrpc": "2.0", "id": 1, "result": { "content": "[MASKED]" } } }
reasonis optional and operator-facing only — never sent to the client, never audited.maskrequiresbody: it replaces the envelope that gets forwarded.
A ~20-line Python endpoint that rejects bodies containing a US-SSN-shaped string (plain HTTP on loopback — development only):
#!/usr/bin/env python3
import json
import re
from http.server import BaseHTTPRequestHandler, HTTPServer
SSN = re.compile(r"\b\d{3}-\d{2}-\d{4}\b")
class Guardrail(BaseHTTPRequestHandler):
def do_GET(self):
self.send_response(200) # liveness for `mcpgw doctor`'s GET probe
self.end_headers()
def do_POST(self):
raw = self.rfile.read(int(self.headers["Content-Length"]))
req = json.loads(raw)
hit = SSN.search(json.dumps(req.get("body", "")))
verdict = {"action": "reject", "reason": "ssn detected"} if hit else {"action": "pass"}
out = json.dumps(verdict).encode()
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(out)))
self.end_headers()
self.wfile.write(out)
def log_message(self, *args):
pass
HTTPServer(("127.0.0.1", 8585), Guardrail).serve_forever()
Verify offline
policy test never makes network calls. Without a stub it reports which endpoint the rule would call; --guardrail-stub folds in the outcome a verdict would produce:
$ mcpgw policy test --config=mcpgw.yaml --tool=fs_write --guardrail-stub=reject
action: deny
rule: guard-pii-on-writes
guardrail: pii-scanner (stub verdict: reject)
A stubbed reject exits 1 like any deny; pass and mask exit 0. With --json, the output carries "guardrail": "pii-scanner" and a "guardrail_stub": "reject" provenance field so a stub-forced outcome is never mistaken for an engine-derived one. A stubbed mask passes the body through unchanged and notes on stderr that the webhook’s mutated body is unknowable offline.
Verify live
Start the reference server and the gateway, then send a matching call:
$ python3 guardrail_server.py &
$ mcpgw --config=mcpgw.yaml &
$ curl -i http://127.0.0.1:7332/mcp -H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"fs_write","arguments":{"note":"ssn 123-45-6789"}}}'
HTTP/1.1 403 Forbidden
{"jsonrpc":"2.0","error":{"code":-32001,"message":"policy_denied"}}
The client stays generic — the webhook’s reason never reaches it. The audit line carries the provenance:
{ "ts": "2026-07-29T00:14:24.281Z", "session_id": "demo", "method": "tools/call", "tool_name": "fs_write", "action": "deny", "rule_id": "guard-pii-on-writes", "latency_ms": 6, "guardrail_name": "pii-scanner", "guardrail_verdict": "reject", "guardrail_latency_ms": 4 }
Stop the webhook and resend: with the default fail_open: false the call is still denied, now with error_kind: "guardrail_unreachable" (and no guardrail_verdict — the failure itself is the signal). With fail_open: true the call flows and the audit line shows guardrail_verdict: "error_failopen", so the bypass is visible rather than silent.
mcpgw doctor probes each endpoint with a plain GET: any HTTP status below 500 counts as reachable and the check reports TLS validity and round-trip latency. An unreachable endpoint is an error when fail_open: false, a warning when fail_open: true. The probe does not attach configured static headers, so a 401 from an auth-gated endpoint still counts as reachable.
Production notes
- HTTPS unless loopback. Endpoint URLs must be
https://for any non-loopback host (same posture asjwks_url). A MITM on the webhook path could forgepassverdicts and disable the guardrail entirely.tls.ca_file/cert_file/key_filecover private CAs and mTLS;tls.skip_verifyexists for dev and logs a loud startup warning. - Credentials via env.
headersvalues expand${VAR}at config load:Authorization: "Bearer ${GUARDRAIL_TOKEN}". - Size the timeout honestly. The call is synchronous: every matched request pays the round trip, and when the endpoint hangs matched requests pay the full
timeoutuntil the breaker opens (below). Default1s, hard cap10s. Keepwhenmatchers narrow so only traffic that needs the check pays for it; watchguardrail_latency_msin audit. - No retries. One attempt per matched request. Retrying on the hot path would double worst-case latency on a sick endpoint. (The async audit webhook sink does retry — different tradeoff.)
- Bound what leaves the box.
max_body_bytes(default 1 MiB) caps the envelope an endpoint can be sent. Oversize bodies are refused before the callout — never truncated, since a verdict on a partial envelope approves content the guardrail never saw — and take the fail posture witherror_kind: guardrail_body_too_large. The refusal does not count toward the breaker. Set0to remove the cap and rely only on the gateway’s 16 MiB inbound limit. - A sustained outage trips the breaker, not every request. After 5 consecutive failures (
breaker.failure_threshold) calls short-circuit for 5s (breaker.open_duration) and apply the fail posture immediately, then one probe tests recovery. The outcome is identical to an attempted-and-failed call — you keep the enforcement, you stop paying the timeout and stop hammering a sick endpoint. Watch forcircuit breaker openin the logs;guardrail_latency_msnear0on denied requests is the audit-side tell. Setfailure_threshold: 0to disable it. - SSE streams are not intercepted in v1. A
direction: server_to_clientguardrail rule fires on buffered JSON responses only, never per SSE frame;policy lintwarns (guardrail-sse-direction) so you do not believe SSE is covered. Regexredact/denyrules remain the SSE mechanism.
Replay caveat
policy replay cannot re-derive a guardrail decision offline — the verdict is not in the audit log and request bodies never are. Records whose audited decision came from a callout (guardrail_name is set) are counted as guardrail_skipped, reported on stderr, shown inline in the text summary and as guardrail_skipped in --json, and excluded from the divergence gate. Divergences on any other record still fail the gate.
Related
- Explanation: guardrail webhooks — fail posture, latency budget, engine purity
- Reference: configuration — full
guardrails:schema - Reference: policy rule schema
- Reference: audit log schema — the
guardrail_*fields