Configuration reference
Full schema for mcpgw.yaml. The default search path is /etc/mcpgw/mcpgw.yaml; override with --config=/path/to/file.yaml on mcpgw server.
mcpgw validates the entire config at startup and on every SIGHUP. Validation errors are fatal at startup and rejected on reload (the previous config remains in effect).
Top-level keys
| Key | Type | Default | Required | Hot-reload? |
|---|---|---|---|---|
listen |
string host:port |
— | yes | no |
tls |
object | nil |
no | partial |
input_rate_limit |
object | disabled | no | yes |
trusted_proxies |
list of CIDR strings | [] |
no | no |
license |
object | — | yes | no |
default_upstream |
string | "" |
recommended | no |
upstreams |
list of upstream |
[] |
yes | no |
routes |
list of route |
[] |
no | no |
auth |
object | disabled | no | partial |
identity |
object | defaults | no | yes |
policy |
object | — | yes | yes |
telemetry |
object | — | no | partial |
audit |
object | — | yes | yes |
tool_search |
object | disabled | no | no |
rate_limit_store |
object | in_memory |
no | no |
guardrails |
list of guardrail endpoint |
[] |
no | yes |
allowed_protocol_versions |
list of string | all versions | no | no |
task_affinity |
object | defaults | no | no |
subscriptions |
object | defaults | no | no |
Hot-reload column: “yes” means SIGHUP re-applies; “no” means restart required; “partial” means some sub-keys reload and others do not (called out per-section).
listen
listen: 0.0.0.0:7332
host:port the gateway binds to. Use 0.0.0.0 to accept from any interface; use 127.0.0.1 to restrict to localhost. The gateway exits non-zero if the port is in use.
tls
tls:
cert_file: /etc/mcpgw/tls/fullchain.pem
key_file: /etc/mcpgw/tls/privkey.pem
client_ca: /etc/mcpgw/tls/client-ca.pem # optional — enables mTLS
When cert_file and key_file are set, listen serves HTTPS. When tls is absent, listen serves plaintext HTTP. cert_file and key_file must be configured together.
client_ca enables mutual TLS: every client connection must present a certificate signed by the listed CA file. client_ca requires cert_file and key_file.
On SIGHUP, mcpgw re-reads cert_file, key_file, and client_ca from the current config. If the reload fails, the previous certificate remains active. Changing TLS from disabled to enabled, or enabled to disabled, requires a restart.
input_rate_limit
input_rate_limit:
enabled: true
requests_per_second: 50
burst: 100
Protects the gateway itself from request floods before body read, auth, JSON-RPC parse, routing, policy, or upstream forwarding. Buckets use the TCP peer address from RemoteAddr unless that peer matches trusted_proxies.
This limiter is separate from policy.rules[].action: rate_limit. Use input_rate_limit for gateway resource protection and policy rate limits for per-tool quotas.
| Field | Type | Default | Notes |
|---|---|---|---|
enabled |
bool | false |
When false, no input throttle is applied. |
requests_per_second |
float | — | Token refill rate. Required and must be > 0 when enabled. |
burst |
int | — | Bucket capacity. Required and must be > 0 when enabled. |
On SIGHUP, mcpgw re-applies the input limiter config (rate/burst). The underlying rate-limit store is reused across reloads, so existing buckets are preserved — a reload no longer hands every client a fresh burst.
trusted_proxies
trusted_proxies:
- 10.42.0.0/16
- 2001:db8:42::/48
CIDRs for reverse proxies allowed to provide the effective client IP used by input_rate_limit and anonymous policy rate-limit buckets. The default is empty: Forwarded and X-Forwarded-For are ignored and the TCP peer from RemoteAddr is used.
When the immediate peer is trusted, mcpgw prefers RFC Forwarded, falls back to X-Forwarded-For, and walks the chain from right to left until it finds the first untrusted address. This prevents a client-controlled leftmost value from selecting an arbitrary bucket. Invalid or absent forwarding headers fall back to the TCP peer.
Only list the CIDRs from which the gateway actually receives proxy connections, and enforce the same boundary with firewall or security-group rules. A broad trusted CIDR lets any host in that range assert client identity. Changes require a restart and appear in /readyz.restart_pending after SIGHUP.
rate_limit_store
Selects the backend shared by both input_rate_limit and policy.rules[].action: rate_limit. The default (in_memory) keeps per-process token buckets. With multiple replicas behind a load balancer, per-process buckets mean the effective limit is N × configured (each replica has its own). redis shares buckets across replicas so the deployment enforces one global limit.
rate_limit_store:
type: redis # in_memory (default) | redis
url: redis://:${REDIS_PASSWORD}@redis:6379/0
key_prefix: "mcpgw:rl:" # default; namespaces keys in a shared Redis
operation_timeout: 50ms # per-call cap on the request hot path
fail_open: false # default: Redis unreachable → deny
tls:
enabled: false
ca_file: /etc/mcpgw/redis-ca.pem
| Field | Type | Default | Notes |
|---|---|---|---|
type |
string | in_memory |
in_memory or redis. |
url |
string | — | Required for redis. redis:// or rediss://; supports host:port, auth, db index. |
key_prefix |
string | mcpgw:rl: |
Prefix for all bucket keys; lets mcpgw share a Redis with other apps. |
operation_timeout |
duration | 50ms |
Max time per Redis call. Exceeding it is treated as a store failure (see fail_open), protecting the request hot path from a slow Redis. |
fail_open |
bool | false |
On Redis error/timeout: false denies the request so the configured limit remains fail-closed; true explicitly prioritizes availability and allows it. After a failure, a 1s window short-circuits further calls to avoid thundering-herd on a sick Redis. |
tls.enabled |
bool | false |
Use TLS to Redis. (rediss:// URLs also enable TLS.) |
tls.ca_file |
string | — | PEM CA bundle to verify the Redis server certificate. |
type is not hot-reloadable — switching in_memory ↔ redis requires a restart (it would otherwise drop all live buckets). The shared store instance is reused across SIGHUP, so reloads do not reset buckets.
See How-to: configure a shared rate-limit store.
guardrails
Names outbound guardrail webhook endpoints. A policy rule with action: guardrail references one of these by name; when the rule matches, the proxy POSTs the full JSON-RPC envelope to the endpoint synchronously and maps its verdict (pass → allow, mask → forward the mutated body, reject → deny). See policy rule schema: action guardrail for the verdict contract and How-to: enable a guardrail webhook for an end-to-end setup.
SECURITY: the webhook receives the full JSON-RPC body (tool arguments included). This is the first mcpgw feature that ships request bodies off-box — the audit log is metadata-only. Treat every endpoint as a trusted data recipient.
guardrails:
- name: pii-scanner # required, unique, [a-z0-9_-]+
url: https://guardrails.internal/check # https:// required unless loopback
timeout: 1s # default 1s; must be > 0 and <= 10s
fail_open: false # default: callout failure → deny
max_body_bytes: 1048576 # default 1 MiB; 0 removes the cap
headers: # optional static headers
Authorization: "Bearer ${GUARDRAIL_TOKEN}"
breaker: # optional; circuit breaker
failure_threshold: 5 # default 5; 0 disables the breaker
open_duration: 5s # default 5s; must be <= 5m
tls: # optional
enabled: false
ca_file: /etc/mcpgw/guardrail-ca.pem
cert_file: /etc/mcpgw/guardrail-client.pem
key_file: /etc/mcpgw/guardrail-client.key
| Field | Type | Default | Notes |
|---|---|---|---|
name |
string | — | Required, unique, must match ^[a-z0-9_-]+$. Appears in audit (guardrail_name), spans (mcp.guardrail.name), and rule references. |
url |
string | — | Required. Must be https:// unless the host is loopback (same posture as auth.oauth.jwks_url) — a MITM on this path could forge pass verdicts and disable the guardrail. |
timeout |
duration | 1s |
Per-call deadline covering the whole round trip. Must be > 0 and <= 10s — the call is synchronous on the request hot path. One attempt; no retries. |
fail_open |
bool | false |
On callout error (non-200, timeout, malformed verdict): false denies the request (error_kind: guardrail_unreachable); true allows it and records verdict error_failopen plus a warning log, so the bypass is visible. |
max_body_bytes |
int | 1048576 (1 MiB) |
Largest envelope this endpoint will be sent. An oversize body is never truncated — it is refused before the callout and takes the fail posture (error_kind: guardrail_body_too_large). 0 removes the cap, leaving only the gateway’s 16 MiB inbound limit. See Outbound body cap. |
headers |
map[string]string | {} |
Static headers sent on every call. Values expand ${ENV_VAR} at config load. |
breaker.failure_threshold |
int | 5 |
Consecutive callout failures that open the circuit breaker. 0 disables the breaker (every request attempts the endpoint). See Circuit breaker below. |
breaker.open_duration |
duration | 5s |
How long an open breaker short-circuits before admitting one probe. Must be >= 0 and <= 5m. |
tls.enabled |
bool | false |
Enable the TLS block below. |
tls.ca_file |
string | system roots | Optional PEM CA bundle appended to the system pool (private CA). |
tls.cert_file |
string | — | Optional client certificate for mTLS; requires key_file. |
tls.key_file |
string | — | Optional client key for mTLS; requires cert_file. |
tls.skip_verify |
bool | false |
Development escape hatch; logs a loud startup warning. Never use in production. |
Guardrail outbound body cap
The webhook receives the full JSON-RPC envelope, and the gateway accepts requests up to 16 MiB. Shipping a multi-megabyte envelope inside a 1s synchronous budget is a guaranteed timeout that spends the bandwidth first, so max_body_bytes caps what any endpoint can be sent, defaulting to 1 MiB — far above a realistic tool call.
An envelope over the cap is refused, never truncated: a verdict rendered on a partial envelope would approve content the guardrail never saw. The request is treated as a callout failure and takes the endpoint’s fail posture — fail-closed denies with error_kind: guardrail_body_too_large, fail_open: true allows with verdict error_failopen. The webhook is not contacted, and the refusal does not count toward the circuit breaker: it is a property of the caller’s traffic, not the endpoint’s health, so a run of large payloads cannot take a healthy guardrail offline for everyone else.
max_body_bytes: 0 removes the separate cap; only the gateway’s inbound limit then applies.
Guardrail circuit breaker
The callout is synchronous on the request hot path, so without a breaker an endpoint that hangs costs every matched request its full timeout before the fail posture applies — endpoint availability becomes gateway availability for matched traffic.
After breaker.failure_threshold consecutive failures the breaker opens: for breaker.open_duration, matched requests skip the callout and apply the endpoint’s fail posture immediately (fail-closed → deny with error_kind: guardrail_unreachable; fail_open: true → allow with verdict error_failopen). Once the window elapses the breaker is half-open and admits exactly one probe — success closes it, failure re-opens it for another window. A success at any point clears the failure run.
The breaker is decision-preserving: a short-circuited request gets exactly the outcome an attempted-and-failed request would have gotten, so it trades no enforcement for the latency and endpoint load it saves. Audit and span fields are unchanged; guardrail_latency_ms is ~0 on a short-circuited request. Trips and recoveries are logged (circuit breaker open / circuit breaker closed, naming the guardrail), rate-limited to at most one line per open window.
Caller cancellations — a client disconnecting mid-request — are not counted against the endpoint, so a burst of them cannot trip a healthy guardrail. Set failure_threshold: 0 to disable the breaker entirely.
Hot-reloadable: the endpoint registry swaps atomically on SIGHUP alongside the policy engine. A reload whose guardrails: section is invalid is rejected whole and the previous registry stays live. Breaker state lives in the registry, so a reload starts every endpoint closed. mcpgw doctor probes each endpoint (plain GET, status < 500 = reachable, TLS + latency reported; static headers are not attached, and the probe is not breaker-gated).
license
license:
path: /etc/mcpgw/license.jwt
| Field | Type | Default | Notes |
|---|---|---|---|
path |
string | — | Absolute path to the JWT file. Must be 0600 and readable by the mcpgw process owner. |
The grace window (grace_days) is embedded in the license JWT itself, not set in YAML config. See License JWT claims for the JWT structure and Explanation: licensing for fail-open / fail-closed behavior.
default_upstream
default_upstream: filesystem
Logical upstream name (must match an upstreams[].name) used for MCP methods that have no tool_name: initialize, tools/list, resources/list, prompts/list, ping. Without it, spec-compliant clients fail at handshake with 404 no_route.
If unset, methods without a tool name return 404. There is no “broadcast to all upstreams” mode today; multi-upstream tools/list aggregation is on the roadmap.
upstreams
upstreams:
- name: filesystem
url: http://mcp-fs:9000
| Field | Type | Default | Notes |
|---|---|---|---|
name |
string | — | Unique logical name. Appears in mcp.upstream span attribute. Lowercase recommended. |
url |
string | — | Base URL. Scheme must be http or https; host required. Path is the upstream’s MCP endpoint (most servers use /). |
URLs are validated at startup. default_upstream must reference a known name.
routes
routes:
- match: { tool_prefix: "fs_" }
upstream: filesystem
- match: { tool_prefix: "git_" }
upstream: git
- match: { tool_name_in: ["db_query", "db_schema"] }
upstream: database
Evaluated top-down. First match wins. Routes support the same tool-name matcher fields as policy rules: tool_name, tool_prefix, tool_glob, tool_regex, and tool_name_in. Tool calls that match no route and fall outside default_upstream’s remit return 404 no_route.
routes[].upstreammust reference a knownupstreams[].name.routes[].matchmust set exactly one tool matcher:tool_name,tool_prefix,tool_glob,tool_regex, ortool_name_in.routes[].match.methodandroutes[].match.directionare not supported because routing is tool-name based.
auth
auth:
enabled: true
header: Authorization
scheme: Bearer
keys:
- id: claude-desktop-prod
hash: "argon2id$v=19$m=65536,t=3,p=2$..."
scopes: ["*"]
created_at: "2026-05-01T14:23:01Z"
expires_at: "2026-08-01T14:23:01Z"
| Field | Type | Default | Hot-reload? | Notes |
|---|---|---|---|---|
enabled |
bool | false |
yes | When false or absent, /mcp preserves unauthenticated v1 behavior. |
header |
string | Authorization |
no | Header to read. |
scheme |
string | Bearer |
no | Prefix before the key. Set to "" to read the raw header value. |
keys[].id |
string | — | yes | Human-readable key id. Appears in audit and spans. Must be unique. |
keys[].hash |
string | — | yes | Argon2id hash from mcpgw key generate. |
keys[].scopes |
list | [] |
yes | Reserved for future per-key authorization. Parsed and preserved, but not enforced in this release. |
keys[].created_at |
timestamp | zero | yes | Informational. |
keys[].expires_at |
timestamp | zero | yes | Optional. Zero means never expires. |
If auth.enabled: true, keys must be non-empty. Missing, invalid, and expired keys return HTTP 401 with JSON-RPC -32005 unauthorized.
Use How-to: enable API-key authentication for the full recipe.
auth.oauth
auth:
enabled: true
oauth:
enabled: true
issuer: https://idp.example.com
audience: mcpgw
jwks_url: https://idp.example.com/.well-known/jwks.json
required_scopes: ["mcp:read"]
jwks_cache_ttl: 5m
leeway: 30s
metadata_path: /.well-known/oauth-protected-resource
public_url: https://gateway.example.com
When auth.oauth.enabled: true, the gateway accepts RFC 6750 Bearer tokens in addition to (or instead of) API keys. The gateway fetches JWKS from jwks_url, caches keys for jwks_cache_ttl, and validates each token for issuer, audience, expiry, and (optionally) required scopes.
auth.enabled: true is required when auth.oauth.enabled: true.
| Field | Type | Default | Hot-reload? | Notes |
|---|---|---|---|---|
auth.oauth.enabled |
bool | false |
yes | Enables OAuth 2.1 Bearer-token verification. Requires auth.enabled: true. |
auth.oauth.issuer |
string | — | no | Required when enabled. Must be an HTTP or HTTPS URL with a non-empty host. Matched against the JWT iss claim. |
auth.oauth.audience |
string | — | no | Required when enabled. Matched against the JWT aud claim. |
auth.oauth.jwks_url |
string | — | no | Required when enabled. Must be an HTTP or HTTPS URL with a non-empty host. Used to fetch the signing key set. |
auth.oauth.required_scopes |
list | [] |
yes | Optional. Every listed scope must appear in the token’s scope (or scp) claim. An empty list disables scope enforcement. |
auth.oauth.jwks_cache_ttl |
duration | 5m |
yes | How long fetched JWKS are cached before re-fetching. |
auth.oauth.leeway |
duration | 30s |
yes | Clock-skew tolerance applied to exp and nbf checks. |
auth.oauth.metadata_path |
string | /.well-known/oauth-protected-resource |
no | HTTP path where the gateway serves RFC 9728 protected-resource metadata. |
auth.oauth.public_url |
string | derived from listen |
no | Base URL advertised in WWW-Authenticate headers and the metadata document. Required in production when listen binds 0.0.0.0; the gateway logs a warning if unset in that case. Must be an HTTP or HTTPS URL with a non-empty host. |
auth.oauth.cimd.enabled |
bool | false |
yes | Master switch. When false (default), no CIMD fetches occur. |
auth.oauth.cimd.timeout |
duration | 5s |
yes | Per-fetch HTTP timeout. First-request latency: the CIMD fetch is synchronous in the auth path. The first request from a previously-unseen URL-format client_id can block for up to this duration if the endpoint is unresponsive (the request still completes — errors are silenced — but latency is affected). Subsequent requests are served from cache. |
auth.oauth.cimd.max_body_bytes |
int64 | 65536 (64 KiB) |
yes | Maximum response body size. Larger responses are rejected (silently — CIMD failures never block authenticated requests). |
auth.oauth.cimd.cache_ttl |
duration | 1h |
yes | How long fetched metadata is cached. The cache is keyed by URL and bounded by a 1024-entry LRU (least-recently-used eviction on insert; reads promote). The cap is not currently exposed in YAML — change it at the Go API level if a different ceiling is needed. Hot reload (SIGHUP) creates a fresh fetcher with a fresh cache. |
auth.oauth.cimd — Client ID Metadata Documents
When the verified JWT’s client_id (or azp fallback) is an https:// URL, mcpgw fetches that URL and surfaces the resulting client_name field as audit.oauth_client_name and the mcp.auth.oauth_client_name span attribute. CIMD failures are silent: a misbehaving CIMD endpoint cannot block authenticated requests, only degrade audit richness.
auth:
enabled: true
oauth:
enabled: true
# ... other oauth fields ...
cimd:
enabled: true
timeout: 5s
max_body_bytes: 65536
cache_ttl: 1h
Hot reload (SIGHUP) creates a fresh fetcher with a fresh cache. Toggling enabled: false and sending SIGHUP atomically disables CIMD fetching.
identity
identity:
principal_claim: sub
groups_claim: groups
Controls how verified OAuth JWT claims are projected into gateway policy, audit, and rate-limit identity. The block is optional; defaults are applied at startup and on reload.
| Field | Type | Default | Hot-reload? | Notes |
|---|---|---|---|---|
principal_claim |
string | sub |
yes | JWT claim used as the OAuth principal for policy rate_limit, audit principal, and telemetry mcp.principal. If absent in a token, mcpgw falls back to client_id, then to the client IP. |
groups_claim |
string | groups |
yes | JWT claim carrying group memberships for policy.rules[].when.claims.groups and groups_in. The verifier accepts either an array of strings or a single string. Common alternatives: roles for Keycloak, groups for Okta and Azure AD. |
identity only affects verified OAuth callers. API-key callers do not have an OAuth claim identity, so policy rules with a non-empty when.claims block do not match them. Use Policy rule schema: when.claims for claim matcher semantics.
tool_search
tool_search:
enabled: true
mode: synthesize
refresh_interval: 60s
max_results: 20
Tool-search synthesis. When enabled in synthesize mode, the gateway intercepts tools/list and tools/call mcp_search, serving them from a refreshed index of upstream tool definitions. In passthrough mode (the default), tools/list is forwarded to the upstream unchanged.
See Enable tool-search synthesis for the operator workflow and Tool-search tradeoffs for design rationale.
| Field | Type | Default | Hot-reload? | Description |
|---|---|---|---|---|
tool_search.enabled |
bool | false |
no | Master switch. When false, tool_search.mode is ignored and the upstream receives all requests unchanged. |
tool_search.mode |
enum | passthrough |
no | passthrough (no interception, default — the gateway forwards tools/list to the upstream as-is. Even with enabled: true, the index is NOT built in this mode; this combination is a no-op preserved for forward compatibility with future per-tenant mode toggling.) or synthesize (the gateway answers tools/list from a synthesized stub and tools/call mcp_search from the index). Mode change requires restart — SIGHUP only refreshes the index. |
tool_search.refresh_interval |
duration | 60s |
no | How often the gateway re-fetches tools/list from each upstream to rebuild the index. Per upstream, a shorter ttlMs on the list result (CacheableResult, 2026-07-28) tightens this — the effective cadence is min(refresh_interval, ttlMs), floored at 1s; a longer ttlMs never slows it down. |
tool_search.max_results |
int | 20 |
no | Hard cap on results returned by mcp_search. Caller-provided limit arguments above this value are clamped silently. |
tool_search is not hot-reloadable. Any change to tool_search.mode, tool_search.enabled, refresh_interval, or max_results requires a full restart.
CacheableResult (2026-07-28). Upstream tools/list results may carry ttlMs and cacheScope. The index honors both: ttlMs tightens the per-upstream refresh cadence (see refresh_interval above), and a cacheScope: "private" list is never merged into the shared index — the index is a shared cache across all clients, which private forbids. Excluded upstreams are logged at startup/refresh and their tools resolve through normal routing (passthrough); the gateway keeps polling them on cadence in case the scope changes.
Failing upstreams: negative-cache backoff. A tools/list fetch that errors puts that upstream into exponential backoff (initial 30s, doubling per consecutive failure, capped at 30m). Subsequent refresh cycles skip backed-off upstreams silently — their previous error is still surfaced in lastErrors so metrics scrapers continue to observe the issue, but no HTTP call is made. A successful fetch resets the backoff. Partial-failure semantics are preserved: a single healthy upstream still populates the merged index, and a stuck upstream does not block the merge. The schedule (30s initial / 30m ceiling) is not exposed in YAML.
allowed_protocol_versions
allowed_protocol_versions:
- "2025-06-18"
- "2026-07-28"
Optional allowlist of MCP protocol revisions the gateway will serve. Each request’s dialect is resolved once — params._meta["io.modelcontextprotocol/protocolVersion"] wins over the MCP-Protocol-Version header — and a request declaring a revision outside the list is rejected with -32022 unsupported_protocol_version before policy or routing run.
Unset (default) means all revisions are accepted: requests declaring no version anywhere resolve to the legacy dialect. When the list IS set, version-less requests are still accepted (they never declared an unsupported version); pin clients by requiring your fleet to send an explicit version if you need strictness. Not hot-reloadable.
task_affinity
task_affinity:
max_entries: 16384
ttl: 24h
Tunes gateway support for the MCP Tasks extension (io.modelcontextprotocol/tasks, SEP-2663). Not hot-reloadable.
| Field | Type | Default | Description |
|---|---|---|---|
task_affinity.max_entries |
int | 16384 |
Capacity of the principal + task-handle → upstream affinity map (LRU-evicted). |
task_affinity.ttl |
duration | 24h |
Entries expire this long after they were learned or last overwritten. |
When an upstream answers a request with a CreateTaskResult (resultType: "task"), the gateway records (principal, taskId) → upstream (from both plain-JSON responses and SSE frames). Principal scoping lets different callers safely receive the same opaque task ID from different upstreams without cross-tenant routing. Follow-up tasks/get/tasks/update/tasks/cancel calls — which carry only params.taskId, so no tool route can match them — are routed back to that caller’s minting upstream. Affinity is best-effort: an evicted, expired, or unknown handle falls back to default_upstream (the pre-affinity behavior). Single-upstream deployments skip the map entirely. The handle is surfaced as task_handle in audit lines and mcp.task.handle on spans for cross-RPC correlation.
Note tasks/* methods remain fail-closed under default_action: deny — affinity affects routing, not authorization; grant the methods with explicit method: allow rules.
subscriptions
subscriptions:
max_frames: 1000000
idle_timeout: 5m
Governs subscriptions/listen (2026-07-28): the single long-lived POST-response stream that replaced the HTTP GET endpoint and resources/subscribe. Listen streams run through the same per-frame policy engine as ordinary SSE responses — server_to_client deny/redact/rate-limit rules apply to each notification — and their audit frame lines carry the notification’s _meta subscription tag as subscription_id. Not hot-reloadable.
| Field | Type | Default | Description |
|---|---|---|---|
subscriptions.max_frames |
int | 1000000 |
Frame cap for a listen stream. Ordinary request streams keep their own 10,000-frame cap. |
subscriptions.idle_timeout |
duration | 5m |
Cuts the stream when no frame arrives for this long (terminal event: error + error_kind: sse_truncated). Listen requests are exempt from the 60s upstream client timeout, so this is the only lifetime bound; clients re-issue subscriptions/listen after a cut. |
Note subscriptions/listen is a data channel, not a handshake method: under default_action: deny it requires an explicit method: allow rule.
policy
policy:
default_action: allow
rules:
- id: deny-shell
action: deny
when: { tool_name: "shell_exec" }
- id: rl-fs-write
action: rate_limit
when: { tool_name: "fs_write" }
tokens_per_second: 10
burst: 20
- id: redact-secrets
action: redact
when: { tool_name: "*" }
redact:
- regex: 'Bearer [A-Za-z0-9._-]+'
replacement: "[REDACTED]"
See Policy rule schema for full grammar. Hot-reloadable: SIGHUP re-applies the entire policy block atomically.
default_action controls unmatched tools/call requests:
| Value | Default | Notes |
|---|---|---|
allow |
no | v1 behavior: unmatched tools/call requests pass through. Must be set explicitly — configs that relied on the old implicit allow need default_action: allow. |
deny |
yes | Explicit allowlist mode; applied when default_action is omitted (mcpgw fails closed). Unmatched tools/call requests return -32001 policy_denied and audit as rule_id: "default_deny". Unmatched non-tool protocol methods still route to default_upstream unless an explicit policy rule matches them. |
Policy rule action values:
| Value | Description |
|---|---|
allow |
Forward the request unchanged. |
deny |
Reject the request. HTTP 403, JSON-RPC -32001 policy_denied. |
redact |
Rewrite matching byte sequences in the request body before forwarding. Requires a redact[] list on the rule. |
rate_limit |
Token-bucket rate limit per (rule, principal). HTTP 429, JSON-RPC -32003 rate_limited when the bucket is empty. Requires tokens_per_second and burst. |
strip_app |
Parses the tool-call response and removes content blocks where type == "ui" or mimeType starts with "application/vnd.mcp-ui". Applies to both JSON and SSE responses. The stripped audit field records how many blocks were removed. |
guardrail |
Delegates the decision to a webhook endpoint named by the rule’s guardrail field (defined in the top-level guardrails: section). The proxy POSTs the full JSON-RPC envelope; pass → allow, mask → forward mutated body (audit redact), reject → HTTP 403 -32001 policy_denied. Fires on client_to_server requests and server_to_client buffered JSON responses — not per SSE frame in v1. |
Policy when fields:
| Field | Type | Notes |
|---|---|---|
method |
string | JSON-RPC method exact match (e.g. elicitation/create, sampling/createMessage). Empty or absent matches any method. |
direction |
enum | "" or absent (any direction, default), "client_to_server" (inbound requests only), "server_to_client" (SSE response frames only). Setting "server_to_client" enables SSE-frame inspection on responses from the upstream. |
claims |
object | OAuth identity claim conditions. Keys: groups, groups_in, scopes_in, subject_in. Applies to policy rules only, not routes. |
tool_name |
string | Exact tool name. "*" matches any tool. |
tool_prefix |
string | Tool name prefix match. |
tool_glob |
string | Glob pattern over tool name. |
tool_regex |
string | Go RE2 regex over tool name. |
tool_name_in |
list | Tool name allowlist; matches if the tool name is in the list. |
Set at most one tool matcher (tool_name, tool_prefix, tool_glob, tool_regex, or tool_name_in) per rule. method and direction may be combined with any tool matcher or used alone.
telemetry
telemetry:
customer:
enabled: true
endpoint: http://datadog-agent:4318
service_name: mcpgw
resource_attrs:
env: production
version: v1.0.0
operator:
enabled: false
endpoint: ""
service_name: ""
customer and operator are independent OTLP/HTTP exporters. Both are off by default. They share one OTel Resource; when customer is enabled its service_name and resource_attrs define that Resource, otherwise the operator block does. A blank service name falls back to mcpgw. The gateway build version is always emitted as service.version.
| Field | Type | Default | Hot-reload? | Description |
|---|---|---|---|---|
customer.enabled |
bool | false |
no | |
customer.endpoint |
string | — | no | |
customer.service_name |
string | mcpgw |
no | |
customer.resource_attrs |
map[string]string | {} |
no | |
operator.enabled |
bool | false |
no | |
operator.endpoint |
string | — | no | |
operator.service_name |
string | mcpgw |
no | |
operator.resource_attrs |
map[string]string | {} |
no | |
trust_incoming_trace_context |
bool | false |
no |
endpoint accepts a bare host:port; mcpgw appends /v1/traces. To override the path (OTLP-compatible proxies mounted under non-default paths), pass the full URL.
W3C trace context over _meta (SEP-414). With trust_incoming_trace_context: true, a client’s params._meta.traceparent/tracestate parents the gateway’s span — the request joins the caller’s trace for end-to-end client→gateway→server views. Default false: any client could otherwise forge trace lineage in your backend. Independently of the flag, when tracing is enabled the gateway injects its own span context into the forwarded body’s _meta so upstream spans parent under the gateway; with telemetry disabled, bodies are forwarded byte-identical.
audit
audit:
path: /var/log/mcpgw/audit.jsonl
max_size_mb: 100
compress_rotated: true
sinks:
- type: gcs
bucket: acme-mcpgw-audit
prefix: prod/
flush_interval: 60s
flush_batch_lines: 10000
flush_batch_bytes: 5000000
compress: true
- type: s3
bucket: acme-mcpgw-audit
region: us-east-1
prefix: prod/
flush_interval: 60s
flush_batch_lines: 10000
flush_batch_bytes: 5000000
compress: true
object_lock: true
retention_days: 2555
- type: kafka
brokers: ["kafka-1:9092", "kafka-2:9092"]
topic: mcpgw-audit
partition_strategy: hash
partition_key: session_id
acks: all
compression: zstd
flush_interval: 5s
flush_batch_lines: 5000
flush_batch_bytes: 1000000
tls:
enabled: true
sasl:
mechanism: scram-sha-512
username_env: KAFKA_USERNAME
password_env: KAFKA_PASSWORD
- type: webhook
url: https://siem.acme.com/ingest
method: POST
headers:
Authorization: "Bearer ${SIEM_TOKEN}"
retry:
max_attempts: 5
backoff: exponential
initial_interval: 1s
max_interval: 60s
| Field | Type | Default | Hot-reload? |
|---|---|---|---|
path |
string | — | yes |
max_size_mb |
int | 100 |
yes |
compress_rotated |
bool | true |
yes |
sinks[].type |
string | — | yes |
sinks[].bucket |
string | — | yes |
sinks[].region |
string | — | yes |
sinks[].prefix |
string | "" |
yes |
sinks[].object_lock |
bool | false |
yes |
sinks[].retention_days |
int | 2555 |
yes |
sinks[].brokers |
list | — | yes |
sinks[].topic |
string | — | yes |
sinks[].partition_strategy |
string | hash |
yes |
sinks[].partition_key |
string | session_id |
yes |
sinks[].acks |
string | all |
yes |
sinks[].compression |
string | zstd |
yes |
sinks[].tls.enabled |
bool | false |
yes |
sinks[].tls.ca_file |
string | system roots | yes |
sinks[].tls.cert_file |
string | — | yes |
sinks[].tls.key_file |
string | — | yes |
sinks[].tls.skip_verify |
bool | false |
yes |
sinks[].sasl.mechanism |
string | "" |
yes |
sinks[].sasl.username_env |
string | "" |
yes |
sinks[].sasl.password_env |
string | "" |
yes |
sinks[].url |
string | — | yes |
sinks[].headers |
map | {} |
yes |
sinks[].flush_interval |
duration | sink-specific | yes |
sinks[].flush_batch_lines |
int | sink-specific | yes |
sinks[].flush_batch_bytes |
int | sink-specific | yes |
sinks[].retry.max_attempts |
int | 3 |
yes |
sinks[].retry.backoff |
string | exponential |
yes |
sinks[].retry.initial_interval |
duration | 1s |
yes |
sinks[].retry.max_interval |
duration | 30s |
yes |
mcpgw rotates the file when its size exceeds max_size_mb. Rotated files are renamed to <path>.<unix-millis> and asynchronously gzipped if compress_rotated: true. mcpgw never deletes rotated files.
Audit sinks are additive best-effort copies. The local JSONL file remains canonical; sink failures are logged and do not block /mcp requests. GCS uses Google Application Default Credentials. S3 uses the AWS SDK default credential chain with the configured region. Kafka writes one audit JSONL line per record and uses Kafka-native batching/compression.
proxy_protocol (reserved)
proxy_protocol:
enabled: true
trusted_cidrs:
- 10.0.0.0/24
Reserved for a future release. Today this block is parsed but rejected at startup with a clear error. See How-to: behind a load balancer for the supported alternatives.
Validation rules
The startup validator enforces these invariants. Any failure is fatal; on SIGHUP reload, the previous config remains in effect.
upstreams[].nameis uniqueupstreams[].urlparses and has schemehttp/httpsand a non-empty hostdefault_upstream, if set, references a known upstream nameroutes[].upstreamreferences a known upstream nameroutes[].matchsets exactly one tool matcher:tool_name,tool_prefix,tool_glob,tool_regex, ortool_name_inroutes[].match.methodandroutes[].match.directionare unsupported because routing is tool-name basedroutes[].match.claimsis unsupported because routes are not an authorization surfaceroutes[].match.tool_globparses as a valid globroutes[].match.tool_regexcompiles as Go regexproutes[].match.tool_name_inis non-empty when setinput_rate_limit.requests_per_secondis> 0wheninput_rate_limit.enabledis trueinput_rate_limit.burstis> 0wheninput_rate_limit.enabledis true- every
trusted_proxies[]entry is a valid CIDR rate_limit_store.typeisin_memory,redis, or absentrate_limit_store.urlis required whentype: redisrate_limit_store.operation_timeoutis>= 0auth.keysis non-empty whenauth.enabledis trueauth.keys[].idis uniqueauth.keys[].hashparses as Argon2idauth.enabled: trueis required whenauth.oauth.enabled: trueauth.oauth.issuer,auth.oauth.jwks_url, andauth.oauth.public_url(when set) must usehttporhttpsscheme with a non-empty hostaudit.pathis requiredaudit.max_size_mbmust be>= 0;0means use the default100audit.sinks[].typeisgcs,s3,kafka, orwebhookaudit.sinks[type=gcs].bucketis non-emptyaudit.sinks[type=s3].bucketandregionare non-emptyaudit.sinks[type=kafka].brokersis non-empty and every broker ishost:portaudit.sinks[type=kafka].topicis non-empty and has no whitespaceaudit.sinks[type=kafka].partition_strategyisround_robin,hash,sticky, or absentaudit.sinks[type=kafka].partition_keyissession_id,auth_key_id,tool_name,method,request_id, or absentaudit.sinks[type=kafka].acksis0,1,all, or absentaudit.sinks[type=kafka].compressionisnone,gzip,snappy,lz4,zstd, or absentaudit.sinks[type=kafka].tls.cert_fileandkey_fileare set together and parse as an X.509 key pairaudit.sinks[type=kafka].tls.ca_file, if set, contains at least one PEM certificateaudit.sinks[type=kafka].sasl.mechanismisplain,scram-sha-256,scram-sha-512, or absentaudit.sinks[type=webhook].urlis HTTPS unlessallow_insecure: truepolicy.rules[].idis required and uniquepolicy.default_actionisallow,deny, or absentpolicy.rules[].actionis one ofallow,deny,redact,rate_limit,strip_app,guardrailpolicy.rules[].guardrailis required foraction: guardrail, must reference a definedguardrails[].name, and is rejected on any other actionguardrails[].namematches^[a-z0-9_-]+$and is uniqueguardrails[].urlis HTTPS unless the host is loopbackguardrails[].timeoutis> 0and<= 10s(default1swhen omitted)guardrails[].breaker.failure_thresholdis>= 0(default5when omitted;0disables the breaker)guardrails[].breaker.open_durationis>= 0and<= 5m(default5swhen omitted)guardrails[].max_body_bytesis>= 0(default1048576when omitted;0removes the cap)policy.rules[].whensets at most one tool matcherpolicy.rules[].when.directionis"","client_to_server","server_to_client", or absentpolicy.rules[].when.claimskeys are limited togroups,groups_in,scopes_in, andsubject_inpolicy.rules[].when.tool_globparses as a valid globpolicy.rules[].when.tool_regexcompiles as Go regexppolicy.rules[].when.tool_name_inis non-empty when setpolicy.rules[].redact[].regexcompiles as Go regexp (RE2 syntax)policy.rules[].redact[].jsonpathis not set (rejected; reserved for a future release)policy.rules[].tokens_per_secondis> 0foraction: rate_limitpolicy.rules[].burstis> 0foraction: rate_limitaudit.pathis writable by the gateway userlicense.pathexists, is regular, and has mode0600or stricter
Related
- Tutorial 1 — first traced call — minimal config in context
- Policy rule schema — the
policy:block in detail - Explanation: architecture — what the config drives