How to run mcpgw in HA on Kubernetes

Problem: you need mcpgw to survive a node loss and a rolling deploy without dropping requests or diverging on policy/rate-limit enforcement — and you don’t want to hand-assemble a Deployment, anti-affinity rules, a PodDisruptionBudget, drain timings, and secret mounts yourself.

Solution: deploy the bundled Helm chart at helm/mcpgw. It ships an HA-correct topology: N replicas spread across nodes, a PDB, a shared Redis rate-limit store, two-stage drain wired to terminationGracePeriodSeconds, and locked-down pod security.

Reference architecture

                         ┌──────────────────┐
                         │  Traefik (LB /   │
                         │   IngressRoute)  │
                         └────────┬─────────┘
                                  │
            ┌─────────────────────┼─────────────────────┐
            ▼                     ▼                     ▼
      ┌──────────┐          ┌──────────┐          ┌──────────┐
      │ mcpgw-0  │          │ mcpgw-1  │          │ mcpgw-2  │   (pod anti-affinity:
      └────┬─────┘          └────┬─────┘          └────┬─────┘    one per node)
           │     shared state    │                     │
           └──────────┬──────────┴──────────┬──────────┘
                      ▼                      ▼
                ┌───────────┐        ┌──────────────────┐
                │  Redis    │        │ S3 / GCS / Kafka │
                │ (rate-    │        │  (audit sinks)   │
                │  limit)   │        └──────────────────┘
                └───────────┘
                      │
                ┌─────┴───────┐
                ▼             ▼
          [upstream MCP]  [Datadog Agent]  (OTLP, per-cluster)

Without the shared Redis, each replica keeps its own rate-limit buckets, so a “10 req/sec” limit is enforced at up to N × 10. With it, the deployment enforces one global limit.

1. Install

The chart fails to render until you configure gateway authentication. Start with the API-key guide to generate and store a real key hash, or configure OAuth: enable API-key authentication / enable OAuth. Then add an explicit policy allowlist.

# Add this to the production-values.yaml that contains your real config.auth.
config:
  policy:
    default_action: deny
    rules:
      - id: allow-filesystem-read
        action: allow
        when: { tool_prefix: "fs_read" }
helm install mcpgw ./helm/mcpgw \
  --namespace mcpgw --create-namespace \
  --set license.jwt="$(cat license.jwt)" \
  --values production-values.yaml

security.allowUnauthenticated: true bypasses the chart’s render-time guard only. Use it solely for a private test deployment or when a trusted external auth proxy is the enforced boundary; it is not safe for a generally reachable Kubernetes Service or Ingress.

For production, manage the license out-of-band and reference it:

kubectl -n mcpgw create secret generic mcpgw-license --from-file=license.jwt=./license.jwt
helm install mcpgw ./helm/mcpgw -n mcpgw \
  --set license.existingSecret=mcpgw-license \
  --values production-values.yaml

The pods refuse to boot without a valid license; /readyz returns 503 until one is present.

2. Enable the shared rate-limit store

# values.yaml
rateLimitStore:
  type: redis
  redis:
    url: "redis://:${REDIS_PASSWORD}@redis.mcpgw.svc:6379/0"
    password: ""               # or existingSecret/existingSecretKey
    existingSecret: redis-auth
    existingSecretKey: password

The chart injects the password as the REDIS_PASSWORD env var and mcpgw expands ${REDIS_PASSWORD} in the URL at load, so the secret never lands in the ConfigMap. See configure a shared rate-limit store for the full backend reference.

3. Durable audit

The chart buffers the local audit log to an emptyDir (ephemeral, per-replica). For cross-replica retention configure config.audit.sinks to ship to S3, GCS, Kafka, or a SIEM webhook — see ship audit to S3 / GCS / Kafka / SIEM.

4. Front with Traefik (load balancer)

Traefik is the recommended LB for mcpgw. Enable the bundled IngressRoute:

# values.yaml
traefik:
  ingressRoute:
    enabled: true
    entryPoints: [websecure]
    host: mcpgw.example.com
    tls:
      enabled: true
      certResolver: letsencrypt   # or secretName: <your-tls-secret>
    # middlewares: [{ name: mcpgw-ipallowlist }]   # optional

This renders:

apiVersion: traefik.io/v1alpha1
kind: IngressRoute
spec:
  entryPoints: [websecure]
  routes:
    - kind: Rule
      match: "Host(`mcpgw.example.com`)"
      services:
        - name: mcpgw
          port: 7332
  tls:
    certResolver: letsencrypt

Two things matter for mcpgw specifically:

Readiness gating is automatic. Traefik routes to the Kubernetes Service, whose endpoints are readiness-gated. When the two-stage drain flips /readyz to 503 (SIGTERM) — or a pod’s license degrades to Expired — the pod leaves the endpoint set and Traefik stops routing to it. No sticky sessions are needed because rate-limit state lives in Redis, not per-pod.

Tune entrypoint timeouts for SSE. mcpgw streams long-lived SSE responses; Traefik streams them through without buffering, but the entrypoint’s responding/forwarding timeouts must not cut a slow stream. In Traefik’s static config (e.g. the Traefik Helm chart’s values):

# Traefik static configuration
entryPoints:
  websecure:
    address: ":443"
    transport:
      respondingTimeouts:
        readTimeout: "0s"        # don't cap slow SSE reads
        idleTimeout: "180s"      # > your longest idle-between-frames gap
serversTransport:
  forwardingTimeouts:
    responseHeaderTimeout: "0s"  # upstream may take time before first byte
    idleConnTimeout: "180s"

Set idleTimeout comfortably above your longest expected gap between SSE frames; readTimeout: 0s disables the slow-read cap that would otherwise truncate a long stream.

Optionally add a Traefik service health check that probes /readyz directly (belt-and-suspenders on top of endpoint gating) via the IngressRoute service’s healthCheck field or a TraefikService.

5. Rolling deploys without dropped requests

The chart wires the two-stage drain:

  • shutdown.drainSignalWindow (default 5s) — how long /readyz returns 503 (so the LB stops routing) before draining begins.
  • shutdown.grace (default 30s) — how long in-flight requests (including long SSE responses) get to finish.
  • terminationGracePeriodSeconds (default 40) — must be ≥ grace + drainSignalWindow + headroom, or the kubelet SIGKILLs the pod mid-drain. The chart’s default already satisfies this; if you raise grace, raise this too.

A kubectl rollout restart deployment/mcpgw then drains each pod cleanly. The PodDisruptionBudget (minAvailable: 2 of 3) keeps capacity up during voluntary disruptions.

6. Scale with HPA

autoscaling:
  enabled: true
  minReplicas: 3
  maxReplicas: 12
  targetCPUUtilizationPercentage: 70

With the shared Redis store, scaling out does not loosen the global rate limit — buckets live in Redis, not per-pod.

7. Monitor HA-specific signals

  • Rate-limit store health — WARN rate-limit store unreachable / INFO rate-limit store recovered log lines; alert on the former.
  • Audit completeness — compare per-replica audit-line counts against sink-delivered counts; /readyz exposes audit_sink_dropped_records when a sink buffer overflows.
  • Drain correctness — confirm /readyz flips to 503 on SIGTERM before the pod exits (no requests routed to a terminating pod).

Metrics scraping: mcpgw has no Prometheus /metrics endpoint — telemetry ships via OTLP (Datadog Agent / collector). There is no ServiceMonitor in the chart by design.