How to configure a shared rate-limit store (Redis)

Problem: you run mcpgw as multiple replicas behind a load balancer. Each replica keeps its own in-memory token buckets, so a policy rule of “10 req/sec” or an input_rate_limit of “50 req/sec” is actually enforced at up to N × that — one bucket per replica.

Solution: point all replicas at a shared Redis. Buckets then live in Redis and every replica enforces one global limit. A single rate_limit_store block covers both input_rate_limit and policy.rules[].action: rate_limit — you do not configure two stores.

Recipe

1. Provision a Redis (managed — ElastiCache, Memorystore — or self-hosted). A single instance is fine for modest deployments (< ~1000 RPS gateway-wide); use Redis Cluster above that.

2. Add the rate_limit_store block to every replica’s mcpgw.yaml:

rate_limit_store:
  type: redis
  url: redis://:${REDIS_PASSWORD}@redis.internal:6379/0
  key_prefix: "mcpgw:rl:"
  operation_timeout: 50ms
  fail_open: false

Use ${ENV_VAR} references for credentials — do not inline the password. Use a rediss:// URL (or tls.enabled: true with tls.ca_file) for TLS.

3. Restart the gateway (not SIGHUP). Switching the store type requires a restart so live buckets are not silently dropped. On boot the gateway logs:

INFO rate-limit store type=redis

4. Verify the global limit. With burst set low, hammer a single session across replicas and confirm the aggregate caps at tokens_per_second + burst, not N × it.

Failure behavior

If Redis becomes unreachable or a call exceeds operation_timeout, mcpgw degrades per fail_open:

  • fail_open: false (default) — deny the request (treat the bucket as empty), preserving the configured security control during an outage.
  • fail_open: true — allow the request. Set this explicitly when availability matters more than rate enforcement. The gateway logs WARN rate-limit store unreachable; degrading per fail_open (rate-limited to once/minute) and INFO rate-limit store recovered when Redis returns.

After any failure a 1-second window short-circuits further Redis calls to the fail-open/close result, so a sick Redis does not stack up timeouts on the request hot path.

No-Redis alternatives

If you do not want another stateful dependency:

  • Sticky session affinity at the LB so a session always lands on the same replica. Works, but breaks on replica scale-down and constrains your LB.
  • Accept divergence — set tokens_per_second with enough headroom that N × is still acceptable.

Pitfalls

  • operation_timeout is on the hot path. Every rate-limited request makes one Redis round-trip. Keep Redis close (same region/VPC) and the timeout tight (50ms default).
  • Share carefully. If other apps use the same Redis, keep a distinct key_prefix.
  • type changes need a restart, not SIGHUP. url/TLS changes also currently require a restart.