How to rate-limit a tool

Problem: a tool is expensive, dangerous in volume, or both — fs_write, db_query, send_email. You want to cap how often any single client can fire it.

Solution: add an action: rate_limit rule. mcpgw uses a token-bucket per (rule, client_ip) keyed off the TCP peer address (RemoteAddr).

Recipe

policy:
  rules:
    - id: rl-fs-write
      action: rate_limit
      when: { tool_name: "fs_write" }
      tokens_per_second: 10        # steady-state allowance
      burst: 20                    # peak burst before the bucket starts gating

tokens_per_second is the refill rate, burst is the bucket capacity. A request consumes one token. When the bucket is empty, requests are rejected with HTTP 429 and JSON-RPC error -32003 rate_limited until enough time has passed for tokens to refill.

This is separate from top-level input_rate_limit, which protects the gateway before body read and parse. Use input_rate_limit for request-flood protection; use policy rate_limit rules for per-tool quotas.

Picking values

  • One write every few seconds: tokens_per_second: 0.2, burst: 3
  • Modest API quota (e.g. 60/minute): tokens_per_second: 1, burst: 10
  • High-throughput protection only: tokens_per_second: 100, burst: 200

tokens_per_second accepts fractional values. 0.0001 is one token every ~2.7 hours — useful for tool calls you only want to allow a few times per day.

Per-(rule, client IP) semantics

Two clients calling the same rate-limited tool from different source IPs have independent buckets. A burst from client-A does not consume client-B’s allowance. This means:

  • Each connecting client gets its own quota.
  • A single misbehaving client cannot starve others (unless they share a source IP — e.g. behind a NAT or LB).
  • Total system throughput is tokens_per_second × number_of_distinct_client_ips, not a global cap.

Behind a load balancer: without trusted_proxies, all anonymous clients behind the same LB share one IP bucket. Configure the LB source CIDR to resolve RFC Forwarded/X-Forwarded-For safely. Authenticated callers already use their verified principal. See Explanation: rate-limit identity.

Verifying

for i in $(seq 1 25); do
  curl -s -o /dev/null -w "%{http_code}\n" \
    -X POST http://localhost:7332/mcp \
    -H "Content-Type: application/json" \
    -d "{\"jsonrpc\":\"2.0\",\"id\":$i,\"method\":\"tools/call\",\"params\":{\"name\":\"fs_write\",\"arguments\":{\"path\":\"/tmp/x\",\"content\":\"y\"}}}"
done

You should see 200 for the first burst requests, then a stream of 429s as the bucket drains.

Pitfalls

  • Anonymous buckets use the effective client IP. This is RemoteAddr by default, or the first untrusted address in a forwarding chain when the direct peer matches trusted_proxies. Mcp-Session-Id does not affect the bucket.
  • Never trust forwarding headers globally. Use the narrow ingress/LB CIDR and matching network access controls. See Explanation: rate-limit identity.
  • Rate-limit denial counts as policy refusal, not as a transport error. Client retry-with-backoff logic should treat HTTP 429 specifically.
  • First match wins. A wildcard redact rule placed above your rate-limit rule will redact-and-forward instead of throttling. Order matters.