Health endpoints reference

mcpgw exposes two endpoints for orchestrators. They serve distinct purposes — do not point both your liveness and readiness probes at the same one.

Endpoint Probe type Returns 503 when…
/healthz Liveness Never (a hung process answering 200 is the only failure mode)
/readyz Readiness License is invalid or expired beyond grace, or the gateway is draining on shutdown

/healthz

GET /healthz
Response Meaning
200 OK body ok Process is responsive
no response Process is dead, hung, or the network is broken — your liveness probe should restart

Use as your liveness probe. Kubernetes example:

livenessProbe:
  httpGet:
    path: /healthz
    port: 7332
  initialDelaySeconds: 5
  periodSeconds: 10
  failureThreshold: 3

/healthz deliberately does not check the license, the upstreams, or the audit file writability. Its only job is to prove the process is responsive. Coupling liveness to license state turns a slow license-renewal incident into a thundering-herd restart that makes everything worse.


/readyz

GET /readyz
Response Meaning
200 OK, JSON body License is valid or in grace (within exp + grace_days)
503 Service Unavailable, empty body The gateway is draining (SIGTERM received — stage 1 of graceful shutdown), or the license has soft-degraded to Expired (past the grace window, or the hourly re-verification failed: bad signature, unreadable file)

The 200 body is {"status": "ok"} plus two diagnostic fields, each present only when non-zero:

Field Type Meaning
restart_pending string array A SIGHUP reload changed fields the gateway cannot hot-swap (listen, upstreams, routes, …); the running process still uses the old values for the listed fields. Clears once a reload no longer diverges from the applied config.
audit_sink_dropped_records number Cumulative count of audit records dropped because a remote sink buffer was full. Non-zero means delivery to at least one sink is lossy — alert on this if remote audit shipping is compliance-relevant.

Use as your readiness probe. When /readyz returns 503, the orchestrator should remove the pod from the LB rotation but not restart it.

readinessProbe:
  httpGet:
    path: /readyz
    port: 7332
  initialDelaySeconds: 2
  periodSeconds: 30
  failureThreshold: 2

/readyz does not check upstream health. An upstream MCP server being down does not make the gateway “not ready” — the gateway can still receive requests and return useful errors (502 upstream_unreachable). Coupling readiness to upstream health collapses both layers’ availability into one number.


Graceful shutdown drain

On SIGTERM/SIGINT, mcpgw drains in two stages so a rolling restart does not drop in-flight requests — including long-lived SSE responses:

  1. Signal not-ready. /readyz immediately returns 503. The LB’s next health check removes the replica from rotation and stops sending new traffic; in-flight requests keep being served.
  2. Drain. After shutdown.drain_signal_window, the gateway stops accepting new connections and waits up to shutdown.grace for in-flight requests to complete before exiting.
shutdown:
  grace: 30s                # max wait for in-flight requests (incl. SSE). Default 30s.
  drain_signal_window: 5s   # how long /readyz returns 503 before draining begins. Default 0.

Defaults (grace: 30s, drain_signal_window: 0) preserve the historical single-stage behavior. Behind a load balancer, set drain_signal_window to a few seconds so the LB observes the 503 before connections are cut.

Kubernetes: set terminationGracePeriodSeconds to at least grace + drain_signal_window plus headroom, or the kubelet SIGKILLs the pod mid-drain:

spec:
  terminationGracePeriodSeconds: 40   # >= grace(30) + drain_signal_window(5) + headroom

Why split these?

The Kubernetes-style liveness/readiness split exists for a reason: they correspond to different remediations.

  • Liveness fail = restart me. The process is broken and only a restart will fix it.
  • Readiness fail = remove me from rotation. The process is fine but currently can’t serve traffic correctly.

License expiry is a readiness failure: you do not want the orchestrator to restart-loop a binary while you are renewing the JWT. Hung process is a liveness failure: restart immediately.

If your orchestrator only supports one probe type, use /readyz. The downside is slightly slower recovery from process hangs (one fewer fast-fail signal); the upside is correct behavior on license expiry.


Fronting /readyz with a load balancer

LB health checks are typically configured as readiness probes. Point them at /readyz. If you need the LB to detect process death faster, use a TCP-only check on :7332 in addition.

/readyz is excluded from rate-limit identity and policy evaluation. It does not appear in the audit log or in OTLP spans. This is intentional — health-check traffic should not pollute observability data.