Skip to main content

Reliability & Resilience

Webhooks fail. Networks are unreliable, servers go down, and deployments cause outages. These guides cover the patterns and strategies for building webhook systems that handle failure gracefully.

What makes webhooks reliable​

Webhook reliability comes from four mechanisms built around a plain HTTP POST: retries with exponential backoff and jitter, usually about eight attempts over 24 hours; a dead letter queue for deliveries that exhaust them; a stable message ID on every request so receivers can drop duplicates; and per-endpoint timeouts and circuit breakers so one failing receiver cannot delay the rest.

FailureMechanismGuide
Receiver briefly down or deployingRetries with exponential backoff and jitterWebhook retry strategies
Receiver too slow to answerShort delivery timeout, respond 2xx first, process afterWebhook timeout best practices
Same event delivered twiceIdempotent handlers keyed on the message IDIdempotency and deduplication
Retries exhaustedDead letter queue with replayDead letter queues
One endpoint failing for daysCircuit breaker that pauses delivery to itCircuit breakers
Unclear what “delivered” promisesAt-least-once semantics, stated explicitlyWebhook delivery guarantees

For the numbers to expect from a provider, see how reliable are webhooks and the concrete schedule in our retry best practices.