Chaos Engineering for Webhooks: How to Simulate and Test Network Failures
Chaos Engineering for Webhooks: How to Simulate and Test Network Failures Site reliability engineers routinely run chaos experiments on internal services.

Chaos Engineering for Webhooks: How to Simulate and Test Network Failures
Site reliability engineers routinely run chaos experiments on internal services. Tools like Chaos Mesh, Gremlin and LitmusChaos kill pods, sever gRPC connections and inject latency into database pools. Outbound webhooks, though, are often left out of that testing.
The reason is straightforward: a webhook leaves your trust boundary. Inside a service mesh you control timeouts, retries and mutual TLS. A webhook crosses the public internet and lands on a third party's endpoint, with unknown uptime, an unknown HTTP stack and an unknown idea of what "slow" means. When that endpoint degrades, the damage often flows back into the sender: exhausted worker pools, swollen retry queues and synchronized retry storms.
This article walks through how to test for that on purpose. It covers where webhook pipelines break, what real providers do, how to inject Layer 4 and Layer 7 faults with Toxiproxy and Chaos Mesh, what to measure, and which design patterns hold up.
How a webhook pipeline is put together
Most webhook platforms decouple event generation from delivery with an asynchronous producer-consumer design:
- Event source. Internal services (billing, auth, orders) emit state-change events.
- Ingestion and queueing. Events are written to a durable log or queue such as Kafka, RabbitMQ or SQS.
- Dispatch workers. Workers consume events, build the HTTP POST, sign it with an HMAC secret and deliver it to the subscriber's URL.
- Retry handler and dead letter queue (DLQ). Failed deliveries are re-enqueued with a delay. Events that exhaust their retries go to a DLQ.
Every stage has a failure mode, and most only show up under network stress.
Why ordinary tests miss webhook bugs
Unit and integration tests usually confirm that the dispatcher builds the right JSON and handles one mocked 200 OK. They don't reveal the systemic problems:
- Worker pool exhaustion. Slow subscribers hold sockets open until dispatcher threads or goroutines run out.
- Retry storms (thundering herd). Many endpoints fail at once, their retries line up, and the burst overloads both your queues and their servers.
- Head-of-line blocking. If events must be delivered in order, one failing payload stalls everything behind it.
- Unbounded queue growth. A long outage makes the retry queue balloon until broker memory or disk runs out.
What real providers do (and why it matters for your tests)
Your dispatcher is only half the system. The subscribers you deliver to are usually on the receiving end of another platform's rules, and those rules give you realistic targets for your experiments.
| Provider | Response timeout | Retry behavior | Notable consequence |
|---|---|---|---|
| Stripe | A few seconds; the exact value is not publicly documented | In live mode, retries for up to 3 days with exponential backoff. In test mode, 3 retries over a few hours. The exact schedule is not published. | Emails you if an endpoint hasn't returned a 2xx for multiple days, and can disable it |
| Shopify | 5 seconds | 8 retries over 4 hours | Subscriptions created through the Admin API are deleted after 8 consecutive failures |
| GitHub | 10 seconds | No automatic redelivery | Failed deliveries must be redelivered manually or by script through the REST API |
A few observations follow from these numbers:
- Timeouts are short. Shopify gives a subscriber five seconds, and its own docs tell developers to delay processing until after responding. If your dispatcher waits 60 seconds for a slow endpoint, you are far more patient than the systems your customers depend on.
- Retry windows differ wildly. Three days at Stripe, four hours at Shopify and none at GitHub means "how long do we keep trying?" has no universal answer. Your chaos experiments should include outages longer than your retry window to prove events end up in the DLQ or replay buffer instead of vanishing.
- Consequences are real. Shopify deletes subscriptions after persistent failures, which is why a flash sale that overloads a backend can silently break an integration.
Note that GitHub Enterprise Server documents a 30-second timeout rather than 10, so check the docs for the specific product you are targeting.
Failure modes to inject
Test at both Layer 4 (transport) and Layer 7 (application). Each fault stresses a different part of the pipeline.
| Failure mode | Layer | Toxiproxy toxic | What it exercises |
|---|---|---|---|
| Tail latency and jitter | L4 | latency | Client timeouts, worker pool sizing |
| Silent hang (accepts connection, never answers) | L4 | timeout with timeout=0 | Read deadlines, pool isolation |
| Connection reset mid-flight | L4 | reset_peer | Retry on RST, stale keep-alive reuse |
| Truncated response or request | L4 | limit_data | Partial-read handling |
| Slow reads and writes (Slowloris-style) | L4 | bandwidth, slicer | Per-request deadlines |
| Random data loss | L4 | packet_loss | Retries, corruption handling |
| 5xx responses | L7 | (mock receiver) | Retry classification, backoff |
| Full outage | L4 | proxy disabled | Circuit breakers, replay buffer |
1. Tail latency
Subscribers under database pressure often accept the TCP connection and then take 15 to 45 seconds to answer. The vulnerability is in your outbound client: without strict per-request timeouts, worker threads pile up and block healthy traffic.
2. Transient 5xx responses
Reverse proxies, load balancers and restarting pods return 502, 503 and 504. The vulnerability is retry classification. Providers such as Stripe and Shopify treat any non-2xx response as a failed delivery, but your dispatcher can and should be smarter: 503 and 504 are usually transient, 429 should respect Retry-After, and most other 4xx codes are probably permanent.
3. Connection drops and resets
Connections get reset, packets get lost and sockets hang silently. The vulnerability is connection pooling: a long-idle pooled connection may be reused after the far side has already dropped it.
4. Slow reads
A subscriber can accept a request and then read or answer at a crawl. The vulnerability is missing deadlines: you need a total request deadline, not just a connect timeout.
Hands-on: Layer 4 chaos with Toxiproxy
Toxiproxy, built by Shopify, is a TCP proxy for simulating network conditions. It is designed for testing, CI and development, needs no root access and exposes its control API over HTTP on port 8474. That makes it a good fit for local development and staging integration tests.
Toxics apply to one direction of a connection. downstream (the default) affects the server-to-client link, which is the subscriber's response. upstream affects client-to-server, which is your request. Each toxic also has a toxicity value, the probability that it applies to a given connection, defaulting to 1.0. That lets you say "affect 30% of connections" without extra tooling.
Step 1: put a proxy in front of the subscriber
Assume a test receiver at subscriber.local:8080:
# Start the Toxiproxy server (its HTTP API listens on 8474)
toxiproxy-server &
# Create a proxy that listens on 8666 and forwards to the subscriber
toxiproxy-cli create -l 0.0.0.0:8666 -u subscriber.local:8080 webhook-receiver
Point your dispatcher at http://localhost:8666 instead of the real subscriber URL. Toxiproxy's docs recommend keeping proxy ports outside the Linux ephemeral range (32,768 to 61,000 by default) to avoid intermittent connection failures.
Step 2: inject latency
Delay every response by 12 seconds, plus or minus 3 seconds of jitter:
toxiproxy-cli toxic add -t latency -a latency=12000 -a jitter=3000 webhook-receiver
To affect only 30% of connections, use the HTTP API and set toxicity:
curl -X POST http://localhost:8474/proxies/webhook-receiver/toxics \
-H "Content-Type: application/json" \
-d '{
"name": "slow-30pct",
"type": "latency",
"stream": "downstream",
"toxicity": 0.3,
"attributes": { "latency": 15000, "jitter": 0 }
}'
Step 3: hang, reset, truncate and throttle
# Silent hang: swallow data and never close (timeout=0 means the connection stays open)
toxiproxy-cli toxic add -t timeout -a timeout=0 webhook-receiver
# Connection reset (TCP RST) immediately
toxiproxy-cli toxic add -t reset_peer -a timeout=0 webhook-receiver
# Close the connection after 512 bytes to simulate a truncated body
toxiproxy-cli toxic add -t limit_data -a bytes=512 webhook-receiver
# Throttle the link to 2 KB/s to trigger slow read/write deadlines
toxiproxy-cli toxic add -t bandwidth -a rate=2 webhook-receiver
Toxiproxy also has a packet_loss toxic, which the project describes as randomly dropping chunks flowing through the proxy:
curl -X POST http://localhost:8474/proxies/webhook-receiver/toxics \
-H "Content-Type: application/json" \
-d '{
"name": "flaky-network",
"type": "packet_loss",
"attributes": { "loss_rate": 0.25, "correlation": 0.5 }
}'
Two caveats. First, this toxic is documented in the project's README on the main branch, so confirm it exists in your installed version by calling GET /version and trying it. Second, because Toxiproxy sits at the TCP stream level, it drops chunks of a byte stream rather than IP packets, and it never sees TCP retransmission. If you want true packet loss with kernel-level retransmits, use tc netem or Chaos Mesh's NetworkChaos (covered below).
Step 4: verify, then clean up
Send a batch of, say, 1,000 webhooks while the toxics are active and watch your dispatcher:
- Does the HTTP client abort at your configured limit (for example 5 seconds)?
- Do active workers stay bounded, or does the process climb toward an out-of-memory kill?
- Do healthy subscribers keep receiving events at normal speed?
To remove a single toxic, or reset everything:
toxiproxy-cli toxic remove -n latency_downstream webhook-receiver
# Re-enable all proxies and remove all toxics
curl -X POST http://localhost:8474/reset
Simulating a full outage
Bringing a service down is not a toxic. Toxiproxy does it by disabling the proxy:
curl -X POST http://localhost:8474/proxies/webhook-receiver \
-H "Content-Type: application/json" -d '{"enabled": false}'
Set enabled back to true to restore it. Run this longer than your retry window to verify that events land in the DLQ or replay buffer.
Hands-on: Layer 7 chaos on Kubernetes with Chaos Mesh
Chaos Mesh is a CNCF incubating project, and its 2.8.x documentation is current at the time of writing. Two of its fault types are relevant here: HTTPChaos for request and response faults, and NetworkChaos for kernel-level network faults.
Know HTTPChaos's limits before you use it
Chaos Mesh's own documentation lists several constraints that matter for webhook testing:
- HTTPS is not supported. Injection into HTTPS connections doesn't work, so HTTPChaos is for plain-HTTP staging receivers, not for real third-party HTTPS endpoints.
- New connections only. Requests sent over a TCP connection established before the experiment starts are not affected, so clients that reuse connections may appear to ignore the fault.
- Both sides by default. Rules apply to clients and servers in the selected pod unless you restrict the side.
- Be careful with POST. The docs warn that non-idempotent requests, which includes most POSTs, may not recover just by retrying after injection.
Also note three details that commonly trip people up:
abort: trueinterrupts the connection. It does not return a 503.- The top-level
codefield only selects responses that have a given status. It doesn't set one. - The old
schedulerfield is gone. Recurring experiments now use the separateScheduleresource.
Example: abort connections to a staging receiver on a schedule
Target the mock subscriber pod, not the dispatcher:
apiVersion: chaos-mesh.org/v1alpha1
kind: Schedule
metadata:
name: webhook-receiver-abort
namespace: chaos-testing
spec:
schedule: '@every 30m'
type: HTTPChaos
historyLimit: 2
concurrencyPolicy: Forbid
httpChaos:
mode: all
selector:
namespaces:
- staging-subscribers
labelSelectors:
app: mock-subscriber
target: Request
port: 8080
method: POST
path: /webhooks/*
abort: true
duration: 10m
Example: slow responses on half the receiver pods
The mode field supports one, all, fixed, fixed-percent and random-max-percent. Percentages apply to pods, not requests, so use fixed-percent:
apiVersion: chaos-mesh.org/v1alpha1
kind: HTTPChaos
metadata:
name: webhook-receiver-slow
namespace: chaos-testing
spec:
mode: fixed-percent
value: '50'
selector:
namespaces:
- staging-subscribers
labelSelectors:
app: mock-subscriber
target: Response
port: 8080
path: /webhooks/*
delay: 8s
duration: 15m
Example: real packet loss with NetworkChaos
NetworkChaos supports partition, network emulation (delay, loss, reordering, corruption) and bandwidth limits. The netem-based faults require the NET_SCH_NETEM kernel module, which most mainstream Linux distributions include by default. Also make sure the connection between Chaos Mesh's controller manager and its chaos daemon is healthy, or injected faults can't be reverted.
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: webhook-egress-loss
namespace: chaos-testing
spec:
action: loss
mode: all
selector:
namespaces:
- staging-events
labelSelectors:
app: webhook-dispatcher
loss:
loss: '30'
correlation: '25'
duration: '10m'
Returning real 5xx responses: use a mock receiver
For deterministic 503 and Retry-After testing, the simplest approach is a receiver you control. This standard-library Python server fails a configurable fraction of requests:
import os
import random
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
FAIL_RATE = float(os.getenv("FAIL_RATE", "0.8")) # share of requests that get a 503
DELAY_S = float(os.getenv("DELAY_S", "0")) # artificial response delay
class Handler(BaseHTTPRequestHandler):
def do_POST(self):
self.rfile.read(int(self.headers.get("Content-Length", 0)))
time.sleep(DELAY_S)
if random.random() < FAIL_RATE:
self.send_response(503)
self.send_header("Retry-After", "30")
else:
self.send_response(200)
self.end_headers()
ThreadingHTTPServer(("0.0.0.0", 8080), Handler).serve_forever()
This also lets you check whether your dispatcher actually honors Retry-After, something that's easy to assume and rarely verified.
What to observe while the chaos runs
Watch four signals.
1. Worker pool utilization. If latency on one subscriber pushes pool usage to 100% across all subscribers, you lack endpoint isolation.
worker_pool_utilization = active_workers / max_workers * 100
2. Queue depth and backpressure. When delivery success drops from 99.9% to 40%, events must buffer safely.
queue_growth_rate = ingestion_rate - successful_delivery_rate
If that growth threatens broker disk or memory, confirm that backpressure or cold-storage offload actually kicks in.
3. Retry amplification factor. The ratio of delivery attempts to unique events:
retry_factor = total_http_post_attempts / unique_event_ids_ingested
A healthy backoff keeps this stable during an outage. If it climbs sharply, you have a retry storm.
4. Dispatcher tail latency. Track P99 delivery time per destination, not just globally. A global average hides a single bad subscriber.
Design patterns that survive the experiments
Pattern 1: an explicit timeout hierarchy
Never rely on default HTTP client settings. Set layered timeouts. In Go, note the distinction between the overall deadline and the header wait: Client.Timeout caps the entire exchange including reading the body, while ResponseHeaderTimeout only covers waiting for response headers once the request has been written.
var webhookClient = &http.Client{
Timeout: 10 * time.Second, // total request deadline, including body read
Transport: &http.Transport{
DialContext: (&net.Dialer{
Timeout: 2 * time.Second, // TCP connect
KeepAlive: 30 * time.Second,
}).DialContext,
TLSHandshakeTimeout: 3 * time.Second,
ResponseHeaderTimeout: 4 * time.Second, // wait for headers after request is sent
ExpectContinueTimeout: 1 * time.Second,
MaxIdleConns: 1000,
MaxIdleConnsPerHost: 10,
IdleConnTimeout: 90 * time.Second,
},
}
Choose values with the provider limits above in mind. If subscribers on Shopify-like platforms are expected to answer within five seconds, a 10-second dispatcher ceiling is already generous.
Pattern 2: exponential backoff with full jitter
Retrying immediately makes an outage worse, and plain exponential backoff still leaves clusters of synchronized retries. The AWS Architecture Blog's analysis of contended clients found that adding jitter removes those clusters. In its simulation with 100 contending clients, jitter cut the total number of calls by more than half and improved completion time. Full jitter picks a random delay between zero and the capped exponential value:
sleep = random_between(0, min(cap, base * 2^attempt))
import random
def full_jitter_backoff(attempt: int, base: float = 1.0, cap: float = 300.0) -> float:
"""Exponential backoff with full jitter (AWS Architecture Blog)."""
ceiling = min(cap, base * (2 ** attempt))
return random.uniform(0, ceiling)
for attempt in range(1, 6):
print(f"attempt {attempt}: wait {full_jitter_backoff(attempt):.2f}s")
Pattern 3: per-destination circuit breakers and bulkheads
If one subscriber fails continuously, further attempts waste dispatcher capacity and clutter logs. Track health per destination:
- Closed: normal delivery.
- Open: the failure rate crosses a threshold (for example over 50% across 5 minutes). Skip network calls and divert events to a persistent replay buffer.
- Half-open: after a cooldown, send a single probe. On success, close the breaker.
Pair the breaker with a per-destination concurrency cap (a bulkhead) so one slow endpoint can never consume the whole worker pool. That single control is what most often turns "one subscriber is slow" into "nobody receives anything."
stateDiagram-v2
[*] --> Closed
Closed --> Open: failure rate above threshold
Open --> HalfOpen: cooldown expires
HalfOpen --> Closed: probe succeeds
HalfOpen --> Open: probe fails
Pattern 4: replay buffers and at-least-once delivery
When a connection breaks after the subscriber processed the request but before your dispatcher saw the response, you'll retry, and the subscriber will see the event twice. Delivery is therefore at-least-once, so design for duplicates:
- Replay buffers. Keep undelivered events in append-only storage (S3, Postgres, RocksDB) so they can be replayed in bulk after recovery.
- Stable event IDs. Give every event an ID that stays the same across retries, so subscribers can deduplicate.
- Signed timestamps. Include a timestamp in the signed content so subscribers can reject replayed requests.
The Standard Webhooks specification is a good template. It uses three headers: webhook-id, webhook-timestamp and webhook-signature. The signature is an HMAC-SHA256, base64-encoded and prefixed with a version tag like v1,. It is computed over {webhook-id}.{webhook-timestamp}.{raw body}. Implementations describe the ID as consistent across retries of one delivery, which makes it usable for deduplication, while the timestamp changes on every attempt.
POST /hooks/v1/payment-events HTTP/1.1
Host: api.customer.com
Content-Type: application/json
webhook-id: evt_9f8d7a6b5c4d3e2a
webhook-timestamp: 1774684800
webhook-signature: v1,K5oZfzN95Z9UVu1EsfQmfVNQhnkZ2pj9o9NDN/H/pI4=
{"event":"payment.succeeded","amount":4900,"currency":"usd"}
Subscribers must compute the signature over the raw, unparsed body. If a framework parses the JSON first, the signature usually won't match.
A verification checklist for CI and staging
| Experiment | Tooling | Injection | Pass criteria |
|---|---|---|---|
| Sustained 5xx | Mock receiver | 80% of requests return 503 for 15 minutes | Backoff engages, retry factor stays bounded, healthy subscribers unaffected |
| Severe latency | Toxiproxy latency | +15 s on 30% of connections (toxicity: 0.3) | Client timeout enforced, worker use stays under your limit, no cross-subscriber impact |
| Silent hang | Toxiproxy timeout (0) | Connection accepted, no data | Deadlines fire, sockets are released |
| Truncated body | Toxiproxy limit_data | Close after N bytes | Attempt counted as failed and retried, no partial state committed |
| Data loss | Chaos Mesh NetworkChaos (loss) | 30% loss on dispatcher egress | Retries recover events, no duplicate side effects |
| Full outage | Toxiproxy (proxy disabled) | Cut for longer than your retry window | Breaker opens, events reach replay buffer or DLQ with no loss |
Choose thresholds from your own SLOs. The numbers above are starting points, not universal standards.
Conclusion
Webhooks are the part of an event-driven system you control least and test least. The fix is not exotic: inject the failures that real subscribers produce, measure how your dispatcher's pools, queues and retries respond, and harden what breaks. Start with latency and a full outage, since those expose most isolation and retry-window problems, then add resets, truncation and packet loss.
Sources
- Stripe, Receive Stripe events in your webhook endpoint: retry behavior, test versus live mode, and endpoint disabling
- Shopify, Deliver webhooks through HTTPS and Troubleshoot webhooks: 5-second timeout, 8 retries over 4 hours, subscription removal
- GitHub, Handling failed webhook deliveries: 10-second timeout, no automatic redelivery (30 seconds on GitHub Enterprise Server)
- Toxiproxy, README and toxic reference
- Chaos Mesh, Simulate HTTP Faults, Simulate Network Faults and Define Scheduling Rules
- AWS Architecture Blog, Exponential Backoff and Jitter
- Standard Webhooks, specification and libraries