Modern Webhook Engineering: Beyond Tunnels, Race Conditions, and Flaky Pipelines
Learn how to build deterministic, offline webhook test suites in CI/CD pipelines using local wiremocking and mock receivers—no public tunnels required.

Modern Webhook Engineering: Beyond Tunnels, Race Conditions, and Flaky Pipelines
Webhooks are the standard way for payment providers, SaaS platforms, and job systems to tell your application that something happened. Reliable ingestion matters, but the testing and local-development workflows around webhooks are often fragile.
Public tunnels such as ngrok are excellent for a quick demo. They are a poor foundation for automated pipelines, because they add an external dependency, a public ingress path, and usage quotas to every run. Replaying production-shaped payloads against a local container causes a different class of problems: state pollution, stale signatures, and races with hot reloads.
This guide covers three things:
- Deterministic webhook testing in CI without tunnels.
- Correct HMAC signing and verification in tests.
- Isolated replay and queue-based processing in local development.
Part 1: Beyond Localhost Tunnels: Testing Webhooks in CI/CD
Where tunnels hurt in automation
A tunnel exposes a local port through a public URL, which is exactly what you want when a vendor's sandbox must reach your laptop. In CI the same property causes problems:
- Quotas and rate limits. ngrok's free plan currently allows up to 20,000 HTTP requests per month, 1 GB of outbound data transfer, and a single assigned dev domain. A busy pipeline can exhaust these, and a quota failure looks like a flaky test. Limits change, so check the current ngrok limits.
- Shared, non-ephemeral URLs. With one fixed dev domain, parallel jobs would compete for the same public endpoint. Per-job URLs generally require a paid plan.
- External availability. Every run now depends on the tunnel provider's uptime and on routing between the provider and your runner.
- Public ingress to a build environment. Anything reachable from the internet can be probed. A webhook endpoint that verifies signatures is still an unauthenticated public surface. CI runners are often less hardened than production hosts, and you rarely need that exposure.
- Time-sensitive signatures. Many providers sign a timestamp together with the body and reject old requests. Stripe's libraries, for example, reject signatures whose timestamp is more than five minutes old by default. Slow or stalled pipelines can trip such windows, and replaying captured payloads later will always trip them.
None of this makes tunnels wrong. It means they belong in manual, exploratory work, not in the automated test path.
Shift the webhook source inside the test network
In CI you do not need a vendor to call you. You need something that produces a correctly formed, correctly signed request. Run that producer inside the same network as the application under test:
+----------------------------------------------------------------------+
| CI job / local Docker network |
| |
| +-------------------+ signed POST +-----------------------+ |
| | Test runner or | ---------------> | Webhook receiver | |
| | webhook sender | | (application) | |
| +-------------------+ +-----------+-----------+ |
| | | |
| | asserts on state | writes |
| v v |
| +-------------------+ +-----------------------+ |
| | Assertions | <--------------- | Test DB / Redis | |
| +-------------------+ reads +-----------------------+ |
+----------------------------------------------------------------------+
Three tools fit this role:
- WireMock mocks the vendor's HTTP API and can send callbacks. From version 3.1.0, webhook support is part of WireMock core and is enabled by default. Before 3.1.0 it required a separate extension.
- Prism is a mock server driven by an OpenAPI description. It is useful for mocking the vendor's request/response API, not for signing callbacks.
- Vendor CLIs such as the Stripe CLI can forward real sandbox events to a local endpoint (
stripe listen --forward-to ...) and fire test events (stripe trigger ...). This still depends on network access to the vendor, so it suits local development more than hermetic CI.
Signing and verifying correctly
Most of the "my test signature is rejected" bugs come from one mistake: signatures are computed over the raw request bytes, not over a re-serialized JSON object. If your framework parses the body and you re-stringify it, whitespace and key order can differ from what was signed. Capture the raw body and verify against that.
An Express receiver that does this:
// src/app.ts
import express from 'express';
import crypto from 'crypto';
export const app = express();
const WEBHOOK_SECRET = process.env.WEBHOOK_SECRET ?? '';
// Keep the raw bytes for signature verification.
app.use(
express.json({
verify: (req, _res, buf) => {
(req as any).rawBody = buf;
},
})
);
function verifySignature(rawBody: Buffer, header: string | undefined): boolean {
if (!header) return false;
const expected = crypto
.createHmac('sha256', WEBHOOK_SECRET)
.update(rawBody)
.digest('hex');
const a = Buffer.from(expected);
const b = Buffer.from(header);
// timingSafeEqual throws if lengths differ, so compare lengths first.
return a.length === b.length && crypto.timingSafeEqual(a, b);
}
app.post('/api/v1/webhooks/payment', (req, res) => {
const ok = verifySignature((req as any).rawBody, req.header('X-Signature-256'));
if (!ok) return res.status(401).json({ error: 'invalid signature' });
// Enqueue or process the event here (see Part 2).
return res.status(200).json({ received: true });
});
This example uses a simple scheme (HMAC-SHA256 of the body, hex-encoded, in a custom header). Real providers differ. Stripe, for instance, signs timestamp.body and sends it in a Stripe-Signature header containing the timestamp and one or more v1 signatures. Always follow your provider's documented scheme and prefer its official verification helper, such as stripe.webhooks.constructEvent, when one exists.
In-process tests: the fastest feedback loop
For most suites, skip the network and containers for the sender entirely. Call the app in-process with supertest and sign the exact string you send:
// test/integration/webhook.spec.ts
import crypto from 'crypto';
import request from 'supertest';
import { app } from '../../src/app';
const SECRET = process.env.WEBHOOK_SECRET ?? 'ci_secret_key_12345';
function sign(rawBody: string, secret: string): string {
return crypto.createHmac('sha256', secret).update(rawBody).digest('hex');
}
describe('payment webhook', () => {
it('accepts a correctly signed event', async () => {
const body = JSON.stringify({
event: 'payment.succeeded',
id: `evt_test_${crypto.randomUUID()}`,
amount: 2500,
currency: 'usd',
});
const res = await request(app)
.post('/api/v1/webhooks/payment')
.set('Content-Type', 'application/json')
.set('X-Signature-256', sign(body, SECRET))
.send(body); // send the same string that was signed
expect(res.status).toBe(200);
expect(res.body).toEqual({ received: true });
});
it('rejects a tampered body', async () => {
const body = JSON.stringify({ event: 'payment.succeeded', id: 'evt_x', amount: 1 });
const res = await request(app)
.post('/api/v1/webhooks/payment')
.set('Content-Type', 'application/json')
.set('X-Signature-256', sign(body, SECRET))
.send(body.replace('"amount":1', '"amount":999999'));
expect(res.status).toBe(401);
});
});
Two details make this deterministic. The ID is a random UUID rather than a timestamp, so parallel tests never collide. The signed string and the sent string are the same value, so no serialization difference can break verification.
Also add tests for the failure modes that matter in production: duplicate delivery of the same event ID (the handler must be idempotent), out-of-order events, and a timestamp outside your tolerance window if your scheme uses one.
Containerized mock: WireMock as the vendor
Use a containerized mock when you need to test a flow in which your application calls the vendor API and the vendor later calls back. Define it in docker-compose.ci.yml:
services:
app-receiver:
build: .
environment:
NODE_ENV: test
WEBHOOK_SECRET: ci_secret_key_12345
DATABASE_URL: postgres://user:pass@postgres:5432/testdb
VENDOR_API_URL: http://mock-webhook-provider:8080
depends_on:
postgres:
condition: service_healthy
mock-webhook-provider:
condition: service_started
mock-webhook-provider:
image: wiremock/wiremock:3.13.2
volumes:
- ./wiremock/mappings:/home/wiremock/mappings
postgres:
image: postgres:16-alpine
environment:
POSTGRES_USER: user
POSTGRES_PASSWORD: pass
POSTGRES_DB: testdb
healthcheck:
test: ["CMD-SHELL", "pg_isready -U user -d testdb"]
interval: 2s
timeout: 3s
retries: 15
A few notes on this file:
- The top-level
version:key is obsolete in current Docker Compose and can be omitted. - No ports are published, so nothing is exposed on the host. Containers on the same Compose network reach each other by service name.
- Pin the WireMock image tag. At the time of writing,
3.13.2is the 3.x release shown on the WireMock download page. A WireMock 4.x beta line also exists, but it is labeled beta and expected to contain breaking changes.
WireMock fires the callback through a serve event listener named webhook. Put this in ./wiremock/mappings/payment_success.json:
{
"request": {
"method": "POST",
"url": "/mock-vendor/checkout"
},
"response": {
"status": 200,
"headers": { "Content-Type": "application/json" },
"jsonBody": {
"status": "initiated",
"checkout_id": "chk_99887766"
}
},
"serveEventListeners": [
{
"name": "webhook",
"parameters": {
"method": "POST",
"url": "http://app-receiver:3000/api/v1/webhooks/payment",
"headers": {
"Content-Type": "application/json",
"X-Signature-256": "14caa87134ad8255a2b2337f38d95738842e4b35021cff21d664203cbccdfe29"
},
"body": "{\"event\":\"payment.succeeded\",\"id\":\"evt_123\",\"amount\":4999,\"customer_id\":\"cust_abc\"}"
}
}
]
}
Several points matter here:
- WireMock has no built-in HMAC helper. Its templating ships many helpers (JSON/XML, dates, random values,
base64, and others), but none that computes an HMAC. Webhook parameters can be templated, but a signature has to come from somewhere else. - The signature above is precomputed. It is the HMAC-SHA256 of the exact
bodystring using the secretci_secret_key_12345. You can reproduce it withprintf '%s' '<body>' | openssl dgst -sha256 -hmac 'ci_secret_key_12345'. - This only works for schemes that sign the body alone. If your provider signs a timestamp too, a static signature cannot work. In that case generate the request in your test code (as in the in-process example) or register a custom WireMock template helper through its extension mechanism.
- Older tutorials differ. Many online examples use
postServeActionsand the separate webhooks extension from WireMock 2.x. The current documentation usesserveEventListeners. - Template data comes from
originalRequest. If you template the callback from the triggering request, referenceoriginalRequest(for example{{jsonPath originalRequest.body '$.transactionId'}}), notrequest. - Callbacks are asynchronous. WireMock fires them after the triggering response, so tests should poll or wait for the expected state rather than asserting immediately.
Part 2: Developing Webhook Receivers with Hot Reload
Local development brings its own problems. You run the receiver under a watcher such as nodemon (Node.js), air (Go), or next dev, and you want realistic payloads. Replaying captured production payloads causes three recurring problems:
- State pollution. Payloads carry real IDs and keys. Replaying them inserts foreign records into your local database, collides with unique constraints, or leaves orphaned rows after a restart.
- Restarts mid-request. If a file change restarts the process while a delivery is in flight, that request fails or half-completes.
- Signature mismatches. A captured payload has a signature (and possibly a timestamp) computed for the original secret and time. It will not verify against your local secret, and a time-bound scheme rejects it as stale.
Replaying safely: a local sanitizer
A local relay can rewrite a captured payload so it is safe to replay: namespace the IDs, refresh timestamps, and sign with the local secret. The version below fails safe. It runs only when explicitly enabled, instead of whenever NODE_ENV is not production, so a misconfigured environment cannot silently accept unsigned traffic.
captured payload relay / sanitizer local receiver
+---------------------+ +---------------------------+ +----------------------+
| id: evt_prod_99 | --> | 1. namespace the ID | --> | verifies with local |
| created: 1600000000 | | 2. refresh timestamp | | secret, writes to |
| (old signature) | | 3. re-sign with dev secret| | local DB only |
+---------------------+ +---------------------------+ +----------------------+
The simplest design is a standalone replay script rather than middleware inside the receiver. That keeps test-only code out of the production code path and lets the receiver verify signatures exactly as it does in production:
// scripts/replay-webhook.ts
import crypto from 'crypto';
import fs from 'fs';
const [, , file, url = 'http://localhost:3000/api/v1/webhooks/payment'] = process.argv;
const secret = process.env.WEBHOOK_SECRET;
if (!file || !secret) {
console.error('usage: WEBHOOK_SECRET=... ts-node replay-webhook.ts <payload.json> [url]');
process.exit(1);
}
const runId = process.env.DEV_RUN_ID ?? 'local_dev';
const payload = JSON.parse(fs.readFileSync(file, 'utf8'));
// 1. Namespace the ID so replays never collide with real or previous data.
payload.original_id = payload.id;
payload.id = `${runId}_${payload.id}_${Date.now()}`;
// 2. Refresh timestamps so freshness checks pass.
if (payload.created) payload.created = Math.floor(Date.now() / 1000);
// 3. Sign the exact bytes we send, using the local dev secret.
const body = JSON.stringify(payload);
const signature = crypto.createHmac('sha256', secret).update(body).digest('hex');
const res = await fetch(url, {
method: 'POST',
headers: { 'Content-Type': 'application/json', 'X-Signature-256': signature },
body,
});
console.log(res.status, await res.text());
Two cautions apply to any replay tooling:
- Scrub personal data. Production payloads usually contain customer data. Redact or synthesize it before it lands on a laptop or in a repository.
- Never copy production secrets locally. Use a dedicated dev secret. If you must receive real vendor events locally, use the vendor's test mode and its CLI forwarder.
Surviving hot reloads: acknowledge fast, process in a worker
When a watcher restarts your process, any work still running in the HTTP handler is lost. The standard fix is the same one used in production: verify the signature, persist or enqueue the event, return a 2xx quickly, and do the real work in a separate worker. Most providers retry on timeouts or non-2xx responses, and Stripe documents that you should return a 2xx quickly and that it may deliver the same event more than once. So your processing must be idempotent regardless.
Using BullMQ (a Redis-backed queue for Node.js), use the event ID as the job ID. BullMQ ignores a second job added with a job ID that already exists, which gives you cheap deduplication:
// src/queue/webhookQueue.ts
import { Queue } from 'bullmq';
import IORedis from 'ioredis';
// BullMQ workers require maxRetriesPerRequest to be null on their connection.
export const connection = new IORedis(process.env.REDIS_URL ?? 'redis://localhost:6379', {
maxRetriesPerRequest: null,
});
export const webhookQueue = new Queue('incoming-webhooks', { connection });
export async function enqueueWebhook(event: { id: string; type: string; data: unknown }) {
// Re-adding the same jobId is ignored, so duplicate deliveries collapse into one job.
await webhookQueue.add(event.type, event, {
jobId: event.id,
attempts: 5,
backoff: { type: 'exponential', delay: 1000 },
removeOnComplete: 1000,
removeOnFail: 5000,
});
}
// src/queue/webhookWorker.ts
import { Worker } from 'bullmq';
import { connection } from './webhookQueue';
const worker = new Worker(
'incoming-webhooks',
async (job) => {
await handleWebhookBusinessLogic(job.name, job.data); // your logic, written to be idempotent
},
{ connection, concurrency: 1 } // serialize work locally to avoid races on a shared dev DB
);
async function shutdown() {
// close() waits for jobs that are currently running to finish.
await worker.close();
await connection.quit();
process.exit(0);
}
process.on('SIGTERM', shutdown);
process.on('SIGINT', shutdown);
A few operational notes:
- Setting
concurrency: 1is a local-development convenience. In production you will usually run higher concurrency and rely on idempotent handlers and database constraints instead. - Job IDs deduplicate only while the job still exists in Redis. Keep a durable record of processed event IDs (for example, a unique constraint in your database) if you need long-term idempotency.
- Docker sends
SIGTERMwhen a container stops, followed bySIGKILLafter a grace period. Make sure your process receives the signal (running it directly, or usinginit: truein Compose) and setstop_grace_periodlong enough for in-flight jobs to finish.
Approach Comparison
| Approach | Determinism | Exposure | Network dependency | State isolation | Setup effort |
|---|---|---|---|---|---|
| Public tunnel (ngrok, etc.) | Low | Public ingress to a local port | High (provider and internet) | Poor unless you add it | Very low |
| WireMock in the CI network | High | Low (internal network only) | None at test time | Good (fresh containers per run) | Moderate |
| In-process integration tests | High | None | None | Excellent (per-test data) | Low |
| Replay script plus queue | High | Low | None for local replays | Good (namespaced IDs) | Moderate |
Vendor CLI forwarding (e.g. stripe listen) | Medium | Low (outbound connection) | Yes (vendor sandbox) | Depends on your setup | Low |
Summary and Best Practices
- Keep tunnels out of automated pipelines. Use them for manual exploration and vendor sandbox demos. For CI, produce the webhook inside the test network.
- Sign and verify the raw bytes. Capture the raw request body for verification, compare signatures in constant time, and follow your provider's exact scheme, including timestamp tolerance.
- Start with in-process tests. They are the fastest and most deterministic option. Add a WireMock layer when you need to test vendor API calls and callbacks end to end.
- Know WireMock's limits. Use
serveEventListenerson WireMock 3.1 and later, template fromoriginalRequest, and remember there is no built-in HMAC helper. - Sanitize replays, and make it opt-in. Namespace IDs, refresh timestamps, re-sign with a dev secret, and scrub personal data. Never copy production secrets to a development machine.
- Acknowledge fast, process asynchronously, and be idempotent. Deduplicate by event ID, return a 2xx quickly, and let a worker with graceful shutdown do the work so hot reloads do not lose events.