Async Event Tracking: Collect Analytics Without Slowing Down Your App
Async event tracking keeps analytics off the critical path. Build client batching, retries, a server buffer, backpressure and a graceful shutdown flush.
The fastest way to make your own analytics hurt your product is to do the work inside the request. A tracking call that waits for a database insert adds database latency to every page. When the database slows down, your checkout slows down with it. Async event tracking breaks that link, so that collecting an event never costs the user a single frame or the server a single blocked request.
Asynchronous does not mean careless. Moving work off the critical path creates new failure modes: events lost on page unload, duplicates from retries, a server buffer that grows until the process dies, and a deploy that throws away the last second of data. Each one is real, and each one has a known fix.
In this article you build the non-blocking path end to end for Acme Shop. The browser batches events and retries safely. The Express collector acknowledges fast and buffers in memory. A flusher writes batches to PostgreSQL with one statement, and a shutdown hook drains the buffer before the process exits. The table is the one defined in PostgreSQL schema design for analytics, and the endpoint is POST /v1/events from the earlier collector article.
You will also learn where this design stops being enough. An in-process buffer is the right first step, and I will show you the signals that tell you it is time for a real queue.
Why Synchronous Tracking Fails
Consider the naive collector. It receives an event, runs an INSERT, waits for the database, and then answers. Under light load that works. Under a traffic spike, three things go wrong at once.
First, each request holds a database connection while it waits, so the connection pool empties. Second, requests queue behind the pool, so response times climb. Third, browsers time out and retry, which doubles the load at the worst moment. A tracking endpoint that was meant to be invisible becomes the outage.
The browser side has a mirror problem. A tracker that awaits the response before letting navigation continue delays the next page. A tracker that fires and forgets, on the other hand, loses events when the page unloads. You need a design that is fast and reliable, and the two goals pull in different directions.
The wrong collector
// WRONG: the request waits for the database
app.post('/v1/events', async (req, res) => {
await pool.query(
'INSERT INTO events (event_id, event_name, occurred_at, anonymous_id, session_id, properties) VALUES ($1, $2, $3, $4, $5, $6)',
[req.body.event_id, req.body.event_name, req.body.occurred_at,
req.body.anonymous_id, req.body.session_id, req.body.properties]
);
res.status(201).json({ ok: true });
});
This handler has no validation, one round trip per event, no batching, and latency tied to the database. It also returns 201 only after the insert, so a slow insert means a slow user.
The right shape
The right collector validates, enqueues in memory and answers right away. A separate loop persists. The delta is that response time no longer depends on database time, and one multi-row insert replaces hundreds of single inserts. The rest of this article builds that shape piece by piece.
The Architecture in One Picture
BROWSER SERVER (one Node.js process)
------- ----------------------------
track() --> queue (memory)
| every 5 s, 20 events, or page hide
v
POST /v1/events ------> validate --> buffer (bounded)
(batch array) | |
^ | 202 Accepted | every 1 s or 500 events
| v v
retry with backoff full? 503 + INSERT ... unnest(...)
same event_ids Retry-After ON CONFLICT DO NOTHING
|
PostgreSQL events
SIGTERM: stop accepting, drain buffer, exit
Two queues exist, one in each tier, and each has a job. The browser queue protects the user’s page from the network. The server buffer protects the database from bursts. Both are bounded, and both rely on event_id for safety.
Client Batching and Retries
Capturing events, listeners and delegation belong to the capture article. The tracker below keeps a minimal track(name, properties) so it runs on its own, and focuses on how events leave the browser. It reuses the 30 minute session rule from that article.
The batching rules
Send a batch when any of three things happens: the queue reaches 20 events, five seconds pass since the first queued event, or the page becomes hidden. Those numbers are starting points, not laws. Larger batches use fewer requests but widen the loss window. Smaller batches do the opposite.
Generate event_id when the event is created, not when it is sent. That single decision is what makes retries safe. If the network drops the response after the server stored the batch, the retry carries the same identifiers, and PostgreSQL discards the copies.
Which failures to retry
Retry on network errors, on 429 and on any 5xx status. Do not retry on 400, 401, 403 or 413, because the same payload will fail again. Drop those events and log a counter, or you will build a loop that never ends. For 429 and 503, honour the Retry-After header defined in RFC 9110. Otherwise use exponential backoff with jitter, so a thousand clients that failed together do not all return together.
Sending on page hide
The page lifecycle is where analytics loses the most events. The unload event is unreliable, especially on mobile, and it blocks the back/forward cache. MDN recommends sending final data on visibilitychange when the state becomes hidden, using navigator.sendBeacon(). The browser queues the request and delivers it after the page is gone. MDN documents a 64 KiB limit on the total queued data, so keep each beacon batch well below it.
One cross-origin detail catches people. A beacon with a JSON content type can trigger a CORS preflight, and a beacon cannot wait for one reliably. Send the batch as a text/plain Blob and let the server parse it as JSON. The collector below accepts both content types.
The complete client
// tracker.js - vanilla ES2022, no dependencies
const ENDPOINT = 'https://analytics.acme-shop.example/v1/events';
const MAX_BATCH = 20;
const FLUSH_MS = 5000;
const MAX_QUEUE = 500;
const MAX_ATTEMPTS = 6;
let queue = [];
let timer = null;
let sending = false;
const anonymousId = localStorage.getItem('aid') ?? crypto.randomUUID();
localStorage.setItem('aid', anonymousId);
const SESSION_MS = 30 * 60 * 1000;
function getSessionId() {
const now = Date.now();
let state = null;
try { state = JSON.parse(localStorage.getItem('session_state')); } catch { /* ignore */ }
if (!state || now - state.last > SESSION_MS) {
state = { session_id: crypto.randomUUID(), last: now };
}
state.last = now;
localStorage.setItem('session_state', JSON.stringify(state));
return state.session_id;
}
export function track(eventName, properties = {}) {
if (queue.length >= MAX_QUEUE) queue.shift(); // drop oldest, never grow forever
queue.push({
event_id: crypto.randomUUID(), // created once, reused on retry
event_name: eventName,
occurred_at: new Date().toISOString(),
anonymous_id: anonymousId,
user_id: null,
session_id: getSessionId(),
page_url: location.href,
referrer: document.referrer || null,
user_agent: navigator.userAgent,
properties,
});
if (queue.length >= MAX_BATCH) flush();
else if (!timer) timer = setTimeout(flush, FLUSH_MS);
}
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
async function flush() {
clearTimeout(timer);
timer = null;
if (sending || queue.length === 0) return;
sending = true;
const batch = queue.splice(0, MAX_BATCH);
for (let attempt = 0; attempt < MAX_ATTEMPTS; attempt++) {
try {
const res = await fetch(ENDPOINT, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(batch),
});
if (res.ok) break;
if (res.status < 500 && res.status !== 429) break; // permanent: drop
const retryAfter = Number(res.headers.get('Retry-After'));
await sleep(retryAfter ? retryAfter * 1000 : backoff(attempt));
} catch {
await sleep(backoff(attempt)); // network error
}
}
sending = false;
if (queue.length > 0) flush();
}
function backoff(attempt) {
const base = Math.min(1000 * 2 ** attempt, 30000);
return base / 2 + Math.random() * (base / 2); // jitter
}
document.addEventListener('visibilitychange', () => {
if (document.visibilityState !== 'hidden' || queue.length === 0) return;
while (queue.length > 0) {
const chunk = queue.splice(0, MAX_BATCH);
const body = new Blob([JSON.stringify(chunk)], { type: 'text/plain;charset=UTF-8' });
if (!navigator.sendBeacon(ENDPOINT, body)) {
queue.unshift(...chunk); // browser refused (queue limit): keep for later
break;
}
}
});
Read the loop carefully. A page hide empties the whole queue in chunks, because a single beacon of 20 events would strand the rest. The batch leaves the queue before the first attempt, and retries reuse the same array. If all six attempts fail, the batch is dropped, which is a deliberate bound. An unbounded client queue on a dead network eats memory on the user’s device, and your analytics should never do that.
The trade-off is explicit. A small share of events from users on very bad connections will be lost. You accept that in exchange for a tracker that can never harm the page. If those users matter, persist the queue to localStorage and replay it on the next visit, at the price of extra code and a privacy review.
The Server Buffer
The collector’s job is to move an event from the socket into memory as quickly as possible. It validates, pushes to an array, and answers 202 Accepted, which says the request was received but processing is not complete. That status is honest, and 200 would be a small lie.
Why the buffer must be bounded
An unbounded array is a memory leak waiting for a traffic spike. If PostgreSQL stalls for two minutes and events keep arriving, the process grows until the operating system kills it, taking every buffered event with it. A cap converts that disaster into a controlled signal.
When the buffer is full, return 503 with Retry-After. This is backpressure: the server tells clients to slow down, and well-behaved clients such as the tracker above comply. Rejecting early costs the client a retry, while accepting everything costs you the process.
Batched, idempotent inserts
One INSERT per event wastes round trips. One INSERT for 500 events uses a single round trip and a single transaction. The cleanest way to pass many rows through a parameterized query with plain pg is to send one array per column and expand them with unnest. The statement text never changes, and no value is ever concatenated into SQL.
Add ON CONFLICT (event_id) DO NOTHING, described in the PostgreSQL INSERT reference. Duplicates from client retries disappear without an error. With DO NOTHING, two rows that share an event_id inside the same batch are also fine. A DO UPDATE clause would raise an error in that case, so do not use it for events. One caveat: this conflict target stops working once you partition the events table by time, because the key must then include the partition column. The tuning article shows the change.
The complete collector
// collector.mjs (npm install express cors pg)
import express from 'express';
import cors from 'cors';
import pg from 'pg';
const pool = new pg.Pool({
connectionString: process.env.DATABASE_URL,
max: 5,
});
const ALLOWED_ORIGINS = ['https://acme-shop.example'];
const MAX_BUFFER = 10000; // events held in memory
const FLUSH_SIZE = 500; // rows per INSERT
const FLUSH_INTERVAL_MS = 1000;
const UUID_RE = /^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$/i;
const NAME_RE = /^[a-z][a-z0-9_]{0,63}$/; // use the allow-list from the collector article in production
let buffer = [];
let flushing = false;
let accepting = true;
// Strict on purpose: only the exact output of Date.prototype.toISOString().
// Date.parse alone accepts "2025-02-30" and "Oct 2 2025"; PostgreSQL rejects the first.
function isIsoUtc(v) {
if (typeof v !== 'string') return false;
const t = Date.parse(v);
return !Number.isNaN(t) && new Date(t).toISOString() === v;
}
// PostgreSQL text cannot hold a NUL byte, so reject it here.
const isText = (v, max) => typeof v === 'string' && v.length > 0 && v.length <= max && !v.includes('\u0000');
const isOptText = (v, max) => v === undefined || v === null || isText(v, max);
function isValid(e) {
return (
e && typeof e === 'object' &&
typeof e.event_id === 'string' && UUID_RE.test(e.event_id) &&
typeof e.event_name === 'string' && NAME_RE.test(e.event_name) &&
isIsoUtc(e.occurred_at) &&
isText(e.anonymous_id, 64) && isText(e.session_id, 64) &&
isOptText(e.user_id, 64) && isOptText(e.page_url, 2048) &&
isOptText(e.referrer, 2048) && isOptText(e.user_agent, 512) &&
(e.properties === undefined ||
(typeof e.properties === 'object' && e.properties !== null && !Array.isArray(e.properties)))
);
}
const app = express();
app.use(cors({ origin: ALLOWED_ORIGINS, methods: ['POST'] }));
app.use(express.json({ limit: '100kb', type: ['application/json', 'text/plain'] }));
app.post('/v1/events', (req, res) => {
if (!accepting) return res.set('Retry-After', '5').sendStatus(503);
const batch = Array.isArray(req.body) ? req.body : [req.body];
if (batch.length > 100) return res.sendStatus(413);
const valid = batch.filter(isValid);
if (valid.length === 0) return res.sendStatus(400);
if (buffer.length + valid.length > MAX_BUFFER) {
return res.set('Retry-After', '5').sendStatus(503); // backpressure
}
buffer.push(...valid);
res.status(202).json({ accepted: valid.length, rejected: batch.length - valid.length });
});
async function insertBatch(rows) {
const col = (f) => rows.map(f);
const { rowCount } = await pool.query(
`INSERT INTO events
(event_id, event_name, occurred_at, anonymous_id, user_id,
session_id, page_url, referrer, user_agent, properties)
SELECT * FROM unnest(
$1::uuid[], $2::text[], $3::timestamptz[], $4::text[], $5::text[],
$6::text[], $7::text[], $8::text[], $9::text[], $10::jsonb[]
)
ON CONFLICT (event_id) DO NOTHING`,
[
col((e) => e.event_id), col((e) => e.event_name),
col((e) => e.occurred_at), col((e) => e.anonymous_id),
col((e) => e.user_id ?? null), col((e) => e.session_id),
col((e) => e.page_url ?? null), col((e) => e.referrer ?? null),
col((e) => e.user_agent ?? null),
col((e) => JSON.stringify(e.properties ?? {})),
]
);
return rowCount;
}
async function insertOneByOne(rows) {
for (const row of rows) {
try {
await insertBatch([row]);
} catch (err) {
console.error('dropping bad event', row.event_id, err.message);
}
}
}
async function flush() {
if (flushing || buffer.length === 0) return;
flushing = true;
const rows = buffer.splice(0, FLUSH_SIZE);
try {
await insertBatch(rows);
} catch (err) {
if (/^2[23]/.test(err.code ?? '')) {
// Data or constraint error: the rows are bad, not the database.
await insertOneByOne(rows);
} else {
console.error('flush failed, requeueing', err.message);
buffer = rows.concat(buffer); // connection problem: keep order, retry
}
} finally {
flushing = false;
}
}
const ticker = setInterval(flush, FLUSH_INTERVAL_MS);
const server = app.listen(3000, () => console.log('collector on :3000'));
async function shutdown(signal) {
console.log(signal, 'received, draining', buffer.length, 'events');
accepting = false;
clearInterval(ticker);
server.close();
const deadline = Date.now() + 10000;
while (buffer.length > 0 && Date.now() < deadline) {
await flush();
if (flushing) await new Promise((r) => setTimeout(r, 50));
}
await pool.end();
process.exit(buffer.length === 0 ? 0 : 1);
}
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));
Set DATABASE_URL and start it with node collector.mjs. Test with one request and a repeat:
curl -i -X POST http://localhost:3000/v1/events \
-H "Content-Type: application/json" \
-d '{"event_id":"5b30857f-0bfa-48b5-ac0b-5c64e28078d1","event_name":"page_view","occurred_at":"2025-10-02T09:15:30.123Z","anonymous_id":"anon_7f3a","session_id":"sess_91bc"}'
Send it twice, wait a second, and count rows. You get one row, because the second copy hits the primary key and is skipped. To see the poison fix, change occurred_at to 2025-02-30T09:15:30.123Z and confirm the collector answers 400 instead of stalling.
The poison event
The first version of my collector requeued the batch on any error. It looked careful, and it contained a trap. One malformed event could wedge the whole pipeline.
// WRONG: any error puts the same rows back
} catch (err) {
buffer = rows.concat(buffer);
}
I reproduced the failure with a validator that used Date.parse. V8 parses 2025-02-30T09:15:30.123Z as March 2, so the event passed validation. PostgreSQL rejects the same string with “date/time field value out of range”. The insert failed, the rows went back to the front of the buffer, and the next flush failed identically. Every later event waited behind it, and the buffer filled until the collector returned 503 to everyone.
The fix has two layers, both in the code above. Validation now requires the exact toISOString() form, which rejects impossible dates before they enter the buffer. The flusher also separates error classes. Errors in PostgreSQL class 22 (data exception) or 23 (integrity violation) mean the data is bad, so it retries row by row and drops the offenders with a log line. Connection errors still requeue, because the database is the problem there. The delta is that one bad row now costs one event instead of the whole stream.
Idempotency Is the Contract
Retries only work if repeating a request has no extra effect. That is idempotency, and in this design it has one source: the UUID the browser created before the first send. The server never invents an identifier for an event it already received, and it never trusts a counter.
We once double-counted purchases because a server assigned its own row ID on arrival, while the client retried after a timeout that had in fact succeeded. Each retry became a new row. Moving the identifier to the client, and making it the primary key, fixed the duplicates without any deduplication job. If you have only inherited data with this problem, DISTINCT ON over a stable key in a cleanup query helps, but prevention is far cheaper.
This guarantee is at-least-once delivery with deduplication at the end, which behaves like exactly-once for counting. The broader treatment of delivery guarantees belongs to the event-driven architecture article.
Graceful Shutdown
Deploys and autoscaling stop your process regularly. Orchestrators send SIGTERM, wait a grace period, and then send SIGKILL. A process that ignores SIGTERM loses everything in its buffer at every deploy. With a one-second flush interval and a few deploys a day, that is a small but constant, and nearly invisible, leak.
The Node.js documentation on process signal events shows how to register handlers. The order in the code above matters:
- Stop accepting new events by flipping a flag, so health checks and load balancers see 503 and route elsewhere.
- Stop the periodic timer, and close the HTTP server so in-flight requests finish.
- Flush in a loop until the buffer is empty or a deadline passes. Pick a deadline shorter than the platform’s grace period.
- Close the pool and exit with a non-zero code if events remain, so you can alert on it.
SIGKILL, an out-of-memory kill and a power loss cannot be handled. The events in memory at that moment are gone. This is the honest limit of an in-process buffer, and it is why the buffer interval is short.
Sizing and Measuring
Do not guess numbers; measure them against your own database. Create the table, start the collector and run a load tool such as autocannon against it with a realistic batch body. Record three things: p99 response time, buffer length at peak, and the time a flush takes. Then change FLUSH_SIZE and the pool size one at a time.
Use these relationships as a guide. The buffer limit divided by your peak ingest rate is how many seconds of database outage you can survive. If that is shorter than your typical failover time, raise the cap or add a queue. If a flush regularly takes longer than the flush interval, increase the batch size before you increase concurrency, because batch size reduces round trips, while concurrency adds contention on the same table.
For observability, expose two counters: buffer length and total rejected events. A buffer that ends every minute near zero is healthy. One that climbs steadily is a stalled database, and a flat line at the cap means you are already shedding load.
Async Event Tracking Options Compared
| Option | Durability on crash | Complexity | Use when |
|---|---|---|---|
| Synchronous insert per request | Highest (committed before reply) | Lowest | Prototype, under a few events per second |
| In-process buffer (this article) | Loses the buffer on a hard kill | Low | Hundreds to low thousands of events per second on one or a few processes |
| Redis Streams | Persistent if configured, replayable | Medium | You need durability across restarts and several consumers |
| Kafka or similar log | Replicated and replayable | High | High volume, many consumers, long retention |
| Managed queue (SQS, Pub/Sub) | Provider-managed durability | Medium, plus provider cost | You want durability without running the broker |
Move down the table only when a measured problem forces you to. Signals include buffer losses you can no longer tolerate, several services that need the same events, and a need to replay history. I will not build a Kafka cluster here. The architecture article explains when and why.
How Real Systems Do This
Segment’s analytics libraries queue events on the client and flush them in batches, with retries for failed deliveries. Snowplow’s JavaScript tracker has a configurable buffer and a batch size for the same reason, and its collector hands events to a stream rather than writing to a warehouse directly. Plausible, which is simpler, uses a lightweight script that sends one request per event. Check each product’s current documentation before copying a detail, because these libraries change.
The shared principle is that the collector is a thin front door. It acknowledges quickly and passes the heavy work to something that can fall behind without hurting visitors. Your in-memory buffer is the smallest version of that door.
Decision Framework
- Can you afford to lose the events in flight during a crash? If not, skip the in-process buffer and write to a durable queue first.
- Does one process handle your peak rate? If yes, a single buffer works. If not, run several collectors, each with its own buffer, and rely on
event_idfor deduplication. - Do other services need the same events? If yes, put a broker between collection and storage.
- Is database latency visible to users today? If yes, make the endpoint async before anything else.
- Do you need to replay history after a bug? If yes, plan for a log of raw events, not just the final table.
When NOT to Use This
- Revenue or audit events that must never be lost. Write purchases from your order service in the same transaction as the order. Analytics events are a copy, not the ledger.
- Very low traffic. Under a handful of events per second, a synchronous insert is simpler and has no loss window. Add the buffer when you can measure a reason.
- You would rather buy than operate this. A hosted collector or a customer data platform handles queues, retries and replay for you. If reliability engineering is not what you want to own, paying for it is rational.
Common Mistakes
- Generating
event_idat send time, which turns every retry into a new row and inflates counts. - Using an unbounded buffer, which lets a database stall grow the process until the kernel kills it.
- Retrying on 400-class errors, which creates an endless loop of requests that can never succeed.
- Relying on
unloadinstead ofvisibilitychangeandsendBeacon, which silently drops the last events of many sessions. - Ignoring SIGTERM, which discards the buffer at every deploy.
- Requeueing a failed batch on every error, which lets one malformed event block all later ones.
Key Takeaways
- Acknowledge fast and persist later: respond 202 after validation, then write in batches.
- Create
event_idonce on the client, and make it the primary key so retries are safe. - Bound both queues, and answer overload with 503 and
Retry-Afterinstead of growing. - Send on
visibilitychangewithsendBeacon, and keep each beacon well under 64 KiB. - Insert with one parameterized
unneststatement andON CONFLICT DO NOTHING, and drop bad rows instead of requeueing them forever. - Handle SIGTERM by refusing new work, draining the buffer and exiting with a clear status.
- Know your loss window, and move to Redis Streams or a managed queue only when you measure the need.
FAQ
How do I track analytics events without slowing down my website?
Queue events in the browser, send them in batches in the background, and use sendBeacon when the page hides. Never make a user interaction wait for the network response from your tracking endpoint.
Is sendBeacon reliable for analytics?
It is the most reliable browser option for sending data as a page closes, because the browser delivers the request after the page is gone. It returns only whether the request was queued, not whether the server stored it, and it has a 64 KiB total queue limit. Pair it with server-side deduplication.
How do I prevent duplicate events when retrying?
Generate a UUID when the event is created and keep it through every retry. The server inserts with ON CONFLICT (event_id) DO NOTHING, so a repeated event is ignored.
What happens to buffered events when the server restarts?
If the process receives SIGTERM, a shutdown handler can flush the buffer before it exits. If it is killed with SIGKILL or crashes, events still in memory are lost. A durable queue removes that risk at the cost of more infrastructure.
When should I use Kafka instead of an in-memory buffer?
Move to a log or managed queue when you cannot tolerate the loss window, when several consumers need the same events, or when you need to replay history. For one collector and one database, an in-memory buffer is usually enough.
Conclusion
Async event tracking is a handful of small rules: batch at the edge, give every event a stable identifier, bound every queue and flush on the way out. With those in place, async event tracking lets your analytics fail and recover without your users noticing. The next article tackles a different problem, identifying visitors when you cannot or should not use cookies.
Rule of thumb: accept fast, store later, and let the event_id make every retry free.
Last updated on 9 October 2026.
