How Web Analytics Works: Tracking Pixels, HTTP Requests and Cookies

Learn how web analytics turns a page view into a database row: tracking pixels, HTTP requests, first-party cookies, and why your visitor counts never match.

A visitor opens a product page on Acme Shop, a small online store. Two seconds later a row appears in your database with their page, their referrer and a random identifier. Nothing magical happened in between. Web analytics is a chain of ordinary HTTP requests, and once you can read that chain you can build, debug and trust your own numbers.

Most developers treat web analytics as a script tag they paste into a template. That works until the numbers disagree with your payment provider, a browser update halves your returning visitors, or an ad blocker hides a third of your traffic. At that point you need a mental model, not a dashboard.

This article is the foundation of the series. You will see what a browser sends, what a server can learn from it, how a tracking pixel works, and how cookies turn anonymous requests into repeat visitors. You will also run a tiny pixel server yourself and watch the first events arrive.

Later articles build on it. You will write the browser code in how to track button clicks and form submissions with JavaScript and use server logs as a second data source in how to centralize server logs for a web app.

Executive Summary: Web analytics works by having the browser send an HTTP request to your server whenever something happens, carrying the page, referrer, user agent, and usually a cookie. This post traces a page view to a database row through tracking pixels and first-party cookies, and explains why blockers and privacy rules make visitor counts estimates.

The Chain From Page View to Database Row

Every analytics system, from a hobby project to a commercial platform, follows the same four steps. The browser detects something. It sends a request. A server receives and validates it. A database stores it. The details differ, but the order never does.

Browser                         Analytics server              Database
   |                                   |                          |
   | 1. page loads, script runs        |                          |
   |                                   |                          |
   | 2. GET /p.gif?e=page_view&...     |                          |
   |    Cookie: anonymous_id=abc       |                          |
   |--------------------------------->|                          |
   |                                   | 3. validate, add         |
   |                                   |    server timestamp      |
   |                                   |------------------------> |
   |                                   |        4. INSERT row     |
   |  <---- 200 OK, 1x1 GIF -----------|                          |

Notice what is missing: there is no response the user cares about. The browser sends the request for the server’s benefit and throws the answer away. This one-way shape explains most of the quirks you will meet later, such as lost events on page unload and requests that ad blockers silently drop.

Throughout this series, the stored row follows one canonical event shape, shown below. Every article uses the same fields, the same snake_case names and the same fictional store, so the examples connect.

{
  "event_id": "3f0c2a52-6a3e-4c0b-9b8e-1d7d1f5e7a10",
  "event_name": "page_view",
  "occurred_at": "2025-10-02T09:15:30.123Z",
  "anonymous_id": "b7e1c9d2-48a0-4b52-8f4a-0c5a9a1d3e77",
  "user_id": null,
  "session_id": "e2d4a6c8-91b3-4f55-a7c0-5b8d3e9f1a22",
  "page_url": "https://acme-shop.example/products/42",
  "referrer": "https://www.google.com/",
  "user_agent": "Mozilla/5.0 (...)",
  "properties": {}
}

The six canonical event names are page_view, button_click, form_submit, signup_completed, add_to_cart and purchase_completed. This article only needs the first one.

Anatomy of a Tracking Request

An HTTP request is plain text. Understanding it line by line shows you exactly what a tracker can and cannot know. Here is the request a browser sends when the page on Acme Shop loads a tracking pixel from analytics.acme-shop.example.

GET /p.gif?e=page_view&u=https%3A%2F%2Facme-shop.example%2Fproducts%2F42&r=https%3A%2F%2Fwww.google.com%2F HTTP/1.1
Host: analytics.acme-shop.example
User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) ...
Referer: https://acme-shop.example/products/42
Accept: image/avif,image/webp,*/*
Accept-Language: en-US,en;q=0.9
Cookie: anonymous_id=b7e1c9d2-48a0-4b52-8f4a-0c5a9a1d3e77

Each line carries analytics value. Therefore it helps to sort them into what the browser volunteers, what your own code adds, and what the network reveals.

What the browser volunteers

  • User-Agent: the browser, version and operating system as a raw string. You store it raw and parse it later.
  • Referer: the page that made the request, which for a pixel is your own page, not the site the visitor came from. The header name is misspelled in the HTTP specification, and the misspelling stuck. Browsers often trim it to the origin only.
  • Accept-Language: a rough language signal, useful for audience reports.
  • Cookie: any cookies previously set by this exact host or a parent domain you control.

What your own code adds

The query string is yours. You choose the event name, the page URL and any custom property. In the example above the page script encodes the current address into u and the previous site into r, taken from document.referrer. Because a pixel is a GET request, everything has to fit in a URL, and servers and proxies commonly cap URLs at a few thousand characters. For richer payloads you switch to a POST, which the collector in article 6 accepts at POST /v1/events.

What the network reveals

The server also sees the client’s IP address. An IP address is personal data under most privacy laws, so a careful system uses it briefly, perhaps to look up a country, and does not store it. The series follows that rule: no raw IP addresses in the events table.

The server also owns the clock. A client timestamp can be wrong by minutes or years, because users change their system time. In practice you store the client’s occurred_at for ordering and a server receive time for trust, and you flag large gaps.

The Tracking Pixel

A tracking pixel is an image request used as a message. The classic form is a 1×1 transparent GIF. The page includes it with an img tag, the browser fetches it, and the server logs the fetch. The image itself is irrelevant. The request is the data.

Pixels survive for one reason: every browser, mail client and rendering engine knows how to load an image. A pixel works inside an HTML email, where JavaScript never runs. It also works when a page script fails. That makes it the most portable tracking method ever built.

It also has real costs. A pixel can only send GET requests, so payloads stay small. It cannot read a response. And many ad blockers and mail clients block known tracker hosts, or prefetch images through a proxy that hides the real reader. Email open rates are inflated for exactly that reason.

A complete pixel server you can run

The following example is a demonstration of the mechanism, not the production collector. It uses Node.js (current LTS) and Express. Install with npm install express, save the file as pixel.js, and add "type": "module" to your package.json.

import express from 'express';
import { randomUUID } from 'node:crypto';

// A 1x1 transparent GIF, decoded once at startup.
const GIF = Buffer.from(
  'R0lGODlhAQABAIAAAAAAAP///yH5BAEAAAAALAAAAAABAAEAAAIBRAA7',
  'base64'
);

const ALLOWED_EVENTS = new Set([
  'page_view', 'button_click', 'form_submit',
  'signup_completed', 'add_to_cart', 'purchase_completed',
]);

function readCookie(header, name) {
  if (!header) return null;
  for (const part of header.split(';')) {
    const [key, ...rest] = part.trim().split('=');
    if (key === name) return rest.join('=');
  }
  return null;
}

const app = express();

app.get('/p.gif', (req, res) => {
  const eventName = String(req.query.e ?? '');
  const pageUrl = String(req.query.u ?? '').slice(0, 2048);
  const referrer = String(req.query.r ?? '').slice(0, 2048) || null;

  // Validate every event before it touches storage.
  if (!ALLOWED_EVENTS.has(eventName)) {
    return res.status(400).end();
  }

  let anonymousId = readCookie(req.headers.cookie, 'anonymous_id');
  if (!anonymousId) {
    anonymousId = randomUUID();
    res.append(
      'Set-Cookie',
      `anonymous_id=${anonymousId}; Max-Age=31536000; Path=/; Secure; HttpOnly; SameSite=Lax`
    );
  }

  const row = {
    event_id: randomUUID(),
    event_name: eventName,
    occurred_at: new Date().toISOString(),
    anonymous_id: anonymousId,
    page_url: pageUrl,
    referrer,
    user_agent: req.get('user-agent') ?? null,
  };
  console.log(JSON.stringify(row)); // swap for a database insert later

  res.set({
    'Content-Type': 'image/gif',
    'Content-Length': GIF.length,
    'Cache-Control': 'no-store',
  });
  res.status(200).send(GIF);
});

app.listen(3000, () => console.log('pixel server on :3000'));

Run it with node pixel.js, then open http://localhost:3000/p.gif?e=page_view&u=https%3A%2F%2Facme-shop.example%2Fproducts%2F42 in a browser. Refresh twice. The terminal prints one JSON line per request, and the anonymous_id stays the same after the first hit. (The Secure flag is accepted on localhost by current Chrome and Firefox. If your browser rejects it, remove the flag for local testing only.)

Three details deserve attention. First, Cache-Control: no-store forces the browser to ask again on every page view. Without it, a cached pixel makes a returning page view invisible. Second, the event name passes through an allow-list before anything else happens. Third, the code never reads the client IP. That is deliberate.

On the page itself, a single line of vanilla JavaScript fires the pixel:

const params = new URLSearchParams({
  e: 'page_view',
  u: location.href,
  r: document.referrer,
});
new Image().src = 'https://analytics.acme-shop.example/p.gif?' + params;

Creating an Image object in memory triggers the request without touching the DOM. That is the whole client side of a pixel tracker. Note that the server row carries only the fields a GET request can supply, so session_id and user_id are absent here.

Beacons: The Modern Pixel

Pixels are limited to GET. Modern tracking code therefore uses navigator.sendBeacon() or fetch() with keepalive, both of which send POST bodies and survive page unload. The MDN documentation for sendBeacon describes the exact contract: the browser queues the request and sends it even if the page closes.

The trade-off is simple. Beacons give you structured JSON and reliability on exit, but they need JavaScript and a server that accepts POST with CORS. Pixels give you reach and simplicity, but little else. Article 2 covers the browser code in depth, so this article stops here.

Cookies: How a Browser Becomes a Visitor

HTTP is stateless. Two requests from the same person look unrelated unless something links them. A cookie is that link. The server sends a Set-Cookie response header, the browser stores the value, and it attaches the value to every later request to the matching host. The MDN guide to HTTP cookies and RFC 6265 define the behavior.

An analytics cookie stores one thing: a random identifier. In this series it is the anonymous_id. It contains no name, email or product data. Its only job is to say “these requests came from the same browser profile”.

The attributes that matter

Attribute What it does Analytics recommendation
Max-Age Lifetime in seconds One year is common. Shorter is more privacy-friendly.
Path URL prefix where the cookie is sent Use / so every page is covered.
Secure Send only over HTTPS Always set it in production.
HttpOnly Hide from document.cookie Set it when the server issues the cookie. Page scripts then cannot read it.
SameSite Controls cross-site sending Lax works for a first-party tracker.
Domain Which hosts receive the cookie Omit it to bind to the exact host, or set it to share across subdomains.

A mistake I have seen in production

The wrong version of a cookie looks harmless:

// WRONG: no attributes, set from page JavaScript
document.cookie = 'anonymous_id=' + crypto.randomUUID();

That line creates a session cookie, which vanishes when the browser closes. In my experience building event pipelines, a team shipped exactly this and watched “new visitors” jump every morning, because nobody ever came back as a returning visitor. The fix is explicit lifetime and scope:

// RIGHT: explicit lifetime, path, and SameSite
const maxAge = 60 * 60 * 24 * 365;
document.cookie =
  `anonymous_id=${crypto.randomUUID()}; Max-Age=${maxAge}; Path=/; Secure; SameSite=Lax`;

The difference is the attributes. Without Max-Age the cookie dies with the browser session. Without Path=/ it binds to the current directory path only, so a visitor on /products/42 and one on /cart could receive different identities. Also, Safari limits how long it keeps cookies written by JavaScript, so a script-set cookie often lives far shorter than the lifetime you asked for. A server-set cookie, like the one in the pixel example, avoids that cap in most cases. Check the current WebKit tracking prevention notes before relying on any specific number.

First-Party vs Third-Party Tracking

A cookie is first-party when the host that sets it matches the site in the address bar, and third-party when it belongs to another site embedded in the page. The distinction decides whether the cookie survives.

Suppose Acme Shop embeds a pixel from a vendor domain, tracker.example. The browser treats that cookie as third-party. Safari and Firefox block it by default, and Chrome gives users controls to do the same. Now suppose Acme Shop serves the pixel from analytics.acme-shop.example. That host shares the registrable domain acme-shop.example, so browsers treat it as same-site and first-party.

Property Third-party tracker First-party tracker
Host Vendor domain Your own subdomain
Cookie survival Often blocked by default Kept, within browser lifetime caps
Ad blocker exposure High, hosts appear on block lists Lower, but a recognizable path can still match a filter
Cross-site profiles Possible Not possible by design
Who owns the data The vendor You

Owning your data is the point of this series, so every example uses a first-party host. Be aware that a subdomain pointing at a third-party service through a CNAME can still be flagged by browsers that inspect DNS. A subdomain that runs your own server does not have that problem.

Identity in Web Analytics: Visitor, Session, User

New analytics builders confuse three layers of identity. Keep them separate from day one.

  • Anonymous id: one browser profile on one device. Clearing cookies creates a new one.
  • Session id: a burst of activity from one anonymous id, closed after 30 minutes of inactivity.
  • User id: your own account identifier, null until the person signs up or logs in.

A “unique visitor” is therefore really a unique cookie. One person with a phone and a laptop counts twice. One person who clears cookies weekly counts as many. By contrast, a user id only exists after login, so it undercounts anonymous traffic. Pick the unit that matches your question and name it plainly in every report.

Why Your Numbers Will Never Match

Beginners expect web analytics to equal ground truth. It does not. Every link in the chain loses or distorts data, and you should know where.

  • Ad blockers and tracking protection: the request never leaves the browser, so no row exists.
  • Script failures: a JavaScript error before your snippet runs means no event.
  • Page unload: a plain request started as the tab closes may be cancelled. Beacons reduce this loss.
  • Bots: crawlers and scrapers produce page views. Some execute JavaScript, so they look human.
  • Cookie loss: private windows, expiry and manual clearing split one person into several visitors.
  • Time zones: a day in the report depends on whose midnight you use. Store UTC and convert at query time.

Payment records are the cleanest check. If your payment provider shows 100 orders and analytics shows 85 purchase_completed events, the gap is the client-side loss rate for that event. Moving the purchase event to the server, where no blocker can interfere, closes most of it. I once helped debug a store that double-counted purchases because the thank-you page fired the event on every reload. The fix was an event_id generated once and de-duplicated by the server, a pattern the collector article implements.

How Real Systems Do This

Real analytics products differ in packaging, but they share this request chain. Google Analytics 4 sends measurement requests from a JavaScript library to Google collection endpoints, and its cookies are set by that library. Plausible and Matomo document similar script-plus-endpoint designs, and Matomo also supports a pixel-style image tracker for cases where JavaScript is unavailable. PostHog and Snowplow ship SDKs that batch events and send them to a collector you can host yourself.

The common pattern is a thin client that gathers a few fields, a stateless collector that validates and stores, and everything else done later in SQL. Privacy-focused tools such as Plausible go a step further and avoid cookies entirely, a technique covered in cookieless analytics.

Buying remains a legitimate option. A hosted tool gives you bot filtering, identity stitching and dashboards on day one. Building gives you ownership, custom events and the ability to join analytics to your own tables. This series assumes you want the second.

Decision Framework

Use these questions, in order, to pick a tracking method for a new event.

  1. Does the event happen where JavaScript cannot run, such as an email open? Use a pixel.
  2. Does the event matter for revenue, such as a purchase? Send it from the server, and add a client event only for comparison.
  3. Does the event need a payload larger than a short URL? Use a POST beacon.
  4. Must the event survive page navigation or tab close? Use sendBeacon or fetch with keepalive.
  5. Do you need to recognize returning browsers? Issue a first-party cookie from your own host, and record what the visitor consented to.
  6. Can you answer the question without identity? If yes, skip the cookie.

When NOT to Use This

  • You need reliable revenue numbers. Client-side tracking loses events to blockers. Count money from your order database and treat analytics as the funnel view around it.
  • You need no-effort bot filtering, attribution modeling or ad-platform integrations. A hosted product already does this, and rebuilding it costs months. Buy instead of build.
  • Your site has a few hundred visits a month. A simple server log summary answers your questions, as the article on centralizing server logs shows.

Common Mistakes

  • Setting the cookie with no Max-Age. Every visit looks like a new visitor, so retention reads as zero.
  • Letting the browser cache the pixel. Repeat page views vanish and traffic looks flat.
  • Trusting client timestamps alone. Wrong device clocks scatter events into the wrong days.
  • Storing raw IP addresses by default. You collect personal data you never needed and widen your legal exposure.
  • Accepting any event name. Typos create names like pageview next to page_view and split every report.
  • Counting cookies as people. Reports claim more “users” than you have customers.

Key Takeaways

  • A page view becomes a database row through four steps: detect, send, validate, store.
  • A tracking pixel is an image request used as a message, and it works anywhere images load.
  • Cookies link stateless requests by carrying a random anonymous_id, nothing more.
  • Serve the tracker from your own subdomain to keep it first-party and keep the data yours.
  • Always validate event names and set cookie attributes explicitly.
  • Treat visitor counts as estimates, and cross-check money against your order records.
  • Do not store raw IP addresses unless you have a specific, documented reason.

FAQ

How does web analytics work?

The browser sends an HTTP request to an analytics server when a page loads or a user acts. The request carries the page address, referrer, user agent and usually a cookie with a random identifier. The server validates the request and stores a row, and reports are built from those rows.

What is a tracking pixel?

A tracking pixel is a tiny image, usually a 1×1 transparent GIF, whose only purpose is to make the browser send a request. The server logs the request and returns the image. It works in emails and pages where scripts cannot run.

What is the difference between first-party and third-party cookies?

A first-party cookie is set by the site the user is visiting, or by a subdomain of it. A third-party cookie is set by another domain embedded in the page. Browsers block third-party cookies far more aggressively, so first-party tracking keeps more data.

Why do my analytics numbers differ from my sales numbers?

Client-side tracking loses events to ad blockers, script errors, page unloads and cookie clearing. It also gains false events from bots and reloads. Sales records come from your order system, so they stay more accurate.

Conclusion

Web analytics is a request, a server and a table, held together by a random identifier in a cookie. Once you see that chain, every odd number has an explanation you can test. The next step is writing the browser code that decides which requests to send.

Rule of thumb: an analytics number is only as honest as the request chain behind it, so know where the chain can break.

Last updated on 9 October 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *