How to Centralize Server Logs for a Web App

Learn how to centralize server logs for a web app and turn nginx and app logs into analytics events with JSON logging, Vector and a runnable Node script.

Suppose your tracking script says Acme Shop had 8,200 page views on Tuesday, while your nginx access log shows 11,400 successful page requests. Neither number is wrong. The script missed visitors with ad blockers, and the log counted bots. Server logs are a second, independent record of what happened, and once you centralize server logs you can compare the two and trust both more.

Most teams treat logs as a debugging tool. You grep them when something breaks and ignore them otherwise. That wastes a complete event stream: every request your app served is already written to disk, with a timestamp, a URL, a status code and a user agent. No JavaScript has to run, and no blocker can hide it.

This article shows how to treat logs as event data. You will switch nginx to JSON access logs, emit structured application logs, ship both to one place with Vector, and convert them into the same event shape that the browser tracker from the previous article produces. You will also run a complete Node.js converter on sample data.

This is not an observability tutorial. Metrics, tracing and alerting are out of scope. The goal is narrower: get request and business events out of log files and into your analytics store.

Executive Summary: Server logs capture every request even without JavaScript, which makes them a reliable second source for page views and purchases. This post shows how to write JSON logs from nginx and your app, ship them with Vector, map each line to the standard event shape, filter bots and static files, and keep client IPs out of the pipeline.

What Logs Can and Cannot Tell You

Before building anything, be clear about what a log gives you. As article 1 showed, an HTTP request carries a path, headers and cookies. A log line is a record of one such request, written by the server after it responded.

Question Browser tracker Server logs
Did the request reach my server? Only if the script ran Yes, always
Is the visitor using an ad blocker? Hides the visit Visit still logged
Which button did the user click? Yes No, clicks make no request
Screen size and viewport? Yes No
Did the purchase succeed? Only the attempt Yes, if the app logs it
Is this a bot? Often invisible Visible, and must be filtered
Is it the same person as yesterday? Cookie or storage id Only if the cookie is logged

The two sources complement each other. Logs excel at completeness and server-side truth. Browsers excel at interaction detail. Therefore a mature setup records both and uses logs to audit the browser data.

Step 1: Stop Writing Logs for Humans

The default nginx access log is built for people reading a terminal. It looks like this:

203.0.113.9 - - [02/Oct/2025:09:15:30 +0000] "GET /products/42 HTTP/1.1" 200 5120 "https://www.google.com/" "Mozilla/5.0 (Windows NT 10.0)"

Parsing that line requires a regular expression that must handle quotes, brackets and missing fields. It also contains a raw IP address, which you do not want in an analytics pipeline. Worse, one unusual user agent with a quote inside can break the parser silently, and you only notice when the numbers drift.

The fix is structured logging: one JSON object per line, with named fields and a fixed schema. Machines parse it without guessing, and you control exactly which fields exist.

An nginx log format built for analytics

The nginx log module supports an escape=json parameter on log_format. It escapes quotes and control characters in variable values so the output stays valid JSON. Add this to your http block:

log_format analytics escape=json
  '{'
    '"time":"$time_iso8601",'
    '"request_id":"$request_id",'
    '"host":"$host",'
    '"method":"$request_method",'
    '"path":"$uri",'
    '"query":"$args",'
    '"status":$status,'
    '"bytes":$body_bytes_sent,'
    '"request_time":$request_time,'
    '"referrer":"$http_referer",'
    '"user_agent":"$http_user_agent",'
    '"anonymous_id":"$cookie_anonymous_id"'
  '}';

access_log /var/log/nginx/analytics.json analytics;

Several choices here are deliberate.

  • No $remote_addr. The series stores no raw IP addresses, so the analytics log never writes one. Your default access log may still contain IPs, so set a short retention for it or disable it.
  • $cookie_anonymous_id. Nginx exposes any request cookie as a variable. If your site writes the anonymous_id cookie on its own host, as the browser tracker does, the log carries the same identifier as the browser events, which lets you join the two sources. A cookie set only by the analytics subdomain is not sent to your app host.
  • $request_id. Nginx generates a unique value per request. You can pass it to your app in a header and find the log line for any user complaint.
  • $uri and $args separately. Splitting path from query string keeps URLs groupable and stops tracking parameters from fragmenting your page counts.

Numeric variables like $status and $request_time appear without quotes so they stay numbers in JSON. Reload with nginx -s reload, load a page, and check that tail -n 1 /var/log/nginx/analytics.json | jq . prints clean JSON. If jq complains, your format string has a missing quote.

Step 2: Log Business Events From the App

Access logs know a request happened. They do not know that a customer paid. That fact lives in your application, and the app should log it explicitly. A mistake I have seen in production is relying on the thank-you page URL as the purchase signal. A reload, a bookmark or a prefetch hits that URL again, so purchases get double-counted. A log line written once, at the moment the payment is confirmed, is far more reliable.

Use pino, a JSON logger for Node.js. Install it with npm install pino. By default pino writes newline-delimited JSON to standard output, which is exactly the format the shipper wants.

import pino from 'pino';
import { randomUUID } from 'node:crypto';

// Remove pid and hostname noise, and write ISO 8601 UTC timestamps.
export const logger = pino({
  base: undefined,
  timestamp: pino.stdTimeFunctions.isoTime,
});

export function logPurchase({ orderId, totalCents, anonymousId, userId }) {
  logger.info({
    event_id: randomUUID(),
    event_name: 'purchase_completed',
    anonymous_id: anonymousId,
    user_id: userId,
    properties: { order_id: orderId, total_cents: totalCents },
  }, 'purchase completed');
}

Call logPurchase() once, after your payment provider confirms. The log line is an event: it already has event_name, event_id, and user_id. Notice that pino’s time field is called time, so the converter below reads time and writes occurred_at.

The trade-off is coupling. Business events in logs mix with ordinary log traffic, and a careless refactor that renames a field can corrupt your revenue data. Treat these lines as a contract. Keep the event names from the canonical list, and add a test that fails when the shape changes.

Step 3: Centralize Server Logs in One Place

You now have JSON on disk on one or more servers. Centralizing means a single agent on each server reads those files and forwards lines to one destination. Doing this with cron and scp works for exactly one server and fails on the second. The fix is a log shipper.

This article uses Vector, an open-source observability data pipeline written in Rust. Vector reads files, transforms records with a small language called VRL, and writes to many destinations. Its file source tracks position in each file, so it resumes after a restart and handles log rotation. The file source documentation lists the exact options. Note that its read_from option defaults to beginning, which means a new Vector install replays old logs. That is useful for backfill and dangerous if you did not expect it.

I ran the configuration below against the official timberio/vector Docker image: it passes vector validate and posts a JSON array of events to an HTTP endpoint. Option names change between releases, so run vector validate against your installed version before trusting it. Note that drop_on_error = true matters: without it, a line that fails to parse is forwarded unchanged instead of dropped. This transform maps nginx access lines only. Pino business events need their own source or a second branch, which the Node converter below shows in plain code.

# vector.toml
[sources.nginx]
type = "file"
include = ["/var/log/nginx/analytics.json"]
read_from = "end"            # change to "beginning" for a one-time backfill

[transforms.to_events]
type = "remap"
inputs = ["nginx"]
drop_on_error = true
source = '''
  raw = object!(parse_json!(.message))

  # Keep only successful GET page requests from real users.
  if raw.method != "GET" { abort }
  status = int!(raw.status)
  if status < 200 || status >= 300 { abort }
  if match(string!(raw.user_agent), r'(?i)bot|crawl|spider|slurp|headless|monitor') { abort }
  if match(string!(raw.path), r'(?i)\.(css|js|map|png|jpe?g|gif|svg|ico|webp|woff2?|txt|xml)$') { abort }

  anonymous_id = string!(raw.anonymous_id)
  if anonymous_id == "" { anonymous_id = "unknown" }

  referrer = string!(raw.referrer)
  referrer = if referrer == "" { null } else { referrer }

  # Stable id: the same request always maps to the same event_id.
  h = slice!(sha2(string!(raw.request_id), variant: "SHA-256"), 0, 32)
  event_id = slice!(h, 0, 8) + "-" + slice!(h, 8, 12) + "-4" + slice!(h, 13, 16) + "-a" + slice!(h, 17, 20) + "-" + slice!(h, 20, 32)

  . = {
    "event_id": event_id,
    "event_name": "page_view",
    "occurred_at": parse_timestamp!(raw.time, format: "%+"),
    "anonymous_id": anonymous_id,
    "user_id": null,
    "session_id": null,
    "page_url": "https://" + string!(raw.host) + string!(raw.path),
    "referrer": referrer,
    "user_agent": raw.user_agent,
    "properties": {
      "source": "access_log",
      "request_id": raw.request_id,
      "status": status
    }
  }
'''

[sinks.collector]
type = "http"
inputs = ["to_events"]
uri = "https://analytics.acme-shop.example/v1/events"
encoding.codec = "json"
batch.max_events = 100
batch.timeout_secs = 5

Here the destination is the collector endpoint, POST /v1/events, which accepts a batch array. You build that collector in the Express API article. In my test the http sink sent each batch as one JSON array, but confirm that with your version and your collector. Add an authentication header with request.headers, because a server-to-server sender does not need browser CORS but does need a credential.

Pushing logs through your own collector has a real advantage: validation, deduplication and storage live in one place. Browser events and log events pass the same checks. By contrast, writing directly into the database from the shipper bypasses those rules, and the first malformed line pollutes your table.

The ASCII picture

 nginx (JSON access log) --+
                           |   file source      remap          http sink
 Node app (pino JSON) -----+--> Vector agent --> filter + map --> POST /v1/events
                           |                                         |
 second web server --------+                                         v
                                                              events table

A Runnable Converter You Can Test Today

Vector is one way to do the mapping. The logic itself is small enough to read, test and run in plain Node.js. The following script reads log lines from standard input and prints canonical events. It handles both nginx access lines and pino business events, and it skips malformed lines instead of crashing. Save it as logs-to-events.js with "type": "module" in your package.json.

import { createInterface } from 'node:readline';
import { randomUUID, createHash } from 'node:crypto';

// Same request_id always yields the same UUID-shaped id, so replays dedupe.
function stableId(text) {
  const h = createHash('sha256').update(text).digest('hex');
  return `${h.slice(0, 8)}-${h.slice(8, 12)}-4${h.slice(13, 16)}-a${h.slice(17, 20)}-${h.slice(20, 32)}`;
}

const BOT = /bot|crawl|spider|slurp|headless|monitor/i;
const STATIC = /\.(css|js|map|png|jpe?g|gif|svg|ico|webp|woff2?|txt|xml)$/i;
const ALLOWED = new Set([
  'page_view', 'button_click', 'form_submit',
  'signup_completed', 'add_to_cart', 'purchase_completed',
]);

function fromAccessLog(line) {
  if (line.method !== 'GET') return null;
  if (line.status < 200 || line.status >= 300) return null;
  if (STATIC.test(line.path) || BOT.test(line.user_agent ?? '')) return null;
  return {
    event_id: stableId(line.request_id),
    event_name: 'page_view',
    occurred_at: new Date(line.time).toISOString(),
    anonymous_id: line.anonymous_id || 'unknown',
    user_id: null,
    session_id: null,
    page_url: `https://${line.host}${line.path}`,
    referrer: line.referrer || null,
    user_agent: line.user_agent ?? null,
    properties: {
      source: 'access_log',
      request_id: line.request_id,
      status: line.status,
      request_time_ms: Math.round(line.request_time * 1000),
    },
  };
}

function fromAppLog(line) {
  if (!ALLOWED.has(line.event_name)) return null;
  return {
    event_id: line.event_id ?? randomUUID(),
    event_name: line.event_name,
    occurred_at: new Date(line.time).toISOString(),
    anonymous_id: line.anonymous_id ?? 'unknown',
    user_id: line.user_id ?? null,
    session_id: line.session_id ?? null,
    page_url: line.page_url ?? null,
    referrer: null,
    user_agent: null,
    properties: { source: 'app_log', ...(line.properties ?? {}) },
  };
}

const rl = createInterface({ input: process.stdin });
let kept = 0;
let skipped = 0;
for await (const raw of rl) {
  let line;
  try {
    line = JSON.parse(raw);
  } catch {
    skipped += 1; // malformed line: count it, never crash
    continue;
  }
  const event = line.event_name ? fromAppLog(line) : fromAccessLog(line);
  if (!event) { skipped += 1; continue; }
  kept += 1;
  console.log(JSON.stringify(event));
}
console.error(`kept=${kept} skipped=${skipped}`);

Test it with five sample lines: one normal page request, one JavaScript file, one Googlebot hit, one pino purchase event and one broken line.

{"time":"2025-10-02T09:15:30+00:00","request_id":"a1b2c3","host":"acme-shop.example","method":"GET","path":"/products/42","query":"","status":200,"bytes":5120,"request_time":0.042,"referrer":"https://www.google.com/","user_agent":"Mozilla/5.0 (Windows NT 10.0)","anonymous_id":"b7e1c9d2-48a0-4b52-8f4a-0c5a9a1d3e77"}
{"time":"2025-10-02T09:15:31+00:00","request_id":"a1b2c4","host":"acme-shop.example","method":"GET","path":"/assets/app.js","query":"","status":200,"bytes":9000,"request_time":0.003,"referrer":"","user_agent":"Mozilla/5.0","anonymous_id":""}
{"time":"2025-10-02T09:16:02+00:00","request_id":"a1b2c5","host":"acme-shop.example","method":"GET","path":"/products/42","query":"","status":200,"bytes":5120,"request_time":0.05,"referrer":"","user_agent":"Googlebot/2.1","anonymous_id":""}
{"time":"2025-10-02T09:20:00.500Z","level":30,"event_name":"purchase_completed","user_id":"u_981","anonymous_id":"b7e1c9d2-48a0-4b52-8f4a-0c5a9a1d3e77","properties":{"order_id":"ord_1001","total_cents":4999}}
not json

Save those lines as sample.log and run node logs-to-events.js < sample.log. You get two events on standard output (one page_view and one purchase_completed) and the line kept=2 skipped=3 on standard error. The static file, the bot and the broken line are all dropped, which is the behavior you want.

Handling the Hard Parts

Bots

Access logs contain every crawler, uptime monitor and vulnerability scanner that touches your server. The user-agent filter above removes the polite ones that identify themselves. It misses scrapers that pretend to be Chrome. Expect a residue of fake traffic, and review the top user agents monthly. Keep the filter in one place so you can tighten it without touching the shipper.

Sessions and identity

Notice session_id: null in the output. A log line has no 30-minute session concept, and the server cannot invent one reliably. You derive sessions later in SQL by grouping events per anonymous_id with a 30-minute gap, which the SQL article demonstrates.

Identity is the weaker spot. If the request carries no anonymous_id cookie, as with the first page view of a new visitor, the converter writes unknown. Exclude those rows from unique-visitor counts. If you want an identifier without cookies, cookieless analytics covers daily salted hashes and their limits.

Duplicates

Shippers deliver at least once. After a crash, Vector may resend lines it already sent. A random event_id would make every resend look like a new event. Both the Vector config and the Node converter therefore hash request_id into a UUID-shaped string, so the same request always gets the same id. The server-side deduplication that uses it appears in the async ingestion article.

Privacy and retention

Logs accumulate personal data quietly. Query strings may contain emails from password reset links. Referrers may contain search terms. The pipeline above drops query strings from page_url and never logs IP addresses, but your default nginx log still does. Rotate it quickly, restrict access to it, and document the retention period.

Choosing a Log Pipeline

Vector is a reasonable default, not the only option. The right choice depends on what you already run.

Approach Best for Cost
Vector (this article) One agent with built-in transforms, many destinations You write VRL and manage the agent
Fluent Bit Very small footprint, Kubernetes setups Transform logic is less expressive
Filebeat with Logstash or Elasticsearch ingest Teams already on the Elastic stack Heavier to run, more memory
rsyslog forwarding Simple line forwarding on Linux, no transforms You still need a parser downstream
Cloud logging (for example CloudWatch Logs) Apps already hosted on that cloud Per-GB ingest pricing, and lock-in
Scheduled script on log files One server, low volume No rotation or restart handling

Pick on operational fit, not feature lists. If your team already operates Fluent Bit, its output plugins can reach the same collector, and adding Vector buys little. The mapping logic in this article transfers to any of them, because the hard part is the event contract, not the transport.

How Real Systems Do This

Log-based analytics is older than JavaScript tracking. Early tools such as AWStats read web server logs in batch. GoAccess still does, and it can render a terminal or HTML report from an nginx log in seconds, which makes it a good sanity check for the numbers your pipeline produces. Matomo ships a log import tool that replays server logs into its tracking API, so teams can backfill history or track sites where they cannot add a script.

Content delivery networks follow the same pattern at larger scale. They write request logs for every edge hit, and customers load those logs into a warehouse for traffic analysis. The approach works because a request log is the one record nobody can opt out of at the browser layer.

Commercial tools mostly default to browser scripts, because logs lack interaction data. However, the better ones accept server-side events through an API, which is the same idea you implemented with logPurchase(). Combining both sources, with logs as the audit trail, is a mature design.

Decision Framework

  1. Do you need numbers that survive ad blockers? Use logs as a source or audit trail for page views.
  2. Do you need button clicks, scroll depth or form field friction? Logs cannot see them, so keep the browser tracker.
  3. Do you run more than one server or container? Centralize with a shipper before logs vanish with the instance.
  4. Is the event a business fact, such as a purchase? Log it from the application, once, at the moment it is confirmed.
  5. Do you already operate a log stack? Reuse its agent and port the mapping logic.
  6. Does the line contain personal data? Drop or hash it before it leaves the server.

When NOT to Use This

  • You need interaction analytics. Clicks, hovers and scrolls never reach the server. Use the browser tracker from the previous article.
  • Your site sits behind a cache or CDN that serves most pages. Cached responses never reach your origin, so your origin logs undercount. Pull the CDN’s own logs instead.
  • You run one small site and want a quick report. Run GoAccess on the log file and stop. Building a shipper for 500 visits a month is overkill, and a hosted analytics product costs less than your time.

Common Mistakes

  • Using the default log format. Regex parsing breaks on odd user agents, and your counts drift without any error.
  • Forgetting escape=json. A single quote in a user agent produces invalid JSON and the shipper drops the line.
  • Counting static assets as page views. One page load becomes twenty rows, and page views inflate by an order of magnitude.
  • Leaving read_from at its default. A fresh agent replays the whole file and duplicates yesterday’s events.
  • Treating log lines as exactly-once. Shippers retry, so duplicates appear without a stable event id.
  • Keeping IP addresses in the analytics log. You store personal data you never analyzed and extend your compliance burden.

Key Takeaways

  • Server logs record every request, so they are a second event source that blockers cannot hide.
  • Write logs as one JSON object per line, and use escape=json in nginx.
  • Log business facts like purchases from the application, once, at confirmation time.
  • Use one shipper (Vector here) to read files and forward to a single collector.
  • Filter bots, static files and non-2xx responses before they become events.
  • Expect no session data and weak identity in logs, and derive sessions later in SQL.
  • Keep IP addresses and query strings out of the analytics pipeline.

FAQ

How do I centralize server logs for a web app?

Write logs as JSON, install one log shipper such as Vector on each server, and forward the lines to a single destination. The shipper tracks file positions, survives restarts and handles rotation. Choose the destination based on how you will query the data.

Can server logs replace JavaScript analytics?

Not fully. Logs capture every request but miss clicks, scrolling and other browser interactions. They also include bot traffic. They work best as an audit source next to a browser tracker, or as the main source for simple page-view reporting.

What is the best nginx log format for analytics?

A JSON format with escape=json, containing time, request id, host, method, path, query, status, bytes, request time, referrer, user agent and the analytics cookie. Leave out the client IP address unless you have a documented need for it.

Should I use Vector, Fluent Bit or Logstash?

Use the agent your team already runs. Vector is a good default for new setups because it combines collection and transformation in one binary. Fluent Bit suits tight memory budgets, and Logstash suits teams already invested in Elasticsearch.

Conclusion

Logs already hold your traffic history. When you centralize server logs, write them as JSON, ship them through one agent and map them into your event shape, you get a second view of the truth for very little code. Use it to audit the browser, and use the browser to fill the gaps logs cannot see.

Rule of thumb: the browser tells you what users did, the server tells you what actually happened, and you need both to know which numbers to believe.

Last updated on 9 October 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *