Building log0: a multi-tenant incident platform on Kafka
A working incident-management backend, built solo on one laptop, explained from the first measured number to the last honest caveat. This post is the map for the whole series.

Search for a command to run...

Series
Building log0, a multi-tenant incident-management platform on Kafka, ClickHouse and PostgreSQL. Seven Spring Boot services, built solo on one laptop. Each post takes one box on the architecture diagram or one core design decision, and goes deep: the naive version, why it broke (with a measured number), the fix, the real code, and the tradeoff. Honest load-test results, including the unflattering ones.
A working incident-management backend, built solo on one laptop, explained from the first measured number to the last honest caveat. This post is the map for the whole series.

A liveness probe and a real request take different paths through a service, and the gap between them hides a whole class of bug. The first real POST to log0's gateway hung with no error and no log lin

The core job of an incident platform is deciding that a flood of log lines is one bug, not ten thousand. log0 does it with no machine learning, deterministically, in O(1) per event. The whole trick is

A fingerprint names a bug; ten of them inside five minutes is an incident. How log0's clustering service counts in tumbling event-time windows, emits exactly once per window, and the two failure modes of an in-memory store with no eviction.

A logging endpoint has one job it can never fail at: being available to accept logs. log0's ingestion gateway returns 202 Accepted the instant the event is handed to a Kafka producer future-before the broker acks-and never waits for clustering, a database write, or Slack.
