Posted on September 30, 2026 by Deepak Biswas, Senior Engineering Manager, Central Monitoring and Disaster Recovery, Atlassian Every SaaS company has the same uncomfortable question after a major incident: who noticed first, the monitoring or the customers? For a long time our honest answer was "it depends".
This post describes how a small team rebuilt the detection platform behind our automated incident creation ("AutoHOT", where HOT is our internal name for a high-severity incident) on Apache Kafka, Apache Flink on Kubernetes and OpenTelemetry , what we measured month by month for eighteen months, and what we would do differently. It is not a success story with a bow on it. In-scope recall went from about 60% to a peak of 86%, then fell back to 64% in a bad month. Precision is still below where we want it. We think the numbers are more useful than the slogans.
We run more than ten cloud products for millions of tenants. Each product emits client-side operational telemetry for every user action: an event when a task starts, and one when it succeeds, fails or is abandoned, tagged with tenant, user, the named "experience" (view issue, edit page, view board) and the HTTP status. Across all products, there are billions of events a day, and it is the closest thing we have to ground truth about whether customers can actually do their work. The job of central monitoring is to turn that stream into three answers, fast: The first-generation system did this with a Node.
js aggregator fed from the event bus through a cloud queue, with an in-memory cache for de-duplication and task pairing, running on roughly 90 virtual machines. It worked, and it taught us the domain. It also had three problems we could not tune away: We also had a correctness problem that only became visible once we had better data. Client-side telemetry only exists when a page loads. A hard-down database shard produces silence , not errors. Any design that only looks for failures will read silence as health.
