Observe · Observatory

See through the storm.

Observatory is Tracston's observability platform. It watches your infrastructure, networks, applications and AI agents in one place, turns a flood of alerts into a handful of incidents, and traces each one from the symptom to its root cause, with the evidence attached.

It runs inside your own infrastructure, on its own or next to the rest of the Tracston platform.

  • One viewhosts, network, apps, AI flows
  • Many alerts → one incidentcorrelated, with its working shown
  • Symptom → causewalked through the dependency graph

The problem

Modern systems are getting harder to navigate.

A single customer request now crosses AI agents, models, APIs, microservices, Kubernetes, databases, cloud services and the network underneath all of it. Every layer has its own monitoring tool, and each one raises its own alarms.

AI agentsLLMsMCP serversApplicationsAPIsMicroservicesKubernetesCloudSaaSDatabasesNetworksUsersAutomation

More signals don't create more clarity.

Twelve alarms, one fault

Classic monitoring opens an incident per failing check. A host with twelve failing services produces twelve headlines, and the person on call has to work out that they are one problem.

Nobody says the obvious thing

When a whole site goes dark, ten hosts each report their own outage. Ten quiet alerts are not one loud one, and the sentence the operator needs ("the site is off") is never written.

AI work is invisible

An AI agent can return a clean HTTP 200 while the task it was given fails. Application monitoring sees the request, not the prompt, model, tool calls and policy decisions behind it.

How it works

Turn telemetry into direction.

Observatory collects from everything you run, builds a map of how it depends on itself, and uses that map to decide which alarms belong together and where they started.

Observatory data flow Sources feed collectors, listeners, host agents and the OpenTelemetry receiver. The control plane evaluates checks, builds the dependency graph and correlates incidents, then serves the console, alerts, the assistant and an API. Your estateCollectionControl planeOutput Hosts & VMs Network devices Apps & services Kubernetes & cloud AI agents (AFPM) Collectors · checks Agentless · SSH / SNMP Listeners · syslog, traps Host agent · OTLP Imperium AI events Observatory Evaluatechecks · baselines Mapestate · dependency graph Correlatealerts → incidents Explainroot cause · evidence Console & maps Alerts & on-call Ask the lighthouse API · status pages
Your estateHosts, network devices, apps, Kubernetes and cloud, AI agents
CollectionCollectors and checks, agentless SSH/SNMP, listeners for syslog and traps, host agent, OpenTelemetry, Imperium AI events
ObservatoryEvaluate · map · correlate · explain
OutputConsole and maps, alerts and on-call, the assistant, API and status pages
Collectors go and fetch on a schedule; listeners take whatever arrives unsolicited; everything lands in one control plane.
  1. Collect, with or without an agent

    Collectors run scheduled checks, and reach hosts that will never run an agent over SSH and SNMP instead. Listeners accept syslog, SNMP traps, webhooks and network flows, buffering to disk so an outage of the control plane loses nothing. An optional host agent sends inventory and vitals, and any OpenTelemetry SDK can send traces straight in.

  2. Discover and map the estate

    Scheduled discovery keeps every run as a version, so you see what is new, what changed and what is gone. Nothing becomes monitored just by being found: findings wait for a person to promote them. The estate is modelled as a hierarchy (site, hypervisor, VM, container, cluster) and as a dependency graph that records which way a failure travels.

  3. Evaluate against what is normal

    Every check moves through soft and hard states before it opens an incident. Seasonal baselines, kept per hour of the week, tell a nightly batch job apart from a real spike at two in the afternoon.

  4. Correlate alerts into incidents

    Independent correlators each propose links between incidents: connected in the dependency graph, the same host, followed the same deployment, the same product, the same subnet, or started at the same moment. Links add up, and no single weak signal can group two incidents on its own. Every grouping shows which correlator did the work.

  5. Explain the root cause

    The dependency graph is walked both ways: forward for the blast radius (what else breaks) and backward for the root-cause candidates. The investigation gathers the state changes, metrics, logs and related incidents behind it, so the conclusion arrives with its evidence.

  6. Tell the right people, then verify

    Alerts route to the owning team through email, chat, SMS, voice or webhooks, with storm protection that folds bursts into a digest. Remediation runs only through approved playbooks, and the same checks that raised the alarm show the recovery.

The storm, then the incident

Don't count alerts. Understand incidents.

One component fails and the failure travels: database, API, payments, checkout, AI assistant, customer. Every layer raises its own alarm. Observatory folds the storm into one incident and follows it down to where it started.

57alerts
18related events
5affected components
1incident
1root cause

Illustrative scenario, not customer data.

Root cause is a dive, not a list.

Observatory starts where the customer feels it and follows the dependency graph down, layer by layer, until it reaches the component that broke first.

Ask the lighthouseWhy is checkout slow?

Checkout latency increased at 14:32. Primary cause: PostgreSQL connection pool exhaustion. Contributing factor: the AI reconciliation agent's concurrency rose from 4 to 32. Affected: Checkout, Payments, Customer Portal, AI Assistant.

  • Observation
  • Fact
  • Correlation
  • Inference
  • Recommendation
  • Confidence

AI with operational evidence, not "AI knows everything". Every statement is labelled by what kind of claim it is.

  1. Business serviceCheckoutresponse time up; customers affected
  2. Applicationcheckout-weberrors follow the API
  3. APIInventory APIlatency correlated with checkout
  4. AI agentreconciliation agentconcurrency 8× its baseline
  5. MCP · toolinventory-mcpcalls queue behind the database
  6. Databasepostgres-primaryconnection wait +430%, new deployment 4 minutes earlier
  7. Root causeConnection pool exhaustion6 services affected · evidence attached

Capabilities

Visibility at every depth.

AI doesn't run in isolation. Observatory sees every layer underneath it, from the people and business services at the surface to the network and physical infrastructure on the ocean floor.

Surface

Humans · business services · AI agents

Near surface

Applications · APIs · AI services

Middle

Containers · Kubernetes · databases

Deep ocean

VMs · hosts · storage

Ocean floor

Network · physical infrastructure

01Estate, discovery and topology

Versioned discovery

Scheduled network scans keep every run, so the question "what is new, changed or gone since last time" has a direct answer. Found assets stay pending until someone promotes them.

Estate hierarchy

Site, hypervisor, VM, container and cluster, built from your CMDB, from what the container runtime reports, and from inference, with each source kept distinct.

Maps that are earned

The network map is derived from the addresses your checks actually probe, and a host takes the worst state among its services. Sites appear on a geographic map only when they have real coordinates.

Network topology

LLDP, CDP, ARP, forwarding and route tables over SNMP build physical, access and routed layers, with run-to-run changes recorded as events.

Inventory and vulnerabilities

Host inventory is snapshotted and diffed. Installed packages are matched against OSV advisories, with CISA's known-exploited list on top.

Kubernetes

Nodes, namespaces, workloads and the pods that are not healthy, read through a credential you supply once and keep sealed.

02Metrics, logs and traces

Checks and metrics

Built-in check plugins for ping, TCP, HTTP, DNS, certificates, SSH, SNMP, Docker, disk, PostgreSQL, MySQL, Redis, MongoDB, Prometheus and node exporters, and Nagios-style scripts, plus a metrics explorer and seasonal baselines.

Logs

Search logs with a Loki-style query language. Logs stay in PostgreSQL by default, or move to a ClickHouse tier for estate-wide volume by configuration, not by code change.

Traces and APM

An OpenTelemetry (OTLP/HTTP) receiver accepts spans from standard SDKs. Observatory builds the service map, request rate, errors, duration and Apdex per service.

Deep trace

One trace document from request spans down through protocol, connection, process and kernel activity on the host, where the deep-tracing layer is deployed.

Network flows

NetFlow v5 and v9, IPFIX and sFlow feed a traffic explorer: top talkers, conversations and per-interface views, enriched with device and site.

Synthetic checks

HTTP synthetic tests run from built-in monitoring locations, so you learn about a broken journey before a customer does.

03AI flows: AFPM

AI Flow & Performance Monitoring follows AI execution from the first prompt through models, agents, tools, MCP servers, retrieval systems, applications and infrastructure.

  • Each step of a flow on one timeline, so the slow hop is obvious.
  • Policy decisions next to the execution: allowed, warned or denied.
  • The infrastructure under the AI in the same view as the AI itself.
  • When Imperium is deployed, AFPM reads its governed AI events. Observatory itself does not require Imperium.
Read: Who is watching the AI? Introducing AFPM →

Illustrative tracetotal 2.33 s

  1. Prompt12 ms
  2. Imperium8 ms
  3. Routing4 ms
  4. Model842 ms
  5. Agent184 ms
  6. MCP38 ms
  7. Tool420 ms
  8. Database92 ms
  9. Model730 ms

The lighthouse finds the bottleneck: model routing, 842 ms.

04Incidents, alerting and response

Correlation and stories

Alerts become incidents, and incidents that share a cause become one persistent story, such as "collector lost, 25 targets unobserved" rather than 25 cards. A whole site going dark is said once.

War room

During an incident the response is recorded as structured events (roles, tasks, decisions, updates), so the post-incident review is generated rather than rewritten from memory.

Notifications and on-call

Email, Slack, Microsoft Teams, Telegram, SMS, voice, push, webhooks and n8n, with on-call directories and escalation. Maintenance windows and suppressions are logged with their reason.

Storm protection

Rate limits per channel and per recipient. When a limit is hit, messages are gathered into one digest rather than dropped, and a global kill switch stands behind it.

Business services and status pages

Roll technical services up to the business services and customers that depend on them. Status page components can follow live incidents automatically, and a separate customer portal shows status without exposing infrastructure details.

Security detections

A detection engine runs rules over host events, inventory changes and vulnerability matches, de-duplicating repeats into one finding with a count. Sigma rules can be imported; anything that cannot be expressed is imported disabled, with the reason.

05AI assistant and controlled action

Ask the lighthouse

Ask in plain language: what changed, why did this host alert, is anything suppressed right now. The question is routed in code to a fixed set of read-only tools over live data, and every answer carries the rows it came from.

Works with no model at all

A language model only rephrases what the tools found. With the model off, expired or unreachable you still get the facts, and during an incident many operators prefer them. You can use local models or a hosted provider.

Playbooks under control

Each playbook runs in one of four modes: observe, recommend, require approval, or automatic. Automatic mode is refused for anything destructive, and the person who approves cannot be the person who asked.

Inside the console

Observatory War Room: the lead story groups eight incidents on three services into one, with targets, services, sites and SLA risk, the service dependency graph, priority score history, why it is a top priority, the next best action and live activity.

One lighthouse

Humans and AI agents share one picture.

People get incidents, root causes, evidence and recommended actions they can read. AI agents get the same operational reality as structured context they can safely use to investigate. Nothing here replaces your judgement: Observatory explains and recommends, and an approval-gated action is the only thing it may execute.

Engineer

  1. Observatory console
  2. Investigation
  3. Decision

AI agent

  1. Observatory API · context
  2. Operational evidence
  3. Safe action

Where Observatory sits in the Tracston story

Someone notices checkout is slow. Observatory finds the root cause. Otopia prepares the fix, Imperium checks it against policy, a person approves, and Observatory confirms the recovery.

  1. DetectedObservatory
  2. UnderstoodObservatory
  3. GovernedImperium
  4. ApprovedA person
  5. ExecutedApproval-gated action
  6. VerifiedObservatory

Deployment and security

Runs where your estate runs.

Observatory is deployed inside your infrastructure or a private environment. The monitoring that matters most is the monitoring that still works when the internet does not.

Standalone

On its own

A complete product with its own console, users and API. It needs no other Tracston product.

Together

With the platform

Feeds customer and management views in ResolveSystem, and connects to Imperium for AI flows when both are deployed.

Customer-managed

Inside your walls

Installed on-premises or in your private cloud, including sites without internet access.

Small to start, room to grow

  • An all-in-one container with an embedded PostgreSQL runs a complete instance.
  • Every component (control plane, collectors, listeners, agents) can also run as its own container.
  • Logs, traces and flows can move to ClickHouse, and metric history to VictoriaMetrics, by configuration.
  • Collectors sit at the edge of each network and talk outbound to the control plane.
  • Maps and the assistant are built to work offline: map code is bundled, and answers do not depend on a model.

Built to be believed

  • The host agent runs as an unprivileged user, opens no listening socket, and accepts no shell or plugin download from the server.
  • Collectors enrol with a one-time token that becomes a permanent credential stored locally with owner-only permissions.
  • Credentials are sealed and write-only in the console; keys are never shown back or logged.
  • Sign-in with local accounts or single sign-on (OIDC), with changes recorded in an audit log.
  • The customer portal cannot query addresses, credentials or check parameters: redaction is structural, not a filter.
  • The self-healer may reset, re-queue or re-derive, and never delete. Its capabilities are a fixed list in code, not rows in a database.

Capabilities are enabled according to deployment and licensing.

Integrations

Speaks the protocols you already run.

Standards first, so you can point what you already have at Observatory instead of re-instrumenting.

Telemetry in

OpenTelemetry OTLP/HTTP (JSON & protobuf)Prometheus & node exporterNagios-style pluginsAlertmanager webhooksGeneric JSON webhooksSyslog RFC 3164 / 5424

Network

SNMP v1 / v2c / v3SNMP trapsNetFlow v5 / v9IPFIXsFlow v5LLDP · CDP

Platforms and data

Kubernetes APIDockerPostgreSQLMySQLRedisMongoDBClickHouseVictoriaMetrics

Notify and act

EmailSlackMicrosoft TeamsTelegramSMSVoiceWebhooksn8n

Security feeds

OSV advisoriesCISA KEVSigma rules

AI and Tracston

Local models (Ollama, vLLM)OpenAIAnthropicGeminiTracston ImperiumResolveSystem

Who it is for

For the people who keep the lights on.

Operations and NOC teams

One console for hosts, services, network and sites, with alerts already grouped into incidents and a war room for the bad days.

SRE and platform engineers

OpenTelemetry traces, a service map, logs and metrics next to the infrastructure they run on, and a dependency graph for blast radius.

Network engineers

SNMP polling and traps, discovered L1–L3 topology, flow analytics and a network event timeline.

Teams running AI in production

AFPM shows what agents and models actually did, how long each step took, and which policy decisions applied.

Security operations

Detections, imported Sigma rules, vulnerability matches on real inventory, and approval-gated playbooks with separation of duties.

Service owners and management

Business services, SLA state, status pages and a customer portal that says what is affected in words the business uses.

Questions

What is Observatory, in one sentence?

An observability platform that monitors infrastructure, networks, applications and AI agents in one place, groups related alerts into incidents, and explains each incident's root cause with the evidence behind it.

Do I have to install an agent everywhere?

No. Collectors check hosts and devices from outside, over protocols such as ping, HTTP, SSH and SNMP, and listeners accept syslog, traps, webhooks and flows. The host agent is optional, for when you want local inventory and vitals from a Linux machine.

Does it work with OpenTelemetry?

Yes. Observatory runs an OTLP/HTTP trace receiver that accepts JSON and protobuf, so a standard SDK can send spans by pointing its OTLP endpoint at Observatory. Ingest keys carry the organisation, product and environment the spans belong to.

How does it decide that alerts belong together?

Several independent correlators each propose weighted links: a real dependency between the components, the same host, the same deployment, the same product, the same subnet, or the same moment. The weights add up, a single weak signal is never enough on its own, and each group shows which correlator was decisive.

How does it find the root cause?

It keeps a dependency graph with direction: if A depends on B, a fault in B travels to A. Walking forward gives the blast radius; walking backward gives the root-cause candidates. The investigation then attaches the state changes, metrics, logs and related incidents that support the conclusion.

Does the AI take actions on its own?

No. The assistant uses read-only tools and cannot write to the database. Remediation goes through playbooks with explicit modes; destructive actions can never run automatically, and an approver cannot approve their own request.

Does Observatory need Imperium or Otopia?

No. Observatory runs standalone. When Imperium is also deployed, AFPM can read its governed AI events so AI activity appears alongside the rest of the estate.

What is AFPM?

AI Flow & Performance Monitoring: observing the complete execution path of AI-driven work across prompts, agents, models, tools, infrastructure and outcomes. Our Insights article explains the idea in depth.

Where does it run, and can it run without internet access?

Inside your own infrastructure or a private environment. It starts as a single all-in-one container and can be split into separate components. Map code is bundled rather than fetched from a CDN, and the assistant answers from live data even with no model available.

Autonomy needs visibility.

Find the cause before the storm finds your customers.

See Observatory on its own site, or ask us to walk you through it on your estate.