Observatory Self-Healing

Detect. Recover. Verify.

Bring monitoring, investigation, and recovery into a connected operational workflow. Observatory helps teams respond through approved actions, service-aware policies, and clear evidence of recovery.

  1. 01DetectOperational symptoms, checked against normal.
  2. 02DiagnoseCorrelated signals and a likely cause, with evidence.
  3. 03AuthorizePolicy, permissions and approval where required.
  4. 04RecoverThe authorized workflow runs, step by step.
  5. 05VerifyChecks confirm the service works again.

A complete recovery journey

See every step toward service restoration.

Each recovery follows the same path, from the first symptom to the evidence that the service is back, and every step is recorded as it happens.

  1. Detect

    Detect operational symptoms: checks move through soft and hard states against seasonal baselines before an incident opens.

  2. Investigate

    Correlate signals and investigate likely causes by walking the dependency graph, with the evidence attached.

  3. Select

    Select an applicable recovery workflow for the problem and the affected component.

  4. Authorize

    Apply permissions and approval requirements from the recovery policy that covers the service.

  5. Execute

    Execute the authorized response. Milestone: action completed.

  6. Verify

    Verify functionality and keep observing stability. Milestone: recovery verified.

  7. Retain

    Retain the evidence, the decisions and the approvals, and keep reviewed recovery knowledge for next time.

Milestone one

The action completed.

A restart finished, a workload was re-queued, a replica was scaled. That says the step ran. It does not yet say the service is back.

Milestone two

The recovery is verified.

The step's verification checks passed: the resource answers, the application responds and the check that raised the alarm is healthy again. If verification fails, the step fails.

Choose your level of automation

Automation that follows your operating policy.

Set different policies for development, testing, and production services. Keep operators informed about what is permitted, what is running, and what needs attention.

Mode 1

Observe

Collect evidence and identify recovery opportunities. Nothing is changed.

Observatory watches

Mode 2

Recommend

Prepare a proposed response for review. A person decides whether and how to act.

A person acts

Mode 3

Approve

Execute after authorized approval. The person who approves cannot be the person who asked.

A person approves

Mode 4

Automate

Run approved recovery workflows within defined limits. Destructive steps still wait for a person.

Within your limits

  • Policy can be set broadly and narrowed down to a single host, with more specific settings taking precedence.
  • Limits include a risk ceiling, allowed and denied action categories, freeze windows and a cap on retries.
  • Every run is recorded: each step, the decision behind it, the approval and the result of its verification.

Recovery built around services

Understand dependencies before making changes.

Infrastructure actions can affect multiple applications and teams. Connect recovery decisions to service dependencies, the blast radius of a change, the healthy capacity that remains, and other changes already in flight.

  • The dependency graph shows which applications and business services sit on the component you are about to touch.
  • An action can be held back when it would take out the last healthy replica, or when another action is already running on the same host.
  • Protected infrastructure can be excluded from automatic recovery altogether.

Conceptual diagram

One restart, several dependants A shared database host supports three applications, which support two business services. Before a restart of the host, the diagram marks every application and service in its blast radius, a healthy replica that must stay up, and a change already in flight elsewhere. Business services Online orders Customer portal Applications Web front end Order service Accounts API Infrastructure Replica Ahealthy · keep Database hostproposed: restart Cache nodechange in flight In the blast radius of the restart Must stay up Concurrent change
Conceptual diagram, not a product screen: how a recovery decision takes dependants, remaining capacity and concurrent changes into account.

Human and AI collaboration

Use proven workflows. Bring in AI when investigation needs more help.

Known incidents can follow approved recovery workflows. For unfamiliar problems, AI-assisted investigation can help analyze evidence and propose next steps within configured permissions.

Known problem

  1. Incident matches a known pattern
  2. Approved recovery workflow
  3. Policy, approval if required
  4. Execute and verify

Deterministic. No language model involved.

Unfamiliar problem

  1. AI-assisted investigation of the evidence
  2. Proposed next steps, within configured permissions
  3. A person reviews
  4. Approved action, then verification

Optional, and set per environment. In production it recommends by default.

New recovery knowledge is reviewed before it becomes an approved automatic response.

A workflow drafted with AI starts disabled. It earns trust through testing and review, and only a trusted workflow may start automatically. Any edit sends it back for review.

Recovery visibility

Know what is running, what is waiting, and what worked.

Operators follow each recovery as it happens: the proposed action, who approved it, every step, and whether verification passed. The scene below is a conceptual stage view, not a product recording. We show the real screens in a demo.

  1. 01DetectA fault appears
  2. 02DiagnoseEvidence, likely cause
  3. 03AuthorizePolicy and approval
  4. 04RecoverThe approved action runs
  5. 05VerifyThe service works again
  1. A dependency develops a visible fault.
  2. The lighthouse beam reveals the affected service path.
  3. An investigation card sets out the evidence and the likely cause.
  4. A recovery-policy checkpoint shows what is permitted and who must approve.
  5. The approved action progresses and completes.
  6. Verification checks run against the resource, the application and the user journey.
  7. The service returns to a stable state, and stays under observation.
Illustrative recovery workflow: a conceptual stage view, not a product screen or customer data.

What a demo walks through

  • Recovery operations in progress, in one automation view.
  • Approval requests waiting for a person.
  • Each run step by step, with its decision and progress.
  • The verification result of every step.
  • Recovery history and the audit trail behind it.

We would rather show you the product than a mock-up of it. Bring a recovery you run by hand today, and we will walk through how it would look in Observatory.

Ask for a demo

Infrastructure and AI workflows

Recovery for what you run today, and for the AI you are adding.

Available

Infrastructure and applications

Service failures, dependency issues, configuration problems, and approved operational recovery: restarting services and workloads, re-queuing stuck work, scaling, and running your existing automation, each under recovery policy.

In development

AI workflows

Model endpoint failures, tool errors, stalled flows, and policy-approved recovery paths through AFPM. Today, AFPM gives AI flows the visibility that any recovery starts from. Recovery for AI workflows is in development and is not yet available.

Any model change must follow configured policy and remain visible to the operator.

That is the rule recovery for AI workflows is being built to, and the same rule that already governs recovery for infrastructure: policy first, and nothing hidden from the operator.

Fit into your environment

Recover through what is already there.

Recovery runs through the Spectient agent where it is installed, and through supported agentless integrations where it is not.

Available recovery actions depend on the connected platform, installed capabilities, and granted permissions.

With Spectient

Actions run through the agent on the host, within the capabilities it has been granted.

Agentless

SSH and WinRM with allow-listed command templates, plus supported platform integrations for containers, Kubernetes, virtualisation, automation and out-of-band management.

Observatory itself

A watchdog checks that Observatory's own work is progressing, not only that its processes are up, and applies a fixed list of non-destructive fixes. It never deletes.

Questions

Does every recovery require AI?

No. Known recovery workflows and playbooks run without a language model. AI-assisted investigation can be used when it is enabled and appropriate; it is set per environment, and in production it only recommends by default.

Can production changes require approval?

Yes. Recovery policy can require approval and restrict which actions are permitted for production services. Destructive steps always wait for a person, and the person who approves cannot be the person who requested the action.

How is recovery verified?

Each recovery step can carry verification checks, such as reachability, an HTTP or DNS response, the state of a service or process, a metric, a log pattern, or the state of an existing monitoring check. If verification fails, the step fails. What is checked depends on the configured service, and the same monitoring that raised the alarm keeps watching afterwards.

Can Observatory work without Imperium?

Yes. Observatory has its own recovery policies, approvals and audit trail, and needs no other Tracston product. Imperium integration today covers AI flow events in AFPM; governing AI-assisted recovery through Imperium is not part of the current release.

Does recovery permanently fix every problem?

No. Some actions restore availability while a permanent correction still requires investigation. The incident, the actions taken and their verification stay on record, so a recurring problem remains visible and can be followed up.

Recovery needs evidence.

Bring us a recovery you still run by hand.

We will walk through how Observatory would detect it, who would approve it, and how you would know it worked.