Mode 1
Observe
Collect evidence and identify recovery opportunities. Nothing is changed.
Observatory watches
Observatory Self-Healing
Bring monitoring, investigation, and recovery into a connected operational workflow. Observatory helps teams respond through approved actions, service-aware policies, and clear evidence of recovery.
A complete recovery journey
Each recovery follows the same path, from the first symptom to the evidence that the service is back, and every step is recorded as it happens.
Detect operational symptoms: checks move through soft and hard states against seasonal baselines before an incident opens.
Correlate signals and investigate likely causes by walking the dependency graph, with the evidence attached.
Select an applicable recovery workflow for the problem and the affected component.
Apply permissions and approval requirements from the recovery policy that covers the service.
Execute the authorized response. Milestone: action completed.
Verify functionality and keep observing stability. Milestone: recovery verified.
Retain the evidence, the decisions and the approvals, and keep reviewed recovery knowledge for next time.
Milestone one
A restart finished, a workload was re-queued, a replica was scaled. That says the step ran. It does not yet say the service is back.
Milestone two
The step's verification checks passed: the resource answers, the application responds and the check that raised the alarm is healthy again. If verification fails, the step fails.
Choose your level of automation
Set different policies for development, testing, and production services. Keep operators informed about what is permitted, what is running, and what needs attention.
Mode 1
Collect evidence and identify recovery opportunities. Nothing is changed.
Observatory watches
Mode 2
Prepare a proposed response for review. A person decides whether and how to act.
A person acts
Mode 3
Execute after authorized approval. The person who approves cannot be the person who asked.
A person approves
Mode 4
Run approved recovery workflows within defined limits. Destructive steps still wait for a person.
Within your limits
Recovery built around services
Infrastructure actions can affect multiple applications and teams. Connect recovery decisions to service dependencies, the blast radius of a change, the healthy capacity that remains, and other changes already in flight.
Conceptual diagram
Human and AI collaboration
Known incidents can follow approved recovery workflows. For unfamiliar problems, AI-assisted investigation can help analyze evidence and propose next steps within configured permissions.
Known problem
Deterministic. No language model involved.
Unfamiliar problem
Optional, and set per environment. In production it recommends by default.
New recovery knowledge is reviewed before it becomes an approved automatic response.
A workflow drafted with AI starts disabled. It earns trust through testing and review, and only a trusted workflow may start automatically. Any edit sends it back for review.
Recovery visibility
Operators follow each recovery as it happens: the proposed action, who approved it, every step, and whether verification passed. The scene below is a conceptual stage view, not a product recording. We show the real screens in a demo.
We would rather show you the product than a mock-up of it. Bring a recovery you run by hand today, and we will walk through how it would look in Observatory.
Ask for a demoInfrastructure and AI workflows
Available
Service failures, dependency issues, configuration problems, and approved operational recovery: restarting services and workloads, re-queuing stuck work, scaling, and running your existing automation, each under recovery policy.
In development
Model endpoint failures, tool errors, stalled flows, and policy-approved recovery paths through AFPM. Today, AFPM gives AI flows the visibility that any recovery starts from. Recovery for AI workflows is in development and is not yet available.
Any model change must follow configured policy and remain visible to the operator.
That is the rule recovery for AI workflows is being built to, and the same rule that already governs recovery for infrastructure: policy first, and nothing hidden from the operator.
Fit into your environment
Recovery runs through the Spectient agent where it is installed, and through supported agentless integrations where it is not.
Available recovery actions depend on the connected platform, installed capabilities, and granted permissions.
Actions run through the agent on the host, within the capabilities it has been granted.
SSH and WinRM with allow-listed command templates, plus supported platform integrations for containers, Kubernetes, virtualisation, automation and out-of-band management.
A watchdog checks that Observatory's own work is progressing, not only that its processes are up, and applies a fixed list of non-destructive fixes. It never deletes.
No. Known recovery workflows and playbooks run without a language model. AI-assisted investigation can be used when it is enabled and appropriate; it is set per environment, and in production it only recommends by default.
Yes. Recovery policy can require approval and restrict which actions are permitted for production services. Destructive steps always wait for a person, and the person who approves cannot be the person who requested the action.
Each recovery step can carry verification checks, such as reachability, an HTTP or DNS response, the state of a service or process, a metric, a log pattern, or the state of an existing monitoring check. If verification fails, the step fails. What is checked depends on the configured service, and the same monitoring that raised the alarm keeps watching afterwards.
Yes. Observatory has its own recovery policies, approvals and audit trail, and needs no other Tracston product. Imperium integration today covers AI flow events in AFPM; governing AI-assisted recovery through Imperium is not part of the current release.
No. Some actions restore availability while a permanent correction still requires investigation. The incident, the actions taken and their verification stay on record, so a recurring problem remains visible and can be followed up.
Recovery needs evidence.
We will walk through how Observatory would detect it, who would approve it, and how you would know it worked.