BrutlabsLLC Book a consult
Agent capability

Self-healing, with a human in the loop until you say otherwise

Codified runbooks the agent can execute against conditions it recognises — bounded by approval gates, blast-radius policy and an audit trail that survives a post-incident review.

A large share of network incidents have a known cause and a known fix. Someone wakes up, reads a familiar alert, runs a familiar procedure, and goes back to bed. That work is real, repetitive and almost entirely mechanical.

Auto-remediation moves that class of work to the agent. Not the novel faults, not the ones requiring judgement about business risk — the ones with a runbook that already exists.

The reason teams hesitate is control, and they are right to. So control is the design centre here: what the agent may touch, under what conditions, with whose approval, and what it records while doing it.

Deliverables

What we build

01

Runbook codification

Turning the procedures your engineers already follow into tested, versioned actions. The knowledge stops living only in the heads of three people.

02

Autonomy tiers

An explicit ladder: observe only, propose, approve-then-act, act-and-report. Each action starts low and moves up when the evidence supports it.

03

Blast-radius policy

Hard limits on scope — how many devices, which roles, which sites, what time windows — enforced before an action runs, not reviewed afterwards.

04

Approval workflow

Approval where your team already works, with the evidence, the proposed change and the rollback plan in one place.

05

Verification and rollback

Every action is followed by a check that it worked. If the signal does not recover, the agent rolls back and escalates rather than trying again.

06

Audit trail

A complete record of what the agent saw, concluded, proposed, did and verified — exportable for compliance and readable in a post-incident review.

Agent

What the agent does with this layer

This is the agent acting rather than advising. It is also the capability we deliberately enable last.

Recognise

Matching against known conditions

Remediation only fires when the diagnosis matches a codified runbook with sufficient confidence. Novel faults go to a human, clearly labelled as novel.

Act within limits

Executing inside policy

The action runs through the same pipeline your engineers use, subject to the same validation, capped by blast-radius policy.

Verify and report

Confirming the fix landed

Post-action verification against the original signal. Recovery is reported; failure triggers rollback and escalation.

Stack

Tools we use here

LangGraphorchestrationMCPtyped tool accessAnsible / NornirexecutionPolicy engineguardrailsAlertmanagertrigger routingOpenTelemetryagent tracesGitversioned runbooks
Questions

Questions about auto-remediation

Will the agent change our network without asking?

Not unless you explicitly configure it to, for a specific action, after that action has proven itself in propose-only mode. Approval-gated is the default and most clients keep the majority of actions there permanently.

What happens when the agent is wrong?

Three safeguards apply. Blast-radius policy caps the damage any single action can do. Post-action verification detects that the signal did not recover. Automatic rollback returns the previous state and escalates to a human with the full trace attached.

Where should we start with auto-remediation?

With high-frequency, low-risk, well-understood conditions — interface resets, clearing a stuck process, restoring a known-good configuration block. These build the operational trust that later justifies wider scope.

How is this different from the runbook automation we already have?

Traditional runbook automation fires on a trigger and follows a fixed script. The agent diagnoses first, across correlated signals, and only then selects a runbook — so it does not run the interface-reset procedure on a fault that is actually an upstream MTU mismatch.

Book a 30-minute automation readiness consultation

In 30 minutes, we’ll evaluate your infrastructure maturity, identify operational risk areas, and highlight high-impact automation opportunities.

No scripts. No invasive discovery. Just clarity.