Skip to main content
Uptime Isn't RecoveryContinuous Monitoring
5 min readFor Compliance Officers

Uptime Isn't Recovery

The conventional wisdom: Your incident is over when services come back online. You've restored from backup, users can log in again, and the CEO can tell the board that operations have resumed. Recovery complete.

This belief drives nearly every breach response timeline I've reviewed. Organizations measure recovery in hours of downtime, not in evidence of adversary removal. They track service availability, not trust restoration. They celebrate the moment systems return to production, then move the incident into post-mortem review.

That's not recovery. That's resumption. And the difference will cost you.

The Real Recovery Process

Declaring victory when services restart conflates three separate problems that require three different solutions:

Operational recovery gets you back to business. You rebuild servers, restore applications, and re-enable access. This is the visible, urgent work that executives and customers care about most.

Adversary eviction removes the threat actor from your environment. You identify every credential they touched, every persistence mechanism they installed, every identity system they compromised. You validate that their access is gone, not just dormant.

Governance correction fixes the control failure that made the breach possible. You address the unpatched system, the weak segmentation, the unmanaged exception, or the supplier oversight gap that created the opening.

Most organizations complete the first recovery and defer the other two. Not because they don't care, but because they're exhausted, business units are impatient, and external advisers have moved on. The language shifts from "incident response" to "lessons learned," and momentum disappears.

Here's the problem: attackers don't depend on a single point of access. By the time ransomware encrypts your systems or data theft becomes visible, the adversary has likely established multiple ways to return. Dormant accounts, compromised service credentials, cloud tokens, API keys, scheduled tasks, and persistence in identity systems can all survive a rushed restoration.

These mechanisms don't create immediate disruption. They wait until you've relaxed, reopened access, and declared the incident closed. That's when "back to normal" becomes your most fragile state.

Evidence of Comprehensive Recovery

NIST SP 800-53 Rev 5 control IR-4 (Incident Handling) requires organizations to contain incidents and eradicate incident-related artifacts. It doesn't say "restore services and move on." The control family distinguishes between containment, eradication, and recovery as separate phases with separate objectives.

CMMC Level 2 practice IR.2.093 requires testing incident response capability, which includes validating that your recovery process actually removes adversary access. If you're declaring incidents closed based on uptime alone, you're not meeting the practice's intent.

DFARS 252.204-7012 requires you to implement NIST SP 800-171 controls, including IR-4 (Incident Handling) and RA-5 (Vulnerability Monitoring and Scanning). Those controls assume you're verifying that the conditions enabling the breach have been corrected, not just that services are running again.

The governance piece matters even more. Breaches don't occur only because an attacker was capable. They occur because something in your organization made the attack possible or allowed it to progress. That might be a known control gap, weak identity governance, delayed patching, poor segmentation, insufficient logging, or an accepted risk that was never revisited.

If you restore technology without addressing that enabling condition, you haven't recovered. You've resumed operations on the same assumptions that failed.

Steps for True Recovery

Before you declare an incident closed, require clear answers to these questions:

How did the attacker first gain access? Not "phishing" or "credential compromise" as a category, but the specific control gap, misconfiguration, or unmanaged asset that created the opening.

What access did they obtain, and how was it removed? Document every compromised credential, every privileged account they touched, every identity system they accessed. Validate that those paths no longer exist.

What persistence mechanisms were found? Scheduled tasks, service accounts, API keys, remote access configurations, cloud tokens. What evidence shows they're gone?

Which systems were restored from known-good sources? How did you establish that trust? If you restored from backup, how do you know the backup predates the compromise?

Which governance failure has been assigned an owner, a deadline, and executive oversight? Not "we'll improve our patching process," but a specific control gap with a named executive accountable for closure.

These questions don't need to slow recovery unnecessarily. They make recovery honest. In hybrid environments, legacy infrastructure, and cloud platforms where visibility is uneven, certainty is rarely possible. But uncertainty should be documented and governed, not quietly absorbed into the decision to resume operations.

For CMMC assessors reviewing your IR plan under practice IR.2.092 (incident response plan development), this distinction matters. Your plan should define recovery as evidence-based closure, not as service restoration. It should specify what validation occurs before you declare adversary eviction complete.

For FedRAMP continuous monitoring under CA-7, your incident closure criteria should align with your ongoing authorization posture. If you're declaring incidents closed while residual access remains possible, you're creating a gap between your security posture and your authorization boundary.

When Operational Recovery Takes Priority

There are circumstances where operational recovery must proceed before full adversary eviction is complete. Critical infrastructure environments, essential services, and safety-critical systems may need to operate in a degraded or partially trusted state.

In those cases, the conventional approach isn't wrong. It's incomplete. The honest position is that essential capability has resumed, but security recovery remains ongoing. That language matters. It gives executives and boards an accurate view of risk. It prevents operational restoration from being mistaken for security closure.

It also protects security teams from being pressured into declaring confidence they don't yet have.

Organizations that recover well from breaches aren't the ones that restore fastest at any cost. They're the ones that understand the difference between availability, trust, and resilience. Until those three activities are treated as part of the same recovery cycle, you'll continue to declare victory too early.

And you'll only discover the mistake when the same adversary, or another one using the same path, returns.

You Might Also Like