NCSC has written down how to recover, and the hard part is not the plan

· · Security

Most business continuity plans answer the wrong question. They say who to call and where the backups live. They do not say which parts of the business have to work again first, in what order, and what the company does for the six weeks in between.

That gap is what the National Cyber Security Centre's new response and recovery guidance, published on 28 July, is built to close. It is aimed at organisations where losing critical systems would stop the business operating, which covers rather more small firms than the phrase suggests. If your orders, your payroll, or your production scheduling live in one system, you qualify.

It opens somewhere unusual for a government document. People say they felt sick, NCSC writes, or like being punched in the stomach. Anger, despair, and guilt are named as normal first reactions, and organisations are told to recognise the toll on their people from day one. That is worth noting because it is accurate, and because the plan you wrote in a calm room will be executed by people in that state.

The three stages

The first hours and days. Working out what happened, how bad it is, and coordinating a response. NCSC's emphasis here is on swift defensive action, establishing who is actually making decisions, and getting control of communications. It also flags the actions to take immediately because they speed up everything later, which is the same instinct behind protecting your options in the first five minutes. NCSC recommends engaging an assured Cyber Incident Response firm at this point, partly for the technical help and partly, in its own words, for the reassurance of knowing you have the best help available.

Recovery to minimum viable operations. This is the useful bit, and the phrase is worth stealing. Minimum viable operations means the smallest set of business functions you need running to keep delivering to customers, not the full restoration of your technology. NCSC is explicit that the decision about what counts is business-led, not technical, and that temporary workarounds are a legitimate way to get there. You reach the end of this stage when the core of the business is functioning again, possibly held together with manual processes.

The longer rebuild. Getting back to business as usual, or better. That means fixing whatever contributed to the incident, and rebuilding in a way that makes the fundamentals possible: patching, configuration, and access control. NCSC's phrasing is pointed. It talks about designing systems so those things are easier to achieve, which is a quiet acknowledgement that in a lot of organisations they currently are not.

What minimum viable operations means for a smaller firm

If you take one thing from the guidance, make it this exercise. It costs an afternoon and no money.

List what the business does for its customers. Not systems, functions: take orders, make the thing, ship the thing, invoice, pay people. Then, for each one, work out how long you could survive without it and what you would do instead. Not "restore from backup". What you would actually do on day three with no order system: a spreadsheet, a phone, a notebook.

That list is your recovery order. It is also the most useful thing you can hand an incident response firm on day one, and the answer to the question your insurer will ask. Most firms discover, doing it, that two functions matter enormously and the rest can wait a fortnight. That is worth knowing before you are deciding it at 2am.

The bit almost everyone skips

NCSC's analogy is a marathon. Reading about the race, buying the shoes, and writing a training plan are a good start, but it is the running that gets you round the course.

The point is that a documented plan is not a tested plan. Testing failover, rehearsing a shutdown and restart, and actually rebuilding something from a backup all teach you things the document cannot. NCSC says tabletop exercises have their place, but realistic exercises reveal problems that plans alone never will, and build the muscle memory to respond under pressure.

Here is the version of that for a firm without a security team. Once a year, pick your most important system and genuinely restore it from backup onto different hardware, with the person who would really be doing it, timed. Not a test-restore of one file. The whole thing.

Firms that do this find out that the backup covers less than they thought, or the restore takes eleven hours rather than one, or the only person who knows the process left in March. Every one of those is cheap to learn on a quiet Tuesday and expensive to learn during a ransomware incident. This is the same lesson as the museums that ignored the British Library's warning: the report was public, the actions were known, and the practice never happened.

What to do this quarter

  • Write the minimum viable operations list. An afternoon, no budget, most valuable single artefact you will produce.
  • Do one real restore of one important system, timed, this year. Write down what went wrong.
  • Find out now which incident response firm you would call, and whether your insurance policy dictates that choice. Doing this during an incident wastes the hours that matter most.
  • Read the guidance itself. It is written for boards and decision-makers as much as technical teams, which makes it unusually easy to hand upward.

How Steelwise can help

Working out your minimum viable operations, then testing whether the recovery actually works, is exactly the kind of review that pays for itself once. Get in touch.

Further reading

← All filings