How to test that your backups actually restore

·

How to test that your backups actually restore

The question that really matters about your backups isn’t “are they running?”. It’s “when was the last time you actually restored one?”. In audits, the first gets answered with a calm yes and the second with an awkward silence. That silence is the hole: a backup that runs every night without throwing an error looks like a backup that works, and it isn’t always. You only know once you restore it.

This article is about the part almost nobody does: testing backup restores. Setting up the drill, timing it and finding what breaks in cold blood, not on the day of the fire. Because the backup that has never been restored isn’t a backup: it’s an unverified promise.

The backup that has never been restored isn’t a backup

A backup job that finishes green tells you one thing only: that some data was read and written somewhere else. It doesn’t say the data is complete, that the file isn’t corrupt, that the database is consistent or that you can boot a system from it. “Green” measures the write, not your ability to recover. Confusing the two is the most expensive false sense of security in all of IT.

The day you need to restore is never a quiet Tuesday. It’s three in the morning, with ransomware already inside or the server dead and management asking when invoicing comes back. The worst possible moment to discover the database backup was only half done, that a folder was missing or that restoring 2 TB over your connection takes three days. The restore test exists so those discoveries happen today, cold and at no cost, and not then.

Nobody has a backup problem. Everybody has a restore problem. Most just don’t know it yet.

What a restore test is and how often to run it

A restore test is recovering for pretend so you know that on the real day you actually can: you take a backup, spin it up in a controlled environment and verify that the data is there, is consistent and the system works. It’s not reading the report or checking the file is the size you expected; it’s touching the recovered data and confirming it’s usable. And not all tests are equal: each level tests something different and has its own sensible frequency.

  • Restoring a file or folder (monthly). The cheapest drill and the most common in real life: someone deleted something. Recover a file from a week ago and one from a month ago and check that it opens. It validates retention, not just last night’s copy.
  • Restoring a mailbox or email account (quarterly). Test recovering a full mailbox or specific messages, especially with Microsoft 365 or Google Workspace, where many people wrongly believe the provider is already backing them up.
  • Restoring a full server or application (quarterly or half-yearly). Spin up the file server, the ERP or the database in isolation and verify that it boots and is coherent. This is where forgotten dependencies surface: services that won’t start, licences, connections to other machines.
  • Bare-metal or whole-environment recovery (yearly). Rebuild from scratch, as if nothing were left. The real disaster drill: this is where your RTO truly gets measured and the times nobody had ever clocked come out.

The practical rule: the more critical and harder to rebuild, the more often you test. And whenever something relevant changes—a new server, a migration—you repeat the drill. A backup that worked in January may have been broken since the March migration without anyone noticing.

How to set up a serious restore test

A badly done test gives you peace of mind that’s even worse than not testing. “I opened a PDF from the backup and it looked fine” proves nothing. A serious drill is run in an isolated environment, follows a script and produces a measured number, not a feeling. This is our checklist:

  • Isolated environment, never production. You restore on a separate network or machine, without touching live systems. Restoring over production “to see if it works” is the best way to turn a drill into a real incident.
  • Start from the scenario, not the file. Define which disaster you’re simulating: accidental deletion, lost server, ransomware that encrypted production and the reachable backups. Restore the way you would on that real day.
  • Verify the data, not that it “exists”. That the database opens and is consistent, that the figures add up, that the backup date is the one expected. A restored file that won’t open is a lost file with extra steps.
  • Time the RTO for real. Measure from “we decide to restore” to “the service is usable”, including what nobody counts: downloading the backup from offsite, decrypting, reinstalling, reconfiguring. That clock is your real RTO, not the one on the brochure.
  • Confirm the real RPO. Check how many hours of work you lose with the backup you used. If you copy once a day, the worst case is 24 hours. Make sure management knows that number and accepts it, in writing.
  • Document and compare against what was agreed. Record the result and check it against the RTO/RPO the business said it could withstand. If the plan says eight hours and the test says twenty-six, you have a problem to fix today, calmly.

That last point is what makes the exercise useful. Without an RTO and RPO agreed with management, the test has nothing to measure against and stays a technical anecdote. With them, every drill tells you whether you’re inside or outside the margin your business can bear.

Typical failures you only see when restoring

There are failures no backup report catches, because they don’t happen when copying but when recovering. They stay invisible until the day you restore; that’s why the drill flushes them out before they do damage.

  • Incomplete backups. The copy runs, but doesn’t include everything: a database is missing, a shared folder or the server that was added eight months ago and nobody put in the job. Green forever; you recover and exactly what mattered is gone.
  • Forgotten dependencies. You restore the application and it won’t start because it’s missing a service, a version, a certificate or the neighbouring machine it talked to. Recovering a system is almost never recovering a single server.
  • Encryption and lost keys. The backup is encrypted—good—but the key was on the same server that was lost, or nobody knows where it lives. A backup you can’t decrypt is perfectly useless noise.
  • Inconsistent data. The copy was taken with the database running and without a consistent dump: the file is there, but corrupt or mid-transaction. You only find out when you try to mount it.
  • An unbearable real RTO. Everything’s there and everything’s correct, but restoring it over your internet line takes four days and the business can bear one. You had the backup; you didn’t have the time.

None of these gets fixed by buying more disks. They get fixed by testing: only the drill makes them visible while you can still correct them without rush or losses.

How MagicBoxDesk does it

Testing restores by hand, with judgement and on a regular basis, is exactly the kind of task a company busy with its own work never gets round to: important, not urgent, until the day it’s extremely urgent. That’s why at MagicBoxDesk we run it for you. Our managed backup includes restores tested on a regular basis: we don’t hand you software and wish you luck, we give you a service that answers for the result. We spin your backups up in an isolated environment according to their criticality, verify the data is consistent, time the real RTO and hand you a report with what works and what needs adjusting before it becomes a problem.

All of this fits with our 24/7 monitoring, which watches every backup job and alerts if something fails—you hear it from us, not from a disaster—and with the rest of the managed IT services that make up your outsourced IT. The difference is concrete: you go from believing you’re protected to having the proof, with a date and measured times.

Stop trusting that your backup restores and check it. Request a no-obligation quote and we’ll set up your company’s restore tests so that on the bad day you recover for real, not on faith.


Has this raised a question about your own infrastructure?

Book 30 minutes with a MagicBoxDesk engineer. No strings attached.

Book a call