Client context
A German intermediary for state energy-efficiency funding. Property owners hand over their grant and loan applications; the client navigates the process on their behalf — funding with hard government deadlines, strict documentation rules, and no tolerance for administrative failure.
The technology stack is sophisticated and well-built. The team is lean, around 24 people, and optimises relentlessly. The ask that started the engagement was modest: review the environment and recommend improvements to the service offering.
Business problem
The diagnostic found a structural constraint underneath. This lean operator was carrying the risk profile of a business twenty times its size — not because of what it built, but because of who its failures land on.
A technology failure here does not stay internal. It can cost a property owner their grant or loan — funding they may have already committed to a contractor. Third-party liability, not inconvenience. That single fact moves disaster recovery from optional to essential and overdue.
The stakes
- €508K–€1.05M in modelled expected annual loss from the DR/BCP gap, probability-weighted across three failure scenarios.
- One central database as the single source of truth for eight systems — no visible failover.
- Three regulatory frames triggered at once: GDPR Article 32, the NIS2 Directive, and an implicit fiduciary duty to the end clients whose funding depends on the platform.
What was broken
- One central database as SPOF for eight interconnected systems — no redundancy visible
- 90-day CRM recycle bin is not a backup; AI-agent corruption not natively recoverable
- Silent document-loss path in the Power Automate correction flow
- One named engineer owned the majority of open recovery items across all critical systems
- No tested restore for the central system; recovery time likely measured in days
Diagnostic approach
We mapped every system and its specific recovery exposure, then priced it — grounded in documented incident patterns, not guesswork.
Strategy
What is the exposure worth?
The constraint was priced, not just named. Reliability had quietly become the product.
- €508K–€1.05M modelled expected annual loss
- Reliability reframed as a commercial differentiator, not a compliance cost
- The clearest ROI sits in avoided loss, not saved time
Operations
Where does the risk actually live?
Eight systems mapped; recovery requires all of them restored, in the right order.
- One central database as the single point of failure for the whole estate
- Silent-loss paths in the document pipeline
- One-person recovery dependency across all critical systems
Technology
What makes it resilient without disruption?
The architecture was well-designed. The gaps were in recovery, not delivery — so nothing that worked needed to be torn out.
- Additions to the recovery layer, not replacements to the delivery layer
- Independent backups outside every vendor tenant
- Tested restoration, not assumed restoration
What changed
The risk was priced before a solution was proposed
An actuarial expected-loss model — probability × cost across three failure scenarios — put €508K–€1.05M on paper. The board could now make a decision with numbers, not with fear.
Six practical fixes, none replacing what works
Nightly encrypted backup to an independent German VPS, third-party CRM backup to independent storage, dead-letter monitoring on every Power Automate flow, M365 backup outside the tenant, Zapier configs version-controlled in Gitea, and a Developer Runbook any competent engineer can follow.
A runbook replaced the hero dependency
Step-by-step recovery for all five critical systems, pair programming on critical flows. Recovery no longer depends on one person being reachable at 2am on a Friday.
Decisions
- Price the risk, don't assert it
- An actuarial expected-loss model calibrated to the real architecture. Numbers, not fear.
- Additions, not replacements
- Fix the recovery layer; leave the working delivery layer intact. The regulatory workflow stays; the exposure does not.
- Independence over vendor trust
- Backups outside every vendor tenant. A recycle bin is an undo feature, not recovery.
- Runbook over hero
- Recovery depends on a document any competent developer can follow — not on one person being reachable.
Results
| Measure | Before | After |
|---|---|---|
| Central-system recovery time | No tested restore; likely days | < 4 hrs (designed) |
| Annual risk exposure | €508K–€1.05M unpriced and uncovered | Priced, phased, and being designed out |
| CRM / SaaS data restoration | Not possible beyond the recycle bin | Any point in 30 days |
| Key-person recovery dependency | One engineer required | Any developer can restore any system |
| GDPR Art. 32 / NIS2 continuity | Not demonstrable across 8 systems | Demonstrable across all 8 (designed) |