← Back to CV
Case Study — Resilience & Datacentre Architecture

Redesigning Disaster Recovery for a Regulated Trading Platform

Interactive Investor · Infrastructure Architect · 2020 – present

Context

Interactive Investor's platform handles live trading and client funds, so downtime and data loss carry direct regulatory and financial consequences. The existing disaster recovery posture relied on a traditional active/passive datacentre pairing: asynchronous storage replication and a failover process that required manual network reconfiguration to bring services up in the DR site. Recovery was achievable, but slow and operationally risky — exactly the kind of gap that shows up under real incident pressure rather than in a tabletop exercise.

I first worked this environment as Senior Infrastructure Services Engineer, acting as technical lead for formulating the datacentre migration strategy itself — getting the hypervisor, storage and network estate onto a platform capable of supporting a genuinely resilient DR design, rather than retrofitting one onto legacy hardware. Once that migration was delivered, I moved into the Infrastructure Architect role and took ownership of redesigning the DR architecture properly, with a mandate to close the gap between "we can recover" and "we barely notice."

Technologies: VMware NSX · VMware vSphere/SRM/Aria · Pure Storage (synchronous replication) · ServiceNow CSDM · Cisco/HP switching

Approach

The design centred on removing the two things that made failover slow: storage replication lag and network re-addressing. It also had to be deliverable without an outage window a trading platform can't be given.

Delivery sequencing

Redesigning live DR infrastructure under a regulated trading platform means the migration itself carries as much risk as the end state is supposed to remove. The rollout was staged accordingly:

Outcome

Synchronous
storage replication, primary → DR
Zero re-IP
network config unchanged on failover
Metro cluster
VM mobility across live datacentres

The result is a DR position that lowers both RTO and RPO materially against the previous active/passive design, while giving the infrastructure team an SDN foundation (NSX) that is now being extended into micro-segmentation and, separately, into the identity and access governance work described in the IAM/OIG case study.

As part of the same architecture remit, I've since led mapping of the platform's business and technical services into ServiceNow CSDM — giving the organisation a common model that ties infrastructure components to the business services they actually support, which is now the reference point for impact analysis during incidents and change.