Confidential Government FinTech Platform, Saudi Arabia

A stalled DR initiative became a proven recovery capability in under two months.

47% cloud cost reduction. Under 1 hour RTO. Near-zero RPO for critical data. A live production switchover between two Oracle Cloud regions.

GovTech, FinTech, Digital Justice2026

Published with the client's permission.

A national FinTech platform under a confidential engagement

The challenge

The platform sits at the intersection of finance and law. It manages debt instruments and multi-party financial obligations between businesses and individuals. Infrastructure availability, data consistency, and controlled recovery are not technical preferences. They are legal requirements.

The client knew they needed disaster recovery. But implementation had stalled for months. The platform was a complex web of interconnected, stateful services: Oracle Kubernetes Engine, MySQL HeatWave with multiple replicas, self-managed MongoDB, RabbitMQ, Redis, Python/Django, Node.js, and Celery workers. Each with different replication and recovery behaviors.

In a financial system, an improperly executed failover is more dangerous than a temporary outage. Duplicate transactions. Data loss. Inconsistent state between regions. Background workers executing the same scheduled operations twice. Partial recovery where dependent components become active in the wrong sequence.

The existing team understood the production platform well. But the DR initiative lacked a structured recovery process. Workload classification. Dependency mapping. Recovery sequencing. Switchover runbooks. Rollback conditions. None of it existed in a unified framework.

The initiative had remained under investigation for months while the team operated under increasing delivery pressure. The absence of a disciplined recovery methodology made implementation difficult to validate and too risky to execute directly in production.

Our solution

We changed the approach. We stopped trying to "implement DR" and started engineering and proving a complete recovery process.

Before touching production, we conducted a full technical and operational assessment: mapping infrastructure and application dependencies, classifying workloads by business criticality, identifying stateful and stateless components, defining recovery priorities, and establishing RTO and RPO objectives.

AI tools accelerated the assessment. They helped analyze dependencies, cross-check relationships between services, review software versions and backward-compatibility risks, identify gaps, and generate remediation paths. AI was not the decision-maker. It was the accelerator. Final architectural and operational decisions remained under engineering control.

Instead of attempting a full production implementation immediately, each critical layer was validated independently. Proofs of concept for MySQL, MongoDB, Kubernetes workloads, file storage, and cross-region application recovery. Only after each recovery mechanism was understood and tested was it incorporated into the complete DR plan.

We introduced a runbook-first discipline. The recovery process was documented before it was automated. For every critical scenario we defined preconditions, service shutdown sequences, replication checks, database promotion steps, application activation sequences, Celery and background-worker controls, infrastructure validation, business transaction verification, and rollback conditions. AI-assisted review cross-checked the runbooks for missing steps, unclear dependencies, and inconsistent sequencing.

We designed and implemented a Hot-Warm disaster recovery architecture across two Oracle Cloud regions. Oracle Full Stack DR orchestration. Cross-region OKE recovery. MySQL replication and promotion strategy. MongoDB replica-set recovery with custom orchestration. Managed NFS recovery for Kubernetes workloads. Custom automation for services not natively supported. Only the intended region could actively execute business transactions after switchover.

Special attention was given to application state and background processing. Only the intended region could actively execute business transactions after switchover.

The complete design was validated through a planned live switchover. Not documentation. Not theoretical readiness. A demonstration that the production platform could transition between regions while preserving application integrity and business operations.

Alongside DR, we launched a FinOps initiative using Vantage as the visibility layer. AI-assisted cost analysis accelerated the path from raw cloud billing data to actionable decisions. The full initiative, from structured assessment to verified DR capability, was completed in less than two months.

Results

MetricResult
DR initiativeCompleted in less than 2 months, after being stalled for months
RTO achievedUnder 1 hour
RPO achievedNear-zero for critical data tiers
Cloud cost reductionFrom ~$31.5K to ~$16.8K per month (47% reduction)
Annualized savings~$177K
ROI on FinOps tooling47x return
Production switchoverSuccessfully completed live

Business impact

This engagement addressed two risks that directly affected the platform's long-term sustainability: business continuity and uncontrolled infrastructure cost.

The key outcome was not simply the deployment of a secondary cloud region. Our team introduced the planning discipline, recovery processes, technical architecture, AI-assisted analysis, and validation methodology required to turn a stalled DR initiative into a proven operational capability.

By combining deep cloud and SRE expertise with structured process engineering and AI-assisted assessment, we resolved in less than two months a challenge that had remained under investigation for several months.

The FinOps initiative demonstrated that stronger resilience did not have to mean uncontrolled infrastructure expenditure. AI-assisted cost analysis enabled faster and more precise decisions by reducing the manual effort required to collect, correlate, and interpret cloud cost data.

The result: a financial platform with a significantly stronger continuity posture, predictable recovery procedures, substantially lower cloud costs, and a more mature engineering process for future resilience and optimization work.

Services used

Disaster recovery and backup solutionsCloud cost optimization (FinOps)
Next story: Hackers Academy

Ready to survive the bad day?

We'll assess your current DR posture and your last cloud bill, and show you where you're exposed, before the outage finds it.