Migrating 180 services to a multi-account AWS landing zone without taking down a customer-facing system is, as they say, harder than it looks.
Over the past year, our team partnered with Atlas Federal Credit Union on exactly this migration. The bank had grown organically across three cloud providers, with engineering teams operating independently and infrastructure that ranged from well-architected to genuinely scary. We came in to design a unified landing zone, migrate every workload, and leave behind a platform the bank's engineers would actually want to use.
Here is what we learned along the way.
Start with the platform team, not the platform.
The biggest risk on a migration of this scale is not technical — it is organisational. You cannot land a new platform if the people who operate it do not exist yet. We spent the first six weeks embedded with Atlas's existing infrastructure team, mapping their pain, and co-designing the operating model for the new platform team.
By the time we wrote a single line of Terraform, the bank had committed to hiring a dedicated platform team of seven engineers. Without that commitment, the migration would have been an expensive exercise in futility.
Paved roads beat mandates.
We could have mandated that every team use our new platform from day one. We chose not to. Instead, we invested heavily in the experience of the first five teams that adopted it — measuring their time-to-first-deploy, listening to feedback, and iterating on the paved road.
By the time the second wave of teams joined, word had spread. The platform became the path of least resistance rather than a top-down mandate.
The boring infrastructure is the leverage.
Most of the value in a platform like this comes from unsexy work: golden CI/CD pipelines, standardised observability, cost attribution, secrets management. We spent more time on these than on anything flashy.
Two-year defect warranty kept us honest.
Aightify's two-year defect warranty is a contract we wrote into the engagement. If we shipped something that broke in production, we fixed it — same team, no escalation. That made us paranoid about quality in a healthy way, and it gave Atlas confidence to let us touch their most critical systems.
The result.
One year in, Atlas runs 100% of production workloads on the new platform. CI/CD pipeline duration dropped from 90 minutes to 7 minutes. Cloud spend is visible per product line. And the bank's platform team has the muscle to evolve the platform without us in the room.
That last point is the only one that matters.