Finland
The client specializes in smart hunting equipment, with the current iteration of their platform providing both an e-commerce website and the API all the devices connect to. The backend consisted of several dozen microservices all running in a Kubernetes cluster, and the underlying infrastructure was hosted in a private cloud of a regional provider.
Due to the nature of the business, the traffic load is seasonal, and growing with every season – to the point the existing infrastructure could not cope with the increase. The client had to prepare the infrastructure prior to starting a new season by ordering new nodes to increase the resource capacity. Manual scaling was the only option available, and it got quite inefficient cost-wise when the new instances were not in use - which is more than half a day every day during the hunting season. With the growing popularity of the client’s products it became increasingly difficult to predict the load every new season, which led to several outages and ultimately affected end user experience. Client’s engineers had to pull off long shifts during the worst of it just to keep the system stable.
Ensuring scalability became the ultimate goal: as soon as the infrastructure is able to keep up with increasing traffic and shrink back when it’s over, the users are happy and devices will work reliably. The usual tools and methods were already there, since the backend was running in a Kubernetes cluster, so it was only a matter of migrating it to a global provider with a managed Kubernetes to support automatic scaling. This seemed trivial at first, but we had to take into account the efficiency requirement: the instances get spawn right when they are needed, of the right size, and are deleted as soon as the load is gone.
After the initial audit, we came up with a migration plan: two phases to run in parallel, ensuring we can finish within the half a year period between the seasons. During the first phase we wrote and ran load testing that simulated the amount of traffic the system had experienced during the last outage with the goal of reproducing it to see the impact on both the backend and the database. It had been noted by the client that the database hadn’t been stable during some of the outages so we wanted to make sure that it’s tuned as well. Once we created the new infrastructure and deployed the application stack there (which is the goal of the second phase), the same load testing scenarios were used to test the scalability and stability of the new setup.
The infrastructure in many regards became the replica of the existing setup, with some adjustments made from the best practices standpoint. The new stack was deployed to Google Cloud Platform, where the network was peered with MongoDB Atlas to reduce the latency for DB connections as much as possible. All the major pieces of the infrastructure (monitoring, e-commerce, backend) got their own separate per-environment projects, now immutable, defined as Infrastructure as Code with Terraform. In the end, the client got two identical setups with different machine sizes: changes would first roll out to the testing environment, and would reach the production only after the explicit approval from the engineering leads. Whatever happens in either can be reverted to its previous state with Terraform, granting the setup reliable disaster recovery.
The load tests turned out to be their own kind of isolated project as they had to simulate two different types of concurrent connections at the same time with very different behavior patterns: one that resembled IoT devices posting data, and the other that acted the way the humans would, retrieving the same data at semi-random intervals. We made sure to parameterize the load in the scenarios, and it paid off when the infrastructure was ready: it was easy to multiply the worst case load, and we could simulate even the most ambitious traffic expectations before the new season.
At the application level, we unified the CI/CD pipelines for all microservices making the deployments faster and more predictable, with rollbacks and a strict versioning pattern.
The effectiveness of the new setup was proven the following season when the traffic broke a new record of 200,000 concurrent requests per minute. Corewide continues providing on-demand support for the infrastructure.
Some numbers for reference:
provisioning a new environment from scratch takes 35 minutes in a single action
infrastructure updates take 7 minutes on average
full CI/CD pipeline duration: 3 minutes on average per microservice