Menu

We saved 38% on traffic processing costs by using a cloud-native gateway, but we almost ruined the fresh food delivery network in Europe during Black Friday.

The review meeting for last Friday’s big sale just concluded, and our three-person European fresh food delivery backend team is still feeling the fallout: on the pre-sale day, 17% of the cold chain reservation requests resulted in errors (with codes like 429), and the number of user complaints tripled. We almost ruined the annual sale that we had been preparing for three months.

The problem stemmed from using a traditional API gateway for two years: We had set a unified throttling threshold for two core interfaces, address resolution and cold chain reservation. On Black Friday, the address resolution requests surged, immediately exhausting the gateway's quota, preventing any subsequent reservation requests from being processed. We always thought that switching to a new gateway was something only large teams would need to consider. It was only after encountering this issue that we realized a cloud-native gateway could be a more cost-effective choice for smaller teams.

What exactly is a cloud-native gateway?

In one sentence: It serves as the traffic entry point for applications running within a K8s cluster, eliminating the need to rent additional servers for deployment.You just need to focus on routing rules and throttling strategies; the cloud provider takes care of all the rest, including resource elasticity and operational upgrades.The version we are using now can handle 20,000 requests per second with a single instance, so there's no need to plan for capacity in advance.

After the translation, the three actual benefits we obtained are:

After the failure on the pre-sale day, it took us 3 days to switch to the cloud-native gateway. On the actual Black Friday event, there were no more issues with throttling errors. Additionally, there were three unexpected benefits:

  • The cost has been directly reduced: before, we rented two 4-core 8G servers to run traditional gateways. In addition, an operation and maintenance server spent 8 hours a month to upgrade the version. The monthly cost is 1200 euros. Now the cloud native gateway pays based on the number of calls and only costs 740 euros per month.Directly saving 38% on traffic processing costs
  • I finally understand how the throttling strategy works: The previous gateway could only set global throttling thresholds for all interfaces. Now, we can establish separate rules for different business scenarios. For example, we can set a limit of 5,000 requests per second for the address resolution interface. Even if that interface gets overwhelmed with requests, it won’t affect the subsequent payment or reservation processes.
  • The HTTPS handshake time has been reduced by half: Previously, our SSL certificates were stored on our own gateway servers, which caused handshake delays of up to 200ms for users in different regions of Europe. Now that the cloud provider has cached the certificates on edge nodes, the handshake delay for most requests has been reduced to less than 80ms.

Don't rush to get in the car; we've already encountered those two problems before.

Stunning view of the Bosphorus Bridge and Istanbul cityscape, showcasing historic architecture.

It's not the case that all cloud-native gateways have only advantages; during the transition process, we encountered two issues that almost forced us to redo the work:

The first issue is the cold start latency. After we completed the stress testing on the first day, we noticed that the response time for the first batch of requests suddenly increased to over 300 milliseconds after no requests were made for 10 consecutive minutes. It was only later that we realized that the cloud provider's elastic instances are started on demand, and it is necessary to set a minimum number of instances reserved for the core interfaces. Otherwise, if traffic suddenly increases during idle times, timeouts are likely to occur.

The second point concerns the limitations of custom plugins. We previously wrote a plugin for request signature verification that ran on a traditional gateway. Only when we made the switch did we realize that the cloud-native gateway we were using only supported plugins in WebAssembly format. It took us two days to modify the code to make it compatible. If you have a lot of custom logic, it's best to first check carefully what plugins the vendor supports.

Should I change it or not? Just look at these two criteria to decide.

The situation you should change is:

Aerial photo capturing Kwai Tsing Container Terminals, showing vibrant shipping activity in Hong Kong.

  • The team has fewer than 5 backend developers, and there is no dedicated operations and maintenance personnel. We don't want to spend time on the maintenance and upgrading of the gateway servers.
  • High traffic fluctuations occur, for example, during promotional events when traffic can be 3 to 10 times higher than usual. We don't want to rent a bunch of servers in advance, which would be wasteful since they would remain idle most of the time.

Situations where you shouldn't make changes:

  • The compliance requirements are particularly strict; all traffic must pass through servers that you have control over and cannot use the public nodes provided by cloud service providers.
  • Your gateway has a lot of custom, specialized logic that is not supported by any of the cloud-native gateways available on the market. The cost of modifying it yourself would be even higher than the cost of maintaining the gateway.

Three little tips for beginners

  • Don't switch the entire traffic volume at once; first direct 10% of the traffic to the cloud-native gateway for a week of monitoring. Only after confirming no issues should you gradually transition the rest of the traffic. We started by switching the address resolution interface first, and only after that was confirmed to be working without problems did we migrate the entire traffic volume over.
  • The core interfaces must have a minimum number of instances set. Don't save that little money; we have now reserved 2 permanent instances for each of the core interfaces responsible for cold chain reservation and payment. Since then, we haven't encountered any issues with cold start timeouts again.
  • You don't need to buy the highest-tier version; for the traffic volume of most small and medium-sized enterprises, the basic version has all the necessary features. The basic version we are using now can handle up to 20,000 QPS without any additional costs.

The final answer to the question that is asked most frequently within the team: Will it become tied to a cloud service provider?

Our judgment is: For a back-end team with fewer than 10 people,The efficiency gains brought by being tied to cloud providers far outweigh the cost of building all the components from scratch.When the time comes to switch cloud providers on a large scale, you will have the necessary energy to handle the migration. Don't worry about it too much right now.

Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR