Menu

We solved the error that occurred during the big promotion on April 29th using the Kubernetes gateway, and we also saved 30% on server costs.

Last Black Friday, our team almost collapsed:13% of user address resolution requests were directly rejected with a 429 error.After checking in the backend for a long time, it was discovered that the throttling rules previously assigned to each microservice were too strict. As a result, the traffic to the logistics query module has increased by three times, and the other available resources cannot be utilized at all.

Previously, we always thought that the Kubernetes gateway was a complex component used only by large companies. It wasn't until we encountered some issues that we were forced to implement it ourselves. The whole process took less than a week, and the results were much better than we expected.

What exactly is a Kubernetes gateway?

In simple terms, it serves as the unified traffic entry point for your entire Kubernetes cluster. All external requests first reach this point before being forwarded to the corresponding backend services. The open-source version we use can handle up to 12,000 requests per second with a single instance, so small teams don’t need to worry about performance bottlenecks at all.

The 3 real benefits we've obtained

A yellow traffic light showing red against a clear blue sky background.

The first and most obvious benefit is that we no longer need to configure throttling rules for each service individually. Previously, during major promotions, we had to adjust the throttling thresholds for seven or eight microservices. If even one of them was not correctly configured, it would cause issues. Now, we implement dynamic throttling at the gateway level, allowing any available server resources to be automatically allocated to the modules with higher traffic. As a result, the number of 429 errors during this Christmas promotion has been reduced to zero.

The second benefit is cost savings. Previously, to handle peak loads, we allocated two additional redundant servers for each of the three core services, resulting in a resource utilization rate of only 20% most of the time. Now that the gateway can perform automatic load balancing, we were able to reduce the number of servers by six. We calculated that the monthly cloud costs have decreased by 30%.

The third benefit is that we don't need to have the backend repeatedly modify common logic related to cross-domain interactions and authentication. Previously, whenever a new service was launched, the backend had to write the JWT validation code from scratch. Now, by simply configuring a rule at the gateway level, we can solve this issue. As a result, the three of us in the backend team can save at least half a day each week to work on developing business-related code.

Don't just listen to the benefits; we've definitely stumbled into these two pitfalls.

The first issue is that the default timeout setting is extremely problematic. On the very first day of the service going live, we noticed that 5% of the address resolution requests timed out. It took us a long time to discover that the default gateway timeout was set to 30 seconds, while our address resolution service itself actually takes up to 40 seconds to complete its tasks. Once we adjusted the timeout settings, the issue was immediately resolved. Before going live, it is essential to carefully check the timeout thresholds for each of your business interfaces.

The second mistake is not to overwhelm the system with all features from the start. Initially, we wanted to enable logging, throttling, WAF (Web Application Firewall), and grayscale release at the same time, but due to a configuration error, the entire cluster was down for 20 minutes. Later, we started by only enabling routing and throttling, and after running it for 3 days without any issues, we gradually added more features. Since then, we haven’t encountered any problems again.

Should we use it or not? The criteria we have for making this decision are very simple.

Use cases:

Red traffic light set against a bright, cloudy sky. Perfect for themes of travel, technology, and safety.

  • You have more than 5 microservices, and every time you need to change the general configuration, you have to make adjustments in several places.
  • Frequent encounters with traffic peaks, making it difficult to flexibly allocate resources between different services.
  • The backend team consists of less than 5 people, and we don't want to waste time on repeatedly writing common logic.

Situations where you shouldn't mess around:

  • You only have two microservices, and the traffic has been stable all year round. There are no issues with using NGINX for forwarding at the moment.
  • No one in the entire team understands the basics of Kubernetes; don't force the adoption of new technologies just to keep up with the trend.

2 practical suggestions for small teams starting out for the first time

No need to worry about choosing a solution; for small teams, using open-source APIs like APISIX or Kong is perfectly fine. There’s a wealth of supporting documentation available, and if you encounter any issues, you can usually find solutions with a quick search. There’s no need to buy the commercial versions; the free versions have all the necessary features.

Before going live, divert 10% of the traffic for 24 hours to test for any timeouts or misinterceptions of requests. Only switch over the entire traffic once you confirm there are no issues. Don’t transfer all the traffic at once.

Finally, let's address a common question: No one in our team had previously specialized in gateway development for the three backends. We spent two days working with the official documentation and got it set up. It’s really not as complicated as you might think.

Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR