Menu

Our team reduced the cost of address resolution by 62% using GPT-3.5, and at the same time, we solved the issue of traffic throttling during major promotions.

Last month during the Black Friday sales, our three-person backend team, which specializes in cross-border parcel logistics in the Netherlands, didn't have to stay up all night next to the servers for the first time. Last year at this time, the accuracy rate of our self-developed address correction model was less than 80%, and we had to assign a dedicated person to monitor the error logs 24/7. The return rate of parcels alone due to incorrect addresses was as high as 11%.

For those who haven’t come across it before, let me explain what GPT-3.5 is in simple terms: What exactly is GPT-3.5?

It is a large language model launched by OpenAI in 2022, and its stable version, which is still being updated in 2026, supports a context window of up to 16k. The response time for non-stream-based structured output requests is generally within 200ms, which is fully capable of supporting high-concurrency, lightweight semantic processing requirements online.

We chose it to solve the address resolution problem, and there are three real benefits at the core.

Stunning view of the Bosphorus Bridge and Istanbul cityscape, showcasing historic architecture.

  • The first point is the out-of-the-box accuracy: there's no need to manually label millions of addresses from various countries for training. We only used 200 historical abnormal address examples for few-shot learning, and the parsing accuracy improved significantly as a result.98.7%The error rate has been directly reduced to less than 1%.
  • The second advantage is that the cost is much lower than that of developing the model in-house: Previously, the monthly cost for maintaining our in-house model and three GPU servers was 1,200 euros. Now, by using GPT-3.5 to process an average of 500,000 requests per day, the monthly cost has been reduced to only 450 euros, which represents a direct savings of 62%.
  • The third benefit is that there's no need to worry about scaling out at all: During the Black Friday period, the volume of our requests increased by three times. Our previous self-developed models could only handle up to 1.5 times the normal traffic before experiencing a collapse. This time, we simply submitted a temporary request for an increase in capacity to OpenAI in advance, and not once did we encounter a 429 error throughout the process.

Don't just look at the benefits; in the first three months, we encountered more problems than we gained in profits.

At first, we simply fed the original address fields into the model, and we often encountered issues where the addresses of small villages in European countries were misidentified as American cities. It took us a week to realize that the default prompt did not include a rule that prioritized matching the destination country field associated with the logistics request.

Another time, we accidentally enabled the streaming output, which caused the return time for all parsing results to increase by a factor of three. During the peak period, there were over 20,000 requests that timed out. It took us half a day to realize that a colleague had forgotten to switch back to the correct settings while modifying the parameters.

The most troublesome issue was with the billing: At the beginning, we didn’t implement any input truncation, so all the irrelevant notes that users included when entering their addresses were also fed into the model. As a result, the end-of-month bills were 30% higher than expected. Later, we added a prefix truncation rule that limited the input to 100 characters, and the bills returned to normal in the second month.

Ultimately, GPT-3.5 is not a panacea; I advise you not to blindly join these two types of teams.

Aerial photo capturing Kwai Tsing Container Terminals, showing vibrant shipping activity in Hong Kong.

If your team is dealing with highly sensitive data such as medical records or payment information, don't even consider using the public version of GPT-3.5, no matter how affordable and user-friendly it is; you simply can't afford the risks associated with data compliance violations.

If your requirement is to generate professional documents with tens of thousands of words or to perform complex logical reasoning, you should directly choose GPT-4o or Claude 3. The accuracy of GPT-3.5 in such tasks is significantly lower; the savings you might make on cost are not even enough to cover the expenses associated with using it.

On the other hand, if you are working on lightweight semantic tasks such as user intent classification, short-text correction, or structured information extraction, and do not have extremely complex reasoning requirements, GPT-3.5 will still be the most cost-effective choice until 2026, without a doubt.

3 practical tips for first-time users, all based on the mistakes we've made

  • Run shadow traffic for 7 days before going live: Send requests to both your existing solution and GPT-3.5 simultaneously. Only compare the results; do not divert any traffic. Make the switch only after confirming that both the accuracy and speed meet your expectations. Don’t replace everything all at once.
  • Apply a preprocessing filter to all inputs: Remove any extra spaces, irrelevant notes, and special characters in advance. This can save at least 20% in token usage and also reduce the likelihood of misjudgments by the model.
  • Make sure to have a fallback plan: in case the OpenAI API becomes unavailable, you need a less sophisticated alternative that can still be used. Even if it has a lower accuracy rate, it’s better than having the entire service completely unavailable.

FAQ

It's already 2026, with so many new models available, is there still a need to use GPT-3.5?
It's absolutely necessary for lightweight tasks. We have compared it with open-source models at the same price range, and GPT-3.5 has an accuracy rate that is at least 5 percentage points higher. Moreover, you don't need to manage or deploy it yourself, which actually results in a lower total cost.

Do I need to worry about throttling when making the call?
Normal daily traffic generally doesn't increase by more than 10 times suddenly, so the default quota is usually sufficient. If you need to increase the quota for a major promotional event, you just need to submit a ticket 3 days in advance. We have submitted requests for quota increases three times, and the approval was granted in as fast as 2 hours each time.

Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR