After our team switched address recognition to GPT-5, the error rate associated with code 429 decreased from 13% to 0.2%, but we encountered two unexpected issues.
Last month, the Black Friday promotion just ended, and our three backend teams spent a full 72 hours monitoring the systems. Last year, during the promotion, 13% of the address resolution requests encountered the 429 rate-limiting error, which was a serious issue. This year, that problem didn’t happen again. The only major change we made was to replace the previously used general-purpose large model with GPT-5 for the structured processing of addresses.
For those who haven't tried it yet, let me explain clearly: What exactly is GPT-5 for scenarios like ours?
It's the multi-modal large model updated by OpenAI this year, and the only parameter anchor that is most useful to us is:The maximum number of structured processing requests supported per minute per account is 12 times that of the previous generation of the model.And the recognition accuracy for non-standard inputs has increased by an order of magnitude.
We handle cross-border logistics in Europe, processing 500,000 addresses entered by users every day. Many people mix up streets, postal codes, and cities, and sometimes even add emojis or spell half of the information incorrectly. The previous models either made mistakes in recognition or required three API calls to produce a result. During peak promotional periods, our system would immediately encounter rate limits.
After switching to GPT-5, the three actual benefits we obtained are:
The first and most straightforward solution: the rate-limiting issue is directly addressed.During the peak promotional period, the error rate for code 429 decreased to 0.2%.We don't even need to use a backup model queue, which saves 30% of the operational maintenance time that was previously spent on load balancing.
The second improvement is the increased accuracy of address recognition. Previously, about 8% of addresses required manual re-verification, but now that figure has dropped to less than 1%. As a result, the customer service team handles 2,000 fewer address correction tickets per week.
The third one, on the other hand, doesn't get much mention: it supports direct uploading of photos of handwritten receipts for recognition, eliminating the need for us to use a separate OCR service to convert the images into text. This saves two steps in the process, and most requests receive results within 15 milliseconds. Only a very few blurry receipts may take a bit longer to process.
Don't just look at the benefits; the two pitfalls we encountered, you're very likely to face as well.

The first issue is the inadequate adaptation of the spelling for dialects in smaller language regions. We thought its multilingual capabilities were strong enough, but last week, the error rate for recognizing Basque language addresses in the Spanish region suddenly increased to 12%. Upon investigation, we found that there were very few samples of addresses in such niche dialects in the training data. As a result, we have had to add local rules as a fallback for these addresses.
The second issue is cost. We previously calculated that the cost of a single request token was only 20% higher than that of the previous generation, but we forgot that the number of structured fields returned by default had increased by 3. As a result, the actual token overhead was 40% higher. Later on, we restricted the number of fields returned in the prompt to only the necessary 5, which brought the cost back to being 10% higher than before.
Which teams should go all in, and which teams don’t need to be involved at all?
- What's needed: There are over 100,000 requests per day for the structured processing of unstandardized text/images. Teams that have previously been constrained by model throttling and insufficient accuracy, such as those in e-commerce, logistics, and customer service ticket processing in small and medium-sized enterprises, will see a very clear return on investment (ROI) after making the switch.
- What shouldn't be used: For teams that only perform simple question-and-answer tasks, content generation, or have less than 10,000 requests per day, there's absolutely no need to spend this money. The previous generation of models is sufficient and can even save a lot of costs.
Two specific suggestions for beginners
First, use your historical real requests from the past 7 days to create a test set. Don't just test the examples provided by the official source; especially for content related to minority languages or less popular regions, make sure to run the tests first to check the accuracy. Otherwise, discovering missed judgments after the system goes live will be very troublesome.
Secondly, when making the first call, directly specify the return fields you need in the prompt; don't let it generate content arbitrarily, as the additional generated content will consume a lot of unnecessary token overhead.
Someone asked: Do you still need to queue up to apply for GPT-5 nowadays?
We have an enterprise developer account. The submission for the necessary qualifications was completed 3 days ago, but for individual developers, it may take 1-2 weeks to get approved. However, in case of an urgent need, there is an enterprise-specific fast-track process that can be used to activate the account on the same day.
Article link:https://airai.cc/en/ai-news/33/
Was this helpful?