Menu

We maximized the accuracy of logistics address matching using OpenAI o1, but encountered three unexpected issues.

Just after the peak of last Black Friday, our three-person backend team in Europe finally breathed a sigh of relief: last year, during the major promotional event, 13% of address resolution requests encountered a 429 error due to the throttling limits of a single model. This year, not only did we not experience a single instance of throttling, but the error rate for address matching also decreased by 82%. The key change was that we replaced our primary inference model with the o1 model.

For those who haven’t come across it before: What exactly is o1?

It's the large model for reinforcement learning launched by OpenAI. Different from the previous GPT-4o, it takes more time to "think" about complex logical problems. The official basic reasoning ability test score is 87% higher than that of GPT-4o, and this was the main reason we chose it in the first place.

For small and medium-sized teams like ours, the benefits of o1 are truly substantial.

The first benefit directly solves the most troublesome issue for us, which is address matching. When using GPT-4o before, we often encountered confusion between streets with the same name in different European countries and between postal codes with the same name. This was especially problematic when users entered addresses with spelling errors, requiring us to implement three layers of rule validation to achieve an accuracy rate of only around 92%. Now, with o1, the accuracy rate has been increased to 98.7%, and we deleted all that rule code last week.

The second point is that there's no need to create complex prompt engineering anymore. Previously, in order to make the model produce output in a format that matched our backend requirements, we had to write nearly 2000 words of prompts, and the output format often deviated from what we wanted. Now, just one line of instruction is enough to achieve the same result, which has directly reduced the maintenance cost of the prompts to zero.

The third point is that the stability of concurrent requests is better than we expected. We were initially concerned that the long reasoning time might lead to delays in processing, but the monitoring data shows that most requests are completed within 15 milliseconds. Only a very small number of particularly complex cross-country address queries take more than 200 milliseconds, which is still completely within our acceptable range.

But its pitfalls can really catch you off guard.

Scrabble tiles spelling "CHATGPT" on wooden surface, emphasizing AI language models.

The first issue is that the token consumption is much higher than that of GPT-4o. In the first three days after we switched over, the inference cost increased by a factor of 2.3. It was only later that we realized that o1 automatically includes its thought process in the token consumption as well. If you don’t need to see its reasoning logic, you must disable the output of the thought process in the call parameters. After we did that, the cost dropped back to the previous level of 1.2 times, which is completely acceptable.

The second issue is that it's not suitable for simple, high-frequency requests. At first, to save effort, we directed all address query requests to o1. Later, we found that for addresses with no spelling mistakes and in standard format, there was no difference in accuracy between using o1 and GPT-4o; in fact, it even cost more money. Now we have added a simple preliminary validation, and only requests that do not meet the standard format are directed to o1, which has reduced costs by another 30%.

The third issue is that the support for non-Latin languages such as Chinese and Arabic is not yet stable enough. We have a small number of users from the Middle East, and occasionally, address entries entered in Arabic cause parsing errors. After reporting this issue to OpenAI, we are still waiting for a fix. For now, we are using local rules for additional verification.

Who should use o1? Who shouldn’t touch it right now?

If you're dealing with tasks that require logical reasoning—such as complex rule matching, multi-condition validation, or simple code generation—and you've already hit a bottleneck in accuracy when using GPT-4o, then switching to GPT-4o1 would be a very worthwhile decision, as the cost-benefit ratio is quite high.

But if you're doing simple text classification, content generation, or frequent standardized requests, there's absolutely no need to use o1; it's a complete waste of money. GPT-4o or even GPT-3.5 can handle those tasks just fine.

2 tips for beginners

  • Let's start by testing the 1000 pieces of historical data that you previously processed using the old model. Don't make the switch all at once; you can quickly see whether the improvement in accuracy compensates for the increase in costs. We tested it for two days and then decided to make the change.
  • When making a call, the o1-mini version is preferred. For 90% of small and medium-sized teams, the capabilities of o1-mini are more than sufficient, and the price is also 60% lower than that of the full version of o1.

Let's conclude with the most frequently asked question: Will o1 be quickly replaced by other models? At least in the context of address resolution, which we are working on, we have tested all the major inference models on the market, and none of them have surpassed o1 in accuracy. For the next six months at least, it will still be our top choice.

Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR