Menu

We used o3-mini to handle 1.2 million product description generation requests, saving 62% on inference costs.

Last month, just before Black Friday, our content generation service was on the verge of crashing: the general large model we were using suddenly increased its price by 20%, and at the same time, the rate-limiting threshold was halved. During the peak testing period, 17% of the requests received a 429 error code. The operations team kept asking if the descriptions for the 12,000 discounted products that were supposed to be released that day could be completed on time. With a try-it-out attitude, we diverted 30% of the traffic to the o3-mini model, and not only did we manage to handle the entire promotional period, but the cost was also much lower than expected.

Let's get clear first: What exactly is the o3-mini?

It's the lightweight inference model that OpenAI just launched in Q1 of 2026, focusing on structured content generation and low-latency, simple inference tasks. The maximum context window is 128k, and the cost per token is only 1/8 of that of GPT-4o.

The 3 most tangible benefits we have achieved so far

Red traffic light set against a bright, cloudy sky. Perfect for themes of travel, technology, and safety.

  • During the Black Friday period, there were 1.2 million requests for generating product descriptions.The overall inference cost has decreased by 62% compared to using GPT-4o before.Not a single rate limit was triggered, not even during the peak request period 1 hour before the promotion started; all requests were successfully processed.
  • Most requests return results within 18 milliseconds; previously, there were occasional delays of several hundred milliseconds when using large models. We have removed the wait prompts for loading content on the front end, and as a result, the user conversion rate has increased by 0.8 percentage points.
  • The built-in multi-language generation capability eliminates the need for additional adjustments. The accuracy of Portuguese and Spanish descriptions on the Latin American sites has increased by 21% compared to the previously fine-tuned smaller models, eliminating the need for operators to spend time on secondary proofreading.

Don't rush to make a full switch; we've already encountered those pitfalls.

A yellow traffic light showing red against a clear blue sky background.

Don’t use it for complex product recommendation logic reasoning. Initially, we tried feeding the user browsing history into it to generate personalized recommendation texts, but three out of ten times there were errors in the recommendations. For example, a suggestion for an outdoor tent was displayed alongside a camping lamp. In the end, we decided to integrate that part back into the larger model.

Also, if your business needs to invoke tools, it currently only supports up to 3 parallel tool calls. If the number of calls exceeds this limit, either some calls will be missed or incorrect parameter values will be returned. We have encountered several issues where the price display was inaccurate when our inventory query was executed more than 2 times. We have now modified the system so that if the number of tool calls exceeds 2, it will automatically route the request to a general-purpose large model.

Who is suitable for using it, and who really shouldn't even consider it at all.

If your business scenario involves bulk generation of structured content, simple user intent recognition, or regular FAQ responses, especially for small and medium-sized enterprises that don't have the additional computational resources to fine-tune small models on their own, o3-mini is definitely the cost-effective choice.

But if you need to perform complex code debugging, multi-step logical reasoning, or in-depth analysis of long documents, it's better to stick with the general-purpose large models. In that case, the accuracy of o3-mini will decrease significantly, and you'll spend more time verifying the results.

2 Practical Tips for First-Time Users

Abstract representation of a multimodal model with dots and lines on a white background.

First, implement a 7-day traffic grayscale release; do not make the full switch all at once. Initially, we switched the non-core product tag generation scenarios to the new system for 3 days to ensure that the error rate was within an acceptable range before gradually transitioning to the core product description scenarios. This approach helps to prevent any issues that could affect the online business.

We can build a simple routing layer ourselves to categorize requests based on their complexity. Simple requests are processed by o3-mini, while more complex requests are automatically directed to the larger model. This approach not only saves costs but also ensures that the performance in complex scenarios is not affected. That’s exactly what we’re doing right now: 90% of the requests are handled by o3-mini, and the remaining 10% of the more complex requests are processed by the larger model, resulting in a cost reduction of more than 50%.

Common Small Issues

Question: Should we make any fine-tuning to o3-mini? Answer: If your use case involves general content generation, there's absolutely no need; the native performance is sufficient. However, if you're dealing with a scenario that contains a large number of industry-specific terms, feeding in a few hundred samples for fine-tuning could be beneficial. After fine-tuning, the accuracy of our auto parts product descriptions increased by 14%, while the cost only increased by 5%.

Question: Is the data secure? Answer: By default, models are not trained using user input data. If you have compliance requirements, you can opt for a dedicated instance deployment, which will cost 30% more, but the data will be completely isolated, making it suitable for processing content related to user privacy.

Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR