Running supply chain forecasting with o1-mini: We saved 62% on inference costs, but we ran into 2 fatal pitfalls.
Last Black Friday, our team almost failed.
We were previously using large models to predict the demand for warehouse distribution in North America. During the peak period, for three consecutive days, more than 18% of the requests resulted in a 429 error code, directly causing disruptions. As a result, the operations team's inventory planning was in disarray, with best-selling products in three states being delayed by 48 hours.
The requirement to switch models was directly marked as P0 at that time. We spent 7 days testing 4 alternatives, and in the end, we chose o1-mini. It has been running for 22 days now, and all the issues have been resolved. However, we also encountered some pitfalls that no one had mentioned in advance.
For those who haven’t heard of it before: what exactly is o1-mini?
It's the lightweight inference model launched by OpenAI, which is specifically optimized for the processing speed of logical and computational tasks.Basic reasoning capabilities differ by no more than 8% from those of the leading large models, and the cost per token is only 1/7 of that of the leading models.。
We started out with the aim of exploiting this cost difference, considering that our prediction task requires feeding in over 3,000 tokens of historical orders, weather data, and holiday information each time, and it runs nearly 20,000 times a day.
The three most obvious benefits we have actually achieved
- First of all, the throttling issue has been completely resolved: The single-account throttling threshold for o1-mini is 12 times that of the model we used before. Even on Black Friday, when the system was at its peak load, there was not a single instance of a 429 error (a common error indicating throttling). The operational side's planning system ensures zero latency in output.
- The costs have been significantly reduced: According to the billing data from our backend, running the same prediction task...The monthly reasoning cost has decreased from $12,000 to $4,500.The money we saved was used to add real-time data points for three more warehouses.
- The speed is fast enough for real-time adjustments: previously, it took more than 20 seconds for the model to make a single prediction, but now most requests are completed within 15 milliseconds. We have changed the frequency of prediction updates from once a day to once every 2 hours, and as a result, the out-of-stock rate has decreased by 4 percentage points.
But there were two pitfalls, and we almost had to roll back the changes when we encountered them.

The first problem is its extremely poor ability to handle unstructured data.
In our previous inputs, there were some keywords from users' discussions on social media platforms, which we used as a reference for demand fluctuations. Unfortunately, the o1-mini system treated this information as invalid, resulting in forecast data that was 20% lower than the actual demand for the first three days. Fortunately, we have a cross-validation mechanism in place, so we didn't actually proceed with stock replenishment based on those predictions.
In the end, we can only run a small text analysis model separately to process this part of the data first, convert it into structured scores, and then feed it to o1-mini in order to solve the problem.
The second issue is that its multilingual processing has hidden biases.
In our predictions for the Canadian region, there are French place names and product names. By default, the o1-mini system maps some French categories to their English equivalents, which resulted in the ski equipment predictions for Quebec being reduced by half. We only resolved this issue by adding a rule in the prompt that requires the original language to be retained. This problem is not mentioned at all in the official documentation.
Who should use it and who shouldn’t – we’ve already made the distinctions for you.
If you, like us, are doing what we do...Logical reasoning, numerical calculations, and rule-based decision-makingFor tasks that are primarily data-driven, with high demands on cost and concurrency, o1-mini is basically an excellent choice without much need for further consideration.
But if you're working with content generation, multimodal understanding, or complex multilingual processing, don't even try it. The results it produces can be full of unexpected errors, and the time and effort spent troubleshooting will outweigh any savings you might make.
Two specific suggestions for beginners
First of all, before going live, you must conduct at least 7 days of parallel verification tests; do not directly switch to the production environment with traffic.
We initially wanted to save effort and only ran the process for 3 days, which happened to be a period when no French input was encountered, so we almost ran into a big problem. A 7-day cycle should be able to cover most of the edge cases in your business.
Second, do not alter its temperature parameters arbitrarily.
We initially wanted to make the results more flexible by setting the temperature to 0.7, but the resulting predictive data was too volatile to use. It wasn’t until we switched back to the officially recommended range of 0.1-0.3 that the data became stable. The purpose of this system is to perform deterministic reasoning tasks; if you need creativity, it’s better to use a large model directly.
Finally, let's address a question that many people ask: Should we wait for the next version?
At least in the context of supply chain forecasting, the current o1-mini is already sufficient. We've done the calculations, and even if the next generation version were 20% cheaper, it wouldn't make up for the savings in manpower and the reduced number of failures that we're achieving with the current system. The earlier we adopt it, the sooner we can start making profits.
Article link:https://airai.cc/en/ai-news/36/
Was this helpful?