We used GPT-4o to increase the accuracy of our multimodal customer service to 92%, but we really suffered a lot from these three issues.
Last month, our three-person technical team working on cross-border e-commerce in Southeast Asia stayed up for three days to migrate the intelligent customer service system from the old text-based large model to GPT-4o. We thought it was just about adding an image recognition feature, but the changes were much more significant than we anticipated, and we encountered several unexpected issues along the way.
Let me explain clearly for those who haven’t come across it before: What exactly is GPT-4o?
The essence is OpenAI's multimodal generation model, and the most significant difference from its predecessors isA single round can process three types of inputs simultaneously: text, images, and audio, without the need to split the requests and call different models separately.The interface parameters are nearly 60% fewer compared to combined calls, without the need for you to preprocess the input format yourself.
Why do we have to change it at the specific time point of 2026?
Previously, our customer service system could only handle text-based inquiries. When users sent photos of damaged products or misdelivered items, the images had to be manually converted into text descriptions before being fed into the models. During promotional periods, this process alone caused nearly 20% of the support tickets to get stuck, and the user complaint rate doubled. We also tried a solution that combined an image recognition model with a text recognition model, but the accuracy was only 76%. The cost of handling misjudged after-sales requests was even higher than the cost of the models themselves.
The three most immediate benefits have far exceeded our initial expectations.

- The automatic processing rate of after-sales tickets has increased directly from 47% to 92%. The two part-time customer service agents who were previously responsible for transcribing images have been reassigned to handle more complex customer complaints, resulting in a monthly savings of nearly $1,800 in labor costs.
- The dialectal voice consultations sent by users no longer need to go through a separate ASR (Automatic Speech Recognition) interface; they can be directly transmitted to the model for recognition and a response. Most requests are completed within 15 milliseconds, which means the overall response speed has increased by three times.
- Processing the same after-sales request that includes both images and text, the token consumption was reduced by 32% compared to calling the image recognition and text models separately. As a result, the model cost during the peak promotional period is actually lower than before.
Don't just look at the benefits; the three hidden pitfalls we encountered led to real losses when we tried to use them.
The first issue was that the throttling threshold for multimodal inputs was much stricter than that for plain text requests. On the very first day of the service launch, we encountered a store anniversary event on the Southeast Asian site. 17% of the requests that included images resulted in a 429 error, and the users couldn't see any response, leading them to think that customer service had disappeared. As a result, we received over 300 negative reviews that day. We had no choice but to create a separate downgrade queue for image-containing requests, displaying a message stating "Image received, processing" during peak times, which finally resolved the problem.
The second issue is that the accuracy of the image text recognition for non-English minority languages is far from as high as it is advertised. We encountered a case where a user sent us a handwritten Indonesian delivery note, and the model misrecognized three address characters, resulting in the user being assigned the wrong return or exchange address. This led to a loss of nearly $200 in shipping costs. Now, we have added a local OCR (Optical Character Recognition) interface as an additional layer of verification for the recognition of images in minority languages, which has helped reduce the error rate to an acceptable level.
The third issue is that the token counting rules for multimodal outputs are completely different from those for plain text, so you can't use the previous formulas to estimate the cost. In the first week it was launched, we didn't adjust the budget alert thresholds, and the cost for just one weekend exceeded 40% of the monthly budget. We only realized this when the finance department sent us a reminder.
Should we make the change now? We have established clear criteria for making that decision.
Situations where you can just rush through:

- Your business itself requires processing two or more types of inputs, including text, images, and audio, simultaneously.
- We have already been using multiple models for splicing to perform multimodal processing, but the integration cost is high and the accuracy is not satisfactory.
- Your users primarily speak English, or the mainstream languages are those for which GPT-4o provides relatively good support.
Situations where you absolutely don't need to waste money:
- Your business only requires pure text processing, and the previous model was already capable of meeting the accuracy requirements.
- Your business primarily targets small-language markets, involving a large number of requirements for the recognition of local handwritten text and dialectal speech.
- Your team doesn't have any extra resources to develop the necessary support for downgrade strategies and cost monitoring.
2 Practical Tips for Teams Just Getting Started
First, run a shadow traffic test for 7 days before going live; don’t immediately switch 100% of the requests. Send real user requests to both your original system and GPT-4o simultaneously, and compare the accuracy and costs of both. Only gradually switch the traffic after confirming that the results meet your expectations. We made a big mistake by not following this step.
Secondly, set budget alerts specifically for multimodal requests. The threshold can be set to 120% of the estimated cost to prevent cost overruns.
Frequently Asked Questions
Question: Should we wait for the next generation of models to come out before making a change?
Answer: At least in the context of multimodal integration, the current maturity of GPT-4o is sufficient to solve practical problems. The next generation of models is still at least a year away, so there's no need to waste the costs that can be saved now for an uncertain upgrade.
Question: How does it compare in terms of cost-effectiveness compared to open-source multimodal models?
Answer: If your team has the capability to fine-tune and deploy open-source models on their own, and there are high requirements for data privacy, then open-source models are more cost-effective. Otherwise, using GPT-4o directly is much cheaper than setting up and maintaining your own servers.
Article link:https://airai.cc/en/ai-news/30/
Was this helpful?