Menu

We increased the efficiency of handling customer after-sales tickets by a factor of 3 using Claude 3, but we ran into two problems that we shouldn’t have encountered.

When we finished pulling the reports last Black Friday, three developers from our Southeast Asia e-commerce SaaS team were stunned by the backend data: The number of after-sales tickets that used to require the entire customer service team to work non-stop for three days to process was automatically handled by the system this year, with the volume of customer complaints dropping by 42%. There was only one key change: the original ticket semantic recognition logic was completely replaced with Claude 3 Opus.

To be honest with someone who hasn’t encountered this before…

It is the third-generation large language model launched by Anthropic in 2024. The core parameter we used was quite simple: with a context window of around 200k, it handled after-sales tickets containing 3 product screenshots with an accuracy rate 18% higher than the previous models using the same parameter level.

For a small team like ours, the three benefits are tangible and valuable in terms of real financial gains.

Abstract black and white graphic featuring a multimodal model pattern with various shapes.

The first point is that there's no need to perform multi-modal preprocessing separately anymore. For the damaged product images and logistics screenshots uploaded by users before, we had to first use an OCR model to convert the images into text, and then combine the text with the corresponding work orders and feed them into the large model. Just maintaining this process consumed half of the backend resources. With the adoption of Claude 3, we can now send both the images and the text directly, and the number of recognition errors has decreased.

The second point is that the rejection rate is so low that it can be almost ignored. Previous models would often directly return an inability to recognize the tickets when users wrote them in a very messy manner (for example, using a mix of English and local language, along with internet abbreviations), and the tickets had to be transferred to human operators for processing. We recorded that the proportion of tickets being transferred to humans during that period was 27%; now, that number has dropped to 4%.

The third item is that the call costs were much lower than we expected. We pay based on the actual usage; on the peak day of Black Friday, we processed an average of 12,000 tickets per day, and the total cost was less than $800, which is a 90% saving compared to hiring 10 temporary customer service representatives.

But don’t rush into action just yet; the two mistakes we made are enough to make you waste a whole week working for nothing.

The first issue was the problem with multi-language alignment. Half of our users speak Indonesian, and we launched the product without making any fine-tuning adjustments. As a result, when the model encountered local slang, it often misinterpreted phrases like “the product was sent in the wrong color” as “the user wants to change the delivery address,” leading to 17 customer complaints in just one week. We solved this problem only after feeding the model 3,000 locally annotated tickets for fine-tuning.

The second issue is the omission of information within a long context. If a ticket contains more than 5 images, the model occasionally misses information from one of the images. For example, if a user uploads images of damaged packaging and the product itself, the model may only recognize the packaging issue. We later added a simple validation rule: if there are more than 3 images, the model is instructed to output the recognition results for each image individually before summarizing them, and since then, no such errors have occurred.

To be honest, not all teams are suitable for using it.

Abstract representation of a multimodal model with dots and lines on a white background.

The situations you should use it for:

Scrabble tiles spelling "CHATGPT" on wooden surface, emphasizing AI language models.

  • Your business needs to handle both text and images/short audio simultaneously, and you don't want to set up multiple model pipelines.
  • You have high requirements for the accuracy of the output, especially in scenarios where handling errors can result in significant costs, such as with work orders and contract reviews.
  • Your team lacks the manpower to maintain the complex model preprocessing pipeline.

Situations where you shouldn't waste your money:

  • You just need to perform simple tasks such as providing keyword responses and generating marketing copy – lightweight tasks that can be handled by inexpensive, small models.
  • Your business data has strict localization storage requirements, which prevent it from being transmitted to third-party model interfaces.
  • Your request volume is particularly low, less than 1,000 times per month, and the development cost of switching to a different model is higher than the benefits it would bring.

Two practical tips for beginners

The first step is to test the system in edge scenarios for 7 days before deploying it to the core infrastructure. Initially, we used historical tickets from the past 3 months for offline testing, and only after the accuracy rate exceeded 95% did we start diverting 10% of the online traffic to the new system. We gradually increased the proportion of online traffic until the entire system was transitioned, without any major issues occurring.

The second point is not to choose the most expensive Opus version from the start. We have tested it, and for handling ordinary consultation tickets without images, the accuracy of the Sonnet version is less than 2% different from that of Opus, yet the cost is only half. That’s more than sufficient.

Finally, let's address a question that many people ask: Should we wait for the next generation of models? Our answer is that if your current business operations are already being hindered by the limitations of multi-modal processing and insufficient accuracy, using the current models is perfectly suitable. After all, the time and manpower you save by starting to use them now will far outweigh the additional costs associated with calling the next generation of models.

Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR