Menu

We transferred the customer service ticket processing to Claude 3.5 Sonnet, saving 21,000 euros in labor costs over three months.

Last month, our European e-commerce customer service team was on the brink of collapse: a week before the Black Friday sales rush, we were receiving over 1,200 emails per day from multilingual users. Even with three customer service representatives working non-stop, 40% of the tickets still took more than 24 hours to respond to, and the user satisfaction rating dropped by 18%. Hiring temporary part-time staff not only meant that the training process couldn't keep up, but it also cost nearly 20,000 euros more per month.

For those who haven't heard of it before, here's a brief explanation:

Claude 3.5 Sonnet is a multimodal large model launched by Anthropic in 2024.With a context window of up to 200K, the inference speed is twice as fast as the previous generation.It can just accommodate all the information from our single ticket, as well as nearly three months of user history orders and communication records.

The three core benefits we utilize each address a specific pain point.

Hand holding a smartphone with AI chatbot app, emphasizing artificial intelligence and technology.

  • Multilingual understanding accuracy: Our users speak four languages: German, French, Spanish, and Italian. Previously, using smaller models to handle non-English tickets often led to confusion between "refund" and "exchange." With Sonnet, the error rate is less than 2%, allowing us to process these requests directly without the need for additional translation tools.
  • Stable processing of long documents: When users request a return, they often attach 5 or 6 photos of damaged products along with a few hundred words of description. Sonnet can process all the images, text, and past order records at once, and generate a handling plan that complies with our after-sales policies, eliminating the need for manual verification of the information.
  • The cost is low enough: Our tests show that the cost for processing 1,000 tickets is less than 10 euros in tokens. Even including the fees for the gateway and fine-tuning, the total monthly expense is less than one-fifth of the monthly salary of a full-time customer service representative.

In the first month of using Sonnet, our average response time for tickets decreased from 17 hours to 1.5 hours.The customer service no longer needs to deal with duplicate returns or logistics inquiries; they only have to handle about 10% of the more complex disputes. No additional staff were hired during the promotional periods.

Don't just look at the benefits; we actually ran into two real problems in the past two weeks.

The first issue was related to the recognition of complex rules by the algorithm: We had a specific rule that stated "customized products can be refunded even if there are quality issues after more than 7 days." Initially, when we simply fed this rule into the model, it would still reject users according to the general policy of "no refunds or exchanges for customized products." Later, we created a separate trigger prompt for this type of rule, which would be invoked whenever relevant keywords appeared, and this solved the problem.

The second issue is rate limiting: On the day of the major promotion, the number of peak requests was three times higher than usual. Initially, we did not implement any degradation strategies, which led to some requests receiving a 429 error response. Later, we added a simple queuing mechanism that delayed the processing of non-urgent tickets by 10 minutes, and since then, no failures have occurred again.

To put it clearly: some people are suitable for using it, while others absolutely shouldn’t even consider it.

Scrabble tiles spelling "CHATGPT" on wooden surface, emphasizing AI language models.

If your team meets either of these two criteria, it's a good choice to use this approach: First, if you need to handle a large amount of long, multi-modal structured data, such as customer service tickets, contract reviews, or summaries of long documents; second, if your business covers multiple language regions and you don't want to fine-tune the model for each language separately.

If you only need to perform simple keyword responses, short-text classification, or have extremely high requirements for data privacy and cannot transfer data to third-party models at all, then there's no need to choose it; a smaller model or even a rule-based engine would be sufficient.

Two practical tips for first-time users

For the first test, don’t make a full switch all at once. Start by running the historical tickets that have been processed over the past month. Compare the results output by the model with the results from manual processing. Only deploy the system when the accuracy rate is above 90%; this will help avoid about 80% of the rule adaptation issues.

If your request volume fluctuates significantly, it's essential to implement a simple downgrade queue. Sonnet has strict limits on the free tier of its quotas. During peak times, you can prioritize lower-priority tasks and defer them. This way, you can meet the needs of most small and medium-sized enterprises without incurring additional costs for dedicated throughput.

Several questions you might have

Q: Is it necessary to perform any specific fine-tuning? A: We just included the after-sales rules in the prompt; without any fine-tuning, the accuracy is already sufficient. Unless your business rules are particularly complex, there's no need to spend the extra money on it.

Q: Which is more cost-effective compared to GPT-4o? A: Our tests show that Sonnet costs about 30% less to process the same number of tickets, with similar accuracy. For small and medium-sized enterprises that are sensitive to costs, Sonnet is a more suitable choice.

Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR