Menu

We saved 32% on manpower by using GPT-4 to process millions of after-sales tickets. Avoid these three common mistakes.

During the peak period of last month's Black Friday promotions, three of our team members were closely monitoring the control panels. We witnessed the error rate for after-sales ticket requests soar from 1% to 12%—the small model we were using before simply couldn't handle the massive volume of multi-language inquiries, which amounted to 1.2 million per day. As a result, the number of customer complaints increased by four times within 24 hours. We had to temporarily switch to using GPT-4 to reduce the error rate to below 0.3%.

First, let's understand: What exactly can GPT-4 save for small and medium-sized teams?

Scrabble tiles spelling "CHATGPT" on wooden surface, emphasizing AI language models.

Simply put, it is the most balanced large model currently available from OpenAI for enterprises. It supports up to 128k of context input and has a multi-language understanding accuracy that is more than 40% higher than that of most specialized, smaller models. You can use it to handle complex scenarios directly without having to perform much additional fine-tuning.

For our small team of 20 people that operates cross-country home appliance retail, the three most immediate benefits are so obvious that we don’t even need to calculate a complex ROI (Return on Investment):

  • Firstly, multi-language support for tickets eliminates the need for dedicated translators. It can directly understand after-sales inquiries in Spanish, French, and Polish, and provide standard responses in those languages.Our after-sales staff has been directly reduced by 32%.
  • The accuracy of diagnosing complex faults has significantly increased compared to the smaller models used before. Previously, 17% of the tickets required manual re-verification, but now that figure is less than 4%.
  • We don't need to accumulate a large amount of finely tuned training data. Just throw the new product manuals into the system, and it can adapt to the corresponding after-sales support queries in 2 hours. The efficiency improves significantly during major promotional campaigns.

Don't just look at the advantages; we've actually encountered several real pitfalls.

The first issue is the problem of cost out of control. When we switched to GPT-4, we didn't implement any input truncation; the text generated from dozens of pages of fault screenshots uploaded by users was simply fed into the system, resulting in higher daily token costs than in the previous week. Later, we added a preliminary filtering step to remove unnecessary and redundant information, which immediately reduced the costs by half.

The second issue is that rate limiting can be really troublesome. If you directly use the official API during a temporary peak sales period, it's very easy to hit the concurrency limit. We experienced a 10-minute rate limit on Black Friday, and it wasn't until we submitted a temporary capacity expansion request three days in advance to the customer manager that the problem was resolved. Never wait until the traffic has increased before making a request.

The third issue is the misclassification of sensitive content. A user simply asked, "Will the oven explode if the outer shell gets hot while it's heating?" and the content security system blocked the request, returning an error. Later, we added a pre-check for sensitive words to filter out legitimate troubleshooting queries first, which reduced the misclassification rate to an acceptable level.

Should we switch to GPT-4 now? Let’s look at these two criteria first.

Hand holding a smartphone with AI chatbot app, emphasizing artificial intelligence and technology.

Use cases:If you need to handle multilingual content, complex long-text reasoning (such as summarizing documents with dozens of pages, conducting multiple rounds of after-sales consultations, or debugging code), or if your team doesn't have specialized algorithm experts for fine-tuning small models, using GPT-4 directly is much more cost-effective than trying to develop your own models.

Use cases where it should not be used:If your scenario is particularly simple, such as just generating keyword responses, sending messages using fixed templates, or ensuring that all data does not leave the domain, then there's really no need for it. You can achieve the same results with a cheap small model or even a rule-based engine; spending the money on this would be a complete waste.

Two specific suggestions for teams that are starting out for the first time

Close-up of a computer screen displaying ChatGPT interface in a dark setting.

First, use 10% of the traffic for a grayscale test for 7 days to see if the actual accuracy, cost, and response speed meet your expectations. Don’t make a full-scale switch right from the start; we encountered rate limiting issues when we did that. Now, the grayscale test has become a fixed part of our model replacement process.

Secondly, it is essential to add a preliminary traffic scheduling layer that distributes simple requests to cheaper, smaller models. Only complex requests that the smaller models cannot handle are sent to GPT-4. After implementing this, the overall cost of the models was reduced by another 28%, without any compromise on performance.

Two common small issues

Question: Could there be a data breach? Answer: If the Enterprise Edition API is chosen, OpenAI has clearly stated that they will not use the user's input data to train models. We have conducted compliance audits, and it complies with the EU's GDPR requirements, so you can use it with confidence.

Question: Will the response speed be very slow? Answer: Most normal ticket inquiry requests are returned within 150 milliseconds; only very long text processing may take up to 1 second, which is more than sufficient for after-sales scenarios.

Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR