AirAi Ecosystem Cooperation Program: Allocating a budget of $10,000 to sponsor those who are still actively working on their projects
...
0. Background: How did the costs explode?I work on several side projects. In the early stages, I either had to pay the official subscription fee of /month or endure the rejection of my international credit card applications. When I settled the accounts at the end of the month, I discovered two frustrating issues:80% of the requests were for simple tasks such as summarizing, categorizing, and formatting data. Using the most expensive models for these tasks was a complete waste;The subscription fee is a fixed cost – it results in losses when there are few requests, and it's not enough when there are many requests.
As large language models gradually tackle complex tasks such as reasoning, automated research, and cybersecurity, traditional methods of evaluating models are facing new challenges.For a long time, the release of models has been accompanied by a report of results consisting of various benchmark tests in areas such as mathematics, programming, scientific question answering, network security, and knowledge reasoning, which are then compared horizontally with the previous generation of models.
Voice AI represents another narrative that unfolds alongside the development of general large models. While everyone is focused on the general large models, the relatively quieter field of Voice AI is also seeing the emergence of some noteworthy new models. The keyboard is starting to lose its “dominant position.” Over the past two years, OpenAI introduced the Realtime API, Google launched Gemini Live, and domestic large-model companies have almost all begun to invest in Voice AI. More and more people believe that once agents truly integrate into workflows, voice will become a more natural way to interact with systems than using a keyboard. For an agent to truly become part of a workflow, it must first learn to understand human speech. The foundational capability for this is ASR (Automatic Speech Recognition). The most commonly used benchmark for measuring ASR performance is Hugging Face’s Open ASR Leaderboard, which uses the Word Error Rate (WER) as a key indicator. The lower the WER, the more accurate the recognition.
No More