Menu

The rapidly heating Voice AI competition has seen the emergence of a startup team called Hojo.

Voice AI represents another narrative path that unfolds alongside the development of general large models. While everyone is focusing on these general models, the relatively quieter field of Voice AI is also seeing the emergence of some noteworthy new models. The keyboard is starting to lose its “dominant position.” Over the past two years, OpenAI introduced the Realtime API, Google launched Gemini Live, and domestic large-model companies have almost all begun to invest in Voice AI. More and more people are convinced that, once agents truly integrate into workflows, voice will become a more natural way to interact with systems than using a keyboard. For an agent to effectively integrate into workflows, it first needs to learn to understand human speech. The foundational capability for this is ASR (Automatic Speech Recognition). The most commonly used benchmark for measuring ASR performance is Hugging Face’s Open ASR Leaderboard, with the key metric being the Word Error Rate (WER); the lower this rate, the more accurate the recognition.

For a long time, this leaderboard has been dominated by large companies and prestigious research laboratories. The requirements for training data, computing power, and engineering expertise are quite high, making it difficult for smaller teams to compete directly in this area. Recently, a startup called Hojo has made public a set of speech recognition test results and seems to be on the verge of becoming a “dark horse” in the field. According to the official data, Hojo-ASR-V1 has achieved impressive results on several publicly available English speech recognition datasets.

Among them, the Word Error Rate (WER) on LibriSpeech Clean is only 1.74%, and on datasets that are closer to real-life scenarios, such as GigaSpeech and VoxPopuli, it has also been reduced to within 8%. They did not submit their results to the leaderboard, but if the data were included in the Open ASR Leaderboard comparison, it would rank very high. Overall, this is an ASR model that demonstrates competitiveness in public benchmark tests. Moreover, it has been open-sourced on GitHub under the Apache-2.0 license, with relatively few restrictions on its reuse. 🚥 So, we decided to deploy Hojo-ASR-V1 to experience it firsthand and see what level this low-profile startup team has achieved with their speech recognition model. It also leads us to discuss a new trend that is gaining momentum: the Agent era of Voice AI. What exactly is Hojo-ASR? We tested it out ourselves. First of all, Hojo-ASR-V1 has been open-sourced on HuggingFace and GitHub: Hugging Face link: https://huggingface.co/HojoAI/Hojo-ASR-V1[1] GitHub link: https://github.com/HojoAI/Hojo-ASR[2] Let's clarify what kind of model Hojo-ASR-V1 is. After actually reviewing its code and configuration, we found that its structure is quite different from traditional speech recognition models.

The audio is first processed by Whisper’s feature extractor to convert it into acoustic features that can be used by the model, and then it is fed into the Qwen3-Omni audio encoder. In between, a Conformer structure is used for adaptation and compression.

Finally, the output is handed over to a Qwen3-4B language model, which generates the final text. In other words: Hojo-ASR-V1 is a combination of an “encoder, an adapter, and a large language model.” The idea of using a large language model for ASR decoding is not unique to Hojo; it represents the mainstream trend at the top of the current OpenASR rankings. Several models that rank high on the list, such as NVIDIA’s Canary-Qwen-2.5B, IBM’s Granite-Speech-3.3-8B, and Microsoft’s Phi-4-Multimodal, all integrate the capabilities of language models into speech recognition, reducing the average WER (Word Error Rate) to the range of 5.6% to 5.9%. The advantage of incorporating a language model is that recognition is not just about “hearing what is said and writing it down”; the model can also use semantics to make judgments. When dealing with noise, colloquial speech, a mix of English and Chinese, or specialized terminology from a particular field, a language model that understands semantics can better understand what the speaker intends to convey. Below is the results section: we deployed the model locally and ran multiple tests.

Test Cases, Divided into Two Layers: 【1】The ASR (Automatic Speech Recognition) capabilities of the Hojo-ASR-V1 model. 【2】Hojo-TTS (Text-to-Speech). ASR (Automatic Speech Recognition) ASR stands for Automatic Speech Recognition, which is the process of converting speech into text. It is the first step in the entire Voice AI ecosystem. Whether it's mobile voice input methods, meeting transcripts, real-time subtitles, or the increasingly popular AI Agents, all rely on ASR. I used to be a user of AI voice input tools like Wispr Flow and Typeless. These products provided a good experience, but there were occasional delays when the network was unstable, and there was still room for improvement in their adaptation to Chinese language scenarios. Later on, I gradually switched to locally deployed solutions, with OpenAI's open-source Whisper being one of the most widely used. After the release of Hojo-ASR-V1, I replaced Whisper with Hojo-ASR-V1 to get a firsthand experience. Overall, the recognition speed and accuracy were quite good, making it sufficient for daily voice input tasks. Of course, this comes at the cost of requiring certain local hardware resources and having specific memory requirements. The entire solution is not complicated; it essentially uses Hojo-ASR-V1 as the underlying recognition engine and then takes over voice input through system-level permissions. Once configured, it can work in almost all text input scenarios, such as browsers, ChatGPT, Claude, Notion, and more. During the test, I orally described a segment about a "crossroads," including the location of the feature, the direction of the content, and the origin of its name. I didn't prepare the text in advance, nor did I deliberately slow down my speech speed. What I found interesting is that this type of workflow can be further expanded. After the recognition is complete, it's possible to integrate with DeepSeek, GPT, or other models to refine the text, organize its format, correct spelling errors, and optimize its structure.

The general process is as follows: Sound input → ASR (Automatic Speech Recognition) transcription → optimization by a large model → direct integration into the workflow. For people who frequently write, attend meetings, or collaborate with agents, this experience is more comfortable than using traditional input methods. In addition to Hojo-ASR-V1, we have also noticed that Hojo has made investments in the field of TTS (Text-to-Speech). This time, we tested its TTS speech synthesis model, and the team has also open-sourced a lighter version called Hojo-TTS-Light. Together, these components constitute Hojo’s solution for agent workflows. Another aspect of Voice AI is TTS. The path to test the capabilities of the TTS model can be found on their official website: [hojoai.com](hojoai.com). **Multilingual Speech Synthesis – Cheerful Female Voice** We started with testing the multilingual speech synthesis feature, which is a commonly used function of TTS models. With the World Cup approaching, we asked the model to use a cheerful female voice to read a segment introducing the Japanese team players: The approximately 30-second segment of Chinese sounded quite smooth, with proper enunciation and handling of details; the AI presence was not very noticeable. The female voice’s tone and cheerful tone were clearly audible. Usually, we judge whether a voice sounds unnatural by listening to its enunciation and the tone at the end of each word, and this segment was in line with the initial style set. Hojo supports several languages, including Japanese, French, and Cantonese. The same introduction to the Japanese team was read in Japanese, and the result was also satisfactory; it really sounded like an NHK broadcast, giving a sense of authenticity. **Multilingual Speech Synthesis – Magnetic, Steady Male Voice** This TTS model can also switch to different male voices, and the male voice sounded quite good as well. The following segment was also an introduction to the 2026 World Cup, and I set the voice to a magnetic, steady male tone. This segment matched the style I specified. **Multispectral Speech Synthesis – Audiobooks** The TTS model can use a variety of voices. This time, we asked it to read a segment in the voice used for audiobooks. The text was a theme introduction from “Naruto,” and the resulting audio sounded very natural. The handling of word endings, such as the word “shōnen” (boy), ensured a smooth transition between phrases. Even for conjunctions like “to” (which we use in everyday speech), the pauses were precise, which contributed to the overall smoothness of the audio. **Multispectral Speech Synthesis – Whisper** During the testing, I wanted to mention a particular voice style: whisper. In the following segment, I used a female whisper voice to read a movie introduction for “The Silent Lamb.” You can clearly hear what is being said, and the model successfully simulated the soft airflow, subtle breathing, and gentle sounds that would be produced when someone speaks close to your ear. The whisper voice from the TTS model didn’t just reduce the volume; the tone, rhythm, and emotion were also well-handled. This voice style is suitable for suspenseful narrations, story narration, and ASMR (Autonomous Sensory Meridian Response) content. In addition to the standard broadcast voice, I also tested some more recognizable character voices. **Kang Hui** Since it’s the college entrance examination season, we can hear many blessings and news broadcasts on television, radio, and short videos. The voice of CCTV announcers is quite representative in Chinese broadcasts, with standard pronunciation and clear enunciation. Therefore, I used a segment of Kang Hui’s audio as a reference for this test. The following segment is Kang Hui’s original audio. File Copying Timbre - Kanghui I sent this original audio to a TTS (Text-to-Speech) model, asking it to replicate the timbre used in the audio and read a script that is about 40 seconds long, wishing the students taking the college entrance examination success.

Effectively, the TTS model has retained the characteristics of the original voice quite well, especially Kang Hui’s tone, pauses, and delivery rhythm. As soon as the first sentence is spoken, it’s clear that it sounds very much like Kang Hui himself. When phrases like “Live up to our youth and strive for the future” are mentioned, the familiar style of news broadcasting and the sense of hosting a large event are evident. Nezha The ability of the TTS model to replicate the voice of a character is quite representative of its capabilities. I used it to recreate the voice of several IP characters, starting with Nezha from the movie “Nezha: The Demonic Child’s Birth.” The resulting voice is quite close to Nezha’s unrestrained and lazy demeanor, with pauses that match that character’s style. Peppa Pig The TTS model also does a decent job of replicating the voice of cartoon characters. The voices of cartoon characters are different from those of real people or movie characters, so there are some differences in how they need to be processed. Here’s a segment of Peppa Pig synthesized by the TTS model: MacArthur In fact, I even recorded a popular “MacArthur” commentary voice from TikTok and had it read out: “Crossroads” is a Chinese technology media platform that focuses on the AI trend. Its main format is a podcast, with extensions including videos, text versions on official accounts, and offline community events. “Crossroads” is a metaphor coined by Steve Jobs for Apple, describing the company as standing at the intersection of technology and humanities, where great products often emerge. The platform focuses on the changes and opportunities that AI brings to various industries, seeks out, interviews, and brings together “active actors” of the AI era, and explores and embraces new possibilities with them. Who is the Hojo team? Returning to the Hojo team itself, it’s not a company that only works on large-scale voice models.

Hojo was established about two years ago, and before this information was made public, there was very little knowledge about it outside of its circle of insiders. We managed to get some first-hand information by contacting its team. According to Hojo's own positioning, its main goal is to create a Personal Agent OS for knowledge workers. In simple terms, the aim is to integrate the Agent into everyday work processes such as office work, communication, creation, and collaboration. ASR (Automatic Speech Recognition) and TTS (Text-to-Speech) are two crucial components of this system. This is also key to understanding Hojo's approach. ASR is essentially just the first layer that enables the Agent to understand real work scenarios. Only by reliably capturing information from meetings, phone calls, voice memos, and group discussions can subsequent processes of understanding, planning, and execution be possible. One unique aspect of the Hojo team is that its core members have long experience working with "real-world voice scenarios." Many of the team members come from companies specializing in in-vehicle voice systems, intelligent cockpits, and voice interactions, including Xpeng, NIO Nomi, and SenseTime. To be more specific, the founder, Zhao Hengyi, previously worked on voice interactions, intelligent cockpits, and AgentOS at companies like LeEco, Xpeng, NIO, and SenseTime, focusing on in-vehicle voice systems, intelligent cockpits, multi-language interactions, AI-driven applications, and cross-device services. The Chief Product Officer, Li Ou, was in charge of AgentOS, intelligent cockpits, and digital product planning at SenseTime and NIO, focusing on how voice capabilities can be integrated into a complete product experience, rather than just developing individual model capabilities. The COO, Sun Xiaogang, has experience in commercializing Agent solutions and implementing voice interactions in various industries, including automotive, catering, and transportation sectors. Considering these backgrounds, the collective experience of the Hojo team provides a solid foundation for addressing real-world voice challenges on a long-term basis. This is mainly because the team members have all focused on ASR in real, practical environments, which are quite different from the clean audio conditions typically found in laboratories. In real-world scenarios, there is noise, dialects, multiple speakers, interruptions, and a lot of informal, incomplete speech.

Users do not speak according to standard prompts, nor do they wait for the system to respond slowly. Those who have worked in such scenarios can better understand the real challenges faced by speech recognition models. ASR (Automatic Speech Recognition) needs to reliably interpret people's intentions in complex environments. Hojo’s chief scientist, Sudan,’s background makes this all the more understandable. In his early years at Baidu, he was involved in speech recognition research and development, witnessing the commercialization of ASR through deep learning. Later, at Tencent AI Lab, he led the speech-related efforts, coinciding with the period when large models revolutionized voice interactions. The value of his resume lies in his firsthand experience of the evolution of speech technology: from a focus on recognition accuracy to practical applications in real-world scenarios, and then to the reconfiguration of voice interactions in the era of large models. For a company aiming to develop a Personal Agent OS, this experience is crucial, as it ensures that speech is not merely treated as an input component but is integrated into the overall workflow for a redesign. From a capital perspective, Hojo has raised nearly 100 million yuan in four rounds of financing even before its product was officially launched, and it is currently in the Pre-A round. While this does not guarantee success, it indicates that the primary market has begun to recognize the company’s technical expertise in Voice AI and Agent OS. In terms of commercialization, according to Hojo, its proprietary “large model + Agentic OS” combination has already provided foundational voice capabilities for tens of millions of AI-powered devices. If this figure is accurate, it means the technology has moved beyond the stage of research papers, demos, and rankings, and has been implemented in real devices and user scenarios. Hojo’s own plans indicate that ASR is just the first step. The company is also working on TTS (Text-to-Speech) and plans to launch a full-duplex voice model by the end of June. The ultimate goal is to build a comprehensive voice interaction framework for Personal Agent OSs, enabling agents to hear, understand, and respond, thereby integrating into the real workstreams of knowledge workers. Whether these plans will be realized remains to be seen with future products. However, based on the performance of Hojo-ASR-V1, it has already made a noteworthy contribution to the Voice AI workflow. So, what exactly makes the data released by Hojo-ASR so compelling? Or, in other words, what is the context in which Hojo-ASR is developed? Voice AI represents another narrative within the broader realm of large models. As agents begin to integrate into real workflows, the context they need to process becomes increasingly complex: discussions in meetings, requests over phone calls, sounds in the environment, and spontaneous additions from users, as well as constantly changing instructions during collaborative tasks. In the past, this information was mainly captured through typing. However, in many work scenarios, the information that matters is conveyed verbally, not written down.

This is also why large companies have been increasing their investment in Voice AI over the past two years. OpenAI has launched the Realtime API, Google has developed Gemini Live, and domestic companies such as iFlytek, DouBao, MiniMax, and JieYue have also invested heavily in the voice technology sector. The popularity of this field can be measured by the amount of capital flowing into it. According to AssemblyAI statistics, voice-related AI startups received approximately $2.1 billion in venture capital in 2025; from June 2025 to May 2026, there were 36 publicly disclosed financings in the dialogue AI sector, totaling around $2.58 billion. In short, the industry is extremely hot. Specifically, ElevenLabs, which specializes in voice synthesis, raised $500 million in its Series D round and valued itself at $11 billion; PolyAI, which focuses on phone customer service, received $86 million in its Series D round; Retell AI, which handles over 40 million AI calls per month, saw a quarterly growth of more than 300%. There are also companies like Cartesia, which aim to reduce voice latency to less than 100 milliseconds to make conversations feel smoother. Large companies are also very active in this area. OpenAI’s Realtime API provides an end-to-end solution for voice processing, eliminating the need for the traditional steps of “converting to text, passing through a model, and then synthesizing sound” in one go. Google’s Gemini Live follows a similar approach, allowing for seamless responses even when a conversation is interrupted. In recent days, it also introduced Gamma4, which enhances Voice AI capabilities. It’s clear that voice is evolving from a “more natural mode of interaction” to a primary entry point for agents to obtain context. This trend is even more evident in popular workspaces like Vibe Working, where users don’t necessarily need to organize their requests into complete prompts. They can think, speak, and adjust in real-time, allowing agents to understand their intentions, complete the context, and progress tasks more naturally. In this process, ASR (Automatic Speech Recognition) is the first step; it determines whether the agent can accurately understand the real world. If the recognition is inaccurate, subsequent understanding, planning, and execution will be flawed. Only with accurate recognition can voice truly become part of the workflow. In summary: As agents begin to integrate into daily life, voice, or more specifically, Voice AI, is likely to be the first point of contact between humans and AI. While there are many ways to interact with AI—such as typing, using interfaces, or writing code to call APIs—speaking is still the most akin to natural human-to-human communication. This is why more and more people refer to voice as the “context entry” for agents, meaning the initial point for providing context. If large models are responsible for understanding the world, then voice serves as the channel through which information from the real world flows into the models. In this context, ASR plays a crucial role in the perception layer, determining whether agents can accurately capture information from the external environment and convert sounds from the real world into understandable and usable context for the models. From this perspective, the competition in ASR has elevated to a new level: it’s about competing for the next generation of interfaces that connect agents with the real world. This is also why the disclosure data for Hojo-ASR-V1 is more noteworthy. What it reflects is not just a change in the performance of a single model.

Voice AI is transitioning from being a supplementary capability to becoming an essential infrastructure, playing a increasingly crucial role in the era of Agents. 🚥 Returning to Voice AI, while most companies are focusing on general large models, the relatively quieter field of speech recognition is actually evolving rapidly. Even in areas like ASR (Automatic Speech Recognition), which has long been dominated by giants such as OpenAI, Microsoft, NVIDIA, and IBM, entrepreneurial teams are starting to achieve competitive results. The emergence of Hojo-ASR-V1 may not necessarily indicate a shift in the overall landscape, but it does at least show that Voice AI is moving from a peripheral role to a more central one, attracting more and more attention. Speech recognition is just the first step in Voice AI. The real challenge lies in whether machines can understand speech accurately and respond in a way that is close to human behavior – this is the longer, more valuable, and potentially more impactful aspect of the technology. References: [1] https://huggingface.co/HojoAI/Hojo-ASR-V1: https://huggingface.co/HojoAI/Hojo-ASR-V1 [2] https://github.com/HojoAI/Hojo-ASR: https://github.com/HojoAI/Hojo-ASR This article is from the WeChat official account “Crossing,” written by “Crossing.”

Was this helpful?

Technical SupportLive Support
侧栏
Back to Top
简体中文ZH-CNDefault繁體中文ZH-TWEnglishEN日本語JA한국어KOภาษาไทยTHTiếng ViệtVIBahasa IndonesiaIDEspañolESFrançaisFRDeutschDEРусскийRUPortuguêsPTItalianoITالعربيةAR