NetEase Youdao launched its first AI-native voice agent on Monday, named NetEase Bage Shuo, designed to bypass traditional transcription tools and convert natural spoken language directly into structured, publication-ready text.
The tool addresses a growing industry bottleneck where traditional keyboard input speeds struggle to keep pace with human thought in the generative AI era. While public Google data indicates voice input averages 2.5 times the speed of typing, NetEase Youdao reports that Bage Shuo achieves up to five times the efficiency of typing through advanced automatic speech recognition and semantic understanding pipelines.
Unlike conventional dictation software that records every hesitation, stumble, and redundant phrase verbatim, Bage Shuo focuses on intent recognition. For example, if a user dictates a revised schedule or a correction mid-sentence, the agent filters out errors and structures the final output for immediate use in emails, instant messaging, or document creation.
The underlying technology relies on Youdao’s proprietary framework for hearing clearly, understanding well, and writing well. Powered by the Confucius4-R2T2 streaming speech recognition model, the system handles background noise, low-volume speech, and contextual hotwords with high accuracy, while a separate polishing model organizes punctuation, removes redundancies, and normalizes grammar.
In addition to document creation, the platform supports real-time translation across 124 languages with up to 98 percent accuracy in specialized fields such as economics, physics, and medicine. Addressing privacy concerns associated with voice data, developers noted that the cloud system stores no audio or text data, returning data sovereignty directly to the user.
NetEase Bage Shuo is already available for download starting and will be provided free for life to lower the barrier to entry for AI-native voice interactions.


