WeChat grayscale tests hold-to-talk voice-to-text in input box; five voice input entries now available
WeChat is grayscale testing a new voice input feature that lets users hold the input box to convert speech to text, adding to existing voice input methods. The trend is also seen in other apps like WeChat Mac and Feishu. Advances in voice recognition technology and falling costs are driving wider adoption of voice input, which also helps apps bypass third-party input methods.
WeChat's recently grayscale-tested voice input feature has drawn attention. In text input mode, the input box now shows a prompt saying "hold to convert to text," allowing users to press and hold the input box to speak and have the speech directly converted to text and sent, without switching to voice mode. Additionally, a new microphone button has been added to the right of the input box, which also supports voice-to-text when tapped. Combined with the existing "hold to talk" function in voice mode, as well as the dictation features built into input methods and the system, WeChat's lower section now offers five voice input entries in total.
This trend is not unique to WeChat. WeChat Mac desktop version prompts users to hold the Fn key for voice input, and Feishu's desktop input box similarly notes "hold Fn to talk," with voice-to-text also available in Feishu documents and spreadsheets. Currently, multiple mainstream applications have placed voice input entries in prominent positions, with some using strong text reminders or enlarged buttons to guide users.
Voice recognition technology is not new, as voice assistants have supported dictation for over a decade. In recent years, with the development of large voice models, the accuracy and response speed of voice recognition have improved significantly. For example, the streaming voice recognition service from Volcano Engine's Doubao is priced at less than 1 RMB per hour at retail, and major companies using self-developed models face even lower costs. Technological maturity and falling costs have provided the conditions for widespread promotion of voice input in applications.
The dense appearance of voice input entries is related to the special position of input methods in the application ecosystem. Input methods hold user input data, and historically there have been cases where input methods guided users to their own search products through candidate words. The voice-to-text function allows applications to interact directly with users, bypassing input methods. Currently, some AI-native applications can already complete operations such as ticket booking directly through voice commands, and mainstream applications are gradually cultivating users' voice input habits in preparation for further feature expansion.
Why this event matters
The event has a measured impact on 4 industrys. The strongest current signal is positive for Artificial Intelligence, with intensity 70/100 and 80% confidence over a short term horizon.
Artificial Intelligence
- Direction
- positive
- Intensity
- 70
- Confidence
- 80%
- Horizon
- Short term
Diversified Internet Platforms
- Direction
- positive
- Intensity
- 65
- Confidence
- 80%
- Horizon
- Short term
Cloud Services & Data Centres
- Direction
- positive
- Intensity
- 60
- Confidence
- 75%
- Horizon
- Short term
Financial Technology
- Direction
- positive
- Intensity
- 50
- Confidence
- 70%
- Horizon
- Medium term
Impact figures are analytical estimates that combine direction, intensity, confidence and event importance. They are not investment advice.