Voiceover
Generate crisp narration for product demos, explainers, and support flows.
Audio models
Explore Qwen Audio 3.0, LiveTranslate, Fun ASR, CosyVoice, music, and Qwen Omni routes for synthesis, transcription, translation, and realtime audio workflows.
Audio workflows
Example-only previews show the workflows these Alibaba routes can support.
Generate crisp narration for product demos, explainers, and support flows.
Create lightweight launch loops, interface ambience, and campaign sound beds.
Turn calls, interviews, and demo narration into searchable text workflows.
10 routes
Speech, music, and multimodal audio routes with settings, estimates, and Alibaba Cloud access paths.
Alibaba Cloud
CosyVoice V3.5 Plus is Alibaba's text-to-speech route for generated voice output. Nfero presents it as part of the audio stack alongside ASR, music, and Qwen Omni models.
Best for: Speech, transcription, and audio workflows
Alibaba Cloud
Fun ASR is the Alibaba speech-recognition route for turning audio into text. Nfero connects it to transcription, audio understanding, RAG, and agent workflow planning.
Best for: Speech, transcription, and audio workflows
Alibaba Cloud
Fun ASR Flash 2026-06-15 is Alibaba's current non-realtime recognition snapshot for multilingual audio files. Nfero records the exact model ID and official access route for implementation planning.
Best for: Speech, transcription, and audio workflows
Alibaba Cloud
Fun ASR Realtime is Alibaba's recommended route for continuous multilingual speech recognition and supported dialects. Nfero documents the planning path while public transcription remains disabled.
Best for: Speech, transcription, and audio workflows
Alibaba Cloud
Fun Music V1 is the Alibaba text-to-music route in the Nfero audio catalog. It is positioned for creative exploration, music ideation, and audio workflow planning.
Best for: Speech, transcription, and audio workflows
Alibaba Cloud
Qwen Audio 3.0 TTS Flash is the low-latency synthesis route for interactive speech and voice-cloning workflows. Alibaba reports time to first audio below 200 ms under supported conditions.
Best for: Speech, transcription, and audio workflows
Alibaba Cloud
Qwen Audio 3.0 TTS Plus is the high-quality speech synthesis route for professional voice output and instruction-controlled delivery. Nfero documents its international access and list-price basis without exposing unverified public synthesis.
Best for: Speech, transcription, and audio workflows
Alibaba Cloud
Qwen3.5 LiveTranslate Flash Realtime combines audio and visual context for multilingual translation, with latency as low as 2.8 seconds under supported conditions. Nfero documents the route without claiming public access.
Best for: Speech, transcription, and audio workflows
Alibaba Cloud
Qwen3.5 Omni Plus is Alibaba's multimodal route for combining text, image, audio, and video understanding or generation. Nfero places it in the broader Omni and agent-planning surface.
Best for: Speech, transcription, and audio workflows
Alibaba Cloud
Qwen3.5 Omni Plus Realtime is the low-latency Omni route for interactive multimodal experiences. Nfero uses it for planning realtime, audio, and agent-facing Alibaba workflows.
Best for: Speech, transcription, and audio workflows
Open-source proof
Qwen3-TTS, Qwen3-ASR, and Qwen3-Omni make Alibaba audio and multimodal workflows easier for buyers to understand.