Privacy settings

Privacy and GDPR notice

Nfero uses essential storage for language preferences and demo access controls. With your consent, Google Analytics and lightweight usage events help improve model demos and outbound subscription routing. Live demo prompts may be sent to configured Alibaba-hosted model endpoints.

Audio workflows

Voice, music, and speech understanding.

Example-only previews show the workflows these Alibaba routes can support.

Text to speech

Voiceover

Generate crisp narration for product demos, explainers, and support flows.

Text to music

Music bed

Create lightweight launch loops, interface ambience, and campaign sound beds.

Speech to text

Transcription

Turn calls, interviews, and demo narration into searchable text workflows.

10 routes

Audio model catalog

Speech, music, and multimodal audio routes with settings, estimates, and Alibaba Cloud access paths.

Alibaba Cloud

CosyVoice V3.5 Plus

Demo preview

CosyVoice V3.5 Plus is Alibaba's text-to-speech route for generated voice output. Nfero presents it as part of the audio stack alongside ASR, music, and Qwen Omni models.

Best for: Speech, transcription, and audio workflows

CosyVoiceText to speechTerms apply
Availability
Guide and examples
Estimated standard demo
Confirm with provider

Alibaba Cloud

Fun ASR

Demo preview

Fun ASR is the Alibaba speech-recognition route for turning audio into text. Nfero connects it to transcription, audio understanding, RAG, and agent workflow planning.

Best for: Speech, transcription, and audio workflows

FunSpeech to textTerms apply
Availability
Guide and examples
Estimated standard demo
Confirm with provider
Demo preview

Fun ASR Flash 2026-06-15 is Alibaba's current non-realtime recognition snapshot for multilingual audio files. Nfero records the exact model ID and official access route for implementation planning.

Best for: Speech, transcription, and audio workflows

Fun ASRSpeech to textTerms apply
Availability
Guide and examples
Estimated standard demo
Confirm with provider

Alibaba Cloud

Fun ASR Realtime

Demo preview

Fun ASR Realtime is Alibaba's recommended route for continuous multilingual speech recognition and supported dialects. Nfero documents the planning path while public transcription remains disabled.

Best for: Speech, transcription, and audio workflows

Fun ASRSpeech to textTerms apply
Availability
Guide and examples
Estimated standard demo
Confirm with provider

Alibaba Cloud

Fun Music V1

Demo preview

Fun Music V1 is the Alibaba text-to-music route in the Nfero audio catalog. It is positioned for creative exploration, music ideation, and audio workflow planning.

Best for: Speech, transcription, and audio workflows

FunAudio generationTerms apply
Availability
Guide and examples
Estimated standard demo
Confirm with provider
Demo preview

Qwen Audio 3.0 TTS Flash is the low-latency synthesis route for interactive speech and voice-cloning workflows. Alibaba reports time to first audio below 200 ms under supported conditions.

Best for: Speech, transcription, and audio workflows

Qwen AudioText to speechTerms apply
Availability
Guide and examples
Estimated standard demo
Confirm with provider
Demo preview

Qwen Audio 3.0 TTS Plus is the high-quality speech synthesis route for professional voice output and instruction-controlled delivery. Nfero documents its international access and list-price basis without exposing unverified public synthesis.

Best for: Speech, transcription, and audio workflows

Qwen AudioText to speechTerms apply
Availability
Guide and examples
Estimated standard demo
Confirm with provider
Demo preview

Qwen3.5 LiveTranslate Flash Realtime combines audio and visual context for multilingual translation, with latency as low as 2.8 seconds under supported conditions. Nfero documents the route without claiming public access.

Best for: Speech, transcription, and audio workflows

QwenSpeech to speechTerms apply
Availability
Guide and examples
Estimated standard demo
Confirm with provider

Alibaba Cloud

Qwen3.5 Omni Plus

Demo preview

Qwen3.5 Omni Plus is Alibaba's multimodal route for combining text, image, audio, and video understanding or generation. Nfero places it in the broader Omni and agent-planning surface.

Best for: Speech, transcription, and audio workflows

QwenOmniTerms apply
Availability
Guide and examples
Estimated standard demo
Confirm with provider
Demo preview

Qwen3.5 Omni Plus Realtime is the low-latency Omni route for interactive multimodal experiences. Nfero uses it for planning realtime, audio, and agent-facing Alibaba workflows.

Best for: Speech, transcription, and audio workflows

QwenOmniTerms apply
Availability
Guide and examples
Estimated standard demo
Confirm with provider

Open-source proof

Connect Qwen speech projects to audio demos.

Qwen3-TTS, Qwen3-ASR, and Qwen3-Omni make Alibaba audio and multimodal workflows easier for buyers to understand.

Open Source