One family, three different starting points
HappyHorse 1.1 is not one generic video switch. QwenCloud documents three exact model IDs: happyhorse-1.1-t2v for text-to-video, happyhorse-1.1-i2v for image-to-video and happyhorse-1.1-r2v for reference-to-video. All three appear in the Token Plan Individual allowlist at verification time.
Each route produces an MP4 at 24 frames per second, supports 720P or 1080P output and accepts durations from 3 to 15 seconds. The useful difference is the source of control. T2V relies on the prompt; I2V anchors motion to one first-frame image; R2V uses up to nine reference images to guide people, characters or objects.
| Model ID | Best starting point | Documented control |
|---|---|---|
| happyhorse-1.1-t2v | A written scene | Prompt and aspect ratio |
| happyhorse-1.1-i2v | One strong opening image | First-frame image and prompt |
| happyhorse-1.1-r2v | Character or object references | One to nine reference images and prompt |
T2V: start with a scene description
Text-to-video is the cleanest option when you do not yet have approved artwork. Describe the subject, action, setting, camera movement, visual style and audio cues in one coherent prompt. Because no source image locks the composition, this route gives the model more freedom and may require more iterations to reach a precise branded result.
The official guide says happyhorse-1.1-t2v supports audio and aspect-ratio control. It is a good fit for concept exploration, social-video ideas and establishing shots. It is less suitable when a product must match an exact pack shot or a character must remain recognisably consistent with existing art.
- Use concrete motion verbs and one clear camera instruction.
- State the intended framing instead of listing conflicting angles.
- Describe audio only when it adds to the scene.
- Keep a reusable prompt skeleton so iterations remain comparable.
I2V: animate a controlled first frame
Image-to-video begins with exactly one first-frame image. The resulting video's aspect ratio follows that input, so prepare the source at the intended orientation before submitting the task. This is often the strongest route for a designed key visual, product image or approved social composition because the opening frame provides more spatial control than words alone.
The image still does not guarantee that every detail will remain fixed throughout motion. Use a clean source with a clear subject, leave visual room for the intended movement and describe how the camera and subject should behave. If you need both first- and last-frame control, QwenCloud points to a Wan 2.7 route rather than HappyHorse 1.1 I2V.
R2V: guide characters and objects with references
Reference-to-video accepts between one and nine reference images. In the documented request format, prompts can identify them as [Image 1], [Image 2] and so on. This route is designed for scenes where preserving the recognisable traits of a person, character or object matters more than matching one exact opening composition.
References should be relevant and visually legible. More images are not automatically better: inconsistent outfits, lighting or identity cues can make the instruction harder to resolve. Curate a small set that covers the needed angles, then explicitly assign each reference a role in the prompt. The model supports aspect-ratio control, 720P or 1080P output, audio and 3–15 second duration.
Generation is asynchronous and Credits are variable
HappyHorse requests use an asynchronous task flow. A create request returns a task ID; the client then checks status until the task succeeds or fails and retrieves the result. Treat that as a job workflow, not an instant chat response. Save task metadata and download successful outputs within the retention period stated by the API response or current documentation.
For Token Plan Individual, multimodal generation uses a separate API and must be integrated through a compatible tool's Skill or extension mechanism. Credits consumption is dynamic, and the HappyHorse API notes that longer videos cost more. Do not promise a fixed number of outputs from a tier until you have recorded actual Credits for your chosen route, duration, resolution and retry pattern.
The Individual subscription's interactive-use restriction still applies even though the video API itself is asynchronous.
Who this is for—and who it is not for
T2V is for rapid scene ideation from words. I2V is for creators with an approved first frame or product visual. R2V is for creators who can supply curated identity references. All three are relevant to individual users exploring short, audio-capable clips through a compatible interactive agent tool.
These models are not a substitute for an editing timeline, and the Individual plan is not permission to run a customer-facing rendering service or unattended batch pipeline. For production automation, check Alibaba's pay-as-you-go or other commercial API terms. For longer narratives, plan multiple shots, continuity review and editing outside the generation call.