Real-Time Long Video Generation (Highest Quality Variant)
Most AI video tools either top out at a few seconds or need an expensive multi-GPU cluster to go longer — and quality drifts the moment a clip runs long. Helios-Base is the flagship, highest-quality member of the Helios family: it generates long-form video (up to ~45 seconds) from a text description, a starting image, or an existing clip, and it does it in real time on a single GPU at 19.5 FPS. When output quality is the priority, this is the variant to reach for. Built by PKU-YuanGroup.
What it does
Helios-Base is the flagship, highest-quality variant of the Helios video generation model family, developed by PKU-YuanGroup. It generates long-form, visually rich video content from a text description, a reference image, or an existing video clip. While the distilled version of Helios prioritizes speed, Helios-Base is the go-to choice when output quality is the top priority, producing the most detailed, coherent, and visually polished results in the family. It still runs on a single GPU and generates video in real time at 19.5 FPS, making it both high-quality and cost-effective compared to larger competing models.
Problem it solves
- Best-in-class output – Among the three Helios variants, Base delivers the highest visual fidelity, making it the right choice for premium creative and production use cases
- Cost-effective infrastructure – Runs on a single GPU rather than expensive multi-GPU server clusters, significantly reducing compute costs
- Flexible creative input – Works from a text description alone, a starting image, or an existing video clip, fitting a wide range of creative and production workflows
- Temporal coherence – Videos remain visually consistent and coherent throughout their full length, without the common "drift" or quality degradation seen in longer AI-generated clips
- Production-ready deployment – Designed to integrate into existing production pipelines and tools from day one
Input/Output
- Input:
- Text to Video: Describe the scene you want in plain language and Helios generates the video
- Image to Video: Provide a starting image and a description of the motion or scene
- Video to Video: Provide an existing video clip and a description to transform or extend it
- Output: A fully generated MP4 video file
- Up to ~45 seconds in length
- High visual quality with smooth, coherent motion throughout
Accuracy & Speed
- Speed & Scale
Real-time generation speed | 19.5 frames per second on a single H100 GPU |
Maximum video length | ~ 45 seconds |
Infrastructure needed | Single GPU (no multi-GPU cluster required) |
Output quality | Highest in the Helios model family |
Helios Model Variants
Version | Best For |
Helios-Base | Highest output quality - ideal for premium and production use cases |
Helios-Mid | Intermediate quality/speed balance |
Helios-Distilled | Fastest generation - best for high-volume or real-time use cases |
Note: Image-to-Video and Video-to-Video modes may produce slightly less consistent results than Text-to-Video, as the model was primarily trained on text-based generation.
Model Source
- HuggingFace:
- License: apache-2.0
Compliance & Provenance
Provider | Open-source (PKU-YuanGroup) |
Provider type | Specialized |
License | Apache 2.0 |
EU AI Act risk class | Limited Risk |
Art. 50 transparency | Required — outputs are marked. See AI Policy §2. |
Region availability | Available globally |
Training data summary | Pending — provider has not yet published per Art. 53(d) |
For more on how we classify models and mark outputs, see our AI Policy.