Real-Time Long Video Generation (Highest Quality Variant)
Most AI video tools either top out at a few seconds or need an expensive multi-GPU cluster to go longer — and quality drifts the moment a clip runs long. Helios Base is the flagship, highest-quality member of the Helios family: it generates long-form video (up to ~60 seconds) from a text description, a starting image, or an existing clip, and it does it in real time on a single GPU at 19.5 FPS. When output quality is the priority, this is the variant to reach for. Built by PKU-YuanGroup.
What it does
Helios Base is the flagship, highest-quality variant of the Helios video generation model family, developed by PKU-YuanGroup. It generates long-form, visually rich video content from a text description, a reference image, or an existing video clip. While the distilled version of Helios prioritizes speed, Helios Base is the go-to choice when output quality is the top priority, producing the most detailed, coherent, and visually polished results in the family. It still runs on a single GPU and generates video in real time at 19.5 FPS, making it both high-quality and cost-effective compared to larger competing models.
Problem it solves
- Best-in-class output – Among the three Helios variants, Base delivers the highest visual fidelity, making it the right choice for premium creative and production use cases
- Cost-effective infrastructure – Runs on a single GPU rather than expensive multi-GPU server clusters, significantly reducing compute costs
- Flexible creative input – Works from a text description alone, a starting image, or an existing video clip, fitting a wide range of creative and production workflows
- Temporal coherence – Videos remain visually consistent and coherent throughout their full length, without the common "drift" or quality degradation seen in longer AI-generated clips
- Production-ready deployment – Designed to integrate into existing production pipelines and tools from day one
Input/Output
- Input:
- Text to Video: Describe the scene you want in plain language and Helios generates the video
- Image to Video: Provide a starting image and a description of the motion or scene
- Video to Video: Provide an existing video clip and a description to transform or extend it
- Output: A fully generated MP4 video file
- Up to ~60 seconds in length
- High visual quality with smooth, coherent motion throughout
- Parameters:
- Resolution (W × H, ratio) (dropdown, default 640 × 384 (5:3)) — pick from 11 aspect-ratio presets:
- Frame count (chunk multiples, default 99 frames (~4s @ 24fps)) — pick from a fixed list that maps to standard clip lengths at 24 fps:
- Output FPS (dropdown, default 24) — choose 16 or 24 fps
24— standard cinematic frame rate16— lighter alternative for shorter or more animation-style output- Advanced Options:
- Diffusion steps (range 20 – 60, default 50) — how many denoising passes the model runs per chunk. More steps = higher quality, slower generation; fewer steps = faster, may lose some detail. This slider is the main quality/speed trade-off for the Base variant (the Distilled variant has this collapsed into a fixed schedule)
- Random seed (-1 = random, default 1) — set a fixed integer to reproduce the same video across runs
Landscape | Square | Portrait |
768 × 320 (12:5) | 512 × 512 (1:1) | 448 × 576 (7:9) |
768 × 384 (2:1) | 512 × 768 (2:3) | |
640 × 384 (5:3) (default) | 384 × 640 (3:5) | |
768 × 512 (3:2) | 384 × 768 (1:2) | |
576 × 448 (9:7) | 320 × 768 (5:12) |
Frames | Duration @ 24 fps |
99 (default) | ~4 s |
198 | ~8 s |
297 | ~12.5 s |
396 | ~16.5 s |
495 | ~21 s |
594 | ~25 s |
693 | ~29 s |
792 | ~33 s |
891 | ~37 s |
990 | ~41 s |
1,089 | ~45 s |
1,188 | ~49.5 s |
1,287 | ~54 s |
1,386 | ~58 s |
1,452 | ~60.5 s |
Frame counts must be chunk multiples — the model generates in 33-frame autoregressive chunks (99 = 3 chunks, 198 = 6 chunks, and so on), which is why the dropdown offers only these specific values.
Accuracy & Speed
- Speed & Scale
Real-time generation speed | 19.5 frames per second on a single H100 GPU |
Maximum video length | ~ 60 seconds |
Infrastructure needed | Single GPU (no multi-GPU cluster required) |
Output quality | Highest in the Helios model family |
Helios Model Variants
Version | Best For |
Helios Base | Highest output quality - ideal for premium and production use cases |
Helios Mid | Intermediate quality/speed balance |
Helios Distilled | Fastest generation - best for high-volume or real-time use cases |
Note: Image-to-Video and Video-to-Video modes may produce slightly less consistent results than Text-to-Video, as the model was primarily trained on text-based generation.
Model Source
- HuggingFace:
Compliance & Provenance
Provider | Open-source (PKU-YuanGroup) |
Provider type | Specialized |
License | |
EU AI Act risk class | Limited Risk |
Art. 50 transparency | Required — outputs are marked. See AI Policy §2. |
Region availability | Available globally |
Training data summary | Pending — provider has not yet published per Art. 53(d) |
For more on how we classify models and mark outputs, see our AI Policy.