Preview Checkpoint — 33B Hybrid-Attention VLM with a Reasoning Dial
Agnes 3.0 Flash Preview is a 33-billion-parameter hybrid-attention vision-language model with a 262,144-token context. Text prompt plus up to two images in, text out. A reasoning effort dial lets you trade latency for depth. Apache 2.0.
⚠️ This is the open-weights preview checkpoint, distinct from the production Agnes 3.0 Flash (which runs a 1M-token context). Behavior and availability may change.
What it does
Agnes-3.0-Flash takes a text prompt and, optionally, one or two images, and returns a text answer — description, comparison, classification, or structured extraction.
What sets it apart from the other multimodal models here is the reasoning-effort dial. Most models in this category are either always thinking or never. Agnes gives you four settings, so you can keep it instant for simple replies and turn thinking up only where it pays.
Its architecture is unusual too: 72 layers alternating delta-rule recurrent and global attention blocks in a 3:1 ratio, which is what lets a 33B model hold a 262K context without the memory cost of full attention throughout.
Language coverage. Agnes AI tags this checkpoint for English and Chinese. Other languages may work but are not claimed — if your flow runs Korean or Japanese prompts, test before committing, or use Qwen3.6-35B-MoE.
About "Preview"
This is the open-weights preview checkpoint, not the production Agnes 3.0 Flash. Treat it as something to evaluate rather than something to build a production flow on: quality, speed, and availability can change without a release note. If you need stability today, use Qwen3.6-35B-MoE or Gemma 4 31B.
Problem it solves
Tune thinking per task | Instant replies for simple lookups, deep reasoning where it matters — one parameter |
Long inputs | 262K context holds a long document plus images in a single call |
Describe & caption | Turn an image into written description for alt text, catalogs, or search |
Compare two images | Feed both image ports and ask what changed or which is better |
Evaluate a new model | Compare against your current VLM node before committing |
Input/Output
Input
- Text (required) — the prompt or question
- Image 1 (optional) — reference image, up to 4096×4096
- Image 2 (optional) — second reference image, for comparison or joint reasoning
Parameters
max_new_tokens | 1–8192, default 2048 (advanced). Cap on response length. Raise this whenever you raise reasoning_effort — see the warning below |
temperature | 0–2, default 0.7 (advanced). 0.0 is greedy and deterministic. Above 1.0 is more divergent |
top_p | 0–1, default 0.8 (advanced). Nucleus sampling threshold. Lower is more focused |
top_k | 1–100, default 20 (advanced). Restricts candidates to the k highest-probability tokens |
reasoning_effort | Off (default) / Low / Medium / XHigh (advanced). How long the model thinks before answering |
reasoning_effort — read this before turning it on
Off is dramatically faster. Measured on the same prompt: 0.7s at Off against 6.9s at Low — roughly ten times. Off is the right setting for ordinary replies.
The three levels above Off run an internal chain-of-thought first. That costs latency and tokens, and the tokens come out of max_new_tokens: at a small cap the model is still thinking when generation stops and you get no answer at all. If you raise reasoning_effort, raise max_new_tokens with it.
Output:
- Text — the generated response
Choosing a multimodal model
What you need | Use |
Graded reasoning depth, evaluating a new checkpoint | Agnes-3.0-Flash (Preview) |
Production stability, flagship quality, fast per token | Qwen3.6-35B-MoE |
Steady dense-model reasoning, 262K context | Qwen3.8-27B |
Compact model for light, high-volume workloads | Gemma 4 E2B |
Long-context flagship alternative | Gemma 4 31B |
Korean or Japanese prompts | Qwen3.6-35B-MoE or Gemma 4 31B — Agnes claims English and Chinese only |
Technical Details
Architecture | Hybrid-attention decoder — 72 layers, 54 delta-rule recurrent + 18 global attention in 3:1 alternation. Hidden size 5,120, 3-axis rotary positioning, 27-layer vision tower |
Parameters | 33B |
Context window | 262,144 tokens (the production Agnes 3.0 Flash runs 1M; this preview does not) |
Precision | bf16 |
Languages | English, Chinese |
Checkpoint | Open-weights preview, distinct from the production Agnes 3.0 Flash |
Runtime | GPU, cloud — no local setup |
Compliance & Provenance
Provider | Agnes AI |
Provider type | Open-weights |
License | Apache 2.0 |
EU AI Act risk class | Limited Risk |
Art. 50 transparency | Applicable — multimodal LLM output carries the AI-generated label |
Region availability | Available globally |
Training data summary | Pending — provider has not yet published per Art. 53(d) |
For more on how we classify models and mark outputs, see our AI Policy.
Model Source
- Model card
- License: Apache 2.0