Read Images and Text Together with a 262K-Context Vision-Language Model
Qwen 3.8 27B is a 27-billion-parameter vision-language model with a 262,144-token context. Give it a text prompt plus up to two images and it answers in text — describing, comparing, classifying, or pulling structured data out of what it sees. Apache 2.0.
What it does
Qwen3.8-27B takes a text prompt and, optionally, one or two images, and returns a text answer. It reasons over the words and the pictures together, so you can ask it to describe a photo, compare two product shots, pull fields out of a screenshot, or classify an image against criteria you write in plain language.
Its 262,144-token context means a long prompt, a long document pasted as text, or a long chain of prior node outputs all fit in one call.
Compared with Qwen3.6-35B-MoE, which activates only part of its network per token, Qwen3.8-27B runs every parameter on every token — steadier on long single-pass reasoning, while the MoE is faster per token. Pick by which matters for your flow.
Problem it solves
Describe & caption | Turn an image into a written description for alt text, catalogs, or search |
Compare two images | Feed both image ports and ask what changed, which is better, or whether they match |
Classify by your own rules | Write the criteria in the prompt instead of training a classifier |
Extract structured data | Pull fields out of a screenshot or photo as JSON for the next node |
Reason over long input | 262K context holds a full document plus the image |
Input/Output
Input
- Text (required) — the prompt or question
- Image 1 (optional) — reference image, up to 4096×4096
- Image 2 (optional) — second reference image. With both connected, the model can compare, contrast, or reason over the pair together with the prompt
The upstream checkpoint also accepts video. This node exposes text and two image ports only — for video understanding use the Video Analysis (Gemini) node.
Parameters
max_new_tokens | 1–8192, default 2048 (advanced). Cap on response length. Lower it for terse answers and cost; raise it for long-form generation. Generation also stops at a natural end-of-sequence regardless of this cap |
temperature | 0–2, default 0.7 (advanced). 0.0 is greedy and deterministic — best for factual Q&A, reading text, and code. Above 1.0 is more divergent. Qwen recommends 1.0 when thinking is on, 0.7 when it is off |
top_p | 0–1, default 0.8 (advanced). Nucleus sampling: keeps the smallest set of tokens whose cumulative probability reaches top_p. Lower is more focused |
top_k | 1–100, default 20 (advanced). Restricts candidates to the k highest-probability tokens. 50 is permissive, 20 is Qwen's own default |
enable_thinking | On / Off (default, advanced). On runs an internal chain-of-thought before answering — better on multi-step reasoning and complex visual analysis, at the cost of latency and tokens. Leave it Off for short factual replies |
Note on the default. Qwen ships this checkpoint with thinking on; this node defaults it off, so that ordinary prompts return quickly. If you are reproducing a benchmark or comparing against upstream results, switch it on — and raise temperature to 1.0 with it.
Output:
- Text — the generated response
Choosing a multimodal model
What you need | Use |
Steady long single-pass reasoning, 262K context | Qwen3.8-27B |
Flagship quality with faster per-token inference | Qwen3.6-35B-MoE |
A compact model for light, high-volume workloads | Gemma 4 E2B |
A long-context flagship alternative | Gemma 4 31B |
Graded reasoning depth, latest open-weights preview | Agnes-3.0-Flash (Preview) |
For pure text extraction from documents, an OCR model (GLM-OCR, PaddleOCR, Sarashina2.2-OCR, DeepSeek OCR) is cheaper and more accurate than a VLM.
Technical Details
Architecture | Causal language model with vision encoder. 64 layers mixing Gated DeltaNet and Gated Attention, hidden size 5,120 |
Parameters | 27B |
Context window | 262,144 tokens natively (extensible upstream to ~1M with YaRN RoPE scaling; this node runs the native window) |
Precision | bf16 |
Runtime | GPU, cloud — no local setup |
Compliance & Provenance
Provider | Alibaba (Qwen) |
Provider type | Open-source |
License | Apache 2.0 |
EU AI Act risk class | Limited Risk |
Art. 50 transparency | Applicable — multimodal LLM output carries the AI-generated label |
Region availability | Available globally |
Training data summary | Pending — provider has not yet published per Art. 53(d) |
For more on how we classify models and mark outputs, see our AI Policy.
Model Source
- Model card
- License: Apache 2.0