🧠

Multimodal Language Models - Qwen3.8-27B

🛠

Read Images and Text Together with a 262K-Context Vision-Language Model

Qwen 3.8 27B is a 27-billion-parameter vision-language model with a 262,144-token context. Give it a text prompt plus up to two images and it answers in text — describing, comparing, classifying, or pulling structured data out of what it sees. Apache 2.0.

What it does

Qwen3.8-27B takes a text prompt and, optionally, one or two images, and returns a text answer. It reasons over the words and the pictures together, so you can ask it to describe a photo, compare two product shots, pull fields out of a screenshot, or classify an image against criteria you write in plain language.

Its 262,144-token context means a long prompt, a long document pasted as text, or a long chain of prior node outputs all fit in one call.

Compared with Qwen3.6-35B-MoE, which activates only part of its network per token, Qwen3.8-27B runs every parameter on every token — steadier on long single-pass reasoning, while the MoE is faster per token. Pick by which matters for your flow.

Problem it solves

Describe & caption
Turn an image into a written description for alt text, catalogs, or search
Compare two images
Feed both image ports and ask what changed, which is better, or whether they match
Classify by your own rules
Write the criteria in the prompt instead of training a classifier
Extract structured data
Pull fields out of a screenshot or photo as JSON for the next node
Reason over long input
262K context holds a full document plus the image

Input/Output

Input

  • Text (required) — the prompt or question
  • Image 1 (optional) — reference image, up to 4096×4096
  • Image 2 (optional) — second reference image. With both connected, the model can compare, contrast, or reason over the pair together with the prompt

The upstream checkpoint also accepts video. This node exposes text and two image ports only — for video understanding use the Video Analysis (Gemini) node.

Parameters

max_new_tokens
1–8192, default 2048 (advanced). Cap on response length. Lower it for terse answers and cost; raise it for long-form generation. Generation also stops at a natural end-of-sequence regardless of this cap
temperature
0–2, default 0.7 (advanced). 0.0 is greedy and deterministic — best for factual Q&A, reading text, and code. Above 1.0 is more divergent. Qwen recommends 1.0 when thinking is on, 0.7 when it is off
top_p
0–1, default 0.8 (advanced). Nucleus sampling: keeps the smallest set of tokens whose cumulative probability reaches top_p. Lower is more focused
top_k
1–100, default 20 (advanced). Restricts candidates to the k highest-probability tokens. 50 is permissive, 20 is Qwen's own default
enable_thinking
On / Off (default, advanced). On runs an internal chain-of-thought before answering — better on multi-step reasoning and complex visual analysis, at the cost of latency and tokens. Leave it Off for short factual replies

Note on the default. Qwen ships this checkpoint with thinking on; this node defaults it off, so that ordinary prompts return quickly. If you are reproducing a benchmark or comparing against upstream results, switch it on — and raise temperature to 1.0 with it.

Output:

  • Text — the generated response

Choosing a multimodal model

What you need
Use
Steady long single-pass reasoning, 262K context
Qwen3.8-27B
Flagship quality with faster per-token inference
Qwen3.6-35B-MoE
A compact model for light, high-volume workloads
Gemma 4 E2B
A long-context flagship alternative
Gemma 4 31B
Graded reasoning depth, latest open-weights preview
Agnes-3.0-Flash (Preview)

For pure text extraction from documents, an OCR model (GLM-OCR, PaddleOCR, Sarashina2.2-OCR, DeepSeek OCR) is cheaper and more accurate than a VLM.

Technical Details

Architecture
Causal language model with vision encoder. 64 layers mixing Gated DeltaNet and Gated Attention, hidden size 5,120
Parameters
27B
Context window
262,144 tokens natively (extensible upstream to ~1M with YaRN RoPE scaling; this node runs the native window)
Precision
bf16
Runtime
GPU, cloud — no local setup

Compliance & Provenance

Provider
Alibaba (Qwen)
Provider type
Open-source
License
Apache 2.0
EU AI Act risk class
Limited Risk
Art. 50 transparency
Applicable — multimodal LLM output carries the AI-generated label
Region availability
Available globally
Training data summary
Pending — provider has not yet published per Art. 53(d)

For more on how we classify models and mark outputs, see our AI Policy.

Model Source

  • Model card
  • License: Apache 2.0