Release Note
Release Note

Release Note

Version [1.1.7]

Version [1.1.6]

Version [1.1.5]

Version [1.1.4]

Version [1.1.3]

Version [1.1.2]

Version [1.1.1]

Version [1.1.0]

Release Overview [1.1.7]

CNAPS Studio v1.1.7 is our biggest creator-toolkit release to date, with four major additions:

  • Ultimate Privacy Protection — our first single-purpose web app, with one-click anonymization of faces, license plates, tattoos, and people in any image
  • Speech-to-Text lands as a first-class category — Whisper Large V3 powering transcription and subtitle workflows across the studio's entire video toolkit
  • Long-Form → Short-Form Video Pipelinenine new tools that cut, tighten, reformat, subtitle, and cold-open highlight clips in a single flow, starting from v1.1.6's Video Analysis
  • Easy UI updates — new My Results page, per-card API connection status, and per-card parameter menus

Two more new models: RF-DETR Keypoint (real-time human pose estimation) and QWEN-Inpaint (mask-based image editing with pixel-identical backgrounds). Polish: 16:9 / 9:16 output on four image generators, cleaner BiRefNet cutouts, and 1.5–2× faster model loading.

✨ What's New

Major Improvements

1️⃣ Long-Form to Short-Form Video Production, End-to-End

CNAPS Studio v1.1.7 ships a full long-form → short-form video pipeline — the "podcast episode / interview / gameplay footage → publishable TikTok / Reels / Shorts" workflow, all inside a single flow. It starts with v1.1.6's Video Analysis node (Gemini identifies highlight segments) and continues through nine new tools that cut, tighten, reformat, subtitle, and cold-open each clip so it's ready to publish. Every stage runs in batch — one long video in, a batch of publishable short clips out.

The pipeline in order:

  • 1. Analyze — Video Analysis (v1.1.6) produces a JSON of highlight segments with titles, hooks, and scores
  • 2. CutVideo Trim turns the JSON into actual video clips; Video List Extract picks one clip when you need to work single-file
  • 3. Tighten pacingRemove Dead Zone cuts silent gaps for snappy YouTube-style pacing
  • 4. ReformatVideo Reframe (9:16) turns horizontal clips into 1080×1920 vertical for social platforms (four strategies including Subject-Aware face tracking)
  • 5. Add text overlaysVideo Caption for static Title/Hook overlays; Video Subtitle (Spoken) for karaoke-style word-synced subtitles
  • 6. Cold-openHook Teaser prepends a visual teaser of each clip's hook moment; Video Narrate (Qwen3) adds an AI-voiced spoken cold-open; Narration Mux (Cold Open) lets you bring your own audio

2️⃣ Speech-to-Text Lands as a First-Class Category

CNAPS Studio now has Whisper Large V3 — one of the strongest open-source speech models available — as native audio understanding, with 99-language support and explicit picker options for English, Korean, Japanese, and Chinese. Two nodes expose it:

  • Speech Recognition (Video) (shipped v1.1.6) — one-shot transcription of a single video
  • Speech Recognition (Clips) (new in v1.1.7) — batch transcription across a video list, with per-clip timed transcripts that feed directly into Video Subtitle (Spoken) for karaoke-style captions

3️⃣ Ultimate Privacy Protection — Our First Single-Purpose Web App

Beyond CNAPS Studio, we're also spinning up focused single-purpose web apps for tasks that don't need a full workflow builder. The first one, Ultimate Privacy Protection, automatically detects and anonymizes sensitive content in any image.

How it works:

  • Upload an image — PNG, JPEG, WebP, AVIF, or BMP (up to 50 MB)
  • Pick what to hideFace, Tattoo, License Plate, Person, or a Custom target you describe in your own words
  • Get an anonymized image back — sensitive regions are detected and blurred in a single pass; the result appears right on the same page

What's shipping today: images only, with 3 free images on sign-in (no account needed). Video privacy protection is coming next.

This is the start of a broader strategy: for the tasks people do most often — where assembling a workflow would be overkill — ship a dedicated web app that does one thing and does it well.

image

4️⃣ Easy UI Updates

  • My Results page: a new page (accessible via the My results button at the top of the Easy UI landing) that lets you browse every previous output you've generated through Easy UI.
  • API connection status shown per card: cards that require a commercial AI API key now show a live connection indicator directly on the card.
  • Per-card parameter menus: each card now exposes a parameter setup menu so you can tune the run without opening the full workflow canvas.

New AI Models & Tools

1️⃣ AI Models - Audio Understanding - Speech Recognition (Clips)

The list-friendly sibling to v1.1.6's Speech Recognition (Video). Feed it a list of video clips instead of a single video, and it returns per-clip timed transcripts — one transcript entry per clip, with timing info tied back to the source. Same Whisper Large V3 model under the hood, same 99-language support and same picker (Auto, English, Korean, Japanese, Chinese).

2️⃣ AI Model - Pose Estimation - RF-DETR Keypoint (Real-Time Human Pose Estimation)

The third sibling in the RF-DETR family (joining RF-DETR Medium and RF-DETR Seg Medium from v1.1.5). Where Medium draws boxes and Seg Medium draws outlines, Keypoint finds the joints — 17 standard body landmarks (nose, eyes, ears, shoulders, elbows, wrists, hips, knees, ankles) for every person in a photo — and renders them as an OpenPose-format skeleton ready to feed into pose-guided workflows.

image

3️⃣ AI Model - Image Edit - QWEN-Inpaint

Paint a mask over the part of the image you want to change, describe the change in text, and the model only edits inside that region. Everything outside the mask stays pixel-identical to the original — no more "just change the shirt color" accidentally shifting the face, lighting, or background

4️⃣ Tool - Video Trim and Video List Extract - Cut a Long Video Into Clips, Then Pick the One You Want

Video Trim takes a source video plus the highlight-segments JSON from v1.1.6's Video Analysis and produces a list of video clips — one per highlight the LLM identified. Pre/post padding sliders keep clip boundaries natural (no cut-off words or reactions), and a minimum-duration filter drops LLM segmentation noise. Video List Extract is the video equivalent of v1.1.6's List Extract: pick a single video out of a list by 0-based index, either for individual processing (e.g., feeding into Narration Mux) or for fan-out per-clip branches.

5️⃣ Tool - Remove Dead Zone - Automatic Jump-Cut Editor

Cuts silent gaps out of every clip in a video list to make pacing feel snappy — the same jump-cut technique YouTubers and podcast editors use to make talking-head content watchable. Three sliders: minimum silence length to trigger a cut (default 0.5s), padding around speech so words aren't clipped, and (Advanced) the loudness threshold that defines what counts as silence. Snappy YouTube pacing with default settings; loosen for lecture content, tighten for TikTok.

6️⃣ Tool - Video Reframe (9:16) - Horizontal → Vertical for Social Platforms

Reformat every clip in a video list to 9:16 vertical (1080×1920) — the aspect ratio TikTok, Reels, Shorts, and Facebook Reels demand. Four strategies: Fit + Blur (whole frame preserved with blurred background bands — the polished podcast-clip look), Fit + Black (traditional letterbox), Crop to Fill (immersive full-vertical), and Subject-Aware (crops to fill but tracks detected faces so people don't drift out of frame — the smart default for talking-head content).

7️⃣ Tool - Video Caption and Video Subtitle (Spoken) - Static Overlays and Karaoke Subtitles

Two overlay tools that stack cleanly on the same clips. Video Caption stamps a static Title or Hook text from the clip's Video Analysis metadata onto the frame — great for title-card intros and "wait for it…" hook overlays. Video Subtitle (Spoken) burns karaoke-style word-synced subtitles driven by a word-timed transcript — each word highlights as it's spoken, the on-screen caption style TikTok/Reels/Shorts creators use to hold attention in silent autoplay. Both offer six named look presets (Minimal, Vibrant, Bold, Newsletter, Clean, Custom) and CJK font support (Noto Sans CJK) for Korean/Japanese/Chinese content. Stack Video Caption at Top and Video Subtitle at Bottom for full text coverage.

8️⃣ Tool - Hook Teaser, Video Narrate (Qwen3), and Narration Mux - Three Ways to Build a Cold-Open

Three complementary tools for the "grab attention in the first 3 seconds" cold-open pattern, each attacking it a different way:

  • Hook Teaser (Cold Open) — visual. Cuts a short preview from each clip's own hook moment (identified by Video Analysis) and prepends it to the front of the clip. Default 3-second teaser (tunable 1–10s). The classic cold-open trick used by TV shows and YouTube creators.
  • Video Narrate (Qwen3) — AI-voiced. Reads each clip's Hook or Title text using Qwen3 TTS and prepends the synthesized voiceover to the front of the clip, playing over a frozen first frame. Auto / Korean / English language handling; consistent narrator voice across a batch.
  • Narration Mux (Cold Open) — bring your own audio. Single-clip, zero-config utility that combines a video with any audio track (recorded voiceover, music sting, sound effect, external TTS) as a frozen-frame cold-open. The audio length determines the intro length.

Layer them together for a full trailer effect: Video Narrate → Hook Teaser → Video Caption on the resulting clip.

Platform Improvements

1️⃣ 16:9 and 9:16 resolution options for image generation

FluxSchnell, Flux2-Klein-4B, OmniGen2, and ZImageTurbo now support 1024×576 (16:9) and 576×1024 (9:16) output resolutions

2️⃣ Custom font upload for Video Caption and Video Subtitle

Upload your own .ttf or .otf font file via the new Font Loader node and connect it directly to Video Caption or Video Subtitle. Previously limited to built-in fonts (Noto Sans CJK, DejaVu, Liberation, Ubuntu)

3️⃣ BiRefNet background removal now supports cleaner cutouts

A new "Extract Foreground" option removes color bleeding at edges (e.g. green fringe from grass backgrounds on hair), producing a clean RGBA cutout directly — no need for the Masker + Multiply chain

image

4️⃣ Performance: Significantly faster model loading

Cold-start loading times are now 1.5–2x faster across most large models (GLMImage, Cosmos3-Nano, QWEN, Wan2.2, OmniGen2, ZImageTurbo, and more), following an optimization to how model weights are transferred to GPU.

🔧 Bug Fixes

  • List Extract — Black Background Preserved on Decomposed Layers

Fixed a bug in List Extract where extracting a layer with a black background (from a QWEN-Layered decomposition, or any other list source with pure-black backgrounds) would replace the black with a tan/brown fill in the extracted image. The extracted layer now preserves its original background color, including pure black.

Fixed a bug where an unintended "AI-Generated" label was being added to the bottom-right corner of output images from List Extract. The watermark is no longer applied

  • Video upscalers now preserve original audio

FlashVSR, SparkVSR, and SeedVR2 no longer strip audio from the input video. The original audio track is now carried through to the upscaled output.

Release Overview [1.1.6]

CNAPS Studio v1.1.6 opens the full AI model library on the free tier — every model is now free to try, no subscription needed. Easy UI picks up mobile support, so you can browse cards, run flows, and view results from your phone. This release also completes the frame-by-frame video pipeline: Video Split and Video Reassemble bookend the loop, List Extract / List Inject enable extract → process → inject on image lists, Layer Compose flattens layered stacks Photoshop-style, and new model Qwen-Image-Layered auto-decomposes any flat photo into RGBA layers. Two new video-to-text capabilities join the lineup, Video Analysis (Gemini API — transcripts, timecodes, highlights) and Speech Recognition (Whisper Large V3 spoken-audio transcripts) — alongside SANA-Sprint 1.6B, an ultra-fast text-to-image model that produces 1024px images in ~0.1s. Video upscaling now reaches 4K.

✨ What's New

Major Improvements

1️⃣ All AI Models Now on Free Tier

The entire lineup of AI models in CNAPS Studio is now accessible on the free tier - no subscription needed to try any model. Explore the full library before deciding what fits your workflow.

2️⃣ Easy UI Now Runs on Mobile

The card-based Easy UI is now mobile-friendly. The layout has been reworked so you can browse the card gallery, drop in inputs, hit Run, and view results from your phone. Especially useful for the task-focused cards like Identity Redacted, Virtual Try-On, and Beyond Upscaling, where you'd often want to use a photo straight from your phone.

New AI Models & Tools

1️⃣ AI Models - Image Edit - Qwen-Image-Layered

Takes any flat image and automatically separates it into layers with transparency (RGBA), so you can pull individual elements — foreground, background, specific subjects — out of a photo without manual masking. Ideal for compositing, background swap, and design workflows where each element needs to be its own editable layer.

2️⃣ AI Models - Audio Understanding - Speech Recognition (Video)

A new speech recognition node powered by OpenAI's Whisper Large V3, one of the strongest open-source speech models available. Feed it a video and the node extracts the audio track and returns a full text transcript of everything spoken. Language auto-detects by default, with explicit picker options for English, Korean, Japanese, and Chinese (the underlying model supports 99 languages).

3️⃣ AI Models - Image Generation - SANA-Sprint 1.6B

An NVIDIA’s new text-to-image model built for speed: SANA-Sprint generates a 1024×1024 image in just 1 to 4 diffusion steps instead of the usual 20+, producing an image in around 0.1 seconds on H100 (~0.31s on RTX 4090).

4️⃣ AI Models - Video Upscaling - SeedVR2 now supports up to 4K output

SeedVR2 3B and 7B now support 1440p and 4K output resolutions. Previously limited to 1080p. And 720p video upscaling with SeedVR2 is now significantly faster (20 ~ 37%) following optimization of the inference pipeline

image

5️⃣ Tool - Layer Compose

A new tool for merging multiple images (typically with transparency) into a single flattened output like Photoshop's "Merge Visible" button. Order the layers, wire them into Layer Compose, pick a fit mode (cover, contain, or stretch for handling mismatched sizes), and the tool alpha-composites everything into one image at the background's resolution.

6️⃣ Tool - List Extract and List Inject

Two new companion tools for working with image lists. List Extract pulls a single image out of a list by its 0-based index; List Inject puts an image back into a list at a specific index, replacing what was there. Together they enable the canonical extract → process → inject pattern — pull one item out, enhance or fix it with any image tool, then slot the improved version back where it belongs.

Caption: "Extract → process → inject pipeline: pull one frame from a video list, upscale it, and slot the enhanced version back at the same index”

7️⃣ Tool - Video Split and Video Reassemble - Open and Close the Video Loop

Two new companion tools that bookend any frame-by-frame video pipeline. Video Split decodes a video into an image list at the start; Video Reassemble encodes an image list back into a video at the end. Together with List Extract and List Inject, they enable full-video processing using any image model in the studio.

8️⃣ External AI Model - Video Analysis (Gemini via Google API) — Turn a Video Into Text

A new external API model connection powered by Google Gemini. Feed it a video, add an optional instruction, and Gemini returns text — transcripts, timecodes for specific events, highlight summaries, or answers to natural-language questions about what happens on screen or on the audio track.

Platform Improvements

1️⃣ Performance: Cosmos3-Nano video generation ~15% faster

Video generation with Cosmos3-Nano is now approximately 15% faster following the application of torch.compile optimization.

2️⃣ FlashVSR and SparkVSR now support higher input resolutions

FlashVSR now accepts up to 1024px input (4K output), and SparkVSR up to 2048px input (4K output). Previously limited to 720px and 480px respectively.

3️⃣ My Flows - Search + Multi-Select (Shift-Click)

My Flows supports both search-by-name (real-time filter as you type) and Shift-click multi-select for bulk operations (bulk delete, bulk share).

4️⃣ Batch Run Email - Deep Link to Results

Batch run completion emails now include a deep link that jumps straight to the results for that batch, instead of navigating through the Batch Runs page manually.

5️⃣ Tool - Fit Option on Image Add / Multiply

The Image Add and Image Multiply nodes now have a fit option that auto-resizes the smaller input to match the larger one before blending - no more manual Image Resize node in between when dimensions don't line up.

image
image

🔧 Bug Fixes

  • Library Search Restored

Fixed a bug where the search bar in the Library wasn't returning results - Fixed

  • Random Seed Fixed for Video Generation Models

Fixed two related bugs affecting all video generation models: the "Generate random seed" button wasn't updating the seed value correctly, and seed values weren't being saved across sessions - Fixed.

  • Gemini 3.1 Flash and Gemini 3 Pro API Connection Fixed

Fixed a bug where Gemini 3.1 Flash and Gemini 3 Pro weren't connecting through the API and couldn't be called from flows - Fixed

  • GLM-Image 540px Resolution Option Fixed

Fixed a bug where selecting 540px resolution on the GLM-Image node returned an error. The option has been replaced with 576px, which also produces accurate 16:9 and 9:16 aspect ratios.

  • Document Help Window — Overflowing Off-Screen

Fixed a bug where the in-app document help window could open partially off-screen. Now always positions correctly within the viewport.

Release Overview [1.1.5]

CNAPS Studio v1.1.5 introduces Easy UI (Beta), a new front door that lets you skip workflow building entirely. Pick a card for a common task, drop in your input, hit Run, and the result lands right back on the same page; 12 cards live in Phase I, more on the way. On the model side: RF-DETR Medium (object detection plus pixel-precise segmentation), and two new video generators, Cosmos3 Nano (by NVIDIA) and Wan2.2 TI2V. New utility tool Text Decorator adds reusable prompt templates. Polish: Gemma 4 faster on multimodal prompts, Helios-Base faster for video generation, and Node Lock to preserve a result across re-runs. Sharing also got smarter: Per-input fork control lets you choose what travels with a forked flow.

✨ What's New

Major Improvements

1️⃣ Easy UI (Beta) - Start with a Single Click, No Workflow Building Required

A brand-new way to use CNAPS Studio, Easy UI lets you skip the workflow building entirely. Pick a card for what you want to do, drop in your input, hit Run, and the result lands right back on the same page. No nodes, no wiring, no canvas.

How it works:

  • Pick a card - 12 starter cards in this beta release, each tuned for a common task
  • Drop in your input - text prompt, image, video, or a combination; the card tells you what it accepts and roughly how long it takes
  • Hit Run - a pre-built workflow runs behind the scenes; you never see the canvas
  • See your result right there on the same page

Task-focused cards run a single best-fit workflow.

image

When you want more control, the "Open Cnaps Studio" button (top-right of any card) drops you into the full workflow canvas with the same flow loaded so advanced users can see exactly how the card works under the hood, tweak it, or build on top of it.

This is Phase I, shipping with 12 cards live in beta. More cards and feature updates will land in future releases.

2️⃣ Per-Input Fork Control — Choose What Travels with a Forked Flow

image

When someone forks your flow, you can now control on a per-input basis whether each input image or text prompt gets copied along with it. A badge appears on every input node showing the current state - click it to toggle between "copied when forked" and "excluded."

Sensible defaults:

  • Input images: NOT copied by default - protects photos that may contain personal or sensitive content (faces, IDs, client deliverables, etc.)
  • Text prompts: copied by default - prompts often define what the flow actually does (recipes, instructions, style descriptions), so they usually belong with the fork

You can override either default on any individual input node anytime.

New AI Models & Tools

1️⃣ AI Models - Object Detection (RF-DETR Medium) & Segmentation (RF-DETR Seg Medium)

Two new models from Roboflow that look at a photo and find everyday things in it - people, animals, vehicles, furniture, food, electronics, and 70+ other common categories. They work the same way under the hood, but come in two flavors:

  • RF-DETR Medium draws a box around each object it finds, along with a label ("person", "car", "dog") and a confidence score. Use this when you just need to know what's in the photo and roughly where.
image
  • RF-DETR Seg Medium goes one step further - it traces the exact outline of every object, pixel by pixel. Use this when you need the actual shape, not just a rectangle (cutouts, background removal, compositing, or photo editing)
image

Both are fast enough for live video, and both handle tricky cases well - two cats sitting together come back as two separate outlines instead of one blob. You can also adjust how strict the models are: a higher confidence setting returns fewer but more certain results; a lower setting catches more, including weaker guesses.

How they compare to YOLO: at the same speed, RF-DETR-Seg-Medium produces noticeably cleaner outlines than comparable YOLO segmentation models. The trade-off is memory - RF-DETR uses a bit more of it for the same speed.

2️⃣ AI Models - Video Generation - Cosmos3 Nano and Wan2.2 TI2V

Two new open-source video generators join the lineup. Both can turn a text prompt into a short cinematic video clip, and both can also animate a starting image with a text description of what should happen next. Both output 720p video at 24 fps.

Cosmos3 Nano (NVIDIA) - designed for realistic motion and physical-world scenes. NVIDIA built it for robotics and autonomous-driving research, which makes it especially strong on motion that "looks right." A standout feature: optional synchronized audio — describe ambient sound in your prompt (e.g., "footsteps on gravel, wind in trees") and the model produces matching audio along with the video.

Wan2.2 TI2V 5B (Wan-AI) - a smaller, faster alternative. Among the fastest open-source 720p video generators at this quality level. Natively supports English and Chinese prompts.

3️⃣ Tool - Text Decorator - Reusable Prompt Templates

A new text tool for wrapping live, changing content inside a fixed prompt template. Write your prompt once with a placeholder where the variable part goes (default placeholder: <<<input>>>), connect the dynamic text from any upstream node (OCR, an LLM, an image-description model), and Text Decorator outputs the finished string ready to feed into whatever comes next.

Platform Improvements

1️⃣ Faster Gemma 4 Responses

Gemma 4 now runs noticeably faster, up to ~3× on multimodal prompts and ~1.9× on text-only, with no quality change. The speedup applies automatically to every existing Gemma 4 flow.

2️⃣ Node Lock - Preserve a Result Across Re-Runs

You can now lock any completed node to keep its output frozen across re-runs of a flow. Locked nodes are skipped during execution, so a result you want to keep such as an LLM response you liked, a random-seed image, or a slow OCR or upscale pass stays exactly the same while you tweak everything downstream. Click the lock icon in the node header to toggle it on or off.

image

3️⃣ Performance: Helios-Base inference ~25% faster

Video generation with Helios-Base is now approximately 25% faster (51 min → 39 min for a typical 99-frame job) following the application of torch.compile optimization.

4️⃣ Improved: Real-time progress updates during inference

Progress bars now update smoothly step-by-step during inference across 18 models, including image/video generation, super-resolution, and upscaling models. Previously, progress would jump in large increments.

🔧 Bug Fixes

  • SAM3 (Scene) Restored to the Model List

Fixed a bug where SAM3 (Scene Segmentation) sometimes couldn't be called via the API. The model is now correctly registered at startup and works as expected.

  • Manual Crop - Modal Label and Circle/Polygon Shapes Fixed

The Edit Region modal was incorrectly labeled "Draw Blur Region" and the Circle and Polygon shape options weren't drawing or cropping correctly - Fixed

  • Instant Identity Masking Template Restored

Fixed a bug affecting the Instant Identity Masking pre-built template (available from the "Start Your Flow" picker). Since the previous release, the template was producing image-size-mismatch errors at the blending stage and failing to complete - Fixed

Release Overview [1.1.4]

CNAPS Studio v1.1.4 brings ControlNet end-to-end: v1.1.3's Canny Edge and Depth Anything annotators now pair with two new SDXL generators (ControlNet XL Canny and ControlNet XL Union) to lock the structure of any reference photo while changing everything else with text. Three more open-source models join the lineup, Molmo2-8B (multi-image Q&A), OmniGen2 (unified image generation / editing / composition), and SAM 3.1 (Scene + Video, ~7× faster multi-object tracking), plus GPT-5.5 / GPT-5.5 Pro and Gemini 3.5 Flash on the existing API nodes, and a new Text Concatenate utility.

✨ What's New

Major Improvements

1️⃣ ControlNet - Lock the Structure, Change Everything Else

CNAPS Studio now ships the complete ControlNet workflow end-to-end: turn any photo into a structural "control map" with one of our annotators, then use that map plus a text prompt to generate a brand-new image that keeps the original's geometry but takes on whatever style, lighting, or content you describe.

image

Four nodes work together to make this possible:

  • Canny Edge Annotator (shipped v1.1.3): turns a photo into a white-edges-on-black line map that captures the outlines
  • Depth Anything Annotator (shipped v1.1.3): turns a photo into a grayscale depth map (bright = near, dark = far) that captures the 3D layout
  • ControlNet XL Canny (new in v1.1.4): SDXL image generator that follows a canny edge map; dedicated canny model, sharpest structural adherence for edge-based tasks
  • ControlNet XL Union (new in v1.1.4): SDXL image generator that follows 6 different conditioning types in one model (canny, depth, openpose, hed, normal, segment); the all-in-one workhorse

A few things you can do with this:

  • Restyle a scene without losing its layout: feed a room photo through Depth Anything → ControlNet XL Union (depth mode) with a prompt like "same room, sunset light, watercolor" and the geometry stays while the style transforms
  • Turn sketches into rendered images: pass a hand drawing through Canny Edge Annotator → ControlNet XL Canny with a style prompt; the outlines become a fully painted scene
  • Architectural & product visualization: keep the structural lines of a reference exactly intact while exploring different finishes, lighting, materials, eras
  • Pose-driven character generation: supply an openpose skeleton to ControlNet XL Union and place a character in any pose, styled by text
  • Stack conditions: Union was trained to fuse multiple control signals at once (e.g., openpose + depth) so you can layer constraints in a single generation

New AI Models & Tools

1️⃣ AI Model - Multimodal Language Model - Molmo2-8B

A new open-source vision-language model from the Allen Institute for AI. Give it a text prompt plus one or two images, and it answers questions, writes captions, compares the two images, or counts what's in them, all in plain natural language. It runs locally and is the best open model in its size class for tasks like counting and visual Q&A.

image

2️⃣ AI Model - Image Generation - OmniGen2

A single image model that does three jobs you used to need three separate tools for: generating images from text, editing existing images by typing instructions, and combining elements from multiple reference photos into a new scene.

3️⃣ AI Model - Image Control - ControlNet XL Canny / ControlNet XL Union

ControlNet XL Canny: A purpose-built canny-only ControlNet for SDXL. Same idea as Union, but trained exclusively on edge maps — so when the structural lines of your reference really need to be respected (architecture, product shots, clean sketches), this gives sharper adherence than the general Union model. Pair it with the Canny Edge Annotator from v1.1.3

ControlNet XL Union: A single SDXL ControlNet that takes any of six different "control image" types and uses it to lock the structure of the generated image while you change everything else with text. Pair it with the annotator nodes from v1.1.3 (Canny Edge, Depth Anything) or with any other compatible control image, and switch between conditioning modes without loading separate model.

4️⃣ AI Model - SAM 3.1 - Better Segmentation for Images and Videos

The Object Multiplex refresh of the SAM 3 family, now available in two flavors: SAM 3.1 (Scene) for image segmentation and SAM 3.1 (Video) for video segmentation and multi-object tracking. Same workflow you already know from SAM 3 - type a short phrase like "person", "red car", or "cat in a hat" and the model finds and masks every matching object. Under the hood, a new shared-memory design gives better detection accuracy at the same inference cost, plus a dramatic speedup for multi-object video tracking.

5️⃣ External AI Model - GPT-5.5 / GPT-5.5 Pro and Gemini 3.5 Flash

OpenAI's latest two models, GPT-5.5 and GPT-5.5 Pro, are now selectable from the Model variant dropdown on the ChatGPT node.

image

Gemini 3.5 Flash is also selectable from the Gemini node. And Google has scheduled both Gemini 2.5 Pro and Gemini 2.5 Flash for end-of-life on October 16, 2026. They'll keep working until that date, but it's a good time to start planning the migration.

image

6️⃣ Tool - Text - Text Concatenate

A lightweight text tool that takes up to four text streams and joins them into a single output, with a separator you control between each segment.

7️⃣ QWEN-Image-Edit-2509 Has Been Deprecated

As announced in v1.1.3, QWEN-Image-Edit-2509 is now deprecated as of v1.1.4. Existing flows that reference this model will need to be updated to use the newer, better version.

What this means for you:

  • Switch to QWEN-Image-Edit-2511 - same family, improved quality. It's a drop-in replacement for almost every flow that used 2509
  • Existing flows still using 2509 will show a deprecation warning; we recommend swapping the node out at your earliest convenience

Platform Improvements

1️⃣ Simplified Model Names Across the Library

We cleaned up node names so they read faster in the canvas and library. Two things were happening that made names noisy: the model name was duplicating the category it sat under (e.g., the Image Generation category had a node called Image Generation (Z-Image-Turbo)), and we had two competing naming patterns — Role (Model) in some places and Model (Role) in others.

2️⃣ Parameter Sliders on Tool Nodes

We replaced the numeric input boxes on tool node parameters with sliders. Same values, same ranges, but now you can see where the current value sits inside its allowed range, and drag to adjust instead of clicking into a text field and typing

image

3️⃣ Input Resolution Limits Now Shown on Video Upscaler Nodes

We added the input resolution cap to the node description on each video upscaler, so you can see what the model accepts before you hit Run and get an error:

  • SparkVSR node now reads: "Diffusion video upscale (2x/3x/4x) with optional ref image. Input resolution up to 480×480."
  • FlashVSR v1.1 node now reads: "One-step diffusion 4× video super-resolution (FlashVSR Tiny pipeline). Input resolution up to 720×720."
  • image

🔧 Bug Fixes

  • Image Output Nodes Stay Clickable Across Multiple Runs

Fixed a bug where, in flows with multiple Image Output nodes, the middle ones could go unresponsive after running the flow several times. Only the most recently added/run Image Output would still react to clicks. The previous workaround was to close and reopen the flow to get them all working again. Fixed.

Release Overview [1.1.3]

CNAPS Studio v1.1.3 is built around three themes: workflow precision (11 new tools for masking, cropping, blurring, resizing, video-frame extraction, text overlays, edge detection, and side-by-side text comparison), AI model breadth (6 new open-source models spanning agentic coding, on-device chat with audio, real-time and keyframe-guided video upscaling, depth-map generation, and pro-quality background removal - plus GPT Image 2.0 and Claude Opus 4.7 via API), and the expanding Gemma 4 family (a compact E2B variant with audio input, and multi-image input on the existing 31B). Smaller polish includes multi-select bulk delete on My Flows, plus a one-click upgrade path for QWEN-Image-Edit-2509 ahead of its end-of-May deprecation.

✨ What's New

New AI Models & Tools

1️⃣ AI Model - Multimodal Language Model - Qwen3.6-35B-A3B & Gemma 4-E2B-it

Qwen3.6-35B-A3B is the newest open-weights multimodal AI model from the Qwen Team, built especially for coding, software projects, and multi-step agent workflows. It reads and responds to text, images, and video, and can handle conversations long enough to include an entire book, a full codebase, or hours of meeting transcripts in a single chat. A new Thinking Preservation feature lets the model carry its reasoning across multiple back-and-forth messages, which is particularly helpful for iterative work like step-by-step coding, debugging, or planning agents that handle complex tasks.

Gemma 4 E2B-it is the smallest and fastest member of Google DeepMind's Gemma 4 family, designed to run efficiently on phones, laptops, and other lightweight devices. Unlike the larger Gemma 4 31B (added in v1.1.2), which handles only text and images, the E2B variant adds audio input: drop in a voice recording and ask the model to transcribe what was said, or translate spoken language into another language directly. It still reads text and images, handles 35+ languages, and comes with the same step-by-step thinking mode as its bigger sibling. The trade-off is quality on the hardest tasks; For those, the 31B is still the better pick, but for fast, lightweight work like voice transcription, mobile-friendly chat, or quick visual analysis, E2B is plenty capable.

2️⃣ AI Model - Video Upscaling - FlashVSR-v1.1 & SparkVSR

FlashVSR-v1.1 is a new open-source video upscaling model that takes low-resolution video and reconstructs it at much higher resolution, tuned specifically for 4× upscaling of legacy footage, web-compressed clips, and other low-quality video sources. What sets it apart from other video upscalers is speed, up to 12× faster than competing diffusion-based options.

SparkVSR is a new video upscaling model that does something unusual for the category; it lets you guide the upscale with your own reference frames. Most video upscalers are "black boxes": press go, accept whatever the model produces. SparkVSR works differently. If you're not happy with what an automatic upscale gave you, take a single frame, clean it up using any image upscaler you trust (a commercial API like Nano Banana Pro, or an open-source tool like PiSA-SR), and feed that keyframe back to SparkVSR. The model then propagates that look across the entire video while keeping motion coherent with the original footage

3️⃣ AI Model - Subject Segmentation / Background Removal (BiRefNet)

BiRefNet is a new open-source AI model for background removal and subject cutouts. Think of it as the AI-powered, far smarter version of the "remove background" button in photo editing tools. Drop in a photo and BiRefNet automatically finds the main subject and produces a clean, precise outline you can use to lift it off the background. What sets BiRefNet apart is how well it handles the tricky edges that simpler tools usually mangle: individual strands of hair, fur, lace, leaves, mesh, and other fine details. It works especially well on high-resolution images, making it ideal for product photography, e-commerce visuals, marketing assets, and professional photo prep - anywhere you need a clean transparent cutout. The mask it produces can also feed into other tools for background replacement, relighting, or further editing.

image

4️⃣ External AI Model - GPT Image 2.0 & Claude Opus 4.7

Two new flagship external models have joined the External Models lineup. Connect your API keys to access GPT Image 2.0 and Claude Opus 4.7.

5️⃣ Tool - Image Masker - Object Picker

Object Picker is a new interactive masking tool that pairs with SAM2 and SAM3 (Scene Segmentation). When either model segments an image, it returns each detected object with a generic label like object_0, object_1, and so on, figuring out which number corresponds to the cat vs. the chair vs. the lamp used to mean trial and error. Object Picker fixes that: connect it after SAM2 or SAM3, click the pick objects button on the node, and an interactive modal opens showing the segmentation overlay paired with each object's description. Just click the objects you actually want, and Object Picker builds a clean black-and-white mask covering exactly your selection, ready to feed into any downstream masking, inpainting, blending, or color tool.

image
image
image

6️⃣ Tool - Blur & Effects - Lens Blur, Selective Blur, and Image Blur (Simple)

Lens Blur is a new blur tool that mimics how a real camera lens renders out-of-focus areas, rather than the soft "foggy haze" effect of standard Gaussian blurs. The difference: when you blur a photo with a Gaussian kernel, bright highlights smear into a soft glow. With Lens Blur, those bright spots bloom into clean circular bokeh discs, the look photographers chase for portrait, product, and cinematic shots. It's optical and photographic, not digital and smeared.

image

Selective Blur is a new tool that blurs just the region you draw on an image, leaving the rest untouched. Click Edit Region on the node, pick a shape, Rectangle, Circle, or Polygon, draw it directly on the image, and Apply. Then dial in the Blur slider (1–10) for strength and the Feather slider (0–100) for how softly the blurred area fades into the rest of the picture.

image

Image Blur (Simple) is the easiest way to blur an entire image - one slider, no tuning. Drag the Blur intensity slider from 1 (subtle smoothing) to 10 (heavy, near-abstracted blur), and the tool handles everything else under the hood. Sits alongside the more advanced Image Blur (Standard) and Image Blur (Fast) tools as a friendlier entry point for users who don't want to fiddle with kernel sizes, color multipliers, or scale factors. Great for quick privacy blurs, soft backgrounds, glow layers, or noise smoothing.

7️⃣ Tool - Text & Annotation - Draw Text

Draw Text is a new annotation tool that lets you place text anywhere on an image with full control over how it looks. Type your text (multi-line is fine), pick a font with Bold/Italic toggles, set the size, color, and alignment, and optionally add a colored background box behind the text with adjustable opacity for readability over busy images. Then click Edit Text Area and visually drag the text box to exactly where you want it on the image.

image

8️⃣ Tool - Image Resizing - Image Conditional Resize & Image Resize to Match

Image Conditional Resize automatically upscales small images to the next standard resolution tier. Pick a maximum tier (512, 1024, 1536, 2048, 2560, 3072, or 4096 pixels on the short side), and the tool bumps small images up to the smallest tier that fits both dimensions while preserving aspect ratio. Already big enough? The image passes through untouched. Perfect as a "safety net" early in workflows feeding AI models with minimum-resolution requirements: tiny inputs get raised to a safe minimum, while properly sized images aren't unnecessarily processed.

Image Resize to Match takes two image inputs and resizes the first one to exactly match the second's width × height. Plug both images in and the output is a perfectly aligned pair. Ideal for compositing (Image Add, Image Multiply), masking, comparison, and any workflow where two images must share dimensions.

9️⃣ Tool - Comparison - Text Compare

Text Compare is the text-output companion to Image Compare. Connect any text-producing nodes (LLMs like Gemma 4 or Qwen3.6, OCR models like GLM-OCR or DeepSeek-OCR, BLIP image descriptions, Text Inputs) and read all responses in parallel. Perfect for model bake-offs (compare different LLMs on the same prompt), prompt A/B testing (same model, different prompts), OCR quality assessment (compare transcriptions across OCR engines), or any workflow where you want to read multiple text outputs at once without copy-pasting between tabs.

🔟 Tool - Selection & Extraction - Manual Crop & Video Frame Extract

Manual Crop is a new visual cropping tool that lets you draw the region you want to keep directly on the image. Click Edit Region on the node, choose a shape (Rectangle, Circle, or Polygon), drag it into place, and Apply. The output is exactly that region.

image

Video Frame Extract pulls a single frame out of a video and outputs it as a standalone image, bridging video pipelines and the much larger set of image-only tools available in CNAPS Studio. Just enter a 0-based frame index (0 for the first frame, 1 for the second, etc.) and the tool decodes that exact frame as a standard image, ready to feed into any image upscaler, generator, color tool, masking node, or text overlay.

image

1️⃣1️⃣ Model - Image Control - Depth Anything Annotator

Depth Anything Annotator turns any regular photo into a clean depth map, a grayscale image where bright pixels are close to the camera and dark pixels are far away. No depth sensor, stereo camera, or special hardware needed; works from a single ordinary photo. Built on Depth Anything V1 (Large), trained on roughly 62 million images and widely considered the best open-source choice for this task. Particularly good with fragile details like hair, foliage, branches, and mesh where simpler depth models tend to smear.

image

1️⃣2️⃣ Tool - Image Control - Canny Edge Annotator

Canny Edge Annotator turns any image into a clean white-edges-on-black edge map using OpenCV's classical Canny algorithm. The result is exactly the format ControlNet-Canny expects, making this the standard preprocessor for any image generation flow that uses edges as structural guidance.

image

Platform Improvements

1️⃣ Gemma 4 31B - Multi-Image Input Support

Gemma 4 31B can now handle multiple images at once instead of just one per prompt. Drop in several images alongside a text question and ask Gemma to compare them, summarize a sequence, reason across a set of related visuals, or analyze a multi-page document.

2️⃣ My Flows: Multi-Select for Bulk Delete

Cleaning up your workspace is now much faster. The My Flows page now supports multi-select: a new checkbox column lets you tick off as many flows as you want, see a running "X selected" counter at the top, and delete them all in one go with a single click on the trash icon.

image

3️⃣ QWEN-Image-Edit-2509 - Deprecation Notice & Easy Upgrade

⚠️ Heads up: QWEN-Image-Edit-2509 will be deprecated at the end of May 2026. The newer QWEN-Image-Edit-2511 offers improved quality and better detail preservation, and is now the recommended option for all image editing workflows. To make the switch painless, any flow still using QWEN-Image-Edit-2509 now shows an Upgrade banner on the node: Click it and an upgrade dialog appears letting you swap in the newer 2511 version with one click. Compatible parameters (like diffusion steps and seed) are automatically migrated to the new node; settings unique to 2511 are initialized to its defaults. After the deprecation date, the 2509 node will no longer run in workflows, so we recommend upgrading any active flows before then.

image
image

Release Overview [1.1.2]

CNAPS Studio v1.1.2 brings Selective Flow Sharing with Viewer/Editor roles and Link Sharing, the largest AI model expansion (five new models: FLUX.1-schnell, FLUX.2-klein-4B, GLM-Image, Gemma 4 31B, and GLM-OCR), localization in 5 new languages, a Publish button directly inside the Community site, and a resizable Output Image Viewer node. Built for global teams, faster image workflows, and document-heavy use cases.

✨ What's New

Major Improvements

1️⃣ Flow Sharing - Invite Specific People

The Share menu has been redesigned to give owners precise control over flow access. v1.1.2 introduces three access modes and per-person permissions, bringing flow sharing in line with familiar tools like Google Docs :

  • Private: Only the owner can access this flow (unchanged).
  • Link Sharing (new): Anyone with the link can access this flow. Unlike Public, the flow is not publicly discoverable and only users with the direct link can open it.
  • Public: Anyone with the link can view this flow (unchanged).

A new People with access section lets owners add individual collaborators by searching their name or email, and assign each person as a Viewer (can view and run the flow) or Editor (can view, run, and modify the flow). Owners can also enable Require access code to password-protect a shared link, and Notify by email to automatically send the invitee an email with the share link.

image

2️⃣ Multi-Language Support - 5 New Languages Added

CNAPS Studio now supports 7 languages in total. Building on the existing English and Korean (한국어) support introduced in v1.0.2, this release expands the language menu with 5 additional languages: Japanese (日本語), Chinese (中文), Spanish (Español), Arabic (العربية), and French (Français). Users can switch language anytime under Account Settings → Profile → Language.

New AI Models & Tools

1️⃣ AI Model - Multimodal Language Model (Gemma 4 31B)

Gemma 4 31B is a Google DeepMind's newest multimodal language model that reads both text and images and produces text in response. It's especially good at long-form reasoning tasks like reading and summarizing long documents, parsing scanned PDFs, understanding charts and diagrams, or holding conversations with a lot of context. Gemma 4 includes a thinking mode you can switch on, which lets the model reason through a problem step by step before answering. It also handles 35+ languages out of the box and supports tool calling so it can be wired into agents and automated workflows.

2️⃣ AI Model - Image Generation (FLUX.1-schnell)

FLUX.1-schnell is an open-source image generation model from Black Forest Labs, designed for speed without sacrificing quality. It turns text prompts into high-quality images in just a fraction of a second per image, ideal for fast prototyping, rapid iteration, and any workflow where you need lots of images quickly.

3️⃣ AI Model - Image Generation (FLUX.2-klein-4B)

FLUX.2-klein-4B is the newest member of Black Forest Labs' FLUX.2 family and the first model that handles both image generation and image editing in a single tool. It's also their fastest yet, generating an image in under a second, which makes it ideal for live or interactive workflows where you want results in real time. A standout capability is multi-reference editing: drop in several reference images alongside a text prompt, and the model uses all of them together to guide the final composition, style, and content.

4️⃣ New AI Models - Image Generation & Editing (GLM-Image)

GLM-Image is an open-source image model from Z.ai that handles both creating images from scratch and editing existing images. Where it really shines is in getting text right inside images, a notoriously difficult problem for most image AI. It's especially strong on information-dense layouts like infographics, recipes, posters, charts, and product designs, where readable text and accurate details matter as much as the visuals. GLM-Image also handles image-to-image tasks like style transfer, keeping characters consistent across edits, and combining multiple subjects into a single scene.

5️⃣ New AI Models - Document OCR (GLM-OCR)

GLM-OCR is a compact OCR model from Z.ai built for reading and understanding documents. Despite being one of the smallest models in its category, it currently ranks #1 on the leading document-understanding benchmark. It handles the parts of OCR that usually trip up general-purpose AI: complex tables, code-heavy pages, mathematical formulas, official seals and stamps, and dense business layouts. Beyond extracting plain text, GLM-OCR can also pull structured information out of documents (IDs, invoices, forms) directly into a JSON format you can use downstream.

Platform Improvements

1️⃣ CNAPS Community - Publish Button Added

A new Publish button has been added directly to the Community interface, letting users publish their flows without having to switch back to CNAPS Studio first.

2️⃣ Parameter Input Fields - Clear & Retype Support

Based on user feedback, configuration inputs now support natural number editing. Previously, changing a parameter value (e.g., Diffusion steps, Output width) required selecting the existing number with your mouse before typing a new one. Users could not leave the field empty. Now, users can simply delete the current value, clear the field entirely, and type a new number from scratch.

3️⃣ Editor Settings: Auto-connect on drag

Auto-connect makes it quick to wire nodes together by dragging one near another, but on a dense canvas it sometimes connected nodes you didn't mean to link. You can now disable this behavior from Editor Settings (gear icon, bottom right) and create connections manually instead.

image
image

4️⃣ Output Image Viewer - Drag to Resize

The Output Image Viewer node is now resizable. A new drag handle in the bottom-right corner of the node lets users grab and drag to resize the viewer to any size, scaling the displayed image preview along with it. Useful for inspecting fine details on generated images directly within the canvas, without having to open a full-screen view or compare modal. The original aspect ratio of the output image is preserved as the node grows or shrinks.

CNAPS Studio v1.1.1 delivers a major upgrade to Batch Run with full multi-input and output support, making it compatible with real-world workflows for the first time. This release also expands our video AI lineup with two new models, i) Helios-Base for maximum quality video generation and ii) SeedVR2 for video upscaling and restoration, and brings significant improvements to MCP with Claude, enabling smarter pipeline recommendations, context-aware model selection, and end-to-end workflow execution directly in chat. Additional highlights include My Flows workflow thumbnails, node numbering, CNAPS Community video support, and a wave of bug fixes.

Release Overview [1.1.1]

✨ What's New

Major Improvements

1️⃣ Batch Run - Multi-Input & Output Support

Batch Run now fully supports workflows with multiple inputs and outputs. Previously, Batch Run was limited to single input workflows, making it incompatible with most real-world workflows. The redesigned Batch Run UI allows users to configure each input node individually with three flexible modes: Batch (different input per run), Shared (one input applied to all runs), and Flow Value (use the default value from the workflow). Each input node is clearly identified by the new node numbering system (e.g., #1, #3, #8). Once all inputs are configured, users can export them as a CSV file to reorganize the pairing order and re-import, making large batch setups fast and flexible.

image

2️⃣ MCP with Claude - Smarter Workflow Intelligence Claude's MCP integration has been significantly upgraded with an AI intelligence layer that better understands user intent when building workflows. Previously, Claude relied on simple keyword matching which often led to incorrect model recommendations. Key improvements include:

  • Smarter Pipeline Recommendations: Claude now understands the meaning behind requests in any language, not just keywords. Multi-step requests like "colorize and upscale" are automatically broken down into the correct pipeline.
  • Context-Aware Upscale Model Selection: Claude now selects the most appropriate upscaling model based on context; PiSA-SR for best quality, LatticeNet for speed, and Swin2SR for large images or detail preservation.
  • Actionable Error Diagnosis: When a pipeline fails, Claude now explains exactly what went wrong and suggests a fix, instead of just showing a generic failure message.
  • End-to-End Execution in Chat: Users can now complete the full workflow entirely within the Claude chat, without being redirected to CNAPS Studio to run the flow manually.
  • Improved Image Upload Guidance: Claude now instantly generates an upload link instead of simply asking users to upload an image with no clear direction.

New AI Models & Tools

1️⃣ New AI Models - Video Upscaling and Enhancement - SeedVR2 3B & 7B

SeedVR2 is a one-step video restoration and upscaling model by ByteDance, available in two sizes, 3B and 7B. It restores and enhances degraded videos in a single diffusion step using adversarial post-training, delivering high-resolution output with strong temporal consistency and fine-grained detail. SeedVR2 is ideal for enhancing AI-generated or real-world videos that suffer from blur, low resolution, or visual degradation. The 7B model offers higher quality output, while the 3B model provides a faster, more lightweight alternative.

2️⃣ New AI Models - Video Generation (Helios-Base)

Helios-Base is the highest quality model in the Helios family, joining the previously released Helios-Distilled in CNAPS Studio. While Helios-Distilled is optimized for speed and efficiency, Helios-Base prioritizes maximum video quality, generating coherent, high-fidelity video from text, image, or existing video input at 19.5 FPS on a single GPU. Choose Helios-Base when output quality is the top priority, and Helios-Distilled when speed and efficiency matter most.

Platform Improvements

1️⃣ CNAPS Community - Video Support & Bug Fixes

Features:

  • Video support has been expanded across the Community. Video cards now appear on the main page, video uploads are supported in the Publish & Share modal (with cover selection), video thumbnails are shown in My Posts, and video-related node types are now available in the Detail Preview. Video playback is also consistent across all card view settings.

Bug Fixes:

  • Fixed an issue where the sticky search header was intercepting clicks on the logout button when scrolling down the page.
  • Fixed an issue where workflow posts were not visible to users who were not logged in.
  • Fixed community search not returning results correctly.
  • Fixed auto-translation not working consistently across community posts

2️⃣ My Flows - Workflow Thumbnails Added

Based on user feedback, the My Flows page now displays a small thumbnail preview of each workflow's canvas layout. Click on a thumbnail to view it in a larger size, making it easier to identify and navigate to the right workflow at a glance

3️⃣ Helios - Expanded Generation Parameters

The Helios model family now offers greater control over video generation with additional configuration options. On top of the existing frame count and frame rate settings, users can now configure output resolution, pyramid inference stages, per-stage inference steps, and first chunk amplification for finer control over quality and performance.

4️⃣ Node Numbering Added

Each node on the canvas now displays a unique number (e.g., #4, #11) making it easier to identify and reference specific nodes when communicating with teammates, reporting issues, or tracking results in batch runs.

image
image

🔧 Bug Fixes

  • Copy & Paste Workflow - False Save Error

When copying and pasting a workflow from another user, a misleading error message, "Some edits could not be saved — please verify your changes," was incorrectly displayed even though the workflow was saved successfully. Fixed.

  • Nano Banana - NSFW Error Message Not Displaying

When Nano Banana returned an error related to NSFW content, the error message was not being passed through and displayed to the user in CNAPS Studio. Fixed.

  • Image Compare Node - Download Button Fixed

When clicking the download button in the Image Compare node, the image was opening in a full browser window instead of downloading. Fixed.

  • Helios-Distilled - Video Input Temporarily Disabled

The Helios-Distilled model currently accepts both image and video as optional inputs, which allowed unsupported input combinations to be used. To prevent unintended behavior, the video input (Video-to-Video) has been temporarily disabled while a proper input validation solution is being developed. Text-to-Video and Image-to-Video modes remain fully supported.

  • PiSA-SR + Real-World Blur Removal - GPU Memory Error Fixed

When passing an upscaled image from the PiSA-SR model directly into the Real-World Blur Removal model, a "GPU memory exceeded" error was incorrectly triggered. The two models can now be used together in the same workflow without any memory errors.

  • Error Messages - Improved Clarity

Several error messages that were vague or unclear have been updated to provide clearer, more actionable descriptions of what went wrong, making it easier for users to understand and resolve issues.

  • WebP Image Loading & Execution

An issue where WebP images were failing to load or execute correctly in workflows has been resolved.

Release Overview [1.1.0]

CNAPS Studio v1.1.0 is our most significant release to date. This update launches two major product milestones: 1) an Interactive Community where users can browse, discuss, fork, and run community workflows directly in their own workspace, and 2) MCP Support with Claude, enabling natural, conversational interactions with workflows that meaningfully lower the barrier for non-technical users. This release also marks a major step forward in video AI, introducing Helios-Distilled, the first open-source video generation model. On top of these, v1.1.0 adds two powerful new AI models, QWEN-Image-Edit-2511 and FireRed-Image-Edit-1.1, and introduces Flow History for tracking and restoring previous workflow states, and delivers a wave of community-requested features including Quick Add Node, Right-Click to Add Node, Sticky Notes, Split Compare mode, and background workflow execution.

✨ What's New

New Features

1️⃣ Flow History - Track & Restore Previous States

Never lose your work again. A new History panel is now available on the right side of the canvas, recording every edit and modification made to your workflow in real time from adding and connecting nodes to moving and deleting them. You can review your full edit history and restore to any previous state with a click, giving you complete freedom to experiment without worry.

image

2️⃣ Quick Add Node - Search & Add from Canvas

Based on user feedback, you can now drag a connection line from any existing node and release it on an empty area of the canvas to instantly bring up a search menu. Find and add any AI model or tool directly from the canvas without having to browse the sidebar, making workflow building faster and more intuitive.

image

3️⃣ Right-Click to Add Node & Sticky Notes

Users can now right-click anywhere on the canvas to instantly open a node search menu and add any input, output, model, or tool node directly without browsing the sidebar. The menu also includes a new Sticky Note option, allowing users to add color-coded notes anywhere on the canvas to annotate, organize, or document their workflows.

image
image

4️⃣ Image Compare - Split Compare Mode Added

A new Split Compare mode is now available in the Image Compare node. Move your mouse or drag the divider to reveal a side-by-side split view of up to four images on a single canvas, perfect for inspecting before and after results from your workflow at a glance. Toggle between Split Compare and the existing Grid View to suit your comparison needs.

image

5️⃣ Collaborative Flow Editing - Real-Time Multi-User Workspace

CNAPS Studio now supports real-time collaborative workflow editing. Multiple users can join the same workspace and build, modify, and run workflows simultaneously, seeing each other's changes live on the canvas. Perfect for teams working together on complex AI pipelines without the back-and-forth of sharing files or screenshots

New AI Models & Tools

1️⃣ AI Models - Video Generation (Helios-Distilled)

Helios-Distilled is a real-time video generation model that produces high-quality videos up to 45 seconds long from a text description, a starting image, or an existing video clip. As the fastest and most efficient model in the Helios family, it generates at 19.5 frames per second on a single GPU. Unlike most AI video tools that struggle with longer clips, Helios-Distilled maintains full visual coherence throughout the entire length of the video, making it ideal for production-ready creative and commercial workflows.

2️⃣ AI Models - Smart Image Editing (QWEN-2511) - Enhanced Image Editing

QWEN-2511 is an upgraded version of QWEN-2509, bringing significant improvements to image editing quality and consistency. It delivers better subject identity preservation across edits, high-fidelity multi-person scene composition, integrated LoRA support for lighting and viewpoint enhancements, enhanced industrial design and material replacement capabilities, and geometric reasoning for generating construction lines and annotations. All edits are driven by simple natural language instructions, making it a powerful yet accessible tool for portrait editing, character storytelling, style transfer, and design workflows.

image

3️⃣ AI Models - Smart Image Editing (FireRed-Image-Edit-1.1)

FireRed-Image-Edit-1.1 is a production-grade open-source image editing model that sets a new state-of-the-art across three major benchmarks (ImgEdit, GEdit, and RedEdit). It excels at identity consistency, multi-element fusion of 10+ input elements, portrait and makeup styling, text style reference, and high-quality photo restoration, all driven by natural language instructions. With an end-to-end generation time of just 4.5 seconds and an extensive speed optimization suite, it delivers closed-source quality at open-source accessibility, making it ideal for portrait editing, creative production, and design workflows.

image

Platform Improvements

1️⃣ Workspace Invitations - Pending Status Added

Business plan users can now track outstanding team invitations directly from their workspace. Sent invitations will appear as "Pending" until the recipient accepts, giving workspace admins full visibility into who has been invited and who has yet to join.

2️⃣ Image Input Node - AVIF Format Support Added

The Image Input node now supports AVIF format in addition to the existing JPEG, PNG, and WEBP formats.

3️⃣ Workflow Runs in Background

Based on user feedback, workflows were stopping when users switched browser tabs or minimized the window. Workflow execution now continues to completion regardless of whether the browser window is active or in focus. This improvement also reduces unnecessary network overhead, resulting in a more efficient and reliable workflow experience.

4️⃣ Say Goodbye to Sora - VEO is Here

The VEO external model has received a significant update with greater creative control over video generation. Users can now select how to use their input image(s) via a new dropdown, choosing from Text to Video, Image to Video, First & Last Frame, or Reference Image. The Advanced Options panel has also been expanded with additional settings including prompt rewriting for better video quality, audio generation alongside the video, person generation policy, and reference image types for up to 3 reference images when using Reference Image mode. As Sora approaches deprecation, now is the perfect time to make the switch to VEO for your video generation workflows!

5️⃣ Workflow Rerun - Enabled for Text-to-Image Flows Previously, rerunning the same workflow was disabled as it would produce identical results. However, for text-to-image flows using the random seed feature, each run can generate a different output. Rerunning is now enabled for text-to-image workflows, giving users the flexibility to generate multiple variations from the same flow.

Release Overview [1.0.3]

CNAPS Studio v1.0.3 is a feature-packed release bringing a new AI image generation model (DeepGen-1.0), GPT-5.4 and GPT-5.4 Pro to our ChatGPT external model lineup, and a new Video Compare node. This release also introduces Copy, Paste & Undo/Redo support for faster workflow building, a redesigned Run Flow and Batch Run button with animated states, and a 2-week free Pro trial for new users. Additional platform improvements and bug fixes included.

✨ What's New

New Features

1️⃣ Copy, Paste & Undo/Redo Support

Users can now select nodes on the canvas using Shift + Drag, copy and paste them into new workflows with Ctrl+C / Ctrl+V, and undo or redo previous actions with Ctrl+Z / Ctrl+Shift+Z. A shortcut key guide is displayed on the canvas for quick reference, making workflow building faster and more flexible than ever.

image

2️⃣ 2-Week Free Pro Trial

New users can now try CNAPS Studio's Pro plan free for 2 weeks. Experience the full power of Pro features before deciding on a plan.

New AI Models & Tools

1️⃣ New AI Models - Image Generation (DeepGen-1.0)

DeepGen-1.0 is a lightweight yet powerful unified multimodal image generation and editing model built on a hybrid VLM + Diffusion Transformer architecture. At just 5B parameters, it integrates five core capabilities in a single model, general image generation, image editing, reasoning image generation, reasoning image editing, and text rendering, while remaining competitive with or surpassing models up to 16× larger. It's ideal for design, content creation, and multimodal workflows that require both generation and editing in one streamlined pipeline.

image

2️⃣ GPT-5.4 & GPT-5.4 Pro Added to ChatGPT External Models

GPT-5.4 and GPT-5.4 Pro are now available under the ChatGPT external model lineup. GPT-5.4 delivers best-in-class intelligence for agentic, coding, and professional workflows, while GPT-5.4 Pro produces smarter and more precise responses. Connect your OpenAI API key to access these latest models directly within CNAPS Studio.

Platform Improvements

1️⃣ Scene Segmentation (MaskFormer) - Distinct Color Output Fixed

Previously, segmented objects were displayed in visually similar colors, making it difficult to distinguish between different detected classes. Each segmented object is now assigned a clearly distinct color, making results much easier to read and interpret.

2️⃣ Run Flow & Batch Run Buttons - New Design & Animated States

Both the Run Flow and Batch Run buttons have been redesigned with a fresh look. The Run Flow button features 5 distinct animation states; Default, Hover, Running, Done, and Failed, while the Batch Run button introduces 2 animated states; Default and Hover, giving users clearer visual feedback throughout their workflow experience.

🔧 Bug Fixes

  • Fixed a few bugs in image resizing
    • SAM3 Max Resolution Incorrect - The maximum resolution for SAM3 was incorrectly set to 3072 instead of 2160. Fixed.
    • Input Image Size Limit Removed - The Image Resize node previously capped input images at 4096×4096 and threw an error for anything larger. The size limit has been removed and images of any resolution are now accepted.
    • Auto Resize Node Not Inserting on Existing Connections - When a high-resolution image was loaded into an already-connected node, the Image Resize node was not being automatically inserted as expected. Fixed.
  • Workspace Invitation - False Error Message

When accepting a workspace invitation, an error message was incorrectly displayed even though the invitation was successfully accepted. Users can now accept workspace invitations without seeing any misleading error messages. Fixed

  • Image Rotation Metadata Lost After Processing

When input images contained rotation metadata (EXIF orientation), the rotation was not being preserved after running through an AI model, causing output images to appear incorrectly rotated. Image rotation is now correctly maintained throughout the processing pipeline. Fixed

  • Workspace Scroll Zoom - Accidental Screen Resize Fixed

Based on user feedback, scrolling with Ctrl held down inside certain areas of the workspace (such as model configuration panels) was accidentally triggering browser-level screen resize. Ctrl+Scroll inside model nodes no longer resizes the browser screen, preventing any unintended zoom behavior.

Release Overview [1.0.2]

CNAPS Studio v1.0.2 delivers a major infrastructure upgrade that reduces workflow processing times by up to 98%, making this one of our most impactful performance releases to date. This release also introduces Korean language support, activates Video Input/Output nodes as a preview of exciting things to come, and includes key improvements to Template Flow and the Object Removal model.

✨ What's New

New Features

1️⃣ Multi-Language Support - Korean Added

CNAPS Studio now supports Korean (한국어) alongside English. Users can change their language preference in Account Settings under Profile > Language. More languages are on the way soon!

image

2️⃣ Video Input / Output Nodes

We've activated Video Input and Output nodes in CNAPS Studio, now ready to use with our newly added video generation models (Sora 2 and VEO).

image

A new Video Compare node is also available under the Comparison category, allowing users to compare video outputs side by side directly within their workflows.

image

New AI Models & Tools

New External AI Models - Nano Banana 2, Sora 2 & VEO

Expanded our External Models lineup with three new additions. Connect your API keys to access Nano Banana 2, OpenAI's Sora 2, and Google's VEO directly within CNAPS Studio. Use Sora 2 and VEO with the new Video Output node to generate stunning videos directly within your workflows.

Platform Improvements

1️⃣ Workflow Processing Performance Improvement

We've made a major upgrade to our AI server infrastructure; one of the biggest performance improvements we've shipped to date. By dedicating specific AI servers to individual models, models are kept pre-loaded and ready to run, eliminating startup overhead and allowing inference to begin immediately upon request. As a result, workflow processing times are reduced by up to 98%, making your runs significantly faster than before!

2️⃣ Template Flow - Reset Output Nodes on Run

When opening a template and clicking "Run Flow," nothing appeared to happen on the canvas because all nodes were already in a completed state, leaving new users confused. Now, output nodes are cleared when "Run Flow" is clicked on a template (while keeping the input image), so users can watch the flow execute in real time and better understand how each workflow works.

🔧 Bug Fixes

  • Object Removal (LatentDiffusion) - Tile Boundary Mismatch Fixed

Previously, when input images exceeded 512×512 pixels, the model would split them into 512×512 tiles and run inpainting on each tile separately. This caused visible seams/mismatches at tile boundaries when the results were merged. The fix changes the approach to extract only the masked regions and run inpainting on each region individually, eliminating the boundary mismatch issue.

Release Overview [1.0.1]

CNAPS Studio v1.0.1 brings OpenAI's GPT Image to External Models, a new motion deblurring AI model (DeblurDiff), and two ready-to-use workflow templates for logo detection and object removal. This release also introduces the Account Settings page for managing your profile, security, and activity, plus a new FAQ section in Docs. Additional UX improvements and bug fixes included.

✨ What's New

New Features

⚙️ Account Settings - New User Management Hub

Introduced a comprehensive Account Settings page accessible from the user menu. Manage your account profile, security, and preferences in one place:

  • Profile: Update your profile picture, name, and bio
  • Security: Change password and email settings (for non-Google accounts)
  • Activity: View your sign-in history including location, device, IP address, and timestamp
  • Account: Permanently delete your account

Access Account Settings from the user menu alongside Billing, Usage, and Docs.

New AI Models & Tools

🔧 GPT Image - OpenAI's Multimodal Image Generation

Added GPT Image to External Models, enabling OpenAI's natively multimodal image generation and editing capabilities directly within CNAPS Studio. Connect your OpenAI API key to access GPT Image models for generating and editing images using natural language prompts. Configure model version and output image size through the Advanced settings panel.

image

Image Restoration - Motion Deblur Removal (DeblurDiff)

DeblurDiff removes motion blur from photographs using generative diffusion models with Stable Diffusion priors. The model uses a Latent Kernel Prediction Network (LKPN) for robust real-world deblurring with iterative refinement, achieving state-of-the-art results on both benchmark and real-world images.

Platform Improvements

1️⃣ New Workflow Templates - Logo Detection & Object Removal

Added two new ready-to-use workflow templates to the Template Library:

Logo Detection & Extraction: Detects and extracts specific brand logos from images using AI-powered segmentation guided by simple text prompts. Just upload an image, specify the logo name (e.g., "Nike Logo", "Adidas"), and the Meta Segment Anything Model (SAM) automatically isolates and extracts matching logos.

image

Seamless Object Removal: Completely removes unwanted objects from photos and intelligently reconstructs the background. Use the Image Masker tool to paint over vehicles, people, background clutter, or any distracting elements – the AI erases them as if they were never there.

image

Click "Use this template" to fork and customize these workflows for your own projects.

2️⃣ Auto Resize Node Recommendation

: Added intelligent resolution limit detection that warns users when input images exceed an AI model's maximum allowed resolution. When connecting an oversized image to a model, a "Resolution Exceeds Limit" dialog appears showing the current resolution, maximum allowed, and recommended size. Users can click "Insert Resize Node" to automatically add an Image Resize node between the source and target nodes, or choose to skip the warning.

image

3️⃣ Docs - FAQ Section Added

: Added a comprehensive FAQ section to the Docs menu covering common questions about CNAPS Studio

4️⃣ Run Flow Button - Onboarding Prompt for New Users

: Added a helpful prompt that appears when new users open a pre-built flow template for the first time. The prompt highlights the "Run Flow" button with a "is ready" message and "Run Now" option, guiding users to start their first workflow easily.

image

🔧 Bug Fixes

  • ChatGPT API Integration - Connection Errors Fixed

: Users were experiencing errors when connecting their OpenAI API keys for ChatGPT models, including "This is not a chat model" and "Unsupported parameter: 'max_tokens'" messages. Fixed API compatibility issues to ensure smooth connection with ChatGPT models

  • Batch Run - Auto-Broadcast for Mismatched Inputs

: Fixed an issue where Batch Run would fail when workflows had multiple input types with different quantities (e.g., 4 images but only 1 text input). The system now automatically broadcasts single inputs to match the batch size, so one text input will be applied to all images in the batch

  • Batch Run - Flow Name Image Fallback

: Fixed an issue where the Batch Runs list displayed a broken image icon next to flow names when no thumbnail image was available or when image loading failed. Added proper fallback handling for missing images.

  • Image Resize Node - Preserved Settings on Flow Load

: Fixed a bug where the Image Resize node's width/height values were being overwritten with source image dimensions when loading flows or templates, causing saved size settings to be lost. Auto-fill now only triggers when a new connection is made, not when loading existing flows. Saved resize values are preserved when loading flows, while new node connections still auto-fill with source image dimensions as expected.

  • AI Flow Library - Search Not Working

: Fixed a bug where search functionality in the AI Flow Library was not returning results. Search now works as expected.

Release Overview [1.0.0]

CNAPS Studio v1.0.0 is officially here!

Our first production-ready release includes Batch Run for processing hundreds of items at once, external AI model support (Claude, OpenAI, Gemini, Nano Banana) for maximum AI provider flexibility, advanced image segmentation with SAM3, and UX improvements that make building AI workflows faster and more intuitive. Built for scale, designed for ease.

✨ What's New

New Features

1️⃣ Batch Runs - Process Multiple Items at Once

New Batch Run feature enables processing multiple input files or data items in a single operation without building separate flows. Build your flow once, then use it repeatedly for batch processing.

Once you build a base flow, click the 'Batch Run' button, add multiple input items, and Run Batch to start processing all items automatically. Track progress in the Batch Runs menu in real-time, and once complete, check and download all output results.

  • Key capabilities:
    • Process hundreds of items with one click (no repetitive manual runs)
    • Real-time progress tracking without leaving the platform
    • All results organized in one place for easy download
    • Perfect for bulk image processing, batch content analysis, and overnight workflows
    • image
      image
      image
icon

Batch Run limits by plan:

Free
Starter
Pro / Business
: not available
: limited batch runs
: unlimited batch runs

2️⃣ New & Hot Model Indicators - Discover Models Faster

Added visual indicators to help users discover the latest and most-popular AI models:

  • New Icon: Models added within the last 30 days display a 'new' icon, making it easy to spot the latest additions to the platform.
  • Hot Models Category: New category that highlights the top 5 most-used AI models based on actual customer usage patterns. Each model in this category displays a 'hot' icon. Perfect for finding proven models that solve real problems.
  • image

    Both indicators appear directly in the Model Library, making discovery intuitive.

New AI Models & Tools

🔧 External AI Model Integration - Bring Your Own Providers:

Users can now connect preferred Large Language Model providers (Claude, OpenAI, Gemini) and specialty image / video AI model platforms like Nano Banana directly to CNAPS Studio using their own API keys.

Combine different models in the same workflow for maximum flexibility, cost optimization, and no vendor lock-in. Pair LLMs with open-source image models to build customized, powerful workflows.

image
Scene Segmentation (Meta’s SAM3): SAM3 (Segment Anything Model 3) is Meta's advanced segmentation model that understands natural language to automatically detect and extract all instances of any concept in images and videos - simply describe what you want to segment.

Platform Improvements

1️⃣ Active Node Indicator - Glowing Animation

: Added animated glowing effect around model nodes as they process. Users can now easily see which step is running without having to scan the entire flow

image

2️⃣ Share Menu - Make Public & Copy URL Together

: Simplified sharing workflow. The Share menu now includes the option to change a workflow from private to public, eliminating the need to visit the flow library separately before copying a share link. 3️⃣ Image Resize Tool - Preview Input & Result Dimensions

: Improved Image Resize configuration menu to display the original input image size and real-time preview of the result size as you adjust target dimensions. Users no longer have to guess what the output will be, especially when aspect ratio locking is enabled.

image

4️⃣ Multi-Language Text Recognition (DeepSeekOCR) - Formatted Table Output : Improved text recognition table output to display as formatted tables with visible borders and grid lines, instead of raw HTML tags. Tables are now much easier to read and work with.

5️⃣ Multi-Class Support for Color Pick, Crop & Masker Tools : Enhanced Color Pick by Class, Image Crop by Class, and Image Masker by Class to support multiple class selection. Users can now process multiple object classes in a single workflow step, eliminating the need to create separate flows for each class.

6️⃣ Multi-Class Support for Color Pick, Crop & Masker Tools : Enhanced Color Pick by Class, Image Crop by Class, and Image Masker by Class to support multiple class selection. Users can now process multiple object classes in a single workflow step, eliminating the need to create separate flows for each class.

image

7️⃣ Input Image Node - Preview Button Added : Added '+' button to Input Image nodes to preview uploaded images.

image

8️⃣ Workflow Notifications - Less Spam, More Useful : Adjusted email notification threshold from 30 seconds to 120 seconds (2 minutes) based on customer feedback. Users still get notified when workflows take a long time, but with fewer emails for quick-running processes.

9️⃣ Random Seed Options - Generate & Auto-Generate Buttons : Added "Generate Random Seed" button and "Auto Random Generation on Each Run" toggle to the Random Seed configuration. Users no longer need to manually enter seed values – they can generate them with a click or enable automatic generation for each workflow run.

image

🔟 UI Text - More Accessible for Non-Technical Users : We continue to update UI texts throughout CNAPS Studio to be more accessible and friendly for business users and non-developers. Technical jargon replaced with clearer, action-oriented language to improve usability.

🔧 Bug Fixes

  • Segmentation Models - Missing Parameter Data: Segmentation models were missing parameter data during execution, preventing models from functioning correctly. Fixed.

Release Overview [0.9.14]

v0.9.14 introduces the new CNAPS Studio (formerly AI Flow Studio) with 5 AI models, including SOTA QWEN’s Smart Image Editing, Meta’s Scene Segmentation (SAM2), and DeepSeek Text Recognition (OCR) models, and important bug fixes. Updated documentation reflects the rebrand across all guides and templates.

✨ What's New

New AI Models

1️⃣ Smart Image Editing (QWEN 2509)

Smart Image Editing (QWEN 2509) is a vision-language diffusion model that performs image editing based on natural language instructions.

2️⃣ Scene Segmentation (Meta’s SAM2)

SAM2 (Segment Anything Model 2) is Meta's powerful segmentation model that automatically detects and extracts every object in an image without any manual interaction.

3️⃣ Multi-Language Text Recognition (DeepSeek OCR)

Extracts text in multiple languages from images with markdown conversion and text grounding. Vision-language model for document layout understanding.

4️⃣ Face Detector (DETR)

Detects and localizes human faces in images using the DETR (Detection Transformer) model.

5️⃣ License Plate Detector(DETR)

Detects and localizes license plates in vehicle images using Detection Transformer.

Platform Improvements

1️⃣ UI Text - More Accessible for Non-Technical Users:

Updated category names, model names, and general messages throughout CNAPS Studio to be more accessible and friendly for business users and non-developers. Technical jargon replaced with clearer, action-oriented language to improve usability

2️⃣ Image Compare Tool - Extended Zoom to 3000%:

Increased maximum zoom resolution from 2000% to 3000% in the Image Compare tool for improved pixel-level detail inspection.

image

3️⃣ Inference Completion Sound Alert:

Workflows now play a sound notification when inference completes, alerting users without requiring them to monitor the screen.

4️⃣ Configuration Parameters - Removed Duplicate Labels:

Removed duplicate configuration text labels that appeared on models and tools. Configuration panels are now cleaner and less cluttered.

5️⃣ Image Resize Tool - Improved Error Messaging:

Fixed confusing error message in Image Resize tool that displayed image dimensions in the wrong order. Error messages now show 'input image resolution not matched' for clarity.

6️⃣ Connecting Nodes - introducing empty/unfilled dots for optional connections:

The Image Compare node has 4 empty blue dots, meaning you can connect 1, 2, 3, or 4 input images to compare. Other nodes typically have filled dots, indicating required connections.

image

7️⃣ Text Output - Added Markdown Download Format: Added ability to copy text results as .markdown files alongside the existing .txt format option.

image

🔧 Bug Fixes

  • PiSA-SR Input Resolution Crash: PiSA-SR node would crash when input images were smaller than 256x256 pixels, preventing users from upscaling thumbnails and low-resolution photos. Fixed.
  • Object Detection & Segmentation - Text Output File Format: Object Detection and Segmentation models output both images and text results. Text Output files were incorrectly being saved as .jpg format, rendering them unreadable. Fixed.
  • Conditional Node Logic Error: Conditional nodes were incorrectly proceeding to the next workflow step even when the condition evaluated as false. This broke workflow logic and caused unintended processing. Fixed.
  • SegFormer_b3 'random' Module Error: SegFormer_b3 model would fail with error message 'name 'random' is not defined' because the Python random module was being used without being imported. Fixed by adding the required import statement. Fixed.