AI Model Hub
📖

AI Model Hub

Welcome! This page helps you find the right documentation for each AI model in CNAPS Studio.

🔗 Open-Source Models

🎵 Audio Models

Text to Speech (Qwen3)

Turn text into a natural spoken voiceover

🎙️ Audio Understanding

Speech Recognition (Clips)

Transcribe Every Clip in a Video List — Timed Transcripts

🎙️ Audio Understanding

Speech Recognition (Video)

Transcribe Speech in a Video into Timed Text

🏷️ Image Classification

Adult Content Detection

Flag Explicit Images Automatically

🏷️ Image Classification

Gender Recognition

Classify Pedestrian Gender in Photos

🏷️ Image Classification

Object Classification

Identify What's in a Photo (1,000 Categories)

🎛️ Image Control

ControlNet XL Canny

Generate Images That Follow an Edge Map

🎛️ Image Control

ControlNet XL Union

Generate Images from Any Control Map — 6 Modes in One Model

🎛️ Image Control

Depth Anything Annotator

Turn Any Photo into a Depth Map for ControlNet

🎨 Image Edit

FireRed-1.1

Edit with Text Instructions — Best-in-Class Identity Preservation

🎨 Image Edit

LatentDiffusion (Object Removal)

Remove People, Objects, or Blemishes Seamlessly

🎨 Image Edit

QWEN-Image-Edit-2511

Edit Images with Natural-Language Instructions

🎨 Image Edit

QWEN-Inpaint

Edit One Region, Leave the Rest Untouched

🎨 Image Edit

QWEN-Layered

Split a Flat Image into Editable Photoshop-Style Layers

🚀 Image Generation

DeepGen-1.0

Generate & Edit in One Lightweight Model

🚀 Image Generation

FLUX Schnell

High-Quality Images in 1–4 Steps

🚀 Image Generation

FLUX.2 KLEIN 4B

Compact Generation with Multi-Reference Editing

🚀 Image Generation

GLM-Image

Generate Posters & Thumbnails with Readable Text

🚀 Image Generation

OmniGen2

Generate, Edit & Remix Images in One Model

🚀 Image Generation

SANA-Sprint 1.6B

Ultra-Fast 1024px Images in a Single Step

🚀 Image Generation

Z-Image-Turbo

The Fastest Route to Photorealistic Images

🔧 Image Restoration

Image Colorization

Bring Black-and-White Photos to Life in Color

🔧 Image Restoration

Image Denoiser

Remove Noise, Grain & Speckle from Photos

🔧 Image Restoration

JPEG Quality Restoration

Clean Up Blocky, Over-Compressed JPEGs

🔧 Image Restoration

Motion Blur Removal (MSSNet)

Fix Motion Blur — Handles Uneven Blur Across the Frame

🔧 Image Restoration

Real-World Blur Removal

Fix Motion Blur — Highest Quality, Diffusion-Based

📝 Image Understanding

BLIP (Image Description)

Describe Any Image in Natural Language

📝 Image Understanding

ViLT (Image Q&A)

Ask Questions About an Image, Get Answers

📸 Image Upscaling

LatticeNet

Upscale Efficiently — Best Choice for Large (2K+) Inputs

📸 Image Upscaling

PiSA-SR

The Sharpest Results, with Adjustable Enhancement

📸 Image Upscaling

Swin2SR

Balanced Quality & Speed for Everyday Upscaling

📸 Image Upscaling

SwinIR

The Widest Zoom Range — Up to 8× Enlargement

🧠 Multimodal Language Models

Gemma 4 31B

Flagship Vision-Language Reasoning with 256K Context

🧠 Multimodal Language Models

Gemma 4 E2B

Compact Vision-Language Model for Light Workloads

🧠 Multimodal Language Models

Molmo2-8B

Open Vision-Language Model for Captioning & Visual Q&A

🧠 Multimodal Language Models

Qwen3.6-35B-A3B

Coding, Reasoning & Vision with Built-In Thinking

👁️ Object Detection

DETR

Find Everyday Objects in Any Image (91 COCO Classes)

👁️ Object Detection

DETR (Face Detector)

Find Every Face in a Photo

👁️ Object Detection

DETR (License Plate Detector)

Find License Plates — Chain with Blur for Anonymization

👁️ Object Detection

RF-DETR Medium

Real-Time Detection for Speed-Critical Flows

👁️ Object Detection

YOLO-S

Lightweight, Fast Detection (91 COCO Classes)

👁️ Object Detection

YOLO-S (Fashion Item Detector)

Detect Clothing & Fashion Items (46 Categories)

🤸 Pose Estimation

RF-DETR Keypoint

Real-Time Human Pose — 17 Keypoints, OpenPose-Style Skeleton

🧩 Segmentation

BiRefNet (Subject)

High-Quality Background Removal & Subject Cutout

🧩 Segmentation

RF-DETR Seg Medium

Real-Time Instance Cutouts (80 COCO Classes)

🧩 Segmentation

SAM2 (Scene)

Segment Every Object Automatically — No Prompt Needed

🧩 Segmentation

SAM3 (Scene)

Segment Objects by Describing Them in Text

🧩 Segmentation

SAM3.1 (Scene)

Text-Prompted Segmentation — Higher-Accuracy SAM3 Upgrade

📄 Text Recognition(OCR)

DeepSeekOCR

Read Multilingual Documents, Layouts & Handwriting

📄 Text Recognition(OCR)

GLM-OCR

Full Document Understanding — Layout + Text Extraction

📄 Text Recognition(OCR)

PaddleOCR

Fast Korean & English Text Extraction

🎬 Video Generation

Cosmos3 Nano

Text/Image-to-Video with Optional Synchronized Audio

🎬 Video Generation

Helios-Base

Highest-Quality Video Clips — Up to 60 Seconds

🎬 Video Generation

Helios-Distilled

2–3× Faster Helios for Quick Iterations

🎬 Video Generation

Wan2.2 TI2V 5B

Text- & Image-to-Video at 720p on a Single GPU

🎞️ Video Segmentation

SAM3.1 (Video)

Track & Segment Objects Across a Video by Text Prompt

📺 Video Upscaling

FlashVSR-v.1.1

Real-Time 4× Streaming Video Upscaling

📺 Video Upscaling

SeedVR2 3B

Upscale & Restore Video up to 4K

📺 Video Upscaling

SeedVR2 7B

Upscale & Restore Video up to 4K — Highest-Quality Variant

📺 Video Upscaling

SparkVSR

Reference-Guided Upscaling for Short Clips

🔗 External Models

🎵 Audio Model

Gemini TTS

Multimodal AI for text generation and understanding

📄 Language Model

ChatGPT

Multimodal AI for text generation and understanding

📄 Language Model

Claude

Multimodal AI for text generation and understanding

📄 Language Model

Gemini

Multimodal AI for text generation and understanding

🚀 Image Model

GPT Image

Generate high-quality images from text prompts

🚀 Image Model

Nano Banana

Generate high-quality images from text prompts

🎬 Video Model

Sora 2

Generate high-quality videos from text prompts

🎬 Video Model

Veo 3.1

Generate high-quality videos from text prompts

🎬 Video Model

Gemini Video Analysis

Video Understanding & Analysis

🛠️ Tools

Image Editing & Preparation Tools (Non-AI)

Category
Tool
What It Does
Annotation
Turn any image into a clean white-edges-on-black edge map
Blur & Effects
Quick blur with minimal processing for fast results
Blur & Effects
The safest and simplest form of Blur in the family
Blur & Effects
Apply blur effects to images for various intensity levels
Blur & Effects
Camera-Style Bokeh - Circular Disc Blur, Not Gaussian
Blur & Effects
Apply a clean, region-specific blur to any part of an image
Color Adjustments
Convert color images to grayscale
Color Adjustments
Invert colors in images for creative effects
Comparison
Compare up to 4 images side-by-side for quality or differences
Comparison
Compare up to 4 texts side-by-side for quality or differences
Comparison
Compare up to 4 videos side-by-side for quality or differences
Image Blending
Blend images by adding pixel values together
Image Blending
Blend images by multiplying pixel values for overlay effects
Image Masking
Create masks by selecting its assigned class from Object Detection
Image Masking
Create masks by selecting its bounding box coordinates from Object Detection
Image Masking
Create masks and select regions for targeted processing
Image Masking
Select objects for targeted processing
Image Resizing
Auto-upscale small images to the next standard resolution tier
Image Resizing
Resize images to specific dimensions or aspect ratios
Image Resizing
Resize one image to match another's exact dimensions
Layered Tools
Stack Multiple Images Into One - Photoshop-Style Layer Compositing
Selection & Extraction
Extract segmented region by selecting its assigned class from Segmentation
Selection & Extraction
Extract segmented region by RGB color match
Selection & Extraction
Extract segmented region by selecting its color from Segmentation
Selection & Extraction
Crop an image by selecting its assigned class from Object Detection
Selection & Extraction
Crop an image by its bounding box coordinates from Object Detection
Selection & Extraction
Pass through the input image only when specified text matches
Selection & Extraction
Select the first available image from multiple inputs for downstream
Selection & Extraction
Pick One Image Out of an Image List — By Index
Selection & Extraction
Put One Image Back Into a List — Replace by Index
Selection & Extraction
Manually crop any image by drawing the keep-region directly on it
Selection & Extraction
Pull a Single Frame Out of a Video as an Image
Selection & Extraction
Turn an Image List Back Into a Video
Selection & Extraction
Pick One Video Out of a Video List — By Index
Selection & Extraction
Turn a Video Into an Image List — with Sampling and Frame-Count Controls
Selection & Extraction
Cut a video into highlight clips driven by a Video Analysis JSON
Text
Place text anywhere on an image
Text
Merge Up to 4 Text Inputs Into One
Text
Reusable prompt templates — fill the blank with incoming text
Video Editing
Broadcast-standard loudness normalization (EBU R128, -14 LUFS / -1.5 dBTP) + true-peak limiter — publish-ready audio in one node
Video Editing
Prepend a short preview of each clip's best moment to the front — the classic cold-open attention trick
Video Editing
Freeze the first frame while a narration audio plays over it, then join to the front of the clip — raw voiceover cold-open muxer
Video Editing
Auto jump-cut editor — trim silent gaps out of every clip for snappy talking-head pacing
Video Editing
Nudge Video Analysis segment cuts to the nearest word boundary — no more mid-word clips
Video Editing
Burn a static Title or Hook overlay onto each clip — six named look presets, three positions
Video Editing
Trim each clip to its longest continuous shot — removes mid-clip scene jumps automatically
Video Editing
Glue every clip in a list into a single continuous video, auto-normalizing resolution and frame rate
Video Editing
AI-voiced cold-open narration for every clip — reads the clip's Hook or Title via Qwen3 TTS (EN / KO / auto)
Video Editing
Reframe every clip to 9:16 vertical (1080×1920) — four strategies incl. Subject-Aware punch-in for Reels / TikTok / Shorts
Video Editing
Auto-mix emotion-matched sound effects timed to each clip's peak moment — reads Video Analysis EDL metadata or falls back to keyword scan
Video Editing
Retime every clip in a batch by a single speed factor — audio pitch preserved (no chipmunk voices)
Video Editing
Burn karaoke-style word-synced subtitles onto every clip — seven presets, five fonts (CJK-ready), optional bounce animation

Last updated: Jul 27th, 2026

Full Model List