📸

Image Upscaling - Swin2SR

🛠

Professional-Grade Image Restoration

Transforms small, blurry, or low-resolution photos into larger, sharper versions. Works on screenshots, thumbnails, old photos, and any degraded images. 2× faster training convergence than SwinIR with improved stability.

What it does

Swin2SR makes photos bigger and sharper. Whether you need to enlarge a small thumbnail, upscale a screenshot, or improve an old archived photo—Swin2SR restores quality while increasing size. Choose your upscaling factor (2× or 4×) based on your needs. The advanced Swin Transformer V2 technology delivers professional results.

Problem it solves

  • Small images need to be bigger – Thumbnails, screenshots, small photos
  • Low-resolution photos – Downloaded images, old archives, poor quality sources
  • Blurry or degraded images – Need sharpening and enhancement
  • Flexible upscaling needs – Different scale factors for different tasks
  • Fast, reliable restoration – 2× faster training than competing methods
  • Professional-grade results – Business-quality photo enlargement
  • Video frame enhancement – Upscale individual frames from streams/downloads

Input/Output

  • Input: Any RGB photo (small, blurry, or low-resolution)
  • image
  • Output: Sharp, clear, high-resolution version (bigger and clearer) (2x, 4x bigger)
  • image

Performance

  • Performance (Classical Upscaling, ×4 factor)
Dataset
PSNR (Quality)
Notes
Set5
32.93 dB
Excellent quality
Set14
29.10 dB
Very good quality
BSD100
27.72 dB
Good quality
Urban100
26.89 dB
Good on complex scenes
  • Quality by Scale Factor:
    • 2× upscaling: Excellent quality (minimal artifacts)
    • 4× upscaling: Good quality (balanced)

Technical Details

Architecture
Swin Transformer V2 (Hierarchical Vision Transformer)
Key Innovations
Continuous relative position bias enables handling of resolution shifts, with normalized attention for training stability.
Core Components
• Feature Extraction: Initial convolution layer • Deep Processing: 6 Residual Swin Transformer Blocks with shifted window attention • Upsampling: Pixel shuffle layers • Dimensions: 180 embedding channels, 6 attention heads
Key Features
• 12M parameters: Efficient and fast • Multi-scale support: Single model handles 2×, 3×, 4×, 8× upscaling • Flexible input: Any image size • Real-time capable: Fast GPU inference
Training
• Dataset: DIV2K, Flickr2K, BSD500, WED (8,194 images total) • Loss: L1 + perceptual loss • Training time: 2-3 days (33% faster than SwinIR) • Framework: PyTorch
V2 Improvements Over V1
• Scaled cosine attention: More stable training • Residual post-normalization: Better gradient flow • Continuous position bias: Smooth resolution handling

Compliance & Provenance

Provider
Open-source
Provider type
Specialized
License
Apache 2.0
EU AI Act risk class
Minimal Risk
Art. 50 transparency
Not applicable
Region availability
Available globally
Training data summary
Pending — provider has not yet published per Art. 53(d)

For more on how we classify models and mark outputs, see our AI Policy.

Model Source