Professional-Grade Image Restoration
Transforms small, blurry, or low-resolution photos into larger, sharper versions. Works on screenshots, thumbnails, old photos, and any degraded images. 2× faster training convergence than SwinIR with improved stability.
What it does
Swin2SR makes photos bigger and sharper. Whether you need to enlarge a small thumbnail, upscale a screenshot, or improve an old archived photo—Swin2SR restores quality while increasing size. Choose your upscaling factor (2× or 4×) based on your needs. The advanced Swin Transformer V2 technology delivers professional results.
Problem it solves
- Small images need to be bigger – Thumbnails, screenshots, small photos
- Low-resolution photos – Downloaded images, old archives, poor quality sources
- Blurry or degraded images – Need sharpening and enhancement
- Flexible upscaling needs – Different scale factors for different tasks
- Fast, reliable restoration – 2× faster training than competing methods
- Professional-grade results – Business-quality photo enlargement
- Video frame enhancement – Upscale individual frames from streams/downloads
Input/Output
- Input: Any RGB photo (small, blurry, or low-resolution)
- Output: Sharp, clear, high-resolution version (bigger and clearer) (2x, 4x bigger)
Performance
- Performance (Classical Upscaling, ×4 factor)
Dataset | PSNR (Quality) | Notes |
Set5 | 32.93 dB | Excellent quality |
Set14 | 29.10 dB | Very good quality |
BSD100 | 27.72 dB | Good quality |
Urban100 | 26.89 dB | Good on complex scenes |
- Quality by Scale Factor:
- 2× upscaling: Excellent quality (minimal artifacts)
- 4× upscaling: Good quality (balanced)
Technical Details
Architecture | Swin Transformer V2 (Hierarchical Vision Transformer) |
Key Innovations | Continuous relative position bias enables handling of resolution shifts, with normalized attention for training stability. |
Core Components | • Feature Extraction: Initial convolution layer
• Deep Processing: 6 Residual Swin Transformer Blocks with shifted window attention
• Upsampling: Pixel shuffle layers
• Dimensions: 180 embedding channels, 6 attention heads |
Key Features | • 12M parameters: Efficient and fast
• Multi-scale support: Single model handles 2×, 3×, 4×, 8× upscaling
• Flexible input: Any image size
• Real-time capable: Fast GPU inference |
Training | • Dataset: DIV2K, Flickr2K, BSD500, WED (8,194 images total)
• Loss: L1 + perceptual loss
• Training time: 2-3 days (33% faster than SwinIR)
• Framework: PyTorch |
V2 Improvements Over V1 | • Scaled cosine attention: More stable training
• Residual post-normalization: Better gradient flow
• Continuous position bias: Smooth resolution handling |
Compliance & Provenance
Provider | Open-source |
Provider type | Specialized |
License | Apache 2.0 |
EU AI Act risk class | Minimal Risk |
Art. 50 transparency | Not applicable |
Region availability | Available globally |
Training data summary | Pending — provider has not yet published per Art. 53(d) |
For more on how we classify models and mark outputs, see our AI Policy.
Model Source
- GitHub Repository: https://github.com/mv-lab/swin2sr (official, ECCV 2022 AIM Workshop)
- Paper: Conde et al., "Swin2SR: SwinV2 Transformer for Compressed Image Super-Resolution and Restoration" ECCV 2022 Advances in Image Manipulation Workshop, Tel Aviv arXiv:2209.11345 https://arxiv.org/abs/2209.11345
- License: Apache 2.0