Skip to content

Repository files navigation

image-upscale

original result

Image Upscale

A compact C++ image upscaling experiment based on compressed-sensing-inspired optimization, Discrete Cosine Transform (DCT), and L-BFGS / OWL-QN optimization.

The project attempts to reconstruct a higher-resolution image from a lower-resolution input by optimizing a high-resolution DCT representation whose downsampled version matches the original image.

The implementation is inspired by the approach described in:

Compressed Sensing with Python

Overview

Conventional image upscaling interpolates existing pixels. This project takes a different approach:

  1. The input image is split into its BGR channels.
  2. Each channel is centered around its mean.
  3. A high-resolution DCT-domain representation is optimized.
  4. The inverse DCT produces a candidate high-resolution image.
  5. The candidate is constrained so that its local average corresponds to the original low-resolution pixels.
  6. L-BFGS with OWL-QN regularization is used to find a sparse/compact solution.
  7. A selective high-frequency sharpening stage enhances visible details.
  8. The reconstructed channels are combined into the final image.

The current implementation uses a fixed 2× scale factor.

Algorithm

For an input image channel

src : H × W

the optimizer operates on a DCT representation of size

(H × 2) × (W × 2)

The optimization variable x therefore contains:

H × W × 4

floating-point coefficients.

The inverse DCT transforms the optimization vector into a spatial-domain high-resolution image:

cv::idct(x2, Ax2);

For every source pixel, the corresponding 2 × 2 block in the reconstructed image is averaged:

        a b
        c d

average = (a + b + c + d) / 4

The difference between this average and the original pixel is used as the reconstruction error.

Conceptually, the objective is:

minimize || A x - b ||²

where:

  • x is the high-resolution DCT representation;
  • A represents inverse DCT followed by the 2×2 averaging/downsampling operation;
  • b is the original image channel.

The implementation also applies an OWL-QN/L1-style regularization term through:

param.orthantwise_c = param_c;

with:

const double param_c = 5;

This encourages a sparse representation in the optimization domain.

Reconstruction Constraint

The objective function reconstructs each low-resolution pixel from the corresponding high-resolution block.

A small reconstruction error is explicitly ignored:

if (std::abs(Ax) < 0.5)
    Ax = 0;

This effectively introduces a small tolerance around the original pixel values and prevents the optimizer from spending effort on insignificant differences.

The gradient is then computed by transforming the residual back into the DCT domain:

cv::dct(Ax2, AtAxb2);

followed by scaling by two.

This provides L-BFGS with the gradient required to optimize the high-resolution representation.

Optimization

The project uses L-BFGS through the liblbfgs library.

The optimizer is configured for OWL-QN:

lbfgs_parameter_t param;
lbfgs_parameter_init(&param);

param.orthantwise_c = param_c;
param.linesearch = LBFGS_LINESEARCH_BACKTRACKING;

The optimization starts with all coefficients initialized to 1:

for (int i = 0; i < numImgPixels; i++)
    x[i] = 1;

The main optimization call is:

lbfgs(
    numImgPixels,
    x,
    &fx,
    evaluate,
    progress,
    &src,
    &param
);

This is the computationally expensive part of the application.

Mean Removal

Before optimization, the mean intensity of each channel is removed:

const auto mean = cv::mean(src);
src -= mean;

After reconstruction, the mean is restored:

Xa += mean;

This separates the DC component from the optimization and lets the optimizer concentrate on the spatial structure and higher-frequency information.

Selective Sharpening

After the optimized image has been reconstructed, the project applies a Laplacian-like sharpening kernel:

 0 -1  0
-1  4 -1
 0 -1  0

implemented as:

const cv::Mat sharpening_kernel = (cv::Mat_<double>(3, 3)
    << 0, -1, 0,
       -1, 4, -1,
       0, -1, 0);

The resulting high-frequency component is used as a mask.

A triangular threshold is then applied:

cv::threshold(
    contrastMask,
    contrastMask,
    0,
    255,
    cv::THRESH_BINARY | cv::THRESH_TRIANGLE
);

Only sufficiently strong local detail is sharpened:

sharpened += Xa;
sharpened.copyTo(Xa, contrastMask);

This is preferable to applying uniform sharpening across the entire image because it limits amplification of relatively flat regions.

Color Processing

The image is processed channel-by-channel:

std::vector<cv::Mat> bgr;
split(img, bgr);

for (auto& src : bgr)
{
    ...
}

Each BGR channel independently goes through:

BGR channel
    │
    ├── convert to float
    ├── remove mean
    ├── optimize high-resolution DCT representation
    ├── inverse DCT
    ├── restore mean
    ├── selective sharpening
    └── convert to 8-bit

The processed channels are finally recombined:

cv::Mat dst;
merge(bgr, dst);

Scale Factor

The current implementation uses:

enum { SCALE = 2 };

Therefore the output dimensions are:

width  × 2
height × 2

For example:

640 × 480
    ↓
1280 × 960

The scale factor is currently compile-time fixed rather than being supplied as a command-line option.

Usage

The application accepts an optional input and output filename:

image-upscale.exe <input> [output]

For example:

image-upscale.exe original.png result.jpg

If no input filename is supplied, the program attempts to locate OpenCV's sample lena.jpg.

When an output filename is provided, the reconstructed image is written using:

cv::imwrite(argv[2], dst);

The program also displays the original and reconstructed images in OpenCV windows.

Press a key in the image window to finish execution.

Example

The repository contains example images demonstrating the result:

Original

Original

Upscaled result

Result

Dependencies

The project requires:

  • C++ compiler
  • Visual Studio / MSVC
  • OpenCV
  • liblbfgs

OpenCV is used for:

  • image loading and saving;
  • image representation;
  • DCT / inverse DCT;
  • channel splitting and merging;
  • filtering;
  • thresholding;
  • displaying results.

liblbfgs provides the numerical optimizer and OWL-QN functionality.

Build

The repository contains a Visual Studio project:

image-upscale.sln
image-upscale.vcxproj

The original project targets:

  • Visual Studio toolset v141;
  • Windows 10 SDK 10.0.18362.0;
  • Win32;
  • x64;
  • Debug;
  • Release.

Both Win32 and x64 configurations are provided.

The required OpenCV and L-BFGS include/library paths need to be configured for the local development environment.

Project Structure

image-upscale/
│
├── image-upscale.cpp
│   └── Main implementation
│
├── image-upscale.sln
│   └── Visual Studio solution
│
├── image-upscale.vcxproj
│   └── Visual Studio project
│
├── image-upscale.vcxproj.filters
│   └── Visual Studio file filters
│
├── original.png
│   └── Example source image
│
├── result.jpg
│   └── Example upscaled result
│
├── neco.jfif
├── neco_hi4.jpg
│   └── Additional example images
│
├── .gitattributes
└── .gitignore

Implementation Characteristics

This project is intentionally small. The entire algorithm is implemented in a single C++ source file.

Important characteristics include:

  • 2× image upscaling
  • DCT-domain optimization
  • L-BFGS optimization
  • OWL-QN regularization
  • Per-channel BGR processing
  • Mean/DC component separation
  • Reconstruction-error minimization
  • Selective Laplacian sharpening
  • OpenCV-based image I/O and processing
  • Native C++ implementation
  • Visual Studio project included

Computational Cost

The optimization problem grows with the number of output pixels.

For an input image of:

W × H

the optimization contains:

4 × W × H

variables because the output is 2× larger in both dimensions.

Consequently, processing large images can become computationally expensive and memory-intensive. The three color channels are also optimized independently.

This approach therefore trades computational complexity for potentially better reconstruction than simple interpolation.

Why DCT?

The Discrete Cosine Transform is useful for representing images as combinations of spatial frequencies.

Instead of directly optimizing every high-resolution pixel, this project optimizes a frequency-domain representation.

This makes it possible for the optimizer to find a high-resolution image containing frequency components that are not explicitly present in the original low-resolution pixel grid, while still enforcing consistency with the observed pixels.

The approach is therefore closer to model-based reconstruction than conventional interpolation.

Relationship to Super-Resolution

This project can be viewed as a small experimental super-resolution implementation, but it is important to distinguish it from modern learned super-resolution systems.

It does not use:

  • neural networks;
  • CNNs;
  • transformers;
  • pretrained super-resolution models;
  • GANs;
  • diffusion models;
  • learned image priors.

Instead, the prior is implicit in the optimization formulation and DCT representation, with additional sparsity and sharpening constraints.

Limitations

The current implementation is primarily an experimental demonstration.

Notable limitations include:

  • fixed 2× scaling;
  • no configurable optimization parameters from the command line;
  • no batch processing;
  • CPU-only optimization;
  • potentially high computational cost;
  • no explicit noise model;
  • no automatic degradation/blur estimation;
  • no edge-aware reconstruction constraint;
  • no perceptual loss;
  • no learned image prior;
  • limited error reporting from the L-BFGS result;
  • sharpening parameters are hard-coded.

The optimizer return value is currently captured:

int lbfgs_ret = lbfgs(...);

but is not subsequently used to distinguish successful convergence from termination due to iteration limits or other optimizer conditions.

Possible Improvements

Several directions could make the project more practical.

Configurable Scale

Replace:

enum { SCALE = 2 };

with a runtime parameter.

For example:

image-upscale input.png output.png 2

or:

image-upscale input.png output.png --scale 2

Configurable Regularization

Expose:

param.orthantwise_c

as a command-line parameter.

Different images may benefit from different regularization strengths.

Better Initialization

Instead of initializing every DCT coefficient to 1, initialize the high-resolution image using a conventional interpolation method and transform that image into the DCT domain.

This could give the optimizer a substantially better starting point.

Improved Reconstruction Model

The current forward model effectively performs:

high-resolution image
        ↓
2×2 block averaging
        ↓
low-resolution image

A more realistic degradation model could include:

high-resolution image
        ↓
blur / PSF
        ↓
sampling
        ↓
low-resolution image

This would make the method more closely resemble a classical inverse-problem formulation.

Better Regularization

Possible additional priors include:

  • total variation;
  • gradient sparsity;
  • wavelet sparsity;
  • edge-preserving regularization;
  • Huber penalties;
  • learned priors.

GPU Acceleration

The objective function performs many DCT and per-pixel operations. These operations could potentially be accelerated using:

  • CUDA;
  • OpenCL;
  • OpenCV CUDA;
  • FFT/DCT libraries;
  • parallel CPU implementations.

Technical Summary

The core reconstruction pipeline can be summarized as:

             Low-resolution image
                      │
                      ▼
              Split BGR channels
                      │
                      ▼
                Remove mean
                      │
                      ▼
          High-resolution DCT variables
                      │
                      ▼
              L-BFGS / OWL-QN
                      │
             ┌────────┴────────┐
             │                 │
             ▼                 ▼
        inverse DCT       reconstruction
             │                 │
             └────────┬────────┘
                      ▼
                 residual
                      │
                      ▼
                DCT gradient
                      │
                      └──────► optimizer
                                  │
                           until convergence
                                  │
                                  ▼
                       reconstructed image
                                  │
                                  ▼
                       selective sharpening
                                  │
                                  ▼
                         restore channel mean
                                  │
                                  ▼
                           merge BGR channels
                                  │
                                  ▼
                            Output image

Credits

The implementation was inspired by the compressed-sensing image reconstruction approach described by PyRunner:

Compressed Sensing with Python


Image Upscale is a compact experiment in applying numerical optimization and frequency-domain representations to image reconstruction, demonstrating how classical optimization techniques can be used as an alternative to conventional interpolation and modern neural super-resolution methods.

Releases

Packages

Used by

Contributors

Languages