Skip to content
PolyU-VCLabPublic

About

HarnessIR: Harnessing Multimodal Foundation Models for Universal Real-World Image Restoration

Resources

Stars

12 stars

Watchers

0 watching

Forks

Latest commit

ย 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

HarnessIR: Harnessing Multimodal Foundation Models for Universal Real-World Image Restoration

Restoration by harnessing an MFM executor, not by scheduling restoration tools.

Paper HarnessIR-HuggingFace HarnessIR-BaiduDisk ProjectPage

Xiangtao Kong1,2 | Shuaizheng Liu1,2 | Rongyuan Wu1,2 | Lingchen Sun1,2 | Zhengqiang Zhang1,2 | Jinxin Zhao1,2 | Yuhui Wu1,2 | Lei Zhang1,2,โ€ 

1 The Hong Kong Polytechnic University
2 OPPO Research Institute

โ€  Corresponding author.

HarnessIR teaser

๐Ÿ“Œ Quick Links

๐Ÿ“ฐ News

  • 2026-10-07: Released the code, benchmark and visual results of HarnessIR.

๐Ÿงฐ HarnessIR

HarnessIR pipeline

Click to expand the pipeline description

Prior agentic IR methods apply task-specific restoration models in sequence, and are therefore bounded by those tools. HarnessIR keeps one MFM as the executor and varies only the prompt it is given. Five stages:

  1. Perception and diagnosis โ€” a VLM reads the content and its degradations, and names which auxiliary tools to consult.
  2. On-demand tool invocation โ€” only the planned tools run, returning image-specific evidence.
  3. Prompt composition โ€” the evidence is composed into one prompt stating what to treat and what must survive untouched.
  4. Execution โ€” the MFM receives the original LQ image and that prompt, in a single pass.
  5. Verification-driven refinement โ€” the result is judged against the restoration requirements, and re-executed from the original LQ input only when it fails.

Two choices distinguish it from prior pipelines: no restoration tool is scheduled (the auxiliary tools supply evidence for writing the prompt, not links in a restoration chain), and the executor sees only the image and text (diagnostic maps inform the prompt writer but are never passed to the editor, which would otherwise copy their colours into the output). HarnessIR is inference-only.


๐Ÿ“Š Evaluation Protocol and Metrics

Click to expand the evaluation protocol

Fidelity is measured with full-reference metrics (PSNR, SSIM, LPIPS, DISTS) and image quality with no-reference metrics (MANIQA, CLIP-IQA, MUSIQ, TOPIQ, AFINE-NR).

A higher NR-IQA score does not by itself indicate better restoration: a model that repaints text or fabricates structure can still outscore its own ground truth. Therefore, for images with GT, we calculate the absolute difference between their NR-IQA scores and those of the corresponding GT images. We also report an independent VLM evaluator, giving D-Score for degradation removal and F-Score for content preservation (both 0โ€“100), combined as their per-image geometric mean, DF-Score โ€” taken per image and then averaged, so one image's strong D-Score cannot offset another image's broken F-Score:

$$\mathrm{DF\text{-}Score} = \frac{1}{N}\sum_{i=1}^{N}\sqrt{D_i \cdot F_i}$$

The evaluator's criteria come from each image's degradation type, never from the prompt the method was given: a method must not be able to change the yardstick by changing what it asks for.

NR-IQA rewards altered content


๐Ÿ–ผ๏ธ Experimental Results

Click to expand experimental results

HarnessIR is evaluated with two executors under the identical harness โ€” Nano Banana 2 (NB2) and GPT-Image-2.5-Sunburst โ€” against their direct-use baselines, all-in-one restoration models, and prior agentic IR methods.

MiO100

On the synthetic mixed-degradation benchmark, HarnessIR-NB2 and HarnessIR-GPT gain +1.40 and +1.75 dB PSNR over their direct-use baselines, plus +4.1 and +18.1 DF-Score points, mainly through better content preservation.

Results on MiO100

Real-Paired-200

On real-world images with paired ground truth, HarnessIR-NB2 reaches 26.63 dB PSNR and a DF-Score of 62.0, while HarnessIR-GPT improves PSNR by +2.87 dB and lifts its DF-Score from 40.1 to 60.2. Both exceed the previous best DF-Score of 38.8.

Results on Real-Paired-200

Real-NoGT-200

On real-world images without ground truth, HarnessIR yields DF-Scores of 62.9 for NB2 and 49.3 for GPT-Image-2.5, well above the best prior result of 32.0. NR-IQA alone cannot rank these methods, as it rewards hallucinated detail even when the content has been altered.

Results on Real-NoGT-200


๐Ÿ–ผ๏ธ Visual Results

Visual comparisons on the three test sets

Click to expand more visual comparisons

Additional visual comparisons


๐Ÿ‹๏ธ Deployment

Environment

pip install -r requirements.txt

Configuration

1. API endpoint and key. harness/settings.py speaks the Gemini generateContent protocol:

POST {API_BASE}/v1beta/models/{model}:generateContent
headers: {"Authorization": "Bearer <API_KEY>"}

Fill in API_BASE (origin only, no trailing path) and API_KEY there, or set the environment variables โ€” the environment wins:

export HARNESS_API_BASE="https://<your-endpoint-host>"
export HARNESS_API_KEY="<your-api-key>"

Leaving both unset raises at client construction with the name of the variable to set, rather than failing later with an opaque HTTP error.

2. Model ids. Also in harness/settings.py:

VLM_MODEL_NAME = "gemini-3.7-flash"      # the diagnoser / composer / verifier
MFM_GEMINI_ENDPOINTS = {                  # the executor
    "nb2": "gemini-3.1-flash-image-preview",
    "gpt-image-2.5": "gpt-image-2.5-sunburst",
}

Any model id the endpoint accepts works; these are the ones used in the paper.

3. Tool weights. Expected under ./checkpoints, overridable with HARNESS_CKPT_DIR:

checkpoints/
|-- sam3_semantic.pt                                      T4 segmentation (SAM 3)
|-- depth_anything_v2_base/model.safetensors              T3 depth (Depth-Anything-V2-Base)
|-- insightface/models/buffalo_l/det_10g.onnx             T2 faces (SCRFD)
`-- paddleocr/official_models/PP-OCRv6_medium_{det,rec}   T1 text (PP-OCRv6)
export HARNESS_CKPT_DIR=/path/to/weights

4. Run parameters. configs/default.json holds the defaults (executor, round count, metrics, worker count, geometry). run_harness.py --config <file> points at a different one, and the command-line flags override it. Each parameter is documented in the file and in python run_harness.py --help.

Input format: manifest

All four entry points read the same manifest โ€” a JSON array, or JSONL with one object per line:

[
  {
    "lq": "/abs/path/to/input.png",
    "gt": "/abs/path/to/ground_truth.png",
    "type": ["haze"],
    "prompt": "Please remove the haze from the image."
  }
]
Field Required Meaning
lq yes the degraded input image
gt no ground truth; only the full-reference metrics need it
type no what the restoration is asked to remove
prompt no the restoration request for this image; both paths take it as input

A request may be written per image (prompt), or once for the whole run with --intent, which takes precedence.

The four entry points

1. run_baseline.py โ€” direct use of the executor

Executes the manifest's own prompt field. No diagnosis, no tools, no prompt composition. This is the reference point the harness is measured against.

python run_baseline.py \
  --manifest manifest.json \
  --lq-root  /path/to/test-set \
  --out      runs/baseline

2. run_harness.py โ€” the full harness

Runs stages 1โ€“4, and stage 5 with --redo.

python run_harness.py \
  --manifest manifest.json \
  --lq-root  /path/to/test-set \
  --out      runs/harness

Redo is off by default, which is the single-pass configuration.

3. compute_iqa.py โ€” fidelity and quality metrics

python compute_iqa.py --records runs/harness/record.json

Full-reference: PSNR, SSIM, LPIPS, DISTS. No-reference: MANIQA, CLIP-IQA, MUSIQ, TOPIQ, AFINE-NR. The full-reference metrics need a ground truth and are skipped without one.

4. compute_df.py โ€” D-Score, F-Score and DF-Score

python compute_df.py --records runs/harness/record.json --default-type mix

A VLM evaluator scores each (input, result) pair on two axes, 0โ€“100: D for degradation removal and F for content fidelity, combined as their per-image geometric mean, DF-Score.

Output layout

runs/harness/
|-- config.json                  the configuration this run used
|-- record.json                  flat [{id, lq, gt, type, output, error}] for the scorers
|-- summary.json                 per-sample records and aggregate statistics
|-- results/<id>.png             THE DELIVERED IMAGE, one per sample
`-- per_image/<id>/
    |-- json/diagnosis.json      stage 1
    |-- json/tool_evidence.json  stage 2
    |-- json/composition.json    stage 3
    |-- text/text_evidence.txt   the evidence rendered for the prompt writer
    |-- text/prompt.txt          the composed prompt (round 0)
    |-- output.png               round 0, after colour alignment
    |-- output_r1.png            round 1 (only with --redo)
    |-- output_r1_prompt.txt     the prompt round 1 was executed with
    |-- refined_prompt_r1.txt    the revision that produced it
    `-- record.json              this image's full record: every round's prompt,
                                 verdict, scores and measurements

Re-running the same command resumes: samples that already have a delivered image are skipped.


๐Ÿ“ฎ Contact

If you have any questions, please feel free to contact: xiangtao.kong@connect.polyu.hk

๐Ÿ“š Citation

TODO: Add the BibTeX entry for HarnessIR.

About

HarnessIR: Harnessing Multimodal Foundation Models for Universal Real-World Image Restoration

Resources

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages