Restoration by harnessing an MFM executor, not by scheduling restoration tools.
Xiangtao Kong1,2 | Shuaizheng Liu1,2 | Rongyuan Wu1,2 | Lingchen Sun1,2 | Zhengqiang Zhang1,2 | Jinxin Zhao1,2 | Yuhui Wu1,2 | Lei Zhang1,2,โ
1 The Hong Kong Polytechnic University
2 OPPO Research Institute
โ Corresponding author.
- ๐ฐ News
- ๐งฐ HarnessIR
- ๐ Evaluation Protocol and Metrics
- ๐ผ๏ธ Experimental Results
- ๐ Visual Results
- ๐๏ธ Deployment
- ๐ฎ Contact
- ๐ Citation
- 2026-10-07: Released the code, benchmark and visual results of HarnessIR.
Click to expand the pipeline description
Prior agentic IR methods apply task-specific restoration models in sequence, and are therefore bounded by those tools. HarnessIR keeps one MFM as the executor and varies only the prompt it is given. Five stages:
- Perception and diagnosis โ a VLM reads the content and its degradations, and names which auxiliary tools to consult.
- On-demand tool invocation โ only the planned tools run, returning image-specific evidence.
- Prompt composition โ the evidence is composed into one prompt stating what to treat and what must survive untouched.
- Execution โ the MFM receives the original LQ image and that prompt, in a single pass.
- Verification-driven refinement โ the result is judged against the restoration requirements, and re-executed from the original LQ input only when it fails.
Two choices distinguish it from prior pipelines: no restoration tool is scheduled (the auxiliary tools supply evidence for writing the prompt, not links in a restoration chain), and the executor sees only the image and text (diagnostic maps inform the prompt writer but are never passed to the editor, which would otherwise copy their colours into the output). HarnessIR is inference-only.
Click to expand the evaluation protocol
Fidelity is measured with full-reference metrics (PSNR, SSIM, LPIPS, DISTS) and image quality with no-reference metrics (MANIQA, CLIP-IQA, MUSIQ, TOPIQ, AFINE-NR).
A higher NR-IQA score does not by itself indicate better restoration: a model that repaints text or fabricates structure can still outscore its own ground truth. Therefore, for images with GT, we calculate the absolute difference between their NR-IQA scores and those of the corresponding GT images. We also report an independent VLM evaluator, giving D-Score for degradation removal and F-Score for content preservation (both 0โ100), combined as their per-image geometric mean, DF-Score โ taken per image and then averaged, so one image's strong D-Score cannot offset another image's broken F-Score:
The evaluator's criteria come from each image's degradation type, never from the prompt the method was given: a method must not be able to change the yardstick by changing what it asks for.
Click to expand experimental results
HarnessIR is evaluated with two executors under the identical harness โ Nano Banana 2 (NB2) and GPT-Image-2.5-Sunburst โ against their direct-use baselines, all-in-one restoration models, and prior agentic IR methods.
On the synthetic mixed-degradation benchmark, HarnessIR-NB2 and HarnessIR-GPT gain +1.40 and +1.75 dB PSNR over their direct-use baselines, plus +4.1 and +18.1 DF-Score points, mainly through better content preservation.
On real-world images with paired ground truth, HarnessIR-NB2 reaches 26.63 dB PSNR and a DF-Score of 62.0, while HarnessIR-GPT improves PSNR by +2.87 dB and lifts its DF-Score from 40.1 to 60.2. Both exceed the previous best DF-Score of 38.8.
On real-world images without ground truth, HarnessIR yields DF-Scores of 62.9 for NB2 and 49.3 for GPT-Image-2.5, well above the best prior result of 32.0. NR-IQA alone cannot rank these methods, as it rewards hallucinated detail even when the content has been altered.
pip install -r requirements.txt1. API endpoint and key. harness/settings.py speaks the Gemini generateContent protocol:
POST {API_BASE}/v1beta/models/{model}:generateContent
headers: {"Authorization": "Bearer <API_KEY>"}
Fill in API_BASE (origin only, no trailing path) and API_KEY there, or set the environment variables โ the environment wins:
export HARNESS_API_BASE="https://<your-endpoint-host>"
export HARNESS_API_KEY="<your-api-key>"Leaving both unset raises at client construction with the name of the variable to set, rather than failing later with an opaque HTTP error.
2. Model ids. Also in harness/settings.py:
VLM_MODEL_NAME = "gemini-3.7-flash" # the diagnoser / composer / verifier
MFM_GEMINI_ENDPOINTS = { # the executor
"nb2": "gemini-3.1-flash-image-preview",
"gpt-image-2.5": "gpt-image-2.5-sunburst",
}Any model id the endpoint accepts works; these are the ones used in the paper.
3. Tool weights. Expected under ./checkpoints, overridable with HARNESS_CKPT_DIR:
checkpoints/
|-- sam3_semantic.pt T4 segmentation (SAM 3)
|-- depth_anything_v2_base/model.safetensors T3 depth (Depth-Anything-V2-Base)
|-- insightface/models/buffalo_l/det_10g.onnx T2 faces (SCRFD)
`-- paddleocr/official_models/PP-OCRv6_medium_{det,rec} T1 text (PP-OCRv6)
export HARNESS_CKPT_DIR=/path/to/weights4. Run parameters. configs/default.json holds the defaults (executor, round count, metrics, worker count, geometry). run_harness.py --config <file> points at a different one, and the command-line flags override it. Each parameter is documented in the file and in python run_harness.py --help.
All four entry points read the same manifest โ a JSON array, or JSONL with one object per line:
[
{
"lq": "/abs/path/to/input.png",
"gt": "/abs/path/to/ground_truth.png",
"type": ["haze"],
"prompt": "Please remove the haze from the image."
}
]| Field | Required | Meaning |
|---|---|---|
lq |
yes | the degraded input image |
gt |
no | ground truth; only the full-reference metrics need it |
type |
no | what the restoration is asked to remove |
prompt |
no | the restoration request for this image; both paths take it as input |
A request may be written per image (prompt), or once for the whole run with --intent, which takes precedence.
Executes the manifest's own prompt field. No diagnosis, no tools, no prompt composition. This is the reference point the harness is measured against.
python run_baseline.py \
--manifest manifest.json \
--lq-root /path/to/test-set \
--out runs/baselineRuns stages 1โ4, and stage 5 with --redo.
python run_harness.py \
--manifest manifest.json \
--lq-root /path/to/test-set \
--out runs/harnessRedo is off by default, which is the single-pass configuration.
python compute_iqa.py --records runs/harness/record.jsonFull-reference: PSNR, SSIM, LPIPS, DISTS. No-reference: MANIQA, CLIP-IQA, MUSIQ, TOPIQ, AFINE-NR. The full-reference metrics need a ground truth and are skipped without one.
python compute_df.py --records runs/harness/record.json --default-type mixA VLM evaluator scores each (input, result) pair on two axes, 0โ100: D for degradation removal and F for content fidelity, combined as their per-image geometric mean, DF-Score.
runs/harness/
|-- config.json the configuration this run used
|-- record.json flat [{id, lq, gt, type, output, error}] for the scorers
|-- summary.json per-sample records and aggregate statistics
|-- results/<id>.png THE DELIVERED IMAGE, one per sample
`-- per_image/<id>/
|-- json/diagnosis.json stage 1
|-- json/tool_evidence.json stage 2
|-- json/composition.json stage 3
|-- text/text_evidence.txt the evidence rendered for the prompt writer
|-- text/prompt.txt the composed prompt (round 0)
|-- output.png round 0, after colour alignment
|-- output_r1.png round 1 (only with --redo)
|-- output_r1_prompt.txt the prompt round 1 was executed with
|-- refined_prompt_r1.txt the revision that produced it
`-- record.json this image's full record: every round's prompt,
verdict, scores and measurements
Re-running the same command resumes: samples that already have a delivered image are skipped.
If you have any questions, please feel free to contact: xiangtao.kong@connect.polyu.hk
TODO: Add the BibTeX entry for HarnessIR.







