Skip to content

Question about Qwen-Image-Bench evaluation results #6

Description

@hongsexiaotanhua

Hi LLaDA-Image team,

Thanks for releasing such a great work! It provides us with some new ideas and insights for training image generation models.

I recently evaluated LLaDA-Image and Z-Image-Turbo on Qwen-Image-Bench (zh) and noticed some differences from the results reported in the paper:

Model My Evaluation Reported in the Paper
Z-Image-Turbo 49.52 52.71
LLaDA-Image 51.25 53.38

For comparison, I also evaluated three other open-source models using the same pipeline, and their scores were all within around 0.5 points of the official Qwen-Image-Bench results.

Could you please clarify the evaluation setup used for the reported results, particularly:

  • the evaluation code used for Qwen-Image-Bench;

  • whether there were any modifications to the prompts or preprocessing;

  • the inference configuration (e.g., resolution, sampling steps, seed, etc.).

I would like to make sure my evaluation setup is aligned with yours and better understand the difference in scores. If possible, sharing the relevant evaluation configuration or command would be very helpful.

Thanks!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions