How to Benchmark AI Image Generators: A 30-Prompt Test Plan
A repeatable methodology for comparing Midjourney, GPT Image, Nano Banana and FLUX on visual quality, editing, text accuracy, speed and cost.
Build the test set before viewing the outputs
Create 30 prompts covering six categories: human anatomy, product packaging, exact typography, architecture and perspective, reference-image preservation, and a complex scene with several constraints. Avoid relying on a single visually impressive example. Define the requested aspect ratio, number of images and permitted editing steps. Ensure every tool receives as close to the same instructions and reference inputs as its interface allows. When a platform lacks an equivalent editing operation, record that limitation instead of pretending the tests were identical.
Score measurable properties separately
Use a rubric with prompt adherence, textual accuracy, preservation of reference objects, layout and aesthetic quality scored individually. Have two reviewers judge blind outputs without seeing model names; reconcile major disagreements. Track the number of acceptable images rather than just the most attractive image in a batch. For timing, measure end-to-end completion including queue time; for cost, include retries and editing passes. A low headline per-image price can become expensive if the image repeatedly fails to preserve packaging details or spelling.
Publish enough information for others to reproduce
A proper comparison should disclose model versions, date, settings, generation count, editing tools, final output resolution and any excluded failures. Report mean or median scores along with sample size and uncertainty; note that subjective rankings do not establish factual accuracy. Before commercial use, review each service's terms and permissions. This RecoupRev guide is a proposed experimental protocol; it does not claim we executed a 30-prompt benchmark or measured independent winners today.
Reporting sources & references
These links identify the reporting or public materials on which the article is based; they do not imply our newsroom witnessed the events.