SinoTechIntel Academic Portal
Open AccessDOI: 10.1631/FITEE_2400904Original Research

Image generation evaluation: a comprehensive survey of human and automatic evaluations

Qi LIU¹,Shuanglin YANG¹,Zejian LI¹,Lefan HOU¹,Chenye MENG¹,Ying ZHANG¹,Lingyun SUN¹

Zhejiang University (School of Software Technology and College of Computer Science and Technology), China

Read Executive PreviewQuick FAQ
Image generation evaluation: a comprehensive survey of human and automatic evaluations
Graphical Abstract / Figure
Published In
Frontiers of Information Technology & Electronic Engineering
Published:October 5, 2025Edition:Vol. 32, Issue 10 • pp. 881-893Citation:Qi LIU et al. (2025), Frontiers of Information Technology & Electronic Engineering
Impact Factor2.7 (Q2 - Springer)
Sponsored Research Partner
Keywords & Index Terms:image generation evaluationhuman evaluationautomatic evaluationevaluation protocolstext-to-image generationlayout-to-image generationcomputer vision

Key Takeaways & Executive Findings

  • • Provides a comprehensive survey of both human and automatic evaluation methods for image generation, covering evaluation protocols and methods. • Summarizes 10 image generation tasks and proposes a novel evaluation protocol that addresses human and automatic evaluation aspects. • Presents the first exhaustive summary of human evaluation, detailing methods, tools, procedures, and data analysis techniques. • Discusses current challenges and outlines future research directions for reliable image generation evaluation.
Sponsored Research Highlight

Abstract

Image generation models have made remarkable progress, and image evaluation is crucial for explaining and driving the development of these models. Previous studies have extensively explored human and automatic evaluations of image generation. Herein, these studies are comprehensively surveyed, specifically for two main parts: evaluation protocols and evaluation methods. First, 10 image generation tasks are summarized with focus on their differences in evaluation aspects. Based on this, a novel protocol is proposed to cover human and automatic evaluation aspects required for various image generation tasks. Second, the review of automatic evaluation methods in the past five years is highlighted. To our knowledge, this paper presents the first comprehensive summary of human evaluation, encompassing evaluation methods, tools, details, and data analysis methods. Finally, the challenges and potential directions for image generation evaluation are discussed. We hope that this survey will help researchers develop a systematic understanding of image generation evaluation, stay updated with the latest advancements in the field, and encourage further research.

1. Introduction

Image generation models have undergone significant advancement; therefore, the methods used for their evaluation must be continuously updated to ensure reliable results. The advancement of deep learning has facilitated aligning images and generating them from various data types such as texts, sketches, scene graphs, and layout graphs (Elasri et al., 2022). These images have been used in various fields, including medicine, fashion, material design, media, and e-commerce. Quality must be ensured as it directly impacts the recipient’s visual experience.

The performance of image generation models must be reliably evaluated for their further development. However, the evaluation methods have not been further developed in line with the advancement of image generation models, which is not conducive to their iterative improvement. Fig. 1 compares the number of published papers related to image generation and image generation evaluation in the last decade. Research on image generation evaluation has not kept pace with that on image generation; by checking the abstracts and key words, we retained 2399 papers on image generation and 46 papers on image generation evaluation.

Image generation tasks are diverse, with varying evaluation aspects. In text-to-image generation, visual fidelity and semantic alignment are crucial evaluation aspects, because understanding complex text and accurately translating it into visual are challenging. In contrast, fidelity, recognizability, and diversity are important factors in layout-to-image generation, and generating diverse images featuring multiple complex objects is a major challenge. Therefore, image generation evaluation has been extensively studied, and this paper summarizes such studies and presents the latest developments in this field.

SinoTechIntel Interactive Document Reader
Page 1–5 of Preview
100%
Download Full PDF

Loading authentic research manuscript (Pages 1–5)...

Sponsored Research Partner
Cite This Research Paper
Qi LIU, Shuanglin YANG, Zejian LI, Lefan HOU, Chenye MENG, Ying ZHANG, Lingyun SUN (2025). Image generation evaluation: a comprehensive survey of human and automatic evaluations. Frontiers of Information Technology & Electronic Engineering. https://doi.org/10.1631/FITEE_2400904
SinoTechIntel Academic & Legal Disclaimer

Research & Educational Purpose Only:The translations, structured abstracts, analytical annotations, and data reports provided by SinoTechIntel are intended exclusively for academic research, internal corporate R&D, and educational benchmarking. They do not constitute formal engineering, chemical safety, legal, or professional advice.

Copyright & Intellectual Property Notice: Original copyright of the underlying source articles and experimental data remains with the respective authors, institutions, and original publishing journals. SinoTechIntel claims intellectual property only over its proprietary translations, analytical syntheses, and AEO structured enhancements in accordance with international fair use and academic citation principles.

Frequently Asked Questions

What is the paper about?

The paper is a comprehensive survey of human and automatic evaluation methods for image generation, covering evaluation protocols and methods across 10 image generation tasks.

What are the main parts of the survey?

The survey is organized into two main parts: evaluation protocols and evaluation methods.

What is the novel contribution of this paper?

It proposes a novel evaluation protocol that covers human and automatic evaluation aspects for various image generation tasks, and it presents the first comprehensive summary of human evaluation, including methods, tools, details, and data analysis methods.

Why is image generation evaluation important?

Evaluation is crucial for explaining and driving the development of image generation models, ensuring quality and reliable performance across various applications.

What future directions are discussed?

The paper discusses challenges and potential directions for image generation evaluation, aiming to help researchers stay updated and encourage further research.

Recommended Scientific Literature & Research Partners

Related Technical Papers & Translations

Research Paper
Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

Design and optimization of a high-efficiency distillation process for cellulosic fuel ethanol integrated with thermal coupling and molecular sieve adsorption

To address the challenges of high energy consumption and prominent costs in the traditional three-columns distillation process for cellulosic fuel ethanol, a distillation—molecular sieve coupling separation process is proposed. This process integrates a three-column (crude distillation column, first distillation column, second distillation column) system with a 3A molecular sieve adsorption deep dehydration unit. A thermal coupling network is constructed via differential pressure design (steam from medium/high-pressure columns as mutual heat sources, reboiler liquid waste heat for feed preheating), and molecular sieve adsorption conditions are optimized. The study first performs a thermodynamic consistency test on the ethanol—water system, determines optimal non-random two-liquid (NRTL) model binary interaction parameters via experimental data regression for Aspen Plus simulation. Aiming at minimum total annual cost (TAC), Aspen Plus is used to optimize process parameters (theoretical tray number, feed location, reflux ratio, side-draw position, etc.). Economic analysis shows this process reduces CO2 emission costs by 27.56%, TAC by 15.58% (to 5.123 × 106 USD·a-1), and increases ethanol purity to >99.6%, providing an effective solution for green, efficient separation.

Read Abstract & PDF
Research Paper
A cohesion loss model for determining residual strength of deep bedded sandstone

A cohesion loss model for determining residual strength of deep bedded sandstone

Rock residual strength, as an important input parameter, plays an indispensable role in proposing the reasonable and scientific scheme about stope design, underground tunnel excavation and stability evaluation of deep chambers. Therefore, previous residual strength models of rocks established were reviewed. And corresponding related problems were stated. Subsequently, starting from the effects of bedding and whole life-cycle evolution process, series of triaxial mechanical tests of deep bedded s

Read Abstract & PDF
Research Paper
Federated model with contrastive learning and adaptive control variates for human activity recognition

Federated model with contrastive learning and adaptive control variates for human activity recognition

Recent attention to privacy issues demands a communication-safe method for training human activity recognition (HAR) models on client activity data. Federated learning (FL) has become a compelling technique to facilitate model training between the server and clients while preserving data privacy. However, classical FL methods often assume independent and identically distributed (IID) data among clients. This assumption does not hold true in practical scenarios. Human activity in real-world scena

Read Abstract & PDF