跳到正文
HuggingFace Daily Papers(社區熱門論文)·· 1 天前AI 評分41

RISEBench++:推理導向視覺編輯評測基準

Reasoning-Informed Visual Editing

AI 導讀

研究團隊提出 RISEBench++ 視覺編輯評測基準,並提出免訓練的 RISE-Agent 框架。基準涵蓋 6 個推理維度、12 個子類別、65 種任務及 1,000 個人工標註測試案例,支援英文和中文。評估 58 種方法後,表現最佳的 GPT-Image-2.5 Sunburst 準確率為 56.6%。

正文

Published on Oct 8

·

Submitted by

zpy

on Oct 9

Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

Large Multi-modality Models (LMMs) have made significant progress in visual understanding and generation, but still face challenges in visual editing, particularly in following complex instructions, preserving appearance consistency, and supporting flexible input formats. To study this gap, we introduce RISEBench, the first benchmark for evaluating Reasoning-Informed viSual Editing (RISE), and extend it to RISEBench++, a more comprehensive and fine-grained benchmark for this emerging task. RISEBench++ extends the taxonomy into a hierarchical scheme spanning six reasoning dimensions: Temporal, Causal, Spatial, Logical, and Counterfactual Reasoning, together with Hybrid Reasoning integrating multiple reasoning types across multi-turn edits. These dimensions are further decomposed into 12 subcategories and 65 fine-grained task types. We expand input formats to include multi-image conditioning and scale the benchmark to 1000 human-annotated test cases, released in English and Chinese. We also improve our evaluation framework, assessing Instruction Reasoning, Appearance Consistency, and Visual Plausibility with human judges and an LMM-as-a-judge approach for more reliable and calibrated judgements. Beyond benchmarking, we introduce RISE-Agent, a training-free agentic framework integrating reasoning-driven planning, tool-augmented execution, and verifier-guided refinement, outperforming most strong existing approaches across diverse RISE tasks. We evaluate 58 visual editing approaches, including 34 open-source models, 19 closed-source models, and 5 agentic methods. The results reveal substantial challenges in reasoning-based visual editing, with even the strongest evaluated approach, GPT-Image-2.5 Sunburst, achieving only 56.6% accuracy. RISEBench++ highlights the limitations of contemporary editing models, provides insights, and indicates future directions for reasoning-aware visual editing.

View arXiv page View PDF GitHub Add to collection

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.12343 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.12343 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.12343 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co