跳到正文
HuggingFace Daily Papers(社區熱門論文)·· 1 天前AI 評分45

VibeEdit:以畫布指令編輯圖像

VibeEdit: Image Editing with Canvas Instructions

AI 導讀

VibeEdit 提出以畫布指令編輯圖像的介面,用戶可直接在圖像上標記編輯位置及修改內容,無須另寫文字提示詞。在 419 個案例的基準測試中,VibeEdit 的 VLM 評分為 79.9、區域外 PSNR 為 32.8 dB,高於文字指令基線 FireRed 的 67.4 及 24.0 dB。

正文

Published on Oct 8

Authors:

,

,

,

,

,

,

,

,

,

,

Abstract

In text-guided image editing, describing the desired change is often straightforward, but identifying the intended object or region can be cumbersome, especially when several objects look alike. We introduce a new image editing interface that lets users place spatial marks and optional short notes directly on the image. Together, these annotations form a canvas instruction that specifies where to edit and what to change. Our editor, VibeEdit, follows these instructions to perform object addition, removal, replacement, attribute modification, and movement without a separate text prompt. We construct 1.55 million source-target edit pairs with object masks and structured edit descriptions, from which we render canvas instructions during training. We adapt Qwen-Image-Edit with layer-decoupled conditioning that separately encodes source images and canvas instructions for image editing. We train the model with region-weighted supervised fine-tuning, followed by rubric-guided reinforcement learning to improve edit completion, local edit quality, and preservation of unedited regions. We evaluate VibeEdit on an independently constructed, human-curated benchmark of 419 cases emphasizing target selection among similar objects. VibeEdit achieves a VLM rubric score of 79.9 and an outside-region PSNR of 32.8 dB, compared with 67.4 and 24.0 dB for FireRed, the highest-scoring text-instructed baseline in our evaluation.

View arXiv page View PDF Project page GitHub Add to collection

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.12229 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.12229 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.12229 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co