從提示詞到構圖:用於海報生成的 Spatial Canvas Interface
From Prompting to Composing: A Spatial Canvas Interface for Poster Generation
研究提出 Spatial Canvas Interface,讓用戶在空間畫布上編排海報生成意圖,並據此開發模型 Compo。介面提供語義、身份、文字及像素四種綁定方式,並以 Text Specifications 描述個別元素與整體外觀。評測顯示,Compo 的構圖可控性強於通用圖像生成模型及專用海報生成系統,同時維持高視覺質素。
Published on Oct 8
Authors:
,
,
,
,
,
,
Abstract
Text prompting is an indirect interface for poster generation, requiring users to encode inherently two-dimensional composition intent into a one-dimensional sequence of words. We introduce a Spatial Canvas Interface that enables users to directly compose generation intent in space through four complementary binding types: semantic, identity, text, and pixel, together with Text Specifications for individual elements and global appearance. Based on this interface, we develop Compo, a poster generation model adapted from a pretrained image editing model to understand Spatial Canvas inputs and Text Specifications. Compo supports both direct inference, where users explicitly construct the canvas, and agentic mode, where a high-level request is automatically translated into a planned Spatial Canvas. To train Compo, we develop a scalable pipeline that automatically constructs supervision data for different binding types and their combinations, enabling efficient adaptation without training a specialized poster generator from scratch. We further introduce a benchmark that evaluates adherence to individual binding types and their joint composition. Experiments show that Compo achieves stronger compositional controllability than both general-purpose image generation models and dedicated poster generation systems while maintaining high visual quality. By decoupling intent specification from visual generation, our work shifts poster generation from prompting toward composing.
View arXiv page View PDF Project page GitHub Add to collection
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2610.12230 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2610.12230 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2610.12230 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co