LEGO:無需將影片提升至點雲的外部視角至第一人稱視角影片生成方法
LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation
LEGO 提出一種無需將影片提升至點雲的外部視角轉第一人稱影片生成方法,透過經微調的 Transformer 直接合成目標視角,再由影片擴散模型補足細節。系統估算各區域的信心度,用以遮蔽低信心區域,並在去噪早期引導生成器優先處理高信心區域。方法持續勝過最先進的顯式重建流程,並可不經重新訓練泛化至其他數據集。
Published on Oct 8
Authors:
,
,
,
Abstract
Generating an egocentric video from a single exocentric recording is a challenging case of novel view synthesis, as the two cameras share little overlap and much of the target view is unobserved. Current state-of-the-art methods reconstruct the scene explicitly by estimating depth, lifting the video into a point cloud, and re-rendering it from the egocentric camera to condition a video diffusion model. This deterministic mapping assigns each pixel to a single reprojected location, which preserves texture but translates depth errors into misplaced content. We ask what a video diffusion model should receive as its condition and propose a lifting-free answer: a learned view synthesizer, an LVSM-style transformer fine-tuned to render the egocentric view directly without depth, point clouds, or reprojection, resolving cross-view correspondence internally. In contrast, its probabilistic mapping averages each region over candidate source locations according to a learned correspondence distribution, preserving structure while fine texture is averaged away. We argue that this trade-off suits a diffusion generator, whose denoising training excels at restoring detail, so an effective condition should prioritize structural alignment over sharpness. This distribution's concentration also yields a per-region confidence, used both to mask low-confidence regions and to guide the generator toward high-confidence areas during early layout-forming denoising steps. Our approach consistently outperforms the state-of-the-art explicit pipeline and generalizes to other datasets without retraining. The synthesizer thus supplies view structure, and the diffusion model its detail.
View arXiv page View PDF Project page GitHub 14 Add to collection
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2610.12442 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2610.12442 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2610.12442 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co