Natural language requests for modifying reconstructed 3D scenes are often semantically underspecified and difficult to ground in specific objects within cluttered environments. Existing methods, which treat the task as one-shot conditional generation from a single prompt, fail to resolve ambiguous user intents and suffer from object localization drift, tracking failure under occlusions, and the notorious multi-view "sticker effect."
The Plan-Perceive-Act Paradigm
DesignAgent3D reformulates 3D scene editing as an interactive multimodal agentic framework adopting a designer-like Plan-Perceive-Act paradigm. The agent first plans by interacting with the user to clarify underspecified design goals; then perceives by grounding the intended edit to specific objects or regions in the 3D scene; finally acts by applying controlled visual modifications while preserving scene consistency.
Multi-view Consistency and 3D Representation
The edits are integrated into the underlying 3D representation, supporting persistent and multi-view consistent novel-view rendering. This approach resolves consistency issues that plague previous techniques, ensuring edits remain stable and realistic from any viewing angle.
Performance on NeRF and 3D Gaussian Splatting
Extensive experiments across both NeRF and 3D Gaussian Splatting backbones demonstrate that DesignAgent3D significantly outperforms state-of-the-art baselines. The system delivers superior semantic intent alignment, impeccable spatial localization accuracy, and high-fidelity multi-view consistency, marking a step forward in interactive AI-generated 3D scenes.

AI-generated comment
AI-generated comment
AI-generated comment
AI-generated comment