Hierarchical Surgical Anchor
Organizes first-frame appearance and multi-level semantic masks into shared appearance, shape, and region anchors.
Controllable surgical world modeling
Anchor-centered multimodal control for coherent instrument-tissue interaction generation.
Overview
The initial frame and hierarchical masks establish the persistent appearance and structure; subsequently, optional optical flow, depth, and edge information govern the dynamic, interactive evolution process. The result is the generation of high-quality, controllable surgical videos.
Cholec80-SurgWAM
Motion-guided clip selection is followed by expert-curated hierarchical masks, surgical interaction descriptions, and aligned flow, depth, and edge controls.
Surgical Anchor-Relative Control Adapter
Surg-ARCA separates persistent appearance and semantic structure from optional boundary, geometry, and motion evidence before generating stage-wise control hints for Wan2.2.
Organizes first-frame appearance and multi-level semantic masks into shared appearance, shape, and region anchors.
Dedicated experts interpret flow, depth, and edge evidence relative to the same persistent surgical scene state.
Preserves each activated modality increment during stage-wise composition and produces surgical control hints.
Video demonstrations
01 / Cross-method comparison
Mask, depth, flow, and edge conditions are evaluated on aligned surgical cases. The green outline identifies Surg-UniWorld.
Reference and control inputs are shown alongside Cosmos-H-Transfer, VACE-Wan2.2, ControlNet-Wan2.2, and Surg-UniWorld.
02 / Flexible multimodal control
Three aligned cases show how boundary, geometry, and motion evidence complement the shared surgical anchor.
Every column uses the same reference sequence and differs only in the activated modality subset.
03 / Single-control gallery
Switch the active modality to compare its control signal with the reference sequence and generated result.
Hierarchical semantic masks preserve instrument contours, tissue organization, and interaction regions.
Interactive multimodal control
Select one input sequence, then add geometry, motion, or boundary evidence to the shared mask anchor.