Controllable surgical world modeling

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

Anchor-centered multimodal control for coherent instrument-tissue interaction generation.

Wan2.2 · Hierarchical surgical anchor · Flexible multimodal control

Overview

Multimodal Controllable Surgical Video Generation Model

The initial frame and hierarchical masks establish the persistent appearance and structure; subsequently, optional optical flow, depth, and edge information govern the dynamic, interactive evolution process. The result is the generation of high-quality, controllable surgical videos.

Surg-UniWorld overview from multimodal controls to generated surgical video
Overview of multimodal-controlled surgical world generation.

Cholec80-SurgWAM

A benchmark built for multimodal surgical control

Motion-guided clip selection is followed by expert-curated hierarchical masks, surgical interaction descriptions, and aligned flow, depth, and edge controls.

Construction pipeline of the Cholec80-SurgWAM dataset
Construction pipeline of Cholec80-SurgWAM.
6,001
49-frame clips
294,049
Sampled frames
573,721
Curated masks
5,104 / 897
Train / test clips

Surgical Anchor-Relative Control Adapter

Interpret every modality relative to one surgical anchor

Surg-ARCA separates persistent appearance and semantic structure from optional boundary, geometry, and motion evidence before generating stage-wise control hints for Wan2.2.

Architecture of Surg-UniWorld and the Surgical Anchor-Relative Control Adapter
Architecture of Surg-UniWorld and Surg-ARCA.
01

Hierarchical Surgical Anchor

Organizes first-frame appearance and multi-level semantic masks into shared appearance, shape, and region anchors.

02

Anchor-Relative Modality Experts

Dedicated experts interpret flow, depth, and edge evidence relative to the same persistent surgical scene state.

03

Multimodal Control Expert

Preserves each activated modality increment during stage-wise composition and produces surgical control hints.

Video demonstrations

Controllable surgical dynamics

01 / Cross-method comparison

Four controls, six aligned views

Mask, depth, flow, and edge conditions are evaluated on aligned surgical cases. The green outline identifies Surg-UniWorld.

Reference and control inputs are shown alongside Cosmos-H-Transfer, VACE-Wan2.2, ControlNet-Wan2.2, and Surg-UniWorld.

02 / Flexible multimodal control

Eight modality configurations

Three aligned cases show how boundary, geometry, and motion evidence complement the shared surgical anchor.

Every column uses the same reference sequence and differs only in the activated modality subset.

Interactive multimodal control

Compose the evidence.
Inspect the generation.

Select one input sequence, then add geometry, motion, or boundary evidence to the shared mask anchor.

Input case
Input sequence Case 01
Activated controls Mask
Mask
Generated sequence Surg-UniWorld