Under Review

Scanning While Imagining:
A Scene-Graph World Model for Robotic Ultrasound Navigation

aComputer Aided Medical Procedures (CAMP), Technical University of Munich, Germany bMunich Center for Machine Learning (MCML), Germany cDepartment of Mechanical Engineering, The University of Hong Kong, Hong Kong SAR, China
A sonographer anticipates what the next view will show before moving the probe. SonoGraph-WM gives the robot the same ability: a scene-graph world model imagines the anatomy that candidate motions would reveal, and the robot selects, executes and replans.

Overview

In a Nutshell

Ultrasound acquisition depends on anticipating how the view will change as the probe moves. SonoGraph-WM is an action- and goal-conditioned world model that imagines those changes, not as synthetic ultrasound images, but as anatomical scene graphs, and uses them to plan where the probe should go next.

  • Anatomy, not texture. Scene graphs capture visible structures, their geometry and spatial relations, without synthesizing speckle.
  • One model, two predictions. A unified Transformer jointly predicts future scene graphs and probe poses from the history, a candidate action and goal graphs.
  • Imagine, execute one step, replan. The planner rolls out candidate trajectories, picks the shortest one that reaches the goal graph, takes a step and replans from the new observation.
Read the full abstract

Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditioned world model for anticipatory probe navigation. The model represents anatomy as scene graphs (SGs), capturing visible structures, their geometry, and spatial relationships without synthesizing US images. Given a history of SGs and probe poses, a unified Transformer jointly predicts future SGs and poses. A receding-horizon planner recursively imagines candidate trajectories, selects the shortest predicted path reaching a goal graph, and follows it over a short execution horizon before replanning from new observations. To reduce reliance on tracked and anatomically annotated US sequences, we generate aligned SG–pose training data from computed tomography (CT) label maps along surface-constrained probe trajectories. On four held-out CT cases, spatial relation F1 remains above 93% over 20 prediction steps, and closed-loop navigation achieves 77.50% and 75.00% success for the gallbladder and pancreas, respectively, using annotation-derived SGs. In robot–phantom navigation experiments with label-map-derived SGs, the planner reached the target view in 73.7% of trials. These findings support CT-supervised anatomical world modeling for probe planning and highlight the importance of frequent observation updates for reliable navigation.

Method

How It Works

Every candidate probe motion is judged by the anatomy it is predicted to reveal. The robot compares imagined futures before it moves, and trusts only what it actually observes.

1Scene graphs capture anatomy, not image texture

Structures intersecting the imaging plane become nodes carrying component count, bounding box, centroid and visible-area ratio; edges encode adjacent to, left of and superficial to.

2Imagine. Execute one step. Replan.

At every step the world model imagines eight candidate rollouts (purple), keeps those whose predicted scene graph matches the target graph, selects the one with the shortest predicted probe travel (teal), executes a single step (red) and replans from the new observation. Navigation ends only when the observed scene graph confirms the goal. Click to pause, drag the scrubber, or open it full screen.

Real robot

Robot–Phantom Navigation

A KUKA LBR iiwa 7 carries a BK Medical bk3000 system with a 5C1e curved probe over a Kyoto Kagaku US-22 abdominal phantom. The planner receives only goal scene graphs, never goal poses.

Robotic ultrasound platform

The goal is a scene graph, not a probe pose

Navigation succeeds when the observed graph matches one of the target views (here: pancreas goal graph G1). Click a figure to view it full screen.

Pancreas navigation on the phantom

One pancreas trial, all views synchronized: RViz, third-view camera, acquired B-mode, the registered label map and the observable scene graph at each step. Scene graphs come from the registered label map; B-mode is recorded for assessment, not used for planning.

Evaluation

Results at a Glance

>93%
spatial-relation F1 over 20 imagined steps
77.5%
gallbladder navigation success, held-out CT
75.0%
pancreas navigation success, held-out CT
73.7%
robot–phantom trials reach the target view (14/19)
Why replanning after every step matters

Closed-loop success rate (%) on four held-out CT cases, 120 trials per target, for execution horizons of 1, 5 and 10 steps before replanning. Near/far groups start closer to or farther from the nearest goal view.

Target Start Execute 1 Execute 5 Execute 10
GallbladderNear81.2534.3839.06
Far73.2148.2137.50
Overall77.5040.8338.33
PancreasNear80.0029.2332.31
Far69.0916.3610.91
Overall75.0023.3322.50

Citation

BibTeX

@misc{li2026scanning,
  title={Scanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound Navigation},
  author={Li, Xuesong and Chen, Shuai and Li, Feng and Jiang, Zhongliang and Navab, Nassir and Bi, Yuan},
  year={2026},
  note={Under review}
}