Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume Generation

Draw the shape. Describe the geometry. Generate the volume.

Delin An Chaoli Wang

University of Notre Dame CVPR 2026 Main Conference Denver, Colorado · June 5–7

Swipe to explore the full figure

Sketch and text conditioned Sketch2CT results across CHAOS liver CT, AVT aorta CT, Decathlon liver CT, and Decathlon heart MRI, showing seven consecutive slices of each generated volume.
A lightweight sketch and a geometric description control both organ structure and appearance across coherent 3D slices.

Overview

Controllable generation, grounded in anatomy.

Sketch2CT is a multimodal diffusion framework that turns a user-provided 2D sketch and a text description of 3D geometry into an anatomically coherent segmentation mask and its corresponding medical volume.

The sketch provides an intuitive structural blueprint, while text adds depth, topology, symmetry, and surface cues that a single projection cannot capture. A second latent diffusion stage translates the generated mask into realistic CT or MRI appearance while preserving the requested structure.

01

Local feature refinement

TSFE uses text-guided FiLM modulation to strengthen meaningful sketch features and suppress sparse or ambiguous contours.

02

Global multimodal alignment

CGFM combines cross-attention and self-attention to align sketch geometry with high-level textual semantics.

03

Two-stage 3D synthesis

A mask-first pipeline preserves anatomy before synthesizing realistic, spatially continuous volumetric appearance.

Method

Structure first. Appearance second.

The pipeline separates anatomical reasoning from image synthesis, making the generated volume both controllable and structurally faithful.

01

Fuse conditions

Encode the sketch and geometric text, then align local and global cues.

02

Generate structure

Denoise in latent space to reconstruct a coherent 3D organ mask.

03

Synthesize volume

Use the segmentation latent to guide realistic CT or MRI generation.

Framework

Swipe to explore the full figure

Sketch2CT framework with sketch and text fusion, segmentation generation, and medical volume generation stages.
Sketch and text are fused through TSFE and CGFM. The multimodal feature conditions segmentation latent diffusion, and the resulting mask latent guides medical volume synthesis.

Results

Faithful structure across organs and modalities.

Evaluated on liver, aorta, and heart volumes across four public CT and MRI benchmarks at 128³ resolution.

33.7

FID ↓

Best image fidelity on CHAOS liver CT.

0.912

Mask Dice ↑

Highest mask faithfulness on Decathlon liver.

0.893

Downstream Dice ↑

Versus 0.897 when trained on real CHAOS data.

4

Benchmarks

CHAOS, AVT, and Medical Segmentation Decathlon.

Qualitative results

Swipe to explore the full figure

Generated masks and consecutive CT or MRI slices for liver, aorta, and heart datasets.
Seven consecutive slices show stable inter-slice continuity while the generated masks follow the structure specified by each sketch-text pair.
Comparison with baselines

Swipe to explore the full figure

Comparison of Sketch2CT with Med-DDPM, MedGen3D, and Seg-Diff for liver CT, aorta CT, and heart MRI generation, alongside ground truth masks and images.
Across liver CT, aorta CT, and heart MRI, Sketch2CT more closely preserves the requested anatomy while producing images that resemble the corresponding ground truth.

Sketch2CT achieves the best scores on most benchmark settings and produces synthetic data whose downstream segmentation performance approaches models trained on real data. See the paper for complete comparisons with Med-DDPM, MedGen3D, and Seg-Diff.

Paper presentation

Sketch2CT in 4 minutes.

A concise walkthrough of the motivation, multimodal fusion design, two-stage diffusion pipeline, and experimental results.

04:23 CVPR 2026 English

Citation

Build on Sketch2CT.

If this work supports your research, please cite the CVPR 2026 paper.

BibTeX
@InProceedings{An_2026_CVPR,
  author    = {An, Delin and Wang, Chaoli},
  title     = {{Sketch2CT}: Multimodal Diffusion for Structure-Aware
               {3D} Medical Volume Generation},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer
               Vision and Pattern Recognition (CVPR)},
  month     = {June},
  year      = {2026},
  pages     = {37600--37610}
}