Local feature refinement
TSFE uses text-guided FiLM modulation to strengthen meaningful sketch features and suppress sparse or ambiguous contours.
Draw the shape. Describe the geometry. Generate the volume.
University of Notre Dame CVPR 2026 Main Conference Denver, Colorado · June 5–7Overview
Sketch2CT is a multimodal diffusion framework that turns a user-provided 2D sketch and a text description of 3D geometry into an anatomically coherent segmentation mask and its corresponding medical volume.
The sketch provides an intuitive structural blueprint, while text adds depth, topology, symmetry, and surface cues that a single projection cannot capture. A second latent diffusion stage translates the generated mask into realistic CT or MRI appearance while preserving the requested structure.
TSFE uses text-guided FiLM modulation to strengthen meaningful sketch features and suppress sparse or ambiguous contours.
CGFM combines cross-attention and self-attention to align sketch geometry with high-level textual semantics.
A mask-first pipeline preserves anatomy before synthesizing realistic, spatially continuous volumetric appearance.
Method
The pipeline separates anatomical reasoning from image synthesis, making the generated volume both controllable and structurally faithful.
Encode the sketch and geometric text, then align local and global cues.
Denoise in latent space to reconstruct a coherent 3D organ mask.
Use the segmentation latent to guide realistic CT or MRI generation.
Swipe to explore the full figure
Results
Evaluated on liver, aorta, and heart volumes across four public CT and MRI benchmarks at 128³ resolution.
33.7
FID ↓
Best image fidelity on CHAOS liver CT.
0.912
Mask Dice ↑
Highest mask faithfulness on Decathlon liver.
0.893
Downstream Dice ↑
Versus 0.897 when trained on real CHAOS data.
4
Benchmarks
CHAOS, AVT, and Medical Segmentation Decathlon.
Swipe to explore the full figure
Swipe to explore the full figure
Sketch2CT achieves the best scores on most benchmark settings and produces synthetic data whose downstream segmentation performance approaches models trained on real data. See the paper for complete comparisons with Med-DDPM, MedGen3D, and Seg-Diff.
Paper presentation
A concise walkthrough of the motivation, multimodal fusion design, two-stage diffusion pipeline, and experimental results.
Citation
If this work supports your research, please cite the CVPR 2026 paper.
@InProceedings{An_2026_CVPR,
author = {An, Delin and Wang, Chaoli},
title = {{Sketch2CT}: Multimodal Diffusion for Structure-Aware
{3D} Medical Volume Generation},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer
Vision and Pattern Recognition (CVPR)},
month = {June},
year = {2026},
pages = {37600--37610}
}