Unpaired Image-to-Image Translation · CycleGAN · SynthRAD2023

CycleSynth Teaching a network to see soft tissue where CT cannot.

44,882 unpaired brain slices. Two generators locked in a cycle. No paired example ever seen.

Scroll to translate ↓
Input CT scan Generated MRI
S2 · Why

Why translate CT to MRI?

MRI is the clinical gold standard for soft-tissue contrast—stroke, tumour, demyelination become visible. But MRI is expensive, slow, contraindicated with metal implants, and scarce in lower-resource settings. CT is fast, cheap, and everywhere.

CT scan showing bone contrast

What CT sees

Bone, calcification, acute haemorrhage. Soft-tissue contrast is limited.

MRI scan showing soft tissue

What MRI sees

Grey matter, white matter, tumours, oedema—the structures clinicians need most.

Can a network learn the CT → MRI mapping without ever being shown a matched pair?

S3 · The Data

SynthRAD2023

Rigidly registered CT / T1-MRI volumes from a genuinely heterogeneous cohort. Centre A: non-contrast, mixed 1.5T/3T. Centres B & C: gadolinium-enhanced. Three centres, three acquisition protocols, one shared anatomy.

0Patients
0Centres
0Training slices
0Test slices (locked)
CT sample 1 MRI sample 1 CT sample 2 MRI sample 2 CT sample 3 MRI sample 3

Patient-level split (centre-stratified)

SplitPatientsRatioPurpose
Train12670%22,441 CT + 22,441 MRI (unpaired, shuffled)
Validation1810%3,093 filename-matched pairs
Test3620%6,378 slices — unlocked once, after all decisions final

Train

Patients126
Ratio70%
Purpose

22,441 CT + 22,441 MRI (unpaired, shuffled)

Validation

Patients18
Ratio10%
Purpose

3,093 filename-matched pairs

Test

Patients36
Ratio20%
Purpose

6,378 slices — unlocked once, after all decisions final

The secretly paired test set

Training folders are shuffled independently—the model never sees a matched pair. The paired test set stays locked behind a hard flag (Constraint C5) until every hyperparameter, every design choice, and every checkpoint decision is final. One unlock. One evaluation. No second chances.

S4 · Preprocessing

From Volumes to Training Slices

Each raw NIfTI volume walks through a six-stage pipeline before the network ever sees it.

Raw Slice
Axial from .nii.gz
Mask Filter
≥ 2,000 voxels
Zero-pad
to square
Intensity Window
CT: [0, 1752.8] HU (global)
MRI: [0, p99.5] (per-patient)
-1+1
Scale
[-1, 1]
Augment
Upscale → random 128 crop → h-flip

Why per-patient MRI normalisation? Inter-centre intensity ratio was ~3.7×—global min-max normalisation killed contrast for two of three centres. No rotation or intensity jitter: those transforms would corrupt clinically meaningful signal.

S5 · Architecture

The Cycle

Two generators, two discriminators, one closed loop. Cycle-consistency is what replaces paired supervision—“if you translate to MRI and back, you must land on the CT you started from.”

CT real_A G_AB fake MRI fake_B G_BA recon CT rec_A cycle-consistency: ||rec_A - real_A||₁ D_B 70x70 Patch D_A 70x70 Patch + identity loss

LSGAN adversarial

λ = 1

Least-squares objective for stable training. Two PatchGAN discriminators, one per domain.

Cycle-consistency L1

λ = 10

Translate and back—you must recover what you started with. This is what replaces paired data.

Identity L1

λid = 5 (independent)

Feed an MRI to G_AB; it should return the MRI unchanged. Prevents unnecessary hallucination.

Adam 2e-4 β 0.5 / 0.999 batch 8 128 × 128 px 6 ResBlocks InstanceNorm reflection padding image pool 50

6 residual blocks, not the paper’s 9—Zhu et al.’s own documented low-compute configuration.

S6 · Training

58 Epochs on a Free GPU

Kaggle T4, mixed precision, every trick in the book to get from 3 hours per epoch down to 15 minutes.

1.36 → 0.113 s/it
iteration time
~3h → ~15min
per epoch
256→128 + 9→6
resolution & blocks
AMP 1.71×
mixed-precision speedup
batch 2→8
batch size
channels_last
tested, 30% slower, rejected

Validation SSIM

Epoch 56: the model tried to destroy itself.

The unclamped learning-rate schedule went negative. Adam didn’t minimise loss—it maximised it.

1.76
loss at epoch 55
0
loss at epoch 56
0
loss at epoch 57

Resolution: Checkpointing was independent—epoch 49 (the best) was untouched. Fixed with a floor clamp + assertions.

S8 · Results

The Payoff

SSIM
0
98.4% of supervised · no pairs
MAE
0
beats supervised UNet (18.29)
PSNR
0
trails baselines — see honest note

Baseline Comparison

SSIM ↑

CycleSynth (ours)
0.6764
Supervised UNet
0.687
Prior CycleGAN
0.467

MAE ↓

CycleSynth (ours)
16.81
Supervised UNet
18.29
Prior CycleGAN
23.55

Input CT vs Generated MRI

Input CT Generated MRI
CT Generated

Generated vs Ground Truth

No cherry-picking—these are the five cases from the qualitative grid, spanning the full SSIM distribution.

Generated Ground Truth
Gen GT
SSIM 0.438
Generated Ground Truth
Gen GT
SSIM 0.627
Generated Ground Truth
Gen GT
SSIM 0.664
Generated Ground Truth
Gen GT
SSIM 0.716
Generated Ground Truth
Gen GT
SSIM 0.930

A note on PSNR

PSNR is normalisation-sensitive and confounded across studies. A contemporaneous CT→MRI preprint lands at SSIM 0.60 / PSNR 15.3—mid-teens PSNR looks characteristic of this translation direction. CycleSynth beats it on both metrics.

val 0.6780 → test 0.6764 — no leakage, no overfit to validation.

Where It Breaks

The worst test cases cluster around a single patient. 5 of the 6 lowest SSIM slices belong to 1BC056. Generated MRI is systematically darker—a negative intensity shift.

Worst case CT

Input CT

Worst case Generated

Generated

Worst case Ground Truth

Ground Truth

SSIM 0.438 · MAE 26.5

Then we looked closer at the ground truth…

A rectangular, axis-aligned black box—present in the ground-truth MRI only. Never in the CT. Never in our output. A defacing artefact inherited from the source dataset, proven upstream of our pipeline.

The evidence

Padding and clipping cannot manufacture a one-modality rectangle. Intensity statistics ruled out the normalisation hypothesis. The artefact is present in the raw SynthRAD2023 data—our pipeline faithfully reproduces it.

The failed detectors

Two automated approaches were tried. Border-merge detection missed the artefact. Ventricle-based detection produced false positives that raised SSIM when the “cleaned” slices were removed. Sometimes the right answer is to document the problem honestly.

S10 · Limitations & Future

What Comes Next

No radiologist review

Quantitative metrics only. No clinical evaluation of diagnostic utility has been performed.

Single dataset

SynthRAD2023 only. No external validation on out-of-distribution data from other institutions.

128×128 2D slices

No volumetric context. Each slice is translated independently, blind to its neighbours.

Appearance ≠ truth

CycleSynth produces a learned approximation of MRI appearance, not physical ground truth.

Future: perceptual loss

VGG perceptual loss (from the SynthRAD2023 winner’s approach) targeted at the texture/PSNR gap.

Future: bounded continuation

Resume training from epoch 49 with the clamped scheduler. More epochs, safely.