Unpaired Image-to-Image Translation · CycleGAN · SynthRAD2023
44,882 unpaired brain slices. Two generators locked in a cycle. No paired example ever seen.
MRI is the clinical gold standard for soft-tissue contrast—stroke, tumour, demyelination become visible. But MRI is expensive, slow, contraindicated with metal implants, and scarce in lower-resource settings. CT is fast, cheap, and everywhere.
Bone, calcification, acute haemorrhage. Soft-tissue contrast is limited.
Grey matter, white matter, tumours, oedema—the structures clinicians need most.
Can a network learn the CT → MRI mapping without ever being shown a matched pair?
Rigidly registered CT / T1-MRI volumes from a genuinely heterogeneous cohort. Centre A: non-contrast, mixed 1.5T/3T. Centres B & C: gadolinium-enhanced. Three centres, three acquisition protocols, one shared anatomy.
| Split | Patients | Ratio | Purpose |
|---|---|---|---|
| Train | 126 | 70% | 22,441 CT + 22,441 MRI (unpaired, shuffled) |
| Validation | 18 | 10% | 3,093 filename-matched pairs |
| Test | 36 | 20% | 6,378 slices — unlocked once, after all decisions final |
22,441 CT + 22,441 MRI (unpaired, shuffled)
3,093 filename-matched pairs
6,378 slices — unlocked once, after all decisions final
Training folders are shuffled independently—the model never sees a matched pair. The paired test set stays locked behind a hard flag (Constraint C5) until every hyperparameter, every design choice, and every checkpoint decision is final. One unlock. One evaluation. No second chances.
Each raw NIfTI volume walks through a six-stage pipeline before the network ever sees it.
Why per-patient MRI normalisation? Inter-centre intensity ratio was ~3.7×—global min-max normalisation killed contrast for two of three centres. No rotation or intensity jitter: those transforms would corrupt clinically meaningful signal.
Two generators, two discriminators, one closed loop. Cycle-consistency is what replaces paired supervision—“if you translate to MRI and back, you must land on the CT you started from.”
λ = 1
Least-squares objective for stable training. Two PatchGAN discriminators, one per domain.
λ = 10
Translate and back—you must recover what you started with. This is what replaces paired data.
λid = 5 (independent)
Feed an MRI to G_AB; it should return the MRI unchanged. Prevents unnecessary hallucination.
6 residual blocks, not the paper’s 9—Zhu et al.’s own documented low-compute configuration.
Kaggle T4, mixed precision, every trick in the book to get from 3 hours per epoch down to 15 minutes.
The unclamped learning-rate schedule went negative. Adam didn’t minimise loss—it maximised it.
Resolution: Checkpointing was independent—epoch 49 (the best) was untouched. Fixed with a floor clamp + assertions.
No cherry-picking—these are the five cases from the qualitative grid, spanning the full SSIM distribution.
PSNR is normalisation-sensitive and confounded across studies. A contemporaneous CT→MRI preprint lands at SSIM 0.60 / PSNR 15.3—mid-teens PSNR looks characteristic of this translation direction. CycleSynth beats it on both metrics.
The worst test cases cluster around a single patient. 5 of the 6 lowest SSIM slices belong to 1BC056. Generated MRI is systematically darker—a negative intensity shift.
Input CT
Generated
Ground Truth
SSIM 0.438 · MAE 26.5
A rectangular, axis-aligned black box—present in the ground-truth MRI only. Never in the CT. Never in our output. A defacing artefact inherited from the source dataset, proven upstream of our pipeline.
Padding and clipping cannot manufacture a one-modality rectangle. Intensity statistics ruled out the normalisation hypothesis. The artefact is present in the raw SynthRAD2023 data—our pipeline faithfully reproduces it.
Two automated approaches were tried. Border-merge detection missed the artefact. Ventricle-based detection produced false positives that raised SSIM when the “cleaned” slices were removed. Sometimes the right answer is to document the problem honestly.
Quantitative metrics only. No clinical evaluation of diagnostic utility has been performed.
SynthRAD2023 only. No external validation on out-of-distribution data from other institutions.
No volumetric context. Each slice is translated independently, blind to its neighbours.
CycleSynth produces a learned approximation of MRI appearance, not physical ground truth.
VGG perceptual loss (from the SynthRAD2023 winner’s approach) targeted at the texture/PSNR gap.
Resume training from epoch 49 with the clamped scheduler. More epochs, safely.