Two-Fold Patch Perturbation for Efficient Self-Supervised Learning in 3D Medical Imaging
Two-Fold Patch Perturbation for Efficient Self-Supervised Learning in 3D Medical Imaging
Tirthajit Baruah, Kabir Jamadar, Punit Rathore
Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
AI and Health. Pages 6654-6662.
https://doi.org/10.24963/ijcai.2026/740
Self-supervised pre-training has become a key paradigm for reducing annotation costs in 3D medical imaging, yet many recent approaches rely on complex objectives or incur substantial computational overhead. We propose a simple and efficient self-supervised pre-training framework for 3D medical images based on a two-fold patch-wise perturbation strategy. The method applies Bernoulli patch masking and discrete rotations, and trains a shared encoder with a three-head objective for reconstruction, perturbation localization, and rotation prediction. This design encourages spatially aware and transferable representations while remaining computationally lightweight. Experiments across diverse segmentation and classification benchmarks, including modality-shift scenarios, demonstrate consistent improvements over general self-supervised baselines and competitive or superior performance compared to recent medical self-supervised methods, while requiring substantially less memory, computation, and training time than the state-of-the-art pre-training pipelines.
Keywords:
Self-supervised learning: Self-supervised learning
Medical imaging: Medical imaging
Medical diagnosis: Medical diagnosis
Medical knowledge representation: Medical knowledge representation
