Abstract Head computed tomography (CT) imaging is a widely used imaging modality with multitudes of medical indications, particularly in assessing pathology of the brain, skull and cerebrovascular system. It is commonly used as the first-line imaging in neurologic emergencies given its rapidity of image acquisition, safety, cost and ubiquity. Deep learning models may facilitate detection of a wide range of diseases. However, the scarcity of high-quality labels and annotations, particularly among less common conditions, substantially hinders the development of powerful models. To address this challenge, we introduce FM-HCT, a Foundation Model for Head CT for generalizable disease detection, trained using self-supervised learning. Our approach pretrains a deep learning model on a large, diverse dataset of 361,663 non-contrast 3D head CT scans without the need for manual annotations, enabling the model to learn robust, generalizable features. Our results demonstrate that the self-supervised foundation model substantially improves performance on downstream diagnostic tasks compared to models trained from scratch and previous 3D CT foundation models trained on scarce annotated datasets. Similar content being viewed by others Main Head computed tomography (CT) is widely used for rapid evaluation of neurological emergencies such as trauma, haemorrhage and stroke. Although CT is faster, more accessible and less expensive than magnetic resonance imaging (MRI), it provides lower soft-tissue contrast, limiting sensitivity for many neurological conditions. Improving diagnostic capability from CT therefore provides substantial clinical value. Artificial intelligence (AI) has the potential to enhance CT interpretation and support clinical decision-making by enabling earlier and more accurate diagnosis. However, progress in AI-based head CT analysis remains limited by both data availability and model design. Public datasets such as RSNA1 and CQ500 (refs. 2,3) are relatively small and primarily focus on haemorrhage detection, restricting broader clinical applicability. In addition, many existing approaches rely on two-dimensional (2D) convolutional networks that process