Speaker
Description
Vision Transformers (ViTs) are increasingly being adopted for large-scale astronomical imaging, yet discussion of performance often centers on model scale and dataset size rather than on the quality and structure of the training data themselves. In radio astronomy, this is a significant omission: images are frequently sparse, background-dominated, and affected by instrumental artefacts, tiling patterns, and large variation in source extent. Under such conditions, self-attention may be steered toward non-physical or weakly informative structure unless the training distribution is carefully curated. We therefore present a data-centric study of ViT training for radio astronomy imaging, aimed at quantifying how curation strategy influences downstream behaviour.
Our starting point is a strong domain-adapted ViT baseline trained with a two-stage self-supervised procedure and curation-aware view generation that preferentially samples informative source regions. From this reference point, we isolate the effect of data curation through a sequence of controlled interventions. Specifically, we compare the baseline against: (i) training without object-centric cropping and with minimal augmentation, (ii) training on a pre-filtered dataset with substantially fewer empty cutouts, (iii) training on a more tightly curated object-centric dataset that reduces the need for online selection of informative regions, and (iv) training on a curated dataset rebalanced toward larger and more extended sources. These experiments are designed to test whether improved curation quality and morphological representativeness can compensate for weaker online selection strategies or larger nominal data volume.
We assess the resulting models on downstream radio morphology classification tasks, considering predictive performance, calibration, and robustness across heterogeneous imaging regimes. In addition, we examine internal attention behaviour to determine whether stronger curation reduces reliance on spurious non-physical structure such as empty regions, artefacts, and image-boundary effects. The goal is not merely to compare preprocessing choices, but to establish data curation as a first-order design variable in transformer-based scientific imaging. More broadly, this study aims to clarify when training-set quality, rather than dataset scale alone, governs the reliability of ViT models in radio astronomy.