Speaker
Description
Step outside your office with a printed sheet of SDSS galaxy images and ask a random passerby to sort them into categories. Even without scientific training, most people can readily distinguish early- from late-
type galaxies, the same intuition that motivated Galaxy Zoo’s landmark citizen science initiative [Lintott et al., 2011]. This simple observation underscores a profound principle: the comparison between different
images of the same class of objects encodes physical information. It is precisely this idea that drove the machine learning community to develop contrastive learning [e.g., He et al., 2020, Chen et al., 2020, Radford
et al., 2021, ...].
For observational astrophysics, this approach offers a compelling advantage: it bypasses the labelling step that systematically introduces both epistemic and aleatoric errors into supervised frameworks.
The PHANGS collaboration, which surveys nearby galaxies at high angular resolution, has catalogued more than 100,000 stellar clusters observed across a wide range of instruments and wavelengths[Thilker et al., 2021]. Despite this rich dataset, photometry-based parameter inference remains fundamentally limited by degeneracies,
most notably between an intrinsically old red cluster and a young cluster heavily embedded in a dust cloud. Several classical methods and neural networks have been proposed to resolve these degeneracies, including
CNN [Viana et al., 2026], normalizing flows [Walter et al., 2026], and Rule-based decision tree [Thilker et al., 2025]. Yet all of these approaches remain dependent on labeled training sets of variable quality.
In this talk, we present a neural network trained in a self-supervised manner on multi-wavelength images from HST, JWST, and ALMA using VICReg (Variance-Invariance-Covariance Regularization)[Bardes
et al., 2022]. Our network ingests the full spatial information encoded in multi-wavelength image cutouts, simultaneously leveraging morphology, substructure, and color gradients across all available bands. Beyond its training simplicity and stability, VICReg is specifically designed to produce a decorrelated, interpretable
latent space, a particularly valuable property for disentangling the physical parameters underlying the observed cluster population. We first demonstrate how this latent space organizes stellar clusters in a physically meaningful way, naturally revealing the structure of known parameter degeneracies such as the age-extinction degeneracy, without any label supervision. We then show how fine-tuning this pre-trained model with a small set of carefully selected labeled examples enables the inference of key physical parameters including age, stellar mass, color excess E(B-V), and metallicity, with competitive accuracy and significantly reduced dependence on large labeled datasets. Finally, we discuss how this self-supervised representation can serve
as a bridge between simulations and observations, opening new avenues for inferring parameters that are inaccessible through classical SED fitting alone.