Speaker
Description
Upcoming astronomical surveys are projected to produce petabyte-scale datasets, necessitating the development of intelligent, multimodal foundation models to accelerate scientific insight. While traditional data analysis often treats observational products and scientific literature as isolated domains, this work presents a novel contrastive learning pipeline that aligns Chandra X-ray spectra with natural language descriptions from scientific papers. Our framework uses a transformer-based autoencoder to compress high-dimensional spectral count rates into compact representations, which are then aligned with scientific paper summaries via an InfoNCE contrastive loss. This process establishes a shared multimodal latent space that effectively captures critical physical properties, including hardness ratios and column density, while reducing total data dimensionality by 97%. We demonstrate the scientific utility of this alignment through cross-modal retrieval and unsupervised outlier detection. By analyzing anomalies in the aligned latent space, we successfully isolated rare astrophysical phenomena, including a gravitational lens system and a promising candidate pulsating ultraluminous X-ray source (PULX). These results highlight the potential for intelligent multimodal systems to serve as powerful tools for autonomous discovery and language-driven exploration of large-scale survey data in the impending Big Data era.