← Week 37, 2026

2609.11879v1

Learning JWST. I. A Foundation Model for New Population Discoveries and Morphology-Aware Photometric Redshift Measurements in the JADES Survey

Theme match 4/5

Jiani Ding, Minghao Yue, Yongda Zhu, Xiaohui Fan, Yufeng Luo

First listed 2026-09-11 | Last updated 2026-09-10

Abstract

We present FM-JADES-v1, a self-supervised foundation model for James Webb Space Telescope ({\em JWST}) deep-field science, trained with 482,444 objects from the {\em JWST} Advanced Deep Extragalactic Survey (JADES) Data Release 5 using multi-band imaging and the photometric catalog. The shared embedding space is trained without class labels. We demonstrate that FM-JADES-v1 can serve as a powerful tool for object discovery and improving property measurements using two experiments, blind active discovery and few-band photometric redshift. For blind object discovery, FM-JADES-v1 identifies rare object populations such as high-redshift galaxies and Little Red Dots (LRDs) without any prior population labels or population-specific selection criteria. These rare populations emerge as isolated islands in the embedding space, which can be identified without prior astrophysical knowledge. For few-band photometric redshift, FM-JADES-v1's learned embeddings achieve $σ_{\rm NMAD}=0.157$ in a strictly controlled three-band (F115W/F200W/F356W) photo-$z$ benchmark, compared to $σ_{\rm NMAD}= 0.44$ for template fitting. These results demonstrate the potential of self-supervised multi-modal representations as scalable discovery spaces for large astronomical surveys. Applied to ongoing and future wide-field surveys from JWST, Roman, Euclid, and Rubin/LSST, this framework could enable systematic searches for rare populations, as well as enabling multiple downstream tasks such as improving astrophysical property measurements.

Short digest

FM-JADES-v1 is a self-supervised, cross-modal foundation model trained on 482,444 JADES DR5 sources, jointly encoding 16-band NIRCam image cutouts and catalog measurements into a shared morphology-plus-photometry embedding. Without labels or population-specific cuts, its embedding space isolates high-redshift galaxies and an LRD-rich population as distinct islands, providing a data-driven route to retrieve rare candidates rather than beginning from conventional selections. In a controlled three-band F115W/F200W/F356W photo-z test, the learned representation reaches σNMAD = 0.157 versus 0.44 for template fitting, showing that image morphology can materially recover redshift information when photometric coverage is sparse. For LRD work, the paper is chiefly a methods result: it positions foundation-model embeddings as a scalable discovery layer for finding and characterizing unusual early-universe populations in future wide surveys.

Key figures to inspect

  • Figure 1. This schematic establishes the paper's central technical claim: FM-JADES-v1 fuses separately tokenized multi-band imaging and catalog data in a shared Transformer embedding, then deploys that representation for both blind rare-population discovery and morphology-aware photometric redshifts. It is the most direct visual guide to how the LRD/high-redshift island search and the three-band photo-z improvement arise from the same trained model.

Discussion

Log in to view the paper discussion, see votes, and leave your own feedback.