Project leads: Bau-Ching Hsieh (ASIAA), Yen-Ting Lin (ASIAA), and Keiichi Umetsu (ASIAA)
Point of contact: Bau-Ching Hsieh (bchsieh@asiaa.sinica.edu.tw)
Rubin project code: TAI-ASI-S4
Relevant working groups: photo-z
Project status: Active
Accurate estimates of galaxy redshifts, stellar masses, and star-formation rates are essential for nearly every extragalactic science case enabled by LSST, from galaxy evolution and environmental studies to weak lensing, cluster cosmology, and AGN demographics. With spectroscopic follow-up limited to a small fraction of the billions of galaxies LSST will detect, robust machine-learning-based photometric estimators are a core piece of infrastructure for the survey.
The Direct Empirical Photometric method (DEmP; Hsieh & Yee 2014) is a machine-learning code that derives these physical properties directly from multi-band photometry. It is one of the official photometric redshift codes adopted by the Subaru Hyper Suprime-Cam (HSC) survey (Tanaka et al. 2018; Nishizawa et al. 2020), and its HSC-based redshift, stellar mass, and SFR estimates have been used in numerous studies of galaxy environments, AGNs, groups and clusters, and gravitational lensing (e.g., Shirasaki et al. 2020; Chiu et al. 2021; Li et al. 2021; Tadaki et al. 2020; Jaelani et al. 2020; Jian et al. 2020; Sakakibara et al. 2019; Pintos-Castro et al. 2019).
DEmP2 is the next-generation implementation of this method, developed as part of the LSST in-kind contributions for the LSST Galaxy Science Collaboration. While the original DEmP code was written in Fortran, DEmP2 has been completely rewritten in Python, with significantly improved scaling capability, and has already been applied to produce the final HSC photo-z catalog. Beyond redshifts, DEmP2 can also derive additional physical properties such as stellar mass, star formation rate, and rest-frame color. DEmP2 natively outputs full probability distribution functions (PDFs) for each physical property, from which mean, mode, median, and best-fit point estimates, their associated confidence levels and risk flags, standard deviations, and 68%/95% confidence intervals can be derived.
For LSST, DEmP2 photo-z's have been derived for the DESC DC2 simulation using a realistic training set constructed with an HSC-like spectroscopic selection function (SpecSelection_HSC), so that the training sample reflects the kind of spectroscopic coverage actually achievable in real surveys. Integration with RAIL is currently underway, and near-term priorities include completing the RAIL implementation, participating in the DESC photo-z data challenge, and producing photo-z estimates for LSST DP1 and DP2.
DEmP2 photo-z's were derived for the DESC DC2 simulation using a training set of ~2M objects constructed with an HSC-like spectroscopic selection function (SpecSelection_HSC), so that the training sample reflects the kind of spectroscopic coverage achievable in real surveys. Under leave-one-out validation on this training set, DEmP2 achieves a bias of −0.0007, a scatter of σ_NMAD = 0.0182, and an outlier fraction of 0.0103. The corresponding PIT histogram is close to uniform, indicating that the predicted redshift PDFs are well calibrated over this training sample.
To assess performance under more realistic conditions, DEmP2 was also applied to a randomly selected sample of 50K DC2 objects, including all object types (galaxies, stars, and SNe), with no quality cuts other than a simple magnitude cut (r < 25), i.e., without removing objects with missing bands or low signal-to-noise detections. Under these conditions, DEmP2 achieves a bias of −0.0015, a scatter of σ_NMAD = 0.0346, and an outlier fraction of 0.0584. The corresponding PIT histogram shows mild excess near 0 and 1, indicating some miscalibration of the predicted PDFs for the lowest-quality and/or non-galaxy objects included in this unscreened sample.
These results indicate that DEmP2 performs well when applied to well-selected samples resembling realistic spectroscopic training sets, while performance naturally degrades when applied without quality cuts to a heterogeneous sample containing stars, supernovae, and low-S/N detections.
pFOF