Large-scale-structure clustering & ML survey-systematics diagnostics on real DESI DR1 data
research/claims_registry.md. Full scope and limitations:
Limitations.
CosmoScope measures real galaxy clustering from DESI DR1 and tests whether machine-learning correction methods for observational (imaging) systematics can remove contamination from galaxy angular density without erasing genuine large-scale-structure clustering signal, compared to the linear-regression correction standard in the survey literature.
DESI DR1 LSS catalogs (Iron reduction, v1.5), LRG tracer, SGC region. 662,492 objects 3772 deg² (9.14% of sky)
Data: CC BY 4.0. This research used data obtained with the Dark Energy Spectroscopic Instrument (DESI). DESI Collaboration et al. 2026, arXiv:2503.14745.
Real-space two-point correlation function ξ(r), Landy-Szalay estimator, spatial- jackknife uncertainties, fiducial cosmology from Ross et al. 2025 (arXiv:2411.12020; H₀=67.36, Ωₘ=0.3137721).
Power-law fit (ξ(r) = (r₀/r)^γ): over 1.2–30 Mpc, r₀ = 9.14 ± 0.04 Mpc/h, γ = 1.545 ± 0.009. Restricting to the two-halo regime (3.6–30 Mpc, avoiding the known one-halo/two-halo transition) gives γ = 1.673 ± 0.011, closer to but still below the literature γ≈1.8–2.0 for LRG samples (Zehavi et al. 2005, arXiv:astro-ph/0411557) — see limitations.md for the honest discussion of this residual discrepancy.
Controlled injection study on the real LRG SGC footprint: a mock density field with known clustering, contaminated using REAL DESI imaging-systematics templates (EBV, stellar density, imaging depth, PSF size) at known coefficients, corrected with linear regression vs. random forest vs. gradient boosting, all fit via spatial cross-validation. Every realization includes a negative control (every method applied to the never-contaminated true field).
Canonical run: 10 mock realizations, HEALPix nside=128, 5-fold spatial cross-validation. Values are means across realizations.
| Method | Residual template correlation (contaminated, lower=better) | RMS Cℓ distortion (contaminated) | RMS Cℓ distortion (negative control, lower=better) |
|---|---|---|---|
| Gradient boosting | 0.0273 | 0.0785 | 0.0565 |
| Linear (survey-standard) | -0.0030 | 0.0997 | 0.0507 |
| No correction | 0.5998 | 1.9134 | 0.0000 |
| Random forest (deep) | 0.0058 | 0.0902 | 0.0555 |
| Random forest (shallow) | 0.0847 | 0.1585 | 0.0492 |
Headline result: on this injection design, gradient boosting and deep random forest match or modestly outperform linear regression at removing injected contamination, with no corresponding signal-preservation penalty — negative-control distortion is approximately uniform across every method (0.049–0.057). This reverses the impression from a smaller pilot run of the same pipeline, which had suggested flexible methods overcorrect substantially more; see decision_log.md for the full comparison. A separate out-of-distribution test complicates this in a different way: on sky regions with unusual stellar density, gradient boosting's error inflates ~7.9× vs. linear regression's ~1.9× relative to in-distribution performance — tree ensembles cannot extrapolate beyond their training range the way linear regression can. See the paper for the full breakdown.
Robustness (nside=128, 10 realizations, linear correction only — decision_log.md): coarsening systematics templates from native nside=128 to nside=16 degrades correction quality roughly 8×. Dropping either strongly-injected template (EBV, STARDENS) from the correction design matrix degrades correction in every realization; dropping any non-injected template leaves the typical case unchanged but occasionally destabilizes one spatial fold's fit (near-collinearity among real templates). Reducing to only the true contaminating templates gives the best typical-case correction but a real numerical-stability tail risk under spatial cross-validation — reported as a finding, not smoothed over.
Full methodology, literature grounding, and design decisions: see the repository's
research/ directory
(questions.md,
data_provenance.md,
cosmology_assumptions.md,
literature_matrix.csv,
decision_log.md) and the
paper draft.
This is a single-tracer (LRG), single-region (SGC) study — not a full-survey precision measurement. The injection-study mock is a simplified lognormal angular field, not a full N-body/HOD forward model. Injected contamination is linear-in- templates by construction, which structurally favors the linear correction method — ML correction matched or modestly beat it anyway (see above), but whether ML's advantage would be larger on genuinely nonlinear real-world contamination remains untested. The negative control fits on the noiseless true field, not a Poisson-noise-matched one, so its distortion numbers are a lower bound on real-data behavior. Our fitted γ = 1.673 (two-halo regime) remains shallower than the literature's γ≈1.8–2.0 even after excluding the one-halo transition — a genuine, unresolved discrepancy, not silently corrected. Full list: research/limitations.md.