Executive Summary
The Vera C. Rubin Observatory began its ten-year survey in June 2026, the Nancy Grace Roman Space Telescope is set to launch no earlier than August 30, and Euclid is already delivering space-resolution imaging over thousands of square degrees. By 2027 all three will be covering much of the same sky at once. A community perspective paper signed by 27 researchers, posted to arXiv on August 18, argues that what remains unsolved at this point is not the telescopes.
The paper puts it plainly. The observations are no longer the bottleneck, and realizing the joint scientific return is now an engineering and institutional challenge. Combined US and European public investment in the three surveys exceeds $6 billion. Yet processing that data at the pixel level together, calibrating each survey against the others, and serving it through one door falls inside no mission's or agency's mandate. No funded programme exists, the authors write, to do this on real data.
The heavier claim comes next. No institution here can solve the problem alone, and the reason is mandate rather than capability, because each one is already doing exactly what it was funded to do. Budget attaches to collection and not to alignment, and closing that gap later costs far more than aligning now. Both patterns repeat, at a smaller scale, inside industrial data organizations.
Key Numbers
Source: Rau et al., arXiv:2608.19272 (2026-08-18)
The four numbers below measure different things. The first two size what has already been built, and the last two size what has not been built against it.
Over $6 billion
Combined public investment in the three surveys
US and European public funds, spent on the telescopes and on operating the observations
More than half
Share of sources blended at full Rubin depth
Ground-based imaging merges neighbouring objects into one another, and space resolution is what pulls them apart
Zero
Funded programmes to run joint processing on real data
The algorithms are proven in isolated cases and simulated joint products exist, but nothing is resourced to run at survey scale when the data arrive
330 arcmin²
Area of the GOODS precedent in the early 2000s
Even that was an extraordinary computational and human undertaking, and the scale needed now is thousands of square degrees
The telescopes are no longer the bottleneck
The three surveys are good at different things. Rubin is a ground-based telescope built by the US National Science Foundation and the Department of Energy, and it sweeps roughly 18,000 square degrees of sky repeatedly to produce wide, frequent optical coverage. Euclid is a European Space Agency mission already returning imaging at a resolution the ground cannot reach, over thousands of square degrees. Roman is NASA's near-infrared space telescope, and it goes deeper over a smaller area. Optical and near-infrared, breadth and resolution, each covering what the other lacks.
The paper insists on one premise. No project here is failing. Rubin's data ecosystem is built for community access at unprecedented scale, Roman exceeds its technical requirements in many key aspects of the observatory, and Euclid is already delivering the imaging it promised. Each team is producing exactly what it was scoped, funded, and asked to produce. Those designs, though, predate the full scale of the scientific opportunity now emerging where the three datasets meet.
Looking at one object across several wavelengths is not itself new. From the 1960s through the 1980s, multiwavelength astronomy established the need. From the 1990s through the 2010s, the Great Observatories and the deep-field campaigns turned it into a scheduling problem. What the rest of the 2020s and the 2030s demand is different in kind. Not coordinating observations across facilities, but processing petabyte-scale surveys jointly, at the pixel level.
GOODS gives the difference a scale. In the early 2000s it used high-resolution Hubble imaging as a positional and morphological prior to extract consistent photometry from Spitzer's far coarser pixels, and photometric scatter and catastrophic redshift outliers both dropped measurably. That was done over a mere 330 arcmin², and even so it was an extraordinary computational and human undertaking. What is needed now is genuine simultaneous joint processing over thousands of square degrees, not one survey's resolution informing extraction from another. That is why the authors call it not an incremental extension of GOODS but a different class of challenge.
The diagram below locates the gap this article is about. The three boxes on top are built and funded. The dashed box in the middle is the one nobody disputes the need for and nobody owns.
Every expert was in the room, and the work belonged to no one
Start with the most natural fix. Could one survey simply ingest another survey's data into its own pipeline? The paper says this runs into a structural limit. Each survey's data is best understood by the team that built its pipeline. A project processing another's data works without that expertise, and recreates the same gap under a different name.
Instrumental artifacts make this concrete. Persistence, cosmic-ray afterglows, snowballs, stray light: each detector leaves its own signature, and each team knows only its own. No external group can reconstruct that knowledge reliably. Consistent flagging across surveys therefore requires common bit conventions and tools trained on each survey's own known artefacts, which comes down to pooling the expertise in one place.
The case the paper reaches for here is its strongest evidence. DES and KiDS are two weak-lensing collaborations. They performed the first combined catalogue-level analysis between two such collaborations, with both teams involved. Uniform systematics calibration was still ruled out of scope. Not for lack of expertise, the authors write, but because the work sat outside what either collaboration was scoped to do.
What this case describes is a mandate problem rather than a people problem. Put the best experts in one room, and if the work is assigned to nobody in that room, the work stays undone. The same thing is what keeps a cross-department join untouched for years in an industrial data organization. Each department is the expert on its own data, and the seam between departments appears in nobody's objectives.
Funding structures push the same way. Agencies and projects meet the requirements each survey was defined to meet first, and put resources into activity outside that scope second. That is a reasonable priority order. The trouble is that proving the value of joint science requires exactly the out-of-scope work. Research that draws equally on datasets and infrastructure from several agencies fits no single programme's review criteria neatly. The paper describes the result as an area of shared interest but diffuse ownership.
Nor is the gap a lack of demand. With no standardized option available, scientists adapt single-survey pipelines and build custom cross-matches of their own. IRSA publishes tutorials for retrieving matched cutouts from simulated Roman and Rubin fields. The paper's verdict is short. It works, and it works by hand, one dataset at a time.
Leave them uncombined and the discovery never happens
For some science, the absence of joint processing costs a little precision and no more. Cases needing only photometry, lightcurves, or cutouts of matched objects fall here, and the planned public data products or straightforward extensions to them will serve. The problem is everything else.
Weak lensing is the flagship example. Rubin measures the faint distortion of galaxy shapes over roughly 18,000 square degrees, but its photometric redshifts suffer catastrophic outliers exactly where it matters, and ground-based blending biases the shear. Roman's near-infrared photometry significantly improves the photo-z problem, and its stable space-based point spread function offers a route to calibrating critical ground-based systematics at the pixel level. The paper marks this as a flagship case that is impossible from the planned public products, because reaching maximum depth, deblending sources, and obtaining consistent photometry and shapes across two surveys all require joint pixel-level detection and modeling.
One number gives the scale. At full Rubin depth, more than half of all sources appear confused or merged with a neighbour. No amount of ground-based exposure resolves that blur. Euclid's and Roman's space resolution is what separates that half.
The Milky Way tells the same story from the opposite direction. At high Galactic latitude, independent Roman and Rubin catalogs can reasonably be cross-matched after the fact. In the Galactic plane, where most of the Galaxy's stars and dust reside, that approach breaks down. Extreme crowding and differential extinction between the optical and near-infrared bands make catalog cross-matching ambiguous and compromise the underlying crowded Rubin photometry. Taking full advantage of Rubin's optical photometry there requires joint photometry from joint pixel-level processing.
In time-critical observing, the opportunity disappears entirely. A microlensing event happens once and does not repeat, and the parallax baseline between simultaneous observers is the measurement, so nothing is recoverable later by reconciling independent catalogues. Roman and Euclid watching the Galactic bulge together provide a space-based baseline that resolves the lens mass and distance degeneracy directly, enabling mass measurements of free-floating planets. High-redshift supernovae work the same way. Deep Rubin non-detections placed alongside Roman detections are what distinguish the real events from lower-redshift contaminants.
The paper pushes hardest right here. At the limit, the survey combination does not improve a measurement. It is what makes the discovery possible at all.
On the time-domain side, each survey already has its own infrastructure. For Roman, a Caltech and IPAC team is building RAPID, which will deliver prompt image differencing, a public alert stream of transient and variable candidates, and forced-photometry services. For Rubin, the LSST alert stream and the community brokers are already operating. What the paper points at is that both charters are, deliberately, single-survey. The layer that merges the two streams and enriches each alert with the other survey's lightcurve history and host-galaxy context sits inside neither scope. The paper also records how narrow that gap is. Simply exposing RAPID's forced photometry for query by Rubin brokers would let a Rubin alert position be checked against Roman's photometric history and fed directly into classification. The service is being built for Roman alone, and the cross-survey interface is the small, unfunded step.
Pebblous looked in July at what to trust when filtering the alert torrent inside a single survey. That piece was about Rubin's internal problem. This one is about the space between three surveys.
The corpus came first, the models followed
The second of the four pillars is the data substrate for machine learning. The paper treats a uniformly processed, large-scale training substrate as the most valuable by-product of joint processing. Rubin, Roman and Euclid combined could supply billions of sources across optical and near-infrared, variability over time, and different spatial resolutions, all on the same sky. No project is yet scoped to build that corpus.
Half the challenge, standardization, has already been demonstrated. The Multimodal Universe, an open dataset of roughly 100 TB hosted at the Flatiron Institute, brings images, spectra and time series from many surveys into a common schema with benchmark tasks attached. It reproduces the pattern ImageNet showed: when a curated, benchmarked corpus becomes available, a modeling community rapidly forms around it. The dataset also records its own limit. Standardizing existing survey products leaves the cross-matched sample bounded by each survey's depth, footprint, and the degree to which those footprints overlap.
The analogy the paper draws is protein structure. AlphaFold's breakthrough depended on nearly fifty years of experimentally determined structures deposited in a single, standardized, openly accessible repository. The corpus came first, the models followed, and astronomy has yet to build its equivalent.
The substrate is missing for architectural reasons, not scientific ones. There is no front door. Separate archives, separate formats, no common identifier, no benchmark, no clean API. Pulling one matched cutout means authenticating against each archive separately and reconciling coordinate conventions by hand. This is not a storage problem, and physical co-location of the archives is neither realistic nor necessary.
It is also a question of access. Today only well-resourced teams can practically use a multi-petabyte scientific dataset. The paper's proposal is to expose Roman, Rubin and Euclid data in interoperable cloud regions, and to place shared low-cost compute beside it, such as the Open Science Grid or facilities like NOIRLab's Astro Data Lab and ESA Datalabs. That lets researchers at smaller institutions analyze the data without prohibitive computing costs. The pattern is already being shown to work. The Roman High-Latitude Imaging Survey Project Infrastructure Team is demonstrating it with simulated data: low-level data pulled from the cloud, processed at a university, higher-level products returned for archiving.
The paper also carries evidence that a solution is not far off. HATS and LSDB, built by the LINCC Frameworks team, partition catalogs by position on the sky so they can be queried and cross-matched efficiently, and they already work on billion-source catalogs and support Rubin Data Preview 1. On the alert-stream side, the LSST and ZTF streams are being cross-matched in real time. The community workshops concluded that what remains is a question of design, not of technology.
Read a certain way, this passage is almost indistinguishable from the industrial version. Most organizations adopting foundation models pick the model first. The in-house corpus that model would train and be evaluated on, meaning records from different sources sitting on the same identifiers and the same coordinates, often does not exist yet.
Aligning now is cheaper than reconciling later
The paper asks for four things. Each has already been specified in community studies and proven at prototype level, and not one of them is funded or owned as a shared capability.
| Pillar | What exists today | What is missing |
|---|---|---|
| Joint pixel processing and validation | Algorithms, isolated cases, simulated joint products | Survey-scale validation and a budget to run it on real data |
| AI-ready data substrate | Multimodal Universe at ~100 TB, common schema | A corpus processed jointly on the same sky |
| Interoperable data access | IVOA standards, HATS and LSDB, per-survey platforms | A common front door across all three surveys |
| People and career pathways | Software posts inside facilities and projects | Positions that span surveys and agencies |
▲ The four pillars of the paper's Section 3, sorted into what exists and what is missing
For the first pillar the paper names a starting point. A concrete first step could be a joint pixel-level processing demonstration, with Roman and Rubin stamps processed jointly in a shared deep field such as the Euclid Deep Field South, where all three surveys overlap. The distance to that step is shorter than it sounds, because NASA's OpenUniverse2024 has already produced matched Rubin and Roman simulations and shown that such coordination is achievable at scale. That effort is scoped to simulation tools and infrastructure, though, and does not address joint processing of real survey data. A full Rubin, Roman and Euclid simulation suite does not yet exist either. The authors attach a condition: design for extension from the start. SPHEREx, UVEX and future surveys should be able to join the same framework rather than requiring new bilateral reconciliations each time a pair is added.
The fourth pillar is about people. No established career path is dedicated to cross-survey infrastructure. The talent required spans astronomy, software engineering and machine learning, and people who sit across those fields fall between every agency's mandate and every university's incentive structure. Universities have historically rewarded individual publications over tool-building work. Even the field's most successful software-infrastructure efforts are funded as fixed-term projects, so the strongest people at this intersection live appointment to appointment, by design. The authors' warning is explicit. On the present course, the generation of talent this needs leaves for fields that offer a path.
That leaves the question of why now. The answer is an asymmetry in cost. The three surveys are still early in their life cycles, so foundational choices such as data and metadata formats and coadd projections can still be aligned by design. Aligning now costs one agreement on conventions. Letting the choices diverge and harden means reprocessing at survey scale to undo them. The paper puts the difference at a fraction of the cost of reconciling later, and says the window is open now, will not stay open, and will not return.
On whether institutions can catch that window, one precedent is cited. TCAN was established jointly by NSF/AST and NASA/APD in direct response to Astro2010's finding that no mechanism supported sustained multi-institutional theory collaborations. The mechanism itself worked. Its trajectory is also instructive, because the joint solicitation lapsed while NASA continued the program alone. Mechanisms of this kind work, the paper reads, but without durable standing and resourcing they depend on the continued attention of individuals in both agencies.
So Section 4 of the paper splits the work by actor. Policymakers and federal agencies are asked to make cross-agency, PI-scale proposals actually submittable, to remove the barriers that force single-survey defaults, and to fund the software and processing layer as strategic infrastructure protected from diversion when budgets tighten. The example given for the second item is Roman's default photometric redshifts, which are derived from Roman data alone rather than jointly utilizing optical photometry. Observatories and operators are asked for public, jointly analyzable deep fields and for coordinated observation planning. Universities and computing facilities are asked to fund the people and not just the code. The community is asked to coordinate through platforms that already exist instead of building another one. The sharpest sentence sits in the first list. Every capability in Section 3 is currently additional to someone's actual job, and work that is nobody's primary responsibility advances at the pace of spare time.
Editor's Note: what Pebblous sees daily in industrial data has the same shape as this diagnosis. Budgets to collect more data generally get approved. Putting records from different sources onto the same coordinates gets no owner assigned. Collection counts as a department's result, and alignment counts as nobody's. What makes the astronomy case useful is that the cost asymmetry is written in units of $6 billion and petabytes, which makes the gap between aligning now and reconciling later visible at a glance.
The authors call their own text a community perspective article. Instead of presenting new observations, it records why work recommended repeatedly for decades still has not been built. One sentence sits in the conclusions. What is new is not that flagship surveys will observe the same sky concurrently, but that their greatest discoveries will be revealed across the seams. The paper is at arXiv:2608.19272.
References
Primary Source
- 1.Rau, G., et al. (2026). "The Cross-Survey Decade: A Call to Action." arXiv:2608.19272.
Academic Papers
- 2.Abbott, T. M. C., et al. (2023). "DES Y3 + KiDS-1000: Consistent Cosmology Combining Cosmic Shear Surveys." The Open Journal of Astrophysics, 6.
- 3.The Multimodal Universe Collaboration. (2024). "The Multimodal Universe: enabling large-scale machine learning with 100 TB of astronomical scientific data." NeurIPS Datasets and Benchmarks.
- 4.Jumper, J., et al. (2021). "Highly accurate protein structure prediction with AlphaFold." Nature, 596, 583-589.
- 5.The OpenUniverse Collaboration. (2025). "OpenUniverse2024: A shared, simulated view of the sky for Rubin and Roman."
- 6.Caplar, N., et al. (2025). "Using LSDB to enable large-scale catalog distribution, cross-matching, and analytics."
Official Documents & Institutions
- 7.National Science Foundation. "Theoretical and Computational Astrophysics Networks (TCAN)."
- 8.Space Telescope Science Institute. "GOODS Data Products."