Executive Summary

Astronomers have sorted the icy bodies orbiting beyond Neptune by surface composition for decades, and the tool they reached for was broadband optical color. A brightness difference between two filters stood in for what the surface is made of. What nobody had pinned down is how that color relates, in quantitative terms, to the near-infrared spectrum it substitutes for. A paper posted to arXiv on September 1 defines the relation as compression and measures what the compression costs.

Researchers at the University of Michigan, the University of Arizona and elsewhere built a low-dimensional spectral manifold from the 48 objects that have JWST spectroscopy, then used optical color as the condition for a posterior on that surface. Once the posterior exists, every filter set has a measurable value. The g and r bands together carry about 1.8 nats, and adding g to the ri pair moves the information gain from about 1.5 to about 2.3. What sets the worth of a filter is not how many bands you own but which wavelength the new one opens.

This article follows what the paper measured and what it declined to claim, and then adds a data-collection reading in Section 5. The correspondence drawn there is this article's interpretation, not the paper's.

Key Figures

Source: Lin, Markwardt, Napier, Adams, Malhotra, Gerdes, Optical Colors as Compressed Spectral Information for Trans-Neptunian Objects, arXiv:2609.01169 (2026-09-01)

1.8 nats

Information gain from the gr pair

One overall optical slope narrows an object's position on the manifold by this much

1.5 → 2.3

What adding g to ri is worth

The same single extra band pays differently depending on its wavelength

38.4 → 62.0%

Precision on methanol-rich predictions

What the y band bought was substructure, not overall accuracy

48

Spectra behind the manifold

40 TNOs and 8 Neptune Trojans. This is also the ceiling of the framework

1

A Proxy Used for Decades, Never Priced

Trans-Neptunian objects, or TNOs, are everything that orbits beyond Neptune. Most of them are the icy bodies of the Kuiper Belt, and the Trojans that share Neptune's orbit belong to the same discussion. The direct way to learn what covers their surfaces is near-infrared spectroscopy, which spreads reflected light by wavelength and reads the absorption features of ices and organics off the curve. The obstacle is cost. Spectroscopy reached only the brightest objects, and the sample stayed thin for a long time.

So the working instrument became broadband optical color. Measure brightness through a few filters, take the differences, and you learn whether a surface is red or blue. Cheap and plentiful was the whole argument for it. But TNO spectra in the optical are mostly smooth slopes without sharp absorption features, which left the obvious question hard to answer: how much compositional information does one color actually hold?

JWST changed the circumstances. As homogeneous near-infrared spectroscopy accumulated, a classification into three surface compositional classes settled into place, namely water-rich, CO2-rich and organic-rich. The organic-rich class divides further into methanol-rich and methanol-poor spectral subtypes. The original and the compressed copy now exist side by side for the same objects, which is what makes it possible to ask what the compression keeps and what it throws away.

The machinery for asking was also in place. Ferreras and colleagues had treated astronomical spectra as probability distributions and measured their Shannon entropy to quantify information content, and Doorenbos and colleagues had shown that spectroscopic information can be inferred from lower-dimensional observations such as broadband imaging or photometry. This paper brings both lines to the outer Solar System. Instead of pinning a fixed label on an object from its color, it treats the color as an incomplete measurement of an underlying spectrum and solves for the distribution of answers.

2

Forty-Eight Spectra on a Single Surface

The material comes in two bundles: 40 TNOs observed by the DiSCo-TNO survey and eight Neptune Trojans, 48 objects with JWST spectroscopy in total. One condition governed the selection. An object had to have both spectroscopy and homogeneous optical photometry. That means the sample is not every outer Solar System spectrum published to date. Centaurs, with substantially different evolutionary histories, and dwarf planets, with distinct surface processes, were left out from the start.

The published JWST spectra cover roughly 0.75 to 5.1 micrometers. The authors converted optical colors into relative reflectance constraints to build a visible continuum down to 0.4 micrometers, then joined the two across the wavelength region where they overlap using a cubic spline. The result is a single reflectance spectrum running unbroken from 0.4 to 5.1 micrometers. Because the optical stretch occupies such a small share of that range, the joined spectra were resampled onto a logarithmic wavelength grid, which raises the relative weight of the short-wavelength end.

Compression starts here. Principal component analysis places those spectra in a low-dimensional latent space, and kernel density estimation lays an empirical prior over it. The PCA basis and the KDE prior together define the manifold, which is best pictured as one curved surface where observable spectra live. Ten principal components are used to decode a spectrum, and only the first five when entropy and divergence are computed, a choice made to reduce sensitivity to interpolation noise carried in the higher-order components.

How the manifold was built From 48 spectra to a low-dimensional prior JWST spectra 48 objects 0.75–5.1μm Extended to 0.4μm with color cubic spline join Log grid + PCA 5–10 components Manifold (KDE prior) One color observation narrows the wide prior into a tight posterior position on the manifold prior (wide) posterior (narrow, after color)
▲ Original Pebblous diagram (manifold construction, reinterpreted) | Source: arXiv:2609.01169, Section II methodology

The construction itself is carried over unchanged from the framework the same group set up in an earlier paper. What changed is the conditioning data. It was near-infrared photometry then and broadband optical color now, which means a far cheaper and far more abundant observation is being aimed at the same manifold.

With the manifold as prior, optical color becomes the observation. Given one color vector, Bayesian inference returns the posterior distribution over where the object sits on the manifold. Answering with a distribution rather than a point is the part that matters, because photometric uncertainty propagates straight into the width of that distribution. Draw samples from the posterior, push them back through the principal component basis, and the corresponding spectra come out. The authors say up front that these reconstructions are not the primary product of the work. They serve as a visual check and as the machinery behind the generated colors that come later.

"Although high-frequency spectral details are inevitably discarded, the dominant compositional geometry is largely preserved, allowing major spectral classes to remain distinguishable from optical colors alone. In this framework, spectroscopy and broadband photometry are not competing approaches but complementary representations of the same compositional manifold at different information resolutions: JWST spectroscopy defines the manifold, while wide-field optical surveys provide compressed observations that localize objects within it."

Section IV.1 of the paper · arXiv:2609.01169

3

Pricing One Filter in Nats

Once the posterior is in hand, two quantities can be measured. One is its Shannon entropy, which gives the effective volume of latent spectral space that the observed colors still allow, or in plainer terms how far the range of answers has narrowed. The other is information gain, defined as the Kullback-Leibler divergence of the posterior from the prior. It says how strongly that one color pulled the distribution away from what was expected with no observation at all. Narrowing and moving are different questions, and the two metrics answer them separately.

The comparison runs over seven Rubin/LSST filter configurations the authors picked: ri, rz, gr, gri, grz, griz and grizy, in order of increasing band count. Validation works by leaving one object out, training on the rest and inferring the excluded object, with 100 Monte Carlo samples per realization and the reported values averaged over all leave-one-out trials. It is the most honest procedure a 48-object sample supports.

Information gain by filter set, as stated in the paper's text Averaged over leave-one-out trials on 48 objects, in nats ri ~1.5 2 bands, no g gr ~1.8 2 bands, with g gri ~2.3 Both gr and ri use two bands, yet gr scores higher, and one more band takes ri to 2.3. The other four sets (rz, grz, griz, grizy) appear only as curves in the paper, so no values are printed here. The text describes grz as slightly above gri, griz as close to grz, and grizy as lowest in posterior entropy while its information gain stays similar to griz.
▲ Original Pebblous chart | Source: values reported in the text of Section III.1, arXiv:2609.01169

Among the two-band configurations, gr holds about 1.8 nats, primarily through the overall optical slope. The weight of g shows up clearly in the comparison between ri and gri. Adding g to the ri pair produces a substantial reduction in posterior entropy and raises the information gain from about 1.5 to about 2.3 nats. Band count is secondary here. The worth of the added band comes from which direction its wavelength widens the constraints on the manifold.

Among the three-band sets, grz gives slightly lower entropy and slightly higher information gain than gri. The longer optical baseline that z provides acted as the stronger lever in this experiment. The authors head off a misreading at this point: it does not mean the i band is unimportant. The gri set still improves substantially over gr, and which of i and z serves better can depend on the signal-to-noise and observing efficiency of a given survey.

After that the returns shrink. The griz configuration performs similarly to grz, meaning the incremental information from i becomes modest once the g-r-z baseline is established. The five-band grizy set yields the lowest posterior entropy of all, but its information gain stays close to griz. The paper reads y as a band that tightens the distribution where it already sits rather than relocating it.

3.1What the y Band Actually Bought

A classification experiment checks the information-metric result a second way. Four classes are set up, water-rich, CO2-rich, organic-rich and methanol-rich, and each object's posterior is compared with the empirical latent distribution of each class by KL divergence, with the object assigned to the smallest one. That is a comparison between distributions, not a partition of latent space into fixed regions and not a label read off the posterior mean.

Overall accuracy came to 66.9% for grz and 71.7% for the full five-band grizy. Judged on that gap of under five points alone, the extra bands look like a poor bargain. The improvement is concentrated somewhere else, in the substructure inside the organic-rich class. Adding the y band raised the precision of methanol-rich predictions, the probability that something called methanol-rich really is methanol-rich, from 38.4% to 62.0%. Quote the overall figure only and you miss this.

The same table shows where the method struggles. If the organic-rich and methanol-rich subgroups are combined into one broader organic-rich family, both the water-rich and organic-rich classes are recovered with high fidelity, while the CO2-rich class separates less cleanly. That holds in both of the configurations for which confusion matrices were computed, grz and grizy. The conclusion that adding filters fixes everything does not come out of this paper.

4

Where the Four Color Groups Landed

With the retained information measured, the next question concerns the identity of the old taxonomy. Where do the TNO populations that were carved out of optical color alone sit on this manifold? The test case is the set of four groups Bernardinelli and colleagues obtained by fitting a Gaussian mixture model to DES color data: NIRB+, NIRB−, NIRF+ and NIRF−. Those groups were built entirely in optical color space and are reproducible from the published model parameters without the original photometry, which makes them an independent benchmark for this framework.

Projected into latent space, the four groups fall on different branches and transition regions. The sharpest separation is between NIRB− and NIRF+, which barely overlap at all. NIRB− tracks the water-side branch, and NIRF+ sits on the red organic side that includes the methanol-rich objects. A single optical axis running from blue to red picks out one of the dominant compositional axes of the manifold. Sorting each group by which compositional family it is closest to in KL divergence gives the table below.

Optical color group Closest compositional family Position on the manifold
NIRB− Water-rich Follows the water-side branch, with a strong association
NIRB+ Water-rich Closest to the same family, but with a substantially larger KL divergence than NIRB−. Spreads broadly between the center and the water-side branch
NIRF− CO2-rich Red in the optical, but less extreme in near-infrared morphology, sitting near the central red region
NIRF+ Organic-rich and methanol-rich The reddest and most strongly absorbed end of the manifold

Latent-space KL divergence between the optical GMM groups of Bernardinelli et al. and the near-infrared compositional families (Section III.2, Figures 5 to 7). Lower values mean the two distributions agree better. The authors state in Section II.3 that the KL divergence used here serves as a distribution-to-distribution similarity metric rather than a point-to-point distance or a calibrated classification probability. Source: arXiv:2609.01169.

Where the four color groups landed on the manifold Bernardinelli et al.'s optical GMM groups, projected into latent space (Section III.2) Water-rich CO2-rich Organic- and methanol-rich NIRB− strong water association NIRB+ closest to water, spreads to center NIRF− closest to CO2, less clean split NIRF+ closest to organics, reddest end strong association (short KL distance) broad spread (long KL distance)
▲ Original Pebblous diagram (Figs. 5–7, reinterpreted) | Source: arXiv:2609.01169, Section III.2

Everything so far runs from color to spectrum. The paper's next move runs the other way. Samples are drawn from the empirical latent density of each of the four compositional groups, decoded into spectra, and turned back into g−r and r−z colors, generating a whole color space from the manifold. In that generated space the water-rich class populates the bluest region, the CO2-rich class occupies an intermediate and relatively compact locus, and the organic-rich and methanol-rich classes extend toward progressively redder colors. The terrain agrees closely with the measured NIRB and NIRF contours from Bernardinelli.

That agreement is what carries the central claim. The color groups are not independent empirical clusters but low-dimensional manifestations of a continuous compositional manifold. Color-color structures reported across different surveys, filter systems and statistical methods become, under this reading, different observational expressions of one surface. And the fact that the generated distributions partially overlap is itself the explanation for why broadband color alone cannot uniquely recover a detailed near-infrared type. Resolving objects that sit in transition regions takes additional near-infrared observation.

An appendix relates the framework to the linear continuum model proposed recently by Fraser and colleagues. Reduce this manifold to an approximately one-dimensional representation of two major groups, water-rich and CO2–organic, and it corresponds to that linear model. The authors add that this is a conceptual interpretation rather than a formal mathematical derivation joining the two.

4.1Trying It on Follow-Up Targets

Reconstructing the spectrum of a single object is a way of validating that the framework works. The value the authors point to lies at population scale. Once Rubin Observatory's LSST accumulates broadband photometry for millions of objects, the distribution of manifold posteriors can be studied statistically as a function of dynamical class, heliocentric distance, inclination, resonance and other orbital parameters. That is a compositional map of the outer Solar System.

The paper also writes down a few expectations about where spectroscopy should look next. Candidates with weak-water spectra resembling 2006 RJ103 are expected to sit preferentially in the bluest NIRB− population. Spectra resembling the objects known as blue binaries may preferentially appear within the NIRB+ region. Analogs of the unusual spectrum of 2011 SO277 should lie near the overlap between multiple optical populations, which makes those transition regions particularly attractive targets for follow-up spectroscopy.

The framework has already had one run on real observations. Rubin Observatory's Science Validation survey turned up three L5 Neptune Trojans. Feeding in nothing but the broadband colors reported by Schwamb and colleagues put the three in different regions of the manifold. 2025 NN80 and 2025 MD138 are tightly localized within the canonical water-type class, while 2025 MH348 shows a broader posterior extending toward the CO2-rich region.

The third object is interesting because no CO2-rich Neptune Trojan has been confirmed to date. The closest known analog to such a transition spectrum is 2011 SO277, and the paper suggests 2025 MH348 may occupy a similar region, making it an attractive target for near-infrared photometry or spectroscopic follow-up. Using cheap color observations to choose the targets for expensive spectroscopy became, in this case, a matter of drawing a posterior.

5

Letting the Arithmetic Choose the Next Measurement

That is where the paper ends. Read from the data side, the structure looks familiar. Pairs of an expensive original and a cheap proxy are everywhere in our own work: precise physical measurement and sensor estimate, expert labels and heuristic labels, a full census and a log-derived proxy. The relationship usually gets summarized by a correlation coefficient, which asks only how well the proxy tracks the original.

What this paper changes is the definition of that relationship. Treat the proxy as a lossy compression of the original and the answerable questions change. How much information survived the compression, how much the number rises when one more channel is opened, and which channel duplicates what an existing one already carries all become measurable in the same unit. That is how this paper settled whether to add z or i to gr: the answer came out of a calculation instead of a convention.

The three questions below are this article's reading rather than claims in the paper. They can be asked in the same order when a collection plan is on the table.

  • Does the dataset distinguish which of its fields is the original and which is the compressed copy, and how many records hold both for the same entity?
  • On what grounds was the decision made to populate one more column, and has anyone computed how much variation that column explains which the existing columns miss?
  • Is the benefit of the additional collection being judged by a single overall accuracy number? In this paper the y band was worth 23.6 points of precision on one subtype, not 4.8 points of overall accuracy.

That last item is also the most practical thing in the paper. Averages conceal where an improvement happened. A change that lifts one fine-grained tier sharply and leaves everything else where it was shows up in the average as a slight uptick. Choose the next channel by the average alone and you lose sight of the very problem that channel was meant to solve.

The line the authors draw around their own work is worth carrying over too. This manifold is not a fixed taxonomy but a representation that keeps being updated. Rare or unusual surface types are underrepresented in it today, and the surface itself becomes more complete as JWST spectra accumulate. That it was built from 48 objects is both the strength of this framework and its ceiling.

Editor's Note

One question Pebblous often gets during data quality diagnostics is which fields to fill in next. This paper demonstrates one way of answering it, in a different domain. The inference code and the learned manifold are open source at github.com/sevenlin123/color_to_spec, and the repository ships a machine-readable agent skill so that an AI agent can execute the same workflow. The authors also disclose at the end of the paper that they used a code agent for the figure-generating scripts and a language model for English editing.

R

References

Primary Paper

Related Work

  • 2.Bernardinelli, P. H., Bernstein, G. M., et al. (DES Collaboration) (2025). "Photometry of Outer Solar System Objects from the Dark Energy Survey. II." The Astronomical Journal 169(6), 305. arXiv:2501.01551 — source of the optical NIRB/NIRF color groups discussed in Section 4.
  • 3.Lin, H. W., Markwardt, L., Napier, K. J., Adams, F. C., Malhotra, R., Gerdes, D. W. (2026). "Probabilistic Spectral Reconstruction of Trans-Neptunian Objects from Sparse Photometry." The Astronomical Journal 171(6), 356. arXiv:2604.23840 — Section 2, the same authors' earlier manifold framework.
  • 4.Fraser, W. C., Buchanan, L. E., Wong, I., Holler, B., Brown, M. E. (2026). "Linear Continuum Modelling to Explain the Majority of Bulk Features of Kuiper Belt Object Spectra." arXiv:2608.23926 — the linear-continuum model discussed in the Appendix.
  • 5.Schwamb, M. E., Bernardinelli, P. H., et al. (2026). "NSF-DOE Vera C. Rubin Observatory's First View of the Neptune Trojan L5 (Trailing) Cloud." Submitted — source of the three Neptune Trojan colors in Section 4.1.
  • 6.Ferreras, I., Lahav, O., Somerville, R. S., Silk, J. (2023). "The Entropy of Galaxy Spectra: How Much Information is Encoded?" RAS Techniques and Instruments 2(1), 78–90. doi.org/10.1093/rasti/rzad004 — Section 1, precedent for the entropy-based approach.
  • 7.Doorenbos, L., Sextl, E., Heng, K., Cavuoti, S., Brescia, M., Torbaniuk, O., Longo, G., Sznitman, R., Márquez-Neila, P. (2024). "Galaxy Spectroscopy Without Spectra: Galaxy Properties from Photometric Images with Conditional Diffusion Models." The Astrophysical Journal 977(1), 131. arXiv:2406.18175 — Section 1, precedent for inferring spectral information from low-dimensional observations.