Executive Summary
The team behind SPHEREx, NASA's all-sky spectral survey satellite, gathered the low-level electrical artifacts it found in the first year of flight data into a single catalog and posted it to arXiv on August 10, 2026. The word they put in the title is bestiary, a medieval book of beasts. What the team delivered is not a set of cleaned-up images but eight named artifacts and a written recipe for masking each one.
Three of the eight were artifacts that ground testing had anticipated to some degree. The other four showed themselves only after launch, and one of them turned out to be light traveling sideways through the detector material like an optical fiber, reaching more than 270 pixels away from Saturn. The lab had never set up that condition to test it.
The paper gives equal weight to the places where masking does not work. In front of a planet that no star catalog lists, the source mask can fail to cover the full reach of an artifact, and the persistence question is handed to a follow-up paper with no conclusion attached. Follow those eight entries, their recipes, and the remaining gaps from a data quality seat, and one question stays behind: does your own pipeline have a bestiary like this, or only values you have quietly deleted?
Key Figures
SPHEREx splits its detectors into two groups by wavelength. The three short-wavelength ones are called SWIR, the three long-wavelength ones MWIR. Of the four numbers below, the first two show how differently those two groups respond to the same artifact, and the last two show how much of the data a single artifact takes away. The unit e⁻/s counts the electrons a pixel collects per second.
Source: arXiv:2608.09862
800 vs 4,000
Pixels erased by snowball halos
An average per detector, with MWIR running five times heavier than SWIR
1.8334
Type 3 crosstalk coupling in Band 6
Above 1.0, so the transferred signal grows instead of weakening
270 pixels
Extended blooming spreading around Saturn
A behavior ground testing never once produced
About 2%
Images marked for re-observation
The share after a shower of cosmic rays driven by solar activity
SPHEREx Flags Its Artifacts Instead of Erasing Them
SPHEREx is NASA's satellite for sweeping the entire sky in spectra. It launched on March 11, 2025, and by the time the paper was written it had finished the first year of its observing plan, securing two all-sky spectral maps. The instrument carries six HAWAII-2RG detectors that slice the range from 0.75 to 5.0 micrometers into 102 spectral channels.
Why recording artifacts matters on an instrument like this is written into the paper's own framing. Artifacts left in image space are especially awkward for science that has to measure diffuse light at high precision. The problem is not misreading one bright star. It is measuring a signal spread thinly across the whole sky, which means masking the star is not the end of the job, because the mark that star left inside the detector has to be accounted for too.
Recording is meant literally here. The data SPHEREx sends down carries a flag layer alongside the brightness values. Three basic flags are applied on board, TRANSIENT, OVERFLOW and SUR_ERROR, and ground processing adds observation-dependent marks on top, such as known sources or persistence left by a previous exposure. Combining the three yields six named cases such as late transient or early overflow, and the paper prints the table of those combinations as bit sums in the body. Which pixels to drop from an analysis is settled by reading that table next to the brightness image.
The team documents eight artifacts in all. Counted by group, the ones that first appeared in orbit outnumber the ones the lab called in advance.
The Lab Called Only Three of Them in Advance
The first group starts with charge blooming. A pixel that takes in more light than it can hold spills the excess into its neighbors. Ground testing predicted the effect, but the brightness at which it starts had to be measured again in orbit. Non-linearity begins near 850 electrons per second on the SWIR detectors and near 500 on the MWIR ones, and the OVERFLOW flag is raised there. More than ten times brighter than that, past 9,300 and 5,600 electrons per second, too few samples remain to fit a slope, so the brightness value of that pixel is left empty and recorded as NaN. Around bright stars the BLOOM flag is drawn as a circle.
Second is the snowball. When a cosmic particle strikes the detector or a radioactive decay happens inside it, a round blob is left in the image. Even the name is not something SPHEREx coined; it carried over from work on Hubble WFC3/IR and JWST detectors. The blobs come in many sizes, so the team grouped similar ones by the number of pixels in the transient region, and the group the paper shows as its example has an average radius of 5 to 6 pixels in that region. To pick these blobs out, the team removes blooming first and then looks for places where the OVERFLOW and TRANSIENT flags are both raised. The rate jumps when solar activity picks up, and about 2% of the images taken outside the South Atlantic Anomaly were marked for re-observation.
Third is Type 1 crosstalk. SPHEREx reads its arrays skipping rows, and the parameter that sets the skip is fixed at 32. Where there is a bright source, an inverted negative ghost of it appears exactly 32 pixels above. The coupling is strongest in Band 1 and indistinguishable from noise in Bands 4 through 6.
This artifact came with a warning attached. The CIBER-2 experiment, which uses a similar readout scheme, had already seen ghosts spaced 25 pixels apart, and that number matched the row skip of that instrument. One instrument's record of its own artifacts told the next instrument what to go looking for.
Four of Them Appeared Only in Orbit
The strangest of them is extended blooming. A pixel driven into forward bias emits light at the detector's cutoff wavelength, much like an LED. That light stays trapped inside the mercury cadmium telluride material and spreads sideways as if through an optical fiber. The result sits plainly in the images. Wedge-shaped shadows form near clusters of low-efficiency pixels, and in an image of Saturn the effect reached beyond 270 pixels. Bright streaks parallel to the readout direction can also extend like a diffraction pattern.
Second is the snowball halo. Signal spills past the boundary of the blob itself. Whether the cause is inter-pixel capacitance, blooming, or charge sitting below the transient detection threshold, the paper does not settle. What it does instead is stack the images, measure, and fit an empirical form. The intensity at a distance r from the center is written as 0.2 × Asnow ÷ r², where Asnow is the number of pixels in the transient region of the blob. Cutting the mask at 0.1 electrons per second, the noise level, erases an average of 800 pixels per detector on SWIR and 4,000 on MWIR. There is a reason MWIR runs five times heavier. Those detectors have shallower charge wells, so the same incoming energy spreads a blob wider, and detector 6 of the six is hit most often.
The three crosstalk types share one property. The strength of a crosstalk image depends not on the total charge accumulated in a pixel but on the current flowing at that moment. An already saturated central pixel therefore generates almost no crosstalk, while the ring around it, where blooming has raised the current, does. That is why crosstalk images from genuine electrical coupling, Type 1 among them, appear as donuts with an empty middle.
Third and fourth are the two newly found crosstalk types. Type 2 comes from the large current drawn by a bright star, which disturbs neighboring channels read at the same time and replicates the signal on both sides at every multiple of 64 pixels along the fast readout direction. The coupling coefficient is largest in Band 1 at 0.006194 and sits around 0.0005 in the three MWIR bands, a tenth of the SWIR level. Type 3 belongs to MWIR alone. A bright pixel at a channel edge reaches the end of its row and then leaves a smear at the start of the row 32 rows later. Hysteresis in the readout electronics is the suspected cause, but the explanation is not finished.
The Type 3 smear decays exponentially across the first 24 pixels of a row, so the coefficients were fit band by band: (0.1454, 0.1302) for Band 4, (0.1274, 0.1122) for Band 5, and (0.0497, 0.0144) for Band 6. The value that catches the eye is Band 6's average coupling coefficient of 1.8334. Anything above 1.0 means the transferred signal grows rather than weakens. The paper reads that value as evidence that Type 3 is not true crosstalk. If a phantom pixel read between a bright pixel and a dark background creates a large offset in the readout electronics, and that offset decays slowly, the total signal carried in the smear can exceed the signal of the original pixel.
Every Recipe Has a Place Where It Fails
Putting all eight in one place with their recipes attached makes the character of the paper clear. There is no single cleaning rule running through the artifacts. One is cut at a threshold, one is predicted by an inverse-square form, one is fit with an exponential, and one is written down as still open.
| Artifact | Group | Pipeline recipe |
|---|---|---|
| Charge blooming | Called by the lab | OVERFLOW at 850 e⁻/s on SWIR and 500 on MWIR, BLOOM around bright stars |
| Snowball | Called by the lab | Remove blooming first, then identify where OVERFLOW and TRANSIENT overlap |
| Type 1 crosstalk | Called by the lab | Predict with coupling coefficients, then flag pixels above 0.3 e⁻/s |
| Extended blooming | First in orbit | Enlarge the source mask on the brightest sources |
| Snowball halo | First in orbit | Empirical 0.2 × Asnow ÷ r², cut at 0.1 e⁻/s |
| Type 2 crosstalk | First in orbit | Predict with a per-channel coupling matrix, flag what exceeds 0.3 e⁻/s |
| Type 3 crosstalk | First in orbit | MWIR only, per-band exponential decay model over the first 24 pixels of a row |
| Image persistence | In progress | Flagging method in place, in-flight behavior deferred to a follow-up paper |
Compiled by Pebblous. The descriptions in arXiv:2608.09862 rearranged by group and recipe. The grouping follows the paper's own comparison of ground testing against flight data.
The paper does not hide the places where a recipe is written and still does not work. Extended blooming is handled by enlarging the source mask, and that approach assumes an object listed in a star catalog. A body that moves across the sky and therefore is not in the catalog, such as Saturn, may not have the reach of its artifact fully covered, the paper states. That is the very observation the figure of 270 pixels came from.
Type 3 comes with a condition of its own. The artifact is predicted from the current drawn by the last pixel of the preceding row, and the paper notes that when that pixel saturates so its current cannot be measured, no crosstalk is observed in that row. Image persistence is treated even more frankly. It is a well-known residual effect in the HAWAII-2RG family, and a flagging method along with a current model exists, but the analysis of how it actually behaves in flight is handed to a paper in preparation.
Masking carries a price as well. The recipes work when artifacts sit far apart, but in a crowded field or during a particle shower they overlap. Snowball halos and cosmic ray halos on neighboring pixels together raise the count of pixels marked transient by more than a factor of three, leaving little to work with for science. The re-observation mark mentioned earlier is the last recipe available at that point. If it cannot be masked, taking the image again is the better call.
This is where the value of the bestiary comes from. Had it recorded only the places where masking works, the document would have been an advertisement for the team's own pipeline. Writing down how far the masking got and where it stops working is what lets whoever receives this data know what to watch for in their own analysis.
The Next Mission Doesn't Inherit Clean Images
HAWAII-2RG is not a part that only SPHEREx uses. The paper opens by listing CIBER and CIBER-2, JWST, and Euclid side by side as instruments carrying detectors of the same family. Each mission designs its readout to suit its own science goals, though, so the systematic errors that arise take a different shape from mission to mission. That is why what gets handed down is the method of going out to find artifacts rather than the artifact list itself. The sentence the paper aims at other missions in its conclusion lands on exactly that point. Readout strategies and the shape of the systematics may vary, but the detector physics underneath stays relevant to future infrared observatories.
So what the next mission inherits is not a set of clean SPHEREx images. It is a list: where the number 32 and the number 64 came from, how far around a bright star to stay suspicious, what starts increasing in the week after the Sun gets active. In the same way CIBER-2's 25 pixels foretold Type 1 crosstalk, SPHEREx's 64 pixels tell the next instrument what to look for first. The two recommendations the paper leaves in its conclusion point the same way. Plan observing time and readout strategy so that pixels do not sit saturated for long, and dig further into how crosstalk coupling coefficients vary with detector parameters and bias settings. These are the things a team that lived through the artifacts can hand to the next designer.
For anyone whose work is data quality, the form of the document is the first thing that stands out. Most pipelines quietly drop outliers and keep the results. The count of dropped values may sit in a log somewhere, but why those values were anomalous, whether the same cause will return, and where the dropping rule stops working are usually written nowhere at all. When the person changes, that judgment leaves with them.
Three things are worth checking today.
- Do the outliers your pipeline filters out have names? Without a name, the same cause arriving again cannot be counted as the same artifact.
- Does each filtering rule come with the conditions under which it fails? If the places a mask cannot cover are not written down, everything downstream reads those holes as valid values.
- Have you ever put the artifact list you predicted in pre-launch validation next to the artifact list operations actually produced? SPHEREx compared the two and came away with four more.
Editor's Note: What Pebblous runs into most often in data quality work is not data full of defects but data where nobody can answer what was removed and why. The SPHEREx team named eight artifacts and left behind a recipe for each one along with the conditions under which that recipe fails. That the record is in a form the next instrument can pick up is what makes this paper a data quality case rather than an astronomy instrument note.
References
Academic Papers
- 1.Fazar, C. M., Zemcov, M., Dowell, C. D., Crill, B. P., Nguyen, C., & Hui, H. (2026). "A bestiary of low-level electrical artifacts in SPHEREx flight data." arXiv:2608.09862.
- 2.Akeson, R. et al. (2025). "The SPHEREx Image and Spectrophotometry Processing Pipeline." arXiv:2511.15823.
- 3.Crill, B. P. et al. (2025). "The SPHEREx Sky Simulator: Science Data Modeling for the First All-Sky Near-Infrared Spectral Survey." arXiv:2505.24856.
- 4.Crill, B. P. et al. (2024). "SPHEREx: NASA's Near-Infrared Spectrophotometric All-Sky Survey." arXiv:2404.11017.