Executive Summary

Whether a neural network for material properties can produce a physically impossible value is settled before training begins. A paper posted to arXiv on 19 August demonstrates it experimentally. The fork is a single bit: whether the model's internal features carry parity labels, which record whether a quantity flips sign when you invert the coordinates. The authors built NequIP, Allegro and MACE in two arms each, sharing one architecture, one data pipeline and one set of seeds, differing in that bit alone.

The test property is the piezoelectric tensor of centrosymmetric crystals. This is not a quantity expected to be small; symmetry pins it to exactly zero, so a nonzero prediction is not an inaccuracy but a physical impossibility. The labelled arms stayed at the floating-point floor on all 2,000 crystals. The unlabelled arms crossed the threshold on 90 to 96 percent of them, and the size of what they crossed with matched the size of the genuine piezoelectric tensors sitting in the training data. These read as real responses, not as numerical residue.

The trouble starts with what came next. Training on explicit zeros did not recover exactness, and weighting those rows a hundredfold in the loss reduced the violations without removing them. A readout head placed on the frozen features of a public universal potential simply inherited whatever symmetry group the backbone had. Some errors leave nothing further for the data side to do.

Key figures

The four numbers mark the size of the error, the distance between the two arms, the part that data does not close, and how widely this design bit is spread.

Source: Polat et al., arXiv:2608.18714 (19 Aug 2026)

90-96%

Violation rate without parity labels

The share of 2,000 centrosymmetric crystals given a nonzero piezoelectric coefficient, against 0% for the matched labelled arms

2 million×

Ratio of the two NequIP medians

The distance opened between two models sharing their architecture, data and seeds and differing in one parity bit

0.895 → 0.858

Violation rate after 64× more zero labels

Going from 250 to 16,000 zero-labelled crystals shrank the violations 8.6-fold while the share of crystals violating barely moved

6 of 12

Architectures carrying parity labels

An audit of 18 released architectures found labels in half of the 12 that can emit a rank-3 tensor, and in no model card

1

In a centrosymmetric crystal the coefficient is exactly zero

A crystal's macroscopic properties have to survive every symmetry operation the crystal possesses. That is Neumann's principle, and it dates to the nineteenth century. The piezoelectric tensor is the kind of quantity that flips sign when coordinates are inverted through the origin, and a centrosymmetric crystal carries that inversion as one of its own symmetry operations. So the tensor must be unchanged under inversion and sign-reversed at the same time. Only one value satisfies both. This is an algebraic consequence of symmetry, not an empirical trend.

What makes the property useful for validation is that you can check it without computing the answer. Of the 230 space groups, 92 are centrosymmetric, so gathering those crystals gives you a large population where wrong answers can be identified with no labels and no reference calculation. Any nonzero value there is a confirmed error on its own terms.

The symmetry debate in machine learning has mostly run on rotations. Turn a molecule and the prediction should turn with it, and a model that breaks the requirement is somewhat wrong. That is a question of approximation. Inversion asks for something of a different order. Here exactly one answer is permitted, and a model that cannot produce it is not slightly off; it reports a property that cannot exist.

1.1The parity gap names the exposed cases in advance

The paper's theoretical contribution is a quantity that computes, from group theory alone, which properties and which crystals are exposed. The authors call it the parity gap. It is the number of tensor components allowed when you account for rotations only, minus the number allowed once inversion is counted too. Where the gap is zero, rotation symmetry already forces the zero and parity labels are unnecessary. Where it is positive, parity is the only thing that can hold those components down.

The census is clean. Even-rank properties such as permittivity have a gap of zero everywhere, so they are exempt from the outset, and the elasticity tensor is exempt for the same reason. At rank 3, where the piezoelectric tensor lives, only one of the 11 centrosymmetric crystal families has a gap of zero, the cubic class m3̄m, and the other ten are all positive. In triclinic 1̄, the least symmetric case, the gap runs to 18, which leaves every independent component of the piezoelectric tensor resting on parity alone. Of the 2,000 crystals in the evaluation population, 91.7% sat in the positive-gap region and 8.3% in the safe m3̄m class.

11 centrosymmetric families, split by parity m3̄m (cubic) — 1 family parity gap = 0 rotation symmetry already forces zero 8.3% share of 2,000 evaluated crystals the other 10 families parity gap 1 to 18 parity is the only protection 91.7% share of 2,000 evaluated crystals In triclinic 1̄, the least symmetric case, the gap runs to 18 every independent component of the tensor rests on parity alone
▲ The paper's Table 1 parity-gap census across 11 centrosymmetric families, condensed | Pebblous original diagram
2

One label apart, two million times apart

The design is plain. NequIP, Allegro and MACE were each built twice, sharing architecture and hyperparameters, data and splits, and random seeds, with only the parity labelling of the features changed. One arm carries the natural spherical-harmonic parity that comes with each degree. The other declares every degree even. The second arm is still exactly equivariant to rotations; the only thing it has given up is inversion. EquiformerV2 was added unmodified as the representative of a rotation-only model already in circulation, giving seven arms across four targets and three seeds, or 84 training runs.

The evaluation population is 2,000 centrosymmetric insulators drawn from the Materials Project and re-verified with spglib. The results split at an operating threshold of 0.01 C/m2. Among the labelled arms, not one of the 18,000 structure-runs, three models by 2,000 crystals by three seeds, exceeded it. The largest single prediction fell a factor of 7.6 below the threshold, and the medians sat at the arithmetic floor, the level set by floating-point precision. The unlabelled arms crossed the threshold on 90 to 96 percent of the same crystals. Between the two NequIP medians the ratio was 2×106.

The diagram below places that distance on a log scale. The left band is pressed against the floor and the right band sits above the threshold. Sweeping the threshold across four decades, from 10-4 to 1 C/m2, no arm ever entered the space between the two groups.

Same model, one label bit, six orders of magnitude Predicted magnitude (C/m², log) 10⁰ 10⁻³ 10⁻⁶ 10⁻⁹ threshold 0.01 C/m² No parity labels 90-96% violate Parity labels zero violations all 18,000 runs at the arithmetic floor 2 million× The unlabelled arms carried 0.2-39.8% more parameters Capacity cannot account for the gap
▲ The violation distributions of arXiv:2608.18714 Fig.1-2, redrawn as a schematic | Pebblous original diagram

Insufficient expressiveness does not explain the gap. Declaring every degree even opens additional tensor-product paths, so the unlabelled arm carried between 0.2 and 39.8 percent more parameters than its partner. Capacity ran in favour of that arm, and the result went the other way.

Nor is this a measurement that depended on a well-chosen setup. The evaluation crystals were prepared in two coordinate variants, an idealized one snapped onto the exact space group and a raw one as relaxed by density functional theory, and the violation fractions were unchanged between them. Structure-paired Wilcoxon tests rejected equality at every core and seed, and independent architectures flagged almost the same crystals. There was one visible exception. A labelled arm produced a small response on one structure in the raw variant, and that structure turned out to be genuinely non-centrosymmetric below the selection tolerance. The guarantee had not slipped; it had detected a data artefact.

The safeguard was not bought with accuracy either. On the three targets where the constraint is not active, the QM9 internal energy, the dipole moment and the elasticity tensor, the two arms were indistinguishable within seed noise. On the piezoelectric tensor the labelled arm had the lower error in all three matched cores, which is to say the arm with fewer parameters was the more accurate one. The dipole moment makes a useful control here: it is also parity-odd, but symmetry forces a zero for only 7 of the 10,831 evaluation molecules, so there was almost no room for the arms to separate. What divides them is not whether the target is odd or even but whether the parity gap is open where the model is being evaluated. EquiformerV2, included as released, was ordinarily competent when judged on its error on non-centrosymmetric crystals, and that competence sat alongside violation rates in the nineties.

2.1The spurious floor was there before the symmetry broke

To separate parity as the mechanism from parity as something that merely travels alongside one, the authors followed a symmetry as it actually broke. They displaced rutile TiO2 along a polar mode over 33 amplitudes, from the centrosymmetric parent out past the physically plausible distortion. The labelled arms rose smoothly from the arithmetic floor. The unlabelled arms were already sitting on a spurious floor five to six orders of magnitude higher at zero distortion, and that floor dominated the physical signal throughout the small-distortion regime. Rutile is the decisive case because its parity gap is 1, so parity alone is what forbids a response at the parent. The same separation reproduced on rutile-type SnO2 and anatase TiO2. For a search aimed at candidates with weak piezoelectric response, that spurious floor is the candidate list.

Ball-and-stick model of the rutile TiO2 unit cell. Grey titanium and red oxygen atoms form a centrosymmetric arrangement
▲ The unit cell of rutile TiO2, the crystal whose centrosymmetric parent was displaced along a polar mode in this experiment | Source: Wikimedia Commons (Ben Mills, Public Domain)

The sharper check is to ask which crystals an unlabelled model gets right, because group theory names that set in advance: the m3̄m crystals, where the parity gap vanishes and rotation alone already forces the zero. The three matched cores flagged none of the m3̄m crystals and 97.6 to 99.2 percent of the rest. The headline 90 to 96 percent is the figure across the whole population, safe m3̄m crystals included. The derivative draws the same line. Across 20 crystals spanning 19 space groups, with every Jacobian certified against finite differences, the arms again sat five to six orders of magnitude apart. The labelled architecture does not merely output a zero; it encodes which displacements can activate a response.

Individual materials tell it the same way. On corundum Al2O3, where a gap of 2 leaves parity as the only protection, the unlabelled arms returned a full-magnitude tensor. On the rocksalt and diamond-structure compounds, where rotation alone forces the zero, the same arms stayed between 10-10 and 10-6 C/m2. The only model to cross the threshold on those compounds was EquiformerV2, whose equivariance is approximate, and it did so on two of them, diamond and SrTiO3. An exactly rotation-equivariant model cannot produce that violation at all.

3

Training on zeros did not produce zeros

By this point an objection suggests itself. The training data held no centrosymmetric crystals, so surely this is a data problem. The authors held the same question and worked through four responses in turn. They retrained on 1,000 spglib-verified centrosymmetric crystals labelled with exact zeros, then scaled that augmentation up, then raised the weight of those rows in the loss, then moved the intervention out of training and symmetrized at prediction time instead.

The decisive measurement came from the first attempt. Asked to predict the very crystals used in that training, all four cores still returned a nonzero value about nine times out of ten. This is not underfitting: in training error the zero-labelled rows were fit better than the real-tensor rows. Nor is it a coverage problem, since held-out violation fractions barely moved depending on whether the crystal's space group had been seen. Augmentation even improved the regression error for every core. In the paper's own phrasing, the result is better regressors that still predict impossible values.

The other three attempts point the same way. Multiplying the zero labels 64-fold, weighting those rows a hundredfold in the loss, and symmetrizing the prediction after the fact all reduced the violations. None of them removed them.

Attempt What was done Result
Zero-label augmentation Retrained with 1,000 centrosymmetric crystals labelled exactly zero Violation fraction 0.895 to 0.922 even on the crystals used in training
Scaling up Zero labels raised from 250 to 16,000 crystals Violation median fell from 0.666 to 0.077 C/m2, violation fraction from 0.895 to 0.858
Loss reweighting Zero-label rows weighted a hundredfold Trained-on fraction fell by 0.22, leaving two-thirds still flagged
Test-time symmetrization Subtract the prediction on the inverted structure and halve Exactly zero only for exactly equivariant cores; approximately equivariant EquiformerV2 stalled at 0.82

The scaling numbers are worth a second look. Multiplying the zero labels 64-fold made the violations 8.6 times smaller. The fraction of crystals violating went from 0.895 to 0.858, which is to say it stayed where it was. The predictions shrink, and almost none becomes zero. A value near zero and a value of zero are different values, and physics asks for the second one.

The fourth attempt, symmetrizing at prediction time, looks like a success until you read the conditions on it. The correction can only be applied if you already know the parity of the target property, which means hand-imitating after the fact what the parity label was doing all along. And it failed on EquiformerV2, whose equivariance is approximate. None of the training strategies the authors tested closed the gap.

4

The backbone's symmetry group passes to the head

Few practitioners actually touch this design bit. The common workflow now is to take a released universal interatomic potential, put a property readout head on top of it, and train only the head on your own data. The paper's second proposition addresses that arrangement directly. If the backbone's features carry parity labels, no head placed on them can produce a forbidden value. If the backbone has declared every degree even, no amount of head training will bring the guarantee into being.

The authors checked the proposition on two public models, training an identically constructed piezoelectric head on each one's frozen features. The results split along the backbone.

Same head, different backbone MACE-MP-0 features carry parity labels frozen + piezo head same build, same training violations exactly 0 every crystal, every seed eSEN-30M-OAM declares every degree even frozen + piezo head same build, same training violation rate unchanged regression accuracy comparable The design bit is set once upstream and travels down unseen No amount of head training creates a guarantee the backbone lacks
▲ The same piezoelectric head on two frozen universal potentials | Pebblous original diagram

On the parity-labelled MACE-MP-0, the head's violations were exactly zero for every crystal and every seed. On eSEN-30M-OAM, which declares every degree even, the same head produced impossible values at the rate seen earlier, and the regression accuracy of the two cases was comparable. The head is small, linear and trained, but the property under test lives above it.

What breaks here is the assumption that a team needs only to validate the part it built. Now that fine-tuning and transfer learning are the standard route, the decision that separates possible from impossible outputs was taken in a layer nobody on the team touched, before anyone on the team arrived, and it was never written down.

5

The one bit no model card reports

So how many released models lack the label? The authors audited 18 architectures at pinned versions, reading the source and, where reading did not settle it, measuring directly. Six carry parity as a type: NequIP, Allegro, MACE, MatTen, ICTP and GotenNet. Six are rotation-only, indexing degree and order with no parity flag: EquiformerV1 and V2, eSCN, eSEN, UMA and EquiformerV3. PaiNN, TorchMD-Net and ViSNet carry a vector channel but leave it untyped, and SchNet, DimeNet++ and FAENet have no path to an equivariant tensor output at all. Of the 12 architectures able to emit a rank-3 tensor, exactly half carry parity as a type.

Category Architectures Rank-3 tensor
Parity typed (6) NequIP, Allegro, MACE, MatTen, ICTP, GotenNet Yes — parity present
Rotation-only (6) EquiformerV1, EquiformerV2, eSCN, eSEN, UMA, EquiformerV3 Yes — no parity
Vector channel, untyped (3) PaiNN, TorchMD-Net, ViSNet Not applicable
No equivariant path (3) SchNet, DimeNet++, FAENet Not applicable

▲ The 18-architecture audit (the paper's Table 2) — of the 12 able to emit a rank-3 tensor, exactly half carried parity as a type

What catches the eye in the rotation-only column is how many widely used universal potentials it contains. And this information is written down nowhere. The paper reports finding no model card and no validation suite that states the feature symmetry group. Maximum angular degree is usually published; the parity flag is not.

The omission is not an oversight, either. Allegro ships parity-typed, yet its own paper records that the index may simply be dropped, because doing so simplifies the network and reduces its memory footprint. Parity is a documented efficiency choice, and until now its cost had not been measured.

Checking is not hard. No training, no data, no reference calculation. Apply a single reflection to a randomly initialized model and watch how the output transforms, and the question settles in seconds. The paper's ask is correspondingly modest: report the feature symmetry group, not only the maximum angular degree, and report it for the backbone as well as the head.

The scope reaches past piezoelectricity. In the electric-dipole approximation, the second-harmonic-generation tensor used in nonlinear-optical materials discovery shares the rank, the parity and therefore the gap. Pre-filtering the centrosymmetric candidates out and computing only on the rest is not a safe route either, since it inherits every weakness of the symmetry oracle. All 2,000 crystals in the paper's data are centrosymmetric at the selection tolerance, but 2.2% are not at 10-4 and 6.5% are not at 10-5. The more candidates come from generative structure databases, the more of these boundary cases there will be.

Editor's Note: This overlaps with a scene we meet often in data quality work. When results come out wrong, the first place a team reaches for is usually the data. Collect more, clean it further, relabel it. The error in this paper did not close along that route, because the values shrank when what was needed was an exact zero. Before deciding how far to polish the data, it is worth confirming what the representation that data passes through was built to be able to represent. Sometimes that confirmation takes a single reflection.

The paper states its own limits. EquiformerV2 was measured strictly as released, so no conclusion depends on it. All three matched labelled arms use the same implementation library, so an implementation-specific contribution cannot be fully excluded. The loss-reweighting and zero-injection sweeps used a single architecture. And the architecture audit is pinned to particular versions. The original is at arXiv:2608.18714.