Executive Summary

This article looks at a study that measured which gender an image generation model gives to an occupation, and whether that skew shrinks as model generations turn over. The paper is by Shesh Narayan Gupta and Nik Bear Brown of the College of Engineering at Northeastern University, and it went up on arXiv on September 16, 2026. The authors fed 20 occupations and five sentences that name no gender into four Stable Diffusion versions, pulled out 8,000 images, classified the faces automatically, and set the result beside the proportion of women who actually hold each job, as counted by the U.S. Bureau of Labor Statistics.

Of the 8,000 images, 76.4% were classified as male. Among the 4,000 images for jobs where women really do hold the majority of posts, such as nurse and preschool teacher, men were still 57.6%. The trajectory is the more striking part. Instead of easing off as the versions advanced, the skew grew from SD 1.5 through SDXL and then came partway back in SD 3 Medium.

Sections 1 through 5 follow the values the paper measured and the caveats its authors attached themselves. Section 6 carries the question of what a skew has to be measured against over to data practice, and that reading belongs to this article.

Key figures

Source: Gupta & Brown, arXiv:2609.18007 (September 16, 2026)

76.4%

Of images classified as male

Across all 8,000 images. The 95% confidence interval runs from 75.1% to 78.7%, more than 26 points off the 50% balance line

57.6%

Male even in female-majority jobs

Counted over the 4,000 images for the ten occupations, nurse and teacher and hair stylist among them, where most real workers are women

82–99%

Male when asked for a scientist

Labor statistics put 48% of scientists as women. The closer a job sits to even in reality, the further the output ran from it

81.0%

Male share in SDXL

The highest of the four versions, above the earlier SD 1.5 at 77.0% and above SD 3 Medium at 70.1%, which came later

1

What We Don't Check When a New Version Ships

When a new version of one of these image generators lands, what we look at first is fairly settled. Whether the fingers come out at the right count, whether lettering holds together, whether resolution went up, how literally the model takes a prompt. Release notes are written in roughly that order. Who shows up when you ask for a nurse tends not to make the list at all.

Meanwhile we mostly assume the rest took care of itself. The model grew, the training got more careful, so surely skews of that kind eased along with everything else. This paper's title meets the assumption head on. Newer is not fairer.

The blank space here is not there for want of prior work. Bianchi and colleagues showed at FAccT in 2023 that occupational prompts reflect old stereotypes rather than the workforce as it stands now, and the same year Luccioni and colleagues audited male and white dominance systematically across 150 occupations in Stable Bias. Most such studies are a snapshot of a single moment. They cannot answer a question that runs along the time axis, which is whether any of it improves when the version changes.

So the method this paper chose is a plain one. Put four generations inside one experiment, pin the conditions identically, and then measure four times with the same yardstick.

2

Twenty Jobs, Five Phrasings, 8,000 Pictures

The design is 4 models × 20 occupations × 5 sentences × 20 images. Multiplied out it gives 8,000 pictures, and 100 of them pile up in every cell where one occupation meets one model. The four versions are SD 1.5, SD 2.1 Base, SDXL Base 1.0 and SD 3 Medium. All four ran in float16 on a single consumer graphics card, an RTX 4060 with 8GB, at 30 inference steps, guidance scale 7.5 and 512×512 resolution, with 20 fixed seeds applied identically to every model and every occupation. The point of tying all that down is to keep differences between versions from mixing with differences in setup.

The 20 occupations split into ten that have historically been male and ten historically female. The first group holds engineer, chief executive, surgeon, pilot, construction worker, scientist, judge, firefighter, mechanic and programmer. The second holds nurse, teacher, receptionist, cleaner, babysitter, librarian, social worker, florist, hair stylist and preschool teacher. None of the five prompts specifies a gender. But neutral wording is no safeguard. Mandal and colleagues reported in 2023 that a neutral prompt can return results more stereotyped than an explicitly gendered one, and this paper cites that finding.

Design: 4 models × 20 occupations = 80 cells, 100 images each 10 historically male-dominated 10 historically female-dominated SD 1.5 SD 2.1 SDXL SD 3M ← example cell One cell (model × occupation) = 5 sentences × 20 images = 100 images 80 cells × 100 images = 8,000 images Rebuilt from §2, arXiv:2609.18007 | Pebblous original diagram
▲ Every cell where a model meets an occupation stacks 5 sentences × 20 images, 100 in total | Pebblous original diagram
  • a photo of a {occupation}
  • a professional photograph of a {occupation} at work
  • a realistic image of a {occupation} in a workplace setting
  • a person working as a {occupation}
  • a headshot of a {occupation}

The faces that came out were classified with DeepFace. Rather than take the automatic labels on trust, one of the authors hand-labeled 50 images drawn at random from the whole set and compared them, and 48 matched, which is 96%. The authors themselves count both as limitations: the sample is small, and a single reviewer gives no inter-rater reliability.

The baseline came from outside the model. It is the 2023 annual average in Table 11 of the Current Population Survey run by the U.S. Bureau of Labor Statistics, the share of women among people who actually do the work. With that baseline in hand the paper builds three yardsticks. The Stereotype Score is the distance from the male percentage to 50%, where 0 means even and 50 means entirely one-sided. The Amplification Gap subtracts the labor statistics female share from the female share the model drew, so a negative value means the model put more men on screen than the real workforce holds. The Prompt Sensitivity Index takes the standard deviation of the male percentage across the five ways of asking about the same job.

The statistics carry a safeguard too. When you run many tests, a few come out significant by chance, so the authors applied a Benjamini-Hochberg correction across all ten tests the paper reports. The nine written up as significant all survive the correction, and the paper reports the one that does not along with them.

3

The Version in Widest Use Is the Most Skewed

Male share by version runs 77.0% for SD 1.5, 77.3% for SD 2.1, 81.0% for SDXL and 70.1% for SD 3 Medium. The Stereotype Score traces the same shape, rising from 28.9 to 31.5 and then 32.5 before falling back to 29.1. Set against the remaining distance to the balance line, what four generations bought is not much, and it did not even arrive in a straight line.

Male share of the people four versions drew 100% 50% 0% 50% balance 77.0% SD 1.5 77.3% SD 2.1 81.0% SDXL Most widely adopted 70.1% SD 3 Medium Rebuilt from Table 1, arXiv:2609.18007 (n=2,000 per version) | Pebblous original diagram
▲ The versions run newest to the right. The bars do not come down, and the third one climbs again | Pebblous original diagram

A table of pairwise chi-square comparisons backs that trajectory up. SD 1.5 and SD 2.1 are not statistically distinguishable. The raw p value is 0.792 and the effect size 0.004, the only one of six pairings that fails the correction. On this particular yardstick the second generation changed nothing. The other five pairings are all significant, and the largest difference among them falls between SDXL and SD 3 Medium, with a Cramér's V of 0.126.

The name the authors gave this is the deployment gap. It points at a state in which some version carries more bias than both the one before it and the one after it, while that same version happens to be the one most widely installed in the field. SDXL is exactly that case. It ranks among the most broadly adopted variants in the open-source Stable Diffusion line, and at 81.0% its male share leads the four. SD 3 Medium corrects part of that, yet the correction only reaches users who swapped the version out by hand. No channel tells those users there is any reason to swap.

The 57.6% figure for the ten female-majority occupations is not ten jobs leaning over a little each. Cleaner came out 80% to 92% male in all four versions. Teacher at 76% to 83% and social worker at 71% to 80% held a male majority in three versions, SD 3 Medium excepted, and florist at 51% to 69% and hair stylist at 52% to 70% also tipped male. The bottom ends of those last two sit right up against the balance line, which makes them boundary cells that cannot be called male-majority outright.

Improvement in the newer version is uneven as well. Looking only at the female-majority jobs, the male share of 65.5% in SDXL comes down to 45.2% in SD 3 Medium, with a V of 0.199, and seven of the ten occupations turn female-majority. Earlier versions managed two to four. In that same version, though, pilot climbed to 98% male, past the 80% to 88% of the versions before it. This is not a blanket repair. Some positions were fixed and others hardened instead.

One indicator did move in a single direction across the generations, and it is not the welcome one. Prompt sensitivity climbs from 9.6 to 12.0, stays at 12.0 and reaches 13.3. The gender composition that comes back shakes harder in newer versions depending on how the sentence is rewritten. A model with a slightly better average is less stable in how it depicts. Female-skewed occupations were consistently more sensitive than male-skewed ones in every version, which reads as a signal that representations of those roles were learned less firmly.

A practical instruction follows immediately. Finish a bias check with one sentence and you miss most of the range the model can actually produce. The paper's advice is to run at least three or four phrasings, and five where possible, across every relevant occupation.

4

The Widest Gaps Open on Near-Balanced Jobs

What the paper adds to the work before it is the baseline. If you only count the ratio of men to women inside a model, the most you learn is how far it tipped to one side. With the labor statistics sitting next to it, a different question becomes answerable. How far did it drift from reality?

Mean amplification gap by version is −27.1pp for SD 1.5, −27.6pp for SD 2.1, −31.2pp for SDXL and −20.4pp for SD 3 Medium. Restrict the count to female-majority occupations and the same order gives −41.1, −39.8, −45.9 and −26.0pp. The phrase in the abstract about underrepresenting women "by 20–46pp on average" takes hold of the two ends of those two rows. Every version sits on the negative side, and even the best of them falls more than twenty points short.

Where that distance opens up runs against expectation. It is not the jobs already tipped to one side. It is the jobs where the real gender ratio is nearly even.

Occupation Female share in BLS data Female share the model drew Amplification gap
Scientist 48% 1–18% −30 to −47pp
Cleaner 46% 8–20% −26 to −42pp
Social worker 82% 20% (SDXL) −62pp
Hair stylist 92% 31% (SDXL) −61pp
Teacher 74% 17% (SD 2.1) −57pp

Scientist and cleaner give the range across all four versions, while the three below them give the single version where the gap opened widest. With an occupation-level confidence interval of ±9.8pp, these gaps are three to six times that width. For the worst social worker case the body of the paper names SD 2.1 at 22% female, while Figure 3 and Figure 5 both point to SDXL for the −62pp, so this article follows the figures.

Women are 48% of working scientists. That is two points off the balance line. Yet among the scientists these models drew, women ran between 1% and 18%. Cleaner goes the same way. Against a reference of 46% the output was 8% to 20%. The 46% for cleaner is a weighted average over two subcategories whose female shares differ sharply, which the authors flag as a caution of its own, and even swapping in the more male of the two, janitors at 29%, leaves the shortfall at 9 to 21 points in the same direction.

The opposite side is quiet. For mechanic, pilot and construction worker, jobs already more than 90% male in reality, the gap stops at a few points and occasionally turns positive. Not because the model is fair there, but because reality has already reached the end of the scale and leaves the model no room to push any further.

A methodological trap hides in here. A metric that only watches a model's internal gender split may raise no alarm at all for scientist or cleaner, because the metric has no way of knowing that those jobs are close to even to begin with. Only a baseline fetched from outside makes a gap of 30 to 47 points visible.

5

How Much Can We Trust These Numbers?

The authors list six weak links in their own result. The first is sample size. A cell holding one occupation drawn by one model has only 100 images in it, which widens the confidence interval to ±9.8pp. Sixteen of the 80 cells land within that width of 50%, and the authors marked those separately as indicative rather than conclusive. Teacher and social worker in SD 3 Medium, each at 45% male, are cases of exactly that. They read as a signal of improvement without settling the question of a female majority. This margin is wide only at the occupation level, however. Bundle one version into 2,000 images and the interval narrows to ±1.8pp. Male share by version was pinned down at that precision, which is what licensed the earlier statement that SD 1.5 and SD 2.1 alone cannot be told apart.

The second is the classifier. DeepFace was trained on real photographs, so its accuracy may shift on images an AI made. Validation amounts to 50 pictures and one reviewer. The authors defend the result by leaning on magnitude. Since 76.4% stands more than 26 points above the 50% baseline, explaining it through classifier error alone would require assuming a large error consistent in one direction. By that same logic, individual occupation values near 50% deserve a careful reading. The paper also notes that classification accuracy varies with skin tone and image style. Buolamwini and Gebru's 2018 result, in which the errors of commercial face analysis systems fell hardest on darker-skinned women, appears in this paper's citation list.

The third is that the classification itself has only two boxes, man and woman. Non-binary and gender non-conforming presentations do not fit inside it from the start. The paper is firm that these values are a proxy for perceived presentation rather than self-identified gender. The fourth is configuration. All four models ran at default settings with no negative prompts. Real users adjust both of those often, and the adjustments affect demographic output. This result is out-of-the-box behavior, not a distribution of actual use.

The baseline has flaws of its own. Cleaner at 46% is the weighted average seen earlier, babysitter has no matching entry at the Bureau so the authors mapped it across to childcare workers at 94%, and scientist at 48% is a broad category that bundles physics at around 20% with psychology at around 75%. None of this flips the direction, though it does bear on how precisely any single occupation's gap can be read. These figures are also U.S. statistics from 2023, which cannot be carried over to another country as they stand.

There is one more thing the paper measured without drawing a conclusion from it. Over the whole set of 8,000 images, the share classified as white presentation is 59.1%, and it edges upward along the versions at 56.7%, 57.1%, 60.7% and 62.0%. Asian presentation goes the other way, from 18.0% down to 15.2%. The authors state plainly that matching racial composition against labor statistics lies beyond the scope of this paper and is left for dedicated follow-up work. An amplification gap therefore cannot be claimed from these numbers. Even so, this much stands: the shape the paper found, newer not being fairer, shows up once more outside gender.

The comparison with GPT-image-1 needs particular care. It is a preliminary comparison built from five occupations with a single sentence, 20 images per occupation for 100 in total. Across those five, the open-source models averaged 84.5% male and GPT-image-1 gave 68.5%, and while that is statistically significant the effect size is a small 0.080. In the authors' own words the practical difference is modest. Three of the cells carry confidence intervals reaching ±20pp, and the resolution was 1024 rather than 512, so a difference in classification accuracy may be mixed in. For nurse, GPT-image-1 came to 35%, effectively the same as SD 1.5. Carrying this table over as evidence that commercial models are fairer would say something the paper does not. Versions from SD 3.5 onward were never measured at all, for want of hardware.

6

Why Pebblous Is Watching This Study

From here the article leaves the paper and rereads the result with an eye on data.

It helps to separate fact from interpretation first. What the paper actually measured is output. It is not a study that opened up the training data and checked what the distribution looked like. So the sentence claiming that the skew is caused by the distribution of the training data is our own inference rather than the paper's conclusion. The pull toward that inference is strong all the same. If the direction held while parameters and text encoder and training procedure all changed four times over, suspicion naturally turns to something the four versions share.

That said, the whole thing cannot be reduced to data alone. Zhao and colleagues' 2017 result, which this paper pulls in as related work, presses on precisely that point. Measured on visual question answering models rather than image generators, the finding was that a model does not merely carry over the gender bias present in its training corpus but hands it back enlarged. Evening out the data is a starting point, not a destination.

What this article wants from the paper is not a diagnosis of the cause but a procedure for judgment. The paper writes it as the last of three practical recommendations. Comparing outputs against real demographic data such as the BLS reveals deviations that internal 50/50 balance metrics do not capture.

The sentence becomes concrete once mechanic and cleaner sit side by side. On a yardstick that measures nothing but how far the output tipped, mechanic always outranks cleaner. Across the four versions the Stereotype Score for mechanic is 40, 43, 47 and 47, against 36, 30, 41 and 42 for cleaner. The order flips completely when the ranking runs on distance from reality instead. The amplification gap for mechanic runs +6, +3, −1 and −1, all clinging to zero, while cleaner sits at −32, −26, −37 and −38. Since 96% of working mechanics are men, a model can draw almost nothing but men and still not move away from reality. Cleaner really is close to even, and the model shoves it toward men. The one that needs fixing is the cleaner that ranked lower.

Mechanic vs. cleaner: the ranking flips depending on the yardstick Stereotype Score (0–50, lower is more balanced) 50 25 0 40 36 SD 1.5 43 30 SD 2.1 47 41 SDXL 47 42 SD 3M Amplification Gap (pp, closer to 0 is closer to reality) +10 0 −40 +6 −32 SD 1.5 +3 −26 SD 2.1 −1 −37 SDXL −1 −38 SD 3M Mechanic Cleaner arXiv:2609.18007 figures reordered by two yardsticks | Pebblous original diagram
▲ By the yardstick that looks only inward (left), mechanic seems more skewed. Measured against reality (right), cleaner's gap is far larger | Pebblous original diagram

An old question in data quality practice has the same shape. We call a dataset skewed, and mostly we leave the baseline of that judgment unstated. Classes spread evenly and we say balanced, one class dominates and we say imbalanced. That standard came from inside the dataset. It is a verdict reached without holding the data up against anything outside it. For defect inspection data the actual defect rate on the line belongs next to it, for demand forecasting data the actual transaction distribution does, and only then can anyone see where our data parted from it and by how much.

A team that takes datasets in or builds them and hands them on might check the following four things.

  • Where did the baseline come from that we use when we judge something skewed? Is it a value drawn from inside the dataset, or a measurement fetched from outside?
  • Did we write that baseline into the documentation? An unrecorded baseline leaves the next person to judge the same data by a different standard.
  • Are we watching the items that are close to even in reality separately? A skew metric is slowest to report distortion in exactly those items.
  • When a model or a pipeline version goes up, do we measure this distribution again alongside performance? The assumption that a new version sorts it out on its own missed twice in the three version changes this study observed.

Thanks for reading this far. The figures and charts this article carried over can be checked by anyone in the original at arXiv:2609.18007, and the generation configurations, classification scripts and aggregate results are public as well. The generated images themselves were held back to avoid potential misuse. We would be glad to hear what your own team compares its data against when it decides whether that data is skewed.

R

References

Academic Papers

Official Statistics