Executive Summary
Between January 2020 and June 2023 the Swiss federal migration office split about 2,000 refugee cases at random into two arms. A case is one household placed together, or one person where someone arrived alone. One arm had a canton chosen by calculating employment probability shown on the placement officer's screen as a recommendation; the other had a random recommendation that imitated the existing procedure. The two screens looked identical, and neither the officer nor the refugee knew which was which. This article looks at what those three and a half years established and what they did not.
The pre-registered primary outcome is the share of the three years after placement spent in work. The control group came in at 22.3 percent, and the side that received the AI recommendation was 2.2 percentage points higher. Drop the 2020 placements, made while COVID-19 was shaking the labor market, and the gap is 2.7 points; look only at the 2022 and 2023 placements, made after the market had largely returned to normal, and it widens to 3.9. The lower end of the 95 percent confidence interval for the full sample sits at +0.05 points, though, a hair above zero.
Sections 1 through 4 report what the paper and its appendix say. Section 5 rereads the trial's design through the lens of data quality, and that reading is this article's own.
Key Figures
Source: Bansak et al. (2026). arXiv:2609.35448, Table 1 in the main text and appendix Tables S9 and S13.
+2.2pp
Share of months worked in three years
About 10 percent above the control group's 22.3 percent. 95% CI +0.05 to +4.33pp
+3.9pp
2022–2023 placements
Among the 1,212 cases placed after the labor market returned to normal, the effect grows to about 17 percent
97%
Recommendations the officers followed
Final authority rested with a person, and that person did not know which kind of recommendation it was
Half
Realized against the 2018 forecast
The 2018 backtest promised 11 percentage points; measured the same way, the trial returned 5.2
Nobody Knew Who Got the AI Recommendation
Where a refugee settles shapes the work that follows for a long time. In a study across 20 European countries, refugees were about 12 percent less likely to be employed than otherwise comparable migrants, and the gap persisted 10 to 15 years after arrival. That the first months and first few years weigh heavily on everything after them is a long-standing observation in this field. Yet in most countries the place of settlement is fixed by administrative rules such as population share. Who is likely to do well where does not enter the criteria. The person making the assignment has almost no information to base such a judgment on.
Someone who applies for asylum in Switzerland and receives protection status is assigned to one of 26 cantons. Which canton is not the person's choice; a placement officer at the federal migration office, the State Secretariat for Migration, decides. Placements have to respect a distribution key drawn up in proportion to cantonal population, and several nationality groups have to be balanced separately across cantons. Within those constraints the officer decides, day by day, who goes where.
GeoMatch, built by the Immigration Policy Lab at Stanford University and ETH Zurich, lays one recommendation on top of that step. A prediction model trained separately for each canton estimates how much this person would work over three years in each canton, and the tool picks the canton that makes total predicted employment as large as possible while respecting the constraints. Cases are handled one at a time as they arrive, so a decision has to be made now without knowing who comes next. The tool dealt with that by sketching the mix of cases still to come and choosing the canton that disturbs that expected distribution least. It is a rule against spending in advance the places later arrivals will need.
The trial ran by preparing two versions of that recommendation screen. An eligible case had been routed through the fast-track procedure, had already received subsidiary protection or refugee status, was not legally required to go to a particular canton, and included at least one adult. Every case meeting those terms was split as if by a coin toss, half shown the canton the algorithm had chosen and half shown a canton drawn at random in imitation of the existing procedure. The two screens had the same format, and nothing marked which procedure had produced the recommendation. An officer could respond to the recommendation on the screen but not to its source. The refugee did not know either.
The quotas were in two sets from the start. The distribution key and the nationality balancing rules were copied identically for the treatment and control arms, so the two groups were never competing for the same slots. That makes it structurally impossible for the algorithm to raise its score by crowding people into cantons with strong labor markets. The overall spread across cantons and nationalities is the same on both sides, and the only thing that changes is who goes to which canton within that spread. The distribution of realized placements did in fact turn out statistically indistinguishable between the two groups.
A recommendation was not an order. The officer could take it or overturn it, and final authority stayed with a person to the end. The recommendation on the screen was followed 97 percent of the time. Leaving the assigned canton within three years was rare, 5.4 percent in the control group, and the rate was no different in the treatment group. Swiss rules restrict movement between cantons in law. So the canton first assigned was where the person actually lived for the full three years. Attrition at the three-year observation was low and similar on both sides, 2.5 percent of control adults and 2.3 percent of treatment adults, and pre-placement characteristics were evenly spread across the two groups. Outcome data came from employment histories in administrative registers that are collected regardless of the experiment. Nobody was sitting there scoring the results.
This trial is rare less for its algorithm than for how it was recorded. Because the placement officer did not know which arm a case belonged to, the difference that showed up three years later cannot be explained away as officers trusting the AI more, or pushing back against it. The evidence for an effect was not pieced together once the results arrived; it had been accumulating since the first placement.
More Months of Work, and a Gap That Widens
The pre-registered primary outcome is not whether someone worked but how long. It counts how many of the 36 months after placement were spent in work, as a share, averaged over adults where a case holds more than one. The control value was 22.3 percent. Eight months of work out of thirty-six. Against it, the side that received the AI recommendation was 2.19 percentage points higher. In relative terms about 10 percent, with a 95 percent confidence interval from +0.05 to +4.33 points and a p-value of 0.045. The lower end of the interval all but touches zero, so the result sits on the boundary between having an effect and having next to none.
That value is measured against the random assignment itself, regardless of whether the recommendation was actually followed. For the cases actually placed as recommended it comes to 2.49 percentage points. The two calculations do not diverge much because the random assignment split actual placements almost exactly. The probability of ending up in the recommended canton differed by about 88 percentage points between the two groups. Unless stated otherwise, the rest of the figures in this article are also measured against the random assignment.
Pooling everything this way hides circumstances that differ by period. The trial started in January 2020, and the pandemic shook the labor market immediately after. The model had been trained on pre-pandemic administrative data. So the paper reports two more values: one excluding the 2020 placements, and one for the 2022 and 2023 placements alone. The later the window, the larger the effect.
The paper tests that cohort difference on its own. Measuring the gap in effect between the 2022–2023 placements and the 2020–2021 placements directly gives a p-value of 0.039 in a covariate-adjusted test. Broken out by year, the effects for the 2020 and 2021 placements are indistinguishable from zero, then turn clear in 2022 and 2023 and hold there. The paper's diagnostics show why. Moving placements is something the algorithm managed in both periods. Cases that received a recommendation went more often to cantons high in their own predicted ranking, and that showed up during the pandemic as well. What changed is how closely the ranking tracked actual outcomes. During the pandemic the relation between rank and actual employment was nearly flat at the top of the ranking. In the range where recommendations cluster most heavily, the model could not tell cantons apart. Once the labor market returned to normal, that relation climbed steadily with rank. The model's power to move placements held throughout. What shifted for a while was the labor market those placements aimed at.
The effect grew as time passed. One caution, though: the metric changes once here, so the numbers have to be carried over carefully. At the 36-month mark what the paper looks at is not the share of months worked but whether the case has at least one adult in employment. The control value on that measure is 47.7 percent, and the AI-recommended side was 5.2 percentage points higher. At the individual level it is 4.6 points against a control value of 44.8 percent. That is why 22.3 percent and 5.2 percentage points cannot be laid side by side to produce a relative ratio. The denominators are different.
The difference is concentrated in the durable kinds of work. On the measure that asks only whether someone worked at all within three years, the difference came to 2.26 percentage points and the confidence interval crosses zero comfortably. On that measure the two groups are effectively indistinguishable. By contrast, the share of cases holding a job that ran unbroken for more than six months was 5.57 points higher, climbing to 9.56 points among the 2022–2023 placements. The share working on an open-ended contract at the three-year mark was 4.45 points higher, and 7.07 points among the 2022–2023 placements. Time to first employment was shorter by 0.78 months on average, and by 1.45 months among the 2022–2023 placements. So the recommendation moved when work started and how long it held, more than whether it came at all. Of these four, only the first was written into the pre-analysis plan; the three covering durability and speed were added afterwards.
The authors set this size against other policies. When Denmark's 1999 reform added about 430 hours of language training, full-time-equivalent employment rose 4.2 percentage points on an 18-year average, and Germany's 600-hour integration course raised employment about 5 points a year later. On the closest comparable measure, individual employment at the three-year mark, this trial produced 4.6 points. A Swedish program bundling language classes, work practice and job-search support came to 15 points and an Italian one with vocational mentoring and subsidized internships to 10, both larger, but they cost 2,400 and 3,000 Swiss francs per participant and run on a great deal of staff and facilities. On cost the difference is much wider. In the paper's calculation the additional operating cost of laying the tool onto existing work is 50 Swiss francs per adult, and valuing one added month of employment at 1,950 francs gives a benefit-to-cost ratio of roughly 27 to 1. Per adult that is 0.70 extra months of work over three years, leaving a net fiscal benefit of about 1,300 francs per person after operating costs. That calculation counts only the support payments the government saves and the taxes and social insurance contributions it collects, not the income left in the refugee's hands. The authors also state flatly that the comparison should not be read as saying algorithmic placement can stand in for intensive language training.
Why the 2018 Forecast Was Twice as Large
This approach first became known through a 2018 paper in Science. The same researchers trained a model on Swiss register data from 1999 to 2013, held out the 2013 arrival cohort as a test set, and calculated that optimized placement would have lifted that cohort's employment from 15 percent to 26 percent. Eleven percentage points, or 73 percent in relative terms. That figure was quoted as it stood in articles and briefing material for the eight years since.
Measured the same way, the trial reached a little under half of that forecast. The appendix makes the comparison directly. The measure closest to the 2018 calculation is individual employment at 36 months after placement, and the value there is 4.6 percentage points. At the case level it is 5.2. A calculation that promised 11 points landed a little over 5.
| Item | 2018 backtest | 2026 randomized trial |
|---|---|---|
| Compared against | Past records of the 2013 arrival cohort | A control group split off at random in the same period |
| Baseline employment | 15% | 44.8% (individual level) |
| Effect size | +11pp (73% relative) | +4.6pp (about 10% relative) |
| Nationality quotas | None. The advantage of clustering with co-nationals was available | Fixed per canton. That channel was mostly closed |
| Placement method | A whole year optimized at once, after the fact | One case at a time, in order of arrival |
Most of the gap in relative terms comes from the denominator. The 73 percent was calculated on a baseline of 15 percent, and the 10 percent on a baseline somewhere near 45. An identical rise in percentage points looks far larger as a proportion at the lower end. What happened over those eight years is that employment among refugees in Switzerland went up, not that the algorithm shrank to a seventh of what was expected.
In percentage points, though, roughly half the gap remains. Several reasons overlap here. The 2018 calculation had no nationality quotas, so it could freely use the advantage of sending people to cantons where co-nationals were already established or where their mother tongue was spoken, while the actual trial fixed the nationality spread per canton. The calculation matched a full year of cases after the fact, whereas the actual placements went one at a time without knowing who would come next. The deployed model had fewer predictors and had been trained on pre-pandemic data.
Does that make the backtest an untrustworthy instrument, then? Not quite. The appendix reran the same kind of calculation against the policy as actually deployed. For the 2022–2023 placements it forecast a gain of 3.0 percentage points, and the value that came out was 3.9. Line the conditions up and forecast and outcome sit close together. The fault was not in the principle of the calculation but in the distance between the world it assumed and the world the tool was placed in.
There is a sentence in the paper's introduction this article held onto longest. Every published estimate of this approach's gains has rested on calculations that look backward over past records, and such calculations cannot fully carry the frictions of the field, how staff respond, cases where the recommendation is not followed, the crowding out that happens when places are contested, and the conditions that changed between training and deployment. The five items on that list do not shrink because the model got better.
The Model Saw Only Seven Things
About any one person, the model held seven items. Gender, age on arrival, marital status, the number of people in the case, year and month of arrival, nationality, and mother tongue. That was all the migration office had selected out of the Swiss administrative registers for placement use.
- • No education. Whether someone finished secondary school or university, the model does not know
- • No occupation or work history. What the person did back home, it does not know
- • No local language ability. Whether the person speaks German or French, it does not know
The variables best known for predicting employment are missing wholesale, in other words. They are also items most other countries can usually obtain. That is the first of the authors' reasons for calling this a conservative trial. The other reasons run the same way. New methodology had appeared, but the model and the optimization procedure were left untouched to avoid disturbing the definition of the treatment; the distribution key and the nationality balancing rules narrowed the cantons available to choose from; and splitting the quotas in two meant the optimization was in practice working over only half the cases. The pandemic landed on top of that.
The other side needs writing down as well. The pre-analysis plan was registered in March 2023. That is three years after the trial began, but before the employment outcome data arrived. The cohort comparison is not an analysis written into that plan but one added later, and the paper marks it as such. A test comparing the distributions as a whole did not reject the hypothesis that the treatment group is no worse off than the control group, which is not proof that nobody was harmed so much as a statement that there is room to see it that way. A test viewing the same distributions from another angle rejected the hypothesis that the two are identical, in the treatment group's favor, with a p-value of 0.035 across the full sample and 0.007 among the 2022–2023 placements. The paper records that this shift appears across nearly the whole range rather than in one segment of the distribution. The requirement to obtain informed consent was waived by the institutional review board, on the grounds that the placement sits inside an administrative procedure already under way and the analysis runs on administrative records.
Why Pebblous Is Watching This Trial
From here the same trial gets read again as a data quality problem.
What made it possible to look back at the effect three years later was the recording, not the model. Cases that received a recommendation and cases that did not were split at random and recorded as two groups; whether the officer followed the recommendation stayed on the record, which is what allowed the 97 percent compliance rate and the estimate adjusted to it. The outcomes were drawn from administrative registers that run regardless of the experiment. Splitting the quotas into two sets kept the groups from contesting the same slots and clouding each other's results. Take away any one of those four and the numbers produced in 2026 would have stopped at the level of "it seems that way."
Shrink the same structure down and it becomes a question any organization faces. A flow where AI recommends a candidate and a person checks it and decides is already inside résumé screening, customer-service prioritization and equipment inspection order. For that decision to be answerable a few months later when someone asks whether it was any good, the record has to hold the following.
- • Whether cases that received a recommendation are separated from those that did not. If everything got a recommendation, there is nothing to compare against
- • Whether what split the two was random, or whether only the hard cases went to a person. In the second case the difference comes from difficulty, not from the tool
- • Whether it is recorded case by case that the operator followed the recommendation. Without that column, the effect of using the tool cannot be separated from the effect of the tool being there
- • Whether whoever scores the outcome knows which cases got a recommendation. If they do, the scoring becomes part of the outcome
- • Whether retraining is logged with its date. Without it, a change in performance cannot be attributed either to the model or to the world
The last item is also where this trial showed the most. When the training data and the labor market the tool ran in drifted apart, the effect sank, and when the market came back, the effect came back. GeoMatch was retrained every quarter on fresh administrative data, and because each retraining was logged with its date, the cohorts could be pulled apart and examined. The authors recommend watching whether predictions still hold and retraining when conditions change. Doing that means leaving the records the judgment will need before the judgment is due.
The scene Pebblous runs into often in data quality work resembles this one. Months go into choosing a model, and nothing is left in place to check afterwards what that model decided. The valuable part of the Swiss case is not the 2.2 percentage points but the preparation that made it possible to pull that number out three years later. Without the preparation, what remains a few years on is the impression that things have improved since adoption, and that impression is just as plausible pointed the other way.
Thank you for reading this far. The paper itself, appendix included, can be read at arXiv. If your organization has a flow where a person takes an AI recommendation and decides, we would be glad to hear whether the record holds the columns needed to look back at that decision, and which of them turned out to be empty.
References
Primary Sources
- 1.Bansak, K., Hainmueller, J., Hangartner, D., Ferwerda, J., Paulson, E., Delevoye, A., Adams-Cohen, N., Ramaswami, A., Kurer, S., Pianzola, J. & Hotard, M. (2026). "AI-based matching improves refugee employment in a double-blind randomized trial." arXiv:2609.35448. Table 1 in the main text; appendix sections S9, S12, S13.
- 2.Bansak, K., Ferwerda, J., Hainmueller, J., Dillon, A., Hangartner, D., Lawrence, D. & Weinstein, J. (2018). "Improving refugee integration through data-driven algorithmic assignment." Science 359(6373), 325–329.
Critical Reviews
- 3.Achugamonu, S. (2025). "Place Matters: The Possibilities and Pitfalls of Machine Learning Algorithms for Refugee Resettlement." University of Chicago Public Policy Projects.
- 4."GeoMatch/MisMatch: A Critical Investigation of a Refugee Resettlement and Labour Market Integration Algorithm in the Netherlands." Social Inclusion.
Background
- 5.Immigration Policy Lab. "Improving Refugee Integration Through Data-Driven Algorithmic Assignment." Stanford University · ETH Zurich.