Administrator
Published on 2026-08-13 / 5 Visits
0
0

Worker Retraining and AI: Evidence from 56 RCTs

Worker retraining is the default answer to AI job displacement, but the evidence supports a narrower claim. A new review of 56 randomized US studies finds that training raises employment and earnings on average, yet the gains are modest. The few programs with much larger results depend on selection, employer relationships, and skills that can be taught within months, which makes scale the central policy constraint.

Evidence snapshot: 13 August 2026. Dollar values are in 2025 US dollars unless stated otherwise.

Reading time: 9 minutes · About 1,750 words

TL;DR

  • The published review combines 146 impact estimates from 56 US randomized trials conducted from 1973 onward, plus a smaller body of European evidence.
  • For training-primary programs, employment increased 2.8 percentage points in year two and 1.7 points in years three to five. Annual pre-tax earnings rose $1,139 and $791 in those periods.
  • These are intent-to-treat effects per person offered a place, not effects per graduate. Low take-up and control-group access to other training dilute the measured program effect.
  • Programs cost about $13,598 per offer among the 67% of training-primary interventions reporting costs. The break-even result relies on uncertain assumptions about benefits lasting to retirement.
  • Some sector programs produce far larger gains, but they admit roughly one in five applicants and rely on local employer demand. Their success is difficult to copy at crisis scale.

What the 56 studies actually measure

Anthropic researcher Maxim Massenkoff and independent researcher David Roodman published a targeted evidence review and meta-analysis on 12 August 2026. The authors screened for high-quality causal studies and assembled 56 randomized US evaluations. The programs mostly served low-income adults and young people, lasted about six months, and trained people for jobs such as nursing aide, IT support technician, and welder.

This is historical workforce evidence, not an experiment on people displaced by large language models. The distinction matters. The report asks whether past programs offer a useful baseline for a possible future shock; it does not claim that 56 trials have tested AI displacement.

The published PDF reports 146 impact estimates. The linked GitHub repository describes itself as in progress and, on 13 August, its README listed 144 estimates from the same 56 studies. The difference looks like a versioning lag rather than a substantive contradiction. For published counts and conclusions, the frozen report is the controlling artifact; the repository is useful for inspecting methods and exploratory outputs.

The average result is positive and small

The headline depends on the follow-up window:

Outcome for training-primary programs Year 2 Years 3–5
Employment impact +2.8 percentage points +1.7 percentage points
Annual pre-tax earnings impact +$1,139 +$791

The report's executive summary rounds the long-run result to a 1.7-point employment gain and about $800 a year. Anthropic's web page compresses the evidence further to two to three percentage points and roughly $1,000. Those summaries are compatible, but the table above preserves the time dimension that the rounded headline hides.

Large federal programs did not escape this pattern. The Job Training Partnership Act generated adult effects close to the average and no clear gains for low-income youth. Long-term Job Corps follow-up found little effect on employment or earnings beyond a temporary employment bump. The Workforce Investment Act Gold Standard Evaluation also found mostly small or statistically uncertain training effects, partly because many people offered training did not take it while control-group members found training elsewhere.

That last mechanism is crucial for reading the numbers.

An offer is not the same as completed training

The main estimates are intent-to-treat, or ITT: the effect for every person randomly offered access. They measure whether a program works in its real delivery context, including refusal, dropout, waiting, and alternative training.

Across the trials, assignment increased participation in the targeted program by 66 percentage points, but increased participation in any training by only 28 points. Some people in the treatment arm never trained; some controls received training outside the experiment.

The authors estimate that effects on marginal participants may be two to four times the ITT estimates. They also warn that this scaling is likely an upper bound. It answers a different question: what might training do for people whose participation was changed by the offer? It cannot be substituted for the population-level effect of opening a program.

For AI policy, ITT is the more operational metric. A government facing mass displacement needs to know the result per offered place and per dollar appropriated, not only the result among people who complete a course.

The economics are close, not conclusive

Cost data were available for 67% of training-primary interventions. In that subset, average cost was $13,598 per treatment-group member. The authors estimate $14,146 in present-value pre-tax earnings gains, rising to $19,525 after adding employer-paid taxes and fringe benefits.

This looks slightly better than break-even, but the result projects benefits beyond the few years observed in most studies and assumes that gains persist toward retirement. The authors explicitly call those assumptions debatable.

The fiscal calculation is similarly conditional. Higher tax revenue and lower benefit payments are estimated to recover about three quarters of government outlay, leaving a long-run net fiscal cost near 24% of upfront spending. Treating Social Security contributions as future liabilities rather than pure tax revenue reduces the apparent payback.

The defensible conclusion is therefore modest: ordinary retraining can be a reasonable public investment. Its average effect is too small to offset a sustained, large unemployment shock by itself.

Why sector programs look different

Sector programs connect training to actual local demand. They involve employers in curriculum design, teach occupational and workplace skills, provide coaching, and help place graduates into quality jobs. Programs such as Year Up and Per Scholas produced annual earnings gains of $5,000 to $10,000 in some trials.

Across the report's sector-program group, average cost was $11,602, below the $13,598 average for training-primary programs, while the discounted present value of pre-tax earnings gains reached $60,319. The estimated benefit-cost ratio was about seven.

The result is impressive, but the mechanism contains its own scale limits:

  • Programs commonly accept about 20% of applicants. Selection identifies people likely to complete demanding training and succeed at work.
  • Employers help define the skills and often support placement. This requires relationships that vary by city and industry.
  • Effective courses target skills teachable in months. They do not solve a mid-career professional's need for years of education to recover a previous income.
  • Replication has a mixed record. CET failed across 14 replication sites, new WorkAdvance sites underperformed the original, and Year Up was the clear success among the report's three prominent scale attempts.

Scaling the number of seats can therefore move the bottleneck to employer demand, applicant readiness, program operators, or the supply of jobs. Copying the curriculum without those complements copies the visible component and loses the production system.

The AI extrapolation has four hard boundaries

First, the study population differs. Most trial participants were low-income adults or youth, while AI could displace accountants, paralegals, developers, administrators, and other workers with different earnings expectations and constraints.

Second, the training horizon differs. The successful programs usually teach a marketable skill in less than a year. Deep professional conversion may take several years, during which technology and employer demand can change again.

Third, the shock scale differs. A local intermediary can place hundreds or thousands of screened applicants into expanding fields. It has not demonstrated that it can absorb millions of workers arriving at once.

Fourth, the job side of the market is endogenous. Training redistributes access to available jobs; it does not create unlimited demand. A program cannot place every displaced worker into a high-demand sector after that sector's vacancies are filled.

These boundaries separate an evidence-based baseline from a prediction. The 56 trials show that retraining can help. They do not show that today's institutions can absorb rapid AI displacement.

A better policy test: scale and randomize before the shock

The report recommends a fire drill: rapidly expand a promising sector program for a defined population and evaluate the expansion rigorously. That approach tests the part we know least about, which is whether the mechanism survives scale.

A useful demonstration should pre-register four outcomes:

  1. ITT employment and earnings, reported separately from effects per trainee.
  2. Acceptance, take-up, completion, and placement rates across the full applicant funnel.
  3. Employer demand, job quality, and displacement of non-participants.
  4. Program cost, fiscal recovery, and results after three to five years.

This complements current measurement of AI exposure in the labor market and evidence that AI moves tasks before it removes whole jobs. Exposure identifies where pressure may arrive. A scaled retraining trial tests whether the proposed response can carry the load.

FAQ

Do worker retraining programs work?

On average, yes, but the gains are modest. The randomized evidence shows statistically positive employment and earnings effects, especially after the first year, without a transformative population-level impact.

Why are effects larger per trainee than per person offered training?

Not everyone offered a place participates, while some control-group members find training elsewhere. ITT preserves those real-world frictions; per-trainee estimates scale around them and require stronger assumptions.

Is $13,598 per offered place worth it?

The report estimates benefits around or somewhat above costs, but the calculation depends on gains persisting beyond observed follow-up. It supports a cautious public-investment case rather than a guaranteed financial return.

Why are sector programs hard to scale?

Their results depend on selective admissions, strong local employers, staff capability, coaching, placement, and skills that can be taught quickly. Seats can be funded faster than those complements can be built.

Do the 56 trials prove retraining will work for AI-displaced professionals?

No. They provide historical causal evidence, mostly for lower-income populations and short-duration programs. Applying it to rapid, high-skill AI displacement is an explicit extrapolation.

References


Comment