Skip to content
The research library
Scheduling and business performanceModerate evidence

How Do You Judge a Scheduling Study?

A short field guide to reading scheduling research, worked through this library's own sources, including the claims we checked and cut.

Reviewed against primary sources on July 19, 2026 by the Soon operations research team

The evidence in one line

The first thing to check in any scheduling study is not the effect size, it is how people ended up in the groups being compared. The Gap study randomized stores into treatment and control, which is why its 5.1% productivity result supports causal language, while the Shift Project studies surveyed workers at one point in time on a non-probability sample and support association only (Kesavan et al., 2022; Schneider & Harknett, 2019, 2021). Same topic, very different warrant, and almost everything else about reading this evidence well follows from telling those two situations apart.

Start with how the groups were formed

Every claim about scheduling rests on a comparison, and the credibility of the claim is set by how the two sides of that comparison came to exist. In the Gap study, stores were randomized into treatment and control, so the groups were alike on average before anything changed and the difference that followed can be attributed to the treatment. That is the reason its 5.1% productivity result supports causal language such as raised or increased (Kesavan et al., 2022).

The Shift Project studies ask closely related questions and earn a very different warrant. They survey workers at one point in time on a non-probability sample recruited through social media advertising, so nobody assigned anyone to a stable or an unstable schedule (Schneider & Harknett, 2019, 2021). Workers reporting unpredictable schedules differ from workers reporting predictable ones in ways a snapshot cannot fully separate, including which employers hire them and which jobs they are able to accept, so the findings support association and nothing stronger.

The practical rule follows directly. If a write-up never tells you how people ended up in each group, treat the finding as a pattern worth knowing rather than a lever worth pulling. Associations remain useful, because they tell you where to look and what to measure in your own operation, but they do not tell you what will happen when you change the schedule.

Then ask which version of the number you are looking at

A single experiment usually produces several headline numbers, and the biggest one is rarely the one you should carry into a plan. The Gap experiment reports 5.1% as its intent-to-treat estimate and 16.4% as the adherence-adjusted treatment-on-treated estimate (Kesavan et al., 2022). Intent-to-treat counts everyone assigned to treatment whether or not managers followed the practices evenly, so it is often the more policy-relevant estimate for a real rollout. It remains one result from one retailer, not a lower bound or promised return for your operation.

Survey work carries a parallel distinction between what was observed and what was modeled. The hardship gradient in the Shift Project research, where hunger runs from 23% to 50% across the range of schedule instability, is a predicted probability, adjusting for wages, income, and hours, not the observed share of workers in each group (Schneider & Harknett, 2021). The adjustment is a strength, since it answers the objection that the pattern is really about pay, but the resulting figure describes statistically comparable workers rather than a headcount anyone counted on a shop floor.

The strictest version of this check applies to modeled counterfactuals. The 2019 authors model what eliminating on-call shifts would be associated with, and they state plainly that these estimates are not causal (Schneider & Harknett, 2019). A projection like that extends an observed pattern into a hypothetical, and it is not a measured effect of a policy anyone actually changed. When a deck quotes such a figure as an expected result, the number has been promoted a rung it did not earn.

Ask what was in the treatment, and what the comparison group already had

Even a clean randomized result gets misread if you skip what was actually delivered. The Gap experiment assigned stores to a bundle of six practices that moved as one package, so the 5.1% productivity gain belongs to the package rather than to any member of it (Kesavan et al., 2022). Our write-up of the Gap experiment sets out what the bundle contained. The reading lesson is narrower than the contents: an experiment can only speak to the unit it randomized, and here the unit was the whole set.

The companion question is what the comparison group already had, because an effect size measures the distance between two arms and not the distance from zero. Gap's control stores were not sitting at an untouched status quo, as the Gap write-up explains. The habit generalizes further than that one case, so before porting any published number into a plan, ask what the study's comparison group was already doing on the day measurement started.

Diagnosing your own baseline is concrete work rather than a caveat you nod at. Write down how much notice your sites actually give, how often published shifts get canceled or moved, how many people work variable start times from week to week, and how much of the schedule is set by one manager's habit. If that inventory sits below where the study's comparison arm stood, the published estimate is not measuring the move you are about to make. Your early gains could run larger, because the cheapest fixes tend to come first, or smaller, because your operation differs in ways the design never tested.

What we excluded from this library

Two widely circulated talking points did not survive our verification and are deliberately absent. The first is the claim that unstable schedules matter more than low wages. That ranking did not verify against the sources it is usually attributed to, and we do not repeat it.

The second is a bundled set of prevalence statistics about how many workers get short notice or work clopenings. The bundle failed verification as it circulates, though a couple of its components do appear individually in verified sources elsewhere in this library. Statistics that travel as a bundle deserve to be split apart, because bundles pick up a stray figure or a mismatched year as they pass through blog posts and decks.

Apply the same test to anything quoted at you, including by us. Ask for the study rather than the slide, then confirm that the figure on the slide is the figure the study reports, that it is attached to the design that supports it, and that the comparison group is what you assumed. Most failed scheduling claims break at one of those checks rather than because anyone set out to mislead.

What this means for your schedule

  • Ask how people were sorted into groups before you ask how large the effect was.
  • Describe 5.1% as the experiment's observed intent-to-treat estimate and 16.4% as its adherence-adjusted estimate, without turning either into a promised planning range.
  • Inventory your own baseline in writing before porting any published effect size, since a study measures the gap between its two arms and not the gap from where you are standing.
  • Say the word associated out loud whenever you repeat a Shift Project finding, and never upgrade a modeled projection into a promised effect.
  • Trace every borrowed statistic back to its study before it enters a board deck, and drop the ones that will not resolve.

The business case

Scheduling claims circulate far ahead of the evidence behind them, and a business case built on an unverified figure collapses the first time someone checks it.

The defensible way to use the randomized experiment is to label 5.1% as its observed intent-to-treat estimate and validate the local effect with a controlled pilot, while treating survey findings as risk signals rather than promised returns (Kesavan et al., 2022).

Requiring every number in a proposal to name its study and its design costs a review cycle and removes the category of error that quietly erodes trust in the whole initiative.

Frequently asked questions

How should I picture the difference between intent-to-treat and treatment-on-treated?
Picture a chain that assigns a scheduling bundle to a set of stores. Some managers adopt it fully, others keep most of their old routine, and the results get averaged across all of them anyway. That average is intent-to-treat. Treatment-on-treated estimates what happened where the practices were delivered, using assignment to adjust for adherence. In the Gap experiment those figures are 5.1% and 16.4% for productivity (Kesavan et al., 2022). They answer different questions under different assumptions, so the distance between them is not a menu of returns you get to pick from.
Can I conclude anything about a single scheduling practice from the Gap result?
You can conclude that the package is worth testing, since the Gap experiment randomized stores to a bundle of six practices and measured 5.1% higher productivity for the bundle as a whole (Kesavan et al., 2022). You cannot rank the parts, because nothing in the design separated them, and the result does not forecast the return in another operation. If the ranking matters to your budget, generate the evidence locally: stagger practices across comparable sites in a deliberate order, hold the rest of the operation steady, and measure each step before adding the next.
What is an association about scheduling actually good for?
It is good for sizing risk and choosing what to measure, not for promising a return. Schneider and Harknett (2021) report that the predicted probability of hunger, adjusting for wages, income, and hours, runs from 23% among the most predictably scheduled workers to 50% among the least. That gradient tells you the stakes attached to instability among the workers surveyed, and it survives an adjustment for pay, which makes it hard to wave away. It still does not establish that changing a schedule moves the outcome, so it belongs in the section of a proposal about what you plan to watch.
Which check caught the claims you left out?
Attribution caught the first one. The ranking claim that unstable schedules matter more than low wages could not be traced back to a source reporting it, and Schneider and Harknett (2021) treat wages, income, and hours as covariates to adjust for rather than as a rival factor to rank against scheduling. Splitting the bundle caught the second. A packaged set of prevalence statistics about short notice and clopenings did not hold together once each component was chased to its own source and year, even though a couple of the components appear individually in verified sources here. A smaller library that holds up is worth more than a larger one that does not.

Sources

Every figure on this page is drawn from a cited primary source and checked against the original publication.

  1. Kesavan, S., Lambert, S. J., Williams, J. C., & Pendem, P. K. (2022). Doing Well by Doing Good: Improving Retail Store Performance with Responsible Scheduling Practices at the Gap, Inc. Management Science, 68(11), 7818โ€“7836. https://doi.org/10.1287/mnsc.2021.4291

    Randomized controlled field experiment (28 stores, roughly 150,000 shifts, about 1,500 employees)

  2. Schneider, D., & Harknett, K. (2019). Consequences of Routine Work-Schedule Instability for Worker Health and Well-Being. American Sociological Review, 84(1), 82โ€“114. https://doi.org/10.1177/0003122418823184

    Cross-sectional survey (27,792 hourly workers at 80 large firms, 2016-2017, non-probability sample)

  3. Schneider, D., & Harknett, K. (2021). Hard Times: Routine Schedule Unpredictability and Material Hardship among Service Sector Workers. Social Forces, 99(4), 1682โ€“1709. https://doi.org/10.1093/sf/soaa079

    Cross-sectional survey (37,263 hourly workers at 127 large firms, 2017-2019, non-probability sample)

None of the studies cited here evaluated Soon.They examine scheduling practices, shift patterns, and working hours as studied by independent researchers, so their findings describe what those practices are associated with, not what any particular software produces.

This article summarizes published research for scheduling and operations decisions. It is not medical advice. Individual health questions belong with a qualified clinician.

Your next schedule could take 2 minutes.

Import your team, set your rules, hit auto-fill. Most teams are live the same day.

Try Soon free

30 days free ยท No credit card required

Already have an account?Sign in