Pre/post difference-in-differences · quasi-experimental test planning
How big should my test be? How long should it run?
So, you want to test the impact of something? A price, a promo, a layout? And you want to know if it really works? This tool tells you two things before you start: how many locations you need for the test, and how long to run the test. This tool can be used for any test where you're comparing an average outcome between a treatment group and a control group and want to detect a percentage lift.
Here's how it works: You'll make the change for some of your locations (the test group), but leave some others alone (the control group). Then you'll watch how each group changes from before to after. Calculating and comparing the actual change in values against expected change in values is what separates your effect from everything else - seasons, weather, the economy, etc.
Why this exercise is important: If your test has too few locations, a real improvement will be impossible to tell apart from the standard week to week ups and downs. Worse, with too few locations you'll get unreliable results that look real but can't be replicated and can't be trusted. Answer a few questions below and you'll get a good estimate of test size and duration required for a reliable test.
I want to see if my intervention changes a continuous metric
across my
,
by at least %.
Pick a continuous measure - something you can add up and average, like money or a count (revenue, items sold, minutes). This tool is built for those.
It does not fit a percentage or rate (like "% of visits that end in a sale") - those need an extra input and a different formula, and you'd get a number that's too low.
Everything below updates live. Your numbers stay pinned to the bottom of the screen as you go.
How many stores do you have to work with?
This is your assignable population pool - the stores genuinely available to become a test or control store right now. It's what the feasibility check and the finite-population credit are based on, since it's the pool that actually limits your options and sets how big a slice you'd use.
Heads up: this can be smaller than the full set you eventually want your results to apply to. If 5,000 of 40,000 stores are tied up elsewhere, your assignable pool is 35,000 - enter that. Just know the excluded stores may differ from the rest, so generalizing back to all 40,000 is a separate judgment call the math can't make for you.
Count only the stores you could actually assign right now - the ones free to become a test or control store. Leave out any that are unavailable: already in another test, closing, or too unusual to compare fairly.
How different are your stores from each other?
The technical name is variability, usually measured as the coefficient of variation (CV) - how spread out your stores are compared with their average. More spread means you need more of them.
Think about a normal stretch of time. Do your stores do roughly the same amount of business, or are some giants and some tiny?
Do busy stores stay busy?matters most
The technical name is stability over time, also known as autocorrelation (ρ). It's the biggest impact to the number of required store: a before/after test compares each store to its own past, so the steadier each one is, the fewer you need.
If your stores do well one month, can you count on them doing about the same the next? Or do the numbers jump around?
How long can you run the test?
This is your measurement window (or test duration).
A longer test averages out daily ups and downs at each store, so it reduces the number of test stores you need. Pick what's realistic. You'll need to block this many weeks for a clean pre-test baseline period, plus the same number of weeks for the test itself.
weeks
How sure do you need to be?
This is your confidence level - your guard against false alarms (calling something a win when it was really luck). "Very sure (95%)" is the standard choice. Higher is stricter but costs more stores. Separately, note that this tool targets an 80% chance of catching a real effect - a common target which is hard coded here to keep things simple.
How confident do you need to be that your test results are real? Stricter means lower chance of false alarms, but also requires more stores.
You need
-
test stores
+
-
control stores
Looks doable-
Have stores to spare?
Fewer test stores to change is easier and safer to roll out, but the benefits shrink fast. Using 3x as many controls only reduces the test size slightly vs 2x, and the smaller your test group, the more risk that a few odd stores can impact results (a closure, a new competitor, bad data).
How accurate is this recommendation?
-
-
How long to run it
-
You'll want to choose weeks that cover similarly performing periods, or for better accuracy, apply a year-over-year seasonality adjustment. Learn more here.
How to read the numbers
They're estimates from rough descriptions, so treat them as a ballpark, not to-the-digit targets. When in doubt, plan for the higher end.
One thing to check before you trust any of this
Getting the count right doesn't make the test fair. A before/after test assumes your two groups would have moved together if you'd done nothing - the parallel-trends assumption. If your test stores were already pulling ahead, the test will show an impact that was never there.
Easy gut check: before you start, chart both groups' numbers for the last several months. The two lines should rise and dip together. If they don't, swap some stores until they do - then run your test.
The math, briefly. Test stores per group ≈ (1 + 1/k) · (zsure + zcatch)² · CV² · (1 − ρ²) · w / lift² - the standard two-group formula, with a bonus for using each store's own history. Because the average size of your stores cancels out, no real data is needed. CV = how different your stores are; ρ = how steady they are; k = control-per-test ratio; w = window factor (drifts below 1 as you add weeks). You also get credit when the test would use a big slice of everything you have. Aimed at an 80% chance of catching a real effect, two-sided. Decide the size up front and don't stop early the first time it looks like a win - and if your change causes a novelty bump, skip the first few days.
-
Save these settings
Opening this link reloads every setting exactly as you have them.
Finished here? Learn how to setup a diff-in-diff test.
How to set up your periods, calculate impact, and adjust for seasonality.