« andrewvalentine.com
Pre/post difference-in-differences · quasi-experimental test planning

How big should my test be? How long should it run?

So, you want to test the impact of something? A price, a promo, a layout? And you want to know if it really works? This tool tells you two things before you start: how many locations you need for the test, and how long to run the test. This tool can be used for any test where you're comparing an average outcome between a treatment group and a control group and want to detect a percentage lift.

Here's how it works: You'll make the change for some of your locations (the test group), but leave some others alone (the control group). Then you'll watch how each group changes from before to after. Calculating and comparing the actual change in values against expected change in values is what separates your effect from everything else - seasons, weather, the economy, etc.

Why this exercise is important: If your test has too few locations, a real improvement will be impossible to tell apart from the standard week to week ups and downs. Worse, with too few locations you'll get unreliable results that look real but can't be replicated and can't be trusted. Answer a few questions below and you'll get a good estimate of test size and duration required for a reliable test.

I want to see if my intervention changes a continuous metric across my , by at least %.

Everything below updates live. Your numbers stay pinned to the bottom of the screen as you go.

How many stores do you have to work with?
Count only the stores you could actually assign right now - the ones free to become a test or control store. Leave out any that are unavailable: already in another test, closing, or too unusual to compare fairly.
How different are your stores from each other?
Think about a normal stretch of time. Do your stores do roughly the same amount of business, or are some giants and some tiny?
Do busy stores stay busy?matters most
If your stores do well one month, can you count on them doing about the same the next? Or do the numbers jump around?
How long can you run the test?
A longer test averages out daily ups and downs at each store, so it reduces the number of test stores you need. Pick what's realistic. You'll need to block this many weeks for a clean pre-test baseline period, plus the same number of weeks for the test itself.
weeks
How sure do you need to be?
How confident do you need to be that your test results are real? Stricter means lower chance of false alarms, but also requires more stores.
You need
-
test stores
+
-
control stores
 
Looks doable -
Have stores to spare?
Fewer test stores to change is easier and safer to roll out, but the benefits shrink fast. Using 3x as many controls only reduces the test size slightly vs 2x, and the smaller your test group, the more risk that a few odd stores can impact results (a closure, a new competitor, bad data).
How accurate is this recommendation?
-
-

How long to run it

-

You'll want to choose weeks that cover similarly performing periods, or for better accuracy, apply a year-over-year seasonality adjustment. Learn more here.

How to read the numbers

They're estimates from rough descriptions, so treat them as a ballpark, not to-the-digit targets. When in doubt, plan for the higher end.

One thing to check before you trust any of this

Getting the count right doesn't make the test fair. A before/after test assumes your two groups would have moved together if you'd done nothing - the parallel-trends assumption. If your test stores were already pulling ahead, the test will show an impact that was never there.

Easy gut check: before you start, chart both groups' numbers for the last several months. The two lines should rise and dip together. If they don't, swap some stores until they do - then run your test.

The math, briefly. Test stores per group ≈ (1 + 1/k) · (zsure + zcatch)² · CV² · (1 − ρ²) · w / lift² - the standard two-group formula, with a bonus for using each store's own history. Because the average size of your stores cancels out, no real data is needed. CV = how different your stores are; ρ = how steady they are; k = control-per-test ratio; w = window factor (drifts below 1 as you add weeks). You also get credit when the test would use a big slice of everything you have. Aimed at an 80% chance of catching a real effect, two-sided. Decide the size up front and don't stop early the first time it looks like a win - and if your change causes a novelty bump, skip the first few days.
-

Finished here? Learn how to setup a diff-in-diff test.

How to set up your periods, calculate impact, and adjust for seasonality.

Diff-in-Diff Test Design →