With 2 variants running simultaneously, each result needs to reach 97.5% confidence to count as a win — not 95%. The more variants you test at once, the higher that bar gets.
Confidence over time
Traffic sensitivity
Confidence over time at different visitor volumes — because traffic decides how fast evidence piles up
Visitors needed by power
Based on your expected uplift assumption — larger real effect = fewer visitors needed
Time to power target
Assuming your expected uplift is correct — when the test should have enough visitors