Wilson interval
A 95% range for each creative's rate (n impressions, x clicks or plays, z = 1.96).
p = x/n center = (p + z^2/(2n)) / (1 + z^2/n) half = z/(1 + z^2/n) * sqrt(p(1-p)/n + z^2/(4n^2)) CI = [center - half, center + half]
What wemeasure.And whatwe can't.
Every number on Nexframe says where it comes from. Only one of them is a CTR: yours, measured in your own ads, with its uncertainty.
Last reviewed 9 Oct 2026 · rubric v1
Five kinds of evidence
Five kinds of evidence, from strongest to weakest. Only the top two can say CTR, and only because they count real clicks and plays.
This score checks your thumbnail against things that are known to matter at the size Roblox shows it (subject size, text, clutter, message, rules). It is not a CTR prediction: it can't see your game, your icon, your audience or where Roblox places you. Only a live test tells you which thumbnail wins.
In plain TypeScript, no statistics library
The ad test calculator runs these in your browser, in plain TypeScript with no statistics library. Same numbers in, same result out: even the Monte Carlo draws use a fixed seed.
A 95% range for each creative's rate (n impressions, x clicks or plays, z = 1.96).
p = x/n center = (p + z^2/(2n)) / (1 + z^2/n) half = z/(1 + z^2/n) * sqrt(p(1-p)/n + z^2/(4n^2)) CI = [center - half, center + half]
Is the leader different from another creative?
pbar = (xA + xB) / (nA + nB) z = (pB - pA) / sqrt(pbar(1-pbar)(1/nA + 1/nB)) p = 2 * (1 - Phi(|z|))
The difference as people read it (+34%), with its 95% range.
RR = pB / pA se = sqrt(1/xB - 1/nB + 1/xA - 1/nA) CI = [RR * exp(-1.96 se), RR * exp(+1.96 se)] shown as RR - 1
The absolute gap in percentage points, built from the two Wilson intervals.
d = pB - pA L = d - sqrt((pB - lB)^2 + (uA - pA)^2) U = d + sqrt((uB - pB)^2 + (pA - lA)^2)
Comparing the leader with k-1 creatives gives k-1 chances to get lucky; Holm raises the bar to match.
sort p(1) <= ... <= p(m) adjusted(j) = max over i <= j of min(1, (m - i + 1) * p(i)) Winner only if every adjusted p < 0.05
Bayesian: Beta(1 + x, 1 + n - x) per creative, sampled with a fixed seed so the same numbers give the same answer.
P(i is best) = share of draws where creative i is highest loss(i) = E[ max_j p_j - p_i ] draws = 100,000 (fewer above 6 creatives)
Empirical Bayes: splits the spread between creatives into real difference and noise, and shrinks small samples.
m = SUM x / SUM n observed = n-weighted mean of (p_i - m)^2 noise = n-weighted mean of m(1-m)/n_i tau^2 = max(observed - noise, 0) real = tau^2 / (tau^2 + noise)
Did the creatives get very different volumes? The Ads Manager favors some creatives over time.
chi^2 = SUM (n_i - N/k)^2 / (N/k), k - 1 degrees of freedom warn when p < 0.001
Impressions per creative to see a given lift, at 95% confidence and 80% power.
p2 = p1 (1 + lift), pbar = (p1 + p2)/2 n = [za sqrt(2 pbar(1-pbar)) + 0.8416 sqrt(p1(1-p1) + p2(1-p2))]^2 / (p2 - p1)^2 za = 1.96 for 2 creatives, z(1 - 0.025/(k-1)) for k
Lets you check every day without inflating false winners.
V = 2 pbar(1-pbar), tau = 0.2 pbar, n = smaller arm Lambda = sqrt(V/(V + n tau^2)) * exp(n^2 tau^2 (pB-pA)^2 / (2V(V + n tau^2))) p = min(1, 1/Lambda)
Not statistics: a checklist. 13 checks with public weights, summed and shown in steps of 5.
score = round_to_5( SUM points of the 13 checks ) max 100 band = A 85+, B 70-84, C 50-69, D under 50
The decision metric is plays per impression whenever the table has plays; CTR and plays per click are shown as diagnostics.
Real creatives, real counts
22 real creatives from two of our Roblox Ads campaigns. They taught us more about fooling ourselves than about thumbnails.


Same art, same campaign. #11 only got 448 impressions, so its CTR could land almost anywhere. Ranking creatives by raw CTR at this volume ranks luck.


T1 won on clicks by 13%, yet T4 brought 34% more players per impression (95% CI +25% to +44%). Judged on CTR, we would have scaled the worse creative. That is why the calculator decides on plays.
Validation status
0 / 200
0 of 200 verified tests collected
No accuracy number until the rule below is met.
Every full report of the thumbnail tester is recorded with an image fingerprint (SHA-256), the AI reading, the model and the rubric version, before any real test is run. The report shows a short receipt id and the version (for example Receipt 3f9a1c0b7e · rubric v1). Because the record comes first, a score can never be adjusted to fit a result, and scores from one rubric version never count toward another. The image itself is never stored.
What we can't see
Last reviewed 9 Oct 2026. Not affiliated with or endorsed by Roblox Corporation.
Every change, dated
9 Oct 2026 · Rubric v1
Questions
No. The only CTR numbers on Nexframe are the ones you measured yourself, in the ad test calculator, shown with their uncertainty. The Home Readiness Score is a checklist of what matters at Home size, and it says so next to every score. Research on predicting clicks from an image alone shows weak results, so we do not sell it.
Every full report is saved with a fingerprint of the image (a SHA-256 hash), the AI reading, the model name and the rubric version, before you run any real test. The report shows a short receipt id and the version. Because the record exists first, a score can never be adjusted after a test result comes in. The image itself is never stored.
Only after 200 real A/B results collected after the score was recorded, and only if the lower end of the 95% interval is 55% or more. Today the count is 0. Until then this page shows the counter, not a number.
In your browser. The formulas on this page are exactly what it computes, written in plain TypeScript with no statistics library, and the same numbers always give the same result (the Monte Carlo draws use a fixed seed).
Check it, then test it for real
Run the free thumbnail tester before you upload, and the ad test calculator after your ads run.