Roblox thumbnail testing methodology

What wemeasure.And whatwe can't.

Every number on Nexframe says where it comes from. Only one of them is a CTR: yours, measured in your own ads, with its uncertainty.

0/200verified tests collected so far. No accuracy number until the rule on this page is met.

Last reviewed 9 Oct 2026 · rubric v1

  1. 5Home Readiness ScorePixel checks in your browser plus one AI reading of the imageNot CTR
  2. 4Attention mapAn open eye-tracking modelNot CTR
  3. 3Crowd preferenceVotes of adult creators in a simulated feedNot CTR
  4. 2qPTR verdictRoblox's own Thumbnail Personalization numbers, read-onlyPlanned
  5. 1Ad test calculatorYour own Roblox Ads Manager clicks and playsMeasured

Five kinds of evidence

What each number means

Five kinds of evidence, from strongest to weakest. Only the top two can say CTR, and only because they count real clicks and plays.

This score checks your thumbnail against things that are known to matter at the size Roblox shows it (subject size, text, clutter, message, rules). It is not a CTR prediction: it can't see your game, your icon, your audience or where Roblox places you. Only a live test tells you which thumbnail wins.

In plain TypeScript, no statistics library

The formulas, exactly as they run

The ad test calculator runs these in your browser, in plain TypeScript with no statistics library. Same numbers in, same result out: even the Monte Carlo draws use a fixed seed.

01

Wilson interval

A 95% range for each creative's rate (n impressions, x clicks or plays, z = 1.96).

p = x/n
center = (p + z^2/(2n)) / (1 + z^2/n)
half   = z/(1 + z^2/n) * sqrt(p(1-p)/n + z^2/(4n^2))
CI     = [center - half, center + half]
02

Two-proportion z test

Is the leader different from another creative?

pbar = (xA + xB) / (nA + nB)
z    = (pB - pA) / sqrt(pbar(1-pbar)(1/nA + 1/nB))
p    = 2 * (1 - Phi(|z|))
03

Lift with a Katz interval

The difference as people read it (+34%), with its 95% range.

RR = pB / pA
se = sqrt(1/xB - 1/nB + 1/xA - 1/nA)
CI = [RR * exp(-1.96 se), RR * exp(+1.96 se)]   shown as RR - 1
04

Newcombe interval of the difference

The absolute gap in percentage points, built from the two Wilson intervals.

d = pB - pA
L = d - sqrt((pB - lB)^2 + (uA - pA)^2)
U = d + sqrt((uB - pB)^2 + (pA - lA)^2)
05

Holm correction

Comparing the leader with k-1 creatives gives k-1 chances to get lucky; Holm raises the bar to match.

sort p(1) <= ... <= p(m)
adjusted(j) = max over i <= j of min(1, (m - i + 1) * p(i))
Winner only if every adjusted p < 0.05
06

P(best) and expected loss

Bayesian: Beta(1 + x, 1 + n - x) per creative, sampled with a fixed seed so the same numbers give the same answer.

P(i is best) = share of draws where creative i is highest
loss(i)      = E[ max_j p_j - p_i ]
draws        = 100,000 (fewer above 6 creatives)
07

How much of the difference is real

Empirical Bayes: splits the spread between creatives into real difference and noise, and shrinks small samples.

m = SUM x / SUM n
observed = n-weighted mean of (p_i - m)^2
noise    = n-weighted mean of m(1-m)/n_i
tau^2    = max(observed - noise, 0)
real     = tau^2 / (tau^2 + noise)
08

Sample ratio check

Did the creatives get very different volumes? The Ads Manager favors some creatives over time.

chi^2 = SUM (n_i - N/k)^2 / (N/k),  k - 1 degrees of freedom
warn when p < 0.001
09

Sample-size planner

Impressions per creative to see a given lift, at 95% confidence and 80% power.

p2 = p1 (1 + lift),  pbar = (p1 + p2)/2
n  = [za sqrt(2 pbar(1-pbar)) + 0.8416 sqrt(p1(1-p1) + p2(1-p2))]^2 / (p2 - p1)^2
za = 1.96 for 2 creatives, z(1 - 0.025/(k-1)) for k
10

Always-valid p-value (mSPRT)

Lets you check every day without inflating false winners.

V = 2 pbar(1-pbar),  tau = 0.2 pbar,  n = smaller arm
Lambda = sqrt(V/(V + n tau^2)) * exp(n^2 tau^2 (pB-pA)^2 / (2V(V + n tau^2)))
p = min(1, 1/Lambda)
11

Home Readiness Score

Not statistics: a checklist. 13 checks with public weights, summed and shown in steps of 5.

score = round_to_5( SUM points of the 13 checks )   max 100
band  = A 85+, B 70-84, C 50-69, D under 50

How a verdict is decided

The decision metric is plays per impression whenever the table has plays; CTR and plays per click are shown as diagnostics.

  1. 1Winner: every Holm-adjusted p below 0.05, every lift interval above zero, at least 30 plays (or clicks) per creative, and the planned sample reached or an always-valid p below 0.05.
  2. 2Clicks vs plays: the CTR leader brings significantly fewer players per impression than the plays leader.
  3. 3Practical tie: the whole lift interval sits within 10% for every creative.
  4. 4Likely, not proven: P(best) of 80% or more, but a Winner condition is missing.
  5. 5Too early to tell: everything else, with the smallest difference your volume can detect.

Real creatives, real counts

Our own ads, including what went wrong.

22 real creatives from two of our Roblox Ads campaigns. They taught us more about fooling ourselves than about thumbnails.

The same picture, two very different numbers

Ad creative #10: Lucky block monster
#10CTR 1.41%
Ad creative #11: Lucky block monster (same art as #10)
#11CTR 0.89%

Same art, same campaign. #11 only got 448 impressions, so its CTR could land almost anywhere. Ranking creatives by raw CTR at this volume ranks luck.

The click winner lost on players

Ad creative T1: Dark bedroom, REC camera
T1CTR 4.89%plays/impr. 1.69%
Ad creative T4: Before / after house
T4CTR 4.32%plays/impr. 2.26%

T1 won on clicks by 13%, yet T4 brought 34% more players per impression (95% CI +25% to +44%). Judged on CTR, we would have scaled the worse creative. That is why the calculator decides on plays.

0%of the spread between our 16 Lucky Dino creatives is real. They are statistically the same.
84%of simulations where 16 identical creatives still showed a 'significant' winner.
24%fake winners from checking every 1,000 impressions, instead of 4.5% when you look once.
Load this data in the calculator

Validation status

0 / 200

0 of 200 verified tests collected

No accuracy number until the rule below is met.

Written down before we look

  1. 1Primary metric: pairwise accuracy - in real A/B pairs with a significant winner, how often the higher-scored thumbnail was the one that won.
  2. 2Counts only: a pair enters when both versions ran in the same campaign at the same time, with at least 30 plays each and p < 0.05 on plays per impression (Holm-corrected when there were more than two).
  3. 3Prospective: the score, the image fingerprint and the version are recorded before the test result exists, so no score can be adjusted to fit a result.
  4. 4Ties count too: pairs with no significant winner measure whether the tool can say it does not know.
  5. 5Publication rule: no accuracy number before 200 verified pairs from after the freeze date, and only if the lower end of its 95% interval is 55% or more, shown next to coverage, the baselines (coin flip, brightest image) and the cases we got wrong.

Score receipts

Every full report of the thumbnail tester is recorded with an image fingerprint (SHA-256), the AI reading, the model and the rubric version, before any real test is run. The report shows a short receipt id and the version (for example Receipt 3f9a1c0b7e · rubric v1). Because the record comes first, a score can never be adjusted to fit a result, and scores from one rubric version never count toward another. The image itself is never stored.

What we can't see

Limits

Last reviewed 9 Oct 2026. Not affiliated with or endorsed by Roblox Corporation.

Every change, dated

Changelog

  1. 9 Oct 2026 · Rubric v1

    • Home Readiness Score launched: 13 checks with public weights, 15 warnings, compare mode judged in both orders.
    • Every full report recorded with an image fingerprint (SHA-256), the AI reading, the model and the rubric version.
    • Ad test calculator launched, with the formulas on this page.
    • Performance claims about clicks removed from the site, the checkout, the assistant, the Discord bot and the style names: no tool here measures that for generated thumbnails.

Questions

Common questions

Does Nexframe predict thumbnail CTR?

No. The only CTR numbers on Nexframe are the ones you measured yourself, in the ad test calculator, shown with their uncertainty. The Home Readiness Score is a checklist of what matters at Home size, and it says so next to every score. Research on predicting clicks from an image alone shows weak results, so we do not sell it.

How are thumbnail scores recorded?

Every full report is saved with a fingerprint of the image (a SHA-256 hash), the AI reading, the model name and the rubric version, before you run any real test. The report shows a short receipt id and the version. Because the record exists first, a score can never be adjusted after a test result comes in. The image itself is never stored.

When will you publish an accuracy number?

Only after 200 real A/B results collected after the score was recorded, and only if the lower end of the 95% interval is 55% or more. Today the count is 0. Until then this page shows the counter, not a number.

Where does the calculator's math run?

In your browser. The formulas on this page are exactly what it computes, written in plain TypeScript with no statistics library, and the same numbers always give the same result (the Monte Carlo draws use a fixed seed).

Check it, then test it for real

Run the free thumbnail tester before you upload, and the ad test calculator after your ads run.

Nexframe is not affiliated with, endorsed by, or sponsored by Roblox Corporation or Epic Games. Roblox and Fortnite are trademarks of their respective owners.

AboutContactRefundsTermsPrivacy(c) 2026 JAMDEV LTDAcontato@nexframe.art