Readiness check/Methodology
Methodology

How the Step 2 CK score predictor works

The only Step 2 CK score predictor that shows its work. Every number on this page is the live output of the same model that runs in the app and on the readiness check, the same code, no marketing math. Here is exactly how your practice tests become a projected range.

01 What it predicts

A projected Step 2 CK score, as a range.

You enter your NBME and UWorld Self-Assessment scores. The model returns a projected Step 2 CK score as a range, not a single number and not a guarantee. It is free, needs no sign-up, and runs entirely in your browser. A range is the honest output: any single practice test carries roughly 8 to 9 points of form noise, so a lone point estimate would imply a precision the data does not support.

02 The model, step by step

From your scores to a range, in six steps.

We will follow one real set of eight practice tests through all six steps. Every figure below, including the card, is computed on this page by calling the model directly, so what you read is what a student with these scores actually sees.

Projected Step 2 CK
247  236 to 258
anchor 235 · about the 39th percentile vs US first-takers
Anchor235
Range236 to 258
Percentile39th
Tests8
a

Your inputs

Two kinds of practice test, each a three-digit score: NBME forms 9 to 16, and UWorld Self-Assessments 1 to 3. Optionally, how many days before your exam you took each one. Dates make the projection sharper, but the model works without them.

Worked example, undated: NBME 12 = 218, NBME 13 = 235, NBME 14 = 235, UWSA 1 = 232, UWSA 3 = 204, NBME 15 = 240, UWSA 2 = 214, NBME 16 = 245.

b

The anchor: a recency-weighted average

Your scores collapse into one number, the anchor. Not a flat average: each test is weighted by how recent it is, because a test from last week predicts exam day better than one from two months ago. The weight decays exponentially with how many days before the exam you took it.

weight = exp( -daysOut / tau ),   tau = 30

That is a 21 day half-life: a test taken 21 days earlier counts about half as much as your freshest one. The anchor is the weighted average of every score under those weights.

c

No dates? A hybrid recency fallback

When you skip the dates, the model still needs an order. It reads the order you entered your tests, so the one you added last is treated as freshest, which correctly handles adding an older form out of sequence. But there is a trap: 62% of students list a UWSA last, yet a UWSA is actually their most recent test only 36% of the time. So each undated UWSA is pinned to the median NBME recency. A UWSA you happened to type last cannot pose as your freshest test by listing convention alone. Your real NBME climb carries the number.

Anchor weights
How much each test counts, freshest first
NBME 16 (245) is the latest real form, so it carries 40.0%. The three UWSAs sit together at the median NBME recency, so none of them, including the 214, gets to masquerade as freshest.
NBMEUWSA (pinned to median NBME recency)
NBME 16245latest real form
freshest40.0%
NBME 15240
20.5%
NBME 14235
7.6%
UWSA 1232pinned to median
7.6%
UWSA 3204pinned to median
7.6%
UWSA 2214pinned to median
7.6%
NBME 13235
5.4%
NBME 12218
3.9%
The number tracks the NBME line, not the typing order. The latest real NBME is 245 and the NBME trend is 240 then 245, so the anchor lands at 235, and the projection just above the freshest form. A trailing UWSA can still lower the anchor if it genuinely is one of the recent tests, it just no longer wins freshest by convention.
d

Shrink toward the mean, then add the measured gain

The anchor is not the projection. Two corrections turn it into a center.

center = 250 + 0.79 x ( anchor - 250 ) + 9

The 0.79 multiplier is a Thorndike Case II range-restriction correction. Practice tests spread wider than the real exam does, so raw highs and lows are pulled back toward the population mean of 250. An anchor below the mean is nudged up, a high anchor is nudged down. The +9 is the level offset: across the validation cohort, students outscored their recency-decay anchor by about 9 points on the real exam, on average.

Worked example: anchor 235 shrinks to 250 + 0.79 x (235 - 250) = 238.2, plus 9 gives 247.2, rounded to a center of 247.

e

The band width: tight when close, wide when far

The half-width of the range depends on how near your freshest test is to exam day. The closer you are, the more your practice scores have converged, so the tighter the honest interval.

Within 15 days of examplus or minus 7Tight
15 to 30 days outplus or minus 11Moderate
31 to 56 days outplus or minus 15Wide
57 or more days outplus or minus 18Very wide
No exam date givenplus or minus 11Rough

A guardrail widens the band whenever the anchor sits at or below the 218 pass line, where the low tail is least tested. There the band opens to at least 15 points and the read is flagged high uncertainty.

Worked example: the freshest test is undated, so the band is plus or minus 11, giving the 236 to 258 range below. Add your exam date and it tightens.

218 pass
avg 250
247
205236 to 258275
f

Where that lands nationally

The center is turned into a percentile by interpolating the official USMLE Score Interpretation Guidelines norm table (LCME first-takers, N about 67,934). It answers a different question than pass or fail: of every US first-taker, how many scored at or below this.

2101st
2204th
23010th
24024th
25047th
26074th
27094th

Worked example: a center of 247 interpolates to about the 39th percentile, between the 34 at 245 and the 47 at 250.

02b Form scale quirks

Three forms print off the common scale, so they get translated.

Fit on 160 students and 1,043 dated forms with person and timing controls: old NBME 9 prints about 5.5 low (we credit it back), UWSA 2 prints about 4.8 hot (translated down), and UWSA 3 about 6.3 low on thinner data. Modern NBMEs, forms 10 through 16, are equated by NBME itself and enter at face value: a printed NBME score is never counted down here. NBME 16 does print a few points high, but students who print high on 16 really do score high, so translating it away would under-project; it stays at face value. And no, forms do not carry hidden weights beyond their dates: we tested that too, and recency carries all of it.

03 Calibration and its limits

How it is calibrated, and where it stops.

Those constants are not guesses. They were fit against about 223 real paired outcomes: a student's pre-exam practice scores, then their actual Step 2 CK, gathered from public r/step2 score-release threads. Three choices we made on purpose, in the name of not overselling.

Held wide on purpose

The sample is high-performer selected. People who post score releases skew high, the actuals average well above the national mean, and the low tail is thin. So the band is deliberately held wide rather than tightened to fit this rosy sample. A narrower band would look more confident and be less honest.

Ranges, not fake precision

The interval targets roughly the middle 60% of outcomes, not a 95% guarantee. The slope is held at 0.79 rather than the flatter value this restricted sample suggests, because a flatter slope would badly over-project weak students, the ones who most need an honest low read.

Not a guarantee

It is a projection from practice tests. Any individual can land outside the range. Treat it as a compass heading, not a diagnosis.

It will sharpen as real in-app outcomes accrue, students who log a practice score and later log their actual Step 2, a sample that is not high-performer truncated. When that set is large enough to disagree with this one, the constants move, on data, not on vibes.

04 The blind test

Fit on the past, scored on the future.

Most score converters have never been tested on people they were not fit on. This one has. Before the August 2026 score release, the model's projections were locked for 32 studentsfrom the newest release thread, none of whom were in the calibration set. When their real scores dropped: average miss 4.1 points, average signed error +0.97 (essentially unbiased), 88% landed within 7 points and 94% within 8. Those numbers are what the band width claims they should be, which is the whole point of publishing them.

05 The beat curve

How far past their last practice form do people actually land?

Across 150 reports with an exact outcome and a dated final form, actual score minus the freshest practice form has a median of +9. Half of students land between +1 and +14. The full range runs -20 to +35, and about 1 in 6 land below their freshest form, which is why a projection that just adds a fixed bonus to your last NBME is flattering rather than honest. The model earns the middle of this curve through recency weighting and calibration, and the band width covers its tails.

p10-3
p25+1
median+9
p75+14
p90+18
06 Closing the loop

Predictions locked before the exam, paired with real scores after.

Since July 2026 the app freezes every student's projection on their exam day, before any result exists. When they later report their real score, one tap in the app or from an email, the pair becomes a true blind data point: what the model said then, against what actually happened. These app-native pairs have none of the good-news bias of public score threads, and as they accumulate they drive the next refit. Scores are anonymous, linked only to the account, and only ever used in aggregate.

Help make it sharper. When you get your real Step 2 score, log it, every real outcome pulls these numbers closer to the truth.