What actually predicts your Step 2 score
269 r/step2 score reports, one calibrated model, and its out-of-sample record, updated every Monday. Written by a medical student who believed the folklore until he checked it.
I parsed 269 r/step2 score reports into practice-form histories and real Step 2 CK scores. For the 239 with a three-digit NBME or UWSA on file, the real score beat the practice average 92% of the time, by a median of about 11 points, but the size of that climb depends on where you start: about +20 for averages under 235, about +2 at 265 and up. The most recent form predicts best (half-life about three weeks). Three forms read off scale after controlling for who took them and when: NBME 9 about 6 points low, UWSA 2 about 5 high, UWSA 3 about 8 low. A shrinkage-linear model fit to this data lands within 7 points for 70% of students in-sample; on 40 later posters it had never seen, 80% within 7. On the first 7 users who entered scores in the app it is noticeably worse, and I show that too.
The app carries this projection under a readiness gauge, keeps the track record on the same screen, and lets you add a form in ten seconds. Free to start.
1. Why I did this
Every score release thread has the same comment under it: "add 10 to your NBME average." Sometimes it's 8, sometimes 15. UWSA 2 is "inflated." NBME 9 is "useless." I believed all of these during dedicated and repeated some of them. After my own exam I wanted to know which were true, and the data to check them is sitting in public on Reddit: people post their forms, then come back weeks later and post the real number.
2. Methods
Data
Score-report posts from r/step2 (the aggregate scrape plus the 07-29, 08-05 and 08-19 threads), parsed into one row per assessment: form, printed score, days before the exam when available, and the actual three-digit Step 2 CK result. 269 students, 3,737 assessments. Usernames were dropped at parse time; records carry only an id, the forms, the days and the outcome. Free 120 percentages were kept but do not enter the model.
Inclusion
A student enters the fit when at least one NBME or UWSA three-digit score and a real Step 2 score are present (n=239). Implausible actuals (outside 155 to 300) are dropped. Scores recorded fewer than 13 days after the exam date are excluded on principle, since USMLE releases reports on Wednesdays about two to four weeks after the test and a "score" before then cannot be real.
Model
Each student's forms are collapsed into one anchor: a recency-weighted mean with weight exp(minus days / 30), after per-form level corrections. The anchor is then shrunk toward the population mean:
The slope under one is regression to the mean: a low practice average usually contains a bad day, and bad days don't repeat on purpose. The +9 is the average climb at the middle of the distribution. The slope was deliberately held above the best-fit line through this sample (which lands near 0.55), because the sample skews high and I would rather under-promise a 230 than over-promise it. Band half-widths (7, 11, 15, 18 points) follow how far out the freshest form was taken.
Validation
In-sample accuracy by leave-one-out on the 239. Then the constants were frozen, and every later outcome is judged against the frozen model: 40 r/step2 posters from threads parsed after the freeze, and the users who enter real scores in the app. That second set updates weekly and is shown below and inside the app.
3. Results
3.1 You will probably score above your practice average, by an amount that depends on where you start
92% of students beat their recency-weighted practice average, median +11. The fixed number is the part that's wrong:
| Practice average | Students | Climbed by |
|---|---|---|
| under 235 | 35 | +20 |
| 235 to 244 | 67 | +12 |
| 245 to 254 | 75 | +11 |
| 255 to 264 | 38 | +5 |
| 265 and up | 24 | +2 |
If you're averaging 230, "add 10" undersells you. If you're averaging 262, it oversells you by a lot, and that is the group most likely to set a 270 target off it.
3.2 Your most recent form matters far more than your first one
I tried several ways of collapsing forms into one number: plain mean, last form only, best form, and a recency-weighted mean. The weighted version predicted best, and the weight that worked has a half-life of about three weeks.
This also explains the belief that old forms "deflate." They do read lower, but mostly because of when people take them, not which form it is. Once you account for the student and for timing, NBME 10 through 15 collapse to within a point or two of each other.
3.3 Three forms genuinely read off scale, and they're not the ones people complain about
| Form | Reads | Confidence |
|---|---|---|
| NBME 9 | about 6 points low | solid, 71 students |
| NBME 10 to 16 | on scale | treated as equated, 788 sittings |
| UWSA 1 | on scale | 69 students |
| UWSA 2 | about 5 points high | solid, 93 students |
| UWSA 3 | about 8 points low | thin, 22 students |
So "UWSA 2 is inflated" survives, at about five points, not the fifteen people throw around. "NBME 9 is useless" doesn't: it's a fine form that prints about six under the others. The one nobody warned me about is UWSA 3, which reads low by a lot, though only 22 people had taken it, so hold that one loosely.
3.4 How wrong the model is, in sample
Within 7 points for 70%, within 10 for 85%, average miss 5.8. The +10 rule on the same people: within 7 for 62%, average miss 6.8. Using your practice average with no adjustment at all: average miss almost 12.
3.5 Out of sample: students the model had never seen
This is the section a marketer would cut. After the constants were frozen, two groups of outcomes arrived: r/step2 posters from later threads, and students who entered a real score inside the app. Neither touched the fit.
On the Reddit posters the frozen model did about as well as in-sample, which is the result you hope for. On the app users it did worse: 7 students so far, 4 inside their range, one of them a 40-point miss from a single stale form. Seven is not enough to conclude anything. I am posting it anyway, and it updates every Monday (numbers above as of Aug 24, 2026), inside the app under the readiness gauge and on the methodology page. If the app population keeps missing worse than the forum population, the model gets refit for that population, not the other way round.
4. Try it
Your forms, the model's projection
Form offsets applied before the anchor: NBME 9 +5.5, UWSA 2 -4.8, UWSA 3 +6.3. Recency half-life about 21 days. Projection = 250 + 0.79 x (anchor minus 250) + 9.
5. Limitations
- Self-selection. People who did well post more. The fit set averages 258 against a national mean in the high 240s, and the lowest real score in it is 216. The "+20 under 235" figure rests on 35 students.
- Recall. Days-before-exam are what people remembered when they posted; about a third of forms had no usable date and were imputed from form order.
- Small out-of-sample n. 47 students, 7 of them app users. The in-app population reaches lower than Reddit does and may need its own constants.
- One predictor. The model uses forms only. Cards answered, accuracy by system and study pacing are measured in the app but are not in the projection until there are enough real outcomes to fit them honestly.
6. What I'd actually do with this
- Take a form inside your last two weeks. The projection is only as good as your freshest data point, and the 40-point miss above came from a student with exactly one form entered.
- Don't average your forms and don't let a week-one NBME drag the number around.
- If your average is under 240, the folklore is too pessimistic about you. If it's over 260, it's too optimistic. Plan off the band table, not off "+10."
- Read NBME 9 about six points up and UWSA 2 about five points down. Leave the rest alone.
7. Data and code
The parsed, de-identified forum set (record id, form, printed score, days before exam, outcome) and the fitting scripts live in the Step Gunner repository; the out-of-sample record is served at a public endpoint and recomputed weekly. The model constants on this page are imported from the same module the calculator runs, so the text cannot drift from the code. If you post your own score after exam day, it goes into the next fit, and the next version of this page gets a bigger out-of-sample number.
References
- USMLE. Step 2 CK Score Interpretation Guidelines (percentile norms, LCME first-takers, July 2022 to June 2025).
- NBME. Comprehensive Clinical Science Self-Assessment: score interpretation and equating notes.
- Thorndike RL. Personnel Selection: Test and Measurement Techniques. Wiley, 1949 (range-restriction correction, Case II).
- r/step2 score-report threads, 2025 to 2026, parsed with usernames removed at ingest.