Why this method works

"Just take a course." The evidence disagrees.

Simbox is built on field experiments with real workers doing real tasks. Here's what they found, what we built because of it — and what's still unproven.

19 pts

more likely to be wrong when using AI on a task just outside its limits — 758 consultants, Harvard/BCG

19%

slower with AI while believing they were 20% faster — experienced developers, METR RCT

0.05

correlation between using AI a lot and output quality. Skilled behaviors: 0.26 — CollabSkill

Fig. 01 — Three numbers this product is built on.


The findings

What the research says — and what we built because of it.

+40% / −19758 consultants

The biggest skill is knowing AI's limits.

Inside the model's capability, AI users produced 40% higher quality. On a task designed just outside it, they were 19 points more likely to be wrong than people working alone. Dell'Acqua et al., Harvard/BCG

So SimboxEvery scenario plants errors just outside AI's limits. Catching them is the first skill we grade.
Better AI,
worse humansAppropriate-reliance research

Over-trusting is the default, and it grows with fluency.

People routinely accept incorrect output, and more fluent AI increases uncritical acceptance. Recruiters given better AI performed worse — they "fell asleep at the wheel." Microsoft synthesis · field study

So SimboxVerification is the gate skill. Level 3 and above require it by name.
19% ≠ 20%Randomized trial

People can't self-assess this skill.

Experienced developers were 19% slower with AI while believing they were 20% faster. Self-reported AI skill and measured AI skill are essentially uncorrelated. METR RCT · correlation study

So SimboxWe never ask. We watch — and every session shows your prediction against what actually happened.
r = 0.26
vs 0.05Real worker sessions

How you work matters. How much doesn't.

Scoring real worker–AI sessions, skilled behaviors predicted deliverable quality; raw message volume predicted almost nothing. CollabSkill 2026

So SimboxScores move on decisions, never on activity. Points and streaks can't touch a skill score.
Magic words
don't transferPrompting studies

Clear requirements beat incantations.

Politeness, personas, and threats produce inconsistent, question-specific effects. What transfers: legible requirements, supplied context, systematic iteration. Wharton · expert–novice studies

So SimboxBriefing is graded on the context and requirements you give — not the phrasing tricks you know.
d ≈ 0.7961-study meta-analysis

Reviewing your own work is the strongest lever.

Structured post-work reviews improve performance by roughly 25% — one of the most reliable effects in training research. The spacing effect, which brings skills back before they fade, is among the most replicated results in learning science. meta-analysis · spacing

So SimboxEvery session ends at the review, quoting your own work. Skills fade without practice, and the profile says so.
+34%Novice workers

The upside is biggest for the least experienced.

AI assistance lifted novice call-center agents 34%, with almost no effect for the most skilled. The biggest writing gains went to the weakest writers. NBER w31161 · Noy & Zhang, Science

So SimboxPaths start from zero and gate on demonstrated skill — the people with the most to gain get a real ladder.

Measurement

A quiz can't see the thing that matters.

The skill is behavioral: what you hand over, what you check, what you fix before sending. Only watching real work — under time pressure, with real files — captures it. And behavior over hours of recorded work is categorically harder to fake than a 30-minute test you can take with AI open in the next tab. Scenario details — names, numbers, context — regenerate for every attempt while the rubric stays fixed, so there's nothing to copy, share, or find online.

Fig. 02 — The last row is honest: see below.


The gaps

What the evidence doesn't say. Yet.

A credential is worth what its maker refuses to claim. Four things we won't overstate:

Transfer to the job isn't proven — for anyone.

Realistic simulation with feedback is the best-supported bet in training research, and it's how surgeons and pilots learn. But no one, including us, has yet shown that an AI-skills score predicts job performance. We run before/after studies with partner teams, and the marketing never outruns the data.

Points don't build skill.

Measured effects of gamification on actual competence are small and fragile. That's why our game layer only schedules practice — streaks and points can never buy a score, a level, or a credential.

The findings age.

The METR authors flag their own result as historical: models improve, and yesterday's blind spots close. Every planted error is re-validated against current models, and skills marked "needs practice" expect re-demonstration.

There is no single "AI score."

Effectiveness is task- and domain-dependent; averaging it into one number is statistically indefensible. We report eight skills, each with its own evidence, and refuse to collapse them.

Field experiments · behavioral measurement · honest gaps

Built on evidence. Priced in doubt.

The fastest way to judge the method is to experience it.

See the demo