Simbox is built on field experiments with real workers doing real tasks. Here's what they found, what we built because of it — and what's still unproven.
more likely to be wrong when using AI on a task just outside its limits — 758 consultants, Harvard/BCG
slower with AI while believing they were 20% faster — experienced developers, METR RCT
correlation between using AI a lot and output quality. Skilled behaviors: 0.26 — CollabSkill
Fig. 01 — Three numbers this product is built on.
Inside the model's capability, AI users produced 40% higher quality. On a task designed just outside it, they were 19 points more likely to be wrong than people working alone. Dell'Acqua et al., Harvard/BCG
People routinely accept incorrect output, and more fluent AI increases uncritical acceptance. Recruiters given better AI performed worse — they "fell asleep at the wheel." Microsoft synthesis · field study
Experienced developers were 19% slower with AI while believing they were 20% faster. Self-reported AI skill and measured AI skill are essentially uncorrelated. METR RCT · correlation study
Scoring real worker–AI sessions, skilled behaviors predicted deliverable quality; raw message volume predicted almost nothing. CollabSkill 2026
Politeness, personas, and threats produce inconsistent, question-specific effects. What transfers: legible requirements, supplied context, systematic iteration. Wharton · expert–novice studies
Structured post-work reviews improve performance by roughly 25% — one of the most reliable effects in training research. The spacing effect, which brings skills back before they fade, is among the most replicated results in learning science. meta-analysis · spacing
AI assistance lifted novice call-center agents 34%, with almost no effect for the most skilled. The biggest writing gains went to the weakest writers. NBER w31161 · Noy & Zhang, Science
The skill is behavioral: what you hand over, what you check, what you fix before sending. Only watching real work — under time pressure, with real files — captures it. And behavior over hours of recorded work is categorically harder to fake than a 30-minute test you can take with AI open in the next tab. Scenario details — names, numbers, context — regenerate for every attempt while the rubric stays fixed, so there's nothing to copy, share, or find online.
| Can it measure… | Quiz | Self-report | Course certificate | Work simulation |
|---|---|---|---|---|
| Knowing the facts | ||||
| Real working habits | ||||
| Improvement over time | ||||
| Judgment under pressure | ||||
| Hard to game | ||||
| On-the-job performance |
Fig. 02 — The last row is honest: see below.
A credential is worth what its maker refuses to claim. Four things we won't overstate:
Realistic simulation with feedback is the best-supported bet in training research, and it's how surgeons and pilots learn. But no one, including us, has yet shown that an AI-skills score predicts job performance. We run before/after studies with partner teams, and the marketing never outruns the data.
Measured effects of gamification on actual competence are small and fragile. That's why our game layer only schedules practice — streaks and points can never buy a score, a level, or a credential.
The METR authors flag their own result as historical: models improve, and yesterday's blind spots close. Every planted error is re-validated against current models, and skills marked "needs practice" expect re-demonstration.
Effectiveness is task- and domain-dependent; averaging it into one number is statistically indefensible. We report eight skills, each with its own evidence, and refuse to collapse them.
The fastest way to judge the method is to experience it.
See the demo