Three arms, one comparison
Human+AI against AI alone, and against people working unaided, run blind inside real games rather than on a static test set.
Models are ranked against other models. Almost nothing measures whether a person actually performed better with one. Signal runs the comparison inside real gameplay: people unaided, AI alone, and people amplified by AI — and reports the lift.
A leaderboard position says a model is good at the benchmark. It says nothing about whether a person got further with it in their hands. Signal measures the difference, and the difference is the thing worth optimising.
Human+AI against AI alone, and against people working unaided, run blind inside real games rather than on a static test set.
Gemini, Claude, GPT and open weight models, scored on goal attainment, cost and efficacy by one neutral scorer.
Bayesian averaged ratings, 95% confidence intervals and Welch's t test. A live benchmark keeps us precise.
The benchmark exists to answer questions a model only leaderboard structurally cannot.
Skillprint Labs ranks frontier models from real gameplay sessions. Mood alignment is one of the boards. The benchmark is open, and it updates as sessions land.

Training data is overwhelmingly what people produced once they had finished thinking. Gameplay captures the process instead: the decisions, the recoveries and the adaptations that produced the output.
Talk to researchGameplay maps to how people actually reason and decide, not to how they describe it afterwards.
relevanceEvery session returns schema constrained output against one ontology, so it is comparable across games.
labelsBenchmarks show which systems genuinely improve people, and by how much.
liftResearch access, dataset scope, benchmark participation and governance. Tell us what you want to measure.