Signal · AI Labs

Train on how humans behave in the real world.

Train on how humans act, decide and adapt. Skillprint turns real gameplay into rights-cleared, model-grade human intelligence data: what happened, joined to how people decided, adapted, felt and performed.

Research, white papers and the live benchmark are at skillprint.co/benchmark.
The missing layer in AI training data

Game state data explains what happened. Skillprint adds context for why.

Most datasets capture outputs, pixels, physics or action traces. Skillprint adds fluid human reasoning and behavior, synchronized to the same timeline.

Today's common data

Everyone has it
TextImagesVideoCodeGame StateActions
What you receive

One session. Up to six synchronized views of the human and the world.

Each stream is keyed to the same session and timeline, and can be delivered together as model-ready bundles.

01

Visual stream

Timestamped screenshots and/or video of real human play, ordered by session and time.

what happened
02

Game-state stream

Game, level, configuration, objects, goals, points, outcomes, physics and real-time difficulty changes.

the world
03

Player state and decisions

Cognitive skill levels, mindset, experience, controller inputs, goals, strategies and outcomes.

the decision
04

Cognition, flow and mood

Assessed cognitive skills plus flow and mood signals, with a human-reviewed golden subset.

human state
05

Human ground truth

Player-reported mood and blind A/B experiment arms, including a non-AI-assisted control.

ground truth
06

Longitudinal profiles

Session- and player-level scores across 14 cognitive skills and 9 moods, tracked over time.

over time
The dataset

Each configured session is a labeled reasoning example.

Gameplay captures the process behind an outcome: the world state, the decision, the player's cognitive and emotional state, what happened next, and how that same player changes over time.

Discuss your data needs
01

Real world relevance

Gameplay maps to how people actually reason and decide under changing conditions, not to how they describe it afterward.

relevance
02

Structured labels at scale

Session data is synchronized against one timeline, making visual state, telemetry, cognition, flow, mood, controls and outcomes directly comparable.

labels
03

Collected from real play

Human gameplay comes from real players across partners’ live games rather than only hired or curated capture sessions.

live play
04

Rights and consent built in

Data streams are sourced from rights-cleared games and are player-consented and anonymized before delivery.

provenance
What it answers

Where do humans and AI perform best together?

The data carries blind A/B arms and a non-AI-assisted control, so it answers questions a model-only leaderboard structurally cannot.

01
Does this model actually improve the person using it?Measured against the same person's unaided baseline, in the same task, on the same day.
02
Which model amplifies which kind of person?Lift is not uniform. A model that helps a cautious planner may do nothing for a fast improviser.
03
Where is the human still better alone?The control arm runs with no AI in the loop, so the cases where assistance hurts are visible rather than assumed away.
04
What does the lift cost?Goal attainment scored beside token cost and latency, so efficacy is judged against what it took to get there.
What labs can do

Train and evaluate how models act, adapt and recover in dynamic environments.

Five things the same corpus supports, from fine-tuning a world model through to proving a system actually helps the person holding it.

01 · adaptation

World-model adaptation

Fine-tune on state to action to next-state pairs across multiple environments and human/agent strata.

StateActionNext state
02 · reasoning

Reasoning and planning

Learn from decisions, strategies, outcomes and cognitive-skill labels rather than action traces alone.

DecisionsStrategiesOutcomes
03 · personalization

Personalization

Study how the same environment and intervention affect people with different cognitive, mood and experience profiles.

CognitionMoodExperience
04 · evaluation

Human+AI evaluation

Use blind A/B arms, unaided controls and human-reported outcomes to measure whether systems genuinely improve performance and experience.

Blind A/BUnaided controlReported outcome
05 · recovery

Failure, recovery and adaptation

Capture unsuccessful attempts, strategy changes, retries and recovery paths so models can learn how people recognize errors and revise their approach.

Failed attemptsRetriesRecovery
Work with us

Tell us what you want to measure.

Three ways to work with us: license the gameplay data we already hold, commission collection built to your spec, or benchmark your models against real human play.

01
LicenseExisting human intelligence gameplay datasets.
02
CommissionRights-cleared collection around a target environment or action vocabulary.
03
BenchmarkBuild a human+AI benchmark around the outcome your team cares about.
The benchmark

See how models score against real human play.

A leaderboard position says a model is good at the benchmark. It says nothing about whether a person got further with it in their hands. Signal scores people unaided, AI alone, and people amplified by AI inside real games. The live leaderboard sits on the benchmark page.