Complete collection
Grade and value distributions
Individual continuation grades
All normalized rubric scores
Credit composition
Exact zero, partial, and full credit
Prefix mean grades
One empirical value per collected prefix
Value across reasoning depth
Mean and interquartile range within normalized-depth deciles
Trajectory
-
-
Problem
Mean grade by prefix
Empirical quartiles across 32 continuations; diamond marks the base endpoint grade
Base model output
Reasoning trace
Selected prefix
Selected prefix
-
Continuation grade distribution
32 samplesRun semantics
How to read the collection
Prefix value
Every partial point is the arithmetic mean of 32 normalized rubric grades from fresh continuations of that exact response prefix. It is not a binary success probability.
Trajectory endpoint
The diamond is the original base response's continuous rubric grade. The downstream PRM dataset separately thresholds eligible endpoints at 0.5 for a hard yes/no label.
Empty prefix
Each problem contributes two selected base traces, so its empty prefix appears once per trace in the raw branch set. The 103,518 rows correspond to 99,354 unique semantic problem-prefix inputs.
PRM filtering
The downstream training build removes duplicate empty prefixes, invalid endpoints, and overlength prompts while preserving provenance counts.