Skip to main content
Forenly AI Lab · Reproduction 001 · Baseline

We ran a published result again and got a different number

Same checkpoint, 500 episodes, one fixed seed.

Running success rate, recomputed every 20 episodes
60%65% 70%75% 80% 100200 300400 500 episodes evaluated published 65.4% After 20 episodes: 80.0% After 40 episodes: 67.5% After 60 episodes: 70.0% After 80 episodes: 67.5% After 100 episodes: 67.0% After 120 episodes: 65.8% After 140 episodes: 66.4% After 160 episodes: 68.1% After 180 episodes: 68.3% After 200 episodes: 67.5% After 220 episodes: 67.3% After 240 episodes: 67.5% After 260 episodes: 67.7% After 280 episodes: 66.8% After 300 episodes: 66.7% After 320 episodes: 65.0% After 340 episodes: 64.7% After 360 episodes: 64.7% After 380 episodes: 63.4% After 400 episodes: 63.2% After 420 episodes: 61.9% After 440 episodes: 62.3% After 460 episodes: 61.7% After 480 episodes: 61.9% After 500 episodes: 62.0% 62.0%

The same run reads 80.0% at 20 episodes and 62.0% at 500. Stopping early would have produced a different headline.

62.0%
Measured · 310 of 500 episodes
65.4%
Published on the model card
1.6SE
The gap · not significant at 95%

What we ran

Task
PushT · gym-pusht
Episodes
500
Seed
1000
Stack
LeRobot 0.3.2
Machines
Xeon 8 vCPU · RTX 4090

Getting it to run was the first finding

0.1.0 0.3.2 0.3.3 0.4.x 0.5.x 0.6.1

The published checkpoint predates the processor pipeline introduced in 0.3.3, and the official migration script crashes. 0.3.2 is the last version that loads it.

What could explain the gap

  1. 1
    Most likely

    Sampling noise

    The gap is 1.6 standard errors. A difference this size arises by chance about one run in eight.

  2. 2
    Possible

    A different draw of episodes

    The model card states no seed. PushT randomises every reset, so another seed is another 500 tasks.

  3. 3
    Least evidence

    The software underneath

    We ran on library versions that postdate the claim. The hypothesis we can say least about — so it is last.

  4. Ruled out

    Hardware

    Both machines shared the seed and agree to within one episode.

The experiment that separates them is the same evaluation across several seeds — left open so the next number comes from someone else’s machine.

The honest part

This is not a claim that the model card is wrong. 1.6 standard errors is neither a confirmation nor a refutation. Separating a real gap from noise would take roughly 1,500 episodes.

We made this study’s own mistake, on ourselves. An interim note called an early 80-episode reading “converging.” It was an optimistic window; the correction stayed in the repo.

PushT is not humanoid work, and we do not present it as any. It is the first step of a chain — reproduce → measure → benchmark → humanoid → real world.

Everything is in the open

All 500 per-episode records, the environment fingerprint and the version bisection — every number here can be recomputed from them.

Forenly-AI-Lab/reproduction-001

The Skill Layer for humanoid robots

Forenly AI turns real scenes into simulation where humanoids learn their skills — then transfers them onto the machine.