WWorldModel GymResearch Benchmark
Research benchmark platform

Benchmark world-model agents in one public surface.

WorldModel Gym turns environments, uploads, traces, and leaderboards into a research instrument you can read at a glance — and share with one clean URL.

Sparse-reward environments with reproducible seeds
Public tracks across test, train, and continual
Browser upload alongside API and CLI submission
LeaderboardTest track
Live
1

mu-planner

GridWorld-9

94%
2

dreamer-v3

GridWorld-9

88%
3

mcts-base

GridWorld-9

81%
4

random+

GridWorld-9

62%

Runs

0

Agents

0

Avg cost

0.0k

0+Benchmark tasks
0Public runs
0Tracks
0%Best success rate

Benchmark surfaces

Rigorous evaluation, not cherry-picked demo polish.

Task framing, run uploads, leaderboard slices, and trace inspection move together as one public story.

Task library

Document environments with clear defaults, precise constraints, and readable benchmark framing.

Browse tasks

Live leaderboards

Compare planning quality, return, and cost in one public surface instead of scattered artifacts.

Open leaderboards

Upload studio

Create a run, attach metrics and traces, and publish from the browser — CLI and API still available.

Publish a run

Product workflow

One workflow: create, evaluate, upload, compare.

Every stage feeds the next, so a research idea becomes a public, reproducible benchmark without leaving the product.

Stage 01 / 04

Workflow Create

Design a benchmark brief that reads clearly from the first click.

Shape sparse-reward tasks, defaults, and success criteria before you ever touch a leaderboard. The fastest path from a research idea to a benchmark someone else can immediately understand.

Prompt

Frame a partially observable benchmark with delayed reward, reproducible seeds, and a planning budget that matches the story you want the leaderboard to tell.

Pick an environment

Choose a task with explicit constraints and failure modes.

Lock the defaults

Set seeds and budgets that make the benchmark reproducible.

Carry it forward

Move the task straight into evaluation and upload.

Task defaultsObservation modeReward design

Live benchmark product

Publish planning research without the visual clutter.

Ship new runs, compare them publicly, and reuse the same surface in your README, interviews, portfolio, or research demo.

FastAPI + Postgres + S3Next.js App RouterReal runs in production

Imagination rollout

Every ranked run traces back to a planner search over an environment model — visible, not hidden behind a screenshot.