The VibeiDE Office Benchmark
Give an AI
150 boxes.
Ask for this.
A round head. A curved chair. A tiny desk lamp. How close can a language model get using only rectangular blocks?
- boxes maximum
- 150
- models
- 6
- attempt per round
- 1
Quality, tokens and time
What did each result take?
Compare score, reported input plus output tokens, and elapsed runtime for the selected round. Reasoning tokens are not added again. Missing measurements are left out, not counted as zero.
View token and runtime measurements
The answers
The difference is in the details.
Every model's result, ranked by score. One attempt at medium effort, using the same frozen brief.
Every sculpture uses the same light, camera and scale. Select any result to compare it with the target, or switch to Difference to inspect the mismatches.
Notes
Measured results
Scores over time
Starting baseline preserves the original first-round results and measurement dates. The next point is a fresh measured round, not a result from a different week.
View the recorded scores
| Round | Claude Opus 5 / Claude Opus 5.5 | Claude Sonnet 5 / Claude Sonnet 5.5 | Claude Fable | GPT-6.1 Sol | GPT-6 Astra | GPT-6 Luna |
|---|---|---|---|---|---|---|
| 2026-W41 | 78.5 (Claude Opus 5.5) | 48.1 (Claude Sonnet 5.5) | 74.6 (Claude Fable) | 74.7 (GPT-6.1 Sol) | 73.5 (GPT-6 Astra) | 30.9 (GPT-6 Luna) |
| Starting baseline | 75.5 (Claude Opus 5) | 34.3 (Claude Sonnet 5) | 71.2 (Claude Fable) | 70.6 (GPT-6.1 Sol) | 77.5 (GPT-6 Astra) | 45.6 (GPT-6 Luna) |
The target
56 shapes: a robot agent on a chair with five slanted legs, a round desk on three slanted legs, a monitor with its cable, a desk lamp, a mug and a potted plant. Scoring compares 1 cm cubes; the pictures use 2 cm cubes.
How it works
The brief lists every shape of the target with exact numbers. Each model answers with JSON only: up to 150 boxes, each a position, a size and one of 11 colors. The model draws nothing itself. When a run is scored we render its boxes and the target from the same small cubes, with the same light and camera, and publish the pictures.
The score is the volume overlap between sculpture and target (intersection over union, measured in 1 cm cubes) times the share of shared cubes with the right color. One box per shape scores under 40; slicing every shape into equal slabs scores about 50. Slanted legs, thin cables and round heads reward models that plan where each box goes, and the Difference view shows every missing, extra and wrongly colored cube.
Fair and repeatable
The starting baseline is the original first round, preserved with its actual measurement dates. A second measured round was requested for launch. Both are shown; neither is backdated or selected as a best-of-two result. Later rounds follow the weekly schedule.
- Same brief, same scorer and same settings every week. Brief v3 has the fingerprint
08fb5a5773fe89e7; any change starts a new version and a new chart. - Each model answers once at medium reasoning effort, through Claude Code or Codex signed in with a subscription, the same way VibeiDE runs them. No tools, no web, no retries, no best-of-N.
- Models run one after another, never in parallel. An answer that does not arrive within 40 minutes scores 0.
- The scorer is a deterministic program. No model grades another model.
Read the full brief (plain text, 7,691 characters). Model names are the aliases each CLI offers; the page records the model identifier the CLI reported.
Questions
Is Claude or GPT getting nerfed?
This page cannot see inside a provider, but it shows whether the same model name scores differently on the same frozen task over time. One run per week is a single sample, so treat a move of a few points as noise and a drop that lasts several weeks as a signal.
Why can no model reach 100?
The target is made of spheres, ellipsoids, cylinders and slanted capsules, and the answer may only use 150 axis-aligned boxes. Boxes cannot follow curves exactly, so every answer loses some volume. Better models spend their boxes where the curves are, and their sculptures look more like the target.
How are the models run?
Each model gets the brief once through its official command-line tool, Claude Code or Codex, signed in with a normal subscription, at medium reasoning effort. Tools, sub-agents, web access and retries are off; a run that uses a tool scores 0. The reply is scored by a deterministic program, not by another model.
Why did the benchmark change to v3?
The first two versions asked models to lay out an office. In their first rounds the best models scored 100 and 92, and every correct office looked the same. A benchmark that top models max out cannot show change, and one where answers look alike cannot show quality, so v3 asks for a sculpture whose quality is visible at a glance.
Can I try the brief myself?
Yes. The complete brief is public as plain text. Paste it into any model and compare its answer with the target.