The VibeiDE Office Benchmark

Give an AI
150 boxes.
Ask for this.

A round head. A curved chair. A tiny desk lamp. How close can a language model get using only rectangular blocks?

boxes maximum
150
models
6
attempt per round
1
See what they built
The target: a white robot on a green chair at a round desk with a monitor, lamp, mug and plant
The sculpture to recreate56 shapes

Quality, tokens and time

What did each result take?

Compare score, reported input plus output tokens, and elapsed runtime for the selected round. Reasoning tokens are not added again. Missing measurements are left out, not counted as zero.

View token and runtime measurements

The answers

The difference is in the details.

Every model's result, ranked by score. One attempt at medium effort, using the same frozen brief.

Week 2026-W41

Every sculpture uses the same light, camera and scale. Select any result to compare it with the target, or switch to Difference to inspect the mismatches.

  1. Claude Opus 5.578.5/100Claude Opus 5.5 sculpture
  2. GPT-6.1 Sol74.7/100GPT-6.1 Sol sculpture
  3. Claude Fable74.6/100Claude Fable sculpture
  4. GPT-6 Astra73.5/100GPT-6 Astra sculpture
  5. Claude Sonnet 5.548.1/100Claude Sonnet 5.5 sculpture
  6. GPT-6 Luna30.9/100GPT-6 Luna sculpture

Measured results

Scores over time

Starting baseline preserves the original first-round results and measurement dates. The next point is a fresh measured round, not a result from a different week.

View the recorded scores
Score out of 100 by round
RoundClaude Opus 5 / Claude Opus 5.5Claude Sonnet 5 / Claude Sonnet 5.5Claude FableGPT-6.1 SolGPT-6 AstraGPT-6 Luna
2026-W4178.5 (Claude Opus 5.5)48.1 (Claude Sonnet 5.5)74.6 (Claude Fable)74.7 (GPT-6.1 Sol)73.5 (GPT-6 Astra)30.9 (GPT-6 Luna)
Starting baseline75.5 (Claude Opus 5)34.3 (Claude Sonnet 5)71.2 (Claude Fable)70.6 (GPT-6.1 Sol)77.5 (GPT-6 Astra)45.6 (GPT-6 Luna)

The target

56 shapes: a robot agent on a chair with five slanted legs, a round desk on three slanted legs, a monitor with its cable, a desk lamp, a mug and a potted plant. Scoring compares 1 cm cubes; the pictures use 2 cm cubes.

The target: a white robot agent on a green chair at a round wooden desk with a monitor, lamp, mug and plant, built from small cubes

How it works

The brief lists every shape of the target with exact numbers. Each model answers with JSON only: up to 150 boxes, each a position, a size and one of 11 colors. The model draws nothing itself. When a run is scored we render its boxes and the target from the same small cubes, with the same light and camera, and publish the pictures.

The score is the volume overlap between sculpture and target (intersection over union, measured in 1 cm cubes) times the share of shared cubes with the right color. One box per shape scores under 40; slicing every shape into equal slabs scores about 50. Slanted legs, thin cables and round heads reward models that plan where each box goes, and the Difference view shows every missing, extra and wrongly colored cube.

Fair and repeatable

The starting baseline is the original first round, preserved with its actual measurement dates. A second measured round was requested for launch. Both are shown; neither is backdated or selected as a best-of-two result. Later rounds follow the weekly schedule.

Read the full brief (plain text, 7,691 characters). Model names are the aliases each CLI offers; the page records the model identifier the CLI reported.

Questions

Is Claude or GPT getting nerfed?

This page cannot see inside a provider, but it shows whether the same model name scores differently on the same frozen task over time. One run per week is a single sample, so treat a move of a few points as noise and a drop that lasts several weeks as a signal.

Why can no model reach 100?

The target is made of spheres, ellipsoids, cylinders and slanted capsules, and the answer may only use 150 axis-aligned boxes. Boxes cannot follow curves exactly, so every answer loses some volume. Better models spend their boxes where the curves are, and their sculptures look more like the target.

How are the models run?

Each model gets the brief once through its official command-line tool, Claude Code or Codex, signed in with a normal subscription, at medium reasoning effort. Tools, sub-agents, web access and retries are off; a run that uses a tool scores 0. The reply is scored by a deterministic program, not by another model.

Why did the benchmark change to v3?

The first two versions asked models to lay out an office. In their first rounds the best models scored 100 and 92, and every correct office looked the same. A benchmark that top models max out cannot show change, and one where answers look alike cannot show quality, so v3 asks for a sculpture whose quality is visible at a glance.

Can I try the brief myself?

Yes. The complete brief is public as plain text. Paste it into any model and compare its answer with the target.

Languages

Necessary cookies support sign-in, security, your language and this choice. Optional categories stay off until accepted.

First-party page views and acquisition measurement, plus Google Analytics 4 (Google Ireland, data may reach the US). Advertising personalization stays off.

With Analytics enabled, retain Google and Microsoft ad click identifiers and share download, signup, activation and purchase conversions with the relevant ad provider. No personalized advertising.

Remember referral credit for later. Links still work on the current page without this cookie.

Privacy · Cookie inventory