Office Benchmark

Which AI builds best?

Six models. One brief. One shot each. Scored from 0 to 100 against the exact target.

Leaderboard.

What each model built.

Raw output, rendered from the JSON each model returned. One shot, no retries.

Week 2026-W41

Every sculpture uses the same light, camera and scale. Select any result to compare it with the target, or switch to Difference to inspect the mismatches.

  1. Claude Opus 5.578.5/100Claude Opus 5.5 sculpture
  2. GPT-6.1 Sol74.7/100GPT-6.1 Sol sculpture
  3. Claude Fable74.6/100Claude Fable sculpture
  4. GPT-6 Astra73.5/100GPT-6 Astra sculpture
  5. Claude Sonnet 5.548.1/100Claude Sonnet 5.5 sculpture
  6. GPT-6 Luna30.9/100GPT-6 Luna sculpture

Quality, tokens and time

What did each result take?

Compare score, reported input plus output tokens, and elapsed runtime for the selected round. Reasoning tokens are not added again. Missing measurements are left out, not counted as zero.

View token and runtime measurements

Measured results

Round over round.

Starting baseline preserves the original first-round results and measurement dates. The next point is a fresh measured round, not a result from a different week.

View the recorded scores
Score out of 100 by round
RoundClaude Opus 5 / Claude Opus 5.5Claude Sonnet 5 / Claude Sonnet 5.5Claude FableGPT-6.1 SolGPT-6 AstraGPT-6 Luna
2026-W4178.5 (Claude Opus 5.5)48.1 (Claude Sonnet 5.5)74.6 (Claude Fable)74.7 (GPT-6.1 Sol)73.5 (GPT-6 Astra)30.9 (GPT-6 Luna)
Starting baseline75.5 (Claude Opus 5)34.3 (Claude Sonnet 5)71.2 (Claude Fable)70.6 (GPT-6.1 Sol)77.5 (GPT-6 Astra)45.6 (GPT-6 Luna)

About the benchmark

Give an AI 150 boxes. Ask for this.

A round head. A curved chair. A tiny desk lamp. How close can a language model get using only rectangular blocks?

The target: a white robot on a green chair at a round desk with a monitor, lamp, mug and plant
The sculpture to recreate. 56 shapes.
boxes maximum
150
models
6
attempt per round
1

How it works.

A frozen brief.

Every model receives the same written brief with target dimensions and colors. The full brief is public.

JSON only.

The reply contains nothing else: up to 150 boxes with positions, sizes and colors.

One shot.

Medium reasoning effort, no tools, no retries. Tokens and runtime are tracked.

A frozen scorer.

Volume overlap against the target, weighted by color accuracy, from 0 to 100.

The formula

volume overlap % × share of correctly colored cubes

Measured in 1 cm³ units. Brief fingerprint 08fb5a5773fe89e7. The starting baseline preserves the original first-round results.

Full methodology and questions

How it works

The brief lists every shape of the target with exact numbers. Each model answers with JSON only: up to 150 boxes, each a position, a size and one of 11 colors. The model draws nothing itself. When a run is scored we render its boxes and the target from the same small cubes, with the same light and camera, and publish the pictures.

The score is the volume overlap between sculpture and target (intersection over union, measured in 1 cm cubes) times the share of shared cubes with the right color. One box per shape scores under 40; slicing every shape into equal slabs scores about 50. Slanted legs, thin cables and round heads reward models that plan where each box goes, and the Difference view shows every missing, extra and wrongly colored cube.

Fair and repeatable

The starting baseline is the original first round, preserved with its actual measurement dates. A second measured round was requested for launch. Both are shown; neither is backdated or selected as a best-of-two result. Later rounds follow the weekly schedule.

  • Same brief, same scorer and same settings every week. Brief v3 has the fingerprint 08fb5a5773fe89e7; any change starts a new version and a new chart.
  • Each model answers once at medium reasoning effort, through Claude Code or Codex signed in with a subscription, the same way VibeiDE runs them. No tools, no web, no retries, no best-of-N.
  • Models run one after another, never in parallel. An answer that does not arrive within 40 minutes scores 0.
  • The scorer is a deterministic program. No model grades another model.

Read the full brief (plain text, 7,691 characters). Model names are the aliases each CLI offers; the page records the model identifier the CLI reported.

Questions

Is Claude or GPT getting nerfed?

This page cannot see inside a provider, but it shows whether the same model name scores differently on the same frozen task over time. One run per week is a single sample, so treat a move of a few points as noise and a drop that lasts several weeks as a signal.

Why can no model reach 100?

The target is made of spheres, ellipsoids, cylinders and slanted capsules, and the answer may only use 150 axis-aligned boxes. Boxes cannot follow curves exactly, so every answer loses some volume. Better models spend their boxes where the curves are, and their sculptures look more like the target.

How are the models run?

Each model gets the brief once through its official command-line tool, Claude Code or Codex, signed in with a normal subscription, at medium reasoning effort. Tools, sub-agents, web access and retries are off; a run that uses a tool scores 0. The reply is scored by a deterministic program, not by another model.

Why did the benchmark change to v3?

The first two versions asked models to lay out an office. In their first rounds the best models scored 100 and 92, and every correct office looked the same. A benchmark that top models max out cannot show change, and one where answers look alike cannot show quality, so v3 asks for a sculpture whose quality is visible at a glance.

Can I try the brief myself?

Yes. The complete brief is public as plain text. Paste it into any model and compare its answer with the target.

Languages

Necessary cookies support sign-in, security, your language and this choice. Optional categories stay off until accepted.

First-party page views and acquisition measurement, plus Google Analytics 4 (Google Ireland, data may reach the US). Advertising personalization stays off.

With Analytics enabled, retain Google and Microsoft ad click identifiers and share download, signup, activation and purchase conversions with the relevant ad provider. No personalized advertising.

Remember referral credit for later. Links still work on the current page without this cookie.

Privacy · Cookie inventory