All benchmarks

LayoutBench

V0.3.0

Can a model see how a page is put together? Twenty atomic tasks ask about text alignment, grids, spacing, wrapping, and nested layout hierarchy. Each question has one answer that can be checked against the visible arrangement.

17 model configurations · 292 shared questions · Human validation has not been performed.

Layout perception and efficiency

Compare one task family at a time. Scores, cost, and response time use the same questions within that family.

Explain Dimensions

Green points mark the best accuracy–cost tradeoffs. Both axes are linear.

Layout perception and efficiency0%20%40%60%80%100%$0$0.005$0.010$0.015$0.020$0.025$0.030Metered API cost per scored input (USD) →Text justification accuracy (%) →GPT-6 Luna: 87.5% ($0.0003 / input)GPT-6 LunaGemini 3.5 Flash-Lite: 25.0% ($0.0004 / input)Gemini 3.5 Flash-LiteGPT-5.6 Luna: 68.8% ($0.0005 / input)GPT-5.6 LunaClaude Haiku 4.5: 25.0% ($0.0019 / input)Claude Haiku 4.5Gemini 3.8 Flash: 100.0% ($0.0034 / input)Gemini 3.8 FlashGPT-6 Sol: 100.0% ($0.0059 / input)GPT-6 SolClaude Sonnet 5: 87.5% ($0.0059 / input)Claude Sonnet 5GPT-5.6 Terra: 62.5% ($0.0060 / input)GPT-5.6 TerraGemini 3.1 Pro Preview: 81.3% ($0.0088 / input)Gemini 3.1 Pro PreviewPareto: 62.5% ($0.0097 / input)ParetoGPT-5.6 Sol: 87.5% ($0.0102 / input)GPT-5.6 SolClaude Opus 5.5: 100.0% ($0.0117 / input)Claude Opus 5.5Muse Spark 1.2: 81.3% ($0.0128 / input)Muse Spark 1.2Muse Spark 1.3: 75.0% ($0.0139 / input)Muse Spark 1.3Claude Opus 5: 100.0% ($0.0152 / input)Claude Opus 5GPT-6 Astra: 100.0% ($0.0250 / input)GPT-6 AstraClaude Fable 5.1: 100.0% ($0.0291 / input)Claude Fable 5.1
Swipe chart horizontally to explore →

Choice accuracy includes invalid answers as incorrect. Uniform random choice gives 25.0% expected accuracy in this family. Each task asks for one visible relationship; there is no combined score across families.

A point is omitted when its selected cost or latency measurement is unavailable. Every scored model remains in the accuracy table below; an unavailable measurement is never treated as zero.

Text justification results

Select a model to highlight it in the chart and inspect its diagnostics.

LayoutBench Text justification on shared questions
ModelAccuracyCorrectScoredInvalidAPI cost / questionResponse time
100.0%16/16160$0.02914.51s
100.0%16/16160$0.01522.88s
100.0%16/16160$0.01172.24s
100.0%16/16160$0.00343.13s
100.0%16/16160$0.02502.07s
100.0%16/16160$0.00592.18s
87.5%14/16160$0.00592.03s
87.5%14/16160$0.01021.81s
87.5%14/16160$0.00032.98s
81.3%13/16160$0.00886.44s
81.3%13/16160$0.012811.49s
75.0%12/16160$0.013925.40s
68.8%11/16160$0.00051.72s
62.5%10/16160$0.00602.40s
62.5%10/16163$0.00973.70s
25.0%4/16160$0.00190.89s
25.0%4/16160$0.00041.01s

Invalid responses count as incorrect. Unknown cost and latency measurements appear as —. API estimates use recorded usage and rates; response time also reflects network and service conditions.

What stands out

Text justification: 215/272 (79.0%); Grid gutter comparison: 163/204 (79.9%); Horizontal versus vertical inset: 164/204 (80.4%). These have the lowest pooled accuracy across this model roster. Pooling describes these responses; the graph and table show each model separately. Families have two to six answer choices, so chance levels differ; stimulus difficulty also limits comparisons between families.

Text justification

79.0%

215/272 model–question answers

Grid gutter comparison

79.9%

163/204 model–question answers

Horizontal versus vertical inset

80.4%

164/204 model–question answers

Immediate flow, Text columns, Grid columns, Grid rows, Wrapped row count, Text block placement, Relative panel width, Top-level hierarchy, Nested section flow reached 100% across every compared configuration. A ceiling on these controlled images shows where this corpus stops distinguishing models. It does not establish general layout understanding.

Right-aligned text mistaken for centering

Right-aligned paragraphs were recognized in 42/68 responses; 16/68 instead called them centered. Claude Fable 5.1 and Claude Opus 5 and Gemini 3.8 Flash and GPT-6 Astra and Claude Opus 5.5 and GPT-6 Sol answered all sixteen text-justification questions correctly. The difficult right-aligned example below shows the exact text and all 17 answers.

Equal spaces draw more errors

Equal spacing produced more errors than unequal spacing in these questions. The table pools the same 17 completed configurations and keeps the two tasks separate.

Accuracy on equal and unequal spacing
TaskEqualUnequal
Grid gutter comparison42/68121/136
Horizontal versus vertical inset38/68126/136

GPT-5.6 Sol and GPT-6 Astra and GPT-6 Luna and GPT-6 Sol answered all twelve grid-gutter and all twelve inset questions correctly. These descriptive differences do not establish why the errors occurred.

A page can run vertically and a row horizontally

Layout changes with the level of the hierarchy being inspected. A page can stack sections vertically while the children of one section form a horizontal row or a wrapped collection. These three tasks name the level to inspect so a correct global description cannot substitute for the local answer.

Across 408 paired outcomes (17 models × 24 shared images), both answers were correct in 408/408; 0 had the outer answer correct and the inner answer wrong. Every compared configuration recognized both levels of flow in these outlined synthetic scenes.

Alignment within a wrapped section produced a different result: nested final-row alignment was correct in 398/408 answers. Claude Haiku 4.5 scored 17/24; Pareto scored 21/24; the other 15 configurations answered all 24 correctly. The corpus includes a left-aligned wrapped collection in the second vertically stacked section, so a local alignment question can differ from the page's overall flow.

In the nested example below, Claude Haiku 4.5 recognizes the vertically stacked page and the wrapped children inside the target section, but calls its left-aligned final row centered. The answers below show how the other configurations handled that same image.

Top-level hierarchy

408/408

Are the page's top-level regions organized vertically first or horizontally first?

Nested section flow

408/408

Within the indicated region, do its children form a row, a column, or multiple wrapped rows?

Nested incomplete row alignment

398/408

Inside the indicated region, is the shorter final row aligned left, centered, or aligned right?

Top-level hierarchy and nested section flow ask two questions about each of the same 24 images. Comparing their paired answers shows where a model recognizes the outer structure but misses its children. Flat and nested stimuli still do not isolate a causal effect of nesting.

Inspect the questions and mistakes

Each example shows a selected question from its family. The expandable answers show successes and errors on the same image. Models receive the image and question, including its answer choices, without the ground truth.

LayoutBench Immediate flow question

Immediate flow

17/17 correct

How are the cards inside the outlined frame arranged? A: Stacked in one vertical column B: Side by side in one horizontal row Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: B

  • Claude Fable 5.1B · correct
  • Claude Haiku 4.5B · correct
  • Claude Opus 5B · correct
  • Claude Sonnet 5B · correct
  • Gemini 3.1 Pro PreviewB · correct
  • Gemini 3.5 Flash-LiteB · correct
  • Gemini 3.8 FlashB · correct
  • GPT-5.6 LunaB · correct
  • GPT-5.6 SolB · correct
  • GPT-5.6 TerraB · correct
  • GPT-6 AstraB · correct
  • Muse Spark 1.2B · correct
  • Muse Spark 1.3B · correct
  • Claude Opus 5.5B · correct
  • GPT-6 LunaB · correct
  • GPT-6 SolB · correct
  • ParetoB · correct

layoutbench-direction-001

LayoutBench Main-axis distribution question

Main-axis distribution

14/17 correct

Consider the sequence of cards along its long direction inside the outlined frame. Start means left for a row and top for a column. Which description fits the empty spaces measured from card edges to the INSIDE of the frame border? A: All interior gaps and both outer spaces are equal B: Packed together in the middle, with much larger equal outer spaces C: Packed at the end, touching that inner edge D: Packed at the start, touching that inner edge E: First and last cards touch opposite inner edges; equal gaps between cards F: Equal outer spaces, each half the size of an interior gap Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: D

  • Claude Fable 5.1D · correct
  • Claude Haiku 4.5B · incorrect
  • Claude Opus 5D · correct
  • Claude Sonnet 5F · incorrect
  • Gemini 3.1 Pro PreviewD · correct
  • Gemini 3.5 Flash-LiteB · incorrect
  • Gemini 3.8 FlashD · correct
  • GPT-5.6 LunaD · correct
  • GPT-5.6 SolD · correct
  • GPT-5.6 TerraD · correct
  • GPT-6 AstraD · correct
  • Muse Spark 1.2D · correct
  • Muse Spark 1.3D · correct
  • Claude Opus 5.5D · correct
  • GPT-6 LunaD · correct
  • GPT-6 SolD · correct
  • ParetoD · correct

layoutbench-distribution-004

LayoutBench Cross-axis alignment question

Cross-axis alignment

16/17 correct

Across the direction perpendicular to the sequence of cards, how do their visible boxes align inside the outlined frame? For a horizontal sequence compare vertical positions; for a vertical sequence compare horizontal positions. A: Every box is centered on that cross axis, with empty space on both sides of the box B: Every box fills the frame along that cross axis C: Every box touches the bottom (horizontal sequence) or right (vertical sequence) inner edge, leaving space before the opposite edge D: Every box touches the top (horizontal sequence) or left (vertical sequence) inner edge, leaving space before the opposite edge Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: A

  • Claude Fable 5.1A · correct
  • Claude Haiku 4.5D · incorrect
  • Claude Opus 5A · correct
  • Claude Sonnet 5A · correct
  • Gemini 3.1 Pro PreviewA · correct
  • Gemini 3.5 Flash-LiteA · correct
  • Gemini 3.8 FlashA · correct
  • GPT-5.6 LunaA · correct
  • GPT-5.6 SolA · correct
  • GPT-5.6 TerraA · correct
  • GPT-6 AstraA · correct
  • Muse Spark 1.2A · correct
  • Muse Spark 1.3A · correct
  • Claude Opus 5.5A · correct
  • GPT-6 LunaA · correct
  • GPT-6 SolA · correct
  • ParetoA · correct

layoutbench-crossalign-005

LayoutBench Text justification question

Text justification

7/17 correct

How are the lines of body text aligned within the outlined text region? For both-edge justification, disregard the short final line. A: Right edges line up; left edges are ragged B: Left edges line up; right edges are ragged C: Both left and right edges line up on the nonfinal lines D: Line centers line up; both edges are ragged Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: A

  • Claude Fable 5.1A · correct
  • Claude Haiku 4.5C · incorrect
  • Claude Opus 5A · correct
  • Claude Sonnet 5D · incorrect
  • Gemini 3.1 Pro PreviewD · incorrect
  • Gemini 3.5 Flash-LiteB · incorrect
  • Gemini 3.8 FlashA · correct
  • GPT-5.6 LunaD · incorrect
  • GPT-5.6 SolD · incorrect
  • GPT-5.6 TerraD · incorrect
  • GPT-6 AstraA · correct
  • Muse Spark 1.2D · incorrect
  • Muse Spark 1.3D · incorrect
  • Claude Opus 5.5A · correct
  • GPT-6 LunaA · correct
  • GPT-6 SolA · correct
  • ParetoD · incorrect

layoutbench-textalign-012

LayoutBench Text columns question

Text columns

17/17 correct

How many side-by-side columns of body text are visible inside the frame? A: One B: Two C: Three Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: A

  • Claude Fable 5.1A · correct
  • Claude Haiku 4.5A · correct
  • Claude Opus 5A · correct
  • Claude Sonnet 5A · correct
  • Gemini 3.1 Pro PreviewA · correct
  • Gemini 3.5 Flash-LiteA · correct
  • Gemini 3.8 FlashA · correct
  • GPT-5.6 LunaA · correct
  • GPT-5.6 SolA · correct
  • GPT-5.6 TerraA · correct
  • GPT-6 AstraA · correct
  • Muse Spark 1.2A · correct
  • Muse Spark 1.3A · correct
  • Claude Opus 5.5A · correct
  • GPT-6 LunaA · correct
  • GPT-6 SolA · correct
  • ParetoA · correct

layoutbench-textcolumns-001

LayoutBench Grid columns question

Grid columns

17/17 correct

How many columns of cards are in the rectangular grid? Count the cards across one complete row. A: Two B: Three C: Four Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: A

  • Claude Fable 5.1A · correct
  • Claude Haiku 4.5A · correct
  • Claude Opus 5A · correct
  • Claude Sonnet 5A · correct
  • Gemini 3.1 Pro PreviewA · correct
  • Gemini 3.5 Flash-LiteA · correct
  • Gemini 3.8 FlashA · correct
  • GPT-5.6 LunaA · correct
  • GPT-5.6 SolA · correct
  • GPT-5.6 TerraA · correct
  • GPT-6 AstraA · correct
  • Muse Spark 1.2A · correct
  • Muse Spark 1.3A · correct
  • Claude Opus 5.5A · correct
  • GPT-6 LunaA · correct
  • GPT-6 SolA · correct
  • ParetoA · correct

layoutbench-gridcols-001

LayoutBench Grid rows question

Grid rows

17/17 correct

How many rows of cards are in the rectangular grid? Count the cards down one complete column. A: Two B: Three C: Four Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: A

  • Claude Fable 5.1A · correct
  • Claude Haiku 4.5A · correct
  • Claude Opus 5A · correct
  • Claude Sonnet 5A · correct
  • Gemini 3.1 Pro PreviewA · correct
  • Gemini 3.5 Flash-LiteA · correct
  • Gemini 3.8 FlashA · correct
  • GPT-5.6 LunaA · correct
  • GPT-5.6 SolA · correct
  • GPT-5.6 TerraA · correct
  • GPT-6 AstraA · correct
  • Muse Spark 1.2A · correct
  • Muse Spark 1.3A · correct
  • Claude Opus 5.5A · correct
  • GPT-6 LunaA · correct
  • GPT-6 SolA · correct
  • ParetoA · correct

layoutbench-gridrows-001

LayoutBench Grid gutter comparison question

Grid gutter comparison

9/17 correct

Compare the horizontal empty gutters between columns with the vertical empty gutters between rows. Measure between the visible card boxes. A: Horizontal gutters are larger B: Vertical gutters are larger C: The two gutter sizes are equal Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: C

  • Claude Fable 5.1B · incorrect
  • Claude Haiku 4.5A · incorrect
  • Claude Opus 5A · incorrect
  • Claude Sonnet 5B · incorrect
  • Gemini 3.1 Pro PreviewB · incorrect
  • Gemini 3.5 Flash-LiteB · incorrect
  • Gemini 3.8 FlashB · incorrect
  • GPT-5.6 LunaC · correct
  • GPT-5.6 SolC · correct
  • GPT-5.6 TerraC · correct
  • GPT-6 AstraC · correct
  • Muse Spark 1.2C · correct
  • Muse Spark 1.3C · correct
  • Claude Opus 5.5C · correct
  • GPT-6 LunaC · correct
  • GPT-6 SolC · correct
  • ParetoInvalid answer

layoutbench-gridgaps-012

LayoutBench Grid track proportions question

Grid track proportions

16/17 correct

Compare the widths of the three card columns in this grid. A: All three columns have equal widths B: The left column is wider; the other two have equal widths C: The right column is wider; the other two have equal widths Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: B

  • Claude Fable 5.1B · correct
  • Claude Haiku 4.5A · incorrect
  • Claude Opus 5B · correct
  • Claude Sonnet 5B · correct
  • Gemini 3.1 Pro PreviewB · correct
  • Gemini 3.5 Flash-LiteB · correct
  • Gemini 3.8 FlashB · correct
  • GPT-5.6 LunaB · correct
  • GPT-5.6 SolB · correct
  • GPT-5.6 TerraB · correct
  • GPT-6 AstraB · correct
  • Muse Spark 1.2B · correct
  • Muse Spark 1.3B · correct
  • Claude Opus 5.5B · correct
  • GPT-6 LunaB · correct
  • GPT-6 SolB · correct
  • ParetoB · correct

layoutbench-gridtracks-005

LayoutBench Column span question

Column span

14/17 correct

Use the three equal-width cards in the lower row as column guides. How many of those column positions does the single card in the upper row cover, including any intervening gutters? A: Three column positions B: One column position C: Two column positions Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: B

  • Claude Fable 5.1B · correct
  • Claude Haiku 4.5A · incorrect
  • Claude Opus 5B · correct
  • Claude Sonnet 5B · correct
  • Gemini 3.1 Pro PreviewB · correct
  • Gemini 3.5 Flash-LiteC · incorrect
  • Gemini 3.8 FlashB · correct
  • GPT-5.6 LunaC · incorrect
  • GPT-5.6 SolB · correct
  • GPT-5.6 TerraB · correct
  • GPT-6 AstraB · correct
  • Muse Spark 1.2B · correct
  • Muse Spark 1.3B · correct
  • Claude Opus 5.5B · correct
  • GPT-6 LunaB · correct
  • GPT-6 SolB · correct
  • ParetoB · correct

layoutbench-gridspan-001

LayoutBench Wrapped row count question

Wrapped row count

17/17 correct

How many horizontal rows of small cards are visible inside the frame? A: One row B: Three rows C: Two rows Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: A

  • Claude Fable 5.1A · correct
  • Claude Haiku 4.5A · correct
  • Claude Opus 5A · correct
  • Claude Sonnet 5A · correct
  • Gemini 3.1 Pro PreviewA · correct
  • Gemini 3.5 Flash-LiteA · correct
  • Gemini 3.8 FlashA · correct
  • GPT-5.6 LunaA · correct
  • GPT-5.6 SolA · correct
  • GPT-5.6 TerraA · correct
  • GPT-6 AstraA · correct
  • Muse Spark 1.2A · correct
  • Muse Spark 1.3A · correct
  • Claude Opus 5.5A · correct
  • GPT-6 LunaA · correct
  • GPT-6 SolA · correct
  • ParetoA · correct

layoutbench-wrapcount-001

LayoutBench Incomplete row alignment question

Incomplete row alignment

16/17 correct

Look only at the shorter final row of cards. How is that row placed horizontally inside the outlined frame? A: Centered, with equal empty space on either side B: Against the right inner edge C: Against the left inner edge Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: C

  • Claude Fable 5.1C · correct
  • Claude Haiku 4.5A · incorrect
  • Claude Opus 5C · correct
  • Claude Sonnet 5C · correct
  • Gemini 3.1 Pro PreviewC · correct
  • Gemini 3.5 Flash-LiteC · correct
  • Gemini 3.8 FlashC · correct
  • GPT-5.6 LunaC · correct
  • GPT-5.6 SolC · correct
  • GPT-5.6 TerraC · correct
  • GPT-6 AstraC · correct
  • Muse Spark 1.2C · correct
  • Muse Spark 1.3C · correct
  • Claude Opus 5.5C · correct
  • GPT-6 LunaC · correct
  • GPT-6 SolC · correct
  • ParetoC · correct

layoutbench-wrapalign-002

LayoutBench Neighboring gap comparison question

Neighboring gap comparison

15/17 correct

Compare the two empty gaps between the three card boxes, reading left to right for a row or top to bottom for a column. A: The gaps are equal B: The second gap is larger C: The first gap is larger Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: C

  • Claude Fable 5.1C · correct
  • Claude Haiku 4.5A · incorrect
  • Claude Opus 5C · correct
  • Claude Sonnet 5C · correct
  • Gemini 3.1 Pro PreviewC · correct
  • Gemini 3.5 Flash-LiteA · incorrect
  • Gemini 3.8 FlashC · correct
  • GPT-5.6 LunaC · correct
  • GPT-5.6 SolC · correct
  • GPT-5.6 TerraC · correct
  • GPT-6 AstraC · correct
  • Muse Spark 1.2C · correct
  • Muse Spark 1.3C · correct
  • Claude Opus 5.5C · correct
  • GPT-6 LunaC · correct
  • GPT-6 SolC · correct
  • ParetoC · correct

layoutbench-gapcompare-001

LayoutBench Horizontal versus vertical inset question

Horizontal versus vertical inset

8/17 correct

The tinted rectangle is centered inside the outlined frame. Compare the horizontal inset (frame to rectangle on either side) with the vertical inset (frame to rectangle above or below). Measure to the rectangle's visible outer edge, not to its text. A: The horizontal inset is larger B: The vertical inset is larger C: The horizontal and vertical insets are equal Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: C

  • Claude Fable 5.1B · incorrect
  • Claude Haiku 4.5A · incorrect
  • Claude Opus 5C · correct
  • Claude Sonnet 5B · incorrect
  • Gemini 3.1 Pro PreviewB · incorrect
  • Gemini 3.5 Flash-LiteB · incorrect
  • Gemini 3.8 FlashB · incorrect
  • GPT-5.6 LunaC · correct
  • GPT-5.6 SolC · correct
  • GPT-5.6 TerraA · incorrect
  • GPT-6 AstraC · correct
  • Muse Spark 1.2B · incorrect
  • Muse Spark 1.3A · incorrect
  • Claude Opus 5.5C · correct
  • GPT-6 LunaC · correct
  • GPT-6 SolC · correct
  • ParetoC · correct

layoutbench-padcompare-012

LayoutBench Text block placement question

Text block placement

17/17 correct

Where is the tinted text block placed horizontally within the outer frame? Judge the block's outer edges, independently of how its text lines align. A: Against the left inner edge B: Centered, with equal empty space on either side C: Against the right inner edge Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: A

  • Claude Fable 5.1A · correct
  • Claude Haiku 4.5A · correct
  • Claude Opus 5A · correct
  • Claude Sonnet 5A · correct
  • Gemini 3.1 Pro PreviewA · correct
  • Gemini 3.5 Flash-LiteA · correct
  • Gemini 3.8 FlashA · correct
  • GPT-5.6 LunaA · correct
  • GPT-5.6 SolA · correct
  • GPT-5.6 TerraA · correct
  • GPT-6 AstraA · correct
  • Muse Spark 1.2A · correct
  • Muse Spark 1.3A · correct
  • Claude Opus 5.5A · correct
  • GPT-6 LunaA · correct
  • GPT-6 SolA · correct
  • ParetoA · correct

layoutbench-blockalign-001

LayoutBench Relative panel width question

Relative panel width

17/17 correct

Compare the widths of the two tinted panels, from their visible left edge to right edge. A: Their widths are equal B: The left panel is wider C: The right panel is wider Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: B

  • Claude Fable 5.1B · correct
  • Claude Haiku 4.5B · correct
  • Claude Opus 5B · correct
  • Claude Sonnet 5B · correct
  • Gemini 3.1 Pro PreviewB · correct
  • Gemini 3.5 Flash-LiteB · correct
  • Gemini 3.8 FlashB · correct
  • GPT-5.6 LunaB · correct
  • GPT-5.6 SolB · correct
  • GPT-5.6 TerraB · correct
  • GPT-6 AstraB · correct
  • Muse Spark 1.2B · correct
  • Muse Spark 1.3B · correct
  • Claude Opus 5.5B · correct
  • GPT-6 LunaB · correct
  • GPT-6 SolB · correct
  • ParetoB · correct

layoutbench-widthcompare-001

LayoutBench Spacing within and between groups question

Spacing within and between groups

12/17 correct

The cards form two groups marked Harbor and Meadow. Compare the gap between the groups' nearest card boxes with the gaps between cards within each group. Ignore the small group captions. A: The within-group gaps are larger B: The between-group and within-group gaps are equal C: The between-group gap is larger Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: B

  • Claude Fable 5.1B · correct
  • Claude Haiku 4.5C · incorrect
  • Claude Opus 5B · correct
  • Claude Sonnet 5C · incorrect
  • Gemini 3.1 Pro PreviewB · correct
  • Gemini 3.5 Flash-LiteC · incorrect
  • Gemini 3.8 FlashB · correct
  • GPT-5.6 LunaC · incorrect
  • GPT-5.6 SolB · correct
  • GPT-5.6 TerraB · correct
  • GPT-6 AstraB · correct
  • Muse Spark 1.2C · incorrect
  • Muse Spark 1.3B · correct
  • Claude Opus 5.5B · correct
  • GPT-6 LunaB · correct
  • GPT-6 SolB · correct
  • ParetoB · correct

layoutbench-groupgap-009

LayoutBench Top-level hierarchy question

Top-level hierarchy

17/17 correct

Look at the two large outlined sections headed Field notes and Harbor log. At the top level, how are these two sections arranged? Ignore the arrangement of cards inside either section. A: Side by side horizontally B: Stacked vertically Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: A

  • Claude Fable 5.1A · correct
  • Claude Haiku 4.5A · correct
  • Claude Opus 5A · correct
  • Claude Sonnet 5A · correct
  • Gemini 3.1 Pro PreviewA · correct
  • Gemini 3.5 Flash-LiteA · correct
  • Gemini 3.8 FlashA · correct
  • GPT-5.6 LunaA · correct
  • GPT-5.6 SolA · correct
  • GPT-5.6 TerraA · correct
  • GPT-6 AstraA · correct
  • Muse Spark 1.2A · correct
  • Muse Spark 1.3A · correct
  • Claude Opus 5.5A · correct
  • GPT-6 LunaA · correct
  • GPT-6 SolA · correct
  • ParetoA · correct

layoutbench-topflow-001

LayoutBench Nested section flow question

Nested section flow

17/17 correct

Look inside the large section headed Field notes. How are its small cards arranged? Ignore how Field notes is placed relative to Harbor log. A: One horizontal row B: One vertical column C: Multiple horizontal rows, with a shorter final row Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: A

  • Claude Fable 5.1A · correct
  • Claude Haiku 4.5A · correct
  • Claude Opus 5A · correct
  • Claude Sonnet 5A · correct
  • Gemini 3.1 Pro PreviewA · correct
  • Gemini 3.5 Flash-LiteA · correct
  • Gemini 3.8 FlashA · correct
  • GPT-5.6 LunaA · correct
  • GPT-5.6 SolA · correct
  • GPT-5.6 TerraA · correct
  • GPT-6 AstraA · correct
  • Muse Spark 1.2A · correct
  • Muse Spark 1.3A · correct
  • Claude Opus 5.5A · correct
  • GPT-6 LunaA · correct
  • GPT-6 SolA · correct
  • ParetoA · correct

layoutbench-nestedflow-001

LayoutBench Nested incomplete row alignment question

Nested incomplete row alignment

16/17 correct

Inside Field notes, look only at the shorter final row of cards. How is it placed horizontally inside its own outlined inner frame? Ignore the surrounding page's arrangement. A: Against that frame's left inner edge B: Centered in that frame C: Against that frame's right inner edge Return only a JSON object with one key, "choice", whose value is the selected option letter.

Ground truth and model answers

Correct choice: A

  • Claude Fable 5.1A · correct
  • Claude Haiku 4.5B · incorrect
  • Claude Opus 5A · correct
  • Claude Sonnet 5A · correct
  • Gemini 3.1 Pro PreviewA · correct
  • Gemini 3.5 Flash-LiteA · correct
  • Gemini 3.8 FlashA · correct
  • GPT-5.6 LunaA · correct
  • GPT-5.6 SolA · correct
  • GPT-5.6 TerraA · correct
  • GPT-6 AstraA · correct
  • Muse Spark 1.2A · correct
  • Muse Spark 1.3A · correct
  • Claude Opus 5.5A · correct
  • GPT-6 LunaA · correct
  • GPT-6 SolA · correct
  • ParetoA · correct

layoutbench-nestedwrap-007

Twenty atomic questions about space

The intended building blocks are qualitative: direction, alignment, counts, relative size, proximity, and explicitly scoped hierarchy. Correctness concerns the rendered arrangement. A screenshot cannot uniquely identify the underlying CSS implementation.

LayoutBench qualitative task definitions
FamilyQuestion
Immediate flowAre the items arranged in a row or a column?
Main-axis distributionAre items packed at the start, center, or end, or spread with space between, around, or evenly among them?
Cross-axis alignmentHow do differently sized items align across their row or column?
Text justificationAre the paragraph's visible line edges left-aligned, centered, right-aligned, or justified?
Text columnsHow many side-by-side columns contain the body text?
Grid columnsHow many columns make up the repeated grid?
Grid rowsHow many rows make up the repeated grid?
Grid gutter comparisonAre the horizontal or vertical spaces between grid items larger, or are they equal?
Grid track proportionsAre grid columns equal in width, or is one visibly wider?
Column spanHow many column positions does the upper card occupy, using the complete lower row as a guide?
Wrapped row countHow many lines do the repeated items wrap onto?
Incomplete row alignmentHow is the incomplete final line of a wrapped collection aligned?
Neighboring gap comparisonWhich of two visible gaps is larger, or are they equal?
Horizontal versus vertical insetAre the horizontal or vertical insets around the centered content block larger, or are they equal?
Text block placementWhere does the tinted text block sit horizontally within its larger frame, independently of its text alignment?
Relative panel widthIs the left or right panel wider, or are their widths equal?
Spacing within and between groupsIs the space between the two labeled groups larger or smaller than the gaps within them, or equal?
Top-level hierarchyAre the page's top-level regions organized vertically first or horizontally first?
Nested section flowWithin the indicated region, do its children form a row, a column, or multiple wrapped rows?
Nested incomplete row alignmentInside the indicated region, is the shorter final row aligned left, centered, or aligned right?

What this measures

All 17 published configurations share 292/292 release questions. The release contains 250 distinct images. Choice accuracy is exact correctness, with invalid responses retained in the denominator. Each family has its own answer set and uniform-choice reference; no combined score is reported.

Each 1600×1200 image renders an 800×600 CSS-pixel canvas with bundled DejaVu Sans text and a light or dark theme. The model receives the image and question, without HTML or measured box coordinates. Sample text brings the questions closer to interface layouts, while controlled boxes isolate particular spatial relationships. The experiment measures the complete image-input and answering pipeline, including question interpretation.

Grid-row, grid-column, and text-column questions also permit counting or repeated-content cues in this finite corpus, so their scores do not isolate spatial grouping from those cues. There is no evidence here that a model used that shortcut.

Human agreement has not been measured. The questions aim for clear judgments by design-skilled readers, but that aim is not a measured agreement result. Related renderings are correlated, and small score differences need more evidence. These atomic skills are plausible ingredients of richer design understanding; this experiment does not test transfer to higher-level tasks.

Reproduce this comparison

4,964 final responses · 4,967 runner attempt records · 3 infrastructure attempts · 15 invalid answers. Recorded API estimate: $42.22. Estimates use recorded usage and dated rates, not invoice totals. The campaign ledger includes 6 attempt records with conservative cost reservations and 6 unmetered request allowances. A populated ledger total does not mean every request was metered. Family cost is unavailable if any scored response in that model–family pair lacks complete metering.

Dataset fc531bdb2e7c9d10e2074519bd3e0d0f66bc8ec970d20f25cdb628054784c6de
Source 811dbd016689132a51ec1ffe80de7225751af095