LayoutBench
V0.3.0
Can a model see how a page is put together? Twenty atomic tasks ask about text alignment, grids, spacing, wrapping, and nested layout hierarchy. Each question has one answer that can be checked against the visible arrangement.
17 model configurations · 292 shared questions · Human validation has not been performed.
Layout perception and efficiency
Compare one task family at a time. Scores, cost, and response time use the same questions within that family.
Green points mark the best accuracy–cost tradeoffs. Both axes are linear.
Choice accuracy includes invalid answers as incorrect. Uniform random choice gives 25.0% expected accuracy in this family. Each task asks for one visible relationship; there is no combined score across families.
A point is omitted when its selected cost or latency measurement is unavailable. Every scored model remains in the accuracy table below; an unavailable measurement is never treated as zero.
Text justification results
Select a model to highlight it in the chart and inspect its diagnostics.
| Model | Accuracy | Correct | Scored | Invalid | API cost / question | Response time |
|---|---|---|---|---|---|---|
| 100.0% | 16/16 | 16 | 0 | $0.0291 | 4.51s | |
| 100.0% | 16/16 | 16 | 0 | $0.0152 | 2.88s | |
| 100.0% | 16/16 | 16 | 0 | $0.0117 | 2.24s | |
| 100.0% | 16/16 | 16 | 0 | $0.0034 | 3.13s | |
| 100.0% | 16/16 | 16 | 0 | $0.0250 | 2.07s | |
| 100.0% | 16/16 | 16 | 0 | $0.0059 | 2.18s | |
| 87.5% | 14/16 | 16 | 0 | $0.0059 | 2.03s | |
| 87.5% | 14/16 | 16 | 0 | $0.0102 | 1.81s | |
| 87.5% | 14/16 | 16 | 0 | $0.0003 | 2.98s | |
| 81.3% | 13/16 | 16 | 0 | $0.0088 | 6.44s | |
| 81.3% | 13/16 | 16 | 0 | $0.0128 | 11.49s | |
| 75.0% | 12/16 | 16 | 0 | $0.0139 | 25.40s | |
| 68.8% | 11/16 | 16 | 0 | $0.0005 | 1.72s | |
| 62.5% | 10/16 | 16 | 0 | $0.0060 | 2.40s | |
| 62.5% | 10/16 | 16 | 3 | $0.0097 | 3.70s | |
| 25.0% | 4/16 | 16 | 0 | $0.0019 | 0.89s | |
| 25.0% | 4/16 | 16 | 0 | $0.0004 | 1.01s |
Invalid responses count as incorrect. Unknown cost and latency measurements appear as —. API estimates use recorded usage and rates; response time also reflects network and service conditions.
What stands out
Text justification: 215/272 (79.0%); Grid gutter comparison: 163/204 (79.9%); Horizontal versus vertical inset: 164/204 (80.4%). These have the lowest pooled accuracy across this model roster. Pooling describes these responses; the graph and table show each model separately. Families have two to six answer choices, so chance levels differ; stimulus difficulty also limits comparisons between families.
Text justification
79.0%
215/272 model–question answers
Grid gutter comparison
79.9%
163/204 model–question answers
Horizontal versus vertical inset
80.4%
164/204 model–question answers
Immediate flow, Text columns, Grid columns, Grid rows, Wrapped row count, Text block placement, Relative panel width, Top-level hierarchy, Nested section flow reached 100% across every compared configuration. A ceiling on these controlled images shows where this corpus stops distinguishing models. It does not establish general layout understanding.
Right-aligned text mistaken for centering
Right-aligned paragraphs were recognized in 42/68 responses; 16/68 instead called them centered. Claude Fable 5.1 and Claude Opus 5 and Gemini 3.8 Flash and GPT-6 Astra and Claude Opus 5.5 and GPT-6 Sol answered all sixteen text-justification questions correctly. The difficult right-aligned example below shows the exact text and all 17 answers.
Equal spaces draw more errors
Equal spacing produced more errors than unequal spacing in these questions. The table pools the same 17 completed configurations and keeps the two tasks separate.
| Task | Equal | Unequal |
|---|---|---|
| Grid gutter comparison | 42/68 | 121/136 |
| Horizontal versus vertical inset | 38/68 | 126/136 |
GPT-5.6 Sol and GPT-6 Astra and GPT-6 Luna and GPT-6 Sol answered all twelve grid-gutter and all twelve inset questions correctly. These descriptive differences do not establish why the errors occurred.
A page can run vertically and a row horizontally
Layout changes with the level of the hierarchy being inspected. A page can stack sections vertically while the children of one section form a horizontal row or a wrapped collection. These three tasks name the level to inspect so a correct global description cannot substitute for the local answer.
Across 408 paired outcomes (17 models × 24 shared images), both answers were correct in 408/408; 0 had the outer answer correct and the inner answer wrong. Every compared configuration recognized both levels of flow in these outlined synthetic scenes.
Alignment within a wrapped section produced a different result: nested final-row alignment was correct in 398/408 answers. Claude Haiku 4.5 scored 17/24; Pareto scored 21/24; the other 15 configurations answered all 24 correctly. The corpus includes a left-aligned wrapped collection in the second vertically stacked section, so a local alignment question can differ from the page's overall flow.
In the nested example below, Claude Haiku 4.5 recognizes the vertically stacked page and the wrapped children inside the target section, but calls its left-aligned final row centered. The answers below show how the other configurations handled that same image.
Top-level hierarchy
408/408
Are the page's top-level regions organized vertically first or horizontally first?
Nested section flow
408/408
Within the indicated region, do its children form a row, a column, or multiple wrapped rows?
Nested incomplete row alignment
398/408
Inside the indicated region, is the shorter final row aligned left, centered, or aligned right?
Top-level hierarchy and nested section flow ask two questions about each of the same 24 images. Comparing their paired answers shows where a model recognizes the outer structure but misses its children. Flat and nested stimuli still do not isolate a causal effect of nesting.
Inspect the questions and mistakes
Each example shows a selected question from its family. The expandable answers show successes and errors on the same image. Models receive the image and question, including its answer choices, without the ground truth.

Immediate flow
17/17 correct
How are the cards inside the outlined frame arranged? A: Stacked in one vertical column B: Side by side in one horizontal row Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: B
- Claude Fable 5.1B · correct
- Claude Haiku 4.5B · correct
- Claude Opus 5B · correct
- Claude Sonnet 5B · correct
- Gemini 3.1 Pro PreviewB · correct
- Gemini 3.5 Flash-LiteB · correct
- Gemini 3.8 FlashB · correct
- GPT-5.6 LunaB · correct
- GPT-5.6 SolB · correct
- GPT-5.6 TerraB · correct
- GPT-6 AstraB · correct
- Muse Spark 1.2B · correct
- Muse Spark 1.3B · correct
- Claude Opus 5.5B · correct
- GPT-6 LunaB · correct
- GPT-6 SolB · correct
- ParetoB · correct
layoutbench-direction-001

Main-axis distribution
14/17 correct
Consider the sequence of cards along its long direction inside the outlined frame. Start means left for a row and top for a column. Which description fits the empty spaces measured from card edges to the INSIDE of the frame border? A: All interior gaps and both outer spaces are equal B: Packed together in the middle, with much larger equal outer spaces C: Packed at the end, touching that inner edge D: Packed at the start, touching that inner edge E: First and last cards touch opposite inner edges; equal gaps between cards F: Equal outer spaces, each half the size of an interior gap Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: D
- Claude Fable 5.1D · correct
- Claude Haiku 4.5B · incorrect
- Claude Opus 5D · correct
- Claude Sonnet 5F · incorrect
- Gemini 3.1 Pro PreviewD · correct
- Gemini 3.5 Flash-LiteB · incorrect
- Gemini 3.8 FlashD · correct
- GPT-5.6 LunaD · correct
- GPT-5.6 SolD · correct
- GPT-5.6 TerraD · correct
- GPT-6 AstraD · correct
- Muse Spark 1.2D · correct
- Muse Spark 1.3D · correct
- Claude Opus 5.5D · correct
- GPT-6 LunaD · correct
- GPT-6 SolD · correct
- ParetoD · correct
layoutbench-distribution-004

Cross-axis alignment
16/17 correct
Across the direction perpendicular to the sequence of cards, how do their visible boxes align inside the outlined frame? For a horizontal sequence compare vertical positions; for a vertical sequence compare horizontal positions. A: Every box is centered on that cross axis, with empty space on both sides of the box B: Every box fills the frame along that cross axis C: Every box touches the bottom (horizontal sequence) or right (vertical sequence) inner edge, leaving space before the opposite edge D: Every box touches the top (horizontal sequence) or left (vertical sequence) inner edge, leaving space before the opposite edge Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: A
- Claude Fable 5.1A · correct
- Claude Haiku 4.5D · incorrect
- Claude Opus 5A · correct
- Claude Sonnet 5A · correct
- Gemini 3.1 Pro PreviewA · correct
- Gemini 3.5 Flash-LiteA · correct
- Gemini 3.8 FlashA · correct
- GPT-5.6 LunaA · correct
- GPT-5.6 SolA · correct
- GPT-5.6 TerraA · correct
- GPT-6 AstraA · correct
- Muse Spark 1.2A · correct
- Muse Spark 1.3A · correct
- Claude Opus 5.5A · correct
- GPT-6 LunaA · correct
- GPT-6 SolA · correct
- ParetoA · correct
layoutbench-crossalign-005

Text justification
7/17 correct
How are the lines of body text aligned within the outlined text region? For both-edge justification, disregard the short final line. A: Right edges line up; left edges are ragged B: Left edges line up; right edges are ragged C: Both left and right edges line up on the nonfinal lines D: Line centers line up; both edges are ragged Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: A
- Claude Fable 5.1A · correct
- Claude Haiku 4.5C · incorrect
- Claude Opus 5A · correct
- Claude Sonnet 5D · incorrect
- Gemini 3.1 Pro PreviewD · incorrect
- Gemini 3.5 Flash-LiteB · incorrect
- Gemini 3.8 FlashA · correct
- GPT-5.6 LunaD · incorrect
- GPT-5.6 SolD · incorrect
- GPT-5.6 TerraD · incorrect
- GPT-6 AstraA · correct
- Muse Spark 1.2D · incorrect
- Muse Spark 1.3D · incorrect
- Claude Opus 5.5A · correct
- GPT-6 LunaA · correct
- GPT-6 SolA · correct
- ParetoD · incorrect
layoutbench-textalign-012

Text columns
17/17 correct
How many side-by-side columns of body text are visible inside the frame? A: One B: Two C: Three Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: A
- Claude Fable 5.1A · correct
- Claude Haiku 4.5A · correct
- Claude Opus 5A · correct
- Claude Sonnet 5A · correct
- Gemini 3.1 Pro PreviewA · correct
- Gemini 3.5 Flash-LiteA · correct
- Gemini 3.8 FlashA · correct
- GPT-5.6 LunaA · correct
- GPT-5.6 SolA · correct
- GPT-5.6 TerraA · correct
- GPT-6 AstraA · correct
- Muse Spark 1.2A · correct
- Muse Spark 1.3A · correct
- Claude Opus 5.5A · correct
- GPT-6 LunaA · correct
- GPT-6 SolA · correct
- ParetoA · correct
layoutbench-textcolumns-001

Grid columns
17/17 correct
How many columns of cards are in the rectangular grid? Count the cards across one complete row. A: Two B: Three C: Four Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: A
- Claude Fable 5.1A · correct
- Claude Haiku 4.5A · correct
- Claude Opus 5A · correct
- Claude Sonnet 5A · correct
- Gemini 3.1 Pro PreviewA · correct
- Gemini 3.5 Flash-LiteA · correct
- Gemini 3.8 FlashA · correct
- GPT-5.6 LunaA · correct
- GPT-5.6 SolA · correct
- GPT-5.6 TerraA · correct
- GPT-6 AstraA · correct
- Muse Spark 1.2A · correct
- Muse Spark 1.3A · correct
- Claude Opus 5.5A · correct
- GPT-6 LunaA · correct
- GPT-6 SolA · correct
- ParetoA · correct
layoutbench-gridcols-001

Grid rows
17/17 correct
How many rows of cards are in the rectangular grid? Count the cards down one complete column. A: Two B: Three C: Four Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: A
- Claude Fable 5.1A · correct
- Claude Haiku 4.5A · correct
- Claude Opus 5A · correct
- Claude Sonnet 5A · correct
- Gemini 3.1 Pro PreviewA · correct
- Gemini 3.5 Flash-LiteA · correct
- Gemini 3.8 FlashA · correct
- GPT-5.6 LunaA · correct
- GPT-5.6 SolA · correct
- GPT-5.6 TerraA · correct
- GPT-6 AstraA · correct
- Muse Spark 1.2A · correct
- Muse Spark 1.3A · correct
- Claude Opus 5.5A · correct
- GPT-6 LunaA · correct
- GPT-6 SolA · correct
- ParetoA · correct
layoutbench-gridrows-001

Grid gutter comparison
9/17 correct
Compare the horizontal empty gutters between columns with the vertical empty gutters between rows. Measure between the visible card boxes. A: Horizontal gutters are larger B: Vertical gutters are larger C: The two gutter sizes are equal Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: C
- Claude Fable 5.1B · incorrect
- Claude Haiku 4.5A · incorrect
- Claude Opus 5A · incorrect
- Claude Sonnet 5B · incorrect
- Gemini 3.1 Pro PreviewB · incorrect
- Gemini 3.5 Flash-LiteB · incorrect
- Gemini 3.8 FlashB · incorrect
- GPT-5.6 LunaC · correct
- GPT-5.6 SolC · correct
- GPT-5.6 TerraC · correct
- GPT-6 AstraC · correct
- Muse Spark 1.2C · correct
- Muse Spark 1.3C · correct
- Claude Opus 5.5C · correct
- GPT-6 LunaC · correct
- GPT-6 SolC · correct
- ParetoInvalid answer
layoutbench-gridgaps-012

Grid track proportions
16/17 correct
Compare the widths of the three card columns in this grid. A: All three columns have equal widths B: The left column is wider; the other two have equal widths C: The right column is wider; the other two have equal widths Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: B
- Claude Fable 5.1B · correct
- Claude Haiku 4.5A · incorrect
- Claude Opus 5B · correct
- Claude Sonnet 5B · correct
- Gemini 3.1 Pro PreviewB · correct
- Gemini 3.5 Flash-LiteB · correct
- Gemini 3.8 FlashB · correct
- GPT-5.6 LunaB · correct
- GPT-5.6 SolB · correct
- GPT-5.6 TerraB · correct
- GPT-6 AstraB · correct
- Muse Spark 1.2B · correct
- Muse Spark 1.3B · correct
- Claude Opus 5.5B · correct
- GPT-6 LunaB · correct
- GPT-6 SolB · correct
- ParetoB · correct
layoutbench-gridtracks-005

Column span
14/17 correct
Use the three equal-width cards in the lower row as column guides. How many of those column positions does the single card in the upper row cover, including any intervening gutters? A: Three column positions B: One column position C: Two column positions Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: B
- Claude Fable 5.1B · correct
- Claude Haiku 4.5A · incorrect
- Claude Opus 5B · correct
- Claude Sonnet 5B · correct
- Gemini 3.1 Pro PreviewB · correct
- Gemini 3.5 Flash-LiteC · incorrect
- Gemini 3.8 FlashB · correct
- GPT-5.6 LunaC · incorrect
- GPT-5.6 SolB · correct
- GPT-5.6 TerraB · correct
- GPT-6 AstraB · correct
- Muse Spark 1.2B · correct
- Muse Spark 1.3B · correct
- Claude Opus 5.5B · correct
- GPT-6 LunaB · correct
- GPT-6 SolB · correct
- ParetoB · correct
layoutbench-gridspan-001

Wrapped row count
17/17 correct
How many horizontal rows of small cards are visible inside the frame? A: One row B: Three rows C: Two rows Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: A
- Claude Fable 5.1A · correct
- Claude Haiku 4.5A · correct
- Claude Opus 5A · correct
- Claude Sonnet 5A · correct
- Gemini 3.1 Pro PreviewA · correct
- Gemini 3.5 Flash-LiteA · correct
- Gemini 3.8 FlashA · correct
- GPT-5.6 LunaA · correct
- GPT-5.6 SolA · correct
- GPT-5.6 TerraA · correct
- GPT-6 AstraA · correct
- Muse Spark 1.2A · correct
- Muse Spark 1.3A · correct
- Claude Opus 5.5A · correct
- GPT-6 LunaA · correct
- GPT-6 SolA · correct
- ParetoA · correct
layoutbench-wrapcount-001

Incomplete row alignment
16/17 correct
Look only at the shorter final row of cards. How is that row placed horizontally inside the outlined frame? A: Centered, with equal empty space on either side B: Against the right inner edge C: Against the left inner edge Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: C
- Claude Fable 5.1C · correct
- Claude Haiku 4.5A · incorrect
- Claude Opus 5C · correct
- Claude Sonnet 5C · correct
- Gemini 3.1 Pro PreviewC · correct
- Gemini 3.5 Flash-LiteC · correct
- Gemini 3.8 FlashC · correct
- GPT-5.6 LunaC · correct
- GPT-5.6 SolC · correct
- GPT-5.6 TerraC · correct
- GPT-6 AstraC · correct
- Muse Spark 1.2C · correct
- Muse Spark 1.3C · correct
- Claude Opus 5.5C · correct
- GPT-6 LunaC · correct
- GPT-6 SolC · correct
- ParetoC · correct
layoutbench-wrapalign-002

Neighboring gap comparison
15/17 correct
Compare the two empty gaps between the three card boxes, reading left to right for a row or top to bottom for a column. A: The gaps are equal B: The second gap is larger C: The first gap is larger Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: C
- Claude Fable 5.1C · correct
- Claude Haiku 4.5A · incorrect
- Claude Opus 5C · correct
- Claude Sonnet 5C · correct
- Gemini 3.1 Pro PreviewC · correct
- Gemini 3.5 Flash-LiteA · incorrect
- Gemini 3.8 FlashC · correct
- GPT-5.6 LunaC · correct
- GPT-5.6 SolC · correct
- GPT-5.6 TerraC · correct
- GPT-6 AstraC · correct
- Muse Spark 1.2C · correct
- Muse Spark 1.3C · correct
- Claude Opus 5.5C · correct
- GPT-6 LunaC · correct
- GPT-6 SolC · correct
- ParetoC · correct
layoutbench-gapcompare-001

Horizontal versus vertical inset
8/17 correct
The tinted rectangle is centered inside the outlined frame. Compare the horizontal inset (frame to rectangle on either side) with the vertical inset (frame to rectangle above or below). Measure to the rectangle's visible outer edge, not to its text. A: The horizontal inset is larger B: The vertical inset is larger C: The horizontal and vertical insets are equal Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: C
- Claude Fable 5.1B · incorrect
- Claude Haiku 4.5A · incorrect
- Claude Opus 5C · correct
- Claude Sonnet 5B · incorrect
- Gemini 3.1 Pro PreviewB · incorrect
- Gemini 3.5 Flash-LiteB · incorrect
- Gemini 3.8 FlashB · incorrect
- GPT-5.6 LunaC · correct
- GPT-5.6 SolC · correct
- GPT-5.6 TerraA · incorrect
- GPT-6 AstraC · correct
- Muse Spark 1.2B · incorrect
- Muse Spark 1.3A · incorrect
- Claude Opus 5.5C · correct
- GPT-6 LunaC · correct
- GPT-6 SolC · correct
- ParetoC · correct
layoutbench-padcompare-012

Text block placement
17/17 correct
Where is the tinted text block placed horizontally within the outer frame? Judge the block's outer edges, independently of how its text lines align. A: Against the left inner edge B: Centered, with equal empty space on either side C: Against the right inner edge Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: A
- Claude Fable 5.1A · correct
- Claude Haiku 4.5A · correct
- Claude Opus 5A · correct
- Claude Sonnet 5A · correct
- Gemini 3.1 Pro PreviewA · correct
- Gemini 3.5 Flash-LiteA · correct
- Gemini 3.8 FlashA · correct
- GPT-5.6 LunaA · correct
- GPT-5.6 SolA · correct
- GPT-5.6 TerraA · correct
- GPT-6 AstraA · correct
- Muse Spark 1.2A · correct
- Muse Spark 1.3A · correct
- Claude Opus 5.5A · correct
- GPT-6 LunaA · correct
- GPT-6 SolA · correct
- ParetoA · correct
layoutbench-blockalign-001

Relative panel width
17/17 correct
Compare the widths of the two tinted panels, from their visible left edge to right edge. A: Their widths are equal B: The left panel is wider C: The right panel is wider Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: B
- Claude Fable 5.1B · correct
- Claude Haiku 4.5B · correct
- Claude Opus 5B · correct
- Claude Sonnet 5B · correct
- Gemini 3.1 Pro PreviewB · correct
- Gemini 3.5 Flash-LiteB · correct
- Gemini 3.8 FlashB · correct
- GPT-5.6 LunaB · correct
- GPT-5.6 SolB · correct
- GPT-5.6 TerraB · correct
- GPT-6 AstraB · correct
- Muse Spark 1.2B · correct
- Muse Spark 1.3B · correct
- Claude Opus 5.5B · correct
- GPT-6 LunaB · correct
- GPT-6 SolB · correct
- ParetoB · correct
layoutbench-widthcompare-001

Spacing within and between groups
12/17 correct
The cards form two groups marked Harbor and Meadow. Compare the gap between the groups' nearest card boxes with the gaps between cards within each group. Ignore the small group captions. A: The within-group gaps are larger B: The between-group and within-group gaps are equal C: The between-group gap is larger Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: B
- Claude Fable 5.1B · correct
- Claude Haiku 4.5C · incorrect
- Claude Opus 5B · correct
- Claude Sonnet 5C · incorrect
- Gemini 3.1 Pro PreviewB · correct
- Gemini 3.5 Flash-LiteC · incorrect
- Gemini 3.8 FlashB · correct
- GPT-5.6 LunaC · incorrect
- GPT-5.6 SolB · correct
- GPT-5.6 TerraB · correct
- GPT-6 AstraB · correct
- Muse Spark 1.2C · incorrect
- Muse Spark 1.3B · correct
- Claude Opus 5.5B · correct
- GPT-6 LunaB · correct
- GPT-6 SolB · correct
- ParetoB · correct
layoutbench-groupgap-009

Top-level hierarchy
17/17 correct
Look at the two large outlined sections headed Field notes and Harbor log. At the top level, how are these two sections arranged? Ignore the arrangement of cards inside either section. A: Side by side horizontally B: Stacked vertically Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: A
- Claude Fable 5.1A · correct
- Claude Haiku 4.5A · correct
- Claude Opus 5A · correct
- Claude Sonnet 5A · correct
- Gemini 3.1 Pro PreviewA · correct
- Gemini 3.5 Flash-LiteA · correct
- Gemini 3.8 FlashA · correct
- GPT-5.6 LunaA · correct
- GPT-5.6 SolA · correct
- GPT-5.6 TerraA · correct
- GPT-6 AstraA · correct
- Muse Spark 1.2A · correct
- Muse Spark 1.3A · correct
- Claude Opus 5.5A · correct
- GPT-6 LunaA · correct
- GPT-6 SolA · correct
- ParetoA · correct
layoutbench-topflow-001

Nested section flow
17/17 correct
Look inside the large section headed Field notes. How are its small cards arranged? Ignore how Field notes is placed relative to Harbor log. A: One horizontal row B: One vertical column C: Multiple horizontal rows, with a shorter final row Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: A
- Claude Fable 5.1A · correct
- Claude Haiku 4.5A · correct
- Claude Opus 5A · correct
- Claude Sonnet 5A · correct
- Gemini 3.1 Pro PreviewA · correct
- Gemini 3.5 Flash-LiteA · correct
- Gemini 3.8 FlashA · correct
- GPT-5.6 LunaA · correct
- GPT-5.6 SolA · correct
- GPT-5.6 TerraA · correct
- GPT-6 AstraA · correct
- Muse Spark 1.2A · correct
- Muse Spark 1.3A · correct
- Claude Opus 5.5A · correct
- GPT-6 LunaA · correct
- GPT-6 SolA · correct
- ParetoA · correct
layoutbench-nestedflow-001

Nested incomplete row alignment
16/17 correct
Inside Field notes, look only at the shorter final row of cards. How is it placed horizontally inside its own outlined inner frame? Ignore the surrounding page's arrangement. A: Against that frame's left inner edge B: Centered in that frame C: Against that frame's right inner edge Return only a JSON object with one key, "choice", whose value is the selected option letter.
Ground truth and model answers
Correct choice: A
- Claude Fable 5.1A · correct
- Claude Haiku 4.5B · incorrect
- Claude Opus 5A · correct
- Claude Sonnet 5A · correct
- Gemini 3.1 Pro PreviewA · correct
- Gemini 3.5 Flash-LiteA · correct
- Gemini 3.8 FlashA · correct
- GPT-5.6 LunaA · correct
- GPT-5.6 SolA · correct
- GPT-5.6 TerraA · correct
- GPT-6 AstraA · correct
- Muse Spark 1.2A · correct
- Muse Spark 1.3A · correct
- Claude Opus 5.5A · correct
- GPT-6 LunaA · correct
- GPT-6 SolA · correct
- ParetoA · correct
layoutbench-nestedwrap-007
Twenty atomic questions about space
The intended building blocks are qualitative: direction, alignment, counts, relative size, proximity, and explicitly scoped hierarchy. Correctness concerns the rendered arrangement. A screenshot cannot uniquely identify the underlying CSS implementation.
| Family | Question |
|---|---|
| Immediate flow | Are the items arranged in a row or a column? |
| Main-axis distribution | Are items packed at the start, center, or end, or spread with space between, around, or evenly among them? |
| Cross-axis alignment | How do differently sized items align across their row or column? |
| Text justification | Are the paragraph's visible line edges left-aligned, centered, right-aligned, or justified? |
| Text columns | How many side-by-side columns contain the body text? |
| Grid columns | How many columns make up the repeated grid? |
| Grid rows | How many rows make up the repeated grid? |
| Grid gutter comparison | Are the horizontal or vertical spaces between grid items larger, or are they equal? |
| Grid track proportions | Are grid columns equal in width, or is one visibly wider? |
| Column span | How many column positions does the upper card occupy, using the complete lower row as a guide? |
| Wrapped row count | How many lines do the repeated items wrap onto? |
| Incomplete row alignment | How is the incomplete final line of a wrapped collection aligned? |
| Neighboring gap comparison | Which of two visible gaps is larger, or are they equal? |
| Horizontal versus vertical inset | Are the horizontal or vertical insets around the centered content block larger, or are they equal? |
| Text block placement | Where does the tinted text block sit horizontally within its larger frame, independently of its text alignment? |
| Relative panel width | Is the left or right panel wider, or are their widths equal? |
| Spacing within and between groups | Is the space between the two labeled groups larger or smaller than the gaps within them, or equal? |
| Top-level hierarchy | Are the page's top-level regions organized vertically first or horizontally first? |
| Nested section flow | Within the indicated region, do its children form a row, a column, or multiple wrapped rows? |
| Nested incomplete row alignment | Inside the indicated region, is the shorter final row aligned left, centered, or aligned right? |
What this measures
All 17 published configurations share 292/292 release questions. The release contains 250 distinct images. Choice accuracy is exact correctness, with invalid responses retained in the denominator. Each family has its own answer set and uniform-choice reference; no combined score is reported.
Each 1600×1200 image renders an 800×600 CSS-pixel canvas with bundled DejaVu Sans text and a light or dark theme. The model receives the image and question, without HTML or measured box coordinates. Sample text brings the questions closer to interface layouts, while controlled boxes isolate particular spatial relationships. The experiment measures the complete image-input and answering pipeline, including question interpretation.
Grid-row, grid-column, and text-column questions also permit counting or repeated-content cues in this finite corpus, so their scores do not isolate spatial grouping from those cues. There is no evidence here that a model used that shortcut.
Human agreement has not been measured. The questions aim for clear judgments by design-skilled readers, but that aim is not a measured agreement result. Related renderings are correlated, and small score differences need more evidence. These atomic skills are plausible ingredients of richer design understanding; this experiment does not test transfer to higher-level tasks.
Reproduce this comparison
4,964 final responses · 4,967 runner attempt records · 3 infrastructure attempts · 15 invalid answers. Recorded API estimate: $42.22. Estimates use recorded usage and dated rates, not invoice totals. The campaign ledger includes 6 attempt records with conservative cost reservations and 6 unmetered request allowances. A populated ledger total does not mean every request was metered. Family cost is unavailable if any scored response in that model–family pair lacks complete metering.
- Source code and dataset
- Methodology
- Original committed run and raw answers
- Additional committed run · Results JSON · Publication seal
- Additional committed run · Results JSON · Publication seal
- Sealed results and diagnostics
- Publication seal
Dataset fc531bdb2e7c9d10e2074519bd3e0d0f66bc8ec970d20f25cdb628054784c6de
Source 811dbd016689132a51ec1ffe80de7225751af095