BorderBench

All benchmarks

BorderBench

V1.2.0

How well can a model read the edges of a UI card? BorderBench tests border presence, sides, line style, width, corner radius, corner uniformity, and shadows using reproducible rendered specimens.

Accuracy and efficiency

Explain Dimensions

Green points mark the best accuracycost tradeoffs. Both axes are linear.

Accuracy and efficiency0%20%40%60%80%100%$0$0.010$0.020$0.030Metered API cost per scored input (USD) →All seven attributes correct (%) →GPT-5.6 Luna: 76.6% ($0.0007 / input)GPT-5.6 LunaGemini 3.5 Flash-Lite: 22.7% ($0.0008 / input)Gemini 3.5 Flash-LiteClaude Haiku 4.5: 10.2% ($0.0029 / input)Claude Haiku 4.5Gemini 3.8 Flash: 53.1% ($0.0064 / input)Gemini 3.8 FlashGPT-5.6 Terra: 70.3% ($0.0065 / input)GPT-5.6 TerraClaude Sonnet 5: 7.8% ($0.0071 / input)Claude Sonnet 5GPT-5.6 Sol: 79.7% ($0.0125 / input)GPT-5.6 SolMuse Spark 1.3: 53.9% ($0.0146 / input)Muse Spark 1.3Muse Spark 1.2: 68.0% ($0.0164 / input)Muse Spark 1.2Gemini 3.1 Pro Preview: 76.6% ($0.0186 / input)Gemini 3.1 Pro PreviewClaude Opus 5: 23.4% ($0.0214 / input)Claude Opus 5GPT-6 Astra: 100.0% ($0.0241 / input)GPT-6 AstraClaude Fable 5.1: 51.6% ($0.0342 / input)Claude Fable 5.1
Swipe chart horizontally to explore →

128

Shared scored inputs

52

Recipes in this comparison

13

Model configurations

477

Inputs in the full release

Results

Ranked by exact match. Select a model to highlight it.

Exact match requires all seven attributes to be correct. Missing measurements appear as —.

BorderBench model results on identical shared inputs. All scores are percentages.
ModelExact matchPresenceSidesStyleWidthRadiusUniformityShadowAPI cost / inputResponse time
GPT · 128 scored100.0%100.0%100.0%100.0%100.0%100.0%100.0%100.0%$0.02413.00s
GPT · 128 scored79.7%100.0%98.4%100.0%98.4%89.1%95.3%97.7%$0.01255.27s
Gemini · 128 scored76.6%98.4%97.7%98.4%94.5%96.9%97.7%84.4%$0.018612.68s
GPT · 128 scored76.6%97.7%93.8%96.1%94.5%90.6%93.0%99.2%$0.00073.97s
GPT · 128 scored70.3%99.2%98.4%97.7%93.8%87.5%89.8%97.7%$0.00653.41s
Muse · 128 scored68.0%99.2%98.4%99.2%96.1%93.8%92.2%83.6%$0.016441.41s
Muse · 128 scored53.9%100.0%99.2%100.0%80.5%97.7%100.0%68.8%$0.014629.00s
Gemini · 128 scored53.1%98.4%98.4%98.4%85.9%75.8%98.4%84.4%$0.00645.46s
Claude · 128 scored51.6%98.4%98.4%98.4%93.0%93.0%97.7%60.9%$0.03423.83s
Claude · 128 scored23.4%98.4%97.7%98.4%88.3%64.1%94.5%50.0%$0.02144.86s
Gemini · 128 scored22.7%83.6%75.8%80.5%53.1%82.8%87.5%67.2%$0.00081.04s
Claude · 128 scored10.2%90.6%83.6%84.4%40.6%68.0%83.6%53.1%$0.00291.82s
Claude · 128 scored7.8%94.5%85.9%94.5%44.5%62.5%85.9%43.8%$0.00712.28s

API cost is the mean metered token cost of scored responses at the recorded rates; it excludes separate infrastructure attempts and spending reservations. Response times are observed API timings, affected by network and service conditions.

The hard parts are scale and subtlety

GPT-6 Astra recovers all seven attributes on 100.0% of the shared sample, followed by GPT-5.6 Sol at 79.7%. A perfect score here applies to these 128 images; it does not establish a ceiling on the full 477-image release or on real interfaces.

The archived confusion matrices reveal errors that a headline average hides. Muse Spark 1.3 underestimates width in all 25 of its width mistakes. Claude Fable 5.1 labels 30 of 43 subtle shadows as absent. Claude Haiku 4.5 predicts uniform corners on every image: that earns 107/128 on this attribute while missing all 21 nonuniform examples.

Read per-class counts alongside aggregate accuracy. The sample contains related renders of the same geometry and only sparse complete matched sets, so it cannot establish independent effects of every border, corner, and shadow setting. Balanced accuracy and matched-set coverage are included in the downloadable results.

What the models see

Original images from the scored sample, selected to show the range of labels. Every 400×240px card contains the same 16px DejaVu Sans text at 24px line height and a 64px horizontal guide. These visible references preserve relative scale when an image is resized. Captions below show the ground truth; models receive the image and the common question.

BorderBench card: no border; medium radius, all-corners; subtle-drop.
No bordermedium · all-corners · subtle-dropborderbench-002
BorderBench card: 2px solid border, all-4; large radius, asymmetric; floating-drop.
2px solid · all-4large · asymmetric · floating-dropborderbench-378
BorderBench card: 4px dotted border, bottom-only; pill radius, all-corners; none.
4px dotted · bottom-onlypill · all-corners · noneborderbench-418
BorderBench card: 4px dashed border, top-only; sharp radius, all-corners; subtle-drop.
4px dashed · top-onlysharp · all-corners · subtle-dropborderbench-461
BorderBench card: 8px solid border, bottom-only; subtle radius, all-corners; none.
8px solid · bottom-onlysubtle · all-corners · noneborderbench-400
BorderBench card: 1px solid border, all-4; medium radius, all-corners; floating-drop.
1px solid · all-4medium · all-corners · floating-dropborderbench-015
BorderBench card: 2px solid border, left-only; medium radius, all-corners; subtle-drop.
2px solid · left-onlymedium · all-corners · subtle-dropborderbench-128
BorderBench card: no border; medium radius, top-only; subtle-drop.
No bordermedium · top-only · subtle-dropborderbench-311

The seven dimensions

BorderBench uses coarse, explicitly defined anchors from Tailwind CSS 3.4.17, with a 16px root size. These are rendered prototypes: the test does not ask models to put arbitrary intermediate values into hidden bins.

Full range of the seven tested border, corner, and shadow attributes
DimensionValues in the release
Border presencetrue or false: a visible stroke along the card boundary. A contrasting surface or a drop shadow alone is not a border.
Border sidesall-4; bottom-only; left-only; top-only; none. Right-only and two-edge borders are outside this release.
Stroke stylesolid; dashed; dotted; none. These name the visible line pattern.
Stroke width0px; 1px; 2px; 4px; 8px. Tailwind anchors: border-0, border, border-2, border-4, and border-8.
Corner radiussharp = 0px (rounded-none); subtle = 4px (rounded); medium = 12px (rounded-xl); large = 24px (rounded-3xl); pill = rounded-full, a CSS-clamped capsule.
Corner uniformityall-corners; top-only (two rounded top corners); asymmetric (three rounded corners and a sharp bottom-left corner). Nonuniform examples use 12px or 24px rounded corners.
Elevationnone = shadow-none; subtle-drop = shadow; floating-drop = shadow-lg. Each shadow level occurs with every geometry and theme in the full release, including borderless cards.

Exact match counts an input as correct only when all seven attributes match. Each attribute’s accuracy is also reported separately. Presence, sides, style, and width share absence semantics, so these are not seven independent measures of perception.

The full release crosses 53 geometry recipes with three light themes and three shadow levels, yielding 477 images. Matched sets vary width, style, sides, radius, or uniformity while holding the other applicable properties fixed. The same geometry repeats across themes; those repeats are correlated observations. This partial comparison includes 52 recipes and every label, with uneven class counts and few complete matched sets.

Method

Models receive the same image and question, returning seven attributes in a strict JSON schema. Invalid answers receive no credit; provider failures are excluded from accuracy until resolved. Each input has equal scoring weight. Model settings and response limits are recorded with the results.

This is a partial comparison: all 13 configurations completed the same 128 inputs from the 477-image release, using a fixed shuffled order. The first batch recorded 1,664 responses with zero invalid answers and zero provider failures, at a recorded cost of $21.28.

Scores describe this shared sample. The task combines visual recognition, interpretation of the supplied scale, and following the response contract. Human agreement has not been measured. Synthetic cards, one font and layout, three light themes, and one browser configuration limit generalization to other interfaces.

Dataset and grading documentation

Reproduce this publication

Download the original responses, scores, and verification hashes for this version. The finalized results include per-class counts, confusion matrices, balanced accuracy, and coverage of complete matched sets. The source and run links below are pinned to the published commit.

Version, source, and verification hashes
Benchmark
BorderBench V1.2.0; grading 3
Source commit
3c646bae10ba663c2901759cf675cc173e859096
Source run
results/runs/1.2.0/2026-09-12-first-campaign
Dataset commit
b8c3a6b84bc250b3a22739bfec5e4fed47bbcffc
Dataset SHA-256
85a9c8d4d23a6b2f9d7a3c15bcc8bc1031980139595d2dcb45df1262478c7451
Protocol SHA-256
d22d04185da1464fe0902cb563bc020c9d3f836d02cdf864062d28db88702be4
Shared cohort SHA-256
4b8399796335fd11f290ae4fc9c1c47c4328ab1668b46607866fe70b2a233ed6
Final results SHA-256
4960780ed478d066311ca3e913474b85a0b02d8dced2c14c5d369ef25915a717
Finalization SHA-256
f14aaee6e9f4ca0282eeea0af6f99ed36b53c6773d589fcdf9ed8e078a0b72a9

Changelog

September 12, 2026
First V1.2.0 publication: thirteen OpenAI, Gemini, Claude, and Muse configurations scored on the same 128 inputs. The revised 477-image release adds explicit scale references, coarse Tailwind anchors, matched geometry sets, and independent shadow levels. Results from earlier releases are not directly comparable.