BorderBench
How well can a model read the edges of a UI card? BorderBench tests border presence, sides, line style, width, corner radius, corner uniformity, and shadows using reproducible rendered specimens.
Accuracy and efficiency
Green points mark the best accuracy–cost tradeoffs. Both axes are linear.
128
Shared scored inputs
52
Recipes in this comparison
13
Model configurations
477
Inputs in the full release
Results
Ranked by exact match. Select a model to highlight it.
Exact match requires all seven attributes to be correct. Missing measurements appear as —.
| Model | Exact match | Presence | Sides | Style | Width | Radius | Uniformity | Shadow | API cost / input | Response time |
|---|---|---|---|---|---|---|---|---|---|---|
| GPT · 128 scored | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | $0.0241 | 3.00s |
| GPT · 128 scored | 79.7% | 100.0% | 98.4% | 100.0% | 98.4% | 89.1% | 95.3% | 97.7% | $0.0125 | 5.27s |
| Gemini · 128 scored | 76.6% | 98.4% | 97.7% | 98.4% | 94.5% | 96.9% | 97.7% | 84.4% | $0.0186 | 12.68s |
| GPT · 128 scored | 76.6% | 97.7% | 93.8% | 96.1% | 94.5% | 90.6% | 93.0% | 99.2% | $0.0007 | 3.97s |
| GPT · 128 scored | 70.3% | 99.2% | 98.4% | 97.7% | 93.8% | 87.5% | 89.8% | 97.7% | $0.0065 | 3.41s |
| Muse · 128 scored | 68.0% | 99.2% | 98.4% | 99.2% | 96.1% | 93.8% | 92.2% | 83.6% | $0.0164 | 41.41s |
| Muse · 128 scored | 53.9% | 100.0% | 99.2% | 100.0% | 80.5% | 97.7% | 100.0% | 68.8% | $0.0146 | 29.00s |
| Gemini · 128 scored | 53.1% | 98.4% | 98.4% | 98.4% | 85.9% | 75.8% | 98.4% | 84.4% | $0.0064 | 5.46s |
| Claude · 128 scored | 51.6% | 98.4% | 98.4% | 98.4% | 93.0% | 93.0% | 97.7% | 60.9% | $0.0342 | 3.83s |
| Claude · 128 scored | 23.4% | 98.4% | 97.7% | 98.4% | 88.3% | 64.1% | 94.5% | 50.0% | $0.0214 | 4.86s |
| Gemini · 128 scored | 22.7% | 83.6% | 75.8% | 80.5% | 53.1% | 82.8% | 87.5% | 67.2% | $0.0008 | 1.04s |
| Claude · 128 scored | 10.2% | 90.6% | 83.6% | 84.4% | 40.6% | 68.0% | 83.6% | 53.1% | $0.0029 | 1.82s |
| Claude · 128 scored | 7.8% | 94.5% | 85.9% | 94.5% | 44.5% | 62.5% | 85.9% | 43.8% | $0.0071 | 2.28s |
API cost is the mean metered token cost of scored responses at the recorded rates; it excludes separate infrastructure attempts and spending reservations. Response times are observed API timings, affected by network and service conditions.
The hard parts are scale and subtlety
GPT-6 Astra recovers all seven attributes on 100.0% of the shared sample, followed by GPT-5.6 Sol at 79.7%. A perfect score here applies to these 128 images; it does not establish a ceiling on the full 477-image release or on real interfaces.
The archived confusion matrices reveal errors that a headline average hides. Muse Spark 1.3 underestimates width in all 25 of its width mistakes. Claude Fable 5.1 labels 30 of 43 subtle shadows as absent. Claude Haiku 4.5 predicts uniform corners on every image: that earns 107/128 on this attribute while missing all 21 nonuniform examples.
Read per-class counts alongside aggregate accuracy. The sample contains related renders of the same geometry and only sparse complete matched sets, so it cannot establish independent effects of every border, corner, and shadow setting. Balanced accuracy and matched-set coverage are included in the downloadable results.
What the models see
Original images from the scored sample, selected to show the range of labels. Every 400×240px card contains the same 16px DejaVu Sans text at 24px line height and a 64px horizontal guide. These visible references preserve relative scale when an image is resized. Captions below show the ground truth; models receive the image and the common question.








The seven dimensions
BorderBench uses coarse, explicitly defined anchors from Tailwind CSS 3.4.17, with a 16px root size. These are rendered prototypes: the test does not ask models to put arbitrary intermediate values into hidden bins.
| Dimension | Values in the release |
|---|---|
| Border presence | true or false: a visible stroke along the card boundary. A contrasting surface or a drop shadow alone is not a border. |
| Border sides | all-4; bottom-only; left-only; top-only; none. Right-only and two-edge borders are outside this release. |
| Stroke style | solid; dashed; dotted; none. These name the visible line pattern. |
| Stroke width | 0px; 1px; 2px; 4px; 8px. Tailwind anchors: border-0, border, border-2, border-4, and border-8. |
| Corner radius | sharp = 0px (rounded-none); subtle = 4px (rounded); medium = 12px (rounded-xl); large = 24px (rounded-3xl); pill = rounded-full, a CSS-clamped capsule. |
| Corner uniformity | all-corners; top-only (two rounded top corners); asymmetric (three rounded corners and a sharp bottom-left corner). Nonuniform examples use 12px or 24px rounded corners. |
| Elevation | none = shadow-none; subtle-drop = shadow; floating-drop = shadow-lg. Each shadow level occurs with every geometry and theme in the full release, including borderless cards. |
Exact match counts an input as correct only when all seven attributes match. Each attribute’s accuracy is also reported separately. Presence, sides, style, and width share absence semantics, so these are not seven independent measures of perception.
The full release crosses 53 geometry recipes with three light themes and three shadow levels, yielding 477 images. Matched sets vary width, style, sides, radius, or uniformity while holding the other applicable properties fixed. The same geometry repeats across themes; those repeats are correlated observations. This partial comparison includes 52 recipes and every label, with uneven class counts and few complete matched sets.
Method
Models receive the same image and question, returning seven attributes in a strict JSON schema. Invalid answers receive no credit; provider failures are excluded from accuracy until resolved. Each input has equal scoring weight. Model settings and response limits are recorded with the results.
This is a partial comparison: all 13 configurations completed the same 128 inputs from the 477-image release, using a fixed shuffled order. The first batch recorded 1,664 responses with zero invalid answers and zero provider failures, at a recorded cost of $21.28.
Scores describe this shared sample. The task combines visual recognition, interpretation of the supplied scale, and following the response contract. Human agreement has not been measured. Synthetic cards, one font and layout, three light themes, and one browser configuration limit generalization to other interfaces.
Dataset and grading documentationReproduce this publication
Download the original responses, scores, and verification hashes for this version. The finalized results include per-class counts, confusion matrices, balanced accuracy, and coverage of complete matched sets. The source and run links below are pinned to the published commit.
Version, source, and verification hashes
- Benchmark
- BorderBench V1.2.0; grading 3
- Source commit
- 3c646bae10ba663c2901759cf675cc173e859096
- Source run
- results/runs/1.2.0/2026-09-12-first-campaign
- Dataset commit
- b8c3a6b84bc250b3a22739bfec5e4fed47bbcffc
- Dataset SHA-256
- 85a9c8d4d23a6b2f9d7a3c15bcc8bc1031980139595d2dcb45df1262478c7451
- Protocol SHA-256
- d22d04185da1464fe0902cb563bc020c9d3f836d02cdf864062d28db88702be4
- Shared cohort SHA-256
- 4b8399796335fd11f290ae4fc9c1c47c4328ab1668b46607866fe70b2a233ed6
- Final results SHA-256
- 4960780ed478d066311ca3e913474b85a0b02d8dced2c14c5d369ef25915a717
- Finalization SHA-256
- f14aaee6e9f4ca0282eeea0af6f99ed36b53c6773d589fcdf9ed8e078a0b72a9
Changelog
- September 12, 2026
- First V1.2.0 publication: thirteen OpenAI, Gemini, Claude, and Muse configurations scored on the same 128 inputs. The revised 477-image release adds explicit scale references, coarse Tailwind anchors, matched geometry sets, and independent shadow levels. Results from earlier releases are not directly comparable.