ColorBench
Research pilotCan a model match a color, notice a subtle difference, and recover its numbers? ColorBench asks those questions directly, using simple rendered colors instead of objects or scenes. Choice questions are scored right or wrong; a numeric estimate counts as correct when it lands within ΔEOK 0.01 of the target color — ΔEOK is Euclidean distance in the OKLab color space, so this is an engineering tolerance, not a human visibility threshold.
Perception and efficiency
Compare one task family at a time. Scores, API cost, and response time use the same inputs within that family.
Green points mark the best accuracy–cost tradeoffs. Both axes are linear. Models without this measurement are omitted.
Accuracy is the share of correct choices, including invalid answers in the denominator. Chance accuracy for this family is 25%.
Color matching results
Ranked within this family. Select a model to highlight it and inspect its diagnostics.
| Model | Accuracy | Scored | Valid | Invalid | API cost / input | Response time |
|---|---|---|---|---|---|---|
| 100.0% | 48 | 48 | 0 | $0.0085 | 2.64s | |
| 95.8% | 48 | 48 | 0 | $0.0042 | 1.85s | |
| 89.6% | 48 | 48 | 0 | $0.0024 | 2.58s | |
| 87.5% | 48 | 48 | 0 | $0.0104 | 4.29s | |
| 87.5% | 48 | 48 | 0 | $0.0060 | 5.30s | |
| 85.4% | 48 | 48 | 0 | $0.0004 | 3.31s | |
| 85.4% | 48 | 48 | 0 | $0.0002 | 3.54s | |
| 83.3% | 48 | 48 | 0 | $0.0033 | 4.65s | |
| 77.1% | 48 | 48 | 0 | $0.0073 | 3.70s | |
| 77.1% | 48 | 48 | 0 | — | 14.46s | |
| 75.0% | 48 | 48 | 0 | $0.0025 | 2.33s | |
| 72.9% | 48 | 48 | 0 | $0.0072 | 9.45s | |
| 68.8% | 48 | 48 | 0 | $0.0055 | 8.85s | |
| 68.8% | 48 | 48 | 0 | $0.0085 | 24.14s | |
| 52.1% | 48 | 47 | 1 | $0.0110 | 6.24s | |
| 47.9% | 48 | 48 | 0 | $0.0009 | 0.86s | |
| 47.9% | 48 | 48 | 0 | $0.0004 | 1.67s |
Unknown measurements appear as —. Invalid answers count as incorrect. API costs use recorded token usage and rates, excluding separate infrastructure attempts. Response times depend on network and service conditions.
All twelve families
Each column measures a different task. There is no combined overall score. A numeric estimate counts as correct when its color lands within ΔEOK 0.01 of the target. Choice columns show accuracy; numeric columns show the share of estimates that meet that rule.
| Model | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 87.5% | 96.9% | 98.4% | 100.0% | 85.4% | 100.0% | 83.3% | 95.0% | 62.5% | 100.0% | 62.5% | 56.3% |
| Claude Haiku 4.5 | 47.9% | 65.6% | 82.8% | 56.3% | 41.7% | 87.5% | 50.0% | 37.5% | 34.4% | 12.5% | 6.3% | 0.0% |
| Claude Opus 5 | 77.1% | 95.3% | 100.0% | 90.6% | 77.1% | 100.0% | 75.0% | 83.8% | 67.2% | 81.3% | 75.0% | 87.5% |
| Claude Sonnet 5 | 75.0% | 89.1% | 92.2% | 81.3% | 68.8% | 100.0% | 58.3% | 65.0% | 39.1% | 75.0% | 12.5% | 12.5% |
| Gemini 3.1 Pro Preview | 77.1% | 96.9% | 100.0% | 93.8% | 75.0% | 100.0% | 50.0% | 76.3% | 67.2% | 93.8% | 68.8% | 18.8% |
| Gemini 3.5 Flash-Lite | 47.9% | 82.8% | 93.8% | 75.0% | 47.9% | 100.0% | 50.0% | 41.3% | 28.1% | 43.8% | 18.8% | 6.3% |
| Gemini 3.8 Flash | 68.8% | 98.4% | 96.9% | 90.6% | 70.8% | 100.0% | 62.5% | 72.5% | 50.0% | 31.3% | 18.8% | 18.8% |
| GPT-5.6 Luna | 85.4% | 100.0% | 100.0% | 93.8% | 81.3% | 100.0% | 81.3% | 91.3% | 65.6% | 43.8% | 43.8% | 50.0% |
| GPT-5.6 Sol | 87.5% | 92.2% | 100.0% | 100.0% | 87.5% | 100.0% | 79.2% | 92.5% | 70.3% | 37.5% | 37.5% | 37.5% |
| GPT-5.6 Terra | 83.3% | 95.3% | 100.0% | 96.9% | 77.1% | 100.0% | 75.0% | 90.0% | 56.3% | 43.8% | 50.0% | 25.0% |
| GPT-6 Astra | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 75.0% | 87.5% | 93.8% |
| Muse Spark 1.2 | 72.9% | 78.1% | 87.5% | 81.3% | 72.9% | 100.0% | 54.2% | 32.5% | 40.6% | 12.5% | 18.8% | 18.8% |
| Muse Spark 1.3 | 68.8% | 81.3% | 89.1% | 78.1% | 70.8% | 100.0% | 54.2% | 41.3% | 46.9% | 25.0% | 31.3% | 6.3% |
| Claude Opus 5.5 | 95.8% | 100.0% | 100.0% | 96.9% | 93.8% | 100.0% | 81.3% | 93.8% | 79.7% | 100.0% | 75.0% | 100.0% |
| GPT-6 Luna | 85.4% | 98.4% | 100.0% | 100.0% | 83.3% | 100.0% | 85.4% | 98.8% | 76.6% | 37.5% | 56.3% | 68.8% |
| GPT-6 Sol | 89.6% | 98.4% | 100.0% | 100.0% | 91.7% | 100.0% | 91.7% | 97.5% | 79.7% | 12.5% | 12.5% | 6.3% |
| Pareto | 52.1% | 85.9% | 100.0% | 84.4% | 52.1% | 100.0% | 58.3% | 53.8% | 48.4% | 25.0% | 18.8% | 25.0% |
Plain swatch matching scores 625/816 across the roster; component-fill matching scores 613/816. The latter changes both frames and question wording, so it does not isolate a frame effect. Gradients are nearly saturated at 270/272: this small family leaves little room to distinguish models, despite its shared endpoints and histograms.
What stands out
464/464
GPT-6 Astra’s correct choices
Every matching, ordering, and same–different question in this run, across all nine choice families.
41/48
Its correct numeric estimates
RGB, HSL, and OKLCH answers within ΔEOK 0.01 of the target: 12/16 RGB, 14/16 HSL, 15/16 OKLCH. Picking the right visible color is easier than naming its coordinates: no model was correct on all 48 numeric questions, and none of its 16 RGB answers matched every channel exactly.
Across the full roster, larger filled squares helped matching; thicker outlines had mixed results. Colored surroundings made matching harder in the pooled results, and several models repeatedly called different colors “the same.” The matched comparisons below show each of those.
The same colors, different geometry
These pairs hold the palette, reference color, option positions, and question fixed. One comparison changes filled squares from 20px to 84px. The other changes outline strokes from 3px to 12px, keeping their outer width at 84px. Larger fills and thicker outlines are separate interventions.
82 wrong-to-right pairs; 28 right-to-wrong pairs. Correct answers pooled across all 17 models.
46 wrong-to-right pairs; 34 right-to-wrong pairs. Correct answers pooled across all 17 models.
For one palette, Claude Opus 5 answered the larger filled square correctly after missing the smaller one. The outline pair went in the opposite direction. These examples were selected after examining the results. They illustrate independent recorded answers, not a guarantee that the pattern will repeat.

20px filled squares
Opus answered D · Incorrect · Target A
colorbench-smallmatch-01

84px filled squares
Opus answered A · Correct · Target A
colorbench-smallmatch-05

3px outlines
Opus answered A · Correct · Target A
colorbench-smallmatch-09

12px outlines
Opus answered D · Incorrect · Target A
colorbench-smallmatch-13
Surroundings matter too
The context questions keep the interior colors and geometry fixed while changing the colored surrounds. Each colored condition uses the same sixteen reference–option combinations per model.
The same neutral answers appear in all four comparisons; they are not four independent replications. Pooled changes do not describe every model. Each condition has one request, so differences can include ordinary response variation. These results do not identify a human-like color-constancy mechanism.
“Same” is an easy answer
Every model receives equal numbers of identical and differing color pairs. The answers are much less balanced. Recognizing identical colors is common; noticing the chosen differences is harder.
Each image also appears twice with the answer letters reversed: A means “same” in one prompt and “different” in the other. Models made a consistent semantic judgment on 387/408 image–model pairs. Consistency is not correctness: a model that always judges the colors “same” is perfectly consistent and gets only half the questions right.
Read the per-model same–different breakdownMatching a color is not the same as naming its numbers
A numeric question shows a color and asks for RGB, HSL, or OKLCH coordinates. Each format uses the same sixteen target images, and an estimate counts as correct within ΔEOK 0.01 of the target. Across models, 136/272 RGB answers were correct, but only 22/272 matched all three channels exactly. The prompt asks for an estimate; exact recovery is a stricter diagnostic.
Channel-level diagnostics
Channel bands are not the ΔEOK band: 83 of 83 answers within ±2 steps on every channel clear the tolerance, 12 correct answers are off by more than 5 steps on some channel, and 28 answers within ±5 on every channel still miss it.
| Format | Correct · within ΔEOK 0.01 | Exact | Mean ΔEOK ↓ | Similarity /100 ↑ |
|---|---|---|---|---|
| RGB | 136/272 · 50.0% | 22/272 | 0.01376 | 93.12 |
| HSL | 111/272 · 40.8% | — | 0.01820 | 90.90 |
| OKLCH | 101/272 · 37.1% | — | 0.02054 | 89.40 |
ΔEOK measures distance in OKLab; lower is closer. The ΔEOK 0.01 correctness band is an engineering tolerance, not a human visibility threshold. Exact recovery is defined only for RGB, whose channels are whole numbers; HSL and OKLCH targets are derived decimals. Similarity rescales the distance, so it is not an exact-answer percentage. Mean distance uses valid responses only; correctness and similarity count invalid answers as incorrect. 15 valid OKLCH estimates fell outside sRGB and were scored without clipping. Format differences can reflect visual estimation, coordinate knowledge, and response variation.
What the models see
Original 800×640 sRGB images from this comparison. Matching, hue, binding, gradient, context, and small-region tasks show a reference R beside the lettered options; lightness and chroma tasks show a labeled dimension strip above two patches; same–different shows only the two patches; numeric tasks ask for an unreferenced color estimate. Options use neutral letter labels; models receive image content, not the underlying color values.

Color matching
colorbench-matching-01
Question and ground truth
Which swatch, A, B, C, or D, has the same flat fill color as reference R? Compare the colored interiors, not labels or borders. Return only a JSON object with one key, "choice", whose value is the selected option letter.
{"choice":"A"}

Lightness
colorbench-lightness-01
Question and ground truth
Which patch is lighter, A or B? Compare their flat interior colors on the shared neutral background. Return only a JSON object with one key, "choice", whose value is the selected option letter.
{"choice":"B"}

Chroma
colorbench-chroma-01
Question and ground truth
Which patch, A or B, is more colorful? The gray-to-color examples show the intended dimension: more colorful means farther from gray along that example, not lighter or darker. Return only a JSON object with one key, "choice", whose value is the selected option letter.
{"choice":"B"}

Hue
colorbench-hue-01
Question and ground truth
Which option, A, B, C, or D, has the same hue as reference R? Its lightness and colorfulness may differ. Compare the kind of color, not how light, dark, or gray it is. Return only a JSON object with one key, "choice", whose value is the selected option letter.
{"choice":"A"}

Color binding
colorbench-binding-01
Question and ground truth
Which component, A, B, C, or D, has the same flat interior fill color as reference R? Ignore its neutral frame, label, and surrounding content. Return only a JSON object with one key, "choice", whose value is the selected option letter.
{"choice":"A"}

Gradients
colorbench-gradient-01
Question and ground truth
Which strip, A, B, C, or D, has exactly the same left-to-right color progression as reference R? Match the entire visible color field, including where transitions occur, not just its set of colors. Return only a JSON object with one key, "choice", whose value is the selected option letter.
{"choice":"A"}

Same–different
colorbench-samediff-01
Question and ground truth
Do patches A and B have exactly the same flat fill color? Answer "A" for same and "B" for different. Compare the colored interiors, not labels or borders. Return only a JSON object with one key, "choice", whose value is the selected option letter.
{"choice":"A"}

Color in context
colorbench-context-01
Question and ground truth
Which option, A, B, C, or D, has the same flat interior fill color as reference R? Compare the colored interiors; ignore the surrounding colors, labels, and borders. Return only a JSON object with one key, "choice", whose value is the selected option letter.
{"choice":"A"}

Small-region matching
colorbench-smallmatch-01
Question and ground truth
Which swatch, A, B, C, or D, has the same color as reference R? Compare the colored regions, including colored outlines; ignore labels and neutral surroundings. Return only a JSON object with one key, "choice", whose value is the selected option letter.
{"choice":"A"}

RGB reconstruction
colorbench-rgb-01
Question and ground truth
Estimate the flat interior color of R as gamma-encoded sRGB. Return only JSON with keys "r", "g", and "b", each a JSON integer from 0 to 255 (no decimal fractions). Do not report alpha or hex. The target is opaque; estimate its visible fill, not its surroundings.
{"rgb":[192,64,80]}

HSL reconstruction
colorbench-hsl-01
Question and ground truth
Estimate the flat interior color of R in HSL derived from gamma-encoded sRGB. Return only JSON with numeric keys "h" (hue in degrees, 0 inclusive to 360 exclusive), "s" (saturation percent, 0–100), and "l" (lightness percent, 0–100). Use h=0 for an achromatic gray. The target is opaque; estimate its visible fill, not its surroundings.
{"rgb":[192,64,80]}

OKLCH reconstruction
colorbench-oklch-01
Question and ground truth
Estimate the flat interior color of R in OKLCH using the D65 white point. Return only JSON with numeric keys "l" (lightness, 0–1), "c" (nonnegative chroma in OKLCH units, not percent), and "h" (hue in degrees, 0 inclusive to 360 exclusive). Use h=0 for an achromatic gray. The target is opaque; estimate its visible fill, not its surroundings.
{"rgb":[192,64,80]}
The twelve task families
| Family | What is tested | Tasks |
|---|---|---|
| Color matching | Match a reference to one of four flat swatches across four distractor separations. Each balanced group repeats identical options and changes the reference so every answer is correct once. | 48 |
| Lightness | Choose the lighter or darker patch across four separations. Intermediate colors appear with both lower and higher partners, so a color cannot always be assigned the same answer. | 64 |
| Chroma | Choose the more or less colorful patch. Intermediate colors appear with both lower and higher partners; visible gray-to-color examples define the intended dimension. | 64 |
| Hue | Match the reference hue across four separations. Sixteen questions keep lightness and chroma fixed, permitting exact color matching; sixteen vary them. These are coordinate-defined targets, not measured human equivalences. | 32 |
| Color binding | Match a component’s interior fill within a neutral UI. Paired swatch tasks share the colors and references, but both the neutral frames and question wording change. | 48 |
| Gradients | Match a left-to-right color progression. Options share endpoints and color histograms while their interior order differs. This small set tests finite permutations, not arbitrary gradients. | 16 |
| Same–different | Judge whether two patches have exactly the same fill. Each image appears under both answer mappings. Same and different trials reuse individual colors so neither patch alone reveals the answer. | 48 |
| Color in context | Match a reference to an option’s interior fill. The same colors and geometry appear on neutral surrounds and four colored surround conditions, with the question held fixed. | 80 |
| Small-region matching | Match colored regions with the palette, centers, and question held fixed: 20px versus 84px filled squares, and 3px versus 12px outlines with the same 84px outer width. The instructions explicitly include colored outlines. | 64 |
| RGB reconstruction | Estimate the reference’s gamma-encoded sRGB components on the 0–255 scale. Correct within ΔEOK 0.01 of the target. | 16 |
| HSL reconstruction | Estimate hue in degrees and saturation/lightness in percent, using HSL derived from gamma-encoded sRGB. Correct within ΔEOK 0.01 of the target. | 16 |
| OKLCH reconstruction | Estimate D65 OKLCH lightness on the 0–1 scale, chroma in OKLCH units, and hue in degrees. Correct within ΔEOK 0.01 of the target. | 16 |
Multiple-choice answers are scored against the frozen task target. Numeric answers are converted to OKLab and compared with the reference using Euclidean distance, ΔEOK; an estimate within ΔEOK 0.01 counts as correct, and RGB answers that match every channel are also counted as exact. The supplementary numeric similarity is 100 × (1 − min(ΔEOK / 0.2, 1)). The 0.2 ceiling is an engineering normalization, not a just-noticeable-difference threshold. Tight similarity replaces that ceiling with 0.05. The source data records hit rates for ΔEOK bands at 0.005, 0.01, 0.02, 0.05; the ΔEOK 0.01 band is the correctness rule stated above.
Invalid answers remain in the scored denominator and receive zero. Mean, median, and 90th-percentile color errors use valid answers only. Hue component errors are omitted for colors below the recorded chroma threshold of 0.02; the color-distance score still applies. Out-of-sRGB estimates are reported separately.
Method and limits
Human validation has not been performed. These are measurements on a fixed synthetic corpus, not a general ranking of color vision.
All 17 configurations are compared on the same 512 questions over 456 images; the family table gives the question counts per model. The finalized campaigns recorded 8,704 responses at a campaign ledger spending estimate of $53.51. The campaigns include 4 runner attempts with estimated or reserved cost; these campaign totals are not an invoice. Per-input chart costs use metered token usage and exclude unmetered reserves. The displayed comparison contains 9 invalid answers; the campaigns recorded 0 separate infrastructure attempts. An absence of those records does not imply that every HTTP request succeeded. Incomplete metering leaves a chart cost unavailable rather than averaging a subset.
This pilot measures responses through each model’s image-input pipeline. Matching and numeric reconstruction also require attention to the question and response format. RGB, HSL, and OKLCH use the same sixteen target images with different prompts; flat matching and UI binding use the same option groups and corresponding references. Related questions are correlated observations. Provider defaults and output limits differ, so this is not an equal-compute comparison.
Small samples per family provide limited coverage of color space. Human agreement has not been measured. Coordinate-defined lightness, chroma, and hue targets are operational labels, not universal measurements of appearance. Synthetic stimuli, one renderer, and model-specific image processing limit generalization. One response per model and question does not measure repeatability. The tasks test direct color judgments; they do not evaluate aesthetic preference or accessibility compliance.
Four-option groups reuse the same options with different references, limiting a reference-blind strategy to 25% expected accuracy over a complete group. Ordering tasks still permit a 62.5% single-patch lookup baseline. Construction checks are not model ablations, and no grayscale model experiment was run.
Reproduce this publication
The source repository contains frozen stimuli, questions, model configurations, original responses, and finalized family statistics. These links are pinned to the commit used for this publication.
Version, source, and verification hashes
- Pilot
- ColorBench V0.4.0; grading 2
- Finalized
- 2026-09-16T16:34:13.565175+00:00
- Source commit
- 32349bbed203d6e02c6d8ecbbdded840ff47b45c
- Source checkpoint commit
- a04d7ec74cfb8cd739208fcbf228228ae6653a42
- Source run
- results/runs/0.4.0/pilot-20260916
- Dataset commit
- c14b9c971bd1fbe1a0126a2683ba4df8b0eeaf6c
- Dataset SHA-256
- 3c6b248bfc25f4810c0083bbc5af3c5905fb1000795b21fdd0a57969a7aa3bc0
- Protocol SHA-256
- 78be5960423e0ea7b850ffb5a5c07dbb41a4658d6c89dd8a43e9dc08c76ab502
- Shared cohort SHA-256
- c6bb84361fdffea7887f62ebdedbfa9c6d24da6466b5c90b750c0712bed7ecfa
- Final results SHA-256
- c29585261d5ad9390968abca4385b547ca9983399ad7663d04fb12e59b3e4555
- Finalization SHA-256
- 70578278488b1d53a0f578579a6327088d07b12daee35660432e795ceaed85f6
- Additional source commit
- 1ef9b2d2a1ffe037083755c4fd4ed6595dee5e14
- Additional source run
- results/runs/0.4.0/september-models-20260923
- Additional final results SHA-256
- 0e95fc27a315b80dff5b344cd7cd84deca0d90cc41c4dc4cc8841b6cd8c452f8
- Additional finalization SHA-256
- 3ec678af0d946d84a0dd42974cd8c3e6000312a36de283130956f56b1c85b3a1
- Additional source commit
- 1f260292535a4dc35475327bad558a4d8bbef62a
- Additional source run
- results/runs/0.4.0/pareto-20260929
- Additional final results SHA-256
- 386c2ce8b69524e758a1e8a502f9615e3ce9091defb18d16a9a099a4833f9693
- Additional finalization SHA-256
- 75c118d75ca3a0df220e15274a9ff0e2b1f762fa0ae16156d232cfd166eedbfa