Benchmarks
What can a vision model actually see?
These experiments isolate small pieces of visual understanding: a typeface, a border, a color, a gap. Each report pairs model results with the images, scoring rules, and source data needed to inspect the experiment.
Can a model read the type?
Can a model see the edge?
Reading the results
These are reproducible experiments on specific synthetic images. They measure a model’s complete image-input and response pipeline, including its ability to follow the question. Human agreement has not been measured, and results do not establish a general ranking of visual intelligence.
Compare models within a report using the same scored inputs. The task families, sample sizes, and scoring rules differ across benchmarks, so their percentages are not interchangeable. Each page explains its coverage and limitations and links to immutable results.

