FontBench
V1.0.0How well can a model identify a typeface and its styling from an image? FontBench tests font family, category, weight, modifier, letter spacing, and line height using reproducible rendered specimens.
Accuracy and efficiency
Green points mark the best accuracy–cost tradeoffs. Both axes are linear.
Results
Ranked by exact match. Select a model to highlight it.
Exact match requires all six attributes to be correct. Missing measurements appear as —.
| Model | Exact match | Font family | Category | Weight | Modifier | Letter spacing | Line height | API cost / input | Response time |
|---|---|---|---|---|---|---|---|---|---|
| openai · 728 scored | 40.1% | 45.1% | 96.8% | 91.5% | 99.3% | 94.6% | 99.3% | $0.0464 | 20.63s |
| anthropic · 728 scored | 38.9% | 65.1% | 97.3% | 90.5% | 99.7% | 82.8% | 70.5% | $0.0067 | 2.20s |
| anthropic · 728 scored | 13.3% | 42.9% | 96.4% | 82.1% | 99.6% | 80.6% | 40.0% | $0.0174 | 4.89s |
| openai · 728 scored | 12.2% | 19.2% | 94.5% | 87.5% | 98.9% | 79.8% | 94.8% | $0.0130 | 18.86s |
| openai · 728 scored | 8.1% | 16.1% | 90.8% | 83.0% | 97.5% | 70.1% | 92.3% | $0.0180 | 18.07s |
| openai · 728 scored | 4.7% | 12.5% | 88.5% | 80.8% | 85.2% | 65.1% | 84.3% | $0.0065 | 7.73s |
| openai · 728 scored | 4.5% | 20.2% | 93.4% | 79.9% | 90.7% | 51.9% | 52.5% | $0.0238 | 81.69s |
| anthropic · 728 scored | 4.5% | 13.5% | 91.3% | 73.1% | 96.2% | 58.0% | 63.3% | $0.0179 | 12.15s |
| anthropic · 728 scored | 4.3% | 14.7% | 95.5% | 76.9% | 99.2% | 70.1% | 61.0% | $0.0090 | 3.07s |
| google · 728 scored | 4.0% | 14.7% | 83.4% | 77.5% | 97.7% | 55.5% | 61.5% | $0.0121 | 12.32s |
| openai · 728 scored | 4.0% | 10.4% | 91.6% | 82.0% | 96.2% | 64.1% | 82.3% | $0.0009 | 9.09s |
| openai · 728 scored | 3.7% | 11.7% | 84.9% | 82.8% | 91.8% | 58.1% | 78.2% | $0.0005 | 15.66s |
| google · 728 scored | 3.2% | 12.5% | 85.0% | 75.3% | 97.7% | 61.5% | 50.0% | $0.0169 | 12.27s |
| openai · 728 scored | 3.2% | 12.1% | 95.1% | 77.3% | 96.4% | 52.2% | 58.7% | $0.0140 | 20.50s |
| anthropic · 728 scored | 1.5% | 7.6% | 80.1% | 73.4% | 99.2% | 46.2% | 58.5% | $0.0033 | 2.08s |
| google · 728 scored | 1.4% | 7.6% | 80.9% | 69.1% | 80.1% | 46.3% | 46.4% | $0.0006 | 0.95s |
| anthropic · 728 scored | 0.4% | 3.4% | 62.1% | 71.3% | 83.7% | 42.9% | 42.2% | $0.0014 | 1.70s |
API cost is the mean metered token cost of scored responses at the recorded rates; it excludes failed attempts and spending reservations. Response times are observed API timings, affected by network and service conditions.
Recognizing a font is still difficult
GPT-6 Astra leads this comparison, identifying all six attributes correctly on 40.1% of inputs. Claude Opus 5.5 follows at 38.9%. Even the leading model identifies the exact font family on only 45.1% of specimens. The joint score is demanding: a single wrong attribute makes that input incorrect.
Use the metric selector to distinguish font recognition from simpler attributes such as category or weight. The cost and latency views expose different tradeoffs; a strong score on one attribute does not imply reliable recovery of the complete typography specification.
These results cover 728 of 1,824 release inputs and 49 of 50 font families. They are descriptive scores on a fixed shared sample. Repeated text, related font variants, synthetic rendering, and model-specific image processing limit generalization. Human agreement has not been measured.
What the models see
Examples from the scored sample, using the same sentence with different fonts and settings.










The six dimensions
FontBench V1.0.0 contains 1,824 inputs across 50 font families. This comparison uses 728 shared inputs from 49 families. The release samples supported combinations; it does not cover every combination.
| Dimension | Values in the release |
|---|---|
| Font family | 50 canonical family names, listed below. Accepted aliases follow the benchmark’s grading rules. |
| Category | serif; non-serif (sans-serif); mono (monospace); handwriting; other (decorative display). Categories follow the benchmark’s family catalog. |
| Weight | CSS numeric weights: thin = 200; regular = 400; bold = 700; black = 900. |
| Modifier | regular (no modifier); native italic; underline; strikethrough; small-caps. Small caps may be synthesized when they visibly change the image. |
| Letter spacing | tight = −0.05em; normal = 0em; loose = +0.12em. The prompt calls this kerning, meaning uniform tracking rather than adjustments to individual letter pairs. |
| Line height | tight = 1.15; normal = 1.45; loose = 1.9 times the fixed 22px font size. |
Exact match counts an input as correct only when all six attributes match. Each attribute’s accuracy is also reported separately.
All 50 font families in V1.0.0
In the full release but outside this comparison: Verdana.
- Arial
- Arvo
- Bebas Neue
- Bitter
- Bodoni Moda
- Bungee
- Caveat
- Cinzel
- Comic Sans MS
- Cormorant Garamond
- Courier New
- Crimson Text
- DM Sans
- Dancing Script
- EB Garamond
- Fira Code
- Fira Sans
- Georgia
- Helvetica
- Impact
- Inter
- JetBrains Mono
- Lato
- Libre Baskerville
- Lobster
- Lora
- Merriweather
- Montserrat
- Noto Sans
- Nunito
- Open Sans
- Oswald
- PT Sans
- PT Serif
- Pacifico
- Playfair Display
- Plus Jakarta Sans
- Poppins
- Raleway
- Roboto
- Roboto Mono
- Rubik
- Shadows Into Light
- Source Code Pro
- Source Sans 3
- Space Mono
- Spectral
- Times New Roman
- Verdana
- Work Sans
Method
Models receive the same image and question, returning six attributes. Font names follow the release’s accepted names and aliases. Invalid answers receive no credit; provider failures are excluded from accuracy scores.
Each input has equal scoring weight. All 17 model configurations use the same 728 completed inputs, covering 49 of the release’s 50 font families. Model settings and response limits are recorded with the results.
Muse Spark 1.3 and 1.2, GPT-6 Sol and Luna, and Claude Opus 5.5 were evaluated on all 1,824 release inputs. Their scores, costs, and response times here use the same 728 inputs as the other models. The complete full-release results remain available below.
Dataset and grading documentationReproduce this publication
Download the original responses, scores, and verification hashes for this version.
Version, source, and verification hashes
- Benchmark
- FontBench V1.0.0; grading 3
- Source commit
- 805e2146e5ee1637d05cd82c0477371f2b9e3228
- Dataset commit
- d69e87e2c206ea75c52f5b8340d677bd14af03e3
- Dataset SHA-256
- d40a3cafa91610bbebaa6c49719f24efcbe754a328e448d3e51f0e1f217c79a6
- Protocol SHA-256
- 3769c8cc383a4959533c4801eec6202035a5a8cd30b76644e8cdfbc11cc859ce
- Shared cohort SHA-256
- 19097ac3b1a2cab662304dcea8f051297779934f2fc35d8d22bc79a1e66ee19e
- Final results SHA-256
- b10fe2532925e7c11ade529bf8048dcc72eb7320876d2cfae6fb65a908f4a50c
- Finalization SHA-256
- b6f38785b9228305cf4b4bb49898177b29034965502f085abf9ea23c85e3a3d0
- Muse source commit
- 234803f81550349b09c3f1ec29e2d285cea6416e
- Muse source run
- results/runs/1.0.0/2026-09-10-muse-spark
- Muse full-release cohort SHA-256 (1,824 inputs)
- d46fb9667177ae3a652f06ded92c09da042a08c3beaee4926ad70778423d44ac
- Muse final results SHA-256
- f9a4329053cb668bb3b507827f614b8aeb58dad55128353ae5a5d6af39212152
- Muse finalization SHA-256
- f874714fbe9cb2aae64bf397bbc1194efcbd05c96cfe991c614d64d10537157c
- September source commit
- 77785942a550502f206ebb04a32a27236c50a660
- September source run
- results/runs/1.0.0/2026-09-23-openai-anthropic
- September full-release cohort SHA-256 (1,824 inputs)
- d46fb9667177ae3a652f06ded92c09da042a08c3beaee4926ad70778423d44ac
- September final results SHA-256
- deb6e941e03bf26b94f2720f3c1aba4e0446eb49e22190c15d1c7d45b2e88d36
- September finalization SHA-256
- 2ac0574683e2936250852591fcd80184d6eb262ff160fbe555c06a166fca42de
Changelog
- September 10, 2026
- Added data for Muse Spark 1.3 and 1.2 to the shared 728-input comparison.
- September 9, 2026
- First V1.0.0 publication: eleven OpenAI, Gemini, and Claude configurations scored on the same 728 inputs.