FontBench

All benchmarks

FontBench

V1.0.0

How well can a model identify a typeface and its styling from an image? FontBench tests font family, category, weight, modifier, letter spacing, and line height using reproducible rendered specimens.

Accuracy and efficiency

Explain Dimensions

Green points mark the best accuracy–cost tradeoffs. Both axes are linear.

Accuracy and efficiency0%20%40%60%80%100%$0$0.010$0.020$0.030$0.040$0.050Metered API cost per scored input (USD) →Exact-match accuracy (%) →GPT-6 Luna: 3.7% ($0.0005 / input)GPT-6 LunaGemini 3.5 Flash-Lite: 1.4% ($0.0006 / input)Gemini 3.5 Flash-LiteGPT-5.6 Luna: 4.0% ($0.0009 / input)GPT-5.6 LunaClaude Haiku 4.5: 0.4% ($0.0014 / input)Claude Haiku 4.5Claude Sonnet 5: 1.5% ($0.0033 / input)Claude Sonnet 5GPT-5.6 Terra: 4.7% ($0.0065 / input)GPT-5.6 TerraClaude Opus 5.5: 38.9% ($0.0067 / input)Claude Opus 5.5Claude Opus 5: 4.3% ($0.0090 / input)Claude Opus 5Gemini 3.8 Flash: 4.0% ($0.0121 / input)Gemini 3.8 FlashGPT-6 Sol: 12.2% ($0.0130 / input)GPT-6 SolMuse Spark 1.2: 3.2% ($0.0140 / input)Muse Spark 1.2Gemini 3.1 Pro Preview: 3.2% ($0.0169 / input)Gemini 3.1 Pro PreviewClaude Fable 5.1: 13.3% ($0.0174 / input)Claude Fable 5.1Pareto: 4.5% ($0.0179 / input)ParetoGPT-5.6 Sol: 8.1% ($0.0180 / input)GPT-5.6 SolMuse Spark 1.3: 4.5% ($0.0238 / input)Muse Spark 1.3GPT-6 Astra: 40.1% ($0.0464 / input)GPT-6 Astra
Swipe chart horizontally to explore →

Results

Ranked by exact match. Select a model to highlight it.

Exact match requires all six attributes to be correct. Missing measurements appear as —.

FontBench model results on identical shared inputs. All scores are percentages.
ModelExact matchFont familyCategoryWeightModifierLetter spacingLine heightAPI cost / inputResponse time
openai · 728 scored40.1%45.1%96.8%91.5%99.3%94.6%99.3%$0.046420.63s
anthropic · 728 scored38.9%65.1%97.3%90.5%99.7%82.8%70.5%$0.00672.20s
anthropic · 728 scored13.3%42.9%96.4%82.1%99.6%80.6%40.0%$0.01744.89s
openai · 728 scored12.2%19.2%94.5%87.5%98.9%79.8%94.8%$0.013018.86s
openai · 728 scored8.1%16.1%90.8%83.0%97.5%70.1%92.3%$0.018018.07s
openai · 728 scored4.7%12.5%88.5%80.8%85.2%65.1%84.3%$0.00657.73s
openai · 728 scored4.5%20.2%93.4%79.9%90.7%51.9%52.5%$0.023881.69s
anthropic · 728 scored4.5%13.5%91.3%73.1%96.2%58.0%63.3%$0.017912.15s
anthropic · 728 scored4.3%14.7%95.5%76.9%99.2%70.1%61.0%$0.00903.07s
google · 728 scored4.0%14.7%83.4%77.5%97.7%55.5%61.5%$0.012112.32s
openai · 728 scored4.0%10.4%91.6%82.0%96.2%64.1%82.3%$0.00099.09s
openai · 728 scored3.7%11.7%84.9%82.8%91.8%58.1%78.2%$0.000515.66s
google · 728 scored3.2%12.5%85.0%75.3%97.7%61.5%50.0%$0.016912.27s
openai · 728 scored3.2%12.1%95.1%77.3%96.4%52.2%58.7%$0.014020.50s
anthropic · 728 scored1.5%7.6%80.1%73.4%99.2%46.2%58.5%$0.00332.08s
google · 728 scored1.4%7.6%80.9%69.1%80.1%46.3%46.4%$0.00060.95s
anthropic · 728 scored0.4%3.4%62.1%71.3%83.7%42.9%42.2%$0.00141.70s

API cost is the mean metered token cost of scored responses at the recorded rates; it excludes failed attempts and spending reservations. Response times are observed API timings, affected by network and service conditions.

Recognizing a font is still difficult

GPT-6 Astra leads this comparison, identifying all six attributes correctly on 40.1% of inputs. Claude Opus 5.5 follows at 38.9%. Even the leading model identifies the exact font family on only 45.1% of specimens. The joint score is demanding: a single wrong attribute makes that input incorrect.

Use the metric selector to distinguish font recognition from simpler attributes such as category or weight. The cost and latency views expose different tradeoffs; a strong score on one attribute does not imply reliable recovery of the complete typography specification.

These results cover 728 of 1,824 release inputs and 49 of 50 font families. They are descriptive scores on a fixed shared sample. Repeated text, related font variants, synthetic rendering, and model-specific image processing limit generalization. Human agreement has not been measured.

What the models see

Examples from the scored sample, using the same sentence with different fonts and settings.

Caveat typography specimen, task font-caveat-v16
Caveathandwriting · 320px layoutfont-caveat-v16
Comic Sans MS typography specimen, task font-comic-sans-ms-v18
Comic Sans MShandwriting · 220px layoutfont-comic-sans-ms-v18
Courier New typography specimen, task font-courier-new-v18
Courier Newmono · 220px layoutfont-courier-new-v18
Fira Code typography specimen, task font-fira-code-v16
Fira Codemono · 320px layoutfont-fira-code-v16
Arial typography specimen, task font-arial-v17
Arialnon-serif · 440px layoutfont-arial-v17
DM Sans typography specimen, task font-dm-sans-v02
DM Sansnon-serif · 320px layoutfont-dm-sans-v02
Bebas Neue typography specimen, task font-bebas-neue-v22
Bebas Neueother · 320px layoutfont-bebas-neue-v22
Bungee typography specimen, task font-bungee-v16
Bungeeother · 320px layoutfont-bungee-v16
Arvo typography specimen, task font-arvo-v21
Arvoserif · 220px layoutfont-arvo-v21
Bitter typography specimen, task font-bitter-v02
Bitterserif · 320px layoutfont-bitter-v02

The six dimensions

FontBench V1.0.0 contains 1,824 inputs across 50 font families. This comparison uses 728 shared inputs from 49 families. The release samples supported combinations; it does not cover every combination.

Full range of the six tested typographic attributes
DimensionValues in the release
Font family50 canonical family names, listed below. Accepted aliases follow the benchmark’s grading rules.
Categoryserif; non-serif (sans-serif); mono (monospace); handwriting; other (decorative display). Categories follow the benchmark’s family catalog.
WeightCSS numeric weights: thin = 200; regular = 400; bold = 700; black = 900.
Modifierregular (no modifier); native italic; underline; strikethrough; small-caps. Small caps may be synthesized when they visibly change the image.
Letter spacingtight = −0.05em; normal = 0em; loose = +0.12em. The prompt calls this kerning, meaning uniform tracking rather than adjustments to individual letter pairs.
Line heighttight = 1.15; normal = 1.45; loose = 1.9 times the fixed 22px font size.

Exact match counts an input as correct only when all six attributes match. Each attribute’s accuracy is also reported separately.

All 50 font families in V1.0.0

In the full release but outside this comparison: Verdana.

  • Arial
  • Arvo
  • Bebas Neue
  • Bitter
  • Bodoni Moda
  • Bungee
  • Caveat
  • Cinzel
  • Comic Sans MS
  • Cormorant Garamond
  • Courier New
  • Crimson Text
  • DM Sans
  • Dancing Script
  • EB Garamond
  • Fira Code
  • Fira Sans
  • Georgia
  • Helvetica
  • Impact
  • Inter
  • JetBrains Mono
  • Lato
  • Libre Baskerville
  • Lobster
  • Lora
  • Merriweather
  • Montserrat
  • Noto Sans
  • Nunito
  • Open Sans
  • Oswald
  • PT Sans
  • PT Serif
  • Pacifico
  • Playfair Display
  • Plus Jakarta Sans
  • Poppins
  • Raleway
  • Roboto
  • Roboto Mono
  • Rubik
  • Shadows Into Light
  • Source Code Pro
  • Source Sans 3
  • Space Mono
  • Spectral
  • Times New Roman
  • Verdana
  • Work Sans

Method

Models receive the same image and question, returning six attributes. Font names follow the release’s accepted names and aliases. Invalid answers receive no credit; provider failures are excluded from accuracy scores.

Each input has equal scoring weight. All 17 model configurations use the same 728 completed inputs, covering 49 of the release’s 50 font families. Model settings and response limits are recorded with the results.

Muse Spark 1.3 and 1.2, GPT-6 Sol and Luna, and Claude Opus 5.5 were evaluated on all 1,824 release inputs. Their scores, costs, and response times here use the same 728 inputs as the other models. The complete full-release results remain available below.

Dataset and grading documentation
Version, source, and verification hashes
Benchmark
FontBench V1.0.0; grading 3
Source commit
805e2146e5ee1637d05cd82c0477371f2b9e3228
Dataset commit
d69e87e2c206ea75c52f5b8340d677bd14af03e3
Dataset SHA-256
d40a3cafa91610bbebaa6c49719f24efcbe754a328e448d3e51f0e1f217c79a6
Protocol SHA-256
3769c8cc383a4959533c4801eec6202035a5a8cd30b76644e8cdfbc11cc859ce
Shared cohort SHA-256
19097ac3b1a2cab662304dcea8f051297779934f2fc35d8d22bc79a1e66ee19e
Final results SHA-256
b10fe2532925e7c11ade529bf8048dcc72eb7320876d2cfae6fb65a908f4a50c
Finalization SHA-256
b6f38785b9228305cf4b4bb49898177b29034965502f085abf9ea23c85e3a3d0
Muse source commit
234803f81550349b09c3f1ec29e2d285cea6416e
Muse source run
results/runs/1.0.0/2026-09-10-muse-spark
Muse full-release cohort SHA-256 (1,824 inputs)
d46fb9667177ae3a652f06ded92c09da042a08c3beaee4926ad70778423d44ac
Muse final results SHA-256
f9a4329053cb668bb3b507827f614b8aeb58dad55128353ae5a5d6af39212152
Muse finalization SHA-256
f874714fbe9cb2aae64bf397bbc1194efcbd05c96cfe991c614d64d10537157c
September source commit
77785942a550502f206ebb04a32a27236c50a660
September source run
results/runs/1.0.0/2026-09-23-openai-anthropic
September full-release cohort SHA-256 (1,824 inputs)
d46fb9667177ae3a652f06ded92c09da042a08c3beaee4926ad70778423d44ac
September final results SHA-256
deb6e941e03bf26b94f2720f3c1aba4e0446eb49e22190c15d1c7d45b2e88d36
September finalization SHA-256
2ac0574683e2936250852591fcd80184d6eb262ff160fbe555c06a166fca42de

Changelog

September 10, 2026
Added data for Muse Spark 1.3 and 1.2 to the shared 728-input comparison.
September 9, 2026
First V1.0.0 publication: eleven OpenAI, Gemini, and Claude configurations scored on the same 728 inputs.