FontBench

Fontbench

V1.0.0

How well can a model identify a typeface and its styling from an image? FontBench tests font family, category, weight, modifier, letter spacing, and line height using reproducible rendered specimens.

Accuracy and efficiency

Explain Dimensions

Green points mark the best accuracy–cost tradeoffs. Both axes are linear.

Accuracy and efficiency0%20%40%60%80%100%$0$0.010$0.020$0.030$0.040$0.050Metered API cost per scored input (USD) →Exact-match accuracy (%) →Gemini 3.5 Flash-Lite: 1.4% ($0.0006 / input)Gemini 3.5 Flash-LiteGPT-5.6 Luna: 4.0% ($0.0009 / input)GPT-5.6 LunaClaude Haiku 4.5: 0.4% ($0.0014 / input)Claude Haiku 4.5Claude Sonnet 5: 1.5% ($0.0033 / input)Claude Sonnet 5GPT-5.6 Terra: 4.7% ($0.0065 / input)GPT-5.6 TerraClaude Opus 5: 4.3% ($0.0090 / input)Claude Opus 5Gemini 3.8 Flash: 4.0% ($0.0121 / input)Gemini 3.8 FlashGemini 3.1 Pro Preview: 3.2% ($0.0169 / input)Gemini 3.1 Pro PreviewClaude Fable 5.1: 13.3% ($0.0174 / input)Claude Fable 5.1GPT-5.6 Sol: 8.1% ($0.0180 / input)GPT-5.6 SolGPT-6 Astra: 40.1% ($0.0464 / input)GPT-6 Astra
Swipe chart horizontally to explore →

728

Shared scored inputs

49

Families in this comparison

11

Model configurations

1,824

Inputs in the full release

Results

Ranked by exact match. Select a model to highlight it.

Exact match requires all six attributes to be correct. Missing measurements appear as —.

FontBench model results on identical shared inputs. All scores are percentages.
ModelExact matchFont familyCategoryWeightModifierLetter spacingLine heightAPI cost / inputResponse time
openai · 728 scored40.1%45.1%96.8%91.5%99.3%94.6%99.3%$0.046420.63s
anthropic · 728 scored13.3%42.9%96.4%82.1%99.6%80.6%40.0%$0.01744.89s
openai · 728 scored8.1%16.1%90.8%83.0%97.5%70.1%92.3%$0.018018.07s
openai · 728 scored4.7%12.5%88.5%80.8%85.2%65.1%84.3%$0.00657.73s
anthropic · 728 scored4.3%14.7%95.5%76.9%99.2%70.1%61.0%$0.00903.07s
google · 728 scored4.0%14.7%83.4%77.5%97.7%55.5%61.5%$0.012112.32s
openai · 728 scored4.0%10.4%91.6%82.0%96.2%64.1%82.3%$0.00099.09s
google · 728 scored3.2%12.5%85.0%75.3%97.7%61.5%50.0%$0.016912.27s
anthropic · 728 scored1.5%7.6%80.1%73.4%99.2%46.2%58.5%$0.00332.08s
google · 728 scored1.4%7.6%80.9%69.1%80.1%46.3%46.4%$0.00060.95s
anthropic · 728 scored0.4%3.4%62.1%71.3%83.7%42.9%42.2%$0.00141.70s

API cost is the mean metered token cost of scored responses at the recorded rates; it excludes failed attempts and spending reservations. Response times are observed API timings, affected by network and service conditions.

What the models see

Examples from the scored sample, using the same sentence with different fonts and settings.

Caveat typography specimen, task font-caveat-v16
Caveathandwriting · 320px layoutfont-caveat-v16
Comic Sans MS typography specimen, task font-comic-sans-ms-v18
Comic Sans MShandwriting · 220px layoutfont-comic-sans-ms-v18
Courier New typography specimen, task font-courier-new-v18
Courier Newmono · 220px layoutfont-courier-new-v18
Fira Code typography specimen, task font-fira-code-v16
Fira Codemono · 320px layoutfont-fira-code-v16
Arial typography specimen, task font-arial-v17
Arialnon-serif · 440px layoutfont-arial-v17
DM Sans typography specimen, task font-dm-sans-v02
DM Sansnon-serif · 320px layoutfont-dm-sans-v02
Bebas Neue typography specimen, task font-bebas-neue-v22
Bebas Neueother · 320px layoutfont-bebas-neue-v22
Bungee typography specimen, task font-bungee-v16
Bungeeother · 320px layoutfont-bungee-v16
Arvo typography specimen, task font-arvo-v21
Arvoserif · 220px layoutfont-arvo-v21
Bitter typography specimen, task font-bitter-v02
Bitterserif · 320px layoutfont-bitter-v02

The six dimensions

FontBench V1.0.0 contains 1,824 inputs across 50 font families. This comparison uses 728 shared inputs from 49 families. The release samples supported combinations; it does not cover every combination.

Full range of the six tested typographic attributes
DimensionValues in the release
Font family50 canonical family names, listed below. Accepted aliases follow the benchmark’s grading rules.
Categoryserif; non-serif (sans-serif); mono (monospace); handwriting; other (decorative display). Categories follow the benchmark’s family catalog.
WeightCSS numeric weights: thin = 200; regular = 400; bold = 700; black = 900.
Modifierregular (no modifier); native italic; underline; strikethrough; small-caps. Small caps may be synthesized when they visibly change the image.
Letter spacingtight = −0.05em; normal = 0em; loose = +0.12em. The prompt calls this kerning, meaning uniform tracking rather than adjustments to individual letter pairs.
Line heighttight = 1.15; normal = 1.45; loose = 1.9 times the fixed 22px font size.

Exact match counts an input as correct only when all six attributes match. Each attribute’s accuracy is also reported separately.

All 50 font families in V1.0.0

In the full release but outside this comparison: Verdana.

  • Arial
  • Arvo
  • Bebas Neue
  • Bitter
  • Bodoni Moda
  • Bungee
  • Caveat
  • Cinzel
  • Comic Sans MS
  • Cormorant Garamond
  • Courier New
  • Crimson Text
  • DM Sans
  • Dancing Script
  • EB Garamond
  • Fira Code
  • Fira Sans
  • Georgia
  • Helvetica
  • Impact
  • Inter
  • JetBrains Mono
  • Lato
  • Libre Baskerville
  • Lobster
  • Lora
  • Merriweather
  • Montserrat
  • Noto Sans
  • Nunito
  • Open Sans
  • Oswald
  • PT Sans
  • PT Serif
  • Pacifico
  • Playfair Display
  • Plus Jakarta Sans
  • Poppins
  • Raleway
  • Roboto
  • Roboto Mono
  • Rubik
  • Shadows Into Light
  • Source Code Pro
  • Source Sans 3
  • Space Mono
  • Spectral
  • Times New Roman
  • Verdana
  • Work Sans

Method

Models receive the same image and question, returning six attributes. Font names follow the release’s accepted names and aliases. Invalid answers receive no credit; provider failures are excluded from accuracy scores.

Each input has equal scoring weight. All comparisons use the same completed inputs, covering 49 of the release’s 50 font families. Model settings and response limits are recorded with the results.

Dataset and grading documentation

Reproduce this publication

Download the original responses, scores, and verification hashes for this version.

Version, source, and verification hashes
Benchmark
FontBench V1.0.0; grading 3
Source commit
805e2146e5ee1637d05cd82c0477371f2b9e3228
Dataset commit
d69e87e2c206ea75c52f5b8340d677bd14af03e3
Dataset SHA-256
d40a3cafa91610bbebaa6c49719f24efcbe754a328e448d3e51f0e1f217c79a6
Protocol SHA-256
3769c8cc383a4959533c4801eec6202035a5a8cd30b76644e8cdfbc11cc859ce
Shared cohort SHA-256
19097ac3b1a2cab662304dcea8f051297779934f2fc35d8d22bc79a1e66ee19e
Final results SHA-256
b10fe2532925e7c11ade529bf8048dcc72eb7320876d2cfae6fb65a908f4a50c
Finalization SHA-256
b6f38785b9228305cf4b4bb49898177b29034965502f085abf9ea23c85e3a3d0