The fiction model benchmark

Which AI model actually writes the best fiction?

Live data · measured as of July 23, 2026

Every model card measures maths, coding, and reasoning. None of them measure whether a model can write a scene a reader would not skip. So we built the benchmark that does. A blind judge panel scores real generated novel prose from the frontier models across nine axes — four for craft, five for the kinds of scene a book is actually made of — and the numbers below are the live result. This is not a marketing chart: it is the exact grid Novelmint reads to choose a model for each beat of your book.

Key takeaways

  • The benchmark scores frontier models from Anthropic, OpenAI, Google, and xAI on real generated novel prose — not multiple-choice tests — across nine axes: four craft (literary, dialogue, fidelity, pacing) and five content (emotion, physicality, conflict, romance, eroticism).
  • Craft axes measure how WELL a model writes; content axes measure what KIND of scene it writes well. A model can have beautiful sentences and still write flat fight scenes, so the two are kept separate and never averaged into a single misleading number.
  • Every score is confidence-shrunk toward a neutral 70 baseline in proportion to how little data stands behind it, so a lucky three-sample result can never masquerade as a settled fact. The sample size behind each number is shown.
  • No single model wins everything. The strongest overall literary craft, the steadiest fidelity, the best action, and the most willing romance and explicit writing sit in different columns — which is the whole argument for routing a book across models rather than picking one.
  • This is the same grid Novelmint uses to select a model per beat while building a chapter — and the same judge that scores a passage during a build feeds it, so the numbers keep updating as authors write with the engine. It is a live mechanism, not a retrospective claim.

The benchmark

Frontier models on real novel prose

Live data · updated July 23, 2026

This is live data, not a one-off test

The table below is not a benchmark we ran once and froze. It is a live projection of the scoring store the platform runs on: the same judge that rates a passage during a build feeds this grid, so every chapter written on Novelmint adds observations, the sample sizes grow, and the numbers sharpen. As more authors build with the engine, the benchmark keeps updating — the date above is simply the most recent recalibration.

Novelmint fiction-model benchmark — measured craft and content-fitness scores (0–100) for frontier AI models on real novel prose. Click a column heading to sort.
Craft — how well it writesContent — what it writes well
Claude Opus 4.8LeadAnthropic85887195868677798057
Fable 5Anthropic75798073708170748357
GPT-5.6 SolOpenAI84868383828592888334
Claude Sonnet 5Anthropic74697679728278757550
Claude Haiku 4.5Anthropic65726064637163686851
GPT-5.6 TerraOpenAI80798280798186848135
GPT-4.1OpenAI68716470686770627063
Gemini 3.1 ProGoogle74667277807380777477
Gemini 3.5 FlashGoogle80777984807073787375
Grok 4.5xAI66664975756378737383
Grok 4.3xAI65676174607378717378
0–100 confidence-adjusted score* = fewer than 6 samples (provisional)Click a column heading to sort · hover a cell for its sample size

How to read the grid

The headline number is overall craft — the average of a model’s four execution axes, which is the closest thing to "how good is its prose". The content axes to the right are not better-or-worse rankings; they say what the model is FIT for. A high physicality score means clean action; a low eroticism score usually means refusal or clinical flatness rather than weak writing — willingness, not skill. Read craft down the column to compare models, and read content across the row to see where a single model is strong and where it is not.

Not sure which to use?

Tell us what your book leans on

Pick what matters most for your story. We’ll rank the models for exactly that — the same way the build router chooses.

Choose one or more above to see your best-fit models.

Methodology

How the numbers are produced

01

Generate real prose, not test answers

Each model writes the same beats — briefed scenes with required elements, forbidden knowledge, a point of view, and a dramatic target — under identical conditions. We score what it produces as fiction, the way an editor reads, not how it performs on a quiz.

02

Score blind, on nine axes

A judge panel rates each passage 0–100 on four craft axes (literary, dialogue, fidelity, pacing) and five content axes (emotion, physicality, conflict, romance, eroticism), without knowing which model wrote it. Craft is execution quality; content is fitness for a type of scene.

03

Shrink toward a neutral prior

A raw average from a handful of samples is noise. Each cell is pulled toward a neutral 70 baseline in proportion to how thin its evidence is, so an early, lucky measurement reads as roughly average until enough observations earn it a place. The published number is this shrunk value; the sample size is shown beside it.

04

Keep folding, never overwrite by hand

New judge observations fold into a running mean and the sample count grows, so the grid sharpens over time. Numbers are regenerated from the observations, never typed in, and when a model version changes underneath an alias its cells reset to a clean re-measure rather than blending two different models.

Where this leads

Why a benchmark like this can even be used

Choosing a different model for one scene only works if a book is built in discrete, briefed units rather than generated in one pass. That unit is the beat, and keeping the story coherent while different models write different beats is the job of an arc-and-entity memory that travels with the build. The benchmark is the scoreboard; beats and arc memory are what let a book actually spend it.

Questions

Frequently asked

Which AI model writes the best fiction?
By overall craft — the average of the literary, dialogue, fidelity, and pacing axes on real novel prose — Claude Opus 4.8 currently leads, with GPT-5.6 Sol close behind. But "best" depends on the scene: the highest scores for action, romance, and explicit content sit with different models. The grid on this page shows the current standings and the sample size behind each number.
How is this benchmark measured?
Each model writes the same briefed beats under identical conditions, and a blind judge panel scores the resulting prose 0–100 across nine axes — four for craft (literary, dialogue, fidelity, pacing) and five for content type (emotion, physicality, conflict, romance, eroticism). Scores are confidence-shrunk toward a neutral 70 baseline so thin data cannot swing a number, and the sample count is shown beside each score.
What is the difference between the craft axes and the content axes?
Craft axes measure how well a model writes, independent of subject — sentence quality, character voice, faithfulness to the brief, and control of pace. Content axes measure fitness for a kind of scene: emotion, physical action, conflict, romance, and explicit content. A model can write gorgeous sentences and still handle a fight scene poorly, so the two families are scored and shown separately.
Why not just pick the single best model and use it for everything?
Because no model wins every axis. The strongest literary craft, the best action, and the most capable romance or explicit writing belong to different models. Novelmint routes a book across models beat by beat — keeping the bulk in one series voice for coherence and deviating to a stronger model on the specific beats that earn it — which is only possible because this grid exists.
Does a low score mean a model is bad?
Not necessarily. A low content score can mean the model is capable but unwilling — most models score low on eroticism because they refuse or turn clinical, not because they lack skill. And a low-confidence score (few samples) is shrunk toward the neutral baseline, so it reads as "not yet proven", not "proven weak". Always read the score together with its sample size.
How often is the benchmark updated?
Continuously. The grid is a live projection of the platform’s scoring store: the same judge that rates a passage while Novelmint builds a chapter folds its observations back in, so the numbers move as authors use the engine and the sample sizes grow. The page shows the date of the most recent recalibration. New model versions are re-measured from scratch rather than blended with the model they replaced, so a version change never quietly corrupts a column.

What this page does not claim

  • This benchmark measures prose quality and content fitness for fiction only. It says nothing about a model’s reasoning, coding, maths, factual accuracy, or safety outside the writing context.
  • Scores are relative to the judged set and the prompts used, not an absolute or official rating endorsed by the model providers. Model names and trademarks belong to their respective owners.
  • A higher overall-craft score does not mean a model is the right choice for every book — voice, genre, content needs, and cost all matter, which is exactly why Novelmint routes across models rather than crowning one.
  • Low-sample cells are provisional. Where a number carries few observations it is shrunk toward the neutral baseline and flagged, and it will move as more measurements are folded in.

Stop guessing which model to use. Let the book choose.

Novelmint routes every beat to the model this grid says fits it best. Your first chapter is free.