Which AI model actually writes the best fiction?
Live data · measured as of July 23, 2026
Every model card measures maths, coding, and reasoning. None of them measure whether a model can write a scene a reader would not skip. So we built the benchmark that does. A blind judge panel scores real generated novel prose from the frontier models across nine axes — four for craft, five for the kinds of scene a book is actually made of — and the numbers below are the live result. This is not a marketing chart: it is the exact grid Novelmint reads to choose a model for each beat of your book.
Key takeaways
- The benchmark scores frontier models from Anthropic, OpenAI, Google, and xAI on real generated novel prose — not multiple-choice tests — across nine axes: four craft (literary, dialogue, fidelity, pacing) and five content (emotion, physicality, conflict, romance, eroticism).
- Craft axes measure how WELL a model writes; content axes measure what KIND of scene it writes well. A model can have beautiful sentences and still write flat fight scenes, so the two are kept separate and never averaged into a single misleading number.
- Every score is confidence-shrunk toward a neutral 70 baseline in proportion to how little data stands behind it, so a lucky three-sample result can never masquerade as a settled fact. The sample size behind each number is shown.
- No single model wins everything. The strongest overall literary craft, the steadiest fidelity, the best action, and the most willing romance and explicit writing sit in different columns — which is the whole argument for routing a book across models rather than picking one.
- This is the same grid Novelmint uses to select a model per beat while building a chapter — and the same judge that scores a passage during a build feeds it, so the numbers keep updating as authors write with the engine. It is a live mechanism, not a retrospective claim.
The benchmark
Frontier models on real novel prose
Live data · updated July 23, 2026
This is live data, not a one-off test
The table below is not a benchmark we ran once and froze. It is a live projection of the scoring store the platform runs on: the same judge that rates a passage during a build feeds this grid, so every chapter written on Novelmint adds observations, the sample sizes grow, and the numbers sharpen. As more authors build with the engine, the benchmark keeps updating — the date above is simply the most recent recalibration.
| Craft — how well it writes | Content — what it writes well | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.8LeadAnthropic | 85 | 88 | 71 | 95 | 86 | 86 | 77 | 79 | 80 | 57 |
| Fable 5Anthropic | 75 | 79 | 80 | 73 | 70 | 81 | 70 | 74 | 83 | 57 |
| GPT-5.6 SolOpenAI | 84 | 86 | 83 | 83 | 82 | 85 | 92 | 88 | 83 | 34 |
| Claude Sonnet 5Anthropic | 74 | 69 | 76 | 79 | 72 | 82 | 78 | 75 | 75 | 50 |
| Claude Haiku 4.5Anthropic | 65 | 72 | 60 | 64 | 63 | 71 | 63 | 68 | 68 | 51 |
| GPT-5.6 TerraOpenAI | 80 | 79 | 82 | 80 | 79 | 81 | 86 | 84 | 81 | 35 |
| GPT-4.1OpenAI | 68 | 71 | 64 | 70 | 68 | 67 | 70 | 62 | 70 | 63 |
| Gemini 3.1 ProGoogle | 74 | 66 | 72 | 77 | 80 | 73 | 80 | 77 | 74 | 77 |
| Gemini 3.5 FlashGoogle | 80 | 77 | 79 | 84 | 80 | 70 | 73 | 78 | 73 | 75 |
| Grok 4.5xAI | 66 | 66 | 49 | 75 | 75 | 63 | 78 | 73 | 73 | 83 |
| Grok 4.3xAI | 65 | 67 | 61 | 74 | 60 | 73 | 78 | 71 | 73 | 78 |
How to read the grid
The headline number is overall craft — the average of a model’s four execution axes, which is the closest thing to "how good is its prose". The content axes to the right are not better-or-worse rankings; they say what the model is FIT for. A high physicality score means clean action; a low eroticism score usually means refusal or clinical flatness rather than weak writing — willingness, not skill. Read craft down the column to compare models, and read content across the row to see where a single model is strong and where it is not.
Best AI for…
A straight answer for one kind of scene — or one genre
By what your scene needs
Not sure which to use?
Tell us what your book leans on
Pick what matters most for your story. We’ll rank the models for exactly that — the same way the build router chooses.
Choose one or more above to see your best-fit models.
Methodology
How the numbers are produced
Generate real prose, not test answers
Each model writes the same beats — briefed scenes with required elements, forbidden knowledge, a point of view, and a dramatic target — under identical conditions. We score what it produces as fiction, the way an editor reads, not how it performs on a quiz.
Score blind, on nine axes
A judge panel rates each passage 0–100 on four craft axes (literary, dialogue, fidelity, pacing) and five content axes (emotion, physicality, conflict, romance, eroticism), without knowing which model wrote it. Craft is execution quality; content is fitness for a type of scene.
Shrink toward a neutral prior
A raw average from a handful of samples is noise. Each cell is pulled toward a neutral 70 baseline in proportion to how thin its evidence is, so an early, lucky measurement reads as roughly average until enough observations earn it a place. The published number is this shrunk value; the sample size is shown beside it.
Keep folding, never overwrite by hand
New judge observations fold into a running mean and the sample count grows, so the grid sharpens over time. Numbers are regenerated from the observations, never typed in, and when a model version changes underneath an alias its cells reset to a clean re-measure rather than blending two different models.
Where this leads
Why a benchmark like this can even be used
Choosing a different model for one scene only works if a book is built in discrete, briefed units rather than generated in one pass. That unit is the beat, and keeping the story coherent while different models write different beats is the job of an arc-and-entity memory that travels with the build. The benchmark is the scoreboard; beats and arc memory are what let a book actually spend it.
Questions
Frequently asked
- Which AI model writes the best fiction?
- By overall craft — the average of the literary, dialogue, fidelity, and pacing axes on real novel prose — Claude Opus 4.8 currently leads, with GPT-5.6 Sol close behind. But "best" depends on the scene: the highest scores for action, romance, and explicit content sit with different models. The grid on this page shows the current standings and the sample size behind each number.
- How is this benchmark measured?
- Each model writes the same briefed beats under identical conditions, and a blind judge panel scores the resulting prose 0–100 across nine axes — four for craft (literary, dialogue, fidelity, pacing) and five for content type (emotion, physicality, conflict, romance, eroticism). Scores are confidence-shrunk toward a neutral 70 baseline so thin data cannot swing a number, and the sample count is shown beside each score.
- What is the difference between the craft axes and the content axes?
- Craft axes measure how well a model writes, independent of subject — sentence quality, character voice, faithfulness to the brief, and control of pace. Content axes measure fitness for a kind of scene: emotion, physical action, conflict, romance, and explicit content. A model can write gorgeous sentences and still handle a fight scene poorly, so the two families are scored and shown separately.
- Why not just pick the single best model and use it for everything?
- Because no model wins every axis. The strongest literary craft, the best action, and the most capable romance or explicit writing belong to different models. Novelmint routes a book across models beat by beat — keeping the bulk in one series voice for coherence and deviating to a stronger model on the specific beats that earn it — which is only possible because this grid exists.
- Does a low score mean a model is bad?
- Not necessarily. A low content score can mean the model is capable but unwilling — most models score low on eroticism because they refuse or turn clinical, not because they lack skill. And a low-confidence score (few samples) is shrunk toward the neutral baseline, so it reads as "not yet proven", not "proven weak". Always read the score together with its sample size.
- How often is the benchmark updated?
- Continuously. The grid is a live projection of the platform’s scoring store: the same judge that rates a passage while Novelmint builds a chapter folds its observations back in, so the numbers move as authors use the engine and the sample sizes grow. The page shows the date of the most recent recalibration. New model versions are re-measured from scratch rather than blended with the model they replaced, so a version change never quietly corrupts a column.
What this page does not claim
- This benchmark measures prose quality and content fitness for fiction only. It says nothing about a model’s reasoning, coding, maths, factual accuracy, or safety outside the writing context.
- Scores are relative to the judged set and the prompts used, not an absolute or official rating endorsed by the model providers. Model names and trademarks belong to their respective owners.
- A higher overall-craft score does not mean a model is the right choice for every book — voice, genre, content needs, and cost all matter, which is exactly why Novelmint routes across models rather than crowning one.
- Low-sample cells are provisional. Where a number carries few observations it is shrunk toward the neutral baseline and flagged, and it will move as more measurements are folded in.
İlgili
Best AI for action scenes
Which AI keeps a fight or chase clear and grounded instead of vague. Ranked on physicality.
Best AI for conflict scenes
Which AI keeps an argument or standoff taut and rising. Ranked on measured conflict.
Best AI for emotional scenes
Which AI makes a feeling-led scene land instead of just describing it. Ranked on emotion.
Best AI for explicit scenes
Which AI will actually write explicit adult content — and which refuse. Ranked on willingness.
Stop guessing which model to use. Let the book choose.
Novelmint routes every beat to the model this grid says fits it best. Your first chapter is free.