Methodik

Wie der Fictions-Modell-Benchmark funktioniert

Updated July 23, 2026 · 6 Min. Lesezeit

Der Benchmark misst genau das, was Modellkarten nie messen: ob ein Modell eine Szene schreiben kann, die ein Leser nicht überspringen würde. Diese Seite erklärt genau, wie die Zahlen entstehen – was gemessen wird, wie Modelle bewertet werden, warum Scores durch Konfidenz korrigiert werden und wie das Raster aktuell bleibt –, damit du die Ergebnisse selbst einordnen und korrekt zitieren kannst.

Key takeaways

  • Frontier-Modelle schreiben dieselben gebrieften Romanszenen, und ein blindes, anbieterseitig diverses Jurorenpanel bewertet die Prosa auf neun Achsen mit 0–100 – ohne zu wissen, welches Modell sie geschrieben hat.
  • Neun Achsen in zwei Kategorien: Vier Handwerks-Achsen (Literarizität, Dialog, Werktreue, Tempo) messen, wie gut ein Modell schreibt; fünf Inhalts-Achsen (Emotion, Körperlichkeit, Konflikt, Romantik, Erotik) messen, welche Art von Szene es gut schreibt. Sie werden nie zusammengemittelt.
  • Jeder Score wird proportional zur Datenmenge auf eine neutrale Baseline von 70 zurückgeführt (ein Prior mit dem Gewicht von fünf Beobachtungen), damit eine glückliche kleine Stichprobe nicht an die Spitze gelangen kann. Stichprobengrößen werden angezeigt; dünne Zellen werden markiert.
  • Das Raster ist live – neue Jurorbeobachtungen fließen fortlaufend in einen laufenden Mittelwert ein, und eine Versionsänderung setzt ein Modell zurück und misst es neu, anstatt es mit seinem Vorgänger zu vermischen.
  • Scores beschreiben Prosa nur auf Fictions-Szenen; sie sind eine unabhängige Messung, keine offizielle Anbieterbewertung, und ein niedriger Erotik-Score spiegelt meist die Bereitschaft wider, nicht das Handwerk.

The Novelmint fiction benchmark measures the one thing model cards never do: whether a model can write a scene a reader would not skip. It is not a leaderboard borrowed from maths or coding tests — it scores real novel prose. Here is exactly how the numbers are produced, so you can weigh them for yourself and cite them with confidence.

What it measures

The benchmark scores how well the frontier models — from Anthropic, OpenAI, Google, and xAI — write fiction. Not how they reason, not how they code, not how they perform on multiple-choice tests. Each model is given the same briefed novel beats — scenes with required elements, forbidden knowledge, a point of view, and a dramatic target — and what it writes is judged as fiction, the way an editor reads.

Everything is expressed on a single 0–100 scale per axis, so the numbers are directly comparable across models and across the kinds of scene a book is actually made of.

The nine axes

Fiction is not one skill, so the benchmark does not collapse it into one number. It scores nine axes, split into two families that mean different things.

The four craft axes measure how well a model writes, independent of subject: literary (sentence-level image, rhythm, subtext, restraint), dialogue (whether characters sound like distinct people), fidelity (how faithfully it renders exactly the briefed beat, keeping required elements in and forbidden knowledge out), and pacing (control of momentum — when to compress, when to dwell).

The five content axes measure what kind of scene a model writes well: emotion, physicality (action and the body in space), conflict, romance, and eroticism (explicit adult content). These are not better-or-worse rankings — they say what a model is fit for. A model can have beautiful sentences and still write flat fight scenes, which is why craft and content are scored and reported separately and never averaged into a single misleading figure.

How each model is scored

Scoring is blind peer review. Each model writes the same set of briefed beats, and a panel of strong, provider-diverse judge models rates every passage 0–100 on each axis — without knowing which model wrote it, and never judging its own work. A provider-diverse panel matters: a judge from outside a model's own family catches the tells that family tends to reward.

Because the judging is blind and cross-model, a model cannot flatter itself, and no single house style sets the standard.

Why scores are shrunk by confidence

A raw average from a handful of passages is noise, and noise dressed up as a score is worse than no score at all. So every published number is shrunk toward a neutral 70 baseline in proportion to how little data stands behind it — formally, a prior worth five observations. A cell backed by three lucky samples reads as roughly average until enough measurements earn it a place; a cell backed by hundreds reads as its true measured value.

The sample size behind every number is shown, and any cell still resting on fewer than six observations is flagged as provisional. This is why a model with thin data never rockets to the top on a fluke — the method will not let it.

How the grid stays current

The benchmark is not a one-time test that was run and frozen. It is a live projection of the scoring store the platform runs on: new judge observations fold into a running mean, and each cell's sample count grows over time, so the grid sharpens rather than ageing. The page always shows the date of the most recent recalibration.

One safeguard matters for accuracy. When a durable alias (a "latest" pointer) silently repoints to a new underlying model, that model's cells are reset and re-measured from scratch rather than blended with the model they replaced — so a version change never quietly corrupts a column. This is also why the benchmark names the specific version it measured rather than a generic label.

Limitations

The scores describe prose on fiction beats only, judged against this set of models and prompts. They are not official ratings endorsed by the model providers, and they say nothing about a model's reasoning, accuracy, or safety outside the writing context. Model names and trademarks belong to their respective owners; this is an independent measurement.

Two reads are easy to get wrong. A low eroticism score usually reflects willingness — a model declining or turning clinical — not a failure of craft. And a low-confidence score means "not yet proven", not "proven weak". Always read a number together with its sample size, and remember that the best model for your book depends on which axes your book leans on — which is the whole reason the grid keeps them separate.

Questions

Frequently asked

Wie bewertet Novelmint KI-Modelle für Fiction?
Jedes Frontier-Modell schreibt dieselben gebrieften Romanszenen, und ein blindes, anbieterseitig diverses Panel starker Juror-Modelle bewertet die entstehende Prosa mit 0–100 auf neun Achsen – ohne den Autor zu kennen. Scores werden proportional zur Datenmenge auf eine neutrale Baseline zurückgeführt, und Beobachtungen fließen laufend in einen Mittelwert ein.
Was sind die neun Achsen?
Vier Handwerks-Achsen – Literarizität, Dialog, Werktreue (Treue zum Brief) und Tempo – messen, wie gut ein Modell schreibt. Fünf Inhalts-Achsen – Emotion, Körperlichkeit (Action), Konflikt, Romantik und Erotik – messen, welche Art von Szene es gut schreibt. Handwerk und Inhalt werden separat ausgewiesen und nie gemittelt.
Warum werden die Scores für Konfidenz angepasst?
Ein roher Mittelwert aus wenigen Passagen ist Rauschen. Jeder Score wird proportional zur Datenmenge auf eine neutrale Baseline von 70 zurückgeführt – ein Prior mit dem Gewicht von fünf Beobachtungen –, damit eine glückliche kleine Stichprobe nicht an die Spitze gelangen kann. Die Stichprobengröße wird angezeigt, und Zellen mit weniger als sechs Beobachtungen werden als vorläufig markiert.
Ist dies ein offizielles Ranking von OpenAI, Anthropic, Google oder xAI?
Nein. Es ist eine unabhängige Messung von Novelmint, ausschließlich für Fiction-Prosa, bewertet anhand dieser Modell- und Prompt-Auswahl. Sie sagt nichts über Denkvermögen, Genauigkeit oder andere Fähigkeiten aus, und Modellnamen sind Marken ihrer jeweiligen Inhaber.
Wie oft wird der Benchmark aktualisiert?
Kontinuierlich – neue Jurorbeobachtungen fließen in einen laufenden Mittelwert ein und Stichprobengrößen wachsen, sodass das Raster mit der Zeit schärfer wird. Ändert sich eine Modellversion unter einem Alias, werden ihre Zellen zurückgesetzt und neu gemessen, anstatt mit der älteren Version vermischt zu werden. Die Seite zeigt das Datum der letzten Rekalibrierung.

What this page does not claim

  • Die Scores sind relativ zur bewerteten Modellmenge und den verwendeten Prompts – keine absolute oder offizielle Bewertung.
  • Ein niedriger Erotik-Score spiegelt die Bereitschaft wider (ein Modell lehnt ab oder wird klinisch), nicht mangelndes Handwerk; ein niedrig-konfidenter Score bedeutet „noch nicht belegt", nicht „als schwach belegt".
  • Der Benchmark misst ausschließlich Fiction-Prosa und sagt nichts über Denkvermögen, Coding, Faktengenauigkeit oder Sicherheit außerhalb des Schreibkontexts aus.

Finde heraus, welches Modell deine Art von Szene am besten schreibt.

Live-Daten über neun Achsen, aktualisiert während mehr Prosa gemessen wird.