Metodología

Cómo funciona el benchmark de modelos de ficción

Updated July 23, 2026 · 6 min de lectura

El benchmark mide lo que las fichas técnicas de los modelos nunca miden: si un modelo puede escribir una escena que el lector no se saltaría. Esta página explica exactamente cómo se producen los números —qué se mide, cómo se evalúan los modelos, por qué las puntuaciones se ajustan por confianza y cómo la tabla se mantiene actualizada— para que puedas valorar los resultados y citarlos con precisión.

Key takeaways

  • Los modelos de frontera escriben los mismos fragmentos de novela con instrucciones, y un panel de jueces ciego y diverso puntúa la prosa de 0 a 100 en nueve ejes, sin saber qué modelo la escribió.
  • Nueve ejes en dos familias: cuatro de oficio (literario, diálogo, fidelidad, ritmo) miden qué tan bien escribe un modelo; cinco de contenido (emoción, corporalidad, conflicto, romance, erotismo) miden qué tipo de escena escribe bien. Nunca se promedian entre sí.
  • Cada puntuación se ajusta hacia una línea base neutral de 70 en proporción a cuántos pocos datos la respaldan (un prior equivalente a cinco observaciones), de modo que una muestra pequeña con suerte no pueda llegar a la cima. Los tamaños de muestra se muestran; las celdas con pocos datos se marcan.
  • La tabla está en vivo: las nuevas observaciones de los jueces se integran en una media acumulada con el tiempo, y un cambio de versión reinicia y vuelve a medir un modelo en lugar de mezclarlo con su predecesor.
  • Las puntuaciones describen prosa solo en fragmentos de ficción; son una medición independiente, no una valoración oficial del proveedor, y una puntuación baja en erotismo suele reflejar disposición, no oficio.

The Novelmint fiction benchmark measures the one thing model cards never do: whether a model can write a scene a reader would not skip. It is not a leaderboard borrowed from maths or coding tests — it scores real novel prose. Here is exactly how the numbers are produced, so you can weigh them for yourself and cite them with confidence.

What it measures

The benchmark scores how well the frontier models — from Anthropic, OpenAI, Google, and xAI — write fiction. Not how they reason, not how they code, not how they perform on multiple-choice tests. Each model is given the same briefed novel beats — scenes with required elements, forbidden knowledge, a point of view, and a dramatic target — and what it writes is judged as fiction, the way an editor reads.

Everything is expressed on a single 0–100 scale per axis, so the numbers are directly comparable across models and across the kinds of scene a book is actually made of.

The nine axes

Fiction is not one skill, so the benchmark does not collapse it into one number. It scores nine axes, split into two families that mean different things.

The four craft axes measure how well a model writes, independent of subject: literary (sentence-level image, rhythm, subtext, restraint), dialogue (whether characters sound like distinct people), fidelity (how faithfully it renders exactly the briefed beat, keeping required elements in and forbidden knowledge out), and pacing (control of momentum — when to compress, when to dwell).

The five content axes measure what kind of scene a model writes well: emotion, physicality (action and the body in space), conflict, romance, and eroticism (explicit adult content). These are not better-or-worse rankings — they say what a model is fit for. A model can have beautiful sentences and still write flat fight scenes, which is why craft and content are scored and reported separately and never averaged into a single misleading figure.

How each model is scored

Scoring is blind peer review. Each model writes the same set of briefed beats, and a panel of strong, provider-diverse judge models rates every passage 0–100 on each axis — without knowing which model wrote it, and never judging its own work. A provider-diverse panel matters: a judge from outside a model's own family catches the tells that family tends to reward.

Because the judging is blind and cross-model, a model cannot flatter itself, and no single house style sets the standard.

Why scores are shrunk by confidence

A raw average from a handful of passages is noise, and noise dressed up as a score is worse than no score at all. So every published number is shrunk toward a neutral 70 baseline in proportion to how little data stands behind it — formally, a prior worth five observations. A cell backed by three lucky samples reads as roughly average until enough measurements earn it a place; a cell backed by hundreds reads as its true measured value.

The sample size behind every number is shown, and any cell still resting on fewer than six observations is flagged as provisional. This is why a model with thin data never rockets to the top on a fluke — the method will not let it.

How the grid stays current

The benchmark is not a one-time test that was run and frozen. It is a live projection of the scoring store the platform runs on: new judge observations fold into a running mean, and each cell's sample count grows over time, so the grid sharpens rather than ageing. The page always shows the date of the most recent recalibration.

One safeguard matters for accuracy. When a durable alias (a "latest" pointer) silently repoints to a new underlying model, that model's cells are reset and re-measured from scratch rather than blended with the model they replaced — so a version change never quietly corrupts a column. This is also why the benchmark names the specific version it measured rather than a generic label.

Limitations

The scores describe prose on fiction beats only, judged against this set of models and prompts. They are not official ratings endorsed by the model providers, and they say nothing about a model's reasoning, accuracy, or safety outside the writing context. Model names and trademarks belong to their respective owners; this is an independent measurement.

Two reads are easy to get wrong. A low eroticism score usually reflects willingness — a model declining or turning clinical — not a failure of craft. And a low-confidence score means "not yet proven", not "proven weak". Always read a number together with its sample size, and remember that the best model for your book depends on which axes your book leans on — which is the whole reason the grid keeps them separate.

Questions

Frequently asked

¿Cómo evalúa Novelmint los modelos de IA para ficción?
Cada modelo de frontera escribe los mismos fragmentos de novela con instrucciones, y un panel ciego y diverso de modelos jueces sólidos puntúa la prosa resultante de 0 a 100 en nueve ejes sin conocer al autor. Las puntuaciones se ajustan hacia una línea base neutral según cuántos datos las respaldan, y las observaciones se integran en una media acumulada con el tiempo.
¿Cuáles son los nueve ejes?
Cuatro ejes de oficio —literario, diálogo, fidelidad (cumplimiento del encargo) y ritmo— miden qué tan bien escribe un modelo. Cinco ejes de contenido —emoción, corporalidad (acción), conflicto, romance y erotismo— miden qué tipo de escena escribe bien. El oficio y el contenido se reportan por separado, nunca se promedian.
¿Por qué se ajustan las puntuaciones por confianza?
Un promedio crudo de pocos fragmentos es ruido. Cada puntuación se ajusta hacia una línea base neutral de 70 en proporción a cuántos pocos datos la respaldan —un prior equivalente a cinco observaciones— para que una muestra pequeña con suerte no pueda llegar a la cima. El tamaño de muestra se muestra, y las celdas con menos de seis observaciones se marcan como provisionales.
¿Es esta una clasificación oficial de OpenAI, Anthropic, Google o xAI?
No. Es una medición independiente de Novelmint, sobre prosa de ficción únicamente, evaluada con este conjunto de modelos y prompts. No dice nada sobre razonamiento, precisión u otras capacidades, y los nombres de los modelos son marcas registradas de sus respectivos propietarios.
¿Con qué frecuencia se actualiza el benchmark?
Continuamente: las nuevas observaciones de los jueces se integran en una media acumulada y los recuentos de muestra crecen, de modo que la tabla gana precisión con el tiempo. Cuando la versión de un modelo cambia bajo un alias, sus celdas se reinician y vuelven a medirse en lugar de mezclarse con la versión anterior. La página muestra la fecha de la recalibración más reciente.

What this page does not claim

  • Las puntuaciones son relativas al conjunto de modelos evaluados y a los prompts utilizados; no son una valoración absoluta ni oficial.
  • Una puntuación baja en erotismo refleja disposición (un modelo que declina o se vuelve clínico), no falta de oficio; una puntuación de baja confianza significa «aún no probado», no «probado como débil».
  • El benchmark mide únicamente prosa de ficción y no dice nada sobre razonamiento, código, precisión factual ni seguridad fuera del contexto de escritura.

Descubre qué modelo escribe mejor tu tipo de escena.

Datos en vivo en nueve ejes, actualizados a medida que se mide más prosa.