On September 17, 2026, Vulsar AI released Creative Writing Bench V1, which compares 24 large language models (LLMs) and human writing across 475 creative writing tasks. Overall, GPT 6 Astra edged past the human reference score, but in the category that uses professional authors' work as its reference, humans pulled far ahead of every model. However, the rankings are predictions from a dedicated model trained on readers' preferences, not the results of votes by people who actually read the stories. To choose an AI for creative writing support, you need to look not only at the scores but also at what was compared, on which tasks, and how it was evaluated.
Top overall, yet a gap remains with professional work
In the official ranking, GPT 6 Astra took first place overall with a predicted win rate of 87.8%. The human reference scored 86.6%. In the amateur category, Astra and GPT 5.6 Sol both beat the human reference, while in the professional category humans lead with 99.9%.
| Model / human reference | Overall predicted win rate | Professional | Amateur |
|---|---|---|---|
| Human work | 86.6% | 99.9% | 79.7% |
| GPT 6 Astra | 87.8% | 76.2% | 93.9% |
| GPT 5.6 Sol | 77.6% | 63.1% | 85.2% |
Source: Vulsar AI, published September 17, 2026; checked September 20. The human reference corresponds to the set of works in each category. The overall figure combines the professional and amateur task sets, and the two categories use different prompts. Refusals and empty responses are excluded from scoring.
In the professional category, there is a 21.4-point gap between humans (99.9%) and the top-ranked AI, Grok 4.6 (78.5%). But this is a difference in predicted win rates against all comparison opponents within the category, not a head-to-head win rate between humans and Grok. The gap was calculated by subtracting the displayed 78.5 from 99.9.
The 95% confidence intervals in the same category are 99.7–99.9% for humans and 76.2–80.6% for Grok. In the overall ranking, by contrast, the intervals overlap: 86.7–88.9% for Astra and 85.5–87.7% for humans. This alone cannot establish whether the difference is statistically significant, but it is best to avoid reading the overall ranking as proof that AI has surpassed human creative ability.
Moreover, the human entry is not the record of a single writer. It is a value aggregated from the reference works in each category. The narrow overall margin and the large professional-category gap coexist because they come from different task sets.
"Win rate" is not the result of reader voting
Vulsar assigns a single score to each work using its own reward model, trained on human preferences about creative writing. A reward model is a model trained to give higher scores to outputs that people tend to prefer. According to the company, it does not use the approach of handing a general-purpose LLM a scoring rubric and having it act as judge.
The calculation starts by using the difference between the scores of two works on the same task to derive the probability that one is preferred over the other. That probability is then averaged with equal weight across all scorable combinations of opponents and tasks, giving the predicted win rate. Because the opponents include other AIs, Astra's 87.8% cannot be read as "an 87.8% chance of beating a human's work."
Under this design, the meaning of the numbers changes if the participating models or tasks change. It is not an absolute score measuring literary ability against a fixed yardstick, but a relative indicator of how likely a work is to be preferred within this particular comparison group.
The handling of refusals also matters in practical use. Vulsar does not count refusals and empty responses as losses; it removes the corresponding comparisons from both the numerator and denominator of the average. For example, GPT 5.6 Luna has a predicted win rate of 49.7% in the professional category, but its refusal/empty-response rate is 11.7%. The former is an evaluation under the condition that a work was returned; the latter concerns how often you don't receive a work at all.
So whether a model scores high and whether it will actually write a story on request should be checked separately. If you are choosing a tool for creative support, keeping only the win rate and dropping the refusal rate leaves out information needed to judge usability.
How prompt detail and generation settings shift the rankings
Grok 4.6 rises from 10th overall to 2nd in the professional category, while ranking 12th in the amateur category. Vulsar explains that the professional-category prompts contain many specific, detailed conditions and that Grok is good at that kind of instruction, which contributed to the climb.
In other words, between the professional and amateur tables, not only the human reference works but also the tasks given to the AIs differ. It is not a controlled experiment that simply swaps the authors' credentials. What Grok's movement suggests is that rankings of creative models can depend on the content of the instructions. It does not show that adding more detailed conditions always improves quality.
The generation settings are also limited. Each model writes one short story per task. Reasoning is disabled for models that allow it, and models that don't allow it use the lowest available reasoning effort. The results therefore do not measure the ceiling of production methods that let a model think longer or choose among multiple drafts.
Work length is not uniform either. In the professional category, the average is 4,528 tokens for humans, 607 for Grok, and 1,745 for Astra. However, tokens are units the model uses to process text, not a common character count. The human works are explicitly tokenized with the Qwen3.8 method, so the table should not be used to derive an exact ratio of text volume between models.
The averages alone also don't reveal whether length produced the higher scores or whether well-developed stories are simply preferred. While keeping in mind that the professional reference works were rated highly, it is important not to reduce the cause to a single explanation such as "human quality."
How to verify the reliability of the AI evaluator
The idea of using an evaluator trained for creative writing has precedent. The LitBench preprint, published in 2025 by Daniel Fein and colleagues at Stanford University, examined whether an evaluator can reproduce human preferences, using 2,480 Reddit-derived story comparison pairs for evaluation and 43,827 for training.
In that study, the best result for an off-the-shelf LLM used directly as a judge was 73% agreement, by Claude 3.7 Sonnet, while a trained reward model reached 78%. The researchers also ran an online human evaluation of newly LLM-generated stories. These are LitBench's results and are not figures showing the accuracy of Vulsar's evaluator.
LitBench also addressed the possibility that length influences evaluations. In the original comparison set, the preferred work was the longer one in 65.25% of pairs, so the researchers thinned out comparison pairs to avoid bias in length differences between works. They did not ignore the relationship between text volume and preference, and instead scrutinized the data used to test the evaluator itself.
Vulsar's public explanation page does not state the specific number of training examples for the reward model, the composition of the raters, or the agreement rate with human evaluation of the works in this benchmark. Detailed criteria for how the professional and amateur reference works were selected are also not apparent. What can be confirmed now is a benchmark announcement on the company's own site, which should be distinguished from results validated in a peer-reviewed paper.
This is not evidence that the evaluation is wrong. That a dedicated reward model has promising precedent, and how well this particular evaluator applies to which reader groups, are separate questions still to be verified. To interpret differences between models strongly as differences in creative quality, we would want to know how reliable the evaluating side is, as well as the generating side.
If you use it for creative support, read the stories after the scores
On its public samples page, Vulsar lets readers view models' works for the same task side by side. You can narrow candidates by ranking and then check in actual prose whether character dialogue and story development suit your purpose. The public examples are in English, so the same ranking will not necessarily hold for Japanese fiction or long-form projects.
Vulsar itself also states that a high score does not guarantee novelty. Although the human preferences used in training include judgments about originality, the company does not independently verify whether wording or ideas are new by checking them against existing texts or the models' training data. "Feels fresh" and "has never existed before" are not the same thing.
It is significant that comparisons of creative AI are moving from whether a model can produce fluent text to whether readers want to keep reading. If details such as the composition of raters and the agreement rate with human evaluation are disclosed, and results can be checked on tasks closer to real use, this ranking could become a more reliable guide for writers choosing which models to try.
