A Creative Writing benchmark was recently published by Vulsar AI, evaluating how various LLMs perform against both amateur and professional human writers. With the proliferation of AI, various reports have mentioned that careers that cover various aspects of writing are at risk of being replaced by superior artificial intelligence models. However, based on 475 prompts, the new benchmark reveals that only frontier models can surpass humans in this regard, and that, too, amateur writers, not professional ones. GPT 6 Astra tops the “Overall” leaderboard, slightly outperforming amateur human writers The benchmark conducted by Vulsar AI evaluates where LLMs stand when […]
Read full article at https://wccftech.com/new-benchmark-pits-human-writers-against-24-llms-across-475-prompts/
