> ## Content Index
> Fetch the complete content index at: https://birchtree.me/llms.txt
> Use this file to discover other available public pages before exploring further.

# Claude tops GPT 4 in blind comparisons
- URL: https://birchtree.me/blog/claude-tops-gpt-4-in-blind-comparisons/
- Published: 2024-03-29T15:00:00.000Z
- Updated: 2024-03-29T15:00:00.000Z
- Author: Matt Birchler
- Tags: Links, #Import 2025-03-29 19:29

Benj Edwards writing for ArsTechnica: [“The King Is Dead”—Claude 3 Surpasses GPT-4 on Chatbot Arena for the First Time](https://arstechnica.com/information-technology/2024/03/the-king-is-dead-claude-3-surpasses-gpt-4-on-chatbot-arena-for-the-first-time/?ref=birchtree.me)

> On Tuesday, Anthropic's [Claude 3 Opus](https://arstechnica.com/information-technology/2024/03/the-ai-wars-heat-up-with-claude-3-claimed-to-have-near-human-abilities/?ref=birchtree.me) large language model (LLM) surpassed OpenAI's GPT-4 (which powers ChatGPT) for the first time on [Chatbot Arena](https://arstechnica.com/ai/2023/12/turing-test-on-steroids-chatbot-arena-crowdsources-ratings-for-45-ai-models/?ref=birchtree.me), a popular crowdsourced [leaderboard](https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboard?ref=birchtree.me) used by AI researchers to gauge the relative capabilities of AI language models.

There are all sorts of chatbot leaderboards on the web, and I don’t take many of them very seriously. I guess it’s academically interesting how well each model does in standardized tests, but I don’t personally find that a super compelling data point that helps me figure out which model is going to be most useful for the sorts of things I do with them.

But Hugging Face’s leaderboard is more interesting and you can contribute to the votes yourself by using their [testing page](https://chat.lmsys.org/?ref=birchtree.me) to ask your own questions and compare the results from two unnamed models. I played with this for a bit today and my big takeaways were:

1. I also preferred Claude’s responses almost every single time it was in the comparison.
2. There’s something called “command-r” that lost in every single comparison it appeared in. Woof.

Oh, and you can cheat if you really want to 😉

![](https://storage.ghost.io/c/25/31/25312b2a-2d5b-4d01-ba92-e6755efc202c/content/images/2024/03/CleanShot-2024-03-29-at-09.16.32.png)