Close Menu
Alpha Leaders
  • Home
  • News
  • Leadership
  • Entrepreneurs
  • Business
  • Living
  • Innovation
  • More
    • Money & Finance
    • Web Stories
    • Global
    • Press Release
What's On
Today’s NYT Connections Answers Explained: Friday, July 31

Today’s NYT Connections Answers Explained: Friday, July 31

30 July 2026
‘I’m just stuck’: Meet the former OpenAI researcher sitting on 0K of equity he says is overvalued

‘I’m just stuck’: Meet the former OpenAI researcher sitting on $700K of equity he says is overvalued

30 July 2026
How AI Is Complicating Federal Reserve Interest-Rate Decisions

How AI Is Complicating Federal Reserve Interest-Rate Decisions

30 July 2026
Facebook X (Twitter) Instagram
Facebook X (Twitter) Instagram
Alpha Leaders
newsletter
  • Home
  • News
  • Leadership
  • Entrepreneurs
  • Business
  • Living
  • Innovation
  • More
    • Money & Finance
    • Web Stories
    • Global
    • Press Release
Alpha Leaders
Home » What AI Is The Best? Chatbot Arena Relies On Millions Of Human Votes
Innovation

What AI Is The Best? Chatbot Arena Relies On Millions Of Human Votes

Press RoomBy Press Room18 July 20244 Mins Read
Facebook Twitter Copy Link Pinterest LinkedIn Tumblr Email WhatsApp
What AI Is The Best? Chatbot Arena Relies On Millions Of Human Votes

Topline

With companies like OpenAI, Google and Meta dropping increasingly sophisticated artificial intelligence products, crowdsourced rankings have emerged as a popular—and virtually only practical—way of determining which tool works best, and LMSYS’s Chatbot Arena has become possibly the most influential real-time gauge.

Key Facts

While most organizations choose to measure their AI models against a set of general capability benchmarks that cover tasks like solving math problems, programming challenges or answering multiple choice questions across an array of university-level disciplines, there is no industry benchmark or standard practice for assessing large language models (LLMs) like OpenAI’s GPT-4o, Meta’s Llama 3, Google’s Gemini and Anthropic’s Claude.

Even small differences to factors like datasets, prompts and formatting can have a huge impact on how a model performs, and when companies choose their own evaluation criteria, it can make it hard to fairly compare LLMs, Jesse Dodge, a senior scientist at the Allen Institute for AI in Seattle, told Forbes.

The difficulty in comparing LLMs is magnified given how closely leading models score on many commonly used benchmarks, with some companies and tech executives claiming victory over rivals with differences as narrow as 0.1%., so close it would likely go unnoticed by everyday users.

Community-built leaderboards deploying human insight have emerged, and in recent years their popularity has exploded in step with the steady boom of new AI tools like ChatGPT, Claude, Gemini and Mistral.

The Chatbot Arena, an open source project built by research group LMSYS and the University of California, Berkeley’s Sky Computing Lab, has proven particularly popular and has built AI leaderboards by asking visitors to compare responses from two anonymous AI models and vote which one is best.

Its scoreboards rank more than 100 AI models based on nearly 1.5 million human votes so far, covering an array of categories including long queries, coding, instruction following, maths, “hard prompts” and a variety of languages including English, French, Chinese, Japanese and Korean.

What’s The Best Ai Model On Chatbot Arena?

The top five AI models on Chatbot Arena’s overall leaderboard are:

  1. GPT-4o
  2. Claude 3.5 Sonnet
  3. Gemini Advanced
  4. Gemini 1.5 Pro
  5. GPT-4 Turbo

What To Watch For

Figuring out how to evaluate AI models is set to become increasingly important as more AI tools are rolled out and adopted across society. While benchmarks are important, Vanessa Parli, director of research at Stanford University’s Institute for Human-Centered AI, told Forbes they are also important as “goals for researchers to hit when developing models.” It is important to remember that “not all human capabilities are quantifiable” in a way that we can accurately measure but are nonetheless desirable to have in AI models, Parli said. There is also a clear need for benchmarks to assess traits like “bias, toxicity, truthfulness and other responsibility aspects,” especially for organizations dealing with sensitive information like healthcare companies, Parli said.

Crucial Quote

“The benchmarks aren’t perfect, but as of right now, that’s the primary mechanism we have to evaluate the models,” Parli told Forbes, cautioning that “researchers can somewhat easily game the system” today, with AI models quickly saturating benchmarks. “I think we need to get creative in the development of new ways to evaluate AI models,” Parli said. “

What We Don’t Know

Measuring intelligence is tricky when we do not know what it is we are supposed to be measuring. There is no universally accepted definition of intelligence in humans, let alone a way to measure it, and the possibility, nature and scope of animal intelligence has divided scientists for centuries. While AI benchmarks have typically focused on the ability to perform a particular task, more general assessments will be required in the near future as researchers make progress towards their goal of creating artificial general intelligence (AGI). AGI is capable of excelling and possibly matching humans across a broad set of domains rather than at just one task such as walking, moving boxes, identifying tumors on scans and playing chess.

How Useful Is Chatbot Arena For Evaluating Ai Models?

“The rankings [Chatbot Arena] gives are something that I trust more than most other rankings,” Dodge told Forbes, “because it uses a real human to say whether they prefer one generation over another.” Parli suggested assessments like Chatbot Arena could “implicitly evaluate factors” we want in our AI but are less quantifiable than something like coding ability. But she did stress that something like Chatbot Arena should not be the only evaluation method used, saying there are “many factors that should be important to organizations when evaluating models” and it “does not cover all of them.”

Get Forbes Breaking News Text Alerts: We’re launching text message alerts so you’ll always know the biggest stories shaping the day’s headlines. Text “Alerts” to (201) 335-0739 or sign up here.

Further Reading

Benchmark best AI chatbot arena ChatGPT Claude Gemini generative artificial intelligence LMSYS openAI
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link

Related Articles

Today’s NYT Connections Answers Explained: Friday, July 31

Today’s NYT Connections Answers Explained: Friday, July 31

30 July 2026
‘I’m just stuck’: Meet the former OpenAI researcher sitting on 0K of equity he says is overvalued

‘I’m just stuck’: Meet the former OpenAI researcher sitting on $700K of equity he says is overvalued

30 July 2026
How AI Is Complicating Federal Reserve Interest-Rate Decisions

How AI Is Complicating Federal Reserve Interest-Rate Decisions

30 July 2026
Has OpenAI already quietly hit pause on some AI development?

Has OpenAI already quietly hit pause on some AI development?

30 July 2026
Friday, July 31 Clues And Answers

Friday, July 31 Clues And Answers

30 July 2026
Claude Makes Five AI Labs Publishing Private Chats To Google

Claude Makes Five AI Labs Publishing Private Chats To Google

30 July 2026
Don't Miss
Exclusive: DeFi platform Azura launches after raising .9 million from Initialized

Exclusive: DeFi platform Azura launches after raising $6.9 million from Initialized

By Press Room22 October 2024

Azura, a new platform for decentralized finance, launched on Tuesday after raising $6.9 million in…

Unwrap Christmas Sustainably: How To Handle Gifts You Don’t Want

Unwrap Christmas Sustainably: How To Handle Gifts You Don’t Want

27 December 2024
Sam Altman’s World Wants To Scan Your Eyes To Prove You’re Human

Sam Altman’s World Wants To Scan Your Eyes To Prove You’re Human

22 October 2024
Stay In Touch
  • Facebook
  • Twitter
  • Pinterest
  • Instagram
  • YouTube
  • Vimeo
Latest Articles
Friday, July 31 Clues And Answers

Friday, July 31 Clues And Answers

30 July 20261 Views
Nearly a third of workers admit to sabotaging their company’s AI—smaller paychecks may explain why

Nearly a third of workers admit to sabotaging their company’s AI—smaller paychecks may explain why

30 July 20261 Views
Claude Makes Five AI Labs Publishing Private Chats To Google

Claude Makes Five AI Labs Publishing Private Chats To Google

30 July 20261 Views
Ultra-rich are buying up  million mansions in London, with ‘Trump unease’ fueling the influx

Ultra-rich are buying up $49 million mansions in London, with ‘Trump unease’ fueling the influx

30 July 20261 Views

Recent Posts

  • Today’s NYT Connections Answers Explained: Friday, July 31
  • ‘I’m just stuck’: Meet the former OpenAI researcher sitting on $700K of equity he says is overvalued
  • How AI Is Complicating Federal Reserve Interest-Rate Decisions
  • Has OpenAI already quietly hit pause on some AI development?
  • Friday, July 31 Clues And Answers

Recent Comments

No comments to show.
About Us
About Us

Alpha Leaders is your one-stop website for the latest Entrepreneurs and Leaders news and updates, follow us now to get the news that matters to you.

Facebook X (Twitter) Pinterest YouTube WhatsApp
Our Picks
Today’s NYT Connections Answers Explained: Friday, July 31

Today’s NYT Connections Answers Explained: Friday, July 31

30 July 2026
‘I’m just stuck’: Meet the former OpenAI researcher sitting on 0K of equity he says is overvalued

‘I’m just stuck’: Meet the former OpenAI researcher sitting on $700K of equity he says is overvalued

30 July 2026
How AI Is Complicating Federal Reserve Interest-Rate Decisions

How AI Is Complicating Federal Reserve Interest-Rate Decisions

30 July 2026
Most Popular
Has OpenAI already quietly hit pause on some AI development?

Has OpenAI already quietly hit pause on some AI development?

30 July 20261 Views
Friday, July 31 Clues And Answers

Friday, July 31 Clues And Answers

30 July 20261 Views
Nearly a third of workers admit to sabotaging their company’s AI—smaller paychecks may explain why

Nearly a third of workers admit to sabotaging their company’s AI—smaller paychecks may explain why

30 July 20261 Views

Archives

  • July 2026
  • June 2026
  • May 2026
  • April 2026
  • March 2026
  • February 2026
  • January 2026
  • December 2025
  • November 2025
  • October 2025
  • September 2025
  • August 2025
  • July 2025
  • June 2025
  • May 2025
  • April 2025
  • March 2025
  • February 2025
  • January 2025
  • December 2024
  • November 2024
  • October 2024
  • September 2024
  • August 2024
  • July 2024
  • June 2024
  • May 2024
  • April 2024
  • March 2024
  • February 2024
  • January 2024
  • December 2023
  • March 2022
  • January 2021
  • March 2020
  • January 2020

Categories

  • Blog
  • Business
  • Entrepreneurs
  • Global
  • Innovation
  • Leadership
  • Living
  • Money & Finance
  • News
  • Press Release
© 2026 Alpha Leaders. All Rights Reserved.
  • Privacy Policy
  • Terms of use
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.