32

Each LLM is given the same 1000 chess puzzles to solve. See puzzles.csv. Benchmarked on Mar 25, 2024.

Model Solved Solved % Illegal Moves Illegal Moves % Adjusted Elo
gpt-4-turbo-preview 229 22.9% 163 16.3% 1144
gpt-4 195 19.5% 183 18.3% 1047
claude-3-opus-20240229 72 7.2% 464 46.4% 521
claude-3-haiku-20240307 38 3.8% 590 59.0% 363
claude-3-sonnet-20240229 23 2.3% 663 66.3% 286
gpt-3.5-turbo 23 2.3% 683 68.3% 269
claude-instant-1.2 10 1.0% 707 66.3% 245
mistral-large-latest 4 0.4% 813 81.3% 149
mixtral-8x7b 9 0.9% 832 83.2% 136
gemini-1.5-pro-latest* FAIL - - - -

Published by the CEO of Kagi!

you are viewing a single comment's thread
view the rest of the comments
[-] General_Effort@lemmy.world 2 points 9 months ago

Depends on circumstances, obviously.

[-] bionicjoey@lemmy.ca 1 points 9 months ago

Okay. What if the circumstance is because I'm just recalling a bunch of chess puzzle solutions I've seen before and regurgitating the one I think is the correct solution for this particular pizzle without really understanding the rules of chess?

[-] General_Effort@lemmy.world 1 points 9 months ago

That's another thing I'm wondering about, but so is anyone. I'd still want to know why GPT-4 does so much better than the others.

this post was submitted on 25 Mar 2024
32 points (75.8% liked)

Technology

60148 readers
2039 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


founded 2 years ago
MODERATORS