Grok-3 Beats DeepSeek-R1 At Reasoning, Is As Capable As OpenAI’s O1 Pro: Karpathy

xAI, the AI model maker headed by Elon Musk, unveiled its latest family of models, the Grok-3.

According to benchmarks, the Grok-3 outperforms several competing models and is also the first to score over 1400 on Chatbot Arena, a platform for comparing and evaluating AI models.

Grok-3 also offers reasoning (Think) capabilities and a deep research feature called DeepSearch.

Andrej Karpathy, founder of Eureka Labs, who was also once a part of OpenAI and Tesla, was given early access to Grok-3.

He shared a post on X detailing his experience. He revealed that the model performed well on complex tasks, such as creating a hex grid for the popular board game Settlers of Catan.

“Few models get this right reliably. The top OpenAI thinking models (e.g. o1-pro, at $200/month) get it too, but all of DeepSeek-R1, Gemini 2.0 Flash Thinking, and Claude do not,” he said.

Karpathy also uploaded OpenAI’s GPT-2 technical paper to estimate the number of flops required to train the model. He revealed that while Grok-3 and GPT-4o failed at this task, Grok-3, with thinking (reasoning), solved it ‘great’, and even OpenAI’s o1 Pro failed at the task.

“The impression overall I got here is that this is somewhere around o1-pro capability, and ahead of DeepSeek-R1, though, of course, we need actual, real evaluations to look at,” he added.

Karpathy also tested Grok-3’s DeepSearch capabilities, which he found comparable to Perplexity’s deep research but not yet at the level of that offered by OpenAI. He found that the model was hallucinating URLs that do not exist and reporting incorrect facts without providing citations.

“When I asked it to create a report on the major LLM labs and their amount of total funding and estimate of employee count, it listed 12 major labs but not itself (xAI),” he added.

After using the model for around 2 hours, he concluded by saying, “Grok 3 + thinking feels somewhere around the state of the art territory of OpenAI’s strongest models (o1-pro, $200/month), and slightly better than DeepSeek-R1 and Gemini 2.0 Flash Thinking.”

Others like Lex Fridman, who also received early access to the model, said, “My mind is blown, very impressive model,” in a post on X.

Source link

What's Hot

This Underrated Artificial Intelligence (AI) Stock Just Posted Triple-Digit AI Growth for an 8th Straight Quarter

OpenAI Hopes Animated ‘Critterz’ Will Prove AI Is Ready for the Big Screen

Google’s former security leads raise $13M to fight email threats before they reach you

Grok-3 Beats DeepSeek-R1 at Reasoning, is as Capable as OpenAI’s o1 Pro: Karpathy

AI researcher Andrej Karpathy says he’s “bearish on reinforcement learning” for LLM training

Ex-OpenAI scientist, Andrej Karpathy, is “bearish on reinforcement learning” in the long-term

Tesla’s Vision-Only Autonomous Driving: Karpathy’s Data-Driven Bet

Christie’s Will Auction The First Calculating Machine In History

The Art Market Isn’t Dying. The Way We Write About It Might Be.

Banksy Mural of Judge Beating Protestor Removed by Courts Service

Death of Matthew Christopher Pietras Ruled a Suicide

This Underrated Artificial Intelligence (AI) Stock Just Posted Triple-Digit AI Growth for an 8th Straight Quarter

OpenAI Hopes Animated ‘Critterz’ Will Prove AI Is Ready for the Big Screen

Google’s former security leads raise $13M to fight email threats before they reach you

What's Hot

Grok-3 Beats DeepSeek-R1 at Reasoning, is as Capable as OpenAI’s o1 Pro: Karpathy

Related Posts

Subscribe to Updates