Raghav Ram: 40 LLMs, One Answer

Jim Griffin

When Language Models Start to Reason

You might have noticed that language models are acting quite a bit smarter now, including responses that look a lot like expert reasoning.

This video starts with a concrete example of this, where a bird species is identified by assembling available evidence and justifying the conclusion from multiple angles.

The video then goes on to identify five key developments that help to explain why language models perform so much better now at reasoning-type tasks.

NeurIPS 2025: Top 3 Highlights

This video covers 3 of the top papers at NeurIPS, 2025. All three of the papers covered won a Best Paper Award, and all three topics could have a direct impact on the AI world right away, so that’s how they were selected.

Willow Quantum Computer, Amazing Milestones

This video explains why this week’s announcement of Google’s Willow Quantum Computer is significant, starting with the fact that Willow was able to solve an exceedingly complex problem in under 5 minutes that the world’s fastest traditional supercomputer could not to solve at all.

AI 2025 Forecast: Agents Dominate

This video is our second annual forecast for key trends or developments most likely to define the coming year for AI in the private sector.

Inner Workings of OpenAI-o1? A First Glimpse

Since the architecture for the powerful o1 reasoning model from OpenAI has not been disclosed, there’s a lot of curiosity about how it works.

Andrew Ng at Snowflake: AI Agent Battle Royale

Andrew Ng was the keynote speaker last week on Day Two of the Snowflake BUILD conference, and in that talk, he shared results from testing different kinds of agentic workflows on the Human Eval benchmark.

Sequoia Capital: Move 37 is Here!

This is a special edition of the ‘AI World’ video series covering the release of OpenAI-o1 and its remarkable reasoning abilities.

How an 8B Model Beat an Industry Giant

This video describes how a system called ‘AgentStore’ was able to gain the top spot on a benchmark for AI agents – beating out a gigantic model with a small one.

Mesh Anything (except a Pink Hippo Ballerina)

The developers at MeshAnything have just released new code that offers an important improvement in how the surface of 3D objects can be encoded.

Can Robots Win at Table Tennis? Take a Look!

Google DeepMind has just achieved a new level of robotic skill – the ability to compete and win at table tennis.

Shark Alert! YOLO AI-Vision in Action

This video shows the capabilities of the AI-vision shark detection program, SharkEye, developed at the University of California, Santa Barbara.

AI Can do That?? Silver Medal in Pure Math

AI has just achieved a silver-medal-level performance in a globally-recognized competition in advanced mathematics: IMO 2004.

Will Open-Source Llama Beat GPT-4o?

Last week Meta launched its newest family of models, Llama 3.1, which includes a new benchmark.

Call a Doctor! --Blue Screen Lessons Learned

This video presents 12 top suspects concerning the largest IT outage in history caused by a faulty sensor configuration update.

Amazing Milestone! Million Experts Model

A top researcher at Google DeepMind has released an important paper detailing the first-known Transformer model with more than a million experts.

Behind the Curtain of Figma AI

This video summarizes the new AI features in Figma AI, discussed in under three minutes.

How a Language Model Aced a Top Leaderboard

This video shares details about an experiment by researchers in Tokyo studying the capabilities of large language models in improving their performance.

New Method Runs Big LLMs on Smartphones

There’s a breakthrough that handles large language models on smartphones called PowerInfer-2.

Nemotron-4 is BIG in More Ways than One

Last week, NVIDIA announced Nemotron-4 which consists of three models: Base, Instruct and Reward.