Raghav Ram: 40 LLMs, One Answer
Jim Griffin
When Language Models Start to Reason
You might have noticed that language models are acting quite a bit smarter now, including responses that look a lot like expert reasoning.
This video starts with a concrete example of this, where a bird species is identified by assembling available evidence and justifying the conclusion from multiple angles.
The video then goes on to identify five key developments that help to explain why language models perform so much better now at reasoning-type tasks.
NeurIPS 2025: Top 3 Highlights
This video covers 3 of the top papers at NeurIPS, 2025. All three of the papers covered won a Best Paper Award, and all three topics could have a direct impact on the AI world right away, so that’s how they were selected.
Willow Quantum Computer, Amazing Milestones
This video explains why this week’s announcement of Google’s Willow Quantum Computer is significant, starting with the fact that Willow was able to solve an exceedingly complex problem in under 5 minutes that the world’s fastest traditional supercomputer could not to solve at all.
AI 2025 Forecast: Agents Dominate
This video is our second annual forecast for key trends or developments most likely to define the coming year for AI in the private sector.
Inner Workings of OpenAI-o1? A First Glimpse
Since the architecture for the powerful o1 reasoning model from OpenAI has not been disclosed, there’s a lot of curiosity about how it works.
Andrew Ng at Snowflake: AI Agent Battle Royale
Andrew Ng was the keynote speaker last week on Day Two of the Snowflake BUILD conference, and in that talk, he shared results from testing different kinds of agentic workflows on the Human Eval benchmark.
Sequoia Capital: Move 37 is Here!
This is a special edition of the ‘AI World’ video series covering the release of OpenAI-o1 and its remarkable reasoning abilities.
How an 8B Model Beat an Industry Giant
This video describes how a system called ‘AgentStore’ was able to gain the top spot on a benchmark for AI agents – beating out a gigantic model with a small one.
Mesh Anything (except a Pink Hippo Ballerina)
The developers at MeshAnything have just released new code that offers an important improvement in how the surface of 3D objects can be encoded.
Can Robots Win at Table Tennis? Take a Look!
Google DeepMind has just achieved a new level of robotic skill – the ability to compete and win at table tennis.
Shark Alert! YOLO AI-Vision in Action
This video shows the capabilities of the AI-vision shark detection program, SharkEye, developed at the University of California, Santa Barbara.
AI Can do That?? Silver Medal in Pure Math
AI has just achieved a silver-medal-level performance in a globally-recognized competition in advanced mathematics: IMO 2004.
Will Open-Source Llama Beat GPT-4o?
Last week Meta launched its newest family of models, Llama 3.1, which includes a new benchmark.
Call a Doctor! --Blue Screen Lessons Learned
This video presents 12 top suspects concerning the largest IT outage in history caused by a faulty sensor configuration update.
Amazing Milestone! Million Experts Model
A top researcher at Google DeepMind has released an important paper detailing the first-known Transformer model with more than a million experts.
Behind the Curtain of Figma AI
This video summarizes the new AI features in Figma AI, discussed in under three minutes.
How a Language Model Aced a Top Leaderboard
This video shares details about an experiment by researchers in Tokyo studying the capabilities of large language models in improving their performance.
New Method Runs Big LLMs on Smartphones
There’s a breakthrough that handles large language models on smartphones called PowerInfer-2.
Nemotron-4 is BIG in More Ways than One
Last week, NVIDIA announced Nemotron-4 which consists of three models: Base, Instruct and Reward.