OpenAI just released two new AI models that can 'think' more like humans—but what does that actually mean?
The episode begins with the hosts acknowledging the rapid pace of AI news and the need for a catch-up. They introduce OpenAI's two new models, O1 Preview and O1 Mini, which are designed for enhanced reasoning in specialized areas like science, programming, and mathematics. The hosts explain that current language models lack true logical reasoning and planning abilities, which these new models aim to address by using more computing power during inference, making responses slower and more expensive. They clarify that these models are not for internet search or file uploads but for deep, focused tasks. A key example compares GPT-4 and O1's responses to a short story about Karl leaving his house during a thunderstorm. Both models fail to deduce that the car must be in a garage or under a carport to stay dry, highlighting a limit in everyday common sense despite O1's deeper textual analysis. The hosts note O1's superior performance on math Olympiad tasks (83% vs. GPT-4's 13%). The discussion shifts to OpenAI's corporate news, including the departure of chief technology officer Mira Murati and reports that OpenAI is restructuring to become more commercially oriented, potentially reversing its original non-profit governance to attract investors. They detail OpenAI's financials: $3.6 billion in annualized revenue but over $5 billion in annual training costs, leading to constant fundraising needs. The hosts compare OpenAI's position to big tech competitors like Google and Meta, noting its 'newcomer' advantage in tolerating mistakes. They mention other AI models like Claude from Anthropic that are also improving. The episode concludes with a historical anecdote from a book recommendation, 'Why Machines Learn,' about a 1958 New York Times article that called an early AI model a 'computer embryo,' illustrating recurring cycles of AI hype and exaggerated expectations. The hosts reflect on the mathematical foundations of AI and its limits, tying it back to the reasoning discussion.