Rainshadow Systems

OpenAI Kills Sora, World Models Rise, and the Inference Hardware Crisis Nobody's Talking About

OpenAI Kills Sora, World Models Rise, and the Inference Hardware Crisis Nobody's Talking About

Three major stories shaped the AI landscape this past week. Each one points to a different — and sometimes contradictory — direction for where this technology is heading.

OpenAI shuts down Sora to bet on "Spud"

OpenAI officially wound down Sora, its AI video generation tool, to free up compute for a new project internally called "Spud." Sam Altman says Spud will be ready in weeks. The Sora team, led by Bill Peebles, is pivoting to "world simulation" for robotics — a signal that OpenAI sees physical AI and real-world simulation as more valuable than consumer video generation. This also puts their $1B Disney partnership on hold, a significant business decision that tells you where the real money is heading.

World Models are the new frontier

Yann LeCun's AMI Labs raised over $1 billion on the thesis that AI needs to understand the physical world, not just generate text about it. Meta's V-JEPA 2 achieved zero-shot robot planning after training on just 62 hours of domain-specific data. World Labs, also backed by over $1B, is approaching the problem from a spatial intelligence angle. Five distinct approaches to world models are emerging: JEPA, spatial intelligence, learned simulation, physical AI infrastructure, and active inference. The lines between them will blur fast, but the takeaway is clear: AI that can reason about physical reality is where the biggest bets are being placed.

The inference hardware crisis

A paper by Google researcher Xiaoyu Ma and Turing Award winner David Patterson laid out a cold reality: the hardware we're using to serve AI models was never designed for inference workloads. Training gets all the attention, but inference — actually running models at scale — is where the real bottleneck sits. SRAM-centric chips from Cerebras and Groq are gaining traction because they minimize latency by keeping data close to compute. Cerebras is now deploying on AWS Bedrock, offering 5x throughput improvements through a disaggregated architecture that pairs AWS Trainium for prefill with Cerebras WSE for decode. For anyone building AI-powered products, inference cost and speed will determine what's commercially viable.

What this means for small businesses: The shift from text generation to physical world understanding might seem abstract, but the downstream effects are concrete. Within 12-18 months, expect AI systems that can understand your warehouse layout, optimize your delivery routes by reasoning about physical space, and automate tasks that require spatial awareness — not just text processing. The inference cost improvements mean these capabilities will be accessible, not just available to companies with GPU clusters.

← All posts Work with us