<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Root Access]]></title><description><![CDATA[Make something people want.]]></description><link>https://www.ycrootaccess.com</link><image><url>https://substackcdn.com/image/fetch/$s_!L33H!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa733dd10-ea8b-4137-9673-83e3dda0fb78_1024x1024.png</url><title>Root Access</title><link>https://www.ycrootaccess.com</link></image><generator>Substack</generator><lastBuildDate>Sun, 26 Jul 2026 00:13:20 GMT</lastBuildDate><atom:link href="https://www.ycrootaccess.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Root Access]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[ycrootaccess@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[ycrootaccess@substack.com]]></itunes:email><itunes:name><![CDATA[Root Access]]></itunes:name></itunes:owner><itunes:author><![CDATA[Root Access]]></itunes:author><googleplay:owner><![CDATA[ycrootaccess@substack.com]]></googleplay:owner><googleplay:email><![CDATA[ycrootaccess@substack.com]]></googleplay:email><googleplay:author><![CDATA[Root Access]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[How OpenCode became the world's most popular open source coding agent]]></title><description><![CDATA[Jay V, OpenCode's CEO, applied to YC nine times and worked on startups for almost 20 years]]></description><link>https://www.ycrootaccess.com/p/how-opencode-became-the-worlds-most</link><guid isPermaLink="false">https://www.ycrootaccess.com/p/how-opencode-became-the-worlds-most</guid><pubDate>Fri, 24 Jul 2026 14:02:45 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/_O6x4ktK6JA" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-_O6x4ktK6JA" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;_O6x4ktK6JA&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/_O6x4ktK6JA?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><br>Since the start of the year, OpenCode &#8212; an open source alternative to Claude Code and Codex that works with any model &#8212; exploded to 4.6 million weekly active users, 13 million monthly actives, and roughly $40M in annualized revenue.</p><p>In this episode of The Lightcone, Harj, Jared, and Diana talk with Jay V, OpenCode&#8217;s CEO, about what&#8217;s driving this wild growth, the Anthropic clampdown that inadvertently fueled it, and the almost 20 year founder journey that led him here. </p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.ycrootaccess.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><a href="https://youtu.be/_O6x4ktK6JA">Watch on YouTube</a></p><h3><strong>Timestamps</strong></h3><p>00:44 &#8212; OpenCode&#8217;s Explosive Growth</p><p>01:16 &#8212; 20x Growth, 13M Users, and 7 Trillion Tokens</p><p>03:39 &#8212; The Anthropic Controversy That Changed Everything</p><p>05:43 &#8212; Bringing AI Coding Agents to the World</p><p>06:39 &#8212; When Open Source Models Became Good Enough</p><p>08:56 &#8212; What Millions of Developers Are Actually Using</p><p>13:31 &#8212; Why OpenCode Is Huge Outside the US</p><p>15:27 &#8212; Why Fortune 500 Companies Choose OpenCode</p><p>16:36 &#8212; The Economics of AI Tokens</p><p>20:02 &#8212; How Enterprises Are Using Coding Agents</p><p>22:58 &#8212; AI&#8217;s New Unit Economics</p><p>24:56 &#8212; Why Model Choice Matters</p><p>29:55 &#8212; The Product Decisions Behind OpenCode</p><p>34:21 &#8212; A 16-Year Overnight Success</p><p>41:16 &#8212; Why Jay Never Gave Up</p><h3><strong>Transcript</strong></h3><p><strong>Harj:</strong> Welcome back to another episode of The Lightcone. Garry&#8217;s out traveling today and will be back next episode. Our guest today is Jay V, founder and CEO of OpenCode, an open source alternative to Claude Code that works with any model you want. OpenCode has been growing at an astounding rate this year. They&#8217;re now at 4.6 million weekly active users, which is actually pretty close to Codex&#8217;s numbers. Today we&#8217;re going to talk about what&#8217;s driving their wild growth and also Jay&#8217;s winding road to get here since the company went through YC back in 2021. Jay, thanks so much for being here.</p><p><strong>Jay:</strong> Thank you. Thank you for having me.</p><p><strong>Harj:</strong> Why don&#8217;t we just start with the crazy scale you guys are at. Maybe tell us about any stats you can share with us.</p><p><strong>Jay:</strong> Yeah. You mentioned the weekly actives. We like to post a lot of our metrics on Twitter. Our monthly actives&#8212;I think we ended June with around 13 million or so, and that&#8217;s around a 20X increase since the beginning of the year. We also recently started processing around seven trillion tokens per day. For context, OpenRouter does a total of around six trillion. I think at the beginning of the year, we were probably at around 300 billion or so. In terms of our revenue from our subscription product, if you were to pay per token with our inference and you took, let&#8217;s say, June&#8217;s data and extrapolated for the year, that would be around 31 to 33 million or so. If you took last week&#8217;s data and extrapolated for the year, that&#8217;s around 38 to almost 40 million. And that&#8217;s just the inference part of our business.</p><p>So the way we make money&#8212;we launched that, call it end of September, early October last year&#8212;so about eight months to getting to around 40 million or so. To round out the numbers, our subscribers&#8212;these are people that pay for a monthly subscription with OpenCode&#8212;we launched that product early March, I think end of February, and that&#8217;s grown to around 160,000 monthly subscribers. That accounts for about 18 million of the annualized revenue.</p><p><strong>Harj:</strong> I saw a tweet from&#8212;I think it&#8217;s Thibault, I think he&#8217;s a lead engineer, at least one of the main engineers on Codex&#8212;saying that something like 5% of all Codex subscribers choose OpenCode as the main harness to actually use the underlying API.</p><p><strong>Jay:</strong> Yeah. So Codex officially supported OpenCode. What that basically means is that you can use Codex&#8217;s subscription in OpenCode, and a bunch of their users use OpenCode directly to take advantage of their subscription plan.</p><p><strong>Harj:</strong> And I think they did that right after some of the back and forth you had with Anthropic and Claude Code. Tell us about what happened and how that seemed like it really fueled your growth&#8212;was an inflection point for you guys.</p><p><strong>Jay:</strong> Yeah. To contextualize this, we were at around 650,000 monthly active users at the beginning of the year. The first week of January, we started to hear some rumblings around Anthropic trying to clamp down on people using OpenCode but using Claude Code subscriptions on there. This was a very common way to use the Claude Code subscription. The way they tried to block it was if the system prompt mentioned literally the word &#8220;OpenCode,&#8221; they would reject the request. From our perspective, that sort of makes sense&#8212;they&#8217;re subsidizing usage, that&#8217;s what they want to do. But of course, a lot of users weren&#8217;t happy. When they did that, what it inadvertently did was put OpenCode and Claude Code on the same pedestal.</p><p>It equated the two products in some ways. Even for the people who weren&#8217;t using OpenCode at the time, they took notice of the fact that Claude Code was taking that kind of action against OpenCode.</p><p><strong>Harj:</strong> Oh, so you think you actually got new users because people heard about you for the first time and they&#8217;re like, oh, Claude Code&#8217;s banning this thing&#8212;</p><p><strong>Jay:</strong> &#8212;or clamping&#8212;</p><p><strong>Harj:</strong> &#8212;down on this thing.</p><p><strong>Jay:</strong> Yeah, or that it&#8217;s worth looking into, that it&#8217;s not just one of the other dozen or so coding agents out there.</p><p><strong>Harj:</strong> It&#8217;s funny how often this happens in startups. Same thing happened with Instacart when Amazon bought Whole Foods. It was like, oh, this is the death of Instacart. And the death of Instacart became this meme, but the meme actually just drove all the grocers to check out what Instacart was. Then they went through this explosive growth of signing up every grocer in America. So it seems like actually a lot of your growth is global across the world. Tell us a bit about that.</p><p><strong>Jay:</strong> Yeah. So the premise of the product is that most people in the world still haven&#8217;t experienced the magic of a coding agent. It&#8217;s been almost a year, and I&#8217;m sure it&#8217;s hard to remember for you guys as well, but the first time you had this experience, it&#8217;s a magical experience. When we felt that for ourselves, we recognized how important that moment is in tech history. It comes along once every generation or so. The way we looked at that was, let&#8217;s take that to as many people in the world as possible, have them experience something similar, because the frontier models and the frontier labs charge so much per token that it&#8217;s going to be hard for a lot of people across the globe to have that experience. We wanted to make sure they had that with us in some ways.</p><p><strong>Harj:</strong> I guess at a point early on, was there a big gap between the open source models and the frontier models? Was that true when you were first launching the product?</p><p><strong>Jay:</strong> Yeah, for sure. I think when we first launched, it was mostly, &#8220;Hey, you&#8217;re using your Claude Code subscription with Claude Code. Come try that with OpenCode.&#8221; This was back in June of last year, and by about August or September, we started to see the first crop of open source models. You had this sense that, &#8220;Oh, they&#8217;re maybe six months behind.&#8221; Of course, there&#8217;s always been a gap between them, but that was the first point when people were like, &#8220;There&#8217;s the GLMs of the world, the Kimis of the world, the MiniMaxes of the world, and it seems like there is now an alternative.&#8221; As that started to happen, it triggered a wave of users coming in to try out OpenCode, because that was probably the only way you could try out some of these multiple models. As that gap started to shrink, or as the open weight models became good enough for real work, it became viable to use OpenCode with them.</p><p><strong>Harj:</strong> I think I remember a few months back you were saying that, let me get this right, even though the open source models are obviously cheaper to use through OpenCode, you still saw more usage from the leading frontier model&#8212;except was it when KIMI 2.5 came out?</p><p><strong>Jay:</strong> Yeah.</p><p><strong>Harj:</strong> That was the first time that it was equivalent or it flipped?</p><p><strong>Jay:</strong> Yeah, that&#8217;s exactly it. I think it was probably 2.4&#8212;I forget the exact model&#8212;but yeah, this was February of this year where for a four-week span, and this had never happened in our data before. I think we&#8217;re fortunate to be able to see this global usage. A lot of times we see the comparison of these open source models versus some of the frontier ones. We noticed for the first time that a bunch of users were just using KIMI way more than they were using the Anthropic models. It was Sonnet plus Opus at the time together. That&#8217;s the point when we were like, &#8220;Oh, I think we should launch a subscription product because now maybe these models actually do make sense for real work.&#8221;</p><p><strong>Harj:</strong> Speaking of that, you guys have this incredibly unique insight. You have the best data on how these models are being used by engineers across the whole world, and you released a bunch of this data. Maybe we could just look through some of it and pluck out some more interesting insights.</p><p><strong>Jay:</strong> Yeah. You can go to <a href="http://opencode.ai/data">opencode.ai/data</a>. We started to publish this about a month ago. This is basically taking all of the usage on OpenCode Go, which is our subscription plan where you pay $10 to use any of these open source models. This breakdown specifically is by token volume per day across the different models. What we see here is that DeepSeek Flash is the one that is used a lot. There are some details here&#8212;I can go into why that is the case. But if you just look at the top three, we&#8217;re seeing the two DeepSeeks plus GLM, and you can see all the hype that GLM has been getting lately being reflected here in these charts.</p><p><strong>Diana:</strong> Which the data says otherwise, as opposed to all the Twitter chatter about GLM taking over DeepSeek. This is telling a different story.</p><p><strong>Jay:</strong> Yeah, it is. I can show you a different breakdown here. This is by unique users, and this was a thing I heard you were alluding to. We get to see actual usage data for each of the users, as opposed to maybe an OpenRouter where you&#8217;re seeing it aggregated across a bunch of services or other products that are internally using it. In this case, each one of these is an actual user. What&#8217;s fascinating is if we look at the market share graph&#8212;this is breaking down for each of these labs the amount of token volume they&#8217;re doing per day, comparing that across them&#8212;you can see the DeepSeek one dip around the time GLM comes out, but it seems to bounce back up after. There are some theories around why that&#8217;s the case, but that&#8217;s an interesting fact.</p><p>And then if I go back up to the users one, this is also kind of fascinating in that it might be a little bit hard to see here, but if you look at the top three unique users per model, you&#8217;ve got Flash at 38K, DeepSeek Pro at 31K, and GLM 5.2 at close to 30K as well. That&#8217;s interesting because GLM is on par with one of the DeepSeek models, but the fact that there are two of them makes it a little bit different. I&#8217;ll caveat this by saying that DeepSeek Flash being as cheap as it is allows a lot of users to extend how much they can use a coding agent, because as they get closer to their daily or weekly limits, they can switch over to one of these very cheap models&#8212;in this case, DeepSeek Flash&#8212;to get the rest of their work done.</p><p>Again, this is very different from the way we think about coding agents and LLMs here in the Valley and SF, and in the West in general.</p><p><strong>Harj:</strong> What are you seeing? We definitely are at the other extreme end of where it&#8217;s like&#8212;</p><p><strong>Jay:</strong> Token maximum max is towing money.</p><p><strong>Harj:</strong> But for the users you see, which to be fair, it does seem like there&#8217;s a general vibe shift, especially in the enterprise world towards more token budgeting. What do you see? Is it as simple as people, once they approach their usage limits, switch over to one of these models, or is there more going on?</p><p><strong>Jay:</strong> Yeah, there&#8217;s a few different things. I think people do try and optimize for things. One of the reasons why originally Kimi had taken off was it was being hosted in a way that the tokens per second was a lot higher than what you would get out of even Opus. So it was just a drastically faster model. It felt like you were working almost in real time. Some of these models tend to exhibit those qualities, which makes it characteristically a little bit different from using some of the frontier ones, for example. So there&#8217;s a little bit of that, but cost is obviously a big driver. Another thing that pops up every once in a while is people get a sense that GLM 5.0, for example, is better at front-end design compared to some of the other models. That ends up driving some usage as well.</p><p><strong>Jared:</strong> You also have a really interesting breakdown by geo to show where the users are. So who are your users? Where are they coming from?</p><p><strong>Jay:</strong> Yeah, so some context here is that we had launched the OpenCode Go plan to be able to serve the global audience. You can see that with China being number one at 17%.</p><p><strong>Jared:</strong> That&#8217;s crazy. I feel like you&#8217;re probably the only YC company in history that has meaningful usage in China.</p><p><strong>Jay:</strong> It&#8217;s also fascinating because a lot of these models are Chinese. In their situation, they&#8217;re trying to use the ones that are being built in their country, and OpenCode gives them the choice to do that. That&#8217;s interesting as well. I think the US is actually interesting to us because when we built this plan originally, we weren&#8217;t thinking about the US. We weren&#8217;t thinking about the States because, as we were saying, people here just throw money at it. But it turns out it&#8217;s growing really well in the States too. Maybe that speaks to the vibe shift of being a little more conscious with tokens. But there&#8217;s also the other side, where people want to use some of these models. So when GLM 512 is getting popular, a lot of people are trying out our subscription plan because it&#8217;s one of the options to do that.</p><p><strong>Jared:</strong> So because it&#8217;s so much cheaper, it&#8217;s huge in developing countries&#8212;like Indonesia is 4% of your traffic, Brazil is 5% of your traffic, places like Vietnam, places where a $200 a month Claude Code subscription is very expensive. That makes sense. But what you were telling me earlier was that in addition to that, there are a lot of large US companies that have basically unlimited budgets for tokens that are also using OpenCode. Can you talk about that&#8212;who&#8217;s using you and why those people are using you?</p><p><strong>Jay:</strong> Yeah, this is what we had seen early on before some of these open source models even took off: a lot of companies would start using OpenCode because they didn&#8217;t want to be locked into using a specific model or a specific harness. Some users just wanted more choice, and this was a good neutral option for them that allowed them the flexibility in the future to switch to whatever they wanted.</p><p><strong>Diana:</strong> And I think you had a crazy stat. It was something like a dozen of the top very forward Fortune 500 companies are using you and have a significant footprint.</p><p><strong>Jay:</strong> Yeah. It&#8217;s funny&#8212;we get DMs, we get emails internally about, &#8220;Hey, XYZ company here. We&#8217;ve got a few thousand people using OpenCode now. Please don&#8217;t share this publicly.&#8221; But I think most of what&#8217;s going on is there&#8217;s just a bunch of choice there. We can probably talk about this later, but we&#8217;re very intentional with our product design. We want something that everybody uses every single day, and we hold that bar fairly high. In those instances, I think that&#8217;s what&#8217;s resonating with people as opposed to just the cheaper tokens thing.</p><p><strong>Diana:</strong> I&#8217;m very curious on a slightly different topic. There&#8217;s been a shift for all these companies like you in terms of the AI token economics. In the old world of companies&#8212;other B2C or B2B companies&#8212;a lot of CAC used to be based on ads. Now the equation is different; it&#8217;s based on tokens to acquire users. But even more weird, there&#8217;s this thing you&#8217;re describing where there are some whales that pay for a lot of it at some point. There might be a lot of churn, but it doesn&#8217;t matter because as long as you have the power users and experts really using it, they convert these large orgs. I think the big labs can subsidize that, which is what effectively Claude Code and Codex have done. They can subsidize it, but you have a magic formula to skip all this.</p><p><strong>Jay:</strong> The broader context here is that for you to use AI&#8212;and especially these coding agents&#8212;because of the amount of tokens they use, to use them well, you have to really understand them. This is the experts thing you&#8217;re talking about. And to get there is fairly expensive from a per-token perspective.</p><p><strong>Diana:</strong> Token maxing is expensive.</p><p><strong>Jay:</strong> Right. It&#8217;s very expensive. And that is a chasm that is very hard for a lot of people and a lot of companies to cross. What the Anthropics and the OpenAIs of the world do is subsidize it so that people are able to cross that. The whales idea here is that a certain percentage of them cross it, get to that point, start spending the crazy amounts that you see, and then it makes sense. The entire funnel then makes sense. When we approached this, we thought about it from a product perspective. We talked about wanting everybody in the world to have that aha moment with a coding agent. So that&#8217;s our free tier. And then, because the open source models were now cheap enough and good enough for real work, you want a subscription plan that allows them to do real work with it.</p><p>And that is the thing that allows somebody to buy into, &#8220;Okay, now I can justify spending so much more to potentially redo some of the processes within my company to take advantage of these coding agents.&#8221; And that&#8217;s when hopefully some of these turn into whales.</p><p><strong>Harj:</strong> And that&#8217;s what you&#8217;re seeing. You&#8217;re seeing people come in through the free tier to try it out and then become advocates, saying, &#8220;Oh yeah, we should adopt this at our Fortune 500 company.&#8221;</p><p><strong>Jay:</strong> Yeah. In the older SaaS enterprise world, you would have this procurement process that a lot of them go through, where someone in a region starts the whole dance. In our case, the inbounds that we get are typically just, &#8220;Hey, there are a bunch of people at our company using you guys. Can you fill out the security questionnaire?&#8221; And we&#8217;re just like, &#8220;Oh, first off, I didn&#8217;t know we had users there. But secondarily, give us a second.&#8221;</p><p><strong>Harj:</strong> That&#8217;s when you really know you have product-market fit&#8212;when enterprises are bugging you</p><p><strong>Jay:</strong> To sign the</p><p><strong>Harj:</strong> Security</p><p><strong>Jay:</strong> Agreement</p><p><strong>Harj:</strong> So they can</p><p><strong>Jay:</strong> Use your product. Just like, &#8220;Please do this because I don&#8217;t know who you guys are, but a bunch of us are using this.&#8221; So it is a little bit different now.</p><p><strong>Harj:</strong> Once you&#8217;re through the procurement and admin side of things, what are the enterprises asking you for? Are they trying to pull the product in a different direction? Because that often tends to be something that open source companies have to think through and be careful about.</p><p><strong>Jay:</strong> What&#8217;s interesting here is these coding agents are at the core of how an LLM does work. A lot of times when we get these enterprises reaching out to us, it&#8217;s because they&#8217;re trying to figure out where else they can use it. There&#8217;s the obvious one: we&#8217;ve got a bunch of developers at our company using OpenCode, just figuring out how we can officially use it. And then on the flip side, there are some non-technical people who would want to be able to use this as well. Secondarily, our product is probably going to be using a coding agent as part of its core loop. Can we use it there? So we get some pull there. The other bit of pull that we see is mostly around just managing tokens a little bit, with people saying, &#8220;Hey, there are certain organizations within our company that don&#8217;t need the frontier models. Can we limit access there or have some more creative ways of managing token spend?&#8221; We see pull there. So there are some questions around that. The strange one we got recently was somebody wanting to have really good visibility of exactly what everybody&#8217;s doing at the company. That gets into questions of whether that&#8217;s something we want to build.</p><p><strong>Harj:</strong> Didn&#8217;t Ramp use you in an interesting way?</p><p><strong>Jay:</strong> Yeah, I think Ramp was very forward-thinking. They published a blog post&#8212;I think this was in December of last year&#8212;but they had reached out prior to that, where a team within Ramp had built this Slack bot that was essentially running OpenCode behind the scenes. It was just incredible to see. It was mind-blowing, partly because we hadn&#8217;t even done that yet internally, and they were showing off a use case that was definitely pointing towards the future.</p><p><strong>Harj:</strong> What does that mean exactly? I think it can be hard for people to get their head around because they think of you as an open source Claude Code. So what does that mean that their Slack bot was running on OpenCode?</p><p><strong>Jay:</strong> Yeah, so you can think of OpenCode as a two-part product. There is the UI and the application part that you see and interact with, but then there is the agent loop&#8212;the thing that&#8217;s actually doing work while calling the LLM. That is behind the scenes. We call it the server. That server can be embedded separately from the UI. In this case, they were embedding that and running their Slack bot off of it.</p><p><strong>Diana:</strong> Can you tell us a bit about these numbers and how the unit economics work? There was a shocking stat mentioned in Dylan Patel&#8217;s podcast that Anthropic is profitable and by a huge margin&#8212;in Q2 they are on track to be doing $50 billion annualized revenue and around 70% margin. But the previous year was not profitable and they crossed this chasm, which sounds like where you&#8217;re heading, but you don&#8217;t have to subsidize it, which is special.</p><p><strong>Jay:</strong> Yeah, we do subsidize a little bit, the subscription plan that we have. And you</p><p><strong>Jared:</strong> Have a free tier too.</p><p><strong>Jay:</strong> And we have a free tier. Yeah. I think that&#8217;s the CAC part you were talking about early on where&#8212;</p><p><strong>Diana:</strong> Rather than paying for ads, the CAC is paying for tokens.</p><p><strong>Jay:</strong> Yeah, it&#8217;s because that&#8217;s how people experience that kind of magic moment. That&#8217;s the way we get them into using the product, understanding what a coding agent is. All this stuff is definitely a little too much for a lot of people. I think the part here in terms of the unit economics that works out is if you&#8217;re actually paying per token&#8212;in the case of Anthropic and in the case of how we operate as well&#8212;we&#8217;re able to get, at least in our scenario, volume discounts because of the amount of tokens that we serve. So when you pay per token, that effectively turns into our margin. Whereas when you&#8217;re subsidizing, obviously you&#8217;re eating the cost there, or in the case of the free tier, you&#8217;re eating the cost as well. But as you were talking about before, as you start to get more and more of these whales, those whales are paying per token and it&#8217;s directly feeding into your margins.</p><p><strong>Harj:</strong> Yeah. The discounts you&#8217;re able to get by aggregating the tokens is interesting because it&#8217;s something we&#8217;ve talked a lot about over the last year in particular. Everybody in the Valley is at this point, so where&#8217;s the value going to accrue? Will it be the frontier models that are going to make all of the money and everything at the app layer is left for dust, or will it go the other way? How have you thought about that? Because you&#8217;re in an interesting spot since you actually really do own the relationship with your end user and you&#8217;re effectively making it easy for them to pick and choose the models that they want to run with. So what are your thoughts on where this all plays out and how it will hopefully play out in a good way for OpenCode?</p><p><strong>Jay:</strong> Yeah, it feels a little bit like a marketplace where a user is able to make the choice of the model that they want to use, and different models have different attributes and different cost characteristics. Our claim here is, look, we want to showcase that diversity as well as possible. That also creates an environment where the labs are aware of each other, and that competition ends up being good for the consumer in this case. The flip side is if you&#8217;re locked into a specific vendor, then you don&#8217;t benefit from the competition that would otherwise come with it, and you are likely helping their margins in some ways.</p><p><strong>Harj:</strong> Yeah, I feel like your growth is a pretty decent proxy for the fact that choices have only improved over time, right?</p><p><strong>Jay:</strong> That&#8217;s exactly it. I think we could probably track back our bump in monthly active users to some sort of bump in the open source model market. The way we think about this is that it&#8217;s not that we&#8217;re picking a winner in terms of a model lab&#8212;we&#8217;re just betting the field. We just think the rest of the field is going to do well.</p><p><strong>Diana:</strong> The next logical conclusion of this is that models are becoming commoditized utilities.</p><p><strong>Jay:</strong> Yeah. I think what&#8217;s interesting here is that the market is so large that it&#8217;s hard to imagine people or model labs not picking off niches and chunks of their own, specializing for specific areas or specific characteristics. In our case, when you see some of the data that we were looking at before, you&#8217;ve got the Frontier Labs&#8212;we know them well&#8212;but when we see the success of DeepSeek, it&#8217;s very clear that they have picked the cost part of the quality-cost-performance axes. They&#8217;re basically saying, look, we want to be very, very good at that part. It&#8217;s hard not to imagine that happening across the board, across all these characteristics.</p><p><strong>Harj:</strong> Yeah, it feels like if you truly believe the sales pitch that the labs themselves make&#8212;which is this is just an unprecedented market, the market for intelligence has not existed before&#8212;everybody should be thinking a positive-sum, grow-the-pie mentality. In which case, they should really want you to grow and succeed because you&#8217;ll end up just being a huge customer for all of them.</p><p><strong>Jay:</strong> Yeah. I think that&#8217;s sort of true now with these open source models. We&#8217;re the largest customer for most of them.</p><p><strong>Jared:</strong> You&#8217;re the largest customer for most of the open source models.</p><p><strong>Jay:</strong> Yeah. Just in terms of the token volume, the amount that we do.</p><p><strong>Diana:</strong> Wow.</p><p><strong>Jared:</strong> So that means that you guys must be locked in this symbiotic relationship now where you both need each other for this machine to work.</p><p><strong>Jay:</strong> Yeah. The other half for us is the ecosystem. We look at the open source ecosystem as a whole and we&#8217;re going, look, we need to make this entire thing work. And again, with all the talk of open source models lately, that&#8217;s essentially our pitch.</p><p><strong>Jared:</strong> Since you&#8217;re the largest driver of most of the open source model companies, you must have this incredible insight into the GPU market and where all this inference is actually happening because you&#8217;re the ones driving the inference. Are you seeing anything interesting in the GPU market? Where&#8217;s all this OpenCode-powered compute happening?</p><p><strong>Jay:</strong> Yeah, so we rent GPUs. We work with providers that provide just straight inference. We work with the model labs themselves. One thing that we started to notice that was interesting at some point because of our global usage was that the peaks and troughs over the course of a day for GPU utilization weren&#8217;t that far off for us. Given the fact that when the East is working, maybe the West is asleep, but when the West is working, the East is&#8230; And so as a result, we have a reasonably stable 24-hour GPU cycle, which allows for pretty good utilization and it helps the unit economics for us in running these models a little bit more efficiently. That ends up being a competitive advantage when we think about some of our counterparts that maybe just serve one part of the world.</p><p><strong>Harj:</strong> Do you think it&#8217;s just such an interesting spot to be in? Do you think there are specific product choices or design decisions you made when you were first launching the product that have led to the fact that you&#8217;re growing and winning so much?</p><p><strong>Jay:</strong> Yeah, it&#8217;s funny, the name OpenCode literally comes from that. I think we&#8217;d done a similar project in the past. It was called OpenNext. The idea was when you&#8217;ve got a dominant, or in this case, two dominant players in the market, the rest of the market coalesces around an open alternative, and picking that position ends up being really valuable because if you pick it, it&#8217;s very hard for somebody else to displace you. And if it&#8217;s open, you should try and become the default as quickly as possible. So when we launched, the name was a very deliberate choice. The fact that we wanted to support, even at the time of launch, we claimed that we supported 70-plus models and providers. Just to make that happen, we had to create a separate open source project called models.dev that built up this entire database that didn&#8217;t exist at the time.</p><p>Now you can contribute to it, and this is probably the best dataset of all the models and providers out there in the world. But again, just to make that happen and to occupy that position, it was a very deliberate choice.</p><p><strong>Harj:</strong> Can you think of any other examples of intentional product or design choices you made that you think have really helped you hit this inflection point?</p><p><strong>Jay:</strong> Yeah, I think the name came a little bit later. It happened within a span of a few weeks. But the first thing that happened was with our past product, we had just hit breakeven. Around that time, this was February or so of 2025, Claude Code comes out. We look at Claude Code and we go, &#8220;This is fundamentally different from using autocomplete, using AI in that form. And this is something that actually does make sense for us as a core developer.&#8221; We were NeoVIM/VIM users at the time, and so Cursor didn&#8217;t necessarily resonate with us as much. But watching Claude Code and using it a little bit, we realized that it wasn&#8217;t up to the standard of some of these other terminal UIs like the NeoVIMs of the world. We wanted to build something like that.</p><p>That was really the first bit: when you open this up, it should feel like a modern terminal experience to a lot of the core developer audience, the ones that we were going after very early on. It should feel like one of these other things.</p><p><strong>Harj:</strong> That&#8217;s so interesting because I feel like a lot of Claude Code users, because you have such low expectations for what you can get out of a terminal UI, people are like, wow, this is going crazy how much you can get done. And these graphics and effects are really cute and cool and that&#8217;s awesome. But the fact that you were terminal UI connoisseurs, it sounds like you were looking at it thinking, oh, you could actually do so much more in the terminal.</p><p><strong>Jay:</strong> Yeah. And we had built a couple of terminal UIs in the past, one as a part of our core product with SST. The other, Dax had built on the side with a couple of his friends. You could buy coffee online. It&#8217;s a terminal UI, it&#8217;s a complete storefront. It was a fun pet project, but it showed off what we could actually do.</p><p><strong>Jared:</strong> Wait, wait. So your product was a terminal UI for buying coffee in case you wanted to buy coffee without leaving your terminal.</p><p><strong>Jay:</strong> Because the idea was you&#8217;re a hardcore developer, you&#8217;re in the terminal all day, you can&#8217;t be bothered to open a browser. So you go to <a href="http://sshterminal.shop/">sshterminal.shop</a> and you order coffee over SSH.</p><p><strong>Harj:</strong> It makes sense. Way easier.</p><p><strong>Jared:</strong> I feel like this is an example of that PG essay where if you&#8217;re a developer and you build things that you want yourself, even if it seems really goofy to other people and a VC would turn up their nose as like, &#8220;Well, that&#8217;s a horrible business. What are you talking about, a terminal UI for buying coffee?&#8221; It has this tendency to pull you in an interesting direction.</p><p><strong>Harj:</strong> Yeah. It generalizes&#8212;just having eccentric tastes does not always lead to something interesting, but it&#8217;s kind of how you get to these outlying things.</p><p><strong>Jay:</strong> Because we were so embedded in the open source community and because most of the people that surrounded that community were people that looked up to products that were really good in the terminal, we knew that if we built one of those, it would resonate with them almost right away.</p><p><strong>Harj:</strong> It&#8217;s pretty easy to think that you guys have just come out of nowhere and have exploded and are this overnight success story. But the company&#8217;s actually been around for a little bit longer than six months. Tell us a bit about that backstory because you went through YC over five years ago now, but even before that.</p><p><strong>Jared:</strong> Yeah, but the story starts way before even that.</p><p><strong>Jay:</strong> So it was second year university and I had just done a co-op term at Waterloo and I was like, &#8220;Wow, I don&#8217;t want to do this.&#8221; So I came back being very naive, found the smartest people I could get ahold of. It ends up being Frank, my college roommate, and the two of us were like, &#8220;Yeah, let&#8217;s just start a company.&#8221; Again, I&#8217;m a 21-year-old, and so I picked the name Anomaly as the name of the company because I thought, &#8220;Ah, I&#8217;m special.&#8221; But yeah, it literally starts off with me reading. How</p><p><strong>Harj:</strong> How long ago was this?</p><p><strong>Jay:</strong> 2006, 2007. I think I started reading PG&#8217;s essays because I wanted to start a company. I wanted to do a startup. This was effectively the best thing you could find on it at the time. I think the first YC application was probably around that time, but the first interview and what brought me the first time to Silicon Valley&#8212;I forget the exact batch, but I think this was the Mixpanel batch because I know the day that I interviewed was literally after Mixpanel. So I was interviewed right after them. I think this was right after Airbnb&#8217;s batch, and one of the Airbnb founders was hanging out in the room when we were waiting to be interviewed. This was obviously with PG back then. That was one, and then I think the first of many interviews and applications to YC. It took more than a decade to get in, let&#8217;s just put it that way.</p><p><strong>Jared:</strong> And you told me something crazy, which is I assumed that during those years, these were different startups, but apparently it was literally the same startup, the same legal entity all of those years.</p><p><strong>Jay:</strong> Yeah. Yeah.</p><p><strong>Diana:</strong> So this legal entity is almost 20?</p><p><strong>Jay:</strong> Yeah. We incorporated in 2010.</p><p><strong>Jared:</strong> So it&#8217;s a 16-year-old legal entity. And you&#8217;ve had successful products before, but OpenCode is the most successful. So it took 16 years since you started the company to have a truly runaway success product.</p><p><strong>Jay:</strong> Yeah. I think part of it is also to do with your level of maturity, your level of ability. We were maybe too young for some of these other waves&#8212;the mobile wave, the cloud wave, whatever. We did build things that did reasonably well. But this time around, it feels a little bit different in that it is a sum total of our experiences. That&#8217;s maybe made things a lot easier to navigate, especially given the chaos in the space that we operate and the competition.</p><p><strong>Diana:</strong> I looked at the numbers and you did nine applications from 2016 up to 2021 when you got accepted.</p><p><strong>Jay:</strong> Wow.</p><p><strong>Diana:</strong> And you did four interviews. For all of these, they were all different ideas, but it was the same legal entity.</p><p><strong>Jay:</strong> It was the same legal entity. Yeah.</p><p><strong>Jared:</strong> And the same founders.</p><p><strong>Diana:</strong> And the same co-founders.</p><p><strong>Jay:</strong> It&#8217;s just been&#8212;</p><p><strong>Jared:</strong> You and Frank doing it the whole time.</p><p><strong>Jay:</strong> Yeah, doing it the whole time. Then after we did YC, our third co-founder joined, and it&#8217;s been the three of us for the last four years. Another chunk of time flew by.</p><p><strong>Harj:</strong> What was the idea you applied to YC with for 2021 that you got in with?</p><p><strong>Jay:</strong> Yeah, so we were building a serverless platform at the time. It was like Heroku, but for AWS and serverless. We wanted to do a better job in that space and maybe grow the market. We built basically a serverless framework. That was our first big open source project and our first move into building open source products, building in public, doing the whole thing. Eventually, all of that ties into OpenCode.</p><p><strong>Harj:</strong> Where does the building in public come from? Because you&#8217;re clearly doing it now. You&#8217;re quite transparent with your metrics and growth, which is awesome. But where does that come from?</p><p><strong>Jay:</strong> A little bit of that came from Dalton. I think he was pushing us, just in general, saying, &#8220;You should probably do things in public because you&#8217;re an open source company.&#8221; Then it sort of dawned on us&#8212;this was maybe 2022 or so&#8212;where it was actually Dax, one of our other co-founders, who said, &#8220;Look, all your code is public. You work basically in public. If you don&#8217;t talk about it publicly, you&#8217;re probably doing yourself and your product a disservice.&#8221; Ever since then, it&#8217;s just been a part of our identity. For a lot of our community, the people that follow us on Twitter, it sounds funny, but to them, it feels like watching a reality TV show of this team and the company and the journey that they&#8217;re on.</p><p><strong>Diana:</strong> I think the thing that&#8217;s fascinating is that even though it sounds like such a winding road, and if people just heard the beginning of the podcast, the company sounds like a lightning-in-the-bottle moment, like you just got caught by Stripe and got so lucky. But the reality is you&#8217;ve been grinding for a good 10 years and never gave up, which is so impressive. All those winding destinations that ended up being dead ends actually taught you different things because you also ran a consumer company. So you got really good at consumer acquisition, tracking numbers, and all of that, which really plays right now and the level of detail you have for OpenCode and, of course, open source. All of these were not wasted, quote unquote. It was really more a journey that took 10 years to get to zero to 30 million in eight months.</p><p><strong>Jay:</strong> Yeah, it&#8217;s crazy when you put it that way. I think what&#8217;s fun is that with OpenCode, we feel like we can go out and address the entire market. That includes consumers, individual users, small teams, obviously the open source community, mid-market, all the way to the enterprise. In our past iterations of all the different products we&#8217;ve worked on, we&#8217;ve probably done one of each. Now it feels like we get to do all of them together in one part. It&#8217;s a lot more fun because you can think about the customer journey through that entire thing, and it makes a lot more sense because you&#8217;re not hyper-optimizing for specific parts. I&#8217;m not trying to build a very specific enterprise company or a very specific consumer company. We&#8217;re trying to just do the whole thing.</p><p><strong>Diana:</strong> What got you to not give up?</p><p><strong>Jay:</strong> Partly maybe being a little stubborn. It&#8217;s funny, maybe we should put a little warning that says, don&#8217;t try this at home. Because I think the prudent thing would&#8217;ve been to shut down your company, go join a high-growth startup, learn a bunch of things. But there was something in the back of my mind&#8212;Frank is probably the same&#8212;in that we felt like we were learning these different things along the way. In our heads, we were figuring out, okay, here&#8217;s what it takes to build not just a product, but also do the marketing, understand the positioning, do the whole thing. That journey felt like progress and positive progress. Part of that was obviously us being fortunate to have the ability to do that, which was basically just living with parents for a while, and we ran out of money. But maybe it was a combination of being a little thick-headed and seeing positive progress.</p><p><strong>Jared:</strong> I&#8217;ll admit I&#8217;ve been kind of a diehard Claude Code user since Garry Tan became addicted to Claude Code. But recently I&#8217;ve been using OpenCode. Very cool. I dropped a bunch of PRs from OpenCode this week and I&#8217;ve been really impressed with how well it works with open source models on our existing codebase, which is a very large, very complex codebase. For folks who are watching who have possibly only used Claude Code or Codex or Cursor, how should they think about trying OpenCode and potentially switching to it?</p><p><strong>Jay:</strong> When you hear about a new model that comes out, especially an open source one, and you want to try it out, you could hack your way into using it with Claude Code or one of the closed source alternatives. But OpenCode is really good for this. You just go in, look at the model picker, and GLM 5.2 or whatever the new open source model is will probably be up there. You can pick it and start using it right away.</p><p><strong>Harj:</strong> Okay. Well, Jay, I think that&#8217;s all we have time for today. I actually learned lots of really interesting stuff about your backstory that I didn&#8217;t know in this episode. I think it&#8217;s a pretty inspiring story, honestly, for anyone that wants to start a company.</p><p><strong>Jay:</strong> Or at least entertaining.</p><p><strong>Harj:</strong> Entertaining and inspiring. It can be. It just hits on so many of the classic startup wisdoms. You should pursue your interests, have eccentric taste, be living the future a little bit. I would argue that trying to order coffee through the terminal is&#8212;I&#8217;m not sure if that&#8217;s living in the future or the past, but it&#8217;s not living in the current time. Maybe there&#8217;s something in there. Also, just the fact that you kept building your taste and kept building things over a long period of time. Then when lightning strikes, you&#8217;re actually in a position to capture it. I think that&#8217;s the thing that doesn&#8217;t get mentioned: to catch lightning in a bottle, you actually have to position the bottle correctly, be ready for it, and know what to do with it. You guys were well positioned to do that. So congrats on all your success. I know it&#8217;s only going to get more explosive from here.</p><p><strong>Jay:</strong> Thank you. Thank you for having me.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.ycrootaccess.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[World Models: An intuitive introduction]]></title><description><![CDATA[Sample efficiency is one of the biggest open problems in AI. World models, which build internal representations of the physical world, may be the way to solve it.]]></description><link>https://www.ycrootaccess.com/p/world-models-an-intuitive-introduction</link><guid isPermaLink="false">https://www.ycrootaccess.com/p/world-models-an-intuitive-introduction</guid><pubDate>Fri, 17 Jul 2026 14:05:04 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/6fa19c64-8be0-40ee-8d42-d4b479ae343d_1920x1080.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div id="youtube2-qz4GQ0zUFRw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;qz4GQ0zUFRw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/qz4GQ0zUFRw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>Why do even our best AI models need tens of thousands of examples to learn skills that a human picks up in a handful of tries? Solving this problem is one of the great open challenges in modern AI. World models, which give AI an internal simulation of its environment, are one of the most promising paths forward. </p><p>In this episode of Decoded, YC General Partner <strong><a href="https://x.com/agupta">Ankit Gupta</a></strong> and Visiting Partner <strong><a href="https://x.com/FrancoisChauba1">Francois Chaubard</a></strong> discuss the intuition and math behind world models, new research, and current applications in self-driving, robotics, and more. </p><p><a href="https://youtu.be/qz4GQ0zUFRw">Watch on Youtube</a></p><h3>Timestamps: </h3><p>00:00 &#8212; Intro</p><p>01:45 &#8212; What would perfect efficiency look like?</p><p>05:10 &#8212; World models in the human brain</p><p>09:20 &#8212; Control theory &amp; the drone example</p><p>14:30 &#8212; When physics breaks down</p><p>17:45 &#8212; Chess, Go &amp; the action space problem</p><p>24:10 &#8212; Why AlphaGo can&#8217;t scale</p><p>28:00 &#8212; Monte Carlo tree search explained</p><p>34:00 &#8212; Self-Driving: state space is infinite</p><p>40:30 &#8212; Model-Free vs. Model-Based RL</p><p>44:00 &#8212; Why robotics is the hardest case</p><p>48:20 &#8212; World models that actually work</p><p>54:10 &#8212; JEPA &amp; latent space tricks</p><p>59:00 &#8212; Open problems remaining</p><p>01:04:30 &#8212; Does this pass the squint test?</p><p>01:08:00 &#8212; Outro</p><h3>Transcript</h3><h4>00:00 &#8212; Intro</h4><p><strong>Ankit Gupta</strong></p><p>One of the biggest open problems in AI right now is how to solve sample efficiency. That is, how do you get models to quickly learn new tasks or skills from relatively small amounts of training data?</p><p><strong>Francois Chaubard</strong></p><p>Humans do this incredibly well. We can learn new games, concepts, and skills, often after just a handful of tries. Our best models, on the other hand, often need tens of thousands of data points just to learn.</p><p><strong>Ankit Gupta</strong></p><p>So today we&#8217;re going to discuss what many top researchers believe is the most promising path to closing that gap, world models.</p><p><strong>Francois Chaubard</strong></p><p>We&#8217;re going to discuss the motivation and math behind world models, current applications, and why this approach might be the key to unlocking AGI.</p><p><strong>Ankit Gupta</strong></p><p>You and I have talked a lot about the various ways people are creating models and the sample efficiency of them. Why don&#8217;t we start by just defining sample efficiency and how we intuitively think about it as humans?</p><p><strong>Francois Chaubard</strong></p><p>Yeah. So I think from my perspective, the two major problems that we haven&#8217;t left to solve is intelligence per watt and intelligence per sample. Intelligence per watt is how many valve perplexity points we get per watt of spend. And then intelligence per sample is basically if I have one additional sample in my dataset, how much more intelligent am I getting? And so if I imagine I have a new tasks like ARC-AGI, For example, I think really Fran&#231;ois Chollet has been on the forefront of this thinking and talking about intelligence as a rate of skill acquisition versus skill acquisition and that&#8217;s very different. And so how fast do we get smarter with more and more samples? And these things are incredibly poor at getting smarter with fewer and fewer samples.</p><p><strong>Ankit Gupta</strong></p><p>And for context, the ARC-AGI test sets are a really good example of cases where humans are intuitively very good at them. Most humans can intuitively solve those puzzles with some amount of thinking and effort. But our current state-of-the-art AI systems, what people consider frontier intelligence, basically can&#8217;t do them.</p><h4>01:45 &#8212; What would perfect efficiency look like?</h4><p><strong>Francois Chaubard</strong></p><p>Right. I mean, we come into new problems with such inductive bias from K through 12, all this math in school that we&#8217;ve had, that these models are kind of getting from compressing the entire internet. And so when we come in, we&#8217;re not coming in tabula rasa or bare bones, but even so that they have... I don&#8217;t know what percent of the internet you&#8217;ve read. I&#8217;ve read very little percent of the internet, but despite that and having read the entire internet, it still can&#8217;t really do well and generalizing to these new tasks.</p><p><strong>Ankit Gupta</strong></p><p>So now let&#8217;s think about this in the extreme cases. In the extreme case where let&#8217;s say we were perfectly sample efficient, we were as sample efficient as possible, what would that mean in terms of a model that is taking a set of actions in the world?</p><p><strong>Francois Chaubard</strong></p><p>Well, I guess the perfect sample efficiency would be zero samples. And there are examples of this, and that sounds absurd to say. And the example, the hypothetical I&#8217;ll give on this is imagine I had a perfect world model, then I should never go to the environment to go and collect samples to train on. And well, that can&#8217;t possibly happen for us. No, it actually can happen. We do it all the time. It&#8217;s called Newton&#8217;s Second Law of Motion. In Newtonian mechanics, we basically know how to get an object from point A to point B with a rocket quite easily just by following Newton&#8217;s laws of motion.</p><p><strong>Ankit Gupta</strong></p><p>Yeah. When NASA plans to intercept an asteroid and is planning it years in advance and can set it off in a trajectory where it just glides to the right thing and intersects to the right point, that is an example of a perfect world model we&#8217;ve built where we&#8217;re then just letting that world model act. And that system does not need to intelligently collect new samples from the environment to decide which direction to go next. It&#8217;s already been pre-programmed and it can perfectly do it.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. Can you imagine if we needed to collect a one million training examples of us shooting spaceships to the moon to know how to do it? We definitely wouldn&#8217;t have the Apollo missions, right? But we do have that ability because the real world is differentiable and we can do something called model predictive control that we&#8217;re going to talk about in a little bit. But even in our own brain, I was just thinking about this on the drive up, but there&#8217;s so many ways that I can basically think about the things that you are going to say or what a VC is going to say when I was pitching them or what-</p><p><strong>Ankit Gupta</strong></p><p>A customer might say.</p><p><strong>Francois Chaubard</strong></p><p>... a customer might say. And even product, having taste, what is taste? It&#8217;s predicting that other people are going to like this thing. And so we&#8217;ve built this world model over years of entrepreneurship, 10 years of getting it wrong, that maybe Bill Gates, Steve Jobs and Jensen have 50 years of world modeling experience to know what people want. And basically this is actually proven in the 1967 Cog-Sci study by Richardson that basically showed that if you take a cohort of three groups of people and you have one go practice layups in basketball and they go and they shoot, for one hour they improve by I think it was like 24% or something like that. And then if you take the other one and they just blindfold them and they imagine laying up a basketball, they improve at 23%-</p><p><strong>Ankit Gupta</strong></p><p>Interesting.</p><h4>05:10 &#8212; World models in the human brain</h4><p><strong>Francois Chaubard</strong></p><p>... against the control. I mean, that&#8217;s insane. It means that we have this crazy good world model and there&#8217;s this neuroscientist at Stanford named Shaul Druckmann who basically is of the view that the entire point of the growing neocortex during the great critical expansion 10 million years ago was to get better and better and better and better at world modeling. And having just like my little VLA, which we&#8217;ll define of predicting the next action is not as good as having a world model to lean on either for training purposes or for test time adaptation.</p><p><strong>Ankit Gupta</strong></p><p>Yeah. What it fundamentally comes down to is we as humans, we think about our intuitive ability to think as coming from some implicit world model we have in our heads, encoded by genetics and our ability to learn and whatever else. It seems like models can do surprisingly intelligent things despite not having an explicit world model when it comes to natural language. When they&#8217;re just talking, it seems like maybe under the hood, deep inside the weight somewhere, there&#8217;s some kind of implicit understanding of the world, but there isn&#8217;t an explicit representation of that. But it seems like in certain domains, especially in robotics and self-driving as we&#8217;ll talk about, that sort of breaks down. And maybe it would be helpful now to just think a little bit about and just sort of define some of the pieces of what makes it challenging in these different domains, and then we can use that to build up to why it&#8217;s particularly hard in things like self-driving and robotics to get these types of predictive models to work.</p><p><strong>Francois Chaubard</strong></p><p>Yeah, let&#8217;s do it. So let&#8217;s actually take a step back and just talk about control, reinforcement learning and define some common terms. So typically we teach a course called Decision Making Under Uncertainty, which is the main reinforcement learning course at Stanford. I like to show a specific example of let&#8217;s say I have some drone and this is my poor little drone here and it has some mass.</p><p><strong>Ankit Gupta</strong></p><p>Helicopter.</p><p><strong>Francois Chaubard</strong></p><p>And we know that gravity G is pulling down on it and it&#8217;s currently at position T with velocity T, which we will collectively call the state. And to be really clear, this is going to be PX, PY, PZ, and VX, VY, VZ.</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s like the six-dimensional state vector.</p><p><strong>Francois Chaubard</strong></p><p>Yep. And we have some thrust vector U that we control and we&#8217;re trying to get to some point P-star and V-star, which is V-star is typically zero. And so you have some platform that I want this drone to land on. So this is this control problem, right? And so let&#8217;s say this is like, and we&#8217;ll go through optimal.</p><p><strong>Ankit Gupta</strong></p><p>Optimal, yeah.</p><p><strong>Francois Chaubard</strong></p><p>Optimal control. So how would I actually solve this? So the first thing I need to know is my transition function. And so this is my state transition function, which is ST+1.</p><p><strong>Ankit Gupta</strong></p><p>Given the previous.</p><p><strong>Francois Chaubard</strong></p><p>Given ST and my action, which I control is UT. And so this is my state transition or dynamics function or a world model. This is a world model.</p><p><strong>Ankit Gupta</strong></p><p>This is a very fundamental for context, this equivalent to a transition function you would think about in RL in general.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. And then what I&#8217;m trying to learn is something called a policy, which is like what UT should I emit given some ST. And so this is the ultimate question, what should I do? What action should I take given some state ST? And so the way that we&#8217;ll solve this and luckily we have a world model that is perfect and it&#8217;s called-</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s Newtonian physics.</p><p><strong>Francois Chaubard</strong></p><p>Newtonian physics. This is new and second law of motion, which is F equals MA. And so we know that the position PT plus one is going to equal PT plus delta TVT plus one half delta T squared. So everyone&#8217;s taking high school physics. And the same thing for the velocity, blah, blah, blah.</p><p><strong>Ankit Gupta</strong></p><p>Delta TA.</p><h4>09:20 &#8212; Control theory &amp; the drone example</h4><p><strong>Francois Chaubard</strong></p><p>And then my acceleration is sum of the forces, which is going to be my UT. I think I divide by the mass and G. And so that&#8217;s it. And now I have my transition function. Now, how do I get to a policy? And I&#8217;m going to apply something called model predictive control or real-time model predictive control, which is the way that SpaceX lands the rocket on some platform in the ocean. And what you&#8217;re going to do is you&#8217;re going to set up your loss function. You&#8217;re going to minimize sum over all T. You have UT to infinity, and I&#8217;m going to minimize my P-star minus PT plus V-star minus VT. And usually you add this little lambda UT, which is how much energy you&#8217;re exerting. And you can&#8217;t have infinite thrust. So you typically will have to say UTU max thrust that can be achieved. And so this is easily solvable with comics optimization. And so this is convex, this is convex, this is convex. The sum of convex functions is convex. This is a convex constraint. And so DCP, discipline convex programming means that I can put this into a CVX pie and it will just give me out my policy, which will be the solution will be the optimal UT plus one all the way to infinity.</p><p><strong>Ankit Gupta</strong></p><p>So we can solve this in closed form basically.</p><p><strong>Francois Chaubard</strong></p><p>Exactly.</p><p><strong>Ankit Gupta</strong></p><p>Because we have this world model of Newtonian physics, we can say at every step exactly how this drone should fly so that it lands on the appropriate thing under a set of constraints like max thrust available.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. You&#8217;ll run your log barrier or enter a point, whatever to some solver on this and it will give me my optimal... Then this will be literally the optimal path that this thing can take to get to this state. And that will minimize and I can increase this if I want it to do the least energy path and I make that zero if I want it to be the fastest. And so that&#8217;s typically the way that you would do what I&#8217;ll call deterministic differentiable control. And why differentiable? Because I can form the Lagrangian by taking this minus this constraint and take the grading of it. And I can do mono-robins.</p><p><strong>Ankit Gupta</strong></p><p>You use the fact that it&#8217;s differentiable to do the optimization.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. If this is non-differentiable, you cannot do convex optimization and you cannot do SGD. Even if it&#8217;s non-convex, you could still solve and get a pretty good solution as we do in deep learning. But if it&#8217;s non-differentiable, you can and can&#8217;t. There&#8217;s nothing you can do.</p><p><strong>Ankit Gupta</strong></p><p>So yeah, let&#8217;s have an example then of how you could make this non-differentiable. What&#8217;s a scenario I guess even in this drone scenario where it now becomes non-differentiable?</p><p><strong>Francois Chaubard</strong></p><p>Yeah. So I&#8217;ll put this adversary named Ankit. And your job is to, you have another drone, let&#8217;s say. Ankit&#8217;s drone is to try to hit me and stop me from getting there.</p><p><strong>Ankit Gupta</strong></p><p>Now from the position of your drone, you don&#8217;t know what actions I&#8217;m going to take.</p><p><strong>Francois Chaubard</strong></p><p>Right. And so now let&#8217;s just call this the, this would be now we&#8217;re definitely not deterministic, we&#8217;re stochastic. And stochastic and non-differentiable. And in this case, my state transition, what is ST plus one? It&#8217;s going to be say I&#8217;m in now, my thrust and what Ankit&#8217;s going to do.</p><p><strong>Ankit Gupta</strong></p><p>Right. And it was all differentiable until this new variable.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. And I can&#8217;t backprop through your brain to say what you&#8217;re going to do with your little drone controller, right? It&#8217;s completely non-differentiable now. And I&#8217;m resorting and I have to resort to this awful area called reinforcement learning, which is just super brutal and it&#8217;s sprawling and there&#8217;s so many different things. And you&#8217;ll hear things like when you study initial reinforcement learning called value iteration or policy iteration. And there&#8217;s DQN or deep Q learning or just Q learning. There&#8217;s actor critic. There&#8217;s all this mega stuff.</p><p><strong>Ankit Gupta</strong></p><p>And all of this stuff ultimately comes down to ways to estimate, to model this non-differentiable stochastic process.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. Yeah. And so that&#8217;s basically the main thing is you&#8217;re going to start talking about this as a model where I&#8217;m going to introduce this SI to say that this is going to be some model that&#8217;s going to take in these things and then output this and that we&#8217;re going to train it over many, many instantiations of this. And that&#8217;s so it&#8217;d get a better and better world model. And then I need to train some policy, AT/ST. And then typically you also need a value function. And that is the value of some state. And to discern between the value of different states. And in this case, I don&#8217;t know what a valid state is, but let&#8217;s just say I was doing the SpaceX with launching rockets.</p><h4>14:30 &#8212; When physics breaks down</h4><p><strong>Ankit Gupta</strong></p><p>Landing the rocket.</p><p><strong>Francois Chaubard</strong></p><p>And landing rockets in Florida. Let&#8217;s just say that there&#8217;s different... If I have my launchpad here and I have a whole bunch of houses here, let&#8217;s just say, the path going from here to here, I may think that doing this and then coming across here and burning all these houses alive may be not highly valued. So I might say as an example, they typically call this some kind of a cone here, and I might say it&#8217;s low value to be here and it&#8217;s very high value to be in this cone or something as an example.</p><p><strong>Ankit Gupta</strong></p><p>Yes, right. But in a sense, a value gives you some expectation of future rewards, like the sum of future rewards you&#8217;re getting. And so if you&#8217;re in a bad space, you would set the value to zero or negative infinity or something like that.</p><p><strong>Francois Chaubard</strong></p><p>Yeah, so we should introduce RT as well. And so typically if you&#8217;re playing go or chess, winning the game, you can say winning the game is plus one minus one for losing, draw zero, that&#8217;s what&#8217;s done in AlphaGo. In chess, we have these heuristics, like a pawn is worth one point, a rook is worth five, et cetera, et cetera. Et cetera. So you can already have reward is the difference in board state. And then this, yes, will be the sum of my discount. Should just do T of RT given. And it&#8217;s important also to use this nomenclature, V pi. And the reason why that&#8217;s important is because what&#8217;s actually happening here is this is the discounted reward following policy pi.</p><p><strong>Ankit Gupta</strong></p><p>Correct.</p><p><strong>Francois Chaubard</strong></p><p>And that means that when I&#8217;m in this state, I will take this action and then I&#8217;ll end up in this to SC plus one and then I&#8217;ll take this action and taking it greedy. And so that&#8217;s the value with respect to pi.</p><p><strong>Ankit Gupta</strong></p><p>And so ultimately what it comes down to is we are trying to still find a new policy pi. And along the way we will use machine learning models in various capacities, this is standard RL, to estimate the value function given the rewards we&#8217;re receiving. And then where world models come in is a way of incorporating all of those into some sort of joint modeling of the state and action distribution so that we can make more intelligent policies off of it.</p><p><strong>Francois Chaubard</strong></p><p>Right. And so your standard setup for this is what I&#8217;m always trying to get to at the end of the day is some joint distribution, which would be ST plus one given where I&#8217;m at now. And then this factorizes with chain rule simply to my pi, my policy, AT given ST, and my world model. And I&#8217;ll give this, this is usually represented with theta. And this is my world model, which would be ST plus one given ST and AT. And these are typically learned separately. And you can imagine, in fact, actually you can actually learn this. This a video generation model and I have the frame ST, and I predict the next frame ST plus one. And we&#8217;ll get into this.</p><p><strong>Ankit Gupta</strong></p><p>Yeah. For those of us who saw our diffusion model series, often people these days use video diffusion for exactly this.</p><h4>17:45 &#8212; Chess, Go &amp; the action space problem</h4><p><strong>Francois Chaubard</strong></p><p>Yeah. And then what you can do, and this is the in vogue thing to do since Danijar and the Dreamer paper series from V1 to V4 is do action conditioning later. Similar to Clip where we will inject this input head or input tail to come into the model to influence and enable the world model to have embodiment. What does that mean? It means that not only can I predict as a plant or tree growing on the side of the building, I can see the world passing by, but I can actually influence it and I can change the world and I can learn that with AT. And it&#8217;s far fewer samples to do this post action conditioning if I already have a really good ST to ST plus one world model.</p><p><strong>Ankit Gupta</strong></p><p>And so here you&#8217;re saying what&#8217;s also in vogue now is jointly training these versus separately training them.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. So that is called the world action model where some of the issues here is one, there&#8217;s all these training dynamics if these things are disparate, training on different sets and things like that. The other issue is plainly obvious what I have to do to actually do test time planning is I&#8217;ll have to sample with model one, invoke theta, and then pass that sampled action into here and then roll it out to ST+1. And it&#8217;s very expensive and it&#8217;s a very not real time. Two major issues and why can&#8217;t we just scale up AlphaGo to solve all the problems is because of this property. If I have one invocation to the model and it gives me both, here&#8217;s the action I should take and here&#8217;s the ST plus one, that&#8217;ll end up much, much cheaper and much, much faster.</p><p><strong>Ankit Gupta</strong></p><p>Okay. So I think that&#8217;s a really good segue. I think why don&#8217;t we now motivate everything we just described through a series of increasingly complex environments? So I&#8217;ll contend that I think the right set of environments for us to consider is chess followed by go, followed by self-driving, followed by robotics.</p><p><strong>Francois Chaubard</strong></p><p>All right, so let&#8217;s go through a couple examples of problems that we want to apply reinforcement learning to. So chess is a pretty easy one. There&#8217;s an 8x8 grid. And so typically when you approach any RL problem, you&#8217;re going to look at star. And so the size of the state, the number of states I can be in. So if I have these eight here and these eight, so this would be 8, 16, 32. So it&#8217;d be 32 to the 64.</p><p><strong>Ankit Gupta</strong></p><p>Yes, quite large.</p><p><strong>Francois Chaubard</strong></p><p>Quite large. Then my transition function is stochastic and non-differentiable because you can-</p><p><strong>Ankit Gupta</strong></p><p>You don&#8217;t know what the other player&#8217;s going to do.</p><p><strong>Francois Chaubard</strong></p><p>I don&#8217;t know what you&#8217;re going to do. So if I&#8217;m playing chess.com at my house, I move and then something happens and it comes back and then now you moved and the board has changed. So I can&#8217;t really differentiate through what the other player is doing. The car line in my action space is actually quite small. Even though there&#8217;s 32 pieces and all that stuff, there&#8217;s only eight possible moves in expectation that are legit moves. So any-</p><p><strong>Ankit Gupta</strong></p><p>In any given state, there&#8217;s only eight-ish moves you could do.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. Let&#8217;s just say in the beginning, I can move all my pawns, I can move my horses. So that&#8217;s 10. That&#8217;s not that much. So this is extremely small. And then my reward, we can use the heuristic-based approach or we can just say plus one, zero, or minus one if I lose, plus one if I win. And so this is very tractable.</p><p><strong>Ankit Gupta</strong></p><p>You say it&#8217;s tractable even though there&#8217;s a really big state space here.</p><p><strong>Francois Chaubard</strong></p><p>Yeah.</p><p><strong>Ankit Gupta</strong></p><p>But why don&#8217;t we talk about that for just a second. I think this is a really important point. I think when you say it&#8217;s tractable, you&#8217;re specifically referring to the action space being small because it affects the combinatorial expansion here. Should we talk about that for just a second?</p><p><strong>Francois Chaubard</strong></p><p>Yeah.</p><p><strong>Ankit Gupta</strong></p><p>Or maybe we can add go and then contrast the two.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. So why don&#8217;t we do that? Because I want to get to the AlphaGo, the way that they solve this. And you&#8217;re right. So if I were to do this naively and I just took my ST plus one and I want to do look aheads, what I would do is I would take all of the actions I can take. So there&#8217;s eight. So I would do action one, action two, action eight. And then each one of these, I need to expand it for all possible states. And so now I need to do carnality S, which we just said is this huge freaking number. And so I have to do that eight times and I have to do it again. I have to do it again. So just looking forward, one move is quite intractable.</p><p><strong>Ankit Gupta</strong></p><p>Although at the same time, everyone starts at the same starting position. And while it is a really large space, there isn&#8217;t an infinity number of potential... There&#8217;s actually a really small number of game boards, even four moves into the game as opposed to a game where you could start in any permutation, for example, of initial game state and what a few states down look like.</p><p><strong>Francois Chaubard</strong></p><p>So this is definitely overdone because it&#8217;s much, much less than this in practice. But just naively looking at what possible game states could be as a rough math here. But this is roughly the idea. And then each one of these leaves, I need to invoke my value function, which is the value of that state T plus one. And so I have to do that all many times. And we&#8217;ll get this with AlphaGo, but this ends up being estimating the leaf node because at the end of the day, my policy AT/ST, I want to pick, I want the ARG max of the value of the following-</p><p><strong>Ankit Gupta</strong></p><p>The ARG max action, I guess it would be an A here.</p><p><strong>Francois Chaubard</strong></p><p>Yeah, A, exactly. Yeah. The ARG max over A of the value of the N state, ST plus N, let&#8217;s say. That&#8217;s the main goal here. And so for me to do that, I need to roll all this out, estimate the value, and then pick the best one. And so this quickly grows. However, and we&#8217;ll see this with AlphaGo, which actually has an even bigger state space. So I think it&#8217;s 19 by 19. Correct me if I&#8217;m wrong.</p><p><strong>Ankit Gupta</strong></p><p>I think it&#8217;s about right now.</p><p><strong>Francois Chaubard</strong></p><p>So yeah, there&#8217;s 19 by 19 grid. In each one, it can be black, white, or nothing there. So I have three. So let&#8217;s do our star again. So the cardinality of the state I think is going to be S3 itinerary thing here.</p><h4>24:10 &#8212; Why AlphaGo can&#8217;t scale</h4><p><strong>Ankit Gupta</strong></p><p>19 squared, I guess.</p><p><strong>Francois Chaubard</strong></p><p>19 squared. I think it&#8217;s 361.</p><p><strong>Ankit Gupta</strong></p><p>381? Yeah, 361.</p><p><strong>Francois Chaubard</strong></p><ol start="361"><li><p>My transition, same issue. I don&#8217;t know. My action space is going to be 361, let&#8217;s say.</p></li></ol><p><strong>Ankit Gupta</strong></p><p>So it&#8217;s a good amount bigger than chess.</p><p><strong>Francois Chaubard</strong></p><p>Much bigger.</p><p><strong>Ankit Gupta</strong></p><p>But it&#8217;s still not enormous.</p><p><strong>Francois Chaubard</strong></p><p>Yeah.</p><p><strong>Ankit Gupta</strong></p><p>As we&#8217;ll see in a second.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. And so basically what they do, they call this Z, which is kind of annoying, but let&#8217;s call it R. And it&#8217;s the terminal when they won the game. And they basically, you have your trajectory, which is S0, A0, R0, then all the way to the end of the game, SN, AN, RN. And if you won, then all the moves that... If black won, all the moves that black did get plus, all the moves that white did were minus one. And that&#8217;s how they create their rollouts.</p><p><strong>Ankit Gupta</strong></p><p>Rollout refers to a taking end steps of play of all players one after another, of moves under a specific policy at the particular instantiation of it.</p><p><strong>Francois Chaubard</strong></p><p>Right, right. So let&#8217;s probably under this policy P, theta T. And we&#8217;re going to overload T, but this is that instantiation.</p><p><strong>Ankit Gupta</strong></p><p>At that model.</p><p><strong>Francois Chaubard</strong></p><p>We froze that model and we play I think it&#8217;s like 70 games and we treat all of those and we&#8217;re going to sub-sample a bunch of these state action results, state action results to update our policy in our world model, our transition model. And what it&#8217;s actually doing is we take in an ST, we give it to some theta, and then it wants to output the probability of ST plus one being played, which is our transition function and the value of the current state, the ST.</p><p><strong>Ankit Gupta</strong></p><p>And how do we get the value.</p><p><strong>Francois Chaubard</strong></p><p>And so the value of the current state, well, both of them are coming out of the model, but basically the loss function, L theta is going to equal, and it&#8217;s going to be eerily close to this control problem one, is we have some V theta minus this Z, which we&#8217;ll just call it R here squared. And then plus... Actually, sorry, it&#8217;s minus this pi, which I&#8217;ll explain in a second log P theta. And I think everyone includes this, but they include it in the paper, so I&#8217;ll include it there as well, which is the weight decay. So this is basically what our loss function is. Then we&#8217;ll play a bunch of these games. Let&#8217;s try to be a little bit organized here. And so this is our setup, this is architecture. And now once we train this thing, we do an insanely expensive task of test time planning. And so this trend in RL is just called test time planning.</p><p><strong>Ankit Gupta</strong></p><p>And the specific algorithm they use here for this is Monte Carlo tree search.</p><p><strong>Francois Chaubard</strong></p><p>It&#8217;s called MCTS. And so this is one of the possible things that you could do. It ends up working extremely well if you have small action spaces.</p><p><strong>Ankit Gupta</strong></p><p>Yeah. So let&#8217;s just very intuitively talk about what MCTS does.</p><p><strong>Francois Chaubard</strong></p><p>Sure.</p><p><strong>Ankit Gupta</strong></p><p>A lot of people have heard about Monte Carlo tree search because AlphaGo was such a big moment, but how exactly does that map into our star in value function and policy?</p><h4>28:00 &#8212; Monte Carlo tree search explained</h4><p><strong>Francois Chaubard</strong></p><p>Yep. So I&#8217;ll take this ST. This will give me 361 numbers that sum to one. And so I&#8217;ll have some probability of where these things are going to go of where my opponent will play here.</p><p><strong>Ankit Gupta</strong></p><p>So these are the sets of actions.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. So I&#8217;m here so that I have all my ST plus ones. I&#8217;ll have 361 of these things. And then-</p><p><strong>Ankit Gupta</strong></p><p>And to be clear, this is action one, action two all the way to action 361.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. Exactly. Yeah. And we have to estimate the value of each one of these. And so then we have to invoke the model all 361 times to give me values for each one of these things. And then I&#8217;ll select it based on the UCB, the upper confidence bound, which is this equation that is roughly something like balancing my value function of ST plus one, which in the literature it&#8217;d be called a Q value because it&#8217;s actually the difference between a value function and a Q value is just that I have the action as well. So it&#8217;d be ST, then AT. So we&#8217;ll just call that Q value, which is my exploitation term. And then my exploration term will be something like it&#8217;s this funky square root of N. So it&#8217;s the ARG max of A of my Q. And then I have this, which is the probability of this move being played, which we have from here of S, let&#8217;s just call it ST plus one. And then I have this term, which is this sum over NSB divided by NSA.</p><p><strong>Ankit Gupta</strong></p><p>What the [inaudible 00:29:59] on this term.</p><p><strong>Francois Chaubard</strong></p><p>So these ends is the visit count during my MCTS process. So this whole tree I&#8217;m going to...</p><p><strong>Ankit Gupta</strong></p><p>So this tree could get really big. It&#8217;s 361 per thing.</p><p><strong>Francois Chaubard</strong></p><p>And it&#8217;s depth of 30.</p><p><strong>Ankit Gupta</strong></p><p>So you can&#8217;t visit every single leaf node.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. And so you want to keep track of which state did you end up in and what action did you take when you were in that state? And you want to make sure that you have good exploration, right? And so the way you ensure that you have good exploration is you want to not just be greedy and always pick the highest value one because that could be very myopic. And so what you&#8217;ll do is during this MCTS process, you&#8217;ll start this dictionary, which will be all zeros of the visit count of being in this state and taking this action. And then once you go through your first rollout, you&#8217;ll go here, all these things will be added to zero, you&#8217;ll have some probability. We&#8217;re going to bias it towards the higher probability of places to go and then we&#8217;ll expand those trees and then we will update the counts that we visited this and that will basically reduce the amount of probability that we&#8217;re going to select it again because this will reduce my exploration term. And if it&#8217;s highly valued, then we&#8217;re going to increase the Q on this because this is the expected value of going down this path.</p><p><strong>Ankit Gupta</strong></p><p>So the gist of it is fundamentally you want to take the optimal-ish path, but have enough exploration in this really expensive step you&#8217;re doing here so that you are making sure you&#8217;re getting a decent chunk of the other potential leaf nodes you could traverse to in these 30-step rollouts.</p><p><strong>Francois Chaubard</strong></p><p>And so I&#8217;m going to do this MCTS simulation 800 times here. And then for all 800, I have to go through this whole process and I have to invoke the model at least 30 times to get through all here. And so that&#8217;s 27,000-</p><p><strong>Ankit Gupta</strong></p><p>800 times 30 invocations.</p><p><strong>Francois Chaubard</strong></p><p>So 24,000 invocations of the model to develop this tree. And then once I have it-</p><p><strong>Ankit Gupta</strong></p><p>That&#8217;s per step.</p><p><strong>Francois Chaubard</strong></p><p>Per step. Just do one action into the game. A lot of people don&#8217;t understand that this is like you don&#8217;t store this MCTS tree, you throw it away after you make the move. But it&#8217;s very expensive to develop this MCTS tree. And once you have it, the probabilities of traversal are actually extremely useful for training. And then you end up biasing it and you train it with the MCTS tree, which is a little bit seems like circular motion or something like that, but you end up treating that as the pie that you&#8217;ll train in your loss function. So we have the R of did we win or lose? We have the pi of what was the end result of this whole expensive process. And then at test time we are going to do these 24,000 steps every single move to pick the ARG max that satisfies both exploration and exploitation.</p><p><strong>Ankit Gupta</strong></p><p>In this case, this still feels somewhat tractable though because the action space is small enough where this kind of works.</p><p><strong>Francois Chaubard</strong></p><p>Exactly.</p><p><strong>Ankit Gupta</strong></p><p>But now let&#8217;s say hypothetically, maybe we can draw an imaginary game of go where it&#8217;s like... Let&#8217;s say this game of go was like a thousand by a thousand. And so now you have A equals more or less a million. And now this tree we&#8217;re drawing here that has to take here, this has cardinal or width, I guess, one million and there&#8217;s S0 through S one million. And the number of steps you would have to take here presumably would have to be way more than 800 in order to get any reasonable kind of sampling of this. And so you&#8217;re probably multiplying the test time cost of doing a rollout or of doing a next step prediction astronomically if the game was even let&#8217;s say this is only a 100X bigger than the current game or not even 50X bigger than the current game.</p><h4>34:00 &#8212; Self-Driving: state space is infinite</h4><p><strong>Francois Chaubard</strong></p><p>Everyone was very excited about AlphaGo and at the time, and what was this, 2017, 2016, everyone&#8217;s very excited about this. And the important thing to pick up is that we did 800 MCTS simulations to cover 361 possible actions on average. So that gives us about two samples roughly on an expectation for every single action.</p><p><strong>Ankit Gupta</strong></p><p>So here you need two million of them for a similar depth.</p><p><strong>Francois Chaubard</strong></p><p>Two million for a similar depth. And then that&#8217;s still to do a depth of 30, I would still have to do this times 30. This had be 60 million invocations of the model. So that better be a small model, right? That&#8217;s a lot. So yeah.</p><p><strong>Ankit Gupta</strong></p><p>That&#8217;s to do a single action to be clear.</p><p><strong>Francois Chaubard</strong></p><p>Yeah, so exactly to do one action. So just imagine, so why AlphaGo doesn&#8217;t scale? To me there&#8217;s one, the cardinality of the action space must be extremely small. If it&#8217;s big, sad. Two, I need a perfect deterministic environment, right? This doesn&#8217;t change. The rules of this game don&#8217;t change, but the rules of the stock market change all the time. The rules to venture change all the time. The real world changes quite often. So homo skedastic and real time. If you saw the movie, the documentary was such an amazing documentary. I&#8217;d highly recommend it to anyone to watch it.</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s really good.</p><p><strong>Francois Chaubard</strong></p><p>The guy&#8217;s sitting there for 60 seconds, maybe five minutes waiting for the computer to decide. And it&#8217;s kind of like imagine that we were driving a car and you took 60 seconds to turn the steering wheel. Everyone&#8217;s dead. The whole car is dead. And so now let&#8217;s talk about robotics and self-driving car and why that approach can&#8217;t scale.</p><p><strong>Ankit Gupta</strong></p><p>Yeah, I think it&#8217;s a really good contrast here because intuitively, I think in thinking through this exact star layout, it actually really changed how I think about the problem space of both of these two. So let&#8217;s take self-driving car as an example. This is one many people have started to experience for the first time because we have some self-driving cars that actually work. You have Waymo and Tesla FSD and whatnot. They seem like they kind of work. So let&#8217;s maybe apply your same star framing here. I would contend that the state space of self-driving car is enormous and it&#8217;s actually not intuitive to me whether it&#8217;s more or less large than this one. I mean, in a sense, the chess and AlphaGo state space is already more than the number of atoms in the universe or something to that effect. But just to emphasize that here, you are considering surroundings, vehicle state, camera details and so [inaudible 00:37:25].</p><p><strong>Francois Chaubard</strong></p><p>Weather.</p><p><strong>Ankit Gupta</strong></p><p>Weather. I guess the point is-</p><p><strong>Francois Chaubard</strong></p><p>Road conditions.</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s massive. This is massive.</p><p><strong>Francois Chaubard</strong></p><p>For all intents and purpose, it&#8217;s infinite.</p><p><strong>Ankit Gupta</strong></p><p>Yeah. For all intents and purpose, it is infinite. Correct. Yeah.</p><p><strong>Francois Chaubard</strong></p><p>And so is the space of pixels. What can I put in an image? I can take an image of anything.</p><p><strong>Ankit Gupta</strong></p><p>Yes, true. True.</p><p><strong>Francois Chaubard</strong></p><p>And so we&#8217;re able to handle it. And same thing here where we compress from the board state. We don&#8217;t represent the board state. We compress it with a comnet. And so they have some deep comnet that actually takes this state and converts it into a latent. And that latent compression is sufficient to do pattern matching, do some type of symmetric equivariance kind of things. And same thing with this. And even better with JEPA, which we can talk about at the end there, which is basically taking some type of state space and doing all of our optimization in the latent space, which Stable Diffusion did, that worked extremely well, which reduces our state space dramatically because I&#8217;m in some latent high-dimensional space.</p><p><strong>Ankit Gupta</strong></p><p>So the key thing there is that despite this state space being effectively infinite, we&#8217;ve actually gotten really good at compressing this. And we&#8217;ll talk more about some of the tricks for how we actually do this in practice here, but the TLDR is where there&#8217;s 10 years of deep learning work that basically makes us extremely good at compressing that very fast.</p><p><strong>Francois Chaubard</strong></p><p>Exactly right. Exactly.</p><p><strong>Ankit Gupta</strong></p><p>T seems to have a similar problem as before. In fact, maybe even more extreme. There&#8217;s infinity other variables around you of things going on.</p><p><strong>Francois Chaubard</strong></p><p>Right. In some ways you&#8217;d think that... This is physics. Newton&#8217;s laws of motion should apply. If I turn the steering wheel like this or I hit the gas, I should be able to really easily model this. But what is non-differentiable is that if I&#8217;m going into a circle, the biggest issue that we faced when I was doing self-driving car is you are imposing your will onto maybe driving in India, I think is one of the [inaudible 00:39:22].</p><p><strong>Ankit Gupta</strong></p><p>Yeah. Exactly, yeah.</p><p><strong>Francois Chaubard</strong></p><p>You&#8217;re imposing your will onto the environment and people just kind of adapt naturally. If you were doing Newton&#8217;s motion, you were going to collide. And so that the optimal policy, if you were being strict Newtonians here would be like, don&#8217;t move because anything you do, you&#8217;re going to crash. But it&#8217;s not true. Then we wouldn&#8217;t function. Cars wouldn&#8217;t go down the road. And so you have to include other people in the environment and understand the embodiment of how your action will change other people&#8217;s actions.</p><p><strong>Speaker 3</strong></p><p>YC&#8217;s next batch is now taking applications. Got a startup in you? Apply at ycombinator.com/apply. It&#8217;s never too early and filling out the app will level up your idea. Okay, back to the video.</p><p><strong>Ankit Gupta</strong></p><p>Now let&#8217;s talk about the action space. One way to look at the action space is that it seems relatively small. It seems like, well, you turn the steering wheel left to right, you hit the brake, you hit the gas. It doesn&#8217;t seem that big, but how big is it actually? How do we actually represent these action spaces when it comes to a realistic self-driving car scenario?</p><p><strong>Francois Chaubard</strong></p><p>Yeah, I don&#8217;t know how they do this nowadays. They&#8217;re doing a whole bunch of bird&#8217;s eye view, different things like that.</p><h4>40:30 &#8212; Model-Free vs. Model-Based RL</h4><p><strong>Ankit Gupta</strong></p><p>Yeah. Let&#8217;s consider even just a very simplified case.</p><p><strong>Francois Chaubard</strong></p><p>But what do you have? You have a steering wheel that you can turn left, right. You have a brake pad and you have the gas. And so-</p><p><strong>Ankit Gupta</strong></p><p>This thing is 365 degrees. So it&#8217;s like a one to 365, let&#8217;s say, or zero to 365.</p><p><strong>Francois Chaubard</strong></p><p>Yep., And let&#8217;s just say you break this up into 10 different severities.</p><p><strong>Ankit Gupta</strong></p><p>Even with just this oversimplified model, your action space cardinality is 365,000. So that&#8217;s 100X bigger than AlphaGo. In fact, it&#8217;s about the size of the example.</p><p><strong>Francois Chaubard</strong></p><p>Yeah, right.</p><p><strong>Ankit Gupta</strong></p><p>In fact, a decent amount smaller than the size we said, which is brake and CPS.</p><p><strong>Francois Chaubard</strong></p><p>Exactly, yeah. And so yeah. So 36,000 action space is very large. And then even worse, unless you&#8217;re Tesla, we have a bunch of video of people driving cars. We don&#8217;t have video of dash cams like that. You actually don&#8217;t have, again, only Tesla has this, of the action as well. And so the things that you have access to, your trajectories are just like ST, ST plus one, ST plus two.</p><p><strong>Ankit Gupta</strong></p><p>So you&#8217;re saying there&#8217;s a decent number of these that&#8217;s from dash cam footage on YouTube or something, but not really that many either relative to complexity.</p><p><strong>Francois Chaubard</strong></p><p>And so if you wanted to do a self-driving car and you didn&#8217;t want to go spend a million dollars, trillion dollars on going collecting all this data, then you want to leverage this data somehow. And this is going to be really applicable for robotics because we have a lot of videos of people doing things, especially with egocentric. We have those videos, but what we don&#8217;t have is-</p><p><strong>Ankit Gupta</strong></p><p>The actions they take.</p><p><strong>Francois Chaubard</strong></p><p>Yeah.</p><p><strong>Ankit Gupta</strong></p><p>So this is a sequence of what you&#8217;re showing here.</p><p><strong>Francois Chaubard</strong></p><p>Unless you&#8217;re Tesla.</p><p><strong>Ankit Gupta</strong></p><p>Unless you&#8217;re Tesla.</p><p><strong>Francois Chaubard</strong></p><p>And Tesla has this. So this is a huge competitive moat of what do people do in that state? And then so you can behavior clone to go from here to here, from here to here, go here to here, et cetera. But even then, it&#8217;s still very, very difficult. It&#8217;s not sufficient. People think that, okay, I have this. We have a self-driving car, right? I mean, the amount of work that they&#8217;re doing at FSD is incredible and it&#8217;s not generally available. It&#8217;s not Waymo level yet.</p><p><strong>Ankit Gupta</strong></p><p>Would this be a good moment to briefly talk about model-free versus model-based RL?</p><p><strong>Francois Chaubard</strong></p><p>Yeah.</p><p><strong>Ankit Gupta</strong></p><p>I think that&#8217;s an important distinction that&#8217;s going to be relevant when you talk about more world models.</p><p><strong>Francois Chaubard</strong></p><p>Yeah, so this is a perfect point. So model-free just means that my policy pi of AT given ST, I have no world model involved. It&#8217;s literally doing what I said. I grab a bunch of these and I go from S to A, S to A, S to A.</p><p><strong>Ankit Gupta</strong></p><p>Yeah. Just predict the next day.</p><p><strong>Francois Chaubard</strong></p><p>That&#8217;s it. And this is logical VLA. This is giving us pretty good results. It&#8217;s behavior cloning. It&#8217;s all the stuff that it&#8217;s not getting us to Rosey the Robot just yet but-</p><p><strong>Ankit Gupta</strong></p><p>In many ways, it&#8217;s the closest thing that just looks like the next token prediction from LLMs that seems to scale pretty well with natural language.</p><p><strong>Francois Chaubard</strong></p><p>Exactly.</p><p><strong>Ankit Gupta</strong></p><p>I mean, it&#8217;s not exactly the same thing because there&#8217;s no action exactly, but picking a token is not exactly the same thing, but it&#8217;s very analogous to that basic thing [inaudible 00:43:33].</p><p><strong>Francois Chaubard</strong></p><p>Yeah. I basically take away the tokenizer head and I give it an action space and I collect a bunch of tele-ops data like this as the self-driving car does in Tesla. And I just take in the state, which is some image or maybe sequence of images and then I&#8217;ll output some action and that&#8217;s it. And this is, let&#8217;s say, model-free because I don&#8217;t have a model for the environment. And then now if I do model-based RL, I have not just some pi, but I have also SI as well here. And so by including this, I can have a much stronger policy, but it would take a lot more time to perform inference because I have to do this full test time planning.</p><h4>44:00 &#8212; Why robotics is the hardest case</h4><p><strong>Ankit Gupta</strong></p><p>Just to remind us, that SI is referring to this specific transition function.</p><p><strong>Francois Chaubard</strong></p><p>Exactly.</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s referring to this. You&#8217;re saying this is specifically referring to a function of ST plus one given ST and action T.</p><p><strong>Francois Chaubard</strong></p><p>Yes, exactly right.</p><p><strong>Ankit Gupta</strong></p><p>So it&#8217;s like your ability to predict the next state you&#8217;ll be in is the crux of it. As opposed to just directly predicting the actions.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. And the main thing that I believe is that this is required for AGI. This is what the human brain has been doing.</p><p><strong>Ankit Gupta</strong></p><p>At least in the way the human brain does it.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. And let me go further in saying that if you look at the billions of years of evolution, basically there&#8217;s this thing called 10 million years ago called the great cortical expansion, which you see the size of a brain just explode, get bigger, bigger, bigger exponentially up until us and it basically stops. And if the entire point of the neocortex is world modeling, what happened is we started from VLAs. This would be like ants-</p><p><strong>Ankit Gupta</strong></p><p>And fish or whatever. Yeah.</p><p><strong>Francois Chaubard</strong></p><p>And fish. Yeah, right. Just very lizard brain, whatever you want to call it. And then we develop this neocortex to go from our motor cortex to actually simulate what&#8217;s going to happen. And that makes us just so much smarter. And then once we get those samples, we can compress it when we sleep or otherwise with this hippocampal, short wave ripple, whatever you want to call it. And then that helps us develop a better policy. And that marriage between the two not only helps us train on hallucinated examples, but it also allows us to test time plan.</p><p><strong>Ankit Gupta</strong></p><p>I guess the extreme case then of self-driving car is general robotics.</p><p><strong>Francois Chaubard</strong></p><p>Yes.</p><p><strong>Ankit Gupta</strong></p><p>Right. So if you&#8217;re a humanoid company like figure or pi or whatever, again, same STAR setup. I guess the gist of it is that A is now even bigger. I guess a very simple robot would be, how would you parameterize the action space? Let&#8217;s take a very basic one.</p><p><strong>Francois Chaubard</strong></p><p>If I take my six-axis arm as your standard here that we&#8217;re actually working on right now in Stanford Robotic Center, you have two degrees of freedom, two degrees of freedom, two degrees of freedom. And then you have another two for the end effector. And so the end effector-</p><p><strong>Ankit Gupta</strong></p><p>That&#8217;s a simple end effector, not even like a fancy one.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. It&#8217;s literally a one-axis. You can rotate, but you have the one-axis Yumi style thing. So this is eight. So you have 16 degrees of freedom. And let&#8217;s just say that you do the 365 divided 10 or whatever kind of thing. I mean-</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s like 10 to the 16.</p><p><strong>Francois Chaubard</strong></p><p>It&#8217;s insane.</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s something like that.</p><p><strong>Francois Chaubard</strong></p><p>It&#8217;s an insane number. And so much bigger than self-driving car. And even worse, getting tele-ops data is extremely painful and expensive. It&#8217;s not just like, &#8220;Oh, we&#8217;ll just get some people in the Philippines, we&#8217;ll give them some things or whatever.&#8221; It totally, totally doesn&#8217;t work.</p><p><strong>Ankit Gupta</strong></p><p>And nor is there yet something like Tesla&#8217;s fleet where there are cars deployed that people are just using and they&#8217;re not even necessarily realizing that every time they turn the steering wheel, they&#8217;re providing this dataset for Tesla to train on.</p><p><strong>Francois Chaubard</strong></p><p>And then even worse, you have this what&#8217;s called cross embodiment gap. And so if I were to train this policy on Tesla Model X and I were to put it on a Tesla Model 3, it wouldn&#8217;t work. It totally wouldn&#8217;t work. So much of this, the way that if I were to brake on a Model 3 versus a Model X, the Model X, it weighs more. It has different dynamics, aerodynamics and things like that. And so what&#8217;s actually going to happen is very different. The degradation you have across embodiments is very, very, very strong.</p><p><strong>Ankit Gupta</strong></p><p>And clearly Tesla&#8217;s figured various ways to get around that. I mean, they have these that roll out, but actually even with Tesla as a new FSD today, they don&#8217;t roll out in all the cars at the same time probably for more or less that reason. And in this case, it&#8217;s even harder now. I mean, you have bigger differences between embodiments than a Model 3 versus Y, and you have way bigger action spaces you have to sell a model.</p><h4>48:20 &#8212; World models that actually work</h4><p><strong>Francois Chaubard</strong></p><p>Yeah. Lane McIntosh, I played hockey with at Stanford who now runs Tesla FSD. I can ask him, but I would bet money that they shard the data per model per car type.</p><p><strong>Ankit Gupta</strong></p><p>Yeah, wouldn&#8217;t be surprised.</p><p><strong>Francois Chaubard</strong></p><p>Because that&#8217;s what I would do. There&#8217;s no way that I would trust data that was collected on a Model X on a Model 3. No way I would trust it.</p><p><strong>Ankit Gupta</strong></p><p>Okay. So now that we understand the basic setup here and why the action space problem is so big, why don&#8217;t we talk a little bit about how world models actually fit into this? Maybe first, I guess what didn&#8217;t work about the naive world models and how do we fix those? And then let&#8217;s talk about some of the newest world modeling techniques.</p><p><strong>Francois Chaubard</strong></p><p>Cool. So in robotics in particular, it&#8217;s very hard to get this kind of trajectories that you want, that you need to train for your VLAs. And people spend up with a whole bunch of tele-ops data. It&#8217;s very expensive, very expensive. Ideally, what we would do is take data like this from someone who just puts a camera on them and just making sushi. I want to make a sushi robot. How do I do it? Give it to all the sushi chefs, don&#8217;t put anything in their hands and just have them start cutting up sushi and making sushi.</p><p><strong>Ankit Gupta</strong></p><p>And ideally, we would train it in that way you were describing of somehow we would train a model just on these two and then later add this.</p><p><strong>Francois Chaubard</strong></p><p>Yes.</p><p><strong>Ankit Gupta</strong></p><p>Afterwards.</p><p><strong>Francois Chaubard</strong></p><p>And so the first real person that went after this was J&#252;rgen Schmidhuber. So he doesn&#8217;t yell at us, we have to make sure we cite him. But he has this really cool paper called World Models, very aptly named. And it&#8217;s basically he took these OpenAI gym classic games, car racing and I think Doom as well. And then just trained a model at that time was an RNN. He had some funky zero order stuff in there or whatever. But basically the key premise was I can take an environment, I can extract a whole bunch of this type of data off of it. I think he actually does actually this data, but we&#8217;ll get into Dreamer where he does it in this way. And then trains a policy on only the synthetic data, the imaginated rollouts. And it actually performs well in the environment. This is the first time in my understanding that that actually happened and it actually works really well. And then-</p><p><strong>Ankit Gupta</strong></p><p>So the key thing there is you can basically use this if you have some predictive model of this in that case and eventually of this, you can use that as basically a synthetic training set to train your policy model and then basically fine-tune it on real data later.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. And which is just a really powerful idea, especially since in robotics, the limiting step is access to large amounts of state action data. And so now the Dreamer series, so basically this publishes in May of 2018. Danijar Hafner publishes Dreamer 1 I think in November of 2018. And then now he&#8217;s been on this rampage for the last seven years publishing these papers. And Dreamer V4 I think is the capstone of it where he basically does the same thing and he focuses on Minecraft and he trains a world model on this type of data and then injects action conditioning on a very small amount of data to get to this type of world model that has the action conditioning as well. And then samples a lot from it and then trains a policy on those synthetic imaginated rollouts. And the policy is so good that it&#8217;s the first paper to mine diamonds in Minecraft. I&#8217;m not a big Minecraft player, but apparently that&#8217;s extremely difficult. That&#8217;s next level difficulty and it did it all on synthetic data, which is kind of crazy.</p><p><strong>Ankit Gupta</strong></p><p>And the key unlock there, yeah, you use synthetic data specifically on a model trained on just this sort of state transition type of thing.</p><p><strong>Francois Chaubard</strong></p><p>Yes.</p><p><strong>Ankit Gupta</strong></p><p>And this ends up being very convenient because it turns out we as a society have a lot of this.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. Yeah, all of YouTube, right? He does do a very small amount of data to enable the action conditioning and that allows you to do this full simulated rollout. But yeah, it&#8217;s true. So we have YouTube, we have Flickr, we have all these datasets online of people doing things. We&#8217;d like to use it and no one has really gotten that to work. And then now that with these video generation models, we can take that data, create a world model out of it, add action conditioning, post-train it with action conditioning for some new task that we want it to do, chopping down wood or making sushi or folding my bed or whatever it is, only a few amount of examples. And then we can train a policy in this neural simulation.</p><p><strong>Ankit Gupta</strong></p><p>And we put out a video about diffusion models very recently and flow matching. I imagine that now ties very closely to this. Ultimately, the current state-of-the-art best way to do this on basically infinity data that we have available and can keep generating is using state-of-the-art video dfusion/flow matching models.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. Yeah. So if you have your CDANCE or your SORA or-</p><p><strong>Ankit Gupta</strong></p><p>ONE.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. All those models, basically the idea is now we have them and they&#8217;re already trained and they&#8217;re great. Let&#8217;s do a small amount of action conditioning on them to get to this world model and then we can sample from it a bunch and then train. And this is exactly what Wayve did with GAIA. And GAIA, I think they&#8217;ve raised $1.5 billion to basically run with this idea for self-driving car. I think a bunch of companies, Nvidia, this paper here is basically talking about doing exactly the same, this Dream 0 for robotics.</p><h4>54:10 &#8212; JEPA &amp; latent space tricks</h4><p><strong>Ankit Gupta</strong></p><p>What I thought was really cool about this paper is that they do exactly this process where they have this joint model of state transitions and actions. They train it by first instantiating it with the open source one video diffusion model. And then it only takes them about 500 hours of tele-op data, which is basically exactly this, to get it to be pretty good. And they have a lot of clever tricks that allowed it to be cross-embodiment and working on scene tasks with relatively small amounts of data. And it really is taking basically the exact concept, I believe, from the Dreamer paper and applying it specifically to these robot embodiments.</p><p><strong>Francois Chaubard</strong></p><p>Exactly.</p><p><strong>Ankit Gupta</strong></p><p>And it turns out it actually works actually better than I would&#8217;ve anticipated it working.</p><p><strong>Francois Chaubard</strong></p><p>Right. So yeah. So I think that this is basically the path, it was the path I believe is the path to get humans to be as good as we are genetically over the last 10, 20 million years of evolution. A bigger world model helps for training and for test time planning. And I think it&#8217;ll be the same thing as true for robotics.</p><p><strong>Ankit Gupta</strong></p><p>What&#8217;s also cool is there&#8217;s a bunch of applications of this to things outside of robotics too. I mean, there was a weather planning paper, for example. We were reading this Gencast paper, which I think applies a relatively similar concept in terms of how they model literally the world&#8217;s weather with something like this.</p><p><strong>Francois Chaubard</strong></p><p>Yeah, we have to talk about the world model for the world. Yeah. So basically they do this exact same thing where the key unlocks for this whole thing was getting diffusion to work in very high-dimensional state spaces like we talked about in the last lecture and then learning to use that to action condition in the way that he&#8217;s done. But they did this for the entire world with this exact same diffusion steps, which go from some... And they go back to two time steps, lag of order two, AR2 for the statisticians there. And then basically predict the next state of the world based on those things with this lingo and diffusion rollouts. My big assertion is that it was necessary for the human brain to develop world modeling. I actually just saw this paper that I wanted to make sure to call out because I though it was so great out of University of Washington where they say explicitly in the abstract, each cortical area estimates both latent sensory states and actions and the cortex as a whole predicts the consequences of those actions. That sounds like a world model to me.</p><p><strong>Ankit Gupta</strong></p><p>Yeah.</p><p><strong>Francois Chaubard</strong></p><p>Right?</p><p><strong>Ankit Gupta</strong></p><p>Yeah.</p><p><strong>Francois Chaubard</strong></p><p>And so-</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s actually describing exactly these two equations here.</p><p><strong>Francois Chaubard</strong></p><p>Right. Exactly, right.</p><p><strong>Ankit Gupta</strong></p><p>Where we&#8217;re estimating both the sensory latent states and actions. I mean, I guess it&#8217;s really the joint model that we showed earlier is what he&#8217;s describing here. It&#8217;s exactly this equation we&#8217;re showing now.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. Exactly right. And so if it works in us, it should work in robotics. And I think that that takes us the rest of the distance.</p><p><strong>Ankit Gupta</strong></p><p>Why don&#8217;t we talk briefly about latent world models, especially the JEPA concept, because I think there&#8217;s been a number of papers that use JEPA as an element of their, I guess, architecture. Why don&#8217;t we just briefly introduce JEPA and how it fits into the current landscape of world modeling?</p><p><strong>Francois Chaubard</strong></p><p>Yeah. In classic RL, if you study Q learning, for example, you basically keep this matrix called the Q matrix and it&#8217;s going to be S by A. And so I have this S by-</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s states and actions.</p><p><strong>Francois Chaubard</strong></p><p>States and actions. And each one I need some amount of counts of being in this state action. And I take the average value of taking that action in this state and that&#8217;s my Q value there. And it&#8217;s a little bit more complicated in that. There&#8217;s Bellman equation, all this backup, all this stuff like that. So this scales horribly because as the cardinality of my state space gets bigger and my cardinal action space gets bigger, stuff, I don&#8217;t have enough, I become less and less sample efficient, right?</p><p><strong>Ankit Gupta</strong></p><p>In the case of robots or whatever, state is like, yeah, it&#8217;s this whole thing we described earlier. It&#8217;s absolutely massive because it has all of these elements in it. You couldn&#8217;t really enumerate a huge grid.</p><p><strong>Francois Chaubard</strong></p><p>And so the classic trick, I mean, since I took 229 with Androung in 2012 is you do this.</p><p><strong>Ankit Gupta</strong></p><p>Stick a neurolab, work on it.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. And you basically are just going to compress that state into some lower dimensional state space. This actually predates deep learning. We were doing stuff like this. I think my first paper was basically doing something like this, basically turning a grid into a bunch of pyramids and the state was how much I&#8217;m in pyramid one or pyramid two or whatever. But anyway, the neural networking can just do this. And so basically the key idea in JEPA, if I have an image one and I have image two and I have image three, I can do my world modeling of ST plus one given ST and AT in pixel space and have, this is let&#8217;s say at time T, T plus one, T plus two, et cetera, et cetera. And I have to actually predict now the full image that&#8217;s extremely expensive from a computation standpoint and also from a sample efficiency standpoint. What I can do instead is put this through some comnet.</p><h4>59:00 &#8212; Open problems remaining</h4><p><strong>Ankit Gupta</strong></p><p>Some encoder.</p><p><strong>Francois Chaubard</strong></p><p>Some encoder. And then I&#8217;ll get a latent for T and I&#8217;ll have a latent for T plus one and I&#8217;ll have a latent for ZT plus two. And then I&#8217;ll have from this, from ZT, I want to predict ZT plus one hat. And my goal is to make this and this and my loss function will be something very simple like I want to minimize this.</p><p><strong>Ankit Gupta</strong></p><p>Minimize that, yeah.</p><p><strong>Francois Chaubard</strong></p><p>That&#8217;s it. Now this doesn&#8217;t work. This collapse is hard. And so what happens is basically if you just predict zero, done.</p><p><strong>Ankit Gupta</strong></p><p>Yeah. It works. Yeah.</p><p><strong>Francois Chaubard</strong></p><p>Just output zero, which the model will learn to do. And I&#8217;m actually incorporating this into my current research right now. And so what you need to do is something called SIG Reg or this is one technique, VIC Reg is another. Where basically I add this another term that basically says over a large enough batch size, I want the distribution of ZT plus one to follow a Gaussian-</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s kind of like a normalized, like a batch norm type of trick. I mean not in the same case, yeah.</p><p><strong>Francois Chaubard</strong></p><p>And if it&#8217;s zero, it can&#8217;t be this because then this is non-zero. And so maybe I think that there&#8217;s probably this or something like that. But basically this prevents it from model collapse and it makes it do something good. And this is the most recent paper for the audience is LEWM, LE World Model, which is super, super great. However, to be completely frank, this is self-supervised learning, super great. It doesn&#8217;t work that well. If you were to not do these techniques and there&#8217;s a bunch of other techniques that you can do, it will actually outperform much better that are, let&#8217;s say for example, if I&#8217;m going to do an LLM. And you have Francois likes sushi, which is definitely true. And I tokenize this into a bunch of different tokens here and this is token ID 6, 19, 28, whatever, and I look up the encoding into this and that&#8217;s going to be E1, E2, E3, et cetera. What you can actually do is have the LLM output. The LLM will take in these things and will output the next token. And so it&#8217;d be like, let&#8217;s call it H. This would be the low jits coming out of it, two plus one. And what you can do is actually have this be close to ET plus one. And a lot of people are playing with this idea and getting rid of the cross-entropy loss entirely. And so if you were to do this, it actually is a proxy for the cross-entropy loss and there is no cross-entropy loss. And the cross-entropy head is actually very expensive. And so this is very cheap and this is literally just grabbing it. So people are playing around with this idea and basically as a cheaper proxy for the cross-entropy loss. So there&#8217;s lots of different ideas on basically taking this JEPA idea to not just pixels, but to LMs as well.</p><p><strong>Ankit Gupta</strong></p><p>Yeah, interesting. Yeah.</p><p><strong>Francois Chaubard</strong></p><p>So just to define what JEPA is, it&#8217;s joint embedding predictive architecture.</p><p><strong>Ankit Gupta</strong></p><p>I think one of the things I find cool about this JEPA idea is it feels like an idea we see over and over in deep learning. There&#8217;s a version of this idea that&#8217;s basically the staple diffusion idea. There&#8217;s a version of this idea that in my company training graph convolutional neural networks to design drugs we use to do latent variable generation, for example. And it&#8217;s an idea that comes back over and over and then has this various tricks that it actually takes to get it to work in practice.</p><p><strong>Francois Chaubard</strong></p><p>Yeah, yeah, yeah.</p><p><strong>Ankit Gupta</strong></p><p>Okay. Now we have a pretty good sense for how world models work. We have a pretty good sense for what the state of the art looks like. If we trust this paper, and it seems like these kind of work on robots too. I mean this paper&#8217;s only from the end of last year into this year, and it seems like they have various methods that allow you to train on relatively small amounts of data that&#8217;s tractable and pre-train on diffusion models. So are we good?</p><p><strong>Francois Chaubard</strong></p><p>We&#8217;re done.</p><p><strong>Ankit Gupta</strong></p><p>Does it all work?</p><p><strong>Francois Chaubard</strong></p><p>Yeah. 2026 will be the year of the robot. We&#8217;re going to have Rosey the Robot in your house. Yeah, no, I don&#8217;t think so.</p><p><strong>Ankit Gupta</strong></p><p>What are one or two, because there&#8217;s lots of open problems remaining, what are a few open problems maybe we can emphasize here that the community can go emphasize working on?</p><h4>01:04:30 &#8212; Does this pass the squint test?</h4><p><strong>Francois Chaubard</strong></p><p>Yeah. So I think the first one is that PINNs doesn&#8217;t really work. What is PINNs? Physics-informed neural networks. So PINNs doesn&#8217;t really work. There&#8217;s physics-informed neural networks. And so basically if almost all of the self-driving car data looks like this, the car is driving down the road. And let&#8217;s just say, for example, I have a house here and I want to train the model on you not driving into the house. And so let&#8217;s say I put it into a state right here to drive into the house. What&#8217;s going to happen is because almost all the data looks like this driving down the road, this will just turn magically into a highway. And they&#8217;re just like, &#8220;Boo, just don&#8217;t worry. You crash it all.&#8221;</p><p><strong>Ankit Gupta</strong></p><p>Basically it needs a ton of data not to do that either from simulation for that to not happen.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. In fact, I actually don&#8217;t even know if because of the data distribution, there&#8217;s no data here. There&#8217;s almost all the data here. And when you&#8217;re training a neural network, it has a tendency to collapse if you don&#8217;t keep the mini batch composition very even over the class space or whatever you want to call it. But you have to be very careful about your data mixing to make sure you get this right to solve this problem that no one really has. But even then, if you take just a simple thing like this, this is the kind of example, and I have some sine wave and I have these as my X. And I have these as my Y. So this is complete interpolation. I may mess this up, but Y, like this. We can&#8217;t get to machine precision. What is it? Minus 16 or whatever it is. The SGD will not get to effectively zero. So we&#8217;ll always have some residual. And for us to be a really good world model, to simulate body interactions, to simulate this, what&#8217;s going to happen when I do this? And let&#8217;s say that I&#8217;m trying to be LeBron James. I saw this one video of Steph Curry dribbling a basketball on a court and he just felt that there was a dead spot in the court because he&#8217;s so good and he knows exactly the physics of what&#8217;s going to happen. If I hit the ball with this force, the ball&#8217;s going to come back exactly this spot and it just didn&#8217;t. And he knew it wasn&#8217;t him, it was the court and he found a dead spot in the court. That&#8217;s how good the human brain is at world modeling. In my opinion, I think it&#8217;s an SGD issue. I think it&#8217;s probably an architecture issue. I think Sam Altman just came and just said that he thinks that there&#8217;s definitely an architecture that&#8217;s going to be more performant than the transformer. I think he&#8217;s right. I think the transformer doesn&#8217;t do compression in the time domain at all. It just keeps around everything. So anyway, so I think that getting higher fidelity in the world model is extremely important, one. I think two-</p><p><strong>Ankit Gupta</strong></p><p>Seems like test time probably is going to be a thing, like adaptation.</p><p><strong>Francois Chaubard</strong></p><p>Exactly. Test time planning, how quickly the human brain can... In sports and things like that, when you&#8217;re playing tennis, when you&#8217;re a tennis player, how quickly we can adapt to what a player is doing and things like that. We&#8217;re not going to sleep and retraining. We&#8217;re very quick to adapt to a new environment.</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s like the out of distribution prediction.</p><p><strong>Francois Chaubard</strong></p><p>Exactly.</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s really challenging.</p><p><strong>Francois Chaubard</strong></p><p>And one little data point we can quickly adapt to that new thing and change. I think there&#8217;s been a lot of papers on basically estimating the friction coefficients. And so those can change over time if you go to a human environment or not, for example, this friction might change and that&#8217;s important in control. And so you need to estimate that very quickly and adapt. And these models just don&#8217;t have a mechanism to do it.</p><h4>01:08:00 &#8212; Outro</h4><p><strong>Ankit Gupta</strong></p><p>And then I guess there&#8217;s the practical speed elements of these. A lot of these are doing some sort of expensive planning step and we&#8217;re doing some sort of... We&#8217;re hacking around it with this pre-training process and synthetic data. But even so, to really get maximum performance right now, you&#8217;d want to do something that&#8217;s closer to the AlphaGo style rollout and that&#8217;s extremely slow.</p><p><strong>Francois Chaubard</strong></p><p>Right. The MCTS process can&#8217;t happen. The other thing that is pretty crazy about the way that the brain works is that everything is kind of running autonomously. And so you might be in the middle of saying sentence one and then be like, &#8220;Oh, actually no, something else.&#8221; And so what does happen there? It&#8217;s like type one and type two thinking are happening at the same time in some way. And so there&#8217;s definitely some really cool mix of these heterogeneous models and some are overriding others and taking control of the motor cortex and commanding the body to do a thing.</p><p><strong>Ankit Gupta</strong></p><p>Okay. But on the flip side now, we talked in the past video about the squint test and how we felt that auto-aggressive LLMs maybe don&#8217;t pass the squint test. Why don&#8217;t we reintroduce what the squint test was for a second? And then maybe let&#8217;s think about whether this passes the squint test despite all those limitations.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. And the squint test for me I think is like, this comes from the Yann LeCun. We didn&#8217;t need flapping wings to achieve flight. And to that I say, &#8220;Well, we did need two wings.&#8221; And if I squint and I look at a bird and I squint and I look at a plane, I&#8217;m like, &#8220;Yeah.&#8221;</p><p><strong>Ankit Gupta</strong></p><p>It&#8217;s kind of similar.</p><p><strong>Francois Chaubard</strong></p><p>It looks right. Similarly, if I squint and I look at the human brain and I squint and I look at all these world models, we have this VLA, this action policy, and that they&#8217;re doing test time planning together and things like that. It&#8217;s getting really close. It&#8217;s much, much closer.</p><p><strong>Ankit Gupta</strong></p><p>It seems closer than an auto-aggressive LLM.</p><p><strong>Francois Chaubard</strong></p><p>100%.</p><p><strong>Ankit Gupta</strong></p><p>And this concept of a world model of implicitly predicting future states and actions feels intuitively like what our brain&#8217;s doing. And it seems like there&#8217;s some neuroscience evidence to support that.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. I mean, I&#8217;m getting to the conclusion that I think that the brain is the optimizer, not the model and that the brain emits, has models that it invokes, but the brain is somehow also the optimizer itself. And so in that way it doesn&#8217;t pass the squint because something magical is happening when you&#8217;re sleeping. There&#8217;s no intelligent species that we&#8217;re aware of that have any amount of intelligence that don&#8217;t sleep. And so octopuses, dolphins, all those, elephants, they all sleep. There&#8217;s some reason for that. And that seems like a really thing about the evolutionary recourse of sleeping, you get eaten when you sleep. So for the benefit of sleeping should be so much better to outperform that. So I think we don&#8217;t have this idea of awake sleep in our current architecture, but I can imagine I&#8217;m simulating compress from the hippocampus some experience in the day. I&#8217;m training on more of those examples. Right?</p><p><strong>Ankit Gupta</strong></p><p>You&#8217;re collecting a whole bunch of these experienced rollouts and then you&#8217;re updating your policy function overnight or something like that. Yeah.</p><p><strong>Francois Chaubard</strong></p><p>There&#8217;s got to be something. There&#8217;s this thing called shortwave ripple where the hippocampus when you&#8217;re sleeping emits these spike trains that are actually reversed from when they actually happen back in through both the hemispheres and for seven times and then it stops.</p><p><strong>Ankit Gupta</strong></p><p>Interesting.</p><p><strong>Francois Chaubard</strong></p><p>So there&#8217;s something happening there that&#8217;s very training something. And if you don&#8217;t sleep, then you don&#8217;t have long-term memory.</p><p><strong>Ankit Gupta</strong></p><p>Right.</p><p><strong>Francois Chaubard</strong></p><p>Right? And so there&#8217;s definitely a reason why we&#8217;re training things that happened into our brain.</p><p><strong>Ankit Gupta</strong></p><p>So where does that put us now? We have all this work happening with world models. How should we think about what&#8217;s coming ahead in these next few years in the research community?</p><p><strong>Francois Chaubard</strong></p><p>Yeah, I think that we&#8217;re going to see a lot more of these world models in robotic policies. I think that&#8217;s going to unlock probably full self-driving would be one of those examples that they can get the real timeness of it.</p><p><strong>Ankit Gupta</strong></p><p>It seems like that&#8217;s coming.</p><p><strong>Francois Chaubard</strong></p><p>They could probably solve it with more compute to have parallel things and you probably don&#8217;t need it for most standard things. Maybe getting out of weird parking jams and things like that would take us some time similar to the Rosey the Robot, which we&#8217;ve always wanted to have a Rosey the Robot to clean up my room for me. I think that this feels like we&#8217;re getting to good enough that we can pay up for data in compute to get to Rosey the Robot. It does feel like that. It&#8217;ll be expensive to collect the data and do the Dreamer sequence of going from state to state and then getting the action conditioning to work. But I feel like it should work.</p><p><strong>Ankit Gupta</strong></p><p>Yeah. I mean, what&#8217;s pretty cool is we see a lot of companies at YC working at every step of this from the collecting egocentric data, collecting the tele-op data, training their own world models and action models, building new embodiments and then making ways of adapting those embodiments. And it feels like this is the first year where you see demos where you&#8217;re like, okay, this actually kind of is starting to look like it&#8217;s going somewhere. And it seems like a very exciting year ahead.</p><p><strong>Francois Chaubard</strong></p><p>Yeah. So anyway, I think that there are real AI problems to solve still. We talked about PINNs, we talked about the real time issues. And then on the robotics side, there&#8217;s real issues. It&#8217;s amazing how effective our epidermis is in terms of we can detect tactile.</p><p><strong>Ankit Gupta</strong></p><p>Oh, epidermis.</p><p><strong>Francois Chaubard</strong></p><p>Yeah, epidermis. Our tactile, we can detect sheer force, we can detect temperature, and it&#8217;s everywhere. And so versus we get one little sensor that only does tactile. We don&#8217;t have the friction component. We don&#8217;t have temperature. We don&#8217;t have all the feeling. We can&#8217;t estimate coefficient of friction very quickly. I can touch something and say, oh, this is smooth, this is rough. We don&#8217;t have any of that. And if I numb your hands, I actually had this experience just recently, if I numb your hands, you actually can&#8217;t tie your shoes.</p><p><strong>Ankit Gupta</strong></p><p>Yeah, interesting.</p><p><strong>Francois Chaubard</strong></p><p>So you can&#8217;t perform control. And so yeah, if you train on enough human data tying your laces, do I think you can do it with no feedback? Maybe. Maybe. But how much would you need if you did actually have the human touch? I think it&#8217;d be so much easier.</p><p><strong>Ankit Gupta</strong></p><p>Yeah. Well, there&#8217;s a lot of more research to do then.</p><p><strong>Francois Chaubard</strong></p><p>Yeah, yeah.</p><p><strong>Ankit Gupta</strong></p><p>Francois, thanks so much for joining us. Thanks so much for watching everyone. We&#8217;ll be back for the next episode of Decoded.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.ycrootaccess.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>