Ep18. Jensen Recap - Competitive Moat, X.AI, Smart Assistant | BG2 w/ Bill Gurley & Brad Gerstner

BG2Pod with Brad Gerstner and Bill Gurley · October 2024 · avg confidence 0.77
Watch original ↗
⚠ 4 span(s) flagged for spot-check (alignment confidence < 0.50) — verify these against the audio
  1. [00:23:11] Bill Gurley (0.24) — Yeah.
  2. [00:15:34] Brad Gerstner (0.37) — Hmm.
  3. [00:35:36] Speaker 1 (0.40) — Right.
  4. [00:18:46] Speaker 1 (0.41) — Yeah.
Bill GurleyBrad GerstnerSunnySpeaker 1
Bill Gurley00:00:00
You may also be running up against the, even for the Mag 7, the size of CapEx deployment where their CFOs start to talk at higher levels. Yes, for sure. Totally. Thank you.
Brad Gerstner00:00:24
Sunny, Bill, great to see you guys.
Bill Gurley00:00:26
Good to see you.
Brad Gerstner00:00:28
Good to be back. Thanks, man. It's great to have you. We literally just finished two days of the Altimeter annual meetings. I mean, we had hundreds of investors, CEOs, founders, and the theme was scaling intelligence to AGI. We had Nikesh talking about enterprise AI. We had Rene Haas talking about AI at the edge. We had Noam Brown talking about, you know, the strawberry and no one model and inference time reasoning. We had Sonny talking about, you know, accelerating inference. And of course, we kicked off with Jensen talking about the future of compute. You know, I did the Jensen talk with my partner, Clark Tang, who covers the compute layer in the public side. We recorded it on Friday. We'll be releasing it as part of this pod.
Brad Gerstner00:01:14
And man, was it dense. I mean, he was, you know, he was on fire. He told me, I asked him at the beginning of the pod, what do you want to do? He said, grip it and rip it. And we did. 90 minutes, we went deep. I shared it with you guys. We've all listened to it. I learned so much playing it back that I just thought it made sense for us to unpack it, right? To really analyze it, see what we agree with, what we may disagree with, things we want to further explore. Sunny, any high-level reactions to it?
Sunny00:01:44
Yeah, you know, first of all, it's the first time I've really seen them in a format where you got all that information out in one setting because you kind of get the tidbits. And the ones that really struck with me was when he said NVIDIA is not a GPU company, they're an accelerated compute company. I think the next one, you know, which you'll touch on is where he's really said the data center is the unit of compute. I thought that was, that was massive. And, you know, sort of just closing out when he talked about he thinks about using and already utilizing so much AI within NVIDIA and how that's a superpower for them to accelerate over everyone they're competing with. I thought those were kind of really awesome points and him, you know, eating the dog food, as they say.
Brad Gerstner00:02:29
It is incredible. There's this thing we'll talk about later, but he said he thinks they can 3X the top line of the business while only adding 25% more humans because they can have 100,000 autonomous agents doing things like building the software, doing the security, and that he becomes really a prompt engineer, not only for his human direct reports, but also for these agents, which really is mind-boggling. Bill, anything stand out for you?
Bill Gurley00:03:00
Well, one, I mean, you should be pleased that you were able to get his time. You know, this is at points in time, the largest market cap company in the world, if not one, two. And so it was so, I think, kind of him to sit down with you for so long. And during the pod, he kept saying, I can stay as long as you want. And I was like, doesn't he have something to be doing?
Brad Gerstner00:03:24
It was incredibly generous.
Bill Gurley00:03:28
It's fantastic. But my other big, I mean, I had two big takeaways. One, I mean, it's obvious that this guy's, you know, rolling on all cylinders here, right? Like you have a company at a 3.3 trillion market cap that's still growing over 100% a year. And the margins are insane. I mean, 65% operating margins. There's only like five companies in the S&P 500 at that level. And they certainly aren't growing at this pace. And when you bring up that point about getting more done on the increment with fewer employees, where is this going to go? Like 80% operating margin? I mean, that would be unprecedented. There's a lot that's already here that's unprecedented. But obviously, Wall Street is fully aware of the unbelievable performance of this company.
Bill Gurley00:04:21
And, you know, the multiples reflected in it. The market cap reflects it, but it's super powerful how they're executing. And you can see the confidence in every answer that he gives.
Brad Gerstner00:04:32
We spent about a third of the pod on NVIDIA's competitive moat, really trying to break it down, really trying to understand this idea of systems-level advantages, the combinatorial advantages that he has in the business. Because I think when I talk to people around the investment community, despite how well it's covered, Bill, right, there's still this idea that it's just a GPU and that somebody is going to build a better chip. They're going to come along and displace the business. And so when he said, again, it can sound like marketing speak, Sunny, when somebody says it's not a GPU company, it's an accelerated compute company. We showed this chart where you can see kind of the NVIDIA full stack.
Brad Gerstner00:05:15
And he talked about how he just built layer after layer after layer of the stack over the course of the last decade and a half. But when he said that, Sunny, I know you had a reaction to it. Even though you know it's not just a GPU company, when he really broke it down, it seemed like he did break new territory here.
Sunny00:05:36
Yeah, what was great to hear from him and really positive for folks thinking about where NVIDIA lives in the stack right now is he kind of got into details and then the sub-details below CUDA. And he really started going into what they're doing very particularly on mathematical operations to accelerate their partners and how they work really closely with their partners, you know, all the cloud service providers, to basically build these functions so that they can further accelerate workloads. The other little nuance that I picked up in there: he didn't focus purely on LLMs. He talked in that particular area about how they're doing that for a lot of traditional models and even newer models that are being deployed for AI.
Sunny00:06:20
And I think just really showed how they are partnering much closer on the software layer than the hardware layer alone, right?
Brad Gerstner00:06:27
Right. I mean, in fact, you know, he talked about, you know, the CUDA library now has over 300 industry-specific acceleration algorithms, right, where they deeply learn the industry, right? So whether this is synthetic biology or this is image generation or this is autonomous driving, they learn the needs of that industry and then they accelerate the particular workloads. And that, for me, was also one of the key things—this idea that every workload is moving from kind of this deterministic, handmade workload to something that's really driven by machine learning and really infused with AI and therefore benefits from acceleration, even something as ubiquitous as data processing.
Sunny00:07:15
Yeah, and I shared this code sample with Bill as we were just preparing for this pod. And I knew Bill processed it right away and then ran it, which was, it really showed every piece of code that's out there now that's related to—or not every piece, many of the pieces have this sort of 'if device equals CUDA, do X, and if it's not, do Y.' And that's the level of impact they're having across the entire ecosystem of services and apps that are being built that are related to AI. Bill, I don't know what you thought when you saw that piece.
Bill Gurley00:07:46
Yeah, I mean, I think there's a question for the long term that relates to CUDA. And I want to go back to the system point you made later, Brad, but while we're on CUDA, is what percentage of developers will touch CUDA and is that number going up or down? And I could see arguments on both sides. You could say the AI models are gonna get more and more hyper-specialized and performance matters so much that the models that matter the most, the deployments that matter the most, they're gonna get as close to the metal as possible and then CUDA is gonna matter. The other side you can make is, those optimizations are going to live in PyTorch. They're going to live in other tools like that. And the marginal developer's not going to need to know that.
Bill Gurley00:08:37
And I could make both arguments, but I think it's an interesting question going forward.
Brad Gerstner00:08:42
I mean, I just asked ChatGPT how many CUDA developers there are today, just to be on top of it. Three million CUDA developers, right? And, you know, a lot more that touch CUDA that, you know, aren't specifically kind of developing on CUDA. So it is one of these things that has become pretty ubiquitous. And his point was, it's not just CUDA, of course. It's, you know, it's really full stack, you know, all the way from data ingestion, all the way through, you know, kind of the post-training.
Sunny00:09:07
I think I'm on the latter of your point, Bill. Like, I think there's going to be fewer people touching that. And I do think that's a point where the moat is not as strong as a longer term, as you say. And think about like, you know, the way, the analogy that I would go with is like, think about the number of iPhone, iOS, like developers working at Apple, building that versus the number of app developers, right? And I think you're going to have a, you know, 10 to one or a hundred to one ratio of people building at layers above versus people building down closer to the bare metal.
Bill Gurley00:09:36
That'll be something to watch. We can ask more people over time. Obviously, it's a big lock today for sure.
Brad Gerstner00:09:43
You know, and I think, Bill, to your point, you know, I reached out to Gavin. Actually, before I did the interview, Gavin Baker, who's a good buddy and who obviously knows the space incredibly well, has followed it at a deeper level for a longer period of time than I have. And, you know, like when I asked him about the competitive advantage, he really said a lot of the competitive advantage is around this algorithmic diversity and innovation and why CUDA matters. He said if the world standardizes on transformers, on PyTorch, then it's less relevant for GPUs. In that environment, if you have a lot of standardization, right, then advantage goes to the custom ASICs. But I'll tell you this, you know, and I've had this conversation with a lot of people.
Brad Gerstner00:10:29
When I asked Jensen, I pushed him on, you know, custom ASICs. I was like, hey, you know, you've got, you know, accelerated inference coming from Meta with their MTIA chip. You know, you got Inferentia and Trainium, you know, coming. He's like, yeah, Brad, like they're, you know, they're my biggest partners. I actually share my three- to five-year roadmap with them. Yes, they're going to have these point solutions that are going to do these very specific tasks. But at the end of the day, the vast majority of the workloads in the world that are machine learning and AI infused are going to run on NVIDIA. And the more people I talk to, the more I'm convinced that that's the case, despite the fact that there'll be a lot of other winners, including Groq and Cerebras, etc.
Bill Gurley00:11:09
And they're acquiring companies, they're moving up the stack, they're trying to do more optimization at higher levels. So they want to extend obviously what CUDA is doing. Don't go to inference yet; that's a whole other story.
Sunny00:11:22
I'm actually on that bit about the deep integrations, right? Because, you know, really that's a playbook that I think Microsoft really had done well for a long time in enterprise software. And you really haven't seen that in hardware ever. You know, if you go back to say Cisco or the PC era or, you know, the cloud era, you didn't see that deep-level integration. Now Microsoft pulled it off with Azure. And when I heard him talking, all I could think about was, man, that was really smart. What he's done is he's gotten together, really understand what the use cases are and build an organization that deeply integrates into his customers and does it so well all the way up into his roadmap that he's much more deeply embedded than anyone else is.
Sunny00:12:04
When I heard that part, I kind of gave him a real tip of the hat on that one. But what did you, you know, Brad, what was your take on that?
Brad Gerstner00:12:12
You and I had this conversation after we first listened to it. And, you know, if you really telescope out, you know, he talks as a systems-level engineer. Even if you hear, people went to Harvard Business School and say, 'How can this guy possibly have 60 direct reports?' But how many direct reports does Elon have? These systems-level—and he said, 'I have situational awareness. I'm a prompt engineer to the best people in the world at these specific tasks.' I think when I look at this, the thing that I deeply underappreciated a year and a half ago about this company was the systems-level thinking, right, that he spent years thinking about how to embed this competitive advantage and how it really, it goes all the way from power, all the way through application.
Brad Gerstner00:13:01
And every day they're launching these new things to further embed themselves in the ecosystem. But I did hear from somebody over the last two days who, you know, Rene Haas, the CEO of Arm—Rene was also at our event, and he's a huge Jensen fan. He worked eight years at NVIDIA before becoming the CEO of Arm in 2013. And he said, 'Listen, nobody is going to assault the NVIDIA castle head-on. The mainframe of AI is entrenched and it's going to become a lot bigger, at least as far as the eye can see.' He said, 'However, if you think about where we're interacting with AI today, on these devices, on edge devices,' he's like, 'our installed base at Arm is 300 billion devices. And increasingly, a lot more of this compute can run closer to the edge.'
Brad Gerstner00:13:58
If you think about an orthogonal competitor, right? Again, if he has a deep competitive moat in the cloud, what's the orthogonal competitor? The orthogonal competitor peels off a lot of the AI on the edge. And I think Arm's incredibly well-positioned to do that. Clearly, NVIDIA's got Arm embedded now in a lot of their, you know, in a lot of their Grace Blackwell, etc. But that to me would be one area, like if you looked out and you said, where can their competitive advantage, you know, be challenged a little bit? I don't think they necessarily have the same level of advantage on the edge as they have in the cloud.
Bill Gurley00:14:34
You started the pod by saying, you know, everyone's heard this in the investment community. It's not a GPU company. It's a systems company. And I, in my brain, I think, had thought, oh, well, they've got four in a box instead of, you know, just one GPU or eight in a box. At the time I was listening to the podcast you did with Jensen, I was reading this NeoCloud playbook and anatomy post by Dylan Patel. Yes, that's a good one. He goes into extreme detail about the architecture of some of the larger systems, you know, like the one that xAI that we're going to talk about that was just deployed, which I think is 100,000 nodes or something like that. And it literally changed my opinion of exactly what's going on in the world and actually answered a lot of questions I had.
Bill Gurley00:15:25
But it appears to me that NVIDIA's competitive advantage is strongest where the size of the system is largest.
Brad Gerstner00:15:34⚠ 0.37
Hmm.
Bill Gurley00:15:34
Which is another way of saying what René said, it's flipping it on its head. It's not to say it's weak on the edge, but it's super powerful when you put a whole bunch of them together. That's when the networking piece thrives. That's where NVLink thrives. That's where CUDA really comes alive in the biggest systems that are out there. And some of the questions that answered for me was, one, why is demand so high at the high end and why are nodes available on the internet, single nodes available on the internet for at or below cost? And this starts to get at that, because you can do things with the large systems that you just can't do with a single node. And so those two things can be simultaneously true.
Bill Gurley00:16:19
Why was NVIDIA so interested in CoreWeave existing? Now I understand. Like, if the biggest systems are where the biggest competitive advantage is, you need as many of these big system companies as you can possibly have. And there may be, if that trajectory remains true, you could have an evolution where customer concentration increases for NVIDIA over time rather than going the other way. Depending on how, you know, if Sam's right that they're going to spend $100 billion or whatever on a single model, there's only so many places they're going to be able to afford that. But a lot of stuff started to make sense to me that didn't before. And I clearly underestimated the scale of what it meant to be a non-GPU company, to be a system company.
Bill Gurley00:17:08
This goes way, way up.
Brad Gerstner00:17:12
Yeah, and again, Bill, you touched on something that I think is really important here. And this is this question of whether their competitive moat is also as powerful in training as it is in inference, right? Because I think that there's a lot of doubt as to whether their competitive moat is as strong in inference. But, you know, let's just... You want to flip to that? Well, no, but I asked him if it was as strong. No, I know you did. He actually said it was greater. To me, when you think about that, in the first instance, I think it didn't make a lot of sense. But then when you really started thinking about it, he said there's a trail of software behind the infrastructure that's already out there that is CUDA compatible and can be amortized.
Brad Gerstner00:18:01
For all this inference. And so he, like, for example, referenced that OpenAI had just decommissioned Volta. So it's like this massive installed base. And when they improve their algorithms, when they improve their frameworks, when they improve their CUDA libraries, it's all backward-compatible. So Hopper gets better and Ampere gets better and Volta gets better. That combined with the fact that he said, everything in the world today is becoming highly machine-learned. Almost everything that we do, he said, almost every single application, Word, Excel, PowerPoint, Photoshop, AutoCAD, it all will run on these modern systems. Sunny, do you buy that? Do you buy that when people go to replace compute, they're going to replace it on these modern systems?
Speaker 100:18:46⚠ 0.41
Yeah.
Sunny00:18:46
So when I was listening to it, I was buying it. But then when I—he said one thing that kept resonating in my mind, which he said, inference is going to be a billion times larger than training. And if you kind of double-click into that, these old systems aren't going to be sufficient enough. If you're going to have that much more demand, that much more workload, which I think we all agree, then how is it that these old systems, which are being decommissioned from training, are going to be sufficient? So I think that's where that argument just didn't hold strong enough for me. If that grows as fast as he says it is, as fast as you guys have seen it in their numbers, then it's going to be a lot more net-new demand.
Sunny00:19:28
Inference-related deployments. And there, I don't think that argument holds on the transfer from older hardware to newer hardware.
Brad Gerstner00:19:38
Well, you said something pretty casually there, right? Let's underscore this, right? We were talking about the Strawberry and the O1 preview, and you said there's a whole new vector of scaling intelligence, inference time reasoning, right? That's not going to be single-shot. but it's going to be lots of agent to agent interactions, thinking time, as Noam Brown likes to say, right? And he said, as a consequence of that, inference is going to a hundred X, thousand X, a million X, maybe even a billion X. And that in and of itself, right, to me was, you know, kind of a wow moment. 40% of their revenues are already inference. And I said, over time, does your inference become a higher percentage of your revenue mix?
Brad Gerstner00:20:24
And he said, of course, right? But again, I think conventional wisdom is all around the size of clusters and the size of training. And if models don't keep getting bigger, then their relevance will dissipate. But he's basically saying every single workload is going to benefit from acceleration, right? It's going to be an inference workload. And the number of inference interactions is going to explode higher.
Sunny00:20:46
Yeah, one technical detail, which is you need bigger clusters if you're training bigger models. But if you're running bigger models, you don't need bigger clusters. It can be distributed, right? And so I think what we're going to see here is that the larger clusters will continue to get deployed. And as Bill said, they'll get deployed for folks, maybe a limited number of folks that need to deploy it for... $100 billion runs or even bigger than that. But you'll see inference clusters be large, but not as large as the training clusters and be a lot more distributed because you don't need it to be all in the same place. And I think that's what will be really interesting.
Bill Gurley00:21:24
It was interesting. He simplified it even more than you did there, Brad. He said, think about a human. How much time do you spend learning versus doing? And he used that analogy as to why this was gonna be so great. But I, in a little different way than Sunny, I thought the argument that the reason we're gonna be great at inference is because there's so much of our old stuff laying around wasn't super solid. In other words, what if some other company, Sunny's or some other one, decided to optimize inference? It wasn't an argument for optimization. It was an argument for cost advantage because it might be fully distributed or whatever. And, of course, if you had maybe poked him back on that, he might have had another answer about why for optimization.
Bill Gurley00:22:18
But there are clearly going to be people, whether it's other chip companies, some of these accelerator companies, there are going to be people working on inference optimization, which may include edge techniques. I think some of the accelerators may look like AI CDNs, if you will, and they're going to be buying stuff closer to the customer. So it... all TBD, but just the argument that you've got it left over didn't seem super solid to me.
Sunny00:22:46
And the three fastest companies in inference right now are not NVIDIA.
Brad Gerstner00:22:51
Right, so who are they, Sunny? We'll post the leaderboard.
Sunny00:22:54
Yeah, it's a combination of Groq, Cerebras, and SambaNova, right? Those are three companies that are not NVIDIA that are on the leaderboards of all the models that they run.
Bill Gurley00:23:05
You're talking about performance. Performance. Performance, yeah. Yeah.
Sunny00:23:09
And I would argue even price.
Bill Gurley00:23:11⚠ 0.24
Yeah.
Brad Gerstner00:23:12
And make the argument, why are they faster? Why are they cheaper in your mind? But yet, notwithstanding that fact, NVIDIA is going to do, let's call it $50 or $60 billion of inference this year. And these companies are, you know, still just getting started, right? Why is their inference business? It's just because of installed base, right?
Sunny00:23:33
Yeah, I think it's a combination of installed base. And I think it's because that inference market is growing so incredibly fast. I think if you're making this decision even 18 months ago, it would be a really difficult decision to buy any of those three companies because your primary workload was training. And the first part of this pod, we talked about how they have such a strong tie-in integration to getting training done properly. I think when it comes to inference, you can see all the non-NVIDIA folks can get the models up and running right away. There is no tie-in to CUDA that's required to go faster, that's required to get the models running, right? Obviously, none of the three companies run CUDA.
Sunny00:24:09
And so that moat doesn't exist around inference.
Bill Gurley00:24:12
Yeah, CUDA is less relevant in inference. That's another point worth making. But I wanted to say one other thing to what Soni just said. If you go back to the early internet days, and this is just an argument that optimization takes a while. All of the startups were running on Oracle and Sun. Every single one of them were running on Oracle and Sun. And five years later, they were all running on Linux and MySQL, like in five years. And it was literally, it went from 100% to 3%. And I'm not making that projection that that's going to happen here, but you did have a wholesale shift as the industry, you know, they went from developing and building it for the first time to optimizing, which are really two separate motions.
Brad Gerstner00:25:03
It seems to me, I pulled up this chart, right, that we shared, we made, Bill, way earlier this year for the pod, which showed the trillion dollars of new AI workloads expected over the next four to five years and the trillion dollars of effectively data center replacement. And I just wanted to get his updated kind of reaction or forecast now that he's had six more months to think about whether or not he thinks that's achievable. And what I heard him say was, yes, the data center replacement is going to look exactly like that. Of course, he's just making his best educated guess. But he seemed to suggest that the AI workloads could be even bigger, right? Like that once he saw Strawberry in 01, that he thought the, you know, the amount of compute that was going to be required to power this.
Brad Gerstner00:25:53
And, you know, the more people I talk to, the more I, you know, I get that same sense. There is this insatiable demand. So maybe we just touch on this. You know, he goes on CNBC and he says the demand is insane. And I kept trying to push on that. I was like, 'Yeah, but what about MTIA? What about custom inference? What about all these other factors? What if models stop getting so big?' I said, 'Well, has any of that changed the equation?' And he consistently pushed back and said, 'You still don't understand the amount of demand in the world because all compute is changing.' Right.
Sunny00:26:32
I thought he nuanced that answer, which was when you asked him that, he said, 'Look, if you have to replace some amount of infrastructure, whatever the number was, was really big and you're part of that and you're a CIO somewhere tasked with doing this. What are you going to do? What are you going to replace it with? It's accelerated compute.' And then immediately once you make that choice, because you're not going to traditional compute, then NVIDIA is your number one choice. So I thought he kind of tied that back together in that, like, are you really going to get yourself in trouble by having something else there? Or are you just going to go to NVIDIA? When he said it, I didn't want to say that, Bill, but it felt like the old IBM argument.
Bill Gurley00:27:14
Yeah, look, I mean, one thing, Brad, is this company's public. When a private company says, 'Oh, the demand's insane,' I immediately get skeptical. This company's doing $30 billion a quarter, growing 122%. Like, the demand is insane. We can see it.
Brad Gerstner00:27:33
There's no doubt about it. And part of that demand was a conversation about Elon and xAI and what they did. And I thought it was also just incredibly fascinating, right? I thought it was funny. I asked him a question about the dinner that he and Elon and Larry Ellison apparently had. And he's like, "You know, just because that dinner occurred and they ended up with 100,000 H100s don't necessarily connect the dots." But listen, he confirmed that his mind was blown by Elon, and he said he has an N-of-1 superhuman that could possibly pull off, that could energize a data center, that could liquid cool a data center. And he said what would take somebody else years to get permitted, to get energized, to get liquid cooled, to get stood up, that xAI did in 19 days.
Brad Gerstner00:28:28
You know, and you could just tell the immense respect that he had for Elon. It's clear, you know, he said it's the single largest coherent supercomputer in the world today, that it's going to get bigger. And if you believe that the future of AI is tied closely together with the systems engineering on the hardware side, you know, what hit me in that moment was that's a huge, huge advantage for Elon.
Sunny00:28:56
Yeah, I think he – I forgot the exact number, but like he talked about how many thousands of miles of cabling that were just in there as part of the task. Look, coming to it from a – doing a lot of that ourselves right now, building data centers, standing them up, racking and stacking our nodes, it's impressive. It's impressive to do something at that scale in 19 days. You know, it doesn't even include how quickly they built that data center. I think it's all happened, you know, within 2024. And so, um, that's part of the advantage. The interesting thing there is he didn't touch on it as much as what, when he talked about it, doing the integration with cloud service providers. What I'd love to kind of double-click into is because, you know, Elon is in a unique situation where he's obviously bought this cluster.
Sunny00:29:46
He has a ton of respect for NVIDIA, but he, you know, is building his own chip, building their own clusters with Tesla. So I wonder how much, you know, cross-correlation or information there is for them to be able to do that at scale. And, you know, you guys look at this, what have you kind of seen on their clusters?
Brad Gerstner00:30:05
I don't really have a lot of data on the non-NVIDIA clusters that they have. I'm sure Frieda on my team does. I just don't have it off the top of my head. If we have it, I'll pull a chart and I'll show it.
Bill Gurley00:30:15
Sunny, you said you now think the xAI cluster is the largest NVIDIA cluster alive today?
Sunny00:30:21
I'm saying because I believe Jensen said it in the pod that he said it's the largest supercomputer in the world.
Bill Gurley00:30:26
Yeah, I mean, I just want to spend 30 seconds on what you said, Brad, about Elon. I'm staring out my window at the Gigafactory in Austin that was also built in record time. Starlink's insane. When we were walking in Diablo, I just kept thinking, "You know who I'd love to reimagine this place? Elon, right?" Yeah. I don't – the world should study how he can do infrastructure fast because if that could be cloned, it would be so valuable. Not really relevant to this podcast, but worth noting. The other thing that I thought about on the Elon thing, and this also – where these pieces coming together in my mind about these large clusters and how important that was to NVIDIA. He got allocation, right? This is supposed to be like the hottest company, the hottest product backed, you know, backed up for years on demand.
Bill Gurley00:31:23
And he walks in and takes what equates, sounds, looks like about 10% of the quarter's availability. And, and, in my mind, I'm thinking that's because, hey, if there's another company that's going to develop these big ones, I'm going to let them to the front of the line. And that speaks to what's happening in Malaysia and the Middle East. And any one of these people that are going to get excited, he's going to spend time with them, put them at the front of the line.
Brad Gerstner00:31:52
You know, I'll tell you, you know, I pushed him on this. I said, you know, Elon's going to, you know, rumor is he's going to get another 100,000, you know, H200s, add them to this cluster. I said, "Are we already at the phase of two- and 300,000 cluster scale?" And he said, "Yes." And then I said, "And will we go to 500,000, a million?" And he's like, "Yes." Now, I think these things, Bill, are already being planned and built. And what he said is, beyond that, beyond that, he said you start bumping up against the limitations of base power. Like, can you find something that can be energized to power a single cluster? And he said, "We're going to have to develop distributed systems." And he said, "But just like with Megatron that we developed to allow to occur what is occurring today, we're working on the distributed stuff because we know we're going to have to decompose these clusters at some point in order to continue scaling them."
Bill Gurley00:32:54
You may also be running up against the, even for the Mag 7, the size of CapEx deployment where their CFOs start to talk in higher levels. Yes. For sure. Totally. And there's a super interesting article in The Information just now where it came out today where Sam Altman is questioning whether Microsoft's willing to put up the money and build a cluster. And that may have been kind of triggered by Elon's comments or Elon's willingness to do it at xAI. Yeah.
Sunny00:33:29
What I will say on the size of the models, we're going to push into this really interesting realm where obviously we can have bigger and bigger training clusters. That naturally imposes that the models are bigger and bigger. But what you can't do is you can't take a single model. You can train a model across a distributed site, and it may just take you a month longer because you have to move traffic around. And so instead of taking three months, it takes you four months. But you can't really run a model across a distributed site because that inference is a real-time thing. And so we do, you know, we're not pushing it there, but when you start to get to models that become way too big to run in single locations, that may be a problem that we want to be aware of.
Sunny00:34:07
And we want to keep on the top of our minds as well.
Brad Gerstner00:34:10
On this question of scaling our way to intelligence, one of the things I asked Noam Brown today in our fireside chat, he made very clear his perspective, although he's working on inference-time reasoning, which is a totally different vector and a breakthrough vector at OpenAI, which we ought to spend a little bit of time talking about. He said, "Now there are these two vectors that, again, are multiplicative in terms of the path to AGI." He's like, "Make no mistake about it. We're still seeing big advantages to scaling bigger models. We have the data. We have the synthetic data. We're going to build those bigger models. And we have an economic engine that can fund it. Don't forget, this company has over $4 billion in revenue, scaling probably, most people think, to $10 billion-plus in revenue over the course of the next year."
Brad Gerstner00:35:04
They just raised $6.5 billion. They got a $4 billion line of credit from Citigroup. So among the independent players, Bill—right. Like Microsoft can choose whether or not they're going to fund it. But I don't think it's a question of whether or not they're going to have the funding. At this point, they've achieved escape velocity. I think for a lot of the other independent players, there's a real question whether they have the economic model to continue to fund the activity. So they have to find a proxy because I don't think a lot of venture capitalists are going to write multi-billion-dollar checks into the players that haven't yet caught lightning in a bottle.
Speaker 100:35:36⚠ 0.40
Right.
Brad Gerstner00:35:37
That would be my guess. I mean, you know, I just think it's hard. You know, listen, at the end of the day, we're economic animals. You know, and I've said before, you know, if you look at the forward multiple, most of us underwrote to on OpenAI, it was about 15 times forward earnings. If ChatGPT wasn't doing what it was doing, if the revenue wasn't doing what it was doing, this would have been massively dilutive to the company. It would have been very hard to raise the money. I think if Mistral or all these other companies want to raise that money, I think it would be very difficult. But there's still a lot of money out there, so it's possible. You said 15 times earnings.
Bill Gurley00:36:15
I think you meant revenue.
Brad Gerstner00:36:16
Or 15 times revenue, for sure. Which I said, you know, when Google went public, it was about 13 or 14 times revenue and Meta was like 13 or 14 times revenue. So I do think we're on the precipice of a lot of this consolidation among the new entrants. What I think is so interesting about xAI is when I was pushing him on this model consolidation, pushing Jensen on it, he was like, listen, with Elon, you have somebody with the ambition, with the capability, with the know-how, with the money, with the brands, with the businesses. So I think a lot of times when we're talking about AI today, we oftentimes talk about OpenAI, but a lot of people quickly then go into all of the other model companies. I think xAI is often left out of the conversation.
Brad Gerstner00:37:01
And one of the things I took away from this conversation with Jensen is, again, if scaling these data centers is a key competitive advantage to winning in AI, you absolutely cannot count out xAI in this battle. They're certainly going to have to figure out something with the consumer that's going to have a flywheel like ChatGPT or something with the enterprise. But in terms of standing it up, building the model, having the compute, I think they're going to be one of the three or four in the game.
Bill Gurley00:37:34
You touched on maybe wanting to close out on the strawberry-like models. One thing we don't have exposure to, but we can guess at, is cost. And that chart that they showed when they released Strawberry, the x-axis was logarithmic. So the cost of a search with the new model preview model, it's probably costing them 20x or 30x what it does to do a normal chat GPT search.
Brad Gerstner00:38:09
Which I think is fractions of a penny.
Bill Gurley00:38:12
But figuring out which, and it also takes longer. So figuring out which problems it's acceptable—and Jensen gave a few examples—for it to take more time and cost more, and to get the cost-benefit right for that type of result, is something we're going to have to figure out, like which problems tilt to that place.
Brad Gerstner00:38:32
Right. And, you know, the one thing I feel good about there, and again, I'm speculating, I don't have information from OpenAI on this, but what we know is that the cost of inference has fallen by 90% over the course of last year. What we, you know, what Soni has told us and other people, you know, in the field have told us that inference is going to drop by another 90% over the course of the next, you know, period of months.
Bill Gurley00:38:56
If you're facing logarithmic curves, you're going to need that tech.
Brad Gerstner00:39:00
Right. And, you know, and here's what I also think happens, Bill, is in this chain of reasoning, you're going to build intelligence into the chain of reasoning. Right. So that, you know, you're going to optimize where you send these, you know, each of these inference interactions. You're going to batch them. You're going to take more time. Because it's just a time money tradeoff. At the end of the day, I also think that we're in the very earliest innings as to how we're going to think about pricing these models. So if we think about this in terms of systems one, systems two level thinking, systems one being what's the capital of France, you're going to be able to do that for fractions of a penny using pretty simple models on chat GPT.
Brad Gerstner00:39:48
When you want to do something more complex, if you're a scientist and you want to use O1 as your research partner, you may end up paying it by the hour. And relative to the cost of an actual research partner, it may be really cheap. So I think there are going to be consumption models for this. I think we haven't even scratched the surface to think about how that's going to be priced, but I totally agree with you that it's going to be priced very differently. Again, I think OpenAI... has suggested that the 01 full model may even be released yet this year. One of the things that I'm kind of waiting to see is, I think, listen, having known Noam Brown for quite a while now, he's an N of 1. And he wasn't the only one working on this, for sure, at OpenAI.
Brad Gerstner00:40:37
But listen, whether it was Pluribus or winning at the game of Diplomacy, he's been thinking about this for a decade. It was his major breakthrough on how to win the game of six-handed poker. And so he brought this to OpenAI. I think they have a real lead here, which leads me back to this question, Bill, you and I talk about all the time, which is memory and actions, right? And so I have to tell you this funny thing that occurred at our Investor Day. So I had Nikesh on stage and, you know, obviously Nikesh, you know, was instrumental at Google for a decade. And so I wanted to talk to him about both consumer AI as well as enterprise AI. And I asked him, I said, I want to make a wager with you. I knew, of course, he would take a bet.
Brad Gerstner00:41:24
And I said, I want to make a wager with you. Over/under, I'll set the line at two years until we have an agent that has memory and can take action. And the canonical use case, of course, that I used was that I could tell my agent, book me the Mercer Hotel next Tuesday in New York at the lowest price. And I said, over/under, you know, two years on getting that done. I said, I'll start 5,000 bucks. I'll take the under. He snap-calls me. He says, I'll take the over. And he said, but only if you 10x the bet. And of course, we're doing it for a good cause. So I had to call him because I can't not step up to a good cause. So we're taking the opposite sides of that trade. Now, what was interesting is over the course of the next couple of days, I asked some other friends who took the stage where they would come down on the same bet, right?
Brad Gerstner00:42:22
Our friend Stanley Tang took the under. A friend from Apple, who will remain nameless, kind of took the over. And then Noam Brown, who was there, pleaded the fifth. He says, I know the answer, so I can't say. And so I was kind of provocative. And I, you know, I texted Nikesh and I said, I think you better get your checkbook ready. You know, so coming back to that bill. Strawberry 01 is an incredible breakthrough, something that thinks of this whole new vector of intelligence. But it kind of makes us forget about the thing you and I focus so much on, which was memory and actions. And I think that we are on the real precipice of not only these models think, can spend more time thinking, not only can they give us less hallucinations and just scaled compute, but I also think...
Brad Gerstner00:43:17
I mean, you already see the makings of this. I mean, use these things today. They already remember quite a bit. So I think they're sliding this into the experience. But I think we're going to have the ability to take simple actions. And I think this metaphor that people had in their minds that they were going to have to build deep APIs and deep integrations to everybody, I don't think is the way this is going to play out. And let me just... What do you think is going to play out? Well, I mean, the Easter egg that I thought got dropped last week is they did this event on, you know, their voice API, right? And it's literally your GPT calling a human on the telephone and placing an order. So why the hell can't my GPT just call up the Mercer Hotel and say, "Brad Gerstner would like to make a reservation"?
Brad Gerstner00:44:02
Here's his credit card number and pass along the information.
Bill Gurley00:44:05
There is a reason for that. I mean, look, scrapers and form fillers have existed for how long, Sunny? 15 years? You could write an agent to go fill out and book at the Mercer Hotel 15 years ago. There's nothing impossible about that. It's the corner cases. And like the hallucination when your credit card gets charged 10 grand, like you just can't have failure and how you architect this so that there's not failure and there's trust. I'm sure you could demo this tomorrow. I have zero doubt you could demo it tomorrow. Could you provide it at scale in a trustworthy way where people are allocating their credit cards to it? That might take a little longer.
Brad Gerstner00:44:48
Okay, so over/under, Bill, on two years?
Bill Gurley00:44:51
I mean, I'm going to get you action either way. But what's the test? The demo? I think you'd do it today.
Brad Gerstner00:44:57
No, not the cheesy demo you just said. I'm talking about a release that allows me, you know, at scale to book a hotel.
Bill Gurley00:45:04
Where it's spending your credit card? And not just you, but everybody, full release?
Brad Gerstner00:45:09
Yeah, we'll call it a full release just because I know that's the only way I can tell you to take the bet.
Bill Gurley00:45:15
Hmm. Which today is October 8th, 2024.
Brad Gerstner00:45:20
I mean, Sunny, you already know what he's going to say. You'll take the over, right, Bill?
Bill Gurley00:45:25
Yeah. Yes.
Brad Gerstner00:45:26
Okay. So Bill's in the cautious camp. Sunny, where do you come down? Over/under on two years? No, don't start hedging, Bill. Don't start hedging. Go ahead, Sunny.
Bill Gurley00:45:32
I already said it. Demo today. It's 15 years ago.
Sunny00:45:36
Let me comment on what you're worried about, Bill. And I think people still are still working their way through it. You don't need a single agent right now to book The Mercer and deal with all the scraping stuff you're talking about. You can have a thousand agents working together. You can have one that's making sure that the credit card charge is not too big. You can have another one to make sure that the address is right. You can have another one checking in your calendar. And so all of that's free. So I'm on the under and Brad, I'll even go under one year.
Brad Gerstner00:46:03
Wow. Wow. Wow. So we got a little side action, you and I, Sonny. I'm not going to go under a year, but I think we could have limited releases in a year. But Sonny, you and I now have action with Bill. What do you want, Bill? A thousand bucks? To a good cause? Okay, $1,000 each to a good cause. And I'll just assume, Sonny, that we'll get action from Nikesh as well. And you know our friend Stanley Tang is definitely in the tank for some. So we're going to give some good money to a good cause. And listen, I think this is the trillion-dollar question. I know we're all focused on scaling models, and I know we're all focused on the compute layer, but what really transforms people's lives, what really disrupts 10 Blue Links—
Brad Gerstner00:46:47
What really disrupts the entire architecture of the app ecosystem is that when we have an intelligent assistant that we can interact with, that gets smarter over time, that has memory and can take actions. And when I see the combination of advanced voice mode, Voice-to-voice API, Strawberry 01, thinking combined with scaling intelligence. I just think this is going to go a lot faster than most of us think. Now, listen, they may pull on the reins, right? They may slow down the release schedule in order, you know, for a lot of business reasons. That's harder to predict. But I think the technology, I mean, even Noam said, I thought it was going to take us much, much longer to see the results that we have seen.
Brad Gerstner00:47:30
Can I hit on one other thing? We started the pod a little bit talking about it. I just want to get your impression, Bill. This idea that Jensen can scale the business two or three times with, you know, increasing the head count by, you know, 20 or 25%, right? We know that Meta's done that over the course of the last two years. And you and I've talked about, are we on the eve of just massive productivity boom and massive margin expansion like we've never seen before, right? Nikesh said we ought to be able to get 20 or 30%, you know, productivity gains out of everybody in the business.
Bill Gurley00:48:07
First of all, I think NVIDIA is a very special company and it's a company that's, even if it's a systems company, it's an IP company and the demand is growing at such a rate that they don't need more designers or more developer engineers to create incremental revenue that's happening on its own. And so their operating margins are record levels. For the majority of companies, I've always just held this belief that you evolve with your tools and the real, the real answer is the companies that don't deploy these things are gonna go out of business. And so I think margins get competed away in many, many cases. I think it's ridiculous to imagine, oh, every company goes to 60% operating margin.
Brad Gerstner00:48:59
I mean, listen, Delta Air Lines is going to do all of these things with AI and immediately because it's in a commodity market, it'll get competed away by Southwest and United. Bad industries remain bad industries.
Bill Gurley00:49:11
Yeah, yeah, yeah. So, so, but, but there might be some, you know, that, that figure it out. And, and I have another theory that, uh, that I always keep in mind, which is hyper-growth tends to delay what you learned in microeconomics class. You know, I, I remember when I was a PC analyst and there were five public PC companies, all growing 100%. And so, in, in moments of hyper-growth, you will have margins that may or may not be durable, and you'll have a number of participants in a market that may or may not be durable during periods of hyper-growth.
Brad Gerstner00:49:48
I have two more things on my mind, Sunny. Do you have any reactions to that? I mean, I just have to get to a couple of these topics. No.
Bill Gurley00:49:56
This is going to be a Lex Fridman-length podcast once you attend the interview.
Sunny00:50:03
No, look. I really, you know, been thinking a lot about Jensen's point in the pod about, you know, how much AI they're using internally for design, design verification, for all those pieces, right? And I think, you know, it's not 30%. I actually think sort of that's an underestimate. I think you're talking, you know, multiple hundreds of percent improvement in productivity gains. And the only issue is that not every company can grasp that that quickly. And so, you know, I think he was kind of holding some cards back at that point when he made that comment. And it really got me thinking about, like, how much are they doing there that they don't want everybody to know about? And you kind of see it now in the model development because they, you know, if you've noticed the last couple of weeks, they've put some models out there that are models trained on their own.
Sunny00:50:51
And they don't get as much noise as, you know, ones from Meta and, you know, the other players that are out there. But they're really doing a lot more than we think. And they, I think they have their arms around a lot of these very, very difficult problems.
Bill Gurley00:51:07
Brad, why did they put their own model out?
Brad Gerstner00:51:09
Well, it's related to this topic of open versus closed. So, Bill, you know, I hope you're proud of me. You know, I went back and I said, I have to ask this question. Right. And, you know, I thought Jensen, you know, I thought he gave a great answer, which is like, "Listen, we're going to have companies that for economic reasons, right, push the boundary toward AGI or whatever they're doing. And it makes sense to have a closed model that can be the best and they can monetize. But the world's not going to develop with just closed models. We're going to..." You know, he's like, "It's both open and closed." And, you know, he said, because open, he's like, "It's absolutely a condition required. It's going to be the vast majority of the models in the industry."
Brad Gerstner00:51:50
He's right. Now, if we didn't have open source, how would you have all these different fields in science, you know, be able to be activated on AI? He talked about Llama models exploding higher. And then with respect to his own open-source model, which I thought was really interesting, he said, "We focused on, like, something that a specific capability, and the capability that we were focused on is how to agentically use this model to make your model smarter, faster." So it's almost like a training, coaching model that he built. And so I think for them, it makes perfect sense why they may put that out into the world. But I also, you know, a lot of times the open versus closed debate, you know, gets hijacked into this conversation about safety and security.
Brad Gerstner00:52:38
And, you know, and I think he said, you know, "Listen, these two things are related, but they're not the same thing." You know, one of the things he commented on that is just, he said, there's so much coordination going on on the safety and security level. Like, we have so many agents and so much activity going on on making sure, you know, just look at what Meta's doing, you know, on this. He's like, "I think that's one thing that's under-celebrated, that even in the absence of any, you know, Platonic guardian sort of regulation, right?" Right. Without any top-down, you already have an extraordinary amount of effort going in by all of these companies into AI safety and security, that I thought was—I thought was a really important comment.
Brad Gerstner00:53:22
Thanks for jumping in, guys, kicking this one around. It was a special one, too.
Bill Gurley00:53:26
Congrats on having that opportunity. That's pretty—that's pretty unique.
Brad Gerstner00:53:31
And now we got a little wager. So, I mean, listen, I am so looking forward to, like, doing a live booking at the Mercer on the pod, right? And then, Sunny, we can just drop the money from the sky. We can just collect. We can just collect. Exactly. Exactly. Good to see you guys. We'll talk soon. All right. Peace. Take care. As a reminder to everybody, just our opinions, not investment advice.