Ep17. Welcome Jensen Huang | BG2 w/ Bill Gurley & Brad Gerstner

BG2Pod with Brad Gerstner and Bill Gurley · October 2024 · avg confidence 0.74
Watch original ↗
⚠ 18 span(s) flagged for spot-check (alignment confidence < 0.50) — verify these against the audio
  1. [01:20:02] Brad Gerstner (0.00) — Yeah.
  2. [01:14:22] Brad Gerstner (0.11) — Right.
  3. [00:34:00] Bill Gurley (0.20) — Right, right.
  4. [00:19:37] Bill Gurley (0.24) — Right.
  5. [00:20:38] Jensen Huang (0.26) — See, now we don't have to worry about the time.
  6. [00:24:12] Bill Gurley (0.28) — Right.
  7. [00:13:29] Brad Gerstner (0.29) — Right.
  8. [01:19:40] Brad Gerstner (0.29) — That's right.
  9. [00:00:26] Brad Gerstner (0.32) — Yes.
  10. [00:40:27] Brad Gerstner (0.32) — Right.
  11. [00:55:51] Jensen Huang (0.34) — Surely an alien intelligence.
  12. [01:11:53] Brad Gerstner (0.40) — Right.
  13. [00:26:00] Bill Gurley (0.42) — It's kind of craziness, right? There's a whole chain.
  14. [01:09:03] Brad Gerstner (0.43) — Right, right.
  15. [01:16:55] Jensen Huang (0.45) — Thank you.
  16. [00:13:14] Brad Gerstner (0.46) — Right.
  17. [00:41:41] Brad Gerstner (0.47) — Right.
  18. [00:20:25] Jensen Huang (0.49) — Don't worry about the time. Hey, guys. Hey, listen. Janine? Yeah. Look. Let's do it until,…
Jensen HuangBrad GerstnerBill Gurley
Jensen Huang00:00:00
What they achieved is singular, never been done before. Just to put in perspective, 100,000 GPUs, that's easily the fastest supercomputer on the planet. That's one cluster. A supercomputer that you would build would take, normally, three years to plan. Right. And then they deliver the equipment, and it takes one year to get it all working.
Brad Gerstner00:00:26⚠ 0.32
Yes.
Jensen Huang00:00:27
We're talking about 19 days.
Brad Gerstner00:00:41
Jensen, nice glasses. Hey, yeah, you too. It's great to be with you. Yeah, I got my ugly glasses on just like you. Come on, those aren't ugly. These are pretty good.
Jensen Huang00:00:51
Do you like the red ones better? There's something only your family could love.
Brad Gerstner00:00:55
Well, it's Friday, October 4th. We're at the NVIDIA headquarters just down the street from Altimeter. Welcome. Thank you. Thank you. And we have our investor meeting, our annual investor meeting on Monday, where we're going to debate all the consequences of AI, how fast we're scaling intelligence. And I couldn't think of anybody better, really, to kick it off with than you. I appreciate that. As both a shareholder, as a thought partner, kicking ideas back and forth. You really make us smarter. And we're just grateful for the friendship. So thanks for being here. Happy to be here. You know, this year, the theme is scaling intelligence to AGI. And it's pretty mind-boggling that when we did this two years ago, we did it on the Age of AI.
Brad Gerstner00:01:35
And that was two months before ChatGPT. And to think about all this change. So I thought we would kick it off with a thought experiment and maybe a prediction. Yeah. If I colloquially think of AGI as that personal assistant in my pocket, if I think of AGI as that colloquial assistant in my pocket, exactly, you know, that knows everything about me, that has perfect memory of me, that can communicate with me, that can book a hotel for me, or maybe book a doctor's appointment for me. When you look at the rate of change in the world today, when do you think we're going to have that personal assistant in our pocket?
Jensen Huang00:02:16
Soon, in some form.
Brad Gerstner00:02:17
Yeah.
Jensen Huang00:02:18
Yeah, soon, in some form. And that assistant will get better over time. That's the beauty of technology as we know it. And so I think in the beginning, it'll be quite useful, but not perfect. And then it gets more and more perfect over time, like all technology.
Brad Gerstner00:02:36
When we look at the rate of change, I think Elon has said the only thing that really matters is rate of change. It sure feels to us like the rate of change has accelerated dramatically, is the fastest rate of change we've ever seen on these questions. Because we've been around the rim, like you, on AI for a decade now—you even longer. Is this the fastest rate of change you've seen in your career?
Jensen Huang00:03:01
It is because we've reinvented computing. A lot of this is happening because we drove the marginal cost of computing down by 100,000x over the course of 10 years. Moore's Law would have been about 100x. Yeah. And we did it in several ways. We did it by, one, introducing accelerated computing, taking what is work that is not very effective on CPUs and put it on top of GPUs. We did it by inventing new numerical precisions. We did it by new architectures, inventing the Tensor Core. The way systems are formulated, NVLink added insanely fast memories, HBM, and scaling things up with NVLink and InfiniBand and working across the entire stack. Basically, everything that I described about how NVIDIA does things,
Jensen Huang00:03:58
That led to a super-Moore's Law rate of innovation. Now, the thing that's really amazing is that as a result of that, we went from human programming to machine learning. And the amazing thing about machine learning is that machine learning can learn pretty fast, as it turns out. And so as we reformulated the way we distribute computing, we did a lot of parallelism of all kinds, right? Tensor parallelism, pipeline parallelism, parallelism of all kinds. And we became good at inventing new algorithms on top of that and new training methods. And all of this invention is compounding on top of each other as a result, right? And back in the old days, if you look at the way Moore's Law was working, the software was static.
Jensen Huang00:04:51
It was precompiled, shrink-wrapped, put into a store. It was static. And the hardware underneath was growing at Moore's Law rate. Now we've got the whole stack growing, innovating across the whole stack. And so I think that that's the—now all of a sudden we're seeing scaling. That is extraordinary, of course. But we used to talk about pre-trained models and scaling at that level, and how we're doubling the model size and doubling, therefore appropriately doubling the data size. And as a result, the computing capacity necessary is increasing by a factor of four every year. Right. That was a big deal. Right. But now, we're seeing scaling with post-training and we're seeing scaling at inference.
Jensen Huang00:05:37
Isn't that right? And so people used to think that pre-training was hard and inference was easy. Now everything is hard, which is kind of sensible. You know, the idea that all of human thinking is one shot is kind of ridiculous. And so there must be a concept of fast thinking and slow thinking, and reasoning and reflection and iteration and simulation and all that. And that now, it's coming in.
Bill Gurley00:06:02
Yeah. I think to that point, one of the most misunderstood things about NVIDIA is how deep the true NVIDIA moat is. I think there's a notion out there that as soon as someone invents a new chip, a better chip, that they've won. But the truth is you've been spending the past decade building the full stack from the GPU to the CPU to the networking, and especially the software and libraries that enable applications to run on NVIDIA. So I think you spoke to that, but when you think about NVIDIA's moat today, do you think NVIDIA's moat today is greater or smaller than it was three to four years ago?
Jensen Huang00:06:47
Well, I appreciate you recognizing how computing has changed. In fact, the reason why people thought, and many still do, that you designed a better chip, it has more FLOPs, has more flips and flops and bits and bytes, you know what I'm saying? Yeah. And you see their keynote slides, and it's got all these flips and flops and bar charts and things like that. And that's all good. I mean, look, horsepower does matter. Yes. So these things fundamentally do matter. However, unfortunately, that's old thinking. It is old thinking in the sense that the software was some application running on Windows, and the software is static. Which means that the best way for you to improve the system is just making faster and faster chips.
Jensen Huang00:07:38
But we realized that machine learning is not human programming. Machine learning is not about just the software, it's about the entire data pipeline. It's about, in fact, the flywheel of machine learning is the most important thing. So how do you think about enabling this flywheel on the one hand and enabling data scientists and researchers to be productive in this flywheel? And that flywheel starts at the very, very beginning. A lot of people don't even realize that it takes AI to curate data to teach an AI. And that AI alone is pretty complicated.
Brad Gerstner00:08:20
And as that AI itself is improving, is it also accelerating, you know, again, when we think about the competitive advantage, right? It's combinatorial of all these systems.
Jensen Huang00:08:30
Exactly, exactly. And I was exactly going to lead to that. Because of smarter AIs to curate the data, we now even have synthetic data generation and all kinds of different ways of curating data, presenting data. And so before you even get the training, you've got massive amounts of data processing involved. And so people think about, oh, PyTorch, that's the beginning and the end of the world, and it was very important. But don't forget, before PyTorch, there's an amount of work. After PyTorch, there's an amount of work. And to think about the flywheel is really the way you ought to think. How do I think about this entire flywheel, and how do I design a computing system, a computing architecture, that helps you take this flywheel and be as effective as possible?
Jensen Huang00:09:16
It's not one slice of an application, training. Does that make sense? That's just one step, okay? Every step along that flywheel is hard. And so the first thing that you should do instead of thinking about, uh, how do I make Excel faster, how do I make, you know, Doom faster—that was kind of the old days, isn't that right?—now you have to think about how do I make this flywheel faster. And this flywheel has a whole bunch of different steps. There's nothing easy about machine learning, as you guys know. There's nothing easy about what OpenAI does or X does or Gemini and the team at DeepMind does. I mean, there's nothing easy about what they do. And so we decided, look, this is really what you ought to be thinking about.
Jensen Huang00:09:58
This is the entire process. You want to accelerate every part of that. You want to respect Amdahl's Law. Amdahl's Law would suggest, well, if this is 30% of the time and I accelerated that by a factor of three, I didn't really accelerate the entire process by that much. Right? Does that make sense? And you really want to create a system that accelerates every single step of that because only in doing the whole thing can you really materially improve that cycle time. And that flywheel, that rate of learning is really in the end what causes the exponential rise. And so what I'm trying to say is that our perspective about, you know, a company's perspective about what you're really doing manifests itself into the product.
Jensen Huang00:10:48
And notice, I've been talking about this flywheel. The entire cycle, yeah. That's right. And we accelerate everything. Right now, the main focus is video. A lot of people are focused on physical AI and video processing. Just imagine that front end. The terabytes per second of data that are coming into the system. Give me an example of a pipeline that is going to ingest all of that data, prepare it for training in the first place. So that entire thing is CUDA-accelerated.
Bill Gurley00:11:23
And people are only thinking about text models today. But the future is video models as well as using some of these text models like o1 to really process a lot of that data before we even get there.
Jensen Huang00:11:38
Yeah, yeah, yeah. Language models are going to be involved in everything. It took the industry enormous technology and effort to train a language model, to train these large language models. Now we're using a large language model in every single step of the way. It's pretty phenomenal.
Brad Gerstner00:11:56
I don't mean to be overly simplistic about this, but again, you know, we hear it all the time from investors, right? Yes, but... What about custom ASICs? Yes, but their competitive moat is going to be pierced by this. What I hear you saying is that in a combinatorial system, the advantage grows over time. So I heard you say that our advantage is greater today than it was three to four years ago because we're improving every component and that's combinatorial. Is that, you know, when you think about, for example, as a business case study, Intel, who had a dominant moat, a dominant position in the stack relative to where you are today. Perhaps just, again, boil it down a little bit. Compare, contrast your competitive advantage to maybe the competitive advantage they had at the peak of their cycle.
Jensen Huang00:12:49
Well, Intel is extraordinary because they were probably the first company that was incredibly good at manufacturing, process engineering, manufacturing. And that one click above manufacturing, which is building the chip.
Brad Gerstner00:13:14⚠ 0.46
Right.
Jensen Huang00:13:15
And designing the chip and architecting the chip in the x86 architecture and building faster and faster x86 chips, that was their brilliance. And they fused that with manufacturing.
Brad Gerstner00:13:29⚠ 0.29
Right.
Jensen Huang00:13:31
Our company is a little different in the sense that, and we recognize this, that in fact, parallel processing doesn't require every transistor to be excellent. Serial processing requires every transistor to be excellent. Parallel processing requires lots and lots of transistors to be more cost-effective. I'd rather have 10 times more transistors, 20% slower, than 10 times less transistors, 20% faster. Does that make sense? They would like the opposite. And so single-threaded performance, single-threaded processing, and parallel processing was very different. And so we observed that, in fact, our world is not about being better going down. We want to be very good, as good as we can be. But our world is really about much better going up.
Jensen Huang00:14:19
Parallel computing, parallel processing is hard because every single algorithm requires a different way of refactoring and re-architecting the algorithm for the architecture. What people don't realize is that you can have three different ISAs, CPU ISAs, they all have their own C compilers. You could take software and compile down to that ISA. That's not possible in accelerated computing. That's not possible in parallel computing. The company who comes up with the architecture has to come up with their own OpenGL. So we revolutionized deep learning because of our domain-specific library called cuDNN. Without cuDNN, nobody talks about cuDNN because it's one layer underneath PyTorch and TensorFlow and back in the old days, Caffe and Theano and now Triton.
Jensen Huang00:15:07
And there's a whole bunch of different frameworks. And so that domain-specific library, cuDNN, a domain-specific library called OptiX, we have a domain-specific library called cuQuantum, RAPIDS, the list of, you know, Aerial for—industry-specific—
Brad Gerstner00:15:26
algorithms that sit below that PyTorch layer that everybody's focused on. Like, I've heard oftentimes, "Well, you know, with LLMs..."
Jensen Huang00:15:33
If we didn't invent that, no application on top could work. Right? You guys understand what I'm saying? So the mathematics is really—what NVIDIA is really good at is algorithms. Right? That fusion between the science above, the architecture on the bottom, that's what we're really good at. Right? Yeah.
Bill Gurley00:15:52
There's all this attention now on inference, finally. But I remember two years ago, Brad and I had dinner with you and we asked you the question, "Do you think your moat will be as strong in inference as it is in training?" And I'm not sure— I said it would be greater. Yeah, yeah. And you touched upon a lot of these elements just now, just the composability between, or we don't know the total mix at one point. And to a customer, it's very important to be able to be flexible in between. That's right. But can you just touch upon, now that we're in this era of inference—
Jensen Huang00:16:34
Training is inferencing at scale. If you train well, it is very likely you'll inference well. If you built it on this architecture without any consideration, it will run on this architecture. You could still go and optimize it for other architectures, but at the very minimum, since it's already been built on NVIDIA, it will run on NVIDIA. Now, the other aspect, of course, is just kind of capital investment aspect, which is when you're training new models, you want your best new gear to be used for training, which leaves behind gear that you used yesterday. Well, that gear is perfect for inference. Mm-hmm. And so there's a trail of free gear. There's a trail of free infrastructure behind the new infrastructure that's CUDA-compatible.
Jensen Huang00:17:30
And so we're very disciplined about making sure that we're compatible throughout so that everything that we leave behind will continue to be excellent. Now we also put a lot of energy into continuously reinventing new algorithms so that when the time comes, the Hopper architecture is two, three, four times better than when they bought it. So that infrastructure continues to be really effective. And so all of the work that we do, improving new algorithms, new frameworks, notice, it helps every single installed base that we have. Hopper is better for it. Ampere is better for it. Even Volta is better for it. And I think Sam was just telling me that they had just decommissioned the Volta infrastructure that they have at OpenAI recently.
Jensen Huang00:18:19
So I think we leave behind this trail of installed base. Just like all computing, installed base matters. And NVIDIA is in every single cloud. We're on-prem and all the way out to the edge. And so the VILA vision-language model that's been created in the cloud works perfectly at the edge on robots. Without modification, it's all CUDA-compatible. And so I think this idea of architecture compatibility was important for large, it's no different for iPhones, no different for anything else. I think the installed base is really important for inference. But the thing that we really benefit from is because we're working on training these large language models and the new architectures of it, we're able to think about how do we create architectures that's excellent at inference someday when the time comes.
Jensen Huang00:19:15
And so we've been thinking about about iterative models for reasoning models and how do we create very interactive inference experiences for this personal agent of yours. You don't want to say something and have to go off and think about it for a while. You want it to interact with you quite quickly. So how do we create such a thing? And what came out of it was NVLink.
Bill Gurley00:19:37⚠ 0.24
Right.
Jensen Huang00:19:37
You know, NVLink so that we could take these systems that are excellent for training, but when you're done with it, the inference performance is exceptional. And so you want to optimize for this time-to-first-token. Right. And time-to-first-token is insanely hard to do, actually, because time-to-first-token requires a lot of bandwidth. But if your context is also rich... then you need a lot of FLOPs. And so you need an infinite amount of bandwidth, infinite amount of FLOPs at the same time in order to achieve just a few millisecond response time. And so that architecture is really hard to do. And we invented Grace, Blackwell, and NVLink for that.
Brad Gerstner00:20:21
Right. In the spirit of time, I have more questions about that.
Jensen Huang00:20:25⚠ 0.49
Don't worry about the time. Hey, guys. Hey, listen. Janine? Yeah. Look. Let's do it until, right?
Brad Gerstner00:20:31
Let's do it until, right. There you go. I love it. I love it. So, you know, I was at dinner with Andy Jassy earlier.
Jensen Huang00:20:38⚠ 0.26
See, now we don't have to worry about the time.
Brad Gerstner00:20:40
With Andy Jassy earlier this week. And Andy said, you know, we've got Trainium coming and Inferentia coming. And I think most people, again, view these as a problem for NVIDIA. But in the very next breath, he said, NVIDIA is a huge and important partner to us and will remain a huge and important partner for us. As far as I can see into the future, the world runs on NVIDIA. Right. So when you think about the custom ASICs that are being built that are going to go after targeted application, maybe the inference accelerator at Meta, maybe, you know, Trainium at Amazon, you know, or Google's TPUs. And then you think about the supply shortage that you have today. Do any of those things change that dynamic?
Brad Gerstner00:21:28
Right. Or are they complements to the systems that they're all buying from you?
Jensen Huang00:21:33
We're just doing different things. Yes. We're trying to accomplish different things. What NVIDIA is trying to do is build a computing platform for this new world, this machine learning world, this generative AI world, this agentic AI world. We're trying to create, as you know, what's just so deeply profound is after 60 years of computing, we reinvented the entire computing stack. Right. The way you write software from programming to machine learning, the way that you process software from CPUs to GPUs, the way that the applications from software to artificial intelligence, right? And so software tools to artificial intelligence. So every aspect of the computing stack and the technology stack has been changed.
Jensen Huang00:22:21
What we would like to do is to create a computing platform that's available everywhere. And this is really the complexity of what we do. The complexity of what we do is, if you think about what we do, we're building an entire AI infrastructure, and we think of it as one computer. I've said before, the data center is now the unit of computing. To me, when I think about a computer, I'm not thinking about that chip. I'm thinking about this thing. That's my mental model. And all the software and all the orchestration, all the machinery that's inside, that's my computer. And we're trying to build a new one every year. That's insane. Nobody has ever done that before. We're trying to build a brand new one every single year.
Jensen Huang00:23:04
And every single year, we deliver two or three times more performance. As a result, every single year, we reduce the cost by two or three times. Every single year, we improve the energy efficiency by two or three times, right? And so we ask our customers, don't buy everything at one time, buy a little every year, okay? And the reason for that, we want them to cost average into the future. All of it's architecturally compatible. So building that alone at the pace that we're doing is incredibly hard. Now, the double part, the double hard part is then we take that, all of that, and instead of selling it as an infrastructure or selling it as a service, we disaggregate all of it and we integrate it into GCP.
Jensen Huang00:23:49
We integrate it into AWS. We integrate it into Azure. We integrate it into X. Does that make sense? Yes. And so everybody's integration is different. We have to get all of our architectural libraries and all of our algorithms and all of our frameworks and integrate into theirs. We get our security system integrated into theirs. We get our networking integrated into theirs. Isn't that right?
Bill Gurley00:24:12⚠ 0.28
Right.
Jensen Huang00:24:12
Then we do basically 10 integrations. And we do this every single year. Now that is the miracle. That is the miracle. Why were you, I mean, it's madness.
Brad Gerstner00:24:24
It's madness that you're trying to do this every year. I'm going insane thinking about it. So what drove you to do it every year? And then related to that, you know, you're just back from Taipei and Korea and Japan, meeting with all your supply partners, yeah, who you have decade-long relationships with. How important are those relationships to, again, the combinatorial math that builds that competitive moat?
Jensen Huang00:24:51
Yeah, when you break it down systematically, the more you guys break it down, the more everybody breaks it down, the more amazed they are. Yes. And how is it possible that the entire ecosystem of electronics today is dedicated to working with us to build ultimately this cube of a computer integrated into all of these different ecosystems and the coordination is so seamless? Yeah. So there's obviously APIs and methodologies and business processes and design rules that we've propagated backwards and methodologies and architectures and APIs that we propagated forward. That have been hardened for decades. Hardened for decades, yeah, and also evolving as we go. But these APIs have to come together.
Jensen Huang00:25:43
Right, right. When the time comes, all these things in Taiwan, you know, all over the world being manufactured, they're going to land somewhere in Azure's data center. They're going to come together. They're click, click, click, click, click.
Bill Gurley00:25:54
Someone just calls it OpenAI API and it just works. That's right.
Jensen Huang00:25:58
Yeah, exactly.
Bill Gurley00:26:00⚠ 0.42
It's kind of craziness, right? There's a whole chain.
Jensen Huang00:26:02
And so that's what we invented. That's what we invented, this massive infrastructure of computing. The whole planet is working with us on it. It's integrated into everywhere. You could sell it through Dell. You could sell it through HPE. It's hosted in the cloud. It's all the way out at the edge. People use it in robotic systems now, humanoid robots. They're in self-driving cars. They're all architecturally compatible. Pretty kind of craziness.
Brad Gerstner00:26:32
It's craziness.
Jensen Huang00:26:33
Brad, I don't want you to leave the impression I didn't answer the question. In fact, I did. What I meant by that when relating to your ASIC is the way to think about we're just doing something different.
Brad Gerstner00:26:46
Yes.
Jensen Huang00:26:48
As a company, as a company... We want to be situationally aware, and I'm very situationally aware of everything around our company and our ecosystem. I'm aware of all the people doing alternative things and what they're doing. And sometimes it's adversarial to us, sometimes it's not. I'm super aware of it. But that doesn't change what the purpose of the company is. The singular purpose of the company is to build an architecture, a platform that could be everywhere. That is our goal. We're not trying to take any share from anybody. NVIDIA is a market maker, not share taker. If you look at our company slides, we don't show, not one day does this company talk about market share, not inside.
Jensen Huang00:27:40
All we're talking about is how do we create the next thing? What's the next problem we can solve in that flywheel? How can we do a better job for people? How do we take that flywheel that used to take about a year? How do we crank it down to about a month? Yes. You know, what's the speed of light of that? Isn't that right? And so we're thinking about all these different things, but the one thing we're not, we're not too, we're situationally aware of everything, but we're certain that what our mission is, is very singular. Yeah. The only question is whether that mission is necessary. Does that make sense? And all companies, all great companies ought to have that at its core. It's about what are you doing?
Jensen Huang00:28:21
The only question, is it necessary? Is it valuable? Is it impactful? Does it help people? And I am certain that you're a developer, you're a generative AI startup, and you're about to decide how to become a company. The one choice that you don't have to make is which one of the A6 do I support? If you just support a CUDA, you know you could go everywhere. You could always change your mind later. But we're the on-ramp to the world of AI, isn't that right? Once you decide to come onto our platform, the other decisions you could defer. You could always build your own ASIC later. We're not against that. We're not offended by any of that. When we work with all the GCPs, the GCPs Azure, we present our roadmap to them years in advance.
Jensen Huang00:29:10
They don't present their ASIC roadmap to us and it doesn't ever offend us. Does that make sense? If you have a sole purpose and your purpose is meaningful, and your mission is dear to you and is dear to everybody else, then you could be transparent. Notice my roadmap is transparent at GTC. My roadmap goes way deeper to our friends at Azure and AWS and others. We have no trouble doing any of that, even as they're building their own ASIC.
Brad Gerstner00:29:40
I think when people observe the business, you said recently that the demand for Blackwell is insane. You said one of the hardest parts of your job is the emotional toll of saying no to people in a world that has a shortage of the compute that you can produce and have on offer. But critics say this is just a moment in time, right? They say, 'This is just like Cisco in 2000. We're overbuilding fiber. It's going to be boom and bust.' You know, I think about the start of '23 when we were having dinner. The forecast for NVIDIA at that dinner in January of '23 was that you would do 26 billion of revenue for the year 2023. You did 60 billion, right?
Jensen Huang00:30:29
The 25 people— let's just let the truth be known. That is the single greatest failure of forecasting the world has ever seen. Right, right, right. Can we all at least admit that? To me— that was my takeaway. I just got—
Brad Gerstner00:30:44
And that was – we got so excited in November of 22 because we had folks like Mustafa from Inflection and Noah from Character coming in our office talking about investing in their companies. And they said, well, if you can't – pencil out investing in our companies than buy Nvidia because everybody in the world is trying to get Nvidia chips to build these applications that are gonna change the world. And of course the Cambrian moment occurred with ChatGPT and not withstanding that fact, these 25 analysts were so focused on the crypto winter that they couldn't get their head around an imagination of what was happening in the world, okay? So it ended up being way bigger. You say, in very plain English, the demand is insane for Blackwell, that it's going to be that way for as far as you can see.
Brad Gerstner00:31:33
Of course, the future is unknown and unknowable. But why are the critics so wrong that this isn't going to be the Cisco-like situation of overbuilding in 2000?
Jensen Huang00:31:45
Yeah. The best way to think about the future is to reason about it from first principles. Correct. Okay, so the question is, what are the first principles of what we're doing? Number one, what are we doing? What are we doing? The first thing that we are doing is we are reinventing computing. Did we not? We just said that. The way that computing will be done in the future will be highly machine-learned. Yes. Highly machine-learned. Almost everything that we do, almost every single application—Word, Excel, PowerPoint, Photoshop, Premiere, AutoCAD—you give me your favorite application that was all hand-engineered, I promise you it will be highly machine-learned in the future. Isn't that right? And so all these tools will be...
Jensen Huang00:32:34
And on top of that, you're going to have machines, agents, that help you use them. Right. Okay. And so we know this for a fact at this point, right? Isn't that right? We've reinvented computing. We're not going back. The entire computing technology stack has been reinvented. Okay. So now that we've done that, we said that software is going to be different. What software can write is going to be different. How we use software will be different. So let's now acknowledge that. Those are my ground truths now.
Bill Gurley00:33:01
Yes.
Jensen Huang00:33:02
Now the question, therefore, is what happens? And so let's go back and let's just take a look at how computing was done in the past. So we have a trillion dollars' worth of computers in the past. We look at it, just open the door, look at the data center, and you look at it and say, 'Are those the computers you want doing that, doing that future?' And the answer is no. You've got all these CPUs back there. We know what they can do and what they can't do. And we just know that we have a trillion dollars' worth of data centers that we have to modernize. And so right now as we speak, if we were to have a trajectory over the next four or five years to modernize that old stuff, that's not unreasonable. Sensible.
Brad Gerstner00:33:37
And you're having those conversations with the people who have to modernize it.
Jensen Huang00:33:41
And they're modernizing it on GPU. That's right. Well, let's make another test. You have $50 billion of CapEx you'd like to spend. Option A, option B: build CapEx for the future or build CapEx like the past. Now, you already have the CapEx of the past.
Bill Gurley00:34:00⚠ 0.20
Right, right.
Jensen Huang00:34:01
It's sitting right there. It's not getting much better anyways. Moore's Law has largely ended. And so why rebuild that? Let's just take $50 billion, put it into generative AI. Isn't that right? And so now your company just got better. Right. Now, how much of that 50 billion would you put in? Well, I would put in 100% of the 50 billion because I've already got four years of infrastructure behind me. That's of the past. And so now I just reasoned about it from the perspective of somebody thinking about it from first principles. And that's what they're doing. Smart people are doing smart things. Now, the second part is this. So now we have a trillion dollars' worth of capacity to go build, right?
Jensen Huang00:34:36
Trillion dollars' worth of infrastructure. What about, call it, $150 billion into it? So we have a trillion dollars in infrastructure to go build over the next four or five years. Well, the second thing that we observe is that the way that software is written is different, but how software is going to be used is different. In the future, we're going to have agents. Isn't that right? We're going to have digital employees in our company. In your inbox, you have all these little dots and these little faces. In the future, there's going to be little icons of AIs, isn't that right? I'm going to be sending them, I'm going to be—I'm no longer going to program computers with C++, I'm going to program AIs with prompting, isn't that right?
Jensen Huang00:35:19
Now, this is no different than me talking to my—you know, this morning, I wrote a bunch of emails before I came here. I was prompting my teams. Of course. Right? Yeah. And I would describe the context. I would describe the fundamental constraints that I know of. And I would describe the mission for them. I would leave it sufficiently—I would be sufficiently directional so that they understand what I need. And I want to be clear about what the outcome should be, as clear as I can be. But I leave enough ambiguous space, a creativity space, so they can surprise me. Isn't that right? Absolutely. It's no different than how I prompt an AI today. Yeah. It's exactly how I prompt an AI. And so what's going to happen is on top of this infrastructure of IT that we're going to modernize, there's going to be a new infrastructure.
Jensen Huang00:36:03
This new infrastructure is going to be AI factories that operate these digital humans. And they're going to be running all the time, 24/7. Mm-hmm. Right. We're going to have them for all of our companies all over the world. We're going to have them in factories. We're going to have them in autonomous systems. Isn't that right? So there's a whole layer of computing fabric, a whole layer of what I call AI factories that the world has to make that doesn't exist today at all. So the question is, how big is that? Right. Unknowable at the moment, probably a few trillion dollars. Unknowable at the moment, but as we're sitting here building into—the beautiful thing is, the architecture for this, modernizing this new data center, and the architecture for the AI factory is the same.
Jensen Huang00:36:49
That's the nice thing.
Brad Gerstner00:36:50
And you made this clear. You've got a trillion of old stuff you've got to modernize. You at least have a trillion of new AI workloads coming on. Give or take, you'll do $125 billion in revenue this year. At one point, somebody told you the company would never be worth more than a billion. As you sit here today, is there any reason, if you're only $125 billion out of a multi-trillion TAM, that you're not going to have 2x the revenue, 3x the revenue in the future that you have today? Is there any reason your revenue doesn't—
Jensen Huang00:37:23
No. Yeah. As you know, it's not about everything. Companies are only limited by the size of the fish pond. A goldfish can only be so big. And so the question is, what is our fish pond? What is our pond? And that requires a little imagination. And this is the reason why market makers think about that future, creating that new fish pond. It's hard to figure this out looking backwards and try to take share. Right. You know, share takers can only be so big. For sure. Market makers can be quite large. For sure. Yeah, and so, you know, I, I think, I think the, the good fortune that our company has is that since the very beginning of our company, we had to invent the market for us to go swim in. That market—and people don't realize this back then anymore, but, you know, we were at the, at the ground zero of creating the 3D gaming PC market. Right. Right. We largely invented this market,
Jensen Huang00:38:26
and all the ecosystem and all the graphics card ecosystem. We invented all that. And so the need to invent a new market to go serve it later is something that's very comfortable for us.
Brad Gerstner00:38:39
Exactly. And speaking to somebody who's invented a new market, let's shift gears a little bit to models and OpenAI. OpenAI raised, as you know, $6.5 billion this week. Yeah, at like $150 billion valuation. We both participated. Yeah, really happy for them. Really, really happy they came together. Right. Yeah, they did a great—Sam and the team did a great job. Reports are that they'll do $5 billion-ish of revenue or run rate revenue this year, maybe going to $10 billion next year. If you look at the business today, it's about twice the revenue as Google was at the time of its IPO. They have 250 million—is that right?—250 million weekly average users, which we estimate is twice the amount Google had at the time of its IPO.
Brad Gerstner00:39:27
And if you look at the multiple of the business, if you believe 10 billion next year, it's about 15 times the forward revenue, which is about the multiple of Google and Meta at the time of their IPO. When you think about a company that had zero revenue, zero weekly average users 22 months ago.
Jensen Huang00:39:44
Brad has an incredible command of history.
Brad Gerstner00:39:47
When you think about that, talk to us about the importance of OpenAI as a partner to you and OpenAI as a force in kind of driving forward, you know, kind of public awareness and usage around AI.
Jensen Huang00:40:04
Well, this is one of the most consequential companies of our time. A pure-play AI company pursuing the vision of AGI and whatever its definition.
Brad Gerstner00:40:27⚠ 0.32
Right.
Jensen Huang00:40:28
I almost don't think it matters fully what the definition is, nor do I really believe that the timing matters. The one thing that I know is that AI is going to have a roadmap of capabilities over time, and that roadmap of capabilities over time is going to be quite spectacular. And along the way, long before it even gets to it, anybody's definition of AGI, we're going to put it to great use. All you have to do is, right now as we speak, go talk to digital biologists, climate tech researchers, material researchers, physical sciences, astrophysicists, quantum chemists. You go ask video game designers, manufacturing engineers, roboticists. Pick your favorite, whatever industry you want to go pick.
Jensen Huang00:41:33
And you go deep in there and you talk to the people that matter and you ask them, 'Has AI revolutionized the way you work?'
Brad Gerstner00:41:41⚠ 0.47
Right.
Jensen Huang00:41:41
And you take those data points and you come back and you then get to ask yourself, how skeptical do you want to be? Right. Right. Because they're not talking about AI as a conceptual benefit, right, someday; they're talking about using AI right now. Right now. Agtech, material tech, climate tech—you pick your tech, you pick your field of science. They are advancing; AI is helping them advance their work right now as we speak. Every single industry, every single company, every university—unbelievable, isn't that right? Right. It is absolutely going to somehow transform business. We know that.
Bill Gurley00:42:28
Right.
Jensen Huang00:42:29
I mean, it's so tangible you could touch it. It's happening today. It's happening today. It's happening today. It's completely incredible. And I love their velocity and their singular purpose of advancing this field. And so really, really consequential.
Brad Gerstner00:42:58
And they build an economic engine that can finance the next-generation, you know, frontier of models, right? And I think there's a growing consensus in Silicon Valley that the whole model layer is commoditizing. Llama is making it very cheap for lots of people to build models. And so early on here, we had a lot of model companies, you know, Character and Inflection and Cohere and Mistral and go through the list. And a lot of people question whether or not those companies can build the escape velocity and the economic engine that can continue funding those next generation. My own sense is that there's going to be—that's why you're seeing the consolidation, right? OpenAI clearly has hit that escape velocity.
Brad Gerstner00:43:43
They can fund their own future. It's not clear to me that many of these other companies can. Is that a fair kind of review of the state of things in the model layer that we're going to have this consolidation like we have in lots of other markets to market leaders who can afford, who have an economic engine and application that allows them to continue to invest?
Jensen Huang00:44:07
First of all, there's a fundamental difference between a model and artificial intelligence. A model is an essential ingredient for artificial intelligence. It's necessary, but not sufficient. And artificial intelligence is a capability, but for what? Then what's the application? The artificial intelligence for self-driving cars is related to the artificial intelligence for humanoid robots, but it's not the same, which is related to the artificial intelligence for a chatbot, but not the same. Correct. And so you have to understand the taxonomy of the stack. And at every layer of the stack, there will be opportunities, but not infinite opportunities for everybody at every single layer of the stack.
Jensen Huang00:44:55
Now, I just said something. All you have to do is replace the word model with GPU. In fact, this was the great observation of our company 32 years ago. that there's a fundamental difference between GPU, graphics chip or GPU, versus accelerated computing. And accelerated computing is a different thing than the work that we do with AI infrastructure. It's related, but it's not exactly the same. It's built on top of each other, it's not exactly the same. And each one of these layers of abstraction requires fundamental different skills. Somebody who's really, really good at building GPUs have no clue how to be an accelerated computing company. There are a whole lot of people who build GPUs. And I don't know which one came.
Jensen Huang00:45:45
We invented the GPU, but you know that we're not the only company that makes GPUs today. Correct. And so there are GPUs everywhere. But they're not accelerated computing companies. And there are a lot of people who, you know, they're accelerators—accelerators that do application acceleration. But that's different than an accelerated computing company. And so, for example, a very specialized AI application could be a very successful thing. Meta's MTIA. That's right. Right. But it might not be the type of company that had broad reach and broad capabilities. And so you've got to decide where you want to be. There's opportunities probably in all these different areas, but like building companies, you have to be mindful of the shifting of the ecosystem and what gets commoditized over time, recognizing what's a feature versus a product versus a company.
Jensen Huang00:46:38
For sure. Okay. I just went through. Okay. Yeah. There's a lot of different ways you can think about this.
Brad Gerstner00:46:44
Of course, there's one new entrant that has the money, the smarts, the ambition. That's xAI. Yeah. Right. And well, there are reports out there that you and Larry and Elon had dinner. They talked you out of 100,000 H100s. They went to Memphis and built a large, coherent supercluster in a matter of months.
Jensen Huang00:47:06
You know, so first, three points don't make a line. Yes, I had dinner with them. Causality.
Brad Gerstner00:47:17
What do you think about their ability to stand up that supercluster? And there's talk out there that they want another 100,000 H200s, right, to expand the size of that supercluster. You know, first talk to us a little bit about xAI and their ambitions and what they've achieved. But also, are we already at the age of clusters of 200,000 and 300,000 GPUs?
Jensen Huang00:47:42
The answer is yes. And then, first of all, acknowledgment of achievement where it's deserved. From the moment of concept to a data center that's ready for NVIDIA to have our gear there, to the moment that we powered it on, had it all hooked up, and it did its first training. Yeah. Okay? Correct. So that first part, just building a massive factory, liquid-cooled, energized, permitted in the short time that was done, I mean, that is like superhuman. And as far as I know, there's only one person in the world who could do that. I mean, Elon is singular in this. Understanding of engineering and construction and large systems and marshaling resources. Incredible. Yeah, it's unbelievable. And of course, then his engineering team is extraordinary.
Jensen Huang00:48:52
I mean, the software team is great. The networking team is great. The infrastructure team is great. Elon understands this deeply. And from the moment that we decided to go, the planning with our engineering team, our networking team, our infrastructure computing team, the software team, all of the preparation in advance. Then all of the infrastructure, all of the logistics and the amount of technology and equipment that came in on that day, NVIDIA's infrastructure and computing infrastructure and all that technology, to training, 19 days. Did anybody sleep 24/7? No question that nobody slept. But first of all... 19 days is incredible, but it's also kind of nice to just take a step back and just, do you know how many days 19 days is?
Jensen Huang00:49:46
It's just a couple of weeks. And the mountain of technology, if you were to see it, is unbelievable. All of the wiring and the networking and, you know, networking NVIDIA gear is very different than networking hyperscale data centers, okay? The number of wires that goes in one node, the back of a computer is all wires. Just getting this mountain of technology integrated and all the software, incredible. So I think what Elon and the xAI team did, and I'm really appreciative that he acknowledges the engineering work that we did with him and the planning work and all that stuff. But what they achieved is singular, never been done before. Just to put it in perspective, 100,000 GPUs, that's easily the fastest supercomputer on the planet as one cluster.
Jensen Huang00:50:37
A supercomputer that you would build would take normally three years to plan. Right. And then they deliver the equipment and it takes one year to get it all working.
Bill Gurley00:50:51
Yes.
Jensen Huang00:50:52
We're talking about 19 days.
Bill Gurley00:50:53
Wow. Well, it's to the credit of the NVIDIA platform, right? That the whole process is hardened.
Jensen Huang00:50:59
That's right. Yeah. Everything's already working. And of course, there's a whole bunch of xAI algorithms and xAI framework and xAI stack and things like that. And we had a ton of integration we had to do. But the planning of it was extraordinary. Just pre-planning of it to, you know.
Brad Gerstner00:51:15
An N-of-1 is right. Elon is an N-of-1. But you answered that question by starting off saying, yes, 200,000 to 300,000 GPU clusters are here, right? Does that scale to 500,000? Does it scale to a million? And does the demand for your products depend on it scaling to millions? Yes.
Jensen Huang00:51:43
The last part is no. My sense is that distributed training will have to work. And my sense is that distributed computing will be invented. And some form of federated learning and asynchronous distributed computing is going to be discovered. And I'm very enthusiastic and very optimistic about that. Of course, the thing to realize is that the scaling law used to be about pre-training. Now we've gone to multimodality. We've gone to synthetic data generation. Post-training has now scaled up incredibly. Synthetic data generation, reward systems, reinforcement learning-based. And then now inference scaling has gone through the roof. Right. The idea that a model, before it answers your question, had already done internal inference.
Jensen Huang00:52:44
Incredible. 10,000 times, it's probably not unreasonable. And it's probably done tree search. It's probably done reinforcement learning on that. It's probably done some simulations. It's surely done a lot of reflection. It probably looked up some data. It looked up some information. Isn't that right? And so its context is probably fairly large. I mean, this type of intelligence is, well, that's what we do. Right. That's what we do, isn't that right? And so the ability, this scaling, if you did that math and you compound that with 4x per year on model size and computing size, and then on the other hand, demand continues to grow in usage. Do we think that we need millions of GPUs? No doubt. Yeah, that is a certainty now.
Jensen Huang00:53:36
And so the question is, how do we architect it from a data center perspective? And that has a lot to do with, you know, are there data centers that are gigawatts at a time or are they 250 megawatts at a time? And my sense is that, you know, you're going to get both.
Bill Gurley00:53:51
I think analysts always focus on the current architectural bet. But I think one of the biggest takeaways from this conversation is that you're thinking about the entire ecosystem and many years out. So the idea that because NVIDIA is just scaling up or scaling out, it's to meet the future. It's not such that you're only dependent on a world where there's a 500,000 or a million GPU cluster. By the time there's distributed training, you'll have written the software to enable that.
Jensen Huang00:54:27
That's right. Remember, without Megatron that we developed some seven years ago now, the scaling of these large training jobs wouldn't have happened. So we invented Megatron, we invented NCCL, GPUDirect, all of the work that we did with RDMA. That made it possible to easily do pipeline parallelism, you know, right? And so, you know, all the model parallelism that's being done, you know, all the breaking of the distributed training and all the batching and all that, all of that stuff is because we did the early work. And now we're doing the early work for the future generation.
Brad Gerstner00:55:07
So let's talk about Strawberry in 01. I wanna be respectful of your time.
Jensen Huang00:55:12
I got all the time in the world, guys.
Brad Gerstner00:55:14
Well, you're very generous. Yeah, I got all the time in the world. But first, I think it's cool that they named 01 after the 01 visa. which is about recruiting the world's best and brightest and bringing them to the United States. It's something I know we're both deeply passionate about. So I love the idea that building a model that thinks and that takes us to the next level of scaling intelligence is an homage to the fact that it's these people who come to the United States by way of immigration that have made us what we are, bring their collective intelligence to the United States.
Jensen Huang00:55:51⚠ 0.34
Surely an alien intelligence.
Brad Gerstner00:55:53
Certainly. It was spearheaded by our friend, Noam Brown, of course. He worked on Pluribus and Cicero when he was at Meta. How big a deal is inference-time reasoning as a totally new vector of scaling intelligence, separate and distinct from just building larger models?
Jensen Huang00:56:12
It's a huge deal. It's a huge deal. I think a lot of intelligence can't be done a priori, right? And a lot of computing, even a lot of computing can't be reordered. Out-of-order execution can't be done a priori. And so a lot of things can only be done in runtime. And so whether you think about it from a computer science perspective or you think about it from an intelligence perspective, too much of it requires context: the circumstance, the type of answer you're looking for. Sometimes just a quick answer is good enough. Depends on the consequential impact of the answer, depending on the nature of the usage of that answer. So some answers, please take a night. Some answers, take a week. Is that right?
Jensen Huang00:57:13
So I could totally imagine me sending off a prompt to my AI and telling it, you know, "Think about it for a night. Think about it overnight. Don't tell me right away. I want you to think about it all night. And then come back and tell me tomorrow what's your best answer and reason about it for me." And so I think the quality, the segmentation of intelligence now from a product perspective, there's going to be one-shot versions of it, right? For sure. Yeah. And then there will be some that take five minutes, you know.
Brad Gerstner00:57:46
And the intelligence layer that routes those questions to the right model for the right use case. I mean, we were using Advanced Voice Mode and o1-preview last night. I was coaching my son for his AP History test. And it was like having the world's best AP History teacher sitting right next to you thinking about these questions. It was truly extraordinary. Again,
Jensen Huang00:58:11
My tutor is an AI today, right?
Brad Gerstner00:58:12
I'm serious. Right, of course. They're here today. Yeah. Which comes back to this, you know, over 40% of your revenue today is inference. But inference is about ready because of chain of reasoning.
Jensen Huang00:58:24
Yeah.
Brad Gerstner00:58:24
Right? It's about ready.
Jensen Huang00:58:25
It's about to go up by a billion times, right?
Brad Gerstner00:58:27
By a million X, by a billion X.
Jensen Huang00:58:30
That's right. That's the part that most people have, you know, haven't completely internalized. This is that industry we were talking about, right? This is the Industrial Revolution, right?
Brad Gerstner00:58:41
That's the production of intelligence.
Jensen Huang00:58:43
That's right.
Brad Gerstner00:58:44
Right? Yeah.
Jensen Huang00:58:45
It's going to go up a billion times.
Brad Gerstner00:58:47
Right. And so everybody's so hyper-focused on NVIDIA as kind of like doing training on bigger models. Yeah. Right? Isn't it the case that your revenue, if it's 50-50 today, you're going to do way
Jensen Huang00:59:01
more inference in the future, yeah, right, than—I mean, training will always be important, but just the growth of inference is going to be way larger than, we hope, than training, we hope. It's almost impossible to conceive otherwise. Yeah, we hope that's right. That's right. Yeah, I mean, it's—it's good—it's good to go to school, yes, but the goal is so that you can be productive in society later. And so it's good that we train these models, but the goal is to inference them, you know? Yeah.
Brad Gerstner00:59:24
Are you already using chain of reasoning and tools like o1 in your own business to improve your own business?
Jensen Huang00:59:33
Yeah. Our cybersecurity system today can't run without our own agents. We have agents helping to design chips. Hopper wouldn't be possible. Blackwell would be possible. Ruben, don't even think about it. We have digital. We have AI chip designers, AI software engineers, AI verification engineers. And we build them all inside because we have the ability and we rather use the opportunity to explore the technology ourselves.
Brad Gerstner01:00:01
When I walked into the building today, somebody came up to me and said, "Ask Jensen about the culture. It's all about the culture." I look at the business. We talk a lot about fitness and efficiency, flat organizations that can execute quickly, smaller teams. You know, NVIDIA is in a league of its own, really, at about $4 million of revenue per employee, about $2 million of profits or free cash flow per employee. You've built a culture of efficiency that really has unleashed creativity and innovation and ownership and responsibility. You've broken the mold on kind of functional management. Everybody likes to talk about all of your direct reports. Is the leveraging of AI the thing that's going to continue to allow you to be hyper-creative while at the same time being efficient?
Jensen Huang01:00:55
No question. I'm hoping that someday—NVIDIA has 32,000 employees today, and we have 4,000 families in Israel. I hope they're well; I'm thinking of you guys. And I'm hoping that NVIDIA someday will be a 50,000-employee company with 100 million AI assistants. "Wow." And they're in every single group. "Right." We'll have a whole directory of AIs that are just generally good at doing things. We'll also have—our inbox is going to be full of directories of AIs that we work with that we know are really, really good, specialized at our skill. And so AIs will recruit other AIs to solve problems. AIs will be in Slack channels with each other. "And with humans." Right, and with humans. And so we'll just be one large employee base, if you will.
Jensen Huang01:01:54
Some of them are digital and AI. Some of them are biological. And I'm hoping some of them are even in mechatronics.
Brad Gerstner01:02:00
I think from a business perspective, it's something that's greatly misunderstood. You just described a company that's producing the output of a company with 150,000 people, but you're doing it with 50,000 people.
Bill Gurley01:02:15
Now, you didn't say, "I was gonna get rid of all my employees."
Brad Gerstner01:02:18
You're still growing the number of employees in the organization, but the output of that organization is gonna be dramatically more.
Jensen Huang01:02:26
This is often misunderstood. AI will change every job. AI will have a seismic impact on how people think about work. Let's acknowledge that. AI has the potential to do incredible good. It has the potential to do harm. We have to build safe AI. Let's just make that foundational. The part that is overlooked is when companies become more productive using artificial intelligence, it is likely that it manifests itself into either better earnings or better growth or both. Right. And when that happens, the next email from the CEO is likely not a layoff announcement. Of course. Because you're growing. Yeah. And the reason for that is because we have more ideas than we can explore. And we need people to help us think through it before we automate it.
Jensen Huang01:03:27
And so the automation part of it, AI can help us do. Obviously, it's going to help us think through it as well. But it's still going to require us to go figure out, what problems do I want to solve? There are a trillion things we can go solve. What problems does this company have to go solve? And select those ideas and figure out a way to automate and scale. And so as a result, we're going to hire more people as we become more productive. People forget that. And if you go back in time, obviously we have more ideas today than 200 years ago. That's the reason why GDPs are larger and more people are employed, even though we're automating like crazy underneath.
Brad Gerstner01:04:04
It's such an important point of this period that we're entering. One, almost all human productivity, almost all human prosperity is the byproduct of the automation and the technology of the last 200 years. I mean, you can look at, from Adam Smith and Schumpeter's creative destruction, you can look at a chart of GDP growth per person over the course of the last 200 years, and it's just accelerated. Which leads me to this question. If you look at the '90s, our productivity growth in the United States was about 2.5% to 3% a year. And then in the 2000s, it slowed down to about 1.8%. And then the last 10 years has been the slowest productivity growth. So that's the amount of labor and capital or the amount of output we have for a fixed amount of labor and capital.
Brad Gerstner01:04:54
The slowest we've had on record, actually. And a lot of people have debated the reasoning for this, but if the world is as you just described and we're going to leverage and manufacture intelligence, then isn't it the case that we're on the verge of a dramatic expansion in terms of human productivity? That's our hope.
Jensen Huang01:05:13
Right. That's our hope. And of course, you know, we live in this world, so we have direct evidence of it. Right. We have direct evidence of it, either as isolated of a case as an individual researcher who is able to, with AI, now explore science at such an extraordinary scale that is unimaginable. That's productivity, a measure of productivity. Or that we're designing chips that are so incredible at such a high pace, and the chip complexities and the computer complexities we're building are going up exponentially while the company's employee base is not—a measure of productivity. Correct. The software that we're developing better and better and better because we're using AI and supercomputers to help us, the number of employees is growing barely linearly. Okay, okay, okay. Um, another demonstration of productivity. So whether it's—I can go into, I can spot-check it in a whole bunch of different industries, I could gut-check it myself.
Brad Gerstner01:06:19
Yes. Your own business.
Jensen Huang01:06:21
That's right. And so I can, you know, and of course you can't—we could be overfit, but the artistry, of course, is to generalize what is it that we're observing and whether this could manifest in other industries. And there's no question that intelligence is the single most valuable commodity the world's ever known. And now we're going to manufacture it at scale. And we, all of us, have to get good at what would happen if you're surrounded by these AIs and they're doing things so incredibly well and so much better than you. Right. And when I reflect on that, that's my life. Right. I have 60 direct reports. Right. The reason why they're on e-staff is because they're world-class at what they do, and they do it better than I do.
Jensen Huang01:07:12
Right. Much better than I do. Right. I have no trouble interacting with them. And I have no trouble prompt-engineering them. I have no trouble programming them. And so I think that that's the thing that people are going to learn is that they're all going to be CEOs. They're all going to be CEOs of AI agents. And their ability to have the creativity, the will, and some knowledge of how to reason, break problems down, so that you can program these AIs to help you achieve something like I do—that's called running companies.
Brad Gerstner01:07:59
Right. Now, you mentioned something, this alignment, the safe AI. Mm-hmm. You mentioned the tragedy going on in the Middle East. We have a lot of autonomy and a lot of AI that's being used in different parts of the world. So let's talk for a second about bad actors, about safe AI, about coordination with Washington. How do you feel today? Are we on the right path? Do we have a sufficient level of coordination? I think Mark Zuckerberg has said the way we beat the bad AIs is we make the good AIs better. Is—how would you characterize your view of how we make sure that this is a positive net benefit for humanity as opposed to leaving us in this dystopian world without purpose?
Jensen Huang01:08:51
The conversation about safety is really important and good. Yes. The abstracted view, this conceptual view of AI being a large, giant neural network—not so good.
Brad Gerstner01:09:03⚠ 0.43
Right, right.
Jensen Huang01:09:04
Okay. And the reason for that is because, as we know, artificial intelligence and large language models are related, not the same. There are many things that are being done that I think are excellent. One: open-sourcing models so that the entire community of researchers and every single industry and every single company can engage AI and go learn how to harness this capability for their application. Excellent. Number two: it is under-celebrated, the amount of technology that is dedicated to inventing AI to keep AI safe.
Brad Gerstner01:09:41
Yes.
Jensen Huang01:09:43
AIs to curate data, to curate information, to train an AI; AI created to align AI; synthetic data generation; AI to expand the knowledge of AI, to cause it to hallucinate less; all of the AIs that are being created for vectorization or graphing or whatever it is to inform an AI; guard-railing AI to monitor other AIs—that the system of AIs to create safe AI is under-celebrated. Right. That we've already built, that we're building, everybody all over the industry—the methodologies, the red-teaming, the process, the model cards, the evaluation systems, the benchmarking systems. All of the harnesses that are being built at the velocity that it's being built is incredible. Under-celebrated. Do you guys understand?
Jensen Huang01:10:39
Yes.
Brad Gerstner01:10:40
And there's no government regulation saying you have to do this. This is—the actors in the space today who are building these AIs are taking seriously and coordinating around best practices with respect to these critical matters.
Jensen Huang01:10:55
That's right. Exactly. And so that's under-celebrated, under-understood.
Brad Gerstner01:10:59
Yes.
Jensen Huang01:11:00
Somebody needs to—well, everybody needs to start talking about AI as a system of AIs and system of engineered systems, engineered systems that are well-engineered, built from first principles, well-tested, so on and so forth. Remember, AI is a capability that can be applied. And it's necessary to have regulation for important technologies. But it's also, don't overreach to the point where some of the regulation ought to be done—most of the regulation ought to be done at the applications. The FAA, NHTSA, FDA, you name it, right? All of the different ecosystems that already regulate applications of technology
Brad Gerstner01:11:53⚠ 0.40
Right.
Jensen Huang01:11:54
now have to regulate the application of technology that is now infused with AI. Don't misunderstand, don't overlook the overwhelming amount of regulation in the world that are going to have to be activated for AI. And don't rely on just one universal galactic AI council that's going to possibly be able to do this. Because there's a reason why all of these different agencies were created. There's a reason why all these different regulatory bodies were created. We'll go back to first principles again.
Brad Gerstner01:12:32
I'd get in trouble by my partner, Bill Gurley, if I didn't go back to the open source point. You guys launched a very important, very large, very capable open source model recently. Obviously, Meta is making significant contributions to open source. I find when I read Twitter, you have this kind of open versus closed, a lot of chatter about it. How do you feel about your own open source models' ability to keep up with frontier? That would be the first question. The second question would be, is that, having that open source model and also having closed source models that are powering commercial operations, is that what you see into the future? And do those two things, does that create the healthy tension for safety?
Jensen Huang01:13:25
Mm-hmm. Open source versus closed source is related to safety, but not only about safety. So, for example, there's absolutely nothing wrong with having closed source models that are the engines of an economic model necessary to sustain innovation. I celebrate that wholeheartedly. Right. It is, I believe, wrong-minded to be closed versus open. It should be closed and open. Because open is necessary for many industries to be activated. Right now, if we didn't have open source, how would all these different fields of science be able to be activated on AI? Because they have to develop their own domain-specific AIs. And they have to develop their own, using open source models, create domain-specific AIs.
Jensen Huang01:14:20
They're related, again, not the same.
Brad Gerstner01:14:22⚠ 0.11
Right.
Jensen Huang01:14:23
Just because you have an open source model doesn't mean you have an AI. And so you have to have that open source model to enable the creation of AIs. So financial services, healthcare, transportation, the list of industries, fields of science that has now been enabled as a result of open source, unbelievable.
Brad Gerstner01:14:38
Are you seeing a lot of demand for your open source models?
Jensen Huang01:14:41
Our open source models, so first of all, Llama downloads, right? Obviously, yeah, Mark and the work that they've done, incredible. Off the charts. And it completely activated and engaged every single industry, every single field of science. Right, right. It's terrific. The reason why we did Nemotron was for synthetic data generation. Intuitively, the idea that one AI would somehow sit there and loop and generate data to learn itself, it sounds brittle. And how many times you can go around that infinite loop, that loop, you know, questionable. However, my mental image is kind of like you get a super smart person, put him into a padded room, close the door for about a month. What comes out is probably not a smarter person.
Jensen Huang01:15:38
But the idea that you could have two or three people sit around, and we have different AIs, we have different distributions of knowledge, and we can go QA back and forth. All three of us can come out smarter. And so the idea that you can have AI models exchanging, interacting, going back and forth, debating, reinforcement learning, synthetic data generation, for example, kind of intuitively suggests it makes sense. And so our model, Nemotron-340B, is the best model in the world for reward systems. And so it is the best critique. Okay. Interesting. Yeah. And so a fantastic model for enhancing everybody else's models. Irrespective of how great somebody else's model is, I'd recommend using Nemotron 340B to enhance and make it better.
Jensen Huang01:16:34
And we've already seen it made Llama better, made all the other models better.
Brad Gerstner01:16:38
Well... we're coming to the end. Thank goodness. As somebody who delivered DGX-1 in 2016, it's really been an incredible journey. Your journey is unlikely and incredible at the same time.
Jensen Huang01:16:55⚠ 0.45
Thank you.
Brad Gerstner01:16:56
You survived. Just surviving the early days was pretty extraordinary. You delivered the first DGX-1 in 2016. Mm-hmm. We had this Cambrian moment in 2022. And so I'm gonna ask you the question I often get asked, which is how long can you sustain what you're doing today? With 60 direct reports, you're everywhere. You're driving this revolution. Are you having fun? And is there something else that you would rather be doing?
Jensen Huang01:17:40
Is this a question about the last hour and a half? The answer is, I had a great time. I had a great time. I couldn't imagine anything else I'd rather be doing. Let's see. I don't think it's right to leave the impression that our job is fun all the time. My job isn't fun all the time, nor do I expect it to be fun all the time. Was that ever an expectation that it was fun all the time? I think it's important all the time. I don't take myself too seriously. I take the work very seriously. I take our responsibility very seriously. I take our contribution and our moment in time very seriously. Is that always fun?
Brad Gerstner01:18:26
No.
Jensen Huang01:18:28
But do I always love it? Yes. Like all things, whether it's family, friends, children, is it always fun? No. Do we always love it? Absolutely, deeply. And so I think the... 'How long can I do this?' The real question is, how long can I be relevant? And that only matters—that piece of information, that question can only be answered with—how am I going to continue to learn? And I am a lot more optimistic today. I'm not saying this simply because of our topic today. I'm a lot more optimistic about my ability to stay relevant and continue to learn because of AI. I use it, I don't know, but I'm sure you guys do. I use it literally every day. There's not one piece of research that I don't involve AI with.
Jensen Huang01:19:28
There's not one question that, even if I know the answer, I double-check on it with AI. And surprisingly, you know, the next two or three questions I ask it reveals something I didn't know.
Brad Gerstner01:19:40⚠ 0.29
That's right.
Jensen Huang01:19:40
You pick your topic. You pick your topic. And I think that AI as a tutor, AI as an assistant, AI as a partner to brainstorm with, double-check my work—boy, you guys, it's completely revolutionary.
Brad Gerstner01:20:02⚠ 0.00
Yeah.
Jensen Huang01:20:02
And that's just – I'm an information worker. My output is information. And so I think the contributions that I'll have on society is pretty extraordinary. So I think if that's the case, if I could stay relevant like this and I can continue to make a contribution, I know that the work is important enough for me to want to continue to pursue it. And my quality of life is incredible.
Brad Gerstner01:20:31
I'll say, I can't imagine. You and I have been at this for a few decades. I can't imagine missing this moment. It's the most consequential moment of our careers. We're deeply grateful for the partnership.
Jensen Huang01:20:42
Don't miss the next 10 years.
Brad Gerstner01:20:43
For the thought partnership. You make us smarter. Thank you. And I think you're really important as part of the leadership, right, that's going to optimistically and safely lead this forward. So thank you for being with us.
Jensen Huang01:20:57
Really enjoyed it. Thanks, Brad. Thanks, Bill. Good job.
Brad Gerstner01:21:09
As a reminder to everybody, just our opinions, not investment advice.