YouSaid · the spoken record
Tuhin Srivastava
- lines on the record
- 54
- first
- 2026-05-01
- most recent
- 2026-05-01
- sittings or episodes
- 1
- sources
- podcast
Every line below is reproduced as it was said and linked to the record it came from. Nothing here is summarised or generated. Directory · Search · Corrections
“Yeah, and I think what that means, that's amazing. I think that's great. And I think the education, same thing you have Clones Education. You get personalized access to everything. I think then you go one step. Back in how FX developers, I think, and companies, I think if you don't embrace this. Extension For a bunch of folks, which is like, you know, everything needs, and I don't think that means that. Designed Figma, I think that's a thing. I think what's more interesting is just like, you know. All these workflow and software companies need to figure out. What is the intelligent or intelligent inserted versions that drive the all that user value for those and consumers that we talked about”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“For consumers, it's the best possible thing, right? Like everything is somewhat smarter. You get better care because your doctors have access to better. Better tools. There's more, you know, like there's all this stuff about there being less software engineered. I think we just build more software. We just build a ton more software and like, you know, I see we're not slowing down hiring of software engineers, we're just building. More things. And that for the consumers, that just means better tools, more software, all those good things.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“More revenue is so. Yeah, I think inference going down just begets more inference. It is truly, I think we're kind of in a world that is, you know, it is the last market, right? Like even if there's AGI, all that's left is inference.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“But if you make it in more cheaper, they'll instead a hell of a lot more intelligent. And you see this with agents, agents are just longer running now. And I think that's what we have seen with the cost of inference going down, which is, you know, folks are just like, okay, we can run this for longer. We can make it do a bit more work and we'll get to a larger influence. I think compute scales from an inference perspective as well. And, you know, I think we are seeing that with almost all our customers, which is they either start with, this is the quality of answer, right? I need to get to. And this is the amount of inference I need to do to get there. Or this is the base level model that I can start with, that I can work with together. And I think the more we drive down the cost, what they realize is more intelligence just means better using a better.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, I think you think about this from a developer perspective and a consumer perspective. I think consumers just want the best answers and the best experience that some want governed by more intelligence to some extent. I think when you go to the developers from the developer's perspective, they would insert more intelligence if you make it cheaper. They will insert more intelligence anyway.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“When we have P0s, everyone on the call, like there's been a joke, there may as well be a siren that goes off in the office.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“I think one, I think if you've worked at infrastructure company, we were once in a meeting with a bunch of AWS execs and this was very senior AWS folks, all their pages went off multiple times during our 45 minute minute. I think it's very much just a cultural thing. But yeah, I don't, you know, our inference can't go down. And like, you know, you learn to like what's like I think Amir, my co-founder, when his pager goes off, his seven-year-old said, is that a P0? Is that a PZR? I think that is, you just have to get used to it. That's the culture you live in. It just changes the speed. But also becomes like a... Know a cultural thing. I think it's very, very, it rejects people that don't fit into it very, very quickly.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Like, is probably not the right place to be. But I think once you have, when you have that clear rubric, the people become very apparent that will fit into it and the people that don't fit into it also become very apparent. I think what's more like we've had amazing people, like you mentioned, but I think what's a lot more interesting is like, I think we haven't had a ton of turnover there unnecessarily. People tend to work because we have very clear on what we want. It took us a while to get there though.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“To like, you know, if you feel like you are micromanaging, if you feel like you have to be involved in everything, I think that's a bit of a cop-out as a founder because you're just like, I just need to be involved in everything. It's like, no, you probably just don't have the right people. I think the second thing is be very, very clear what you're optimizing for. Because I think when you're very, very clear what you're optimizing for, the people and like if it's something generic, like we want the smartest hardworking people, like you can't do much with that. Like with us, what we cared about was, hey, actually, we don't care about a lot of people who have done this before. We care about first people who think from first principles, work has to be a high priority, but they also have to be very kind and nice and care about the collaborative environment. We don't have a hero culture, very low ego. And, you know, if you need a manager,”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Everybody learns it. We're And we rolled on that. But I remember you said it so clearly at the time a lot. And I think that's what we noticed we're just actually having a leadership team that you can trust is so important. The two or three things that I'll say is like you want people where you can give them whole problems.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Are very, very flat until I know 18 months ago. I remember I went on a walk with a lot actually and a lot was like, you just need leaders. And like, it's actually so contrary to everything. As engineers, you're like, oh, everything is overhead. Everything at overhead.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Capacity. I think Yeah, I think capacity, I think the other one is probably just this market's so big and so like it represents a moment when you should be as aggressive as possible and really grown a ton obviously over the last 12 months and the last few months, but the answer is always just go bigger, go faster. And I think that's really, really fun. It's also a little exhausting and it's also like we are all in somewhat uncharted territory in terms of how fast and how big you can go and how things can get. But I think the big one is compute. I think like there's no world in which there's enough compute to get the amount of value that we want to get out of our lens in the next five to 10 years.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“The limitations of the current and the next set of primitives that need to be built from a scale security performance perspective. But I think it's really at the runtime level and assistance level. But the edge cases are, I'd say, a lot more systems level than they are LOM specific.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“You experience them and like, you know, and I'll give you a few examples here. Like, you see, you know, you start seeing yesterday we had for the first time ever, we saw some kernel panic that only happened because some fluent bit worker was creating too many logs and the scale was too big and it was all into one node and it was happening two terms at the same time by two different workers So you see all like the systems level and kernel level problems. But then you start to see I think the craziest stuff is you start to see with LLMs that these runtimes are pretty immature even how we use KV cache is you know you know probably a little less sophisticated than most people see than most people see and we are starting to see the”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah. I mean, I think you just, and again, like I think very, very large companies that run services basically out is probably the same stuff. All the edge cases just become”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“And then I think it's all the primitives that come after that. That just become incredibly marginative, both for us and our customers, which is stuff like sandboxes and the async batch inference, like how do we drive utilization by having a first class batch inference experience. To me, this is like what an inference cloud looks like. It's that you are very good at inference and then you start to do all the things tangential or that loop into inference and partner where necessary and build when necessary. But we really do want to own, like start with our core inference story and then go down to unblock supplier create margin and go up the stack to unlock value.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Loop between inference push reining because we think that just begets more inference. And so we will build a partner in almost everything there. So, you know, we're going to work with the best evals company in the world to make sure that's very well integrated, like Brain Trust into and around base 10. We will partner with or on the sandbox side, build the best sandboxes experience that will exist. And then we'll create the best training APIs to make it so continual learning becomes somewhat of a solve problem. It's not just like a discrete thing. That's, I think, the core base 10 product thesis. It's like, how do we build that loop? And then everything around that becomes how do we make sure that we can do everything we can to ensure that gets the biggest possible, that's access to compute, that's an infrastructure, make sure we can get compute anywhere, make sure we have access to our own compute.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, look, I think for us, all the runtime stuff is obviously very important. And what that means is like what shift we run on, how we run, what kind of workload we support, do we get very good at diffusion transformers? Yes, coding agents need sandboxes. We should go build sandboxes. There's all sorts of new speculation techniques to get faster imprint. We need to do that. Even stuff like KVCAWA routing and that stuff's a bit old now, but like getting continuing to be very good at that. somewhat disentangling pre-fill and decode and starting to treat them as separate problems. I think that's something we're very focused on and we're seeing massive gains. That's at the runtime level. I'd say beyond that, everything we think about is how to create more of that.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“The other players, like what you need to be able to compete here is the ecosystem to form around you. And if you tie up all your supply with one buyer, which a bunch of the other providers have done, it's actually hard for that ecosystem to form. If you think about if you're a big lab and you have a proprietary deal with one chip type where you get 90% of the supply, it's actually in your best interest to make sure you get 95% of supply and everything that's good for you and no one else ever use it.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, I mean, that was a hog rock LP thing. It's like, you know, I think that is very straightforward and makes sense. I think people really, really, really underestimate supply chain stuff in NVIDIA. Like how good they are that CUDA, how good CUDA is, the developer ecosystem around it. And the ability, like to me, like one of the most important things as an infrastructure company in this moment is how fast you can move and you can move faster with NVIDIA today. And I think that is the reality. And like it just like given the scale that they operate at, given the scale that they operate at, it's hard to see. I'm not saying it won't happen, like the short term, like in the next couple years, how anyone's going to be able to compete with that, especially with so much of the other.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, and I think everyone would be sad. I will say to some extent, which is, yeah, and I think there will be inference specific chips. I think you have like decode specific chips. I think, and we're looking at that.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“I think. I think diversification everywhere is a Same way I want to load on many models. And, you know, we want to load many most things. And I think.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“I think GPUs as a service is not sticky. I think that's been seen. Customers generally just see that as commodity. Inference with the software that I included is incredibly sticky. None of our top 30 customers have ever churned. We're talking. Like 400% annual NDR. Around our business. And so it's like very, it's very, very sticky. So I think that software layer is very important. The optimist in me is like, oh, there's so much value in the software. And we will build the best software layer for inference that exists, I think. As I think it's becoming clear now, access to inference computers. A strategic advantage, and I think that is the strategy that even the labs are going after, which is like, if we have all the compute, good luck running infere”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, exactly. Yeah, I think you need, like, I think there was demand for that. But I think the. One of the realizations that we had recently and with software people. And so we don't think like this all the time is that our business has very interesting working capital. Requirements. And I think even. And that as a result of that, it has very interesting financing. Requirement. And we're not, at least right now, we're not even going down to the”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“So, like, actually, what becomes important when acquiring capacity is you need to have enough demand to supply it to server, then you also need a low cost of capital, which is actually changing the dynamic pretty significantly.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah. You could buy that, but you go to also remember how quickly the market is moving. And like, you know, that gets balanced somewhat of like the fact that the H100 is such a great chip. And then it's crazy, it's four years, four and a half years old, and the price is going up still. Maybe he had a useful life nine years. So, you know, that's good. But at the same time, Same time, you know, yeah, yes, you can do that, but you're making a lot like you're making a lot of bets. As part of that. And then in terms of, I think that's the big thing that's changed over the last six months is that the term length that people want has just gone up. So if you wanted. A thousand. Than 24 B2 Which is from a good cloud right now, you're not getting that less than a three to five year. Right now with probably a 20 to 30% TCV prepay.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Suppliers right now it's kind of grifty, you know, like I think, you know, they haven't run data centers before. They don't understand SLAs, especially for inference. And so even when there is capacity available, there's a lot of there's probably we run a lot more and we have redundancy, so it's fine. But if you, you know, there's probably like a dozen good clouds and I probably like put like three or four of them in like the gold tier. And I think that just means that not only are we supply crunch, we're supplier and operationally crunched under people who can run these data centers as well.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“High fabric in half a day, maybe less And that gives us enormous flexibility. Even for us, it is hard for us to grow. We have a, I think it's, yeah, we have a 4 p.m. standing meeting for the company where we basically like, how do we manage capacity for the demand right now? I think the second part, which people don't really, the second part that people don't really understand is that there are also a lot of”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“18 different clouds now. We have 90 clusters around the world across 18 different clouds. And like, you know, initially, we started, we built this technology to be able to create one runtime fabric that spans all these different clouds and try to abstract that away from our customers as a way to think about reliability, latency, failover, all these things that we think are going to be very important for very mission critical use cases. That same technology, like just our ability to get compute wherever humanly possible has been really, really helpful in our ability to get supply. And what I mean by that is we can be introduced to a new provider. In a different country and have it up and running with the whole base 10 inference stack.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“I think There's so much narrative around the supply crunch. No matter, like as much as we hear about it, I don't think people realize how bad Really is like there is, you know, there is very, very little Slack compute available. We run pretty large clusters ourselves and we run them in uncomfortably high utilization. When I'm saying we're like mid 90s utilization. Most of the time, we have, we sit in.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“And that they have something special around that value. And once you have that value, it's like, okay, now how can I do that better, faster, and cheaper with the idea of being that, hey, if you need to be very good at customer support, you maybe don't need to be that good at coding and that a specialized model might be a better fit for that problem and you can do a better, faster, cheaper.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, I think it's, hey, go find to yourself with the best in class model that you have something worth optimizing. And I think, you know, A lot of a customer comes to us. Was that meme which was like, it was like two years ago? It feels like no GPUs. It's like no post-training pre-product market fit is whatever is what I'd say.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“We were doing a lot of research on the performance side and less so on the post reading side. It's interesting as we've started to do a lot more research on the post reading side, you start to see how linked inference and post training are. And like, you know, even when you think about stuff like quantization and when you should do that and like, you know, how training how you train the model affects how you need to quantize for inference and how paired these problems are. Has become very apparent and more and more we rely on the post rain infantile kind of both sides of the same problem. Because inference ideally will get more post-training. Where inference creates data, you do eval, you can now post train on that reward function that you found with those evals and hopefully to set up the entire look.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Just as I said in the opening statement here, which is as more and more post-trained models have come up, we've realized that the demand for people to either For software loops to do post training or for post training expertise is very high and we're really, really investing in that. There are also a bunch of Australians. I like to think that we had a bit of alpha there. But yeah, that's been fantastic. They're working with all sorts of customers. And it's also very interesting when you...”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah. So the restaurant around the acquisition was, you know, we are infrastructure and product people. We are product people. And now our really good infrastructure people. And we didn't have much of a research capability ourselves. And what we saw was the market moving heavily and heavily, like that we could accelerate the market itself with post-training resources, either productized or aren't even just as resources for that market. So PASD was a company that was a base 10 customer. So there were post-training models and running them on base 10. And I think what they realized was that they would eventually need to become an inference company. And what we realized was like, hey, we really needed that expertise because it is, it represents a way for us to get closer to the customer earlier and be able to support them all and just made sense as a like pairing them together.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“No, no. So we have dedicated inference, which is basically custom auto inference. Your SLA is your SLA. Then we have shared inference, which is shared inference endpoint, shared SLAs, and then we have a training business. I'd say 95% of the tokens today are on the first business. And almost all of them. There's probably for almost all of them, the customer is making some modifications to the model with their own data, specialized for the use case. And I think what's even more important is they might be compiling in different ways. No one is just running the vanilla open source weights. You might be customizing it for quality, but you also might be customizing it for performance”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“So, like 95% plus 95%, and I think that's really cool, to be honest. Look, we have two businesses. Three business we have three businesses Should”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“That form, I think it's just a massive loss. And as a country, we won't be able to innovate as fast because the cost of intelligence going down and control over intelligence, what we have seen just means more intelligence. Intelligence being embedded in multiple places.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“And I think the concern also there just becomes like, What happened if we aren't able to, like, if it is fun, like, I think if you think of the economics here, which is deep seek by most, deep seek is a very good model. And you can argue whether it's at the absolute frontier or not, but like, let's go back three months. So, think about everything. We were doing a whole lot of things three months ago. So let's just think about that. You could run Deep Seek, probably 20% of the cost. Running a bandropic models in production with comparable better latency, probably better reliability. If we don't have access to that intelligence,”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“And build with that in mind. It's like, you know, I think you're kind of missing the forest from the trees. There's two scenarios, right? Either. America does not ever come up with good open source models. There's probably a fundamental problem there, or we will get there, and we need to be ready for that work.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“To data, and you know, I don't, and I've never seen any real evidence except from some very early models that I think people picked up on very quickly that there is some agenda or bias built into these. I do think that to some extent is I think there is importance to the US that we develop our own models. I think that that would be a massive loss if there are five companies, five different labs in China that are creating open source models and we're struggling to get one set up so it's necessary. I also think it's inevitable. And, you know, like the deep-seek moment a year ago, I remember someone saying to me and I thought it was like very well said, which is like in the world's change a lot, but they said, hey, you know, we should kind of just forget. This is a Chinese model. We should just act like this came from”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, look, I think these models, firstly, are fantastic. They're amazing. We work with these teams. They're truly awesome. I'd say... It is hard for me to see. And I could be wrong, but if I network bound these models that they're not magically going to be able to cross those network boundaries.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Customers generally want to use whatever's at the frontier. And I think the difference is just being, I think we have a lot more visibility into how to run these and how to run these really well. And secondly, that they're good now.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, I think customers, at least the customers we are serving, are very, and these are like the fastest growing AI companies in the world that are very forward-thinking. They want to use the best model. And they are optimizing, I think there is. There's a subset of TOS which I think is small today, where people really start to start with cost. But everyone comes for capability first because that's really where the economic growth is being unlocked, where the value is being delivered. And then they optimize. And I think that's actually been, you know, and so with that in mind, you name everything from GPT OSS all the way to... Moonshot Molo to deep seeks. Canopy or Orpheus, which is like really good text of speech. Models.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“All these writer gamma, all these companies serve enterprises in math. And what we actually get is like a translational requirements from them, which is like, you know, they're like, hey, we need this sort of data retention. We need models need to be deployed. This is the types of GPUs or the latencies they're okay with. This is the model requirements from like a transparency perspective that they care about. And so I think that is actually the more nuanced answer. It's that if you listen to what their needs are, we actually get a full translation of what the enterprise will require. Like I would say that by serving companies like a bridge and open evidence, we're probably pretty well suited to go serve the healthcare system given that they are selling and latent health, given that they are selling to them.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Yeah, I think mostly you just learn a lot by building with the company's greatest scale, doing the most interesting things. We think of it two ways. I think there's like the most obvious way, which is just build for the highest scale, the customers that will push you the most from technologically and everything kind of will fall into play. I think the stripe evolution as a company showed that we're just like Stripe now serves like so many enterprises, but 12 years ago that wasn't the case. But they just built for the frontier and kind of went with them. And the second way we think about that is to just think about building for companies that are serving enterprises. So yes, we don't serve the enterprises, but our customers serve enterprises. It really serves enterprises, but Evan has Dekagon.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“100%. But what's cool is that we're seeing the transition happen. Before it was like, hey, are they using AI tools? I don't think that was immediately obvious two years ago. I think that's obvious now that yes, they are Are they using closed source model APIs? I think they're starting to get there. And I think once you do that, and then you kind of see what is possible, then comes the whole custom model adoption. I think that is all that is ahead of us.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“It is crazy. The answer is just that it's crazy that the answers do, I think, I think if you look by inference count. It'd be 99% the full. I think that kind of represents the scope of the opportunity here is that the majority of the market hasn't come online and added AI into the”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“Rare to get access to. There will be an application layer. And I think support companies is another example of that where a support task isn't one-shotted. Usually at a company like Bayettown when a ticket comes in, there's like what like 1, 2, 10, 20 actions that get taken. And that is where, you know, someone can develop a specialized model.”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT
“A bridge is an ambient scribe that is used by physicians in Almost all hospitals in the US. I think lots of investors, Greg Shift's amazing, great company, great team, great product. And they've basically got this very, very deep integration into hospitals, into clinician workforce. And my argument would be here is that actually it's very, very hard for a frontier model company to be able to eat a boy at that because they just don't have access to that user signal. And what will happen over time is the folks who have access to that user signal can start to post train models on that reward signal and start to get long horizon agentic models running that. And I think to the extent that that is possible and that signal is differentiated and unique and is somewhat”
2026-05-01 · No Priors · Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud · IDENTIFIED FROM THE TRANSCRIPT