Search for an article…

/

f

Focus:

Off

0

Search for an article…

/

f

Focus:

Off

0

~

/

/

The Prince of Data

Technology

The Prince of Data

An interview with Databricks cofounder and Berkeley professor Ion Stoica

Ion Stoica is a Romanian-American computer scientist and Professor at the University of California, Berkeley. As cofounder of the unicorns Databricks and Anyscale, Stoica is among the most prolific academic-entrepreneurs of the 21st century.

Born in 1965, he was raised in Nicolae Ceaușescu’s Romania, where he served an obligatory nine months in the armed forces. He completed an MS in computer science at the Polytechnic University of Bucharest in 1989, a year of profound change across Europe. As the Warsaw Pact disintegrated, Romania’s particularly harsh totalitarian regime collapsed under the weight of its own contradictions. A combination of extreme austerity and social repression, combined with Ceaușescu’s distinctive policy of self-reliance, had eroded the regime’s reputation across all levels of society. In December, Ceaușescu was overthrown and executed, and Romania began its westward transition. Not long after the revolution, Stoica departed his homeland for the United States.

Stoica has had an enormous impact on computer systems and networking beginning with his early work as a graduate student at Old Dominion University, a public university in Norfolk, Virginia. At ODU, he devised a process-scheduling algorithm that was later adopted as the default scheduler in the Linux kernel. After moving to Carnegie Mellon to complete his PhD in electrical and computer engineering, and with a brief stopover at MIT, Stoica began teaching at Berkeley in 2001. In 2006, he launched Conviva, a streaming analytics company that was born out of his academic research into video streaming over the internet.

After launching Conviva, Stoica’s research became increasingly focused on big data. In 2009, machine learning researchers at Berkeley’s interdisciplinary AMPLab found that Hadoop, which was then the dominant framework for large-scale data processing, was hopelessly slow for algorithms that needed to iterate over the same dataset many times. In response to this, Stoica’s doctoral student Matei Zaharia built a compact system that kept data in memory rather than repeatedly writing it to disk. Open-sourced in 2010, that system, Apache Spark, has become one of the most influential pieces of data infrastructure ever created.

Databricks was founded in 2013 by Stoica, Ali Ghodsi, and five PhD students including Zaharia. Stoica served as the company’s first CEO from 2013 to 2016, when he became executive chairman. In early August, Databricks closed a $5 billion strategic funding round at a $190 billion valuation. In 2019, Stoica co-founded Anyscale, built around Ray, a distributed computing framework developed in his lab to overcome structural limits in Spark itself. Ray now undergirds much of the AI industry’s training and inference workloads. Meanwhile, Anyscale, where Stoica also serves as executive chairman, has reached a private market value of over $1 billion.

As director of Berkeley’s Sky Computing Lab, Stoica is currently preoccupied by questions surrounding the reliability of AI and the mounting complexity of the AI stack itself. Some of his recent research projects have included vLLM, Vicuna, MemGPT, and the model-evaluation platform Arena AI (formerly LMArena), alongside work on inference. Professor Stoica continues to teach Berkeley graduate and undergraduate students.

I sat down with him to understand his journey from communist Romania to the pinnacle of American academia and business, the enormous commercial successes of his research, and his perspective on AI. What follows is a transcript of our conversation.

***

Carson Becker: Tell me about your family and your upbringing in Bucharest, Romania.

Ion Stoica: I was born in Bucharest. Both of my parents were engineers. My father was a geophysicist. His work entailed looking under the surface of the earth, trying to figure out where to find oil and other things like that. My mother was a geologist, so pretty related. My grandparents were out in the countryside, so I was spending, in general, the summers and some of my vacations there.

CB: What was your family’s experience during the many upheavals of the 20th century: the world wars and so on?

IS: My grandparents on my mother’s side were pretty close to Bucharest, like 50 miles. They had a tougher time because of collectivization: the state taking their land and putting it together in a cooperative, as it was called. So that was tougher. Actually, my mother, because my grandparents had some land and so forth, was excluded for two years from college because of that.

My grandparents on my father’s side were a little farther away, near Târgoviște, a city which was the capital a long time ago, and it was a little bit in the hills. The communists did not take the land there, because it was more for growing apple trees and things like that. You couldn’t grow grain. So they still had some land and were doing relatively well compared to others.

My family experienced the Second World War, and I just heard stories, also a little from the First World War. Moldavia was part of Romania, and just before the Second World War, Russia invaded and took it away. That was one of the reasons Romania was initially fighting on the side of Germany: to free that territory. In 1944 there were changes in alliances. But really, I think that from all sides of my family, it is obviously a story that the communists didn’t have a positive impact on them, because fundamentally they took away property and land and things like that.

CB: What was your education like, and how did the communist regime perform in that regard?

IS: They actually did invest in education. Let me take two steps back. Romania had been more or less a constitutional monarchy since around 1850, and it was a pretty democratic country. There was this reform giving land to the peasants after the First World War. Economically, Romania was pretty good, they were manufacturing trains, airplanes, and so forth, kind of in the middle of the European economic rankings. After communism, it was pretty much at the bottom. What I’m trying to say is that communism was not good for the economy in general.

Now, I obviously cannot compare education before and after, but I do think that when I grew up, the education up to college, and maybe including college a bit, was very solid. A lot of math, a lot of STEM, as you call it here. And you had to learn two languages: one starting in the first grade, the other in the fourth or fifth grade. Unfortunately, in some cases one of them was Russian.

They put a lot of value on being good at school and being educated: getting into college and so forth. This was the state, like other communist countries. Romania, like Russia, wanted to use education and scientific progress as a tool of propaganda, to demonstrate the superiority of the system. This is what Russia did after the Second World War, and they were pretty successful for a while.

CB: What was your experience in the Romanian military?

IS: I did nine months. It was a unit that was not really for combat, it was related to electronic warfare: using radar and sensors for discovering the enemy. It wasn’t that bad, except that it wasted one year. It was pretty close to Bucharest, like 40 miles by train. I learned how to shoot and things like that, but in the mornings we also had classes to learn about electronic warfare and so forth.

CB: Did you enjoy it?

IS: No, of course not. At the end of the day, I could have done a lot of more useful things with that time.

CB: You completed your MS in Romania, right around the fall of the Ceaușescu regime in 1989.

IS: Yeah, it was right after the regime fell.

CB: What drove you to then pursue further studies in the United States? I understand you began at Old Dominion University in Virginia.

IS: I started at ODU because there was a Romanian faculty member there, his name was Stephan Olariu. At that time I didn’t know as much about the opportunities in Europe or in the US. After two years at ODU I decided to transfer to Carnegie Mellon University, where I finished my PhD.

CB: How did you formulate your thesis on quality of service in the internet?

IS: At the high level, I was interested in how to better manage resources and scheduling. Actually, at ODU I worked on this in the context of operating systems, and some of the stuff I did there, a scheduler for operating systems, is right now the default scheduler in Linux, as of, I think, one or two years ago. It’s called EEVDF, a pretty complicated name, not a good name, but anyway.

When I moved to CMU, that was ‘96. I changed my thesis a bit to focus on the internet: how to provide quality of service for, say, voice and audio on the internet. As you know, that was a golden age for the internet, when everything was happening. Everything was about the internet, like today is all about AI. Google was founded in ‘98, right? Amazon, I think, ‘96. So there was a lot of excitement. I focused on internet networking, and I graduated in 2000. After I graduated, I spent a few months at MIT, and then I joined Berkeley. I started to teach here, I think, January 2001, and I’ve been here since then.

CB: What drew you to Berkeley, and what makes it so special as a university?

IS: I was lucky and got quite a few offers when I graduated, and I came to Berkeley for a few reasons.

One, I liked the fact that it was on the West Coast, where everything was happening, at least in networking. Cisco was there, a bunch of other companies, startups, even AT&T and Bell Labs had research labs on this coast back then. I wanted to be close to where things were happening on the internet.

The other thing I liked about Berkeley, it did have a good balance, at least at that time, between people going to academia and going to industry or starting companies. When I looked, I liked that balance, and for a long time, even among my students here at Berkeley, half went to academia, at least 40, 50 percent. Now, for the past few years, things have been skewed toward industry, toward these labs. We can discuss that.

Then there are two other things I liked about Berkeley. One was open source. Berkeley was, in some sense, at the start of the battle of open source. Before Linux there was FreeBSD, the Berkeley Software Distribution. And again, remember that when I graduated, it was networking: a big part of the internet protocols, this TCP/IP, was developed at Berkeley, and Berkeley pioneered networking in the ‘80s and the beginning of the ‘90s.

And the final thing was more of an intangible. A lot of times, Berkeley was a pioneer in new domains and new technologies, the first among the big universities: Stanford, MIT, CMU, and so forth. They pushed on databases early on; Mike Stonebraker was here. I mentioned FreeBSD. Around that time, a little bit before I came here, they were doing this Network of Workstations project, which was about building big computers from commodity servers, as opposed to building supercomputers like Crays. This idea of connecting standard servers with fast networks to create a bigger computer is what became the foundation of all these big internet companies, including Google. They didn’t buy supercomputers, they just connected their commodity servers to create this huge compute and storage infrastructure. And there were sensor networks, also very early on. So I liked that pioneering aspect.

To summarize: I wanted to be where things were happening. They were saying at that time, well, we are close enough to Silicon Valley, but not that close. A lot of things happened in the South Bay, around Stanford, and I liked that balance, because I’m also an academic at heart. And I liked the open source, which was in Berkeley’s DNA, and finally that pioneering aspect, maybe more than at other top schools.

CB: During your first few years at Berkeley, what sort of research did you undertake? The first commercialization of your work was in 2006, with Conviva, right?

IS: When I graduated from CMU, I really wanted to make what I proposed in my thesis real. I spent six months or so trying, working with people, pushing on this standardization effort, because you need to standardize in order to be adopted in the internet. It’s very hard, right? There’s only one internet, so the barrier to adoption is pretty high. Anyway, I spent quite a bit of time pushing for the techniques I proposed in my thesis to make it into the internet, and it was very hard – now, looking back, for obvious reasons.

So after that, I started to move up the stack, like they say, where it’s a little bit easier to make an impact. Immediately after I graduated – I told you I spent a few months at MIT – I worked on what was then another hot topic: peer-to-peer networks. You know, it was Napster and Gnutella and so forth. How to make them much more efficient. I worked on that for a few years, then probably in 2006–2007 I moved toward big data, and around 2015–16 I started to work on AI and systems.

So if you want to look at my career: at ODU I did operating systems scheduling; that’s ‘94 to ‘96, something like that. At CMU, my PhD was networking: internet, quality of service, scheduling. Then peer-to-peer, 2001 to maybe 2004–2005. Then from 2006–2007 I did big data, and around 2015 I started on AI and systems. Of course, there is no very strict delineation, right? You start something new, and what you were doing is still going to continue, maybe even for some years.

The first company, Conviva, was about video distribution on the internet. It was based on the peer-to-peer technologies that we developed.

CB: What were the commercial assumptions which led to Conviva, and what mistakes were made, considering this was your first such effort?

IS: When you talk about using peer-to-peer to distribute video and audio and big files at that time, what was the main assumption? The main assumption was that the internet was going to be overwhelmed by this new traffic, and the idea of peer-to-peer is that, to avoid that, you try to localize the traffic at the edge of the network rather than having everything go through the core. It’s like in a city: you try to localize the traffic at the edges instead of having all the cars go through the center. That was one of the main ideas. The other one was about cost: the narrative was that because the internet was going to be very congested, it was also going to be very expensive.

Now, what happened after we started the company is that that assumption proved to be, I wouldn’t say not true, but not that strong. Actually, it turns out that because people laid so much fiber before the dot-com bust, there was enough fiber that wasn’t activated, dark fiber and so forth. So the internet scaled much better than many people expected, and hence the prices also went down. I remember when I started Conviva in 2006, the cost of downloading one gigabyte was something like 40 cents, and in two years it went down to two cents or something like that. Like 20x.

Obviously, when you build a company or you do anything, you try to solve a problem. You build a product to solve a problem, because that’s why people will buy your product. So if the problem disappears, or is no longer as big, then whatever you built is no longer as relevant, no longer as useful. So we pivoted Conviva, and Conviva is still running today; it’s cash-flow positive. Instead of targeting lowering the cost and scaling, we targeted quality: providing the highest quality of video distribution. And then we had these big customers like ESPN and HBO and Disney. NBC was another one, people starting to use the internet to distribute their content.

The lesson was that you’d better do something which is on the right side of the trends. There are these secular trends, and you’d better be on the right side of them, because there’s not much you can do if you are not.

The other thing I learned, and it remains with me to this day, is that consistency is more important than accuracy. We provided the users some metrics, some dashboards, about how their content distribution was doing: the number of users, how much they watched, and so forth. At some point we had a new release, and it was better. It was more accurate. But because it was more accurate, some of these numbers changed. And I remember what a difficult discussion it was with the customers, because you have these numbers changing: the number of simultaneous users, instead of, say, 10,000, is 9,500 or something like that, because you measure more accurately, and they are unhappy about that.

And it makes sense why they were unhappy: they had built their business processes based on these numbers. These are their metrics, which drive their businesses. So now, if these metrics change, you have a problem, right? So we were then trying to provide consistency with these previous numbers. That’s when you learn that consistency is far more important than accuracy. And it’s something general, I’m sure for every business, you have some metrics about success: impressions, clicks, whatever it is. If I measure it in a different way, even if it’s more accurate, but it gives you a lower number, you’ll not be happy, right?

CB: What would you say preempted your shift into data, around the time your lab began working on Spark?

IS: There are many things which coalesced into this, not in any particular order; I’m just going to order them for the narrative.

One thing that was happening: like I mentioned, at Conviva we were providing this dashboard for the customers, and the users could also query the database, to ask questions about the users, about failures in the video streaming, about the quality of the streaming for a particular group of users or particular content. And for that, we used traditional databases back then: MySQL. I remember this particular episode. You anticipate the growth. At that time the cloud was in its infancy, and we ordered a bunch of servers, the most powerful servers we could think of at that time, to run our database. By the time we got the servers, our traffic was already exceeding their capacity. And we struggled so much that it actually led to removing some features from the product; we couldn’t support them with the server capacity we had bought. So that was one angle. I had seen that problem, and I was thinking, there has to be a better solution here.

The other thing happened at Berkeley. We have these five-year labs, and in these labs there is a group of faculty, me included, and their students, working on particular projects. That lab was interdisciplinary: it had people from hardware, databases, systems, networking, and machine learning, very early on. That was around 2007–2008. These labs also worked very closely with industry. You have these retreats where you meet with industry, and at that time, if you remember, big data had already started to be a thing, and the open-source system out there was Hadoop, developed at Yahoo. We knew people there, so we were already talking in these meetings and retreats. We knew about all that.

Now, I told you the lab was interdisciplinary, and one of the groups from machine learning wanted to do the Netflix challenge: $1 million for whichever group could develop the best recommendations. Netflix provided some data to train your algorithms on, and then evaluated your algorithm to see which one was best. So a group from our lab, machine learning people, decided to compete. It’s a lot of data, so they came to us to ask what they should use to process it and implement their machine learning algorithms. And the only thing we had for them then was to use Hadoop.

Now, the problem is Hadoop was very slow, because Hadoop has these stages, and at each stage it reads the data from the disk and writes the data to the disk. It’s doing that also to be fault-tolerant: if something fails, you still have the data on the disk. And they came back and said, “Oh, but this is so slow,” because in machine learning you have many, many iterations to converge on a good model. Again, that was before; it was classic models: random forests, collaborative filtering, linear regression, things like that.

By the way, at that time we had already made some impact in the Hadoop community. One of the students was Matei Zaharia.

CB: Also Romanian, right?

IS: He was born in Romania, but he went to France, and he came here earlier in his life. He came to Berkeley from Waterloo to do networking. I think his first internship was at Facebook, and at that time, 2008, Facebook was small. Facebook was still in downtown Palo Alto, and they had their big data cluster: like 80 nodes, three people running them. And we solved some problems for Hadoop around scheduling. That’s why I’m saying we were plugged into the community pretty strongly.

So then these machine learning people from our lab come back saying, oh, this is too slow, and Matei put together a small system to address the main limitations of Hadoop at that time. And this was the genesis of Spark. This happened in 2009.

It’s complicated, but initially we didn’t think about Spark as being the big thing. Actually, we had another project called Mesos, which was about how you share a cluster between different large workloads. That was quite popular, and Spark was supposed to run on top of it, to demonstrate the generality and flexibility of Mesos.

Spark was a few hundred lines of code initially. In 2010 it was open-sourced, and then it started to get more and more traction. It was Matei and a great group of students, passionate about it, and it was the right time: people wanted, at least for some workloads, something faster than Hadoop. One of the reasons Spark was faster is that it was keeping as much data as possible in memory, so you don’t need to go to the disk. That made it faster both for doing multiple iterations, so for machine learning, and also for these interactive queries, because the data is in memory instead of on disk. As I mentioned to you, interactive queries was the problem I had with the databases at Conviva. So that’s the genesis.

CB: Did you enjoy your role as Databricks CEO, which you relinquished in 2016? Do you enjoy the business executive role, because you said your heart is in academia?

IS: Actually, things in industry and academia are not so different. In both cases, like I mentioned, you solve problems, right? That’s what you do. So choosing what problem to solve is important in both, and being on the right side of these secular trends. I think it’s quite similar, especially if you work on a larger project in academia.

Obviously, there are many differences as well. Academia is more of an exploration, with many things happening early on. What is the good thing about industry? You can put more and more resources behind an idea. You start with a few people, then tens, then hundreds, then thousands. In academia, that’s not possible. You can have a few people, maybe five, ten people at most, working on a project, a system, but no more. But you are going to have a lot more freedom to explore.

So, you asked me, I think I enjoyed it. There are big things I enjoyed. Obviously, you need to make certain bets, especially early on. And I’m always excited to recruit great people, this is what you do in academia too, you recruit students, faculty, and so forth. I really enjoyed that. And seeing something grow quickly is always exciting.

CB: How are Databricks and Anyscale, respectively, navigating the agentic AI wave?

IS: Databricks provides a broad data and AI platform. It enables enterprises to build and consume agentic applications. So, with the Lakehouse, Lakebase, and Unity Catalog, Databricks gives customers governed and secure access to their data, regardless of where it is stored. Then on top of that, it offers agentic products such as Agent Bricks, Genie, and Unity AI Gateway, which adds things like budget, governance, routing, guardrail, and usage-monitoring controls for AI services and agents. I would also mention Omnigent, an open-source meta-harness for AI agents that helps customers leverage and control existing agents such as Claude Code, Codex, and Cursor.

In contrast, Anyscale provides a lower-level platform and tools that help customers, including physical AI companies, build their own AI platforms, including agentic platforms. So, through open-source Ray, Anyscale helps customers scale complex AI workloads, including training, post-training, inference, and multimodal data processing. Many of today’s AI libraries use Ray to scale, including the vast majority of post-training and RL frameworks.

CB: How do you think government, academia, and industry can better coordinate to solve the problems emerging with AI? Is the cooperation efficient and effective at the moment?

IS: It’s absolutely insufficient. It’s important to look at the past. Look at the internet: it was the poster child of this triumvirate, a very strong collaboration between government, industry, and academia. The government, through DARPA, started the internet program. Academia did the first deployments. Then, when the internet was, so to speak, privatized in ‘89, the industry, AT&T and the big companies laid fiber, built and expanded the internet. Even the standards were mixed academia and industry. The standards body was the IETF, the Internet Engineering Task Force. Unlike other domains, which are very industry-heavy.

Compare that with what happens today with AI. It’s almost the opposite, absolutely the opposite. First, let me talk about the industry. Right now you have all these big labs, which are very much siloed, and very little publication; there is not much information they are going to share publicly, and everyone is doing its own stuff, pretty secretive. So it’s not this free exchange of information which happened in the past, in the case of the internet. That’s one.

The other thing is that the cost is so high to build these models and run them at a certain scale that, again, there are no strong relationships between academia and industry. And the government, if anything, should allocate much more money than in the past, because this is a much more expensive enterprise. So this worries me: you do not have this fundamental diffusion of innovation, of knowledge. If these labs don’t publish, if they don’t share, that knowledge is not available for others to build on top of. And you can see maybe a little bit of what happens: the best open-source models are now coming from China, because, ironically, they are more open than the US.

Also, think about the human capital. If AI is so important, so important for our society, for our civilization, maybe, and of course there are problems, there are challenges, you want your best minds to work to address these problems, whatever they are: security, the impact on jobs, whatever it is. But in order for people to work on them, they have to have shared artifacts they can work with, and shared knowledge. And now this doesn’t happen, because each lab is doing more or less the same thing. It’s a very inefficient use of human capital. The only silver lining is that in California, at least, non-competes are not legal, and people can move. So your diffusion of innovation is people moving between the labs.

CB: You were talking at the beginning of this conversation about the best minds flocking to industry at the expense of academia.

IS: This is again something worrisome. So let’s put it this way: you can understand why this is happening. You are doing research in AI, you are in academia, and one of these companies approaches you and basically says: “Hey, if you join me, you have access to resources you don’t even dream of having in academia.” So this is number one: you can do a lot more interesting research here than in academia. Number two, you can get lots of money: seven, eight figures, even nine figures. And on top of that, they are going to say AGI will come in two years, so if you do not come now, then it’s too late.

And in some cases, people say: oh, if AGI comes in two years, you will not have anything to do anyway; your value will not be there. So come now, do it, be part of it, and get lots of money. It’s a very effective sell. And it’s easy for academics to rationalize: I am going to do this because I am going to do much more research and have more impact than I would in academia. It’s now or never. And in two years, when everything is done and AGI arrives, maybe I’ll come back. We’ll see what happens.

And this is effective, even for students. So that’s what happens. It has never happened before, at least from what I’ve seen.

CB: Which of your recent research efforts–I’m thinking of vLLM, Vicuna, MemGPT, and so on–and which of the emerging research fields are the most exciting for you personally? And which has the most potential, in your view?

IS: I think there are two trends happening. We were talking about trends.

One, on the systems side, is basically all these layers collapsing. It used to be that you had different layers: networking, and even networking has different layers, the operating system, the application, and so forth. What layering gives you is this modularity in building a system: it makes it easier to develop, and you can develop one component without affecting the rest. It’s like in a car: I put in a better engine, the rest remains the same, but it’s going to be faster. However, what you lose, you lose in terms of efficiency. If you build a truck, you can go to the parts bins and pick the engine, pick the wheels, all these standardized parts which are already there. But if you build a race car, you integrate everything. The same here: you have to integrate everything, because every inefficiency is so expensive. It costs a billion dollars to train a model; if a modular architecture impacts the efficiency by 10 percent, that’s a hundred million, right?

So that’s the collapsing, and the stack becomes so much more complicated to build, and things are changing fast, faster and faster. Building this stack for AI is an order of magnitude harder than building the stack for previous generations, like the web or big data. That’s one.

The other one, which I think is very important: when you talk about AI right now, what is the hardest part? The hardest part is not about capabilities. We know that AI has fantastic capabilities. It’s about deploying them in production reliability, predictability: I think these are extremely important aspects. How do you do that? And this will become more important, because AI is about trust. If you use AI, you need to trust it to do the right thing. Otherwise, you wouldn’t use it, at least for critical tasks, which are the most valuable tasks.

On that side, I’m excited about Arena AI, about methods to evaluate these models and applications, and how to understand what the gaps are: where they are doing well, where they are not doing well, and things like that. I think this is very important. And the other thing is inference, because there are three major workloads: training, inference, and this post-training, like reinforcement learning and things like that.

CB: If you were starting your PhD over today, what specifically would you choose to focus on?

IS: Good question. I’ll focus on what I’m actually working on today. There are two main directions.

The first is making AI reliable: ensuring that its outputs are predictable and aligned with the user’s intent. This will be critical for AI to reach its full potential and to be used in mission-critical, high-value applications.

The second is using AI to build the AI stack itself, which is growing increasingly complex. Far more complex than previous technology stacks, such as big data or web infrastructure. In other words, using AI to build better AI.

Technology

The Prince of Data

An interview with Databricks cofounder and Berkeley professor Ion Stoica

Ion Stoica is a Romanian-American computer scientist and Professor at the University of California, Berkeley. As cofounder of the unicorns Databricks and Anyscale, Stoica is among the most prolific academic-entrepreneurs of the 21st century.

Born in 1965, he was raised in Nicolae Ceaușescu’s Romania, where he served an obligatory nine months in the armed forces. He completed an MS in computer science at the Polytechnic University of Bucharest in 1989, a year of profound change across Europe. As the Warsaw Pact disintegrated, Romania’s particularly harsh totalitarian regime collapsed under the weight of its own contradictions. A combination of extreme austerity and social repression, combined with Ceaușescu’s distinctive policy of self-reliance, had eroded the regime’s reputation across all levels of society. In December, Ceaușescu was overthrown and executed, and Romania began its westward transition. Not long after the revolution, Stoica departed his homeland for the United States.

Stoica has had an enormous impact on computer systems and networking beginning with his early work as a graduate student at Old Dominion University, a public university in Norfolk, Virginia. At ODU, he devised a process-scheduling algorithm that was later adopted as the default scheduler in the Linux kernel. After moving to Carnegie Mellon to complete his PhD in electrical and computer engineering, and with a brief stopover at MIT, Stoica began teaching at Berkeley in 2001. In 2006, he launched Conviva, a streaming analytics company that was born out of his academic research into video streaming over the internet.

After launching Conviva, Stoica’s research became increasingly focused on big data. In 2009, machine learning researchers at Berkeley’s interdisciplinary AMPLab found that Hadoop, which was then the dominant framework for large-scale data processing, was hopelessly slow for algorithms that needed to iterate over the same dataset many times. In response to this, Stoica’s doctoral student Matei Zaharia built a compact system that kept data in memory rather than repeatedly writing it to disk. Open-sourced in 2010, that system, Apache Spark, has become one of the most influential pieces of data infrastructure ever created.

Databricks was founded in 2013 by Stoica, Ali Ghodsi, and five PhD students including Zaharia. Stoica served as the company’s first CEO from 2013 to 2016, when he became executive chairman. In early August, Databricks closed a $5 billion strategic funding round at a $190 billion valuation. In 2019, Stoica co-founded Anyscale, built around Ray, a distributed computing framework developed in his lab to overcome structural limits in Spark itself. Ray now undergirds much of the AI industry’s training and inference workloads. Meanwhile, Anyscale, where Stoica also serves as executive chairman, has reached a private market value of over $1 billion.

As director of Berkeley’s Sky Computing Lab, Stoica is currently preoccupied by questions surrounding the reliability of AI and the mounting complexity of the AI stack itself. Some of his recent research projects have included vLLM, Vicuna, MemGPT, and the model-evaluation platform Arena AI (formerly LMArena), alongside work on inference. Professor Stoica continues to teach Berkeley graduate and undergraduate students.

I sat down with him to understand his journey from communist Romania to the pinnacle of American academia and business, the enormous commercial successes of his research, and his perspective on AI. What follows is a transcript of our conversation.

***

Carson Becker: Tell me about your family and your upbringing in Bucharest, Romania.

Ion Stoica: I was born in Bucharest. Both of my parents were engineers. My father was a geophysicist. His work entailed looking under the surface of the earth, trying to figure out where to find oil and other things like that. My mother was a geologist, so pretty related. My grandparents were out in the countryside, so I was spending, in general, the summers and some of my vacations there.

CB: What was your family’s experience during the many upheavals of the 20th century: the world wars and so on?

IS: My grandparents on my mother’s side were pretty close to Bucharest, like 50 miles. They had a tougher time because of collectivization: the state taking their land and putting it together in a cooperative, as it was called. So that was tougher. Actually, my mother, because my grandparents had some land and so forth, was excluded for two years from college because of that.

My grandparents on my father’s side were a little farther away, near Târgoviște, a city which was the capital a long time ago, and it was a little bit in the hills. The communists did not take the land there, because it was more for growing apple trees and things like that. You couldn’t grow grain. So they still had some land and were doing relatively well compared to others.

My family experienced the Second World War, and I just heard stories, also a little from the First World War. Moldavia was part of Romania, and just before the Second World War, Russia invaded and took it away. That was one of the reasons Romania was initially fighting on the side of Germany: to free that territory. In 1944 there were changes in alliances. But really, I think that from all sides of my family, it is obviously a story that the communists didn’t have a positive impact on them, because fundamentally they took away property and land and things like that.

CB: What was your education like, and how did the communist regime perform in that regard?

IS: They actually did invest in education. Let me take two steps back. Romania had been more or less a constitutional monarchy since around 1850, and it was a pretty democratic country. There was this reform giving land to the peasants after the First World War. Economically, Romania was pretty good, they were manufacturing trains, airplanes, and so forth, kind of in the middle of the European economic rankings. After communism, it was pretty much at the bottom. What I’m trying to say is that communism was not good for the economy in general.

Now, I obviously cannot compare education before and after, but I do think that when I grew up, the education up to college, and maybe including college a bit, was very solid. A lot of math, a lot of STEM, as you call it here. And you had to learn two languages: one starting in the first grade, the other in the fourth or fifth grade. Unfortunately, in some cases one of them was Russian.

They put a lot of value on being good at school and being educated: getting into college and so forth. This was the state, like other communist countries. Romania, like Russia, wanted to use education and scientific progress as a tool of propaganda, to demonstrate the superiority of the system. This is what Russia did after the Second World War, and they were pretty successful for a while.

CB: What was your experience in the Romanian military?

IS: I did nine months. It was a unit that was not really for combat, it was related to electronic warfare: using radar and sensors for discovering the enemy. It wasn’t that bad, except that it wasted one year. It was pretty close to Bucharest, like 40 miles by train. I learned how to shoot and things like that, but in the mornings we also had classes to learn about electronic warfare and so forth.

CB: Did you enjoy it?

IS: No, of course not. At the end of the day, I could have done a lot of more useful things with that time.

CB: You completed your MS in Romania, right around the fall of the Ceaușescu regime in 1989.

IS: Yeah, it was right after the regime fell.

CB: What drove you to then pursue further studies in the United States? I understand you began at Old Dominion University in Virginia.

IS: I started at ODU because there was a Romanian faculty member there, his name was Stephan Olariu. At that time I didn’t know as much about the opportunities in Europe or in the US. After two years at ODU I decided to transfer to Carnegie Mellon University, where I finished my PhD.

CB: How did you formulate your thesis on quality of service in the internet?

IS: At the high level, I was interested in how to better manage resources and scheduling. Actually, at ODU I worked on this in the context of operating systems, and some of the stuff I did there, a scheduler for operating systems, is right now the default scheduler in Linux, as of, I think, one or two years ago. It’s called EEVDF, a pretty complicated name, not a good name, but anyway.

When I moved to CMU, that was ‘96. I changed my thesis a bit to focus on the internet: how to provide quality of service for, say, voice and audio on the internet. As you know, that was a golden age for the internet, when everything was happening. Everything was about the internet, like today is all about AI. Google was founded in ‘98, right? Amazon, I think, ‘96. So there was a lot of excitement. I focused on internet networking, and I graduated in 2000. After I graduated, I spent a few months at MIT, and then I joined Berkeley. I started to teach here, I think, January 2001, and I’ve been here since then.

CB: What drew you to Berkeley, and what makes it so special as a university?

IS: I was lucky and got quite a few offers when I graduated, and I came to Berkeley for a few reasons.

One, I liked the fact that it was on the West Coast, where everything was happening, at least in networking. Cisco was there, a bunch of other companies, startups, even AT&T and Bell Labs had research labs on this coast back then. I wanted to be close to where things were happening on the internet.

The other thing I liked about Berkeley, it did have a good balance, at least at that time, between people going to academia and going to industry or starting companies. When I looked, I liked that balance, and for a long time, even among my students here at Berkeley, half went to academia, at least 40, 50 percent. Now, for the past few years, things have been skewed toward industry, toward these labs. We can discuss that.

Then there are two other things I liked about Berkeley. One was open source. Berkeley was, in some sense, at the start of the battle of open source. Before Linux there was FreeBSD, the Berkeley Software Distribution. And again, remember that when I graduated, it was networking: a big part of the internet protocols, this TCP/IP, was developed at Berkeley, and Berkeley pioneered networking in the ‘80s and the beginning of the ‘90s.

And the final thing was more of an intangible. A lot of times, Berkeley was a pioneer in new domains and new technologies, the first among the big universities: Stanford, MIT, CMU, and so forth. They pushed on databases early on; Mike Stonebraker was here. I mentioned FreeBSD. Around that time, a little bit before I came here, they were doing this Network of Workstations project, which was about building big computers from commodity servers, as opposed to building supercomputers like Crays. This idea of connecting standard servers with fast networks to create a bigger computer is what became the foundation of all these big internet companies, including Google. They didn’t buy supercomputers, they just connected their commodity servers to create this huge compute and storage infrastructure. And there were sensor networks, also very early on. So I liked that pioneering aspect.

To summarize: I wanted to be where things were happening. They were saying at that time, well, we are close enough to Silicon Valley, but not that close. A lot of things happened in the South Bay, around Stanford, and I liked that balance, because I’m also an academic at heart. And I liked the open source, which was in Berkeley’s DNA, and finally that pioneering aspect, maybe more than at other top schools.

CB: During your first few years at Berkeley, what sort of research did you undertake? The first commercialization of your work was in 2006, with Conviva, right?

IS: When I graduated from CMU, I really wanted to make what I proposed in my thesis real. I spent six months or so trying, working with people, pushing on this standardization effort, because you need to standardize in order to be adopted in the internet. It’s very hard, right? There’s only one internet, so the barrier to adoption is pretty high. Anyway, I spent quite a bit of time pushing for the techniques I proposed in my thesis to make it into the internet, and it was very hard – now, looking back, for obvious reasons.

So after that, I started to move up the stack, like they say, where it’s a little bit easier to make an impact. Immediately after I graduated – I told you I spent a few months at MIT – I worked on what was then another hot topic: peer-to-peer networks. You know, it was Napster and Gnutella and so forth. How to make them much more efficient. I worked on that for a few years, then probably in 2006–2007 I moved toward big data, and around 2015–16 I started to work on AI and systems.

So if you want to look at my career: at ODU I did operating systems scheduling; that’s ‘94 to ‘96, something like that. At CMU, my PhD was networking: internet, quality of service, scheduling. Then peer-to-peer, 2001 to maybe 2004–2005. Then from 2006–2007 I did big data, and around 2015 I started on AI and systems. Of course, there is no very strict delineation, right? You start something new, and what you were doing is still going to continue, maybe even for some years.

The first company, Conviva, was about video distribution on the internet. It was based on the peer-to-peer technologies that we developed.

CB: What were the commercial assumptions which led to Conviva, and what mistakes were made, considering this was your first such effort?

IS: When you talk about using peer-to-peer to distribute video and audio and big files at that time, what was the main assumption? The main assumption was that the internet was going to be overwhelmed by this new traffic, and the idea of peer-to-peer is that, to avoid that, you try to localize the traffic at the edge of the network rather than having everything go through the core. It’s like in a city: you try to localize the traffic at the edges instead of having all the cars go through the center. That was one of the main ideas. The other one was about cost: the narrative was that because the internet was going to be very congested, it was also going to be very expensive.

Now, what happened after we started the company is that that assumption proved to be, I wouldn’t say not true, but not that strong. Actually, it turns out that because people laid so much fiber before the dot-com bust, there was enough fiber that wasn’t activated, dark fiber and so forth. So the internet scaled much better than many people expected, and hence the prices also went down. I remember when I started Conviva in 2006, the cost of downloading one gigabyte was something like 40 cents, and in two years it went down to two cents or something like that. Like 20x.

Obviously, when you build a company or you do anything, you try to solve a problem. You build a product to solve a problem, because that’s why people will buy your product. So if the problem disappears, or is no longer as big, then whatever you built is no longer as relevant, no longer as useful. So we pivoted Conviva, and Conviva is still running today; it’s cash-flow positive. Instead of targeting lowering the cost and scaling, we targeted quality: providing the highest quality of video distribution. And then we had these big customers like ESPN and HBO and Disney. NBC was another one, people starting to use the internet to distribute their content.

The lesson was that you’d better do something which is on the right side of the trends. There are these secular trends, and you’d better be on the right side of them, because there’s not much you can do if you are not.

The other thing I learned, and it remains with me to this day, is that consistency is more important than accuracy. We provided the users some metrics, some dashboards, about how their content distribution was doing: the number of users, how much they watched, and so forth. At some point we had a new release, and it was better. It was more accurate. But because it was more accurate, some of these numbers changed. And I remember what a difficult discussion it was with the customers, because you have these numbers changing: the number of simultaneous users, instead of, say, 10,000, is 9,500 or something like that, because you measure more accurately, and they are unhappy about that.

And it makes sense why they were unhappy: they had built their business processes based on these numbers. These are their metrics, which drive their businesses. So now, if these metrics change, you have a problem, right? So we were then trying to provide consistency with these previous numbers. That’s when you learn that consistency is far more important than accuracy. And it’s something general, I’m sure for every business, you have some metrics about success: impressions, clicks, whatever it is. If I measure it in a different way, even if it’s more accurate, but it gives you a lower number, you’ll not be happy, right?

CB: What would you say preempted your shift into data, around the time your lab began working on Spark?

IS: There are many things which coalesced into this, not in any particular order; I’m just going to order them for the narrative.

One thing that was happening: like I mentioned, at Conviva we were providing this dashboard for the customers, and the users could also query the database, to ask questions about the users, about failures in the video streaming, about the quality of the streaming for a particular group of users or particular content. And for that, we used traditional databases back then: MySQL. I remember this particular episode. You anticipate the growth. At that time the cloud was in its infancy, and we ordered a bunch of servers, the most powerful servers we could think of at that time, to run our database. By the time we got the servers, our traffic was already exceeding their capacity. And we struggled so much that it actually led to removing some features from the product; we couldn’t support them with the server capacity we had bought. So that was one angle. I had seen that problem, and I was thinking, there has to be a better solution here.

The other thing happened at Berkeley. We have these five-year labs, and in these labs there is a group of faculty, me included, and their students, working on particular projects. That lab was interdisciplinary: it had people from hardware, databases, systems, networking, and machine learning, very early on. That was around 2007–2008. These labs also worked very closely with industry. You have these retreats where you meet with industry, and at that time, if you remember, big data had already started to be a thing, and the open-source system out there was Hadoop, developed at Yahoo. We knew people there, so we were already talking in these meetings and retreats. We knew about all that.

Now, I told you the lab was interdisciplinary, and one of the groups from machine learning wanted to do the Netflix challenge: $1 million for whichever group could develop the best recommendations. Netflix provided some data to train your algorithms on, and then evaluated your algorithm to see which one was best. So a group from our lab, machine learning people, decided to compete. It’s a lot of data, so they came to us to ask what they should use to process it and implement their machine learning algorithms. And the only thing we had for them then was to use Hadoop.

Now, the problem is Hadoop was very slow, because Hadoop has these stages, and at each stage it reads the data from the disk and writes the data to the disk. It’s doing that also to be fault-tolerant: if something fails, you still have the data on the disk. And they came back and said, “Oh, but this is so slow,” because in machine learning you have many, many iterations to converge on a good model. Again, that was before; it was classic models: random forests, collaborative filtering, linear regression, things like that.

By the way, at that time we had already made some impact in the Hadoop community. One of the students was Matei Zaharia.

CB: Also Romanian, right?

IS: He was born in Romania, but he went to France, and he came here earlier in his life. He came to Berkeley from Waterloo to do networking. I think his first internship was at Facebook, and at that time, 2008, Facebook was small. Facebook was still in downtown Palo Alto, and they had their big data cluster: like 80 nodes, three people running them. And we solved some problems for Hadoop around scheduling. That’s why I’m saying we were plugged into the community pretty strongly.

So then these machine learning people from our lab come back saying, oh, this is too slow, and Matei put together a small system to address the main limitations of Hadoop at that time. And this was the genesis of Spark. This happened in 2009.

It’s complicated, but initially we didn’t think about Spark as being the big thing. Actually, we had another project called Mesos, which was about how you share a cluster between different large workloads. That was quite popular, and Spark was supposed to run on top of it, to demonstrate the generality and flexibility of Mesos.

Spark was a few hundred lines of code initially. In 2010 it was open-sourced, and then it started to get more and more traction. It was Matei and a great group of students, passionate about it, and it was the right time: people wanted, at least for some workloads, something faster than Hadoop. One of the reasons Spark was faster is that it was keeping as much data as possible in memory, so you don’t need to go to the disk. That made it faster both for doing multiple iterations, so for machine learning, and also for these interactive queries, because the data is in memory instead of on disk. As I mentioned to you, interactive queries was the problem I had with the databases at Conviva. So that’s the genesis.

CB: Did you enjoy your role as Databricks CEO, which you relinquished in 2016? Do you enjoy the business executive role, because you said your heart is in academia?

IS: Actually, things in industry and academia are not so different. In both cases, like I mentioned, you solve problems, right? That’s what you do. So choosing what problem to solve is important in both, and being on the right side of these secular trends. I think it’s quite similar, especially if you work on a larger project in academia.

Obviously, there are many differences as well. Academia is more of an exploration, with many things happening early on. What is the good thing about industry? You can put more and more resources behind an idea. You start with a few people, then tens, then hundreds, then thousands. In academia, that’s not possible. You can have a few people, maybe five, ten people at most, working on a project, a system, but no more. But you are going to have a lot more freedom to explore.

So, you asked me, I think I enjoyed it. There are big things I enjoyed. Obviously, you need to make certain bets, especially early on. And I’m always excited to recruit great people, this is what you do in academia too, you recruit students, faculty, and so forth. I really enjoyed that. And seeing something grow quickly is always exciting.

CB: How are Databricks and Anyscale, respectively, navigating the agentic AI wave?

IS: Databricks provides a broad data and AI platform. It enables enterprises to build and consume agentic applications. So, with the Lakehouse, Lakebase, and Unity Catalog, Databricks gives customers governed and secure access to their data, regardless of where it is stored. Then on top of that, it offers agentic products such as Agent Bricks, Genie, and Unity AI Gateway, which adds things like budget, governance, routing, guardrail, and usage-monitoring controls for AI services and agents. I would also mention Omnigent, an open-source meta-harness for AI agents that helps customers leverage and control existing agents such as Claude Code, Codex, and Cursor.

In contrast, Anyscale provides a lower-level platform and tools that help customers, including physical AI companies, build their own AI platforms, including agentic platforms. So, through open-source Ray, Anyscale helps customers scale complex AI workloads, including training, post-training, inference, and multimodal data processing. Many of today’s AI libraries use Ray to scale, including the vast majority of post-training and RL frameworks.

CB: How do you think government, academia, and industry can better coordinate to solve the problems emerging with AI? Is the cooperation efficient and effective at the moment?

IS: It’s absolutely insufficient. It’s important to look at the past. Look at the internet: it was the poster child of this triumvirate, a very strong collaboration between government, industry, and academia. The government, through DARPA, started the internet program. Academia did the first deployments. Then, when the internet was, so to speak, privatized in ‘89, the industry, AT&T and the big companies laid fiber, built and expanded the internet. Even the standards were mixed academia and industry. The standards body was the IETF, the Internet Engineering Task Force. Unlike other domains, which are very industry-heavy.

Compare that with what happens today with AI. It’s almost the opposite, absolutely the opposite. First, let me talk about the industry. Right now you have all these big labs, which are very much siloed, and very little publication; there is not much information they are going to share publicly, and everyone is doing its own stuff, pretty secretive. So it’s not this free exchange of information which happened in the past, in the case of the internet. That’s one.

The other thing is that the cost is so high to build these models and run them at a certain scale that, again, there are no strong relationships between academia and industry. And the government, if anything, should allocate much more money than in the past, because this is a much more expensive enterprise. So this worries me: you do not have this fundamental diffusion of innovation, of knowledge. If these labs don’t publish, if they don’t share, that knowledge is not available for others to build on top of. And you can see maybe a little bit of what happens: the best open-source models are now coming from China, because, ironically, they are more open than the US.

Also, think about the human capital. If AI is so important, so important for our society, for our civilization, maybe, and of course there are problems, there are challenges, you want your best minds to work to address these problems, whatever they are: security, the impact on jobs, whatever it is. But in order for people to work on them, they have to have shared artifacts they can work with, and shared knowledge. And now this doesn’t happen, because each lab is doing more or less the same thing. It’s a very inefficient use of human capital. The only silver lining is that in California, at least, non-competes are not legal, and people can move. So your diffusion of innovation is people moving between the labs.

CB: You were talking at the beginning of this conversation about the best minds flocking to industry at the expense of academia.

IS: This is again something worrisome. So let’s put it this way: you can understand why this is happening. You are doing research in AI, you are in academia, and one of these companies approaches you and basically says: “Hey, if you join me, you have access to resources you don’t even dream of having in academia.” So this is number one: you can do a lot more interesting research here than in academia. Number two, you can get lots of money: seven, eight figures, even nine figures. And on top of that, they are going to say AGI will come in two years, so if you do not come now, then it’s too late.

And in some cases, people say: oh, if AGI comes in two years, you will not have anything to do anyway; your value will not be there. So come now, do it, be part of it, and get lots of money. It’s a very effective sell. And it’s easy for academics to rationalize: I am going to do this because I am going to do much more research and have more impact than I would in academia. It’s now or never. And in two years, when everything is done and AGI arrives, maybe I’ll come back. We’ll see what happens.

And this is effective, even for students. So that’s what happens. It has never happened before, at least from what I’ve seen.

CB: Which of your recent research efforts–I’m thinking of vLLM, Vicuna, MemGPT, and so on–and which of the emerging research fields are the most exciting for you personally? And which has the most potential, in your view?

IS: I think there are two trends happening. We were talking about trends.

One, on the systems side, is basically all these layers collapsing. It used to be that you had different layers: networking, and even networking has different layers, the operating system, the application, and so forth. What layering gives you is this modularity in building a system: it makes it easier to develop, and you can develop one component without affecting the rest. It’s like in a car: I put in a better engine, the rest remains the same, but it’s going to be faster. However, what you lose, you lose in terms of efficiency. If you build a truck, you can go to the parts bins and pick the engine, pick the wheels, all these standardized parts which are already there. But if you build a race car, you integrate everything. The same here: you have to integrate everything, because every inefficiency is so expensive. It costs a billion dollars to train a model; if a modular architecture impacts the efficiency by 10 percent, that’s a hundred million, right?

So that’s the collapsing, and the stack becomes so much more complicated to build, and things are changing fast, faster and faster. Building this stack for AI is an order of magnitude harder than building the stack for previous generations, like the web or big data. That’s one.

The other one, which I think is very important: when you talk about AI right now, what is the hardest part? The hardest part is not about capabilities. We know that AI has fantastic capabilities. It’s about deploying them in production reliability, predictability: I think these are extremely important aspects. How do you do that? And this will become more important, because AI is about trust. If you use AI, you need to trust it to do the right thing. Otherwise, you wouldn’t use it, at least for critical tasks, which are the most valuable tasks.

On that side, I’m excited about Arena AI, about methods to evaluate these models and applications, and how to understand what the gaps are: where they are doing well, where they are not doing well, and things like that. I think this is very important. And the other thing is inference, because there are three major workloads: training, inference, and this post-training, like reinforcement learning and things like that.

CB: If you were starting your PhD over today, what specifically would you choose to focus on?

IS: Good question. I’ll focus on what I’m actually working on today. There are two main directions.

The first is making AI reliable: ensuring that its outputs are predictable and aligned with the user’s intent. This will be critical for AI to reach its full potential and to be used in mission-critical, high-value applications.

The second is using AI to build the AI stack itself, which is growing increasingly complex. Far more complex than previous technology stacks, such as big data or web infrastructure. In other words, using AI to build better AI.

About the Author

Carson Becker is an American writer. He is on X @carsonjbecker

Copyright © 2026 Intergalactic Media Corporation of America - All rights reserved

Copyright © 2026 Intergalactic Media Corporation of America - All rights reserved

Copyright © 2026

Intergalactic Media Corporation of America

All rights reserved