Ion Stoica is a Romanian-American computer scientist and Professor at the University of California, Berkeley. As cofounder of the unicorns Databricks and Anyscale, Stoica is among the most prolific academic-entrepreneurs of the 21st century.
Born in 1965, he was raised in Nicolae Ceaușescu’s Romania, where he served an obligatory nine months in the armed forces. He completed an MS in computer science at the Polytechnic University of Bucharest in 1989, a year of profound change across Europe. As the Warsaw Pact disintegrated, Romania’s particularly harsh totalitarian regime collapsed under the weight of its own contradictions. A combination of extreme austerity and social repression, combined with Ceaușescu’s distinctive policy of self-reliance, had eroded the regime’s reputation across all levels of society. In December, Ceaușescu was overthrown and executed, and Romania began its westward transition. Not long after the revolution, Stoica departed his homeland for the United States.
Stoica has had an enormous impact on computer systems and networking beginning with his early work as a graduate student at Old Dominion University, a public university in Norfolk, Virginia. At ODU, he devised a process-scheduling algorithm that was later adopted as the default scheduler in the Linux kernel. After moving to Carnegie Mellon to complete his PhD in electrical and computer engineering, and with a brief stopover at MIT, Stoica began teaching at Berkeley in 2001. In 2006, he launched Conviva, a streaming analytics company that was born out of his academic research into video streaming over the internet.
After launching Conviva, Stoica’s research became increasingly focused on big data. In 2009, machine learning researchers at Berkeley’s interdisciplinary AMPLab found that Hadoop, which was then the dominant framework for large-scale data processing, was hopelessly slow for algorithms that needed to iterate over the same dataset many times. In response to this, Stoica’s doctoral student Matei Zaharia built a compact system that kept data in memory rather than repeatedly writing it to disk. Open-sourced in 2010, that system, Apache Spark, has become one of the most influential pieces of data infrastructure ever created.
Databricks was founded in 2013 by Stoica, Ali Ghodsi, and five PhD students including Zaharia. Stoica served as the company’s first CEO from 2013 to 2016, when he became executive chairman. In early August, Databricks closed a $5 billion strategic funding round at a $190 billion valuation. In 2019, Stoica co-founded Anyscale, built around Ray, a distributed computing framework developed in his lab to overcome structural limits in Spark itself. Ray now undergirds much of the AI industry’s training and inference workloads. Meanwhile, Anyscale, where Stoica also serves as executive chairman, has reached a private market value of over $1 billion.
As director of Berkeley’s Sky Computing Lab, Stoica is currently preoccupied by questions surrounding the reliability of AI and the mounting complexity of the AI stack itself. Some of his recent research projects have included vLLM, Vicuna, MemGPT, and the model-evaluation platform Arena AI (formerly LMArena), alongside work on inference. Professor Stoica continues to teach Berkeley graduate and undergraduate students.
I sat down with him to understand his journey from communist Romania to the pinnacle of American academia and business, the enormous commercial successes of his research, and his perspective on AI. What follows is a transcript of our conversation.
***
Carson Becker: Tell me about your family and your upbringing in Bucharest, Romania.
Ion Stoica: I was born in Bucharest. Both of my parents were engineers. My father was a geophysicist. His work entailed looking under the surface of the earth, trying to figure out where to find oil and other things like that. My mother was a geologist, so pretty related. My grandparents were out in the countryside, so I was spending, in general, the summers and some of my vacations there.
CB: What was your family’s experience during the many upheavals of the 20th century: the world wars and so on?
IS: My grandparents on my mother’s side were pretty close to Bucharest, like 50 miles. They had a tougher time because of collectivization: the state taking their land and putting it together in a cooperative, as it was called. So that was tougher. Actually, my mother, because my grandparents had some land and so forth, was excluded for two years from college because of that.
My grandparents on my father’s side were a little farther away, near Târgoviște, a city which was the capital a long time ago, and it was a little bit in the hills. The communists did not take the land there, because it was more for growing apple trees and things like that. You couldn’t grow grain. So they still had some land and were doing relatively well compared to others.
My family experienced the Second World War, and I just heard stories, also a little from the First World War. Moldavia was part of Romania, and just before the Second World War, Russia invaded and took it away. That was one of the reasons Romania was initially fighting on the side of Germany: to free that territory. In 1944 there were changes in alliances. But really, I think that from all sides of my family, it is obviously a story that the communists didn’t have a positive impact on them, because fundamentally they took away property and land and things like that.
CB: What was your education like, and how did the communist regime perform in that regard?
IS: They actually did invest in education. Let me take two steps back. Romania had been more or less a constitutional monarchy since around 1850, and it was a pretty democratic country. There was this reform giving land to the peasants after the First World War. Economically, Romania was pretty good, they were manufacturing trains, airplanes, and so forth, kind of in the middle of the European economic rankings. After communism, it was pretty much at the bottom. What I’m trying to say is that communism was not good for the economy in general.
Now, I obviously cannot compare education before and after, but I do think that when I grew up, the education up to college, and maybe including college a bit, was very solid. A lot of math, a lot of STEM, as you call it here. And you had to learn two languages: one starting in the first grade, the other in the fourth or fifth grade. Unfortunately, in some cases one of them was Russian.
They put a lot of value on being good at school and being educated: getting into college and so forth. This was the state, like other communist countries. Romania, like Russia, wanted to use education and scientific progress as a tool of propaganda, to demonstrate the superiority of the system. This is what Russia did after the Second World War, and they were pretty successful for a while.
CB: What was your experience in the Romanian military?
IS: I did nine months. It was a unit that was not really for combat, it was related to electronic warfare: using radar and sensors for discovering the enemy. It wasn’t that bad, except that it wasted one year. It was pretty close to Bucharest, like 40 miles by train. I learned how to shoot and things like that, but in the mornings we also had classes to learn about electronic warfare and so forth.
CB: Did you enjoy it?
IS: No, of course not. At the end of the day, I could have done a lot of more useful things with that time.
CB: You completed your MS in Romania, right around the fall of the Ceaușescu regime in 1989.
IS: Yeah, it was right after the regime fell.
CB: What drove you to then pursue further studies in the United States? I understand you began at Old Dominion University in Virginia.
IS: I started at ODU because there was a Romanian faculty member there, his name was Stephan Olariu. At that time I didn’t know as much about the opportunities in Europe or in the US. After two years at ODU I decided to transfer to Carnegie Mellon University, where I finished my PhD.
CB: How did you formulate your thesis on quality of service in the internet?
IS: At the high level, I was interested in how to better manage resources and scheduling. Actually, at ODU I worked on this in the context of operating systems, and some of the stuff I did there, a scheduler for operating systems, is right now the default scheduler in Linux, as of, I think, one or two years ago. It’s called EEVDF, a pretty complicated name, not a good name, but anyway.
When I moved to CMU, that was ‘96. I changed my thesis a bit to focus on the internet: how to provide quality of service for, say, voice and audio on the internet. As you know, that was a golden age for the internet, when everything was happening. Everything was about the internet, like today is all about AI. Google was founded in ‘98, right? Amazon, I think, ‘96. So there was a lot of excitement. I focused on internet networking, and I graduated in 2000. After I graduated, I spent a few months at MIT, and then I joined Berkeley. I started to teach here, I think, January 2001, and I’ve been here since then.
CB: What drew you to Berkeley, and what makes it so special as a university?
IS: I was lucky and got quite a few offers when I graduated, and I came to Berkeley for a few reasons.
One, I liked the fact that it was on the West Coast, where everything was happening, at least in networking. Cisco was there, a bunch of other companies, startups, even AT&T and Bell Labs had research labs on this coast back then. I wanted to be close to where things were happening on the internet.
The other thing I liked about Berkeley, it did have a good balance, at least at that time, between people going to academia and going to industry or starting companies. When I looked, I liked that balance, and for a long time, even among my students here at Berkeley, half went to academia, at least 40, 50 percent. Now, for the past few years, things have been skewed toward industry, toward these labs. We can discuss that.
Then there are two other things I liked about Berkeley. One was open source. Berkeley was, in some sense, at the start of the battle of open source. Before Linux there was FreeBSD, the Berkeley Software Distribution. And again, remember that when I graduated, it was networking: a big part of the internet protocols, this TCP/IP, was developed at Berkeley, and Berkeley pioneered networking in the ‘80s and the beginning of the ‘90s.
And the final thing was more of an intangible. A lot of times, Berkeley was a pioneer in new domains and new technologies, the first among the big universities: Stanford, MIT, CMU, and so forth. They pushed on databases early on; Mike Stonebraker was here. I mentioned FreeBSD. Around that time, a little bit before I came here, they were doing this Network of Workstations project, which was about building big computers from commodity servers, as opposed to building supercomputers like Crays. This idea of connecting standard servers with fast networks to create a bigger computer is what became the foundation of all these big internet companies, including Google. They didn’t buy supercomputers, they just connected their commodity servers to create this huge compute and storage infrastructure. And there were sensor networks, also very early on. So I liked that pioneering aspect.
To summarize: I wanted to be where things were happening. They were saying at that time, well, we are close enough to Silicon Valley, but not that close. A lot of things happened in the South Bay, around Stanford, and I liked that balance, because I’m also an academic at heart. And I liked the open source, which was in Berkeley’s DNA, and finally that pioneering aspect, maybe more than at other top schools.
CB: During your first few years at Berkeley, what sort of research did you undertake? The first commercialization of your work was in 2006, with Conviva, right?
IS: When I graduated from CMU, I really wanted to make what I proposed in my thesis real. I spent six months or so trying, working with people, pushing on this standardization effort, because you need to standardize in order to be adopted in the internet. It’s very hard, right? There’s only one internet, so the barrier to adoption is pretty high. Anyway, I spent quite a bit of time pushing for the techniques I proposed in my thesis to make it into the internet, and it was very hard – now, looking back, for obvious reasons.
So after that, I started to move up the stack, like they say, where it’s a little bit easier to make an impact. Immediately after I graduated – I told you I spent a few months at MIT – I worked on what was then another hot topic: peer-to-peer networks. You know, it was Napster and Gnutella and so forth. How to make them much more efficient. I worked on that for a few years, then probably in 2006–2007 I moved toward big data, and around 2015–16 I started to work on AI and systems.
So if you want to look at my career: at ODU I did operating systems scheduling; that’s ‘94 to ‘96, something like that. At CMU, my PhD was networking: internet, quality of service, scheduling. Then peer-to-peer, 2001 to maybe 2004–2005. Then from 2006–2007 I did big data, and around 2015 I started on AI and systems. Of course, there is no very strict delineation, right? You start something new, and what you were doing is still going to continue, maybe even for some years.
The first company, Conviva, was about video distribution on the internet. It was based on the peer-to-peer technologies that we developed.
CB: What were the commercial assumptions which led to Conviva, and what mistakes were made, considering this was your first such effort?
IS: When you talk about using peer-to-peer to distribute video and audio and big files at that time, what was the main assumption? The main assumption was that the internet was going to be overwhelmed by this new traffic, and the idea of peer-to-peer is that, to avoid that, you try to localize the traffic at the edge of the network rather than having everything go through the core. It’s like in a city: you try to localize the traffic at the edges instead of having all the cars go through the center. That was one of the main ideas. The other one was about cost: the narrative was that because the internet was going to be very congested, it was also going to be very expensive.
Now, what happened after we started the company is that that assumption proved to be, I wouldn’t say not true, but not that strong. Actually, it turns out that because people laid so much fiber before the dot-com bust, there was enough fiber that wasn’t activated, dark fiber and so forth. So the internet scaled much better than many people expected, and hence the prices also went down. I remember when I started Conviva in 2006, the cost of downloading one gigabyte was something like 40 cents, and in two years it went down to two cents or something like that. Like 20x.
Obviously, when you build a company or you do anything, you try to solve a problem. You build a product to solve a problem, because that’s why people will buy your product. So if the problem disappears, or is no longer as big, then whatever you built is no longer as relevant, no longer as useful. So we pivoted Conviva, and Conviva is still running today; it’s cash-flow positive. Instead of targeting lowering the cost and scaling, we targeted quality: providing the highest quality of video distribution. And then we had these big customers like ESPN and HBO and Disney. NBC was another one, people starting to use the internet to distribute their content.
The lesson was that you’d better do something which is on the right side of the trends. There are these secular trends, and you’d better be on the right side of them, because there’s not much you can do if you are not.
The other thing I learned, and it remains with me to this day, is that consistency is more important than accuracy. We provided the users some metrics, some dashboards, about how their content distribution was doing: the number of users, how much they watched, and so forth. At some point we had a new release, and it was better. It was more accurate. But because it was more accurate, some of these numbers changed. And I remember what a difficult discussion it was with the customers, because you have these numbers changing: the number of simultaneous users, instead of, say, 10,000, is 9,500 or something like that, because you measure more accurately, and they are unhappy about that.
And it makes sense why they were unhappy: they had built their business processes based on these numbers. These are their metrics, which drive their businesses. So now, if these metrics change, you have a problem, right? So we were then trying to provide consistency with these previous numbers. That’s when you learn that consistency is far more important than accuracy. And it’s something general, I’m sure for every business, you have some metrics about success: impressions, clicks, whatever it is. If I measure it in a different way, even if it’s more accurate, but it gives you a lower number, you’ll not be happy, right?