# 60 Billion Predictions Daily: Inside Credit Karma’s Agentic Data Layer (/blog/60-billion-predictions-daily-inside-credit-karmas-agentic-data-layer)
What does MLOps look like when you are deploying 22,000 models a month?
Maddie Daianu, Head of Data and AI at Credit Karma, joins the Data Bros to pull back the curtain on one of the most high-volume data environments in FinTech. With a 100-person team serving 140 million members, standard data practices break down.
Maddie shares how her team manages terabytes of daily data on Google Cloud and explains the massive strategic pivot they are undertaking right now: The move from "Information" to "Agency."
Listen on [Spotify](https://tinyurl.com/m77csepd) or [Apple Podcasts](https://tinyurl.com/2zsvmmjs)
\[00:00:01] Benjamin: Alright. Hello, everyone, and welcome back to the Data Engineering Show. Today, we're super happy to have Maddie on. Maddie kind of works at Credit Karma. She leads data and AI there. Super great to have you, on the show. Do you wanna quickly introduce yourself, kind of give us some background on what you're working on, yeah, and how you got into data?
\[00:00:23] Maddie: Of course. Thank you so much for having me. It's a pleasure to be here today. Yes. I am Maddie, and I lead data and AI at Credit Karma Intuit. I've been at the company for about two and a half years. Here, the team for data and AI consists of AI scientists, machine learning engineers, data practitioners, including data engineers, assist intelligence engineers, and the experimentation platform. And this team has been at the genesis of how we engage our members across Credit Karma, which is essentially a financial app that helps you not only advance your credit score but also find a marketplace of very useful and relevant highly personalized offers for credit cards, personal loans, or anything that you might be shopping for. And, it's enabling to you to advance that financial journey that you're on. So for that, the key elements and ingredients of making this app successful is data and ai. So I'm very very excited to like be telling you a bit more about how we build the team and really what are some of the competitive advantages but also challenges that that we are seeing. a bit of background about me before being at Credit Karma, I was at smaller companies, but also at Meta at some point. And then my first job out of academia was at Intuit. I kind of came back full circle. Before Intuit, I was an academic working on biomedical engineering, using at the time machine learning to extract biomarkers of disease for Alzheimer's. I really wanted to essentially hopefully see if I can change the path of alzheimer's in my lifetime. I learned that that might not be the case, so I switched into industry where you can see a faster pace of innovation. ultimately though what I'm grounded by is a very consumer facing mission. So I love helping users like yourselves um see the benefit of technology.
\[00:02:28] Benjamin: That's amazing and like what a wide range of things, seriously. I mean like going kind of from biotech, kind of smaller companies, bigger companies. Very, very cool. So maybe before we dive into the data and tech stack and kind of your teams, we have a lot of listeners from kind of across the globe. I'm sure many of our US listeners or virtually everyone, will know about Credit Karma and also Intuit. Do we maybe wanna give, like, a one minute kind of background for every kind of listener who has not heard of these companies in the past? kind of, yeah, what they're all about, kind of what what your products enable.
\[00:03:03] Maddie: Yeah. Absolutely. So maybe starting with Credit Karma, which is, again, a financial product that enables you to check your credit score for free. And that was the genesis of Credit Karma from maybe more than ten years ago, and our founders really wanted to lean in to help you find your your way to advancing your your financial progress and making this easier without having to pay for information that can be readily available at your fingertips. Now, of course, credit scores, you can find them on a wide range of apps. So we've evolved that journey of our products to fully personalizing with highly relevant machine learning models that enable you to make the right decisions given what you need in your financial journey. Say that you might want to change or evolve your credit score, you might need very different types of support either consolidating your debt or finding ways to save money that are different from people who might have a very different financial goal. So that's credit card bond. Now we are part of intuit. So the beauty of of uh this uh essentially journey that we are on is that we can combine credit scores from credit karma with people who want to do their taxes on turbo tax. So turbo tax is also a very successful product that has been providing a lot of customers across the globe, but especially in The US, the ability to do your taxes also in a very personalized way. So think of it this way, credit card is almost a front door to turbo tax, and one thing that we are doing uh at intuit right now as part of what we call this consumer ecosystem is building synergies and personalization and seamlessness across credit karma and turbo tax such that we can fully understand you as a user as you traverse these products. And that is really what's behind what I believe is one of our biggest competitive advantages as a company.
\[00:04:56] Benjamin: Awesome. Cool. So let's let's start chatting about the data stack. Right? Just kind of like, how big are your teams? How are maybe they structured? What types of data do you bring in? Kind of what pieces of software do you have in your tech stack? We'd love to learn more.
\[00:05:12] Maddie: Yeah. Absolutely. So my team specifically is about a 100 people, but we do have, over 800 engineers that are part of Credit Karma. So it's very, very engineering heavy. and we essentially build a, engineering first mindset from the genesis of the company to how we operate our teams. Specifically for my team, as mentioned, we have AI scientists, which are about a third of my team, and then we have machine learning, engineers as well as data experts who build a platform that enables us to run, the training for machine learning, but also facilitate the deployment of those. And then I have a third of the team that is more data practitioners for business intelligence as well as the experimentation platform. Credit Karma is very specifically a Google Cloud, infrastructure. They transitioned to this a while back. So think of it this way, BigQuery is our data warehouse, we use Bigtable for our operational serving layer, and we ingest large information from the financial institutions that we work with, from our partners that we work with every single day. so we have and we process and transform multiple terabytes of information daily for our 140,000,000 members every single day. So like that's a very exciting large data set to like have our hands on and be able to, provide personalized information for. So Yeah.
\[00:06:40] Eldad: Of course The only problem is BigQuery. The only problem here, like, everything is perfect. The stack is amazing. Everyone, like, participants, everything. But just BigQuery, come on. We need more. More spice. Can't have BigQuery solve all problems. Just kidding. Go on.
\[00:06:57] Maddie: It's not. It's it's certainly not solving all problems. It's something that we have learned to essentially, ensure we build holistically for for the company. Nonetheless, we are, very much focused on using Google as as an infrastructure. Having said that, we also are focused on multi cloud integrations because Intuit is an AWS shop. So we are essentially, when you think about the TurboTax products, but also like a wider range of products that Intuit has, they are in AWS. So we have to be oftentimes cloud agnostic to be able to facilitate, the integration of our cloud products and serve our members effectively across, end to end consumer ecosystems. Now
\[00:07:43] Benjamin: And then looking like, one thing I'm curious about looking at kind of your utilization of BigQuery is most of your analytics, like, internal and, like, kind of batch based. So let's say every six hours, you refresh kind of credit scores and do some predictions, and then you expose it back to customer. Or there's also that kind of customer facing, almost like real time analytics challenge, where if I'm a customer of Credit Karma or kind of other Intuit products, I log on to the app, kind of go on to the website, and can kind of consume my data kind of directly in a live way with, like, kind of customized like, kind of take us through that internal versus external part of your data, data stack.
\[00:08:25] Maddie: Yeah. So it's both. We have both streaming and batch pipelines depending on, the financial institutions that we serve or depending on the data sources. The patterns are are different. Let's say that you as a user log into Credit Karma and your credit score has changed. As soon as we learn that the credit score has changed, we update that in as freshly as possible. So you can see the most recent information within the app versus waiting for, say, a batch pipeline that might take weeks or months to update. So when it comes to that, we are trying to be very, timely. We're providing you information that is relevant. Similarly, if our partners change, say, certain aspects about their offerings and APR changes or similar, we also update that information as quickly as we learn so that users, again, get access to the latest, more import most important information to make the best decision for themselves.
\[00:09:19] Benjamin: And then that last mile customer facing journey is also done on Big Query, or you basically start exporting into some other systems at that point?
\[00:09:29] Maddie: So for instance, one of the Bigtable is where we use, as our operational serving layer. And also, we have something called, Alchemy, which is our feature online feature store, which also uses, big BigQuery and Vertex AI so that we can do transformations in real time and aggregations of information, and enable you to, like, see, derivations of the data, in real time within the app. And more importantly, what we use this for is ultimately fueling our machine learning models. So on top of our data, which I believe is not nearly as powerful if you don't really have that AI intelligence layer, we have our models that essentially, lead to almost 80,000,000,000 daily predictions for our 140,000,000 member base every single day. We have more than 22,000 model deploys every single month, so we can refresh and personalize the models so that we can meet our members where they're at. So for that also we use Alchemy, which is again our feature store to be able, to personalize that experience.
\[00:10:38] Benjamin: Okay. Super super impressive scaling and a very, very nice stack. If you I'm curious about the ML side of things, because I assume that kind of here, like, what types of obviously, I'm sure you can't go into all the details, but, like, what types of models do you actually kind of train? Right? Because especially
\[00:10:55] Eldad: in the
\[00:10:55] Benjamin: finance space, things like it's, like, heavily regulated, of course, kind of there's, kind of explainability, fairness of models, and all of that really matters. so it would be super interesting to learn about some of your experience there.
\[00:11:09] Maddie: Absolutely. So let me break that down into ways. One is, like, the machine learning stack, which is our bread and butter and tradition that would be traditional AI. And then we can talk about gen AI too, which I think is is very relevant and also important to discuss. So for machine learning and traditional AI, our bread and butter is recommendation models. And that includes essentially ranking, and personalization of, say, an offer either within the app or a marketing channel. So we have hundreds of those models that enable us to ultimately deliver the right experience for our users. And within those, of course, we have to be very cognizant of information that we can and cannot show based on compliance as we are in a highly regulated space. Now having said that, where compliance, for instance, especially as it pertains to our partners, comes into play is when it comes to gen ai, which, unlike traditional ai is non deterministic. Right? So oftentimes what we do and we've seen a lot of success with especially at product karma is using traditional ml and pairing that with GenAI for contextualization explainability purposes. So when we show you an offer for, say, a credit card or a loan, we have something called c y, which explains why are we showing you this particular offer such that we can build trust with the user but also educate on whether or not they should consider opening or taking that particular offer. Now when it comes to essentially say features like cy that contains more context about the offer, we have to be very thoughtful about the benefits that oftentimes pertain to our partners that they are represented with a 100% accuracy. So that's where we invest tremendous effort into evaluation platforms so that we make sure that everything is a 100% accurate before we, roll it out to users.
\[00:13:02] Benjamin: Awesome. Super interesting. And then if if you look ahead, like, kind of both on the data and then also the ML side, so you already mentioned you folks are kind of pulling the products closer together. It's like if you think about 2026, like, what are the challenges you're most excited about, the things you really wanna drive? yeah. It would be interesting to hear that.
\[00:13:21] Maddie: Yeah. So I will say that one of the biggest things that we are tackling right now especially more than two years into adopting gen AI and seeing some great success in contextualization, we want to take this to the next level. So Intuit as a whole believes and I truly believe this is tremendously valuable in creating done for you experiences for our users. And what that really means is starting to like create and and um alleviate tasks on our users behalf. Last week we announced something called the debt agent which is helping you, consolidate that or or find options to manage that debt and taking some actions on your behalf. Now none of this honestly is easily done without having and building an agentic data layer. So that's one of the biggest, I think, revelations that not only we are seeing as a company, but also many other other companies are seeing where if you don't structure your data in a semantically, well structured way, you are not likely able to provide the most highly relevant and personalized experiences for users. So, especially because we are part of Intuit and we are part of a broader consumer ecosystem group. One thing that we've been building, in the last year or so it's called the unified consumer profile. And we started this as part of a hackathon. We love doing hackathons here at least twice a year where we really wanted to get to the bottom of how can we create a semantic graph that depicts a financial journey for a user in almost think of it as a tree format, such that entities like debt can be well explained by attributes like balance or interest, and being able to essentially contextualize that and serve it in a highly accurate, real time and consistent way across Credit Karma and TurboTax. And that is ultimately what I believe is one of the biggest unlocks to being able to serve a very personalized experience, at least from a user facing perspective when it comes to GenAI, but also for a wide range of other products. AI in general, but also other product capabilities. So I believe, like, if we don't invest in agentic data layers, may that be user facing or maybe context facing, in session facing, then we will not really be able to truly unlock what we can do, especially with generative AI that contains either small or large language large language models that really need that contextualization to be able to be relevant.
\[00:15:52] Benjamin: Right. Okay. And then kind of that basically agentic layer, what is the tech stack for that kind of a may I ask? Because that's obviously an across your product line. Yeah.
\[00:16:03] Eldad:
\[00:16:04] Benjamin: so it would be interesting. Yeah.
\[00:16:06] Maddie: Absolutely. So that is, ultimately based in AWS, and it is something that we are building once more in as possible a cloud agnostic way, but it is part of the broader Intuit stack, which as mentioned earlier is, AWS based. So we are utilizing a lot of the capabilities that Intune has put in place to be able to like ingest the data into AWS to be able to serve the data in real time as both Credit Karma, which is again on Google, and TurboTax, which is on AWS, ingest or serve this information in real time.
\[00:16:44] Benjamin: Awesome. Cool. Eldon, any questions from your side? Anything you're curious about?
\[00:16:50] Eldad: I think it's crazy how, you know, Intuit I know Intuit. Most of us know Intuit back in the days. Right? Like, so there are people who used to manage their taxes on the Intuit platform, switching products, kind of evolving with Intuit. And now kind of AI is kind of gets context out of all of those system into one place and then getting crazy amazing experience no matter which app you're using. Like, to me, Intuit was always about the software, but, apparently, it's not. So it's all now about AI and the data. And and and Intuit knows us. Right? Like, if you've been running your taxes on Intuit for so long, then you probably get a much better experience continuing to use it because they have the right data. So it's super interesting to see kind of how those platforms evolve and and kind of how crazy value you're getting out of them.
\[00:17:47] Maddie: I'm glad that you're you're saying that, and and thank you for being an Intuit user. Indeed, it's been a tremendous journey, and I've seen this for many years. I was back at Intuit more than seventy years ago, and, now here again, one of the core components of what you're describing is something called the generative AI operating system, which Intuit has been investing in tremendously over the last, few years. That also helps fuel and democratize how we use Gen AI across Intuit. So Credit Karma or TurboTax or any other subsidiary of Intuit can tap into this centralized platform and really democratize and adopt Gen AI at scale within its product. And those are the types of examples really that enable us to move fast and continuously disrupt ourselves, especially in the age of AI.
\[00:18:41] Eldad: So, Manny, Benjamin, as we all know now, he just recently moved to the Bay Area. And coming out of Munich, he is not acquainted with the whole world of taxes, tech management, and tax education. So I hope that with AI, Benjamin will get a very different experiences and kind of new onboarded user into the Intuit, into the Intuit, platform, Benjamin.
\[00:19:08] Maddie: That's the goal. I hope I hope, Benjamin, you you use Credit Karma to build your credit score and credit history. And and use that as an entry point into TurboTax where we can serve you no matter what the complexity of your tax background is. So, hopefully, you take advantage of that.
\[00:19:28] Benjamin: Sounds good. I'll let I'll let you know how it goes. And, yeah, kind of excited to learn about this this new world. It's like moving to The US kind of, like, coming from, like, a just different country. Right? Like, it's a lot to take in in the beginning and kind of a lot of things you need to learn about, kind of a lot of new new systems you need to understand. so, yeah, it's gonna be a fun time basically onboarding into all of that.
\[00:19:52] Eldad: I think Now when you look to into it, you're having a conversation. Think about it. Like, you previously, you had to learn how to use the UI. Like, ten years ago, like, that was the challenge on on Intuit was protecting you from writing the wrong stuff into your tax process journey. Now it's planning stuff. It's thinking for you. It's it's finessing all of those rough edges and not like, that no one likes to do. So I think, like, this is one of those, like, the auto zoom is completely broken, by the way. It's a new webcam. As you notice, it's still in experimentation mode. let's zoom out. This is the wrong AI. By the way, this is AI driven, the Zoom. and now I've switched to manual mode. I apologize. so many, it's it's fascinating to hear about it, and, we look forward to see where and how it evolves. And it's crazy to see so many engineers are working fearlessly behind the scenes to get that painful process so smooth. What can I say? Well done.
\[00:20:58] Maddie: Thank you so much. I'm glad, and thank you for the kind words. And Benjamin, let us know how we can help you in your financial journey and tax journey. And it's a pleasure, speaking with you today.
\[00:21:09] Benjamin: Awesome. Thank you. It was great having you on the show. Kind of thank you so much. all the best and looking forward to staying in touch. Thank you. Awesome.
\[00:21:18] Eldad: Take care. Bye.
# A ClickHouse Review from a Practitioner’s Point of View (/blog/a-clickhouse-review-from-a-practitioners-point-of-view)
Sudeep Kumar, Principal Engineer at Salesforce is a ClickHouse fan. He considers the shift to Clickhouse as one of his biggest accomplishments during his eBay days and walks Boaz through his experience with the platform. How on one hand it handled 2B events per minute, but also how it required rollups which compromised granularity when extending time windows. Besides a ClickHouse review from a practitioner's point of view, Sudeep tells us about interesting use-cases he's working on at Salesforce.
Listen on [Spotify](https://open.spotify.com/episode/03LovJ2BxShGsE6Lspy4Ag) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/a-clickhouse-review-from-a-practitioners-point-of-view/id1561927688?i=1000578021028).
**Boaz:** Hello, everybody. Welcome to another episode of a Data Engineering Show. Today with me is, Sudeep Kumar.
Hi, Sudeep. How are you?
**Sudeep:** Hey, I'm doing great. How are you?
**Boaz:** Very good. Very good. You're in Austin, correct?
**Sudeep:** Yes and it's very hot here, right now.
**Boaz:** I'm in Tel Aviv. It's also very hot indeed here. We are at around 30 Celsius, which should be, let's see 86 Fahrenheit.
**Sudeep:** Oh, we at 103.
**Boaz:** Ah, you win! But we have higher humidity which sucks.
**Sudeep:** Really okay. Here it's like 40% right now, maybe it's probably higher. In the studio, I just checked it up here.
**Boaz:** Yeah, here it's higher. You go out, you start sweating.
**Sudeep:** Oh, okay. Probably that gives you more feeling of heat, right?
**Boaz:** Yeah. Unfortunately, not a dried desert heat, but more of annoying city heat.
**Sudeep:** Right.
**Boaz:** Okay, so Sudeep thank you for joining us. Sudeep is a Principal Engineer at Salesforce and has been there for almost a year. Before that spent a long career and especially many years at eBay, doing a variety of engineering and data-related projects and roles. We will spend some time asking you a lot of questions about that. So, are you ready?
**Sudeep:** Yeah! Yes, let's do it.
**Boaz:** Okay. Let's start, Tell us a little bit about your background, how did your career over the years and what kind of data-related things your career end up letting you do?
**Sudeep:** Sure. Actually, I started my career within the telecom domain and this was back in my country, India. And at that time, telecom was like really booming into around like 2006 around that time. That was the most happening place to be. I spent a couple of years there and then I immediately saw a shift towards mobile. Everybody started just working on mobile, mobile applications, developing around that, wanted to be part of that ecosystem.
So, I joined the largest player out there within the mobile, Realm in India, which was Samsung at that time. And I joined Samsung and it got involved there and that's probably where I got exposed to the first real background on the data volumes specifically because we were building applications around social networking, very similar to Facebook that we had today. So, this was like in 2008 around that time. And then I was there for like around four and four and a half years, and immediately saw the shift of the market and the interest towards E-Commerce and around that and I wanted to be in that area then, and I just made that conscious switch from moving from mobile to E-Commerce. And, that's how I ended up on eBay. And I was there for eBay for 9 years. Primarily, within the platform engineering team where we were handling huge volumes of structured, and unstructured data, specifically more on telemetry data. So, we are handling events, logs, and metrics at scale. I moved to the US around 6 years back. I had an opportunity to do that and joined a similar team here, working on things like structured events or creating a platform around OLAP, so worked there for a while. And, then created a metrics platform, which can handle like 20-30 million data points per second. Further to that, I looked at distributed tracing there and got a similar opportunity in Salesforce. And, then I just made the switch at that point in time. It was a long stint there anyway. So, it was a good time to switch. During all this, I moved to Austin. I was in Bay Area initially but then moved to Austin a couple of years back when the pandemic started, before Elon Musk, for sure.
**Boaz:** So, at 9 years you were afraid of hitting the 10-year mark at eBay and decided let's try something new.
**Sudeep:** Right!
**Boaz:** You got scared by saying, staying at 10 years at the same spot. I know that thing that I had earlier in my career as well.
**Sudeep:** Yeah. It's like, everything is so comfortable and then you just want to switch around things.
**Boaz:** Yeah. So, let's talk about those days at eBay. Sounds super interesting. I mean, yes, starting from the early E-Commerce days, eBay probably was one of the first companies that really tackled a lot of data, really at the center of the big data game. Tell us a little bit more about that data, that platform engineering team, which you were a part of, sort of, how big was it? What kind of teams did it have? What kind of data-related projects were you in charge of?
**Sudeep:** Right. The larger umbrella was the data platform team. But then within it, we were specifically looking into monitoring and telemetry platforms, essentially they also handle large volumes of structure and structure data. And within that particular charter of monitoring or observability in general, we had different pillars. We had a pillar for metrics, another one for logs and events and then we have something even alerting, dashboarding, everything. So, they were not really mutually exclusive from each other. There was a lot of overlap, the people are moving around here within this charter. So, my primary focus was on the backend. So, I was completely on the backend side. I had exposure towards this logs, metrics and events in general. And, then towards the end, I did also look into distributed tracing. The team was divided. We had a team in China and we had a team in the US. We like fairly, evenly distributed around. I think we were around 30 folks, overall if you really look at it. I don't recall the exact number, but somewhere around there, overall.
**Boaz:** During those years what did the data stack look like and what part did you have a chance to work on?
**Sudeep:** Yeah. So, initially, I started with C++ there. We had something called a Cal publisher, which accepts all your log data. That was kind of what we were using and internally, behind it had Hadoop, where we were storing the data and so forth. Slowly, we moved on to Java and as part of Java, specifically Flink, we had a lot of streaming jobs that we are writing on Flink to emit useful metrics signals on this particular log data or different kinds of metrics even for doing some aggregations. So around that, and then slowly, when we started looking at or working with Prometheus for metrics, I think the entire team kind of shifted to Golang and we've been like a big-time Golang advocate for the past few years now. But I think, if anything you ask any of our engineers within eBay, specifically the data platform team to write any solution, they would just go to Golang and write something up.
**Boaz:** Interesting! What other major shifts throughout the years changed, I mean, sort of, were there any big projects that replaced Legacy and modernization? What were you part of?
**Sudeep:** Yeah. So, we did have a Legacy logging system, which was not very real-time, it was more from a batch kind of mode. You run a job, find out your matching logs and from there more real-time logs we moved to. There was a contribution that I played there, but more recently the one that I can think of is we completely shifted our OLAP use case from Druid onto ClickHouse, which happened a couple of years back. That was probably one of the bigger events that we did, from what I remember. For metrics also, we moved from HBase to Prometheus. We created our own distributed architecture on Prometheus score and that platform I still pretty well do. So, we say some of the things that over the period that I've seen change.
**Boaz:** So, the OLAP workloads also fell under the same group, under the platform engineering group. So, what kind of OLAP used cases, who are the end users?
**Sudeep:** SRE was a primary end user for most of these OLAP use cases, but there were also some alerting that people had on top of this particular data. There are no business-critical alerts or data there specifically. The way this OLAP data was generated was through logs. So, our logs were pretty structured. So, there was a component that would look into all these logs and emit these as events, like a much more scrape-down version of it. And, then we would write onto some kind of OLAP store, which is Druid in our case. We had a Kafka in between then finally just reached to Druid and that's kind of our...
**Boaz:** What data volumes were pushed to Druid and ClickHouse?
**Sudeep:** Right. I think, I had written a blog around this, and we had a few talks also around this, but from what I remember when we were doing Druid in a minute, at that time, right and the volume had definitely increased from them. It was, I think around 250 million events per minute. But, I think more recently on the newer platform, we were looking upwards of 2 billion events per minute.
**Boaz:** How big was the time window that was accessible?
**Sudeep:** So, you can go up till a year back, but then the data will be rolled up like we just keep rolling up the data. The granularity will be lost, but you can essentially access data up to a year.
**Boaz:** So, all the orchestration and how is the ETL done, the aggregation? What was you around Druid and ClickHouse?
**Sudeep:** Yeah. So, when these raw events were coming in, we really didn't do a lot of, like that was an entry point for us, but we did some level of transformations in terms of, I think removing some PCI data and things around that. But, other than that I think we also structured the data as part of the transformation that we did and essentially rolled them into raw tables initially. And, on these raw tables, you had more aggregations built on top of it. Like, there were other tables that would do aggregations on top of this data and roll it up further. And, that's kind of the model that we followed. So, so those rolled up would happen for every application, for every region. So, those were kind of rolled-ups that we essentially did and this was similar.
**Boaz:** Are we talking On-Premise or are we talking Cloud SaaS?
**Sudeep:** This was On-Premises. So, we have a Managed Kubernetes platform that was hosted On-Prem and all our workloads run over there.
**Boaz:** Data lake, how's the raw data stored? What is being used?
**Sudeep:** For the raw data store? It's the same, we used ClickHouse for the raw data store, for the OLAP, right, specifically.
**Boaz:** No, outside the OLAP.
**Sudeep:** Oh, the raw data, so Hadoop. For the same data, the log data, which will go onto Hadoop. We had a little bit of data, I think also written on to Elastic for some of this past search use cases, I guess.
**Boaz:** Over the years, was there any attempt or push to, go to AWS or something like everyone else?
**Sudeep:** So, eBay and Amazon probably are competitors, right?
So, probably I think AWS may not be a good fit, but I think there was some talk about doing a hybrid and many companies are trying to go that route. But I think, more or less the encouragement from our management has been to run on On-Prem. I'm not sure what is the motivations around that. There's probably more to it than I have visibility into, but I have a feeling that there is something more there but I'm sure I remember the recall that we did explore other clouds like a GCP and all, we did explore it.
**Boaz:** Today, at Salesforce, are you guys doing Premise work or do you do cloud work?
**Sudeep:** At Salesforce, I think it's publicly known that we do, sort of, AWS, so I think that's mostly it's there, but similar to many other companies, we have our own infrastructure as well. But, mostly these days, we are moving towards everything on the cloud.
**Boaz:** What do you miss from On-Premise the most?
**Sudeep:** I think, one thing, I loved about On-Premise, is you are much more control over things. You can SSH and all that with not hampering a lot of security, rules, or policies on that, that was convenient. But, I'm guessing that those things are very important, right? When you get things too broad. So, there is like a Pro and Con there, but definitely, for a developer, like for me to be able to develop things, roll out, test it out, I think, it was much easier to do on a machine where I had complete control on.
**Boaz:** That's an interesting trade-off. For more and more people, it's something they don't even have a chance in their career to try out anymore.
**Sudeep:** That's true.
**Boaz:** I appreciate the differences. Druid where you transitioned to ClickHouse, you guys were probably on Druid for a long time, what was the tipping point or the reasoning to try and replace or try another platform?
**Sudeep:** We've been running with Druid for quite a bit time, and we were very successful running it for like, I think, over like 3 years or so. But, then volumes fell low. What we did see that year over year, the growth was increasing and the amount of resources that we had to add to support something for Druid was also substantial and it was showing up on the cost.
Another one that we saw is that our availability for Druid on our infrastructure was not very reliable. It was very flaky. It had to do a lot with how our infrastructure is hosted. So, not really like, specific to Druid, but we have things that we have done to make things a little harder for systems like Druid to work, very nicely on our infra. So, we started hitting like our DevOps folks, those guys just started getting a lot of pages and probability issues. So, it was just becoming a huge problem. We wanted to look at alternatives, but we were not really seriously looking at them. We had basically tried to add more infrastructure and scale horizontally. But, at some point in time in 2019, I guess, somewhere around that, we heard of this data sort called ClickHouse. It was not very popular then, like, as it is today. I still remember in DB rankings, it was showing up like 180 or somewhere there. It was not popular at all by any standards.
**Boaz:** It was a hidden gem.
**Sudeep:** It was a hidden gem. Yes, that's probably the right way to put it. But, we just tried it out. There was another architect of ours and we just tried it out. We just liked trying out whatever new comes, and we just ran it. Since we were having all this load, also come to Kafka. We had separate consumers running there for this particular, same data pipeline. It was just something very basic for ClickHouse and started writing. And we didn't expect much. Like we just thought, okay, let's see what happens. And, I think we did not see any incidents. It was just running without any issues. And, meanwhile, our current existing platform started having a lot more issues and then it forces us to actually look deep into ClickHouse more because it just was eating up any kind of volume that you really throw at it, which was just very impressive for us, especially on the ingestion side. So, we investigated more, and we saw that this potentially could replace our OLAP pipeline and then, we put in our engineering effort around it. And, I think that really was very fruitful and then, we saved quite a bit of, like substantial infrastructure costs, like around 90% from what I recall.
**Boaz:** Wow.
**Sudeep:** Yeah, that was pretty substantial costs that we saved.
**Boaz:** Because ClickHouse had proven to be more hardware efficient.
**Sudeep:** Yeah, it was more hardware efficient when it comes to really ingestion. And for the search and things around that, we had to put a lot more effort around Druid to make sure that it scales. For ClickHouse, we really didn't have to do that much and that is pretty impressive about ClickHouse.
**Boaz:** Tell me more about the search use case. What kind of queries you're trying to run there?
**Sudeep:** There are queries that we run for SRE specifically, wherein we say that application health-related data, that we get from the log, so it says that how many errors did you see? How many URL counts you have? What has been the application heartbeat, in general? like how many heartbeats you have seen? and how much volumes you have handled, in terms of like, overall? What is a distribution for different kind of HTTP statuses that you have for your application? So, things around that was actually being rendered visually using this particular data and SRE was a big time, like power users of this particular data.
**Boaz:** You guys also do or doing like Joins in the queries or was it all sort of one big factor?
**Sudeep:** Yeah. There was Joins. I don't really recall, what kind of Joins we were really doing but there...
**Boaz:** Because Druid and Joins were never good friends, that became easier in ClickHouse.
**Sudeep:** Yeah, you're right. I do remember we had some Joins across different tables, but I just don't remember what top of my head, what the use case was.
**Boaz:** Did you use that opportunity to also do a proper evaluation and try other tools head to head that ClickHouse came out the winner or was it just ClickHouse, sort of, suddenly came in from the corner and surprised everybody for the better and that was that?
**Sudeep:** Yeah, I think, we were trying a few options, but nothing really was serious. Even ClickHouse was not really serious for us. But, ClickHouse probably made the cut just because with minimal effort, we could see a lot more benefits on our pipeline. So that's the only reason probably we looked a little more deeper into it. I do remember that we looked at things around whether we can write this data offline onto Hadoop itself and have some jobs run there, but it's not really very real-time.
**Boaz:** You guys, I guess, being On-Premise, you self-managed and ran the open source versions of both Druid and ClickHouse.
**Sudeep:** Yes.
**Boaz:** In eBay, did your teams also contribute to their projects along the way or sort of go into the code and then add functionality?
**Sudeep:** We had some contributions on the ClickHouse operator, for the open source operator, from our team. But, we were quite active in the community. We were very active with folks who are actually very involved with ClickHouse as well. And, as part of that, probably I'm guessing some of those requirements that we raised probably translated into features in that product and that community was not that big at the time. It is probably much bigger right now. But, those days people generally knew each other by name and it was very, very small community. And, the way I used to meet is just hop on a call with them and then, they were much more accessible at that point.
**Boaz:** Yeah. Full disclosure, we don't talk about Firebolt in this podcasts typically, but Firebolt is based on part of it. The core execution engine was a hard fork from ClickHouse back then. ClickHouse is essentially amazing when it comes to speed and performance at scale, in sort of, part of the journey for us in Firebolt was how could we adapt it to the modern data stack and then to assess experience and so forth, but absolutely the performance and the scale is top notch.
**Sudeep:** Pretty impressive, right.
**Boaz:** Yeah, and tell me about the metrics platform that you mentioned.
**Sudeep:** What we essentially ended up doing is that we saw the advantage of using Prometheus, and it was integrated very nicely with Kubernetes and we had our own HBase-based solution for metrics, which was fine. But again, it had availability challenges and it was not very rich in feature, as much as we would've liked and some of the things around data locality and all that was very much well addressed in the Kubernetes world. So, we looked at whether we can use Prometheus core and then create a distributed architecture on top of it to be able to scale. So, essentially what we did is, we took the Prometheus core and we created an abstraction layer on top of it, which is essentially we call it like a shard and you can have multiple shards within the cluster and for things around replications and all was taken care by dual rights. And, when you egress a query, you essentially look at things around the health of the shard based on the data points for a particular tenant, how healthy that was from that shard and we would cater to those query based on that.
**Boaz:** Awesome! What are the projects do you remember the most as the ones you enjoyed the most in the last decade working with all these things you've worked on?
**Sudeep:** The one that probably I enjoyed the most was actually on a problem of called discovery of data within large volume sets. And that was pretty interesting because we had the challenge of uniquely identifying different metrics and within metric identifying different keys and values for it.
**Boaz:** We're talking about the eBay days, alright.
**Sudeep:** Yeah, the eBay days will give more context. So, we had a problem around topology discovery for metrics and at large, and these were like huge volumes of data, right? So, essentially those used to take a lot of time when you really run at the backend to find out what are, give me all the metrics, for each metric give me all the key and values for it and specifically when the...
**Boaz:** Sorry to question in the middle, but the topology for metrics, can you elaborate on that a little bit? So, what's the business scenario there? What we're trying to look into?
**Sudeep:** Imagine a case where a user is coming onto your metrics platform and he starts first with a metric name. So, he starts typing in, say, metric name called, say, my tenant dot CPU, something like that, and then you should see all the metrics associated with that particular expression, and then he selects it. It should be like a very interactive kind of mode there.
**Boaz:** And, that's a homegrown eBay metric store.
**Sudeep:** Yes. Right. So, this particular system was created in parallel to the metric store. So, your metric store has all this raw time series data, but the topology itself for each of these metrics, like the metric name and key and value pairs used to reside in a different data store. We used to use Elastic for that but then we had a different data store for it and all that, the metadata discovery on the, give me the metric name, give me the key and values those used to come from Elastics. So, initially what we tried is to save all that from the raw store but that was not very performant. And, we looked at what we can do to make it much more efficient because those need to be super fast. Users cannot really wait for those interactions to complete. So, we wanted to give a very delightful experience there and that's the first kind of interaction that users would have on our platform, start searching. And since we were like the system which is getting all this metric data in one, like single sync, it became more important for us to be able to do that faster.
**Boaz:** Got it! Awesome! We also sometimes like to ask about bad memories. So, give us, sort of, regretful story, something that you learned from over mistake, that we can teach others to avoid. Tell us a horror story.
**Sudeep:** Probably one of the horror stories I can remember is on the distributed tracing. So, when we really look at tracing, you can create a platform, you can create everything around it, but when it comes to adoption for something like tracing, you need end-to-end visibility. So that means, like any hop of your request that goes to say service one to service two to service three, needs instrumentation done on all three services. And, that was a challenge that we probably underestimated. So, what happened, we created the platform. We created everything. It works very nicely, but then if one of the service owners says that I am not ready for the instrument. I'm not ready to do it. So, you had the entire chain break. And then the experience really becomes not that great. You'll see different traces, which are not even like matching with each other. So, that experience was a problem. So, one thing I realized from that is that maybe a good idea to always check how feasible is it also for people to adopt it, like speed, right? And, I think that's something definitely was a little bit of pain also I felt to make people adopt.
**Boaz:** So, the program was discovered way too late.
**Sudeep:** Yeah, it is like, like you were so happy with this particular system that you built and you're thinking this is going to change the world. But then you find that, okay not everybody is excited about this, as you are excited about the same.
**Boaz:** Which is interesting. Something we hear about like in classic product management. How can we keep the end users close to what we're building? And often times in deep engineering projects that can get lost.
**Sudeep:** Yeah, exactly.
**Boaz:** But, for data projects and I myself, I'm a product guy, dealing with data my entire career, but a product guy and you see more and more positions to data that are called a data product manager because there's so much data activity and data projects going on in the software world that, there's this specialty in being a product manager for data related projects only. And, we need to put the end users into the planning ahead of time, even if those are sort of our internal users or analysts or whatever or engineers.
**Sudeep:** Yeah, that's a fair point. Yeah, I agree.
**Boaz:** How different is the data world at Salesforce compared to eBay?
**Sudeep:** The problems are similar. I guess so there are a lot of similarities, like in the concepts and everything, but, each company has, I think, a different way to go about things in terms of execution. So, there are pros and cons, in both places. I think, for me it was more like lift and shift, really to be honest because some of these terms and constructs and even technology stack was all familiar, I guess. So, for me, it didn't really take long and I don't feel a lot of difference within the two.
**Boaz:** What stack are you dealing with now at Salesforce?
**Sudeep:** It's Java, a little bit of Golang and Elastic search, Hadoop, I think a lot of OLAP, ClickHouse is not yet there, not yet. So, let me see what they can do.
**Boaz:** So one is being used for OLAP.
**Sudeep:** I think for OLAP, we do have something like Druid. I'm not so sure actually, because I'm still new in this particular charter, but I've heard that there are events pinned around, I think around Druid, but I'm not sure.
**Boaz:** It's probably a little bit of everything.
**Sudeep:** It's a little bit of everything. This team is also pretty big, and being remote and not being able to like interact with the team closely, I think some of the things I don't have visibility into directly, but, right now, we are focused on distributed tracing, me and our team.
**Boaz:** And that's for yourself, you've gotten used to the desert heat in Austin or are you planning to move back to the Bay Area, now that the pandemic is behind us? I missed the weather. I really do miss the weather in Bay Area and the food in Bay Area also is good. People say food in Texas, is also great but not to my liking probably. The Bay Area food is also good. So, I do miss these two things about Bay Area, but Bay Area is super expensive. And that's probably one of the reasons I moved in the first place, especially if you want to raise a family and you need more space and everything. Texas is pretty accommodating.
**Boaz:** Yeah, well, Sudeep it's been super, super interesting.
**Sudeep:** Same here.
**Boaz:** Thanks for sharing your journey and your experience with these technologies and the data challenges you've done and yeah, that's it, super interesting!
**Sudeep:** Thank you. Thank you for your time. All right.
**Boaz:** Thank you so much.
# A Deep Dive into Slack's Data Architecture (/blog/a-deep-dive-into-slacks-data-architecture)
Growing from a startup to an IPOed and then an acquired company meant that Slack's sales org was scaling rapidly.
Apun Hiran, Slack's Director of Software Engineering explains how the data stack and architecture evolved to support this growth with more reliable and timely metrics.
Listen on [Apple Podcasts](https://podcasts.apple.com/us/podcast/a-deep-dive-into-slacks-data-architecture/id1561927688?i=1000560434630) or [Spotify](https://open.spotify.com/episode/1CMbAfbkK2FTKMGDDW7ET4)
Transcript:
Boaz: Thank you everybody for joining us here in San Francisco. We are super excited to be here. We are Eldad.
Eldad: Hi! everyone.
Boaz: I am Boaz, from Firebolt. Some background for those of you who have not heard of the podcast. What is this event? The backstory is this. This guy here loves data, decided to found the data company called Firebolt. It is pretty good. Check us out. But it is not really about Firebolt. To avoid that, marketing came up with a good idea. Let's do podcasts. We started doing this Data Engineering Show podcast, which is all about bringing great people from the industry, data practitioners, people who work with data on daily basis and just interrogate to help them about their day-to-day, their challenges, their visionaries, their points of happiness, and everything data related.
Eldad: Do not cut anything out today, all the embarrassing moments that we have on LinkedIn.
Boaz: Yeah. You cannot just skip your speaker today and it counts everything is there to say. The podcast took off. We had a good time doing it, have a lot of listeners and then, we decided why not take it on the road. So, this is the first attempt to do the podcast on tour. We are very excited to be in San Francisco during this for the first time. We have great speakers lined up. So, are you ready?
Eldad: I am ready.
Boaz: The first speaker to join us, is Apun Hiran. Come on Apun! Where are you? See how he comes.
Apun: First time doing this.
Eldad: He is doing it every week.
Apun: Hot seat. It is really warm.
Boaz: Apun, we all worked from home for a couple of years. All of us tried a variety of zoom backgrounds. Some of us went for animated backgrounds. Some of us went for weird backgrounds. Apun sat very close to his huge Jeep car which was his zoom background in the garage. The best background I have ever seen.
Apun: Yeah! That is my favorite place in the house. Quiet. My garage is where I worked from home.
Boaz: And he spotless the car also, shiny.
Apun: You need to keep it like that. It is going to be the background.
Eldad: We have known each other for so long. We started to know each other in Covid, and it is the first time we are meeting in person. So there are quite a few people here that met us when we started. So, just want to quickly say "Thank You" for believing in us and being with us, first when we started.
Eldad: It is great to physically see a lot of you here.
Boaz: Apun, is the software engineering director at Slack, heavy, heavy focus on data. He has been on Slack for around a half-year, spent a few years prior at Dropbox, doing data stuff, and also had a long career behind him, in data, Yahoo, and Oracle. He even used to have a DBA title, but that seems to be disappearing.
Eldad: You want to add further things to this.
Boaz: Things you would like to say about yourself that I did not include?
Apun: Yeah. Sure. My name is Apun Hiran and I have been in the data space for about 20 plus years now. I started my career as a DBA right out of college. The first thing I was doing was Oracle DB and did that for a very long time, almost 10-12 years before moving into the big data space and writing fix fixed scripts and then did a stint in a company called AppDynamics, doing software engineering. That was the first time writing code in Java for about two and a half years, building that database monitoring products, and then moved into Dropbox for about two and a half years doing data engineering. And, the last five years for me have been more around functional data engineering, primarily working with business stakeholders, finance, sales, and marketing and that has been an interesting change for me, moving from the platform side of things and building large databases or Hadoop platforms to doing functional data engineering.
Boaz: Awesome! Let us start with, you are in Slack. First day in the office. What is your impression of what is going on in the data and Slack?
Apun: I think, my first impressions of data on Slack were pretty good. I was pretty excited to see the data platform that was already available in place and the way it was scaling and the team was amazing that I was going to lead the company. We already had a pretty good data stack, and I think my job was much easier in Slack than before. From there, we are just trying to get to the next level.
Boaz: You are saying you worked very hard before?
Apun: Yes, it was.
Apun: At Dropbox, when I joined, me and the team, I build out a data stack, analytics data stack over there.
Eldad: Found Dropbox.
Apun: Yeah. We did not get our files on DropBox, but we used to ingest them, but that was it.
Apun: Everybody is amazing. The technology behind it is really impressive. I think we did load data from Dropbox. There were folx who dropped data at DropBox and we would pull the data into the system.
Eldad: But it is hard to see if you are moving back to Dropbox where you went to release a provider on Dropbox needs to cold scan.
Boaz: You brought into the Slack into which position?
Apun: Yes, I joined Slack a year and a half and my role was leading the data engineering team which is the functional data engineering team. Basically, looking at the sales, marketing, finance, customer success, customer experience, HR, all the different analytics that you can do, other than just product and that was my goal, and I had a wonderful team to lead there and we have been building a lot of data products over there and since the last couple of months, my role as a standard.
Boaz: Most speakers on our podcast, get to hold a pretty good team.
Apun: Yeah, now I also lead the data platform team and the enterprise integration team at Slack
Boaz: Tell us about the various teams that do the data at Slack?
Apun: Sure. At Slack, we have a pretty large, self-service data platform based on Hive and Presto where anybody at Slack has access, who can write sequel queries can go there and write that. But, we have our data engineering team which looks purely around the product data, the usage metrics, and stuff like that. And, then we have a data science team that works very closely with the product data engineering team, looking at all the new features that are being released or beta releases, or A/B testing and things like that. And then, we have business analysts spread across all the different business units who are pretty SQL savvy and logged into the system, write their own queries, and sometimes make our lives difficult because we fix those queries. So, we have a lot of data folks. I am a pretty robust sense of platform at Slack.
Boaz: During those years and a half, when you joined and since, what were the priorities at Slack?
Apun: I think the biggest priority for us was to have reliable metrics as Slack sales were growing rapidly and to support the sales organization with reliable metrics, friendly metrics that were like the first...
Boaz: We all give a contribution to the sales.
Eldad: Allowed the free versions.
Apun: There were pretty lofty goals, right from stewards and all the sales organization around growing sales that are at a very fast rate. So, the first thing that we came in is to build away from a robust sales analytics system and build in all the important metrics around and the user experience that come with those metrics. So, that was probably the first biggest challenge that happened. But for us, the other challenge that happened was the acquisitions. So, then there was a second tier of things how do you look at metrics Salesforce space that we can align on metrics. Some of those things that come with accuracy have become pretty critical. At the same time, we were also looking at some gaps in the platform that we had in terms of, not having a very robust data catalog and having proper data quality tools. So, we were having a lot of issues with data quality and we would usually get notified by the end-users that this does not look right, which is...
Boaz: What was a typical example?
Apun: Typical example will be data load that is say coming from CRM application and we load the data and during the data load, a few of the files got missed because of some error during the copy command that happened and by the time that goes beyond call and had an alerting system, it has already been published. The dashboard has already been refreshed and then, there were other data issues with lack of data. A job was supposed to run in 2 hours and for 24 hours and the end-user does not know why there is no data or the data is half picked.
Boaz: How did you go about fixing that?
Apun: We use Slack a lot for alerting, prioritizing, and all that, so we made Slack in the middle of everything that we do like all the alerts started to go down, go through the channel.
Eldad: Who gave you this idea by the way.
Apun: We definitely leveraged Slack and a lot of workflow in Slack to be proactive in messaging that there are delays and things like that. One thing was just around building the ecosystem alerting, better monitoring, and better on-call support for all of that stuff. The other was looking at what is the root cause of these issues? Can you be more proactive in publishing data only after the data loads are completed and we know that there are no errors? We are thinking of the entire ETL pipeline. I must say we are in a much, much better situation right now, and we are at the place where we are actually trying to identify the data issues, which are not easily related which are actual data issues, somebody has the resources from data has problems, building alerts from that. We have come a long way, in one year.
Boaz: Most of that approach to the data engineering team, your team?
Apun: That's correct. There were functional data issues. So, we had to work with functional teams. We need to fix the data upstream. We did identify. So, we became like that team which was identifying issues from source systems and reporting them as well from just being the folks that we get the alerts through our ETL.
Boaz: Tell us about the various stack?
Apun: As I mentioned, we have a data lake, which is on S3 with Rust environment and where people can write their own queries, and publish dashboards as well.
Boaz: Which kind of users typically run their own queries there?
Apun: Almost all users run their queries because that is the data lake, you get all of the regular data available over there. But I would say data science is probably the biggest use case. Almost everybody in the company who needs any kind of product information would go there and look at the data.
Boaz: They are analysts or people embedded within the different departments?
Apun: Yes, they are. Every team product, product analyst, market analyst, and sales analyst are all pretty people savvy. They would go around queries. At times, we have written queries for them and we have given them the lake and they just run it every now and then to get their reports or numbers as well.
Boaz: You mentioned before the lack of data cataloging in a company, like Slack, and analysts embedded across departments. How does knowledge sharing work? How do they know what statics to look at? How about sharing work?
Apun: I mean, that is a problem. It has always been a problem. There is a nice search and people have keywords, and they find a table which looks like the right thing to look at and they write queries and they look at data, and then the data would not match to publish that. Then, they will come back to some data engineering team and why is this component not looking the same and we will go ahead and tell them what is the right table to look at. That is a problem with any self-serve platform that is there. But, there has been a lot of investment made on the metadata management part, on the Presto and Hive as well and people are putting in comments and putting in column definitions, paper definitions, and creating the right schema where you have conformed data stack. Somebody is monitoring those datasets. Those are some of the changes that have happened to the self-serve side of the data analytics department.
Boaz: What data volumes are you dealing with?
Apun: I do not have the number. The whole data warehouse, we have got 30 petabytes in size as such. I do not have the number like what's the ingestion every day but you can imagine every message is captured and put somewhere else, which is more secure, but the event that you know, who sends a message to somebody is an event and we want to see how many messages you sent and how many messages that Firebolt sending. So, all of that comes, there are a lot of messages.
Eldad: This guy is counting your messages.
Apun: And your emojis too!!
Boaz: What do not you do something more popular data sets, those aggregated...?
Apun: Of course. So, all the user-related metrics are daily aggregated. The company, account metrics, company metrics like Firebolt, what is at the company level, how many messages, how many weekly active users, and monthly active users. Those are aggregated on a daily basis.
Boaz: Tell us a little bit about the place for those jobs?
Apun: We use Airflow as our orchestration platform. Again, airflow, and then we create ETL jobs over there. So, basically, that will be the data engineering teams, which are creating these ETL jobs and publishing them as confirmed data sets on the self-serve. That is one part of our data infrastructure, which caters to quite set product-related metrics and then there is the other part, which is all of the business-facing data sets, dashboards in which we use Snowflake Matillion as an ETL platform, and Looker and Tableau, as the BI platform. So, that's all, there be like the functional data engineering happens, maybe over 100 different data sources at this point of time and bringing them all in, that does not happen to the second part of the platform.
Boaz: As for the data warehouse with Snowflake is the data engineering team solely in charge of everything that is going on there or is there also a level of self-serve that people can use Snowflake on their own?
Apun: I think when we initially started, it was pretty held in place by the data engineering team, but as the use cases, as a lot of people wanted to ingest Excel spreadsheets to do analytics and stuff, so we tried to create a separate environment in Snowflake which is self-serve, which there is the infrastructure around, you have a Google sheet and you want to load that and do some analytics, and you can do that by providing some metadata in spreadsheets. So we do that as well.
Eldad: What happens when someone edits a message? Do you keep the same ID? Do you run different dbt scripts that update every day?
Apun: I have no idea, Eldad.
Eldad: Is someone concerned now that Twitter is going to release an edit feature on Twitter, is that something the board is discussing?
Apun: I am sure a lot of people have that them on their minds since yesterday.
Boaz: How layman should write the thing in Slack as you already mentioned, Presto, Snowflake, BI tools like Looker and Tableau? How does that go along with testing?
Apun: You are talking about like a new feature that is being introduced. Slack, the product engineering team, they have its own roadmap in terms of what product features it would want to introduce. There is a component, whether the data teams are involved in terms of what kind of metrics usage information we would want to derive from a particular feature. It is a more collective thing. You have folks from sales, like how do you want to sell a product, involved? And, what kind of metrics would need from that product involved in this kind of discussion before something goes into production. There is a metaphase where people are testing, AB testing with few customers in more comfort and once it goes into production, it is more like collective agreement, you fill those data sizes, and then it propagates all the way for analytics as well.
Boaz: Can you give a recent example of the process to support the current feature?
Apun: Of course. I think Slack introduced the feature of Huddle and Slack Connect, not too long, I think a year and a half ago. Those two features were built particularly that way in terms of because Slack Connect is one of those features that you can connect with anybody who is on the Slack platform as such beyond your organization. And, it was one of the very critical features that was released by Slack. So, we wanted to make sure that we can track the usage, and how much time people are spending. How are permissions and approvals and all of that stuff is being taken care of? And, that was a collective discussion between even with sales, like, how do you want to sell this feature? And, how will you track that and the same with Huddle? Huddle is another most popular new feature of Slack in terms of people using it and we can see the usage goes on every day. All those features would have that way.
Boaz: You guys played the full game and the metrics are agreed upon in advance?
Apun: It is not as perfect as it sounds, but there are different definitely a recreation of that, but we do go with requirements what would you need at bare minimum to launch something like what kind of data requirements are? Because we are talking about software engineers and data engineers, software engineers are building this particular product feature. They think differently sometimes in terms of what is critical from like metrics perspective or logging perspective, from what data engineers and business things. So it is a good idea to come together.
Boaz: Is there any collaboration between software engineering and data engineering, can you please tell me that?
Apun: I have no idea like that. I have not been just there for a year and a half. I definitely see that we work together a lot, even from as simple as you have to update the website with certain information, which is critical to marketing. We do work very closely with the web team - How do you want to present this? How will I get that information from the logs? Where will it arrive and how can I push it to the marketing team and sales? We do a lot of collaboration that way, but not like on the product as such.
Boaz: What is like a nightmare situation you were in? What is the worst complaint you got into the data engineering department? This does not work. This is wrong. What is a horror story?
Apun: One horror story for me was just coming from a different company to Slack and figuring out that there are no emails. You do not get any emails, everything is on Slack, and then slowly in the first week, you are added to some 200 channels across seven workspaces and it is like a total nightmare. You do not know where to look forward.
Boaz: It is a horror story.
Apun: It was like very difficult for me for the first month and a half. I think it is still difficult on days, but from a data perspective, I think we also do a lot of bots, like the Slack bot that I mentioned where we publish data to execs on a daily basis of the company performance from both the sales and personal usage metrics and that is like a very critical bot and the other bot, marketing bot, sales bot, but that's like an exec bot is very critical.
Eldad: You are sending that over an email.
Apun: No, all Slack, everything at Slack. That is something that everybody has eyes on at 9:00 AM like it is published at 9:00 AM. And those numbers are very critical. Everybody is looking at it.
Boaz: Specific that.
Apun: Yeah, specifically. That particular thing, has broken at certain times, which is very critical where everybody is from the CTO, the CIO or the CEO, everybody is pinging why do numbers were wrong? Why these numbers are not printing correctly and then. If it is a simple data issue, it is fine. But if it is an infrastructure issue, it becomes a nightmare. So, that was one of them that happened around publishing business metrics. That is something I check everyday morning.
Boaz: How do you guys go about if something like that happens? How do you go about learning from your mistakes?
Apun: We do maintain a lot of documentation, purely from run books and from an on-call perspective, and certain critical pipelines like this, as you mentioned, execs bots have pretty solid documentation and as part of all our stories, we have one story that we created around documentation. Make sure you are updating those books for documentation. So, it has been through like brown bag sessions, especially something like this kind of visible event happens, we make sure that people are on that. We fit with the whole team, both offshore, and onshore teams to educate, focus on what exactly happened and record those sessions, document them, and then also update on books.
Boaz: What are your plans now for the future? So 12 months forward, what do you want to achieve in data?
Apun: Right now, if you look at it, it is a pretty simple data platform from my perspective. You have a data warehouse, you have a BI tool, you have an ETL platform, but then there are a lot of questions around like, what do you have in your database warehouse like the catalog. That is something very critical that we are working on collectively right now. We are evaluating tools to build a more cohesive data catalog with proper data lineage models, specifical things like Tablaue. Then, the second thing is purely around data quality checks. We are also evaluating and we are in certain final stages of evaluation of certain tools like to do automated data catalog and similarly in Monte Carlo and things like that. Those are critical junctures. And the third piece that we are working on collectively as a team is how do you present the data that you have in the data warehouse on the different platforms? How do you present it through API? How do you present this data born programmatically? And in certain use cases, we have requirements to build microservices for marketing use cases, where we have enrichment processes and all of that stuff. So, these are, I think in my mind, the next 12 months are going to be programmatic data and then data catalog and data quality.
Boaz: Can you give data programmatic examples?
Apun: Yeah. Sure. I will give you an example of marketing events, an event like this, then you have all the email addresses of folks that come here and then you want to make sure that you send email campaigns on these, but you do not know everybody here, who is a manager, who is an engineer. So, folks from entire different platforms generate the same. So people will get this data, like an excel spreadsheet or CSB. Right now, Part of the job that our team is to ingest the data, enrich this data with information about the company you come from, what is the role and responsibility, and then check your marketing data, like our marketing database to see if this is already exist and if it exists, did we get any new information and update that? So, this whole pipeline needs to be fairly quick from the ingestion to the time you can see into a targeting dashboard to see who all I need to target for marketing.
Eldad: You had to be around all street when you got from home?
Apun: It happens every day, it is a daily pipeline. But, like trying to build all of these more programmatically. A similar example would be consent management, from a compliance perspective. If I have to target somebody, I want to make sure that I have the latest consent information for that product. So, that is an API-based, like programmatic data access that I want to provide to the business just in marketing, we have about 40 different platforms that are being used for either email marketing, web marketing, and whatnot. How To integrate, how to build one single source of truth for this kind of information? That would be one example.
Boaz: Who is the central data engineering team, you guys serve the market team, sales team has a dedicated team essentially within your data team?
Apun: Yes, we have folks on the team who have a lot of experience working with sales organizations at Slack, at previous organizations and they have a very rich experience with CRM and things like that. So, we asked if they need the sales data engineering, and finance data engineering roles. Similar with marketing. Marketing is a very complex trade, so we have a lot of folks who come with that background, multi-touch attribution models, and things like that. So yes, we do have dedicated folks around, but also people get bored doing the same things, we have to move people around.
Eldad: What about quality? Do you also analyze quality or is that something that is being done only by engineering, like the quality of product or quality of service.?
Apun: my team does not look at that part, but what we look at is customer experience, we do analytics, are people logging, are people complaining about something and like, managing those bad data sets on the dashboard and presenting it into the right team so that the support folks can do. So yeah, that part is not like the product and it comes up eventually consistency. I am sure that is the product team.
Boaz: How do all the different data teams work together and collaborate?
Apun: We do work a lot independently in there. My team needs product data, for a particular matter, I would go talk to the product managers on the product data team, and provide them the details of the information that I will need. Examples would be Huddle. I want to use stacks, Huddles, and metrics. So, I will provide them with that information, and then they usually work in 2-week sprints depending on how big that particular requirement is, they will prioritize it, and then once that data is available we will use the data and publish it. But, engineers like to collaborate with each other, they have a community for the data engineers to talk to each other. But more formal data managers and DBMs.
Boaz: Awesome.
Eldad: Question coming from the crowd.
Boaz: Now it is time for a few questions from the crowd, and you know, you do not have to be asked Apun anything, and personnel matters too.
Audience Speaker 1: Thank you for being here. I wanted to ask you is that when you joined and now you are transformed into functional database development, what were some of the challenges that you faced? And, why did you choose the stack, you had? Was it just because of the past or the prior or was there any kind of vision involved?
Apun: Yeah, for what I understand like two parts to question, like, what are the challenges in what we do? And why did I move from more towards functional data?
Audience Speaker 1: Yeah.
Apun: I think challenges that I would say is I did prepare some slides but the challenges around that we have is just the three V's, the famous three V's is volume, variety, and amount of data sources, you get data CSV, XML, JSON. You get flat-file spreadsheets, it is just crazy, and whatnot. So those are the kind of challenges that our team works on a daily basis just to figure out how to ingest data in a more effective manner and build scalable data models. For me, I came from a data platform side, like doing databases for a long time and big data. So, I was on the side where I had no idea what business does. All I did was data. I build the platform. I serve the data and that's all. And I was always curious about a fact like how business works, like, how do you sell things? What is important to that? So, that is why I decided to move into the functional data engineering space. And I am definitely loving it. The challenges are very similar but different.
Audience Speaker 2: From an organizational perspective, you mentioned software engineering, and data engineering, they are part of the job that is happening for data engineers and also a lot of business logic, like the example mentioned, we want to some other even more complex operations. So, your data engineering team stays away from the business logic or you make sense too because if we move the data engineer to the software engineer team for the knowledge of the business logic, a lot of engineers are not that capable to handle all the data engineer challenge?
Apun: I think business logic in general usually lies in either the platform that you are using in the sense like when you say business logic, if it is business logic, build in a CRM platform, it is part of the platform. But as data engineers, we are always looking at the business logic to generate metrics. A simple example is weekly active users are a metric. It is a very complex metric. It is a simple and very complex metric and you need to get the definition and you need to build that business logic as a data engineer and publish those metrics. So if you are looking at more like Slack more as a product and the business logic in the product, I think it is more software engineering in my mind. Data engineering is more like an input to that process like, what metrics we see or regenerate, maybe the business logic needs to change, somehow. I do not have a very great example for that, but that is where I think the delineation is typical.
Audience Speaker 2: Thank You.
# A Technical Deep Dive to Yelp's Data Infrastructure — with Steven Moy (/blog/a-technical-deep-dive-to-yelps-data-infrastructure-with-steven-moy)
On The Data Engineering Show, we recently had the pleasure of speaking with Steven Moy, a Software Engineer at Yelp.
Yelp really needs no introduction, but in case you don't know, Yelp is an extremely popular website and app that publishes crowd-sourced reviews about local businesses and provides a table reservation service for restaurants. There are almost 5 million businesses claimed on Yelp. The website and app get around 178 million views combined each month.
As an expert in query engines and performance-related challenges, Steven Moy went on a deep dive with us, talking about how Yelp has handled its huge data growth in the past ten years.
Listen to the full podcast on [Spotify](https://open.spotify.com/episode/2MJmVxw2DGPLtYG3Q0H1sl) or [Apple Podcasts.](https://podcasts.apple.com/il/podcast/technical-deep-dive-to-yelps-data-infrastructure-steven/id1561927688?i=1000521449358)
We all know that the restaurant and hospitality industries were hit hard during the pandemic, so naturally Yelp usage declined as well. Now that the vaccine is being distributed, we are all excited to go back to using Yelp to help us find and review all our favorite restaurants and local businesses.
As data sets go, we think Yelp has a pretty cool one. So let's get into it.
## How big is Yelp's data stack? [#how-big-is-yelps-data-stack]
One of the coolest features of Yelp is its ability to make recommendations for great local businesses for you to check out or even what the best menu items to order are. To provide the best recommendations, Yelp needs to have a strong understanding of how users are interacting with the website and app.
Yelp collects and stores data on all of these user events. With millions of visits each day, this amounts to an enormous amount of data.
The user events are streamed through Kafka and are then forwarded to multiple data lakes. The data lakes are a very efficient way to store lots of data, but they are not very efficient when you want to use the data. Therefore, they also stream their data to an Amazon Web Services (AWS) cloud warehouse.
Yelp has been an outspoken AWS partner for years and has been using them for their cloud data warehouse for many years.
In 2010, Yelp implemented their first data warehouse solution. They started with a leader database and ran a lot of replicas. As the company grew and hired analysts, the replicas started to take too long to run, which slowed the team down. They needed to find a way to scale up and build a MySQL analytics-specific replica data warehouse.
This is when they switched to an AWS Amazon Elastic MapReduce (EMR) stack. This enabled them to speed up the transient clusters, provided good compute compatibility, and plenty of power.
This solution continued to work for them until 2013. They wanted to maintain a high-performance data product so they piloted screening data into Amazon Redshift. This made a world of difference for speed and productivity. Once they made the switch, analysts were able to run things in mere seconds that used to take them an hour.
However, the Yelp data team found that they needed to find another solution yet again in 2017.
Redshift was so popular with the team and it worked extremely well, so they continued to use it a lot. The problem is that Redshift uses a pay-per-use model, so with so much usage the costs were growing. The cost was already too high and was not sustainable for future growth.
So in 2017, the team decided to build a data lake solution instead. They now store all of their event stream data in Parquet and S3, and they use an Amazon data catalog. This solution allows them to ship data to Amazon Redshift and Athena. They also use Spark connectivity to mine directly on S3 via Parquet.
Steven reminds us that there is always a push and pull between innovation and cost. You need to always be innovating, but you also have to do it in a cost-conscious way. One way to do this efficiently is to work backwards. When their data usage was too high, they evaluated why they were scanning so many TB and tracked uses based on event types. This way, they were able to find an efficient solution, drop costs significantly, and continue providing all the same features to their customers.
## How many people work on data initiatives at Yelp? [#how-many-people-work-on-data-initiatives-at-yelp]
Data is huge at Yelp. Almost half of the organization needs to derive information from the data on a daily basis.
They use a lot of different technologies to manage their large data stack, so they have a dedicated team to focus on each technology. They also match data scientists with various other groups at Yelp so teams can work together to explore what is possible.
Much of the data is siloed within the organization because each team needs to build a unique use for the data. Every team that needs to track insights within the organization has their own data sets.
At the time of speaking with Steven, Yelp had an analyst team of 5 people.
## How does Yelp manage so much user-generated content? [#how-does-yelp-manage-so-much-user-generated-content]
Yelp runs off of user-generated content, such as photos and reviews. With millions of visits each day, that's a lot of contributions! So what does Yelp do with all of it?
For example, Yelp uses machine learning to figure out what the most popular dishes are at a restaurant based on what gets photographed and written about the most. Pretty cool, right? So how did they come up with this idea.
The idea actually was created at an annual hackathon. Yelp has 2–3 hackathons per year. One year, they wanted to focus on enhancing the "popular dishes" feature. Some attendees came up with a great idea to analyze the photos.
They immediately engaged the machine learning engineering team on this project, as well as their content team and data scientists to evaluate the new feature. It then went through prototyping and AB testing, and now is the great feature we use today.
# AI and Data Change Management with Chad Sanderson, CEO Gable AI (/blog/ai-and-data-change-management-with-chad-sanderson-ceo-gable-ai)
In this episode of The Data Engineering Show, host Benjamin and co-host Eldad are joined by Chad Sanderson, CEO and co-founder of Gable AI to discuss the revolution of data quality and governance, the importance of understanding data flow and the processes that help organizations manage their data more effectively.
Listen on [Spotify](https://spoti.fi/4h7NGZs) or [Apple Podcasts](https://apple.co/4fXIp5V)
\[00:00:04] Intro/outro: The data engineering show is brought to you by Fireball, the cloud data warehouse for low latency analytics. Get $200 credits and start your free trial at fireball.io.
\[00:00:15] Benjamin: Hi, everyone, and welcome back to the Data Engineering Show. Today, we're super happy to have Chad joining us. Chad is the CEO and cofounder of Gable AI, which is a data change management platform. Chad, how about you just quickly introduce yourself, tell us what you're doing, tell us what Gable is all about, and we're super excited to hear more.
\[00:00:34] Chad: Sure. Well, first of all, thanks for having me on the show, guys. My name is Jan. I live in Seattle. I've been in the data engineering infrastructure space for a pretty long time, well over a decade. As you said, right now, I'm the CEO of a company called Gable AI. We do code scanning. We essentially scan application code, figure out what data is being produced from those application systems, trace where data is flowing within a code base and, ultimately, where it lands. We run during CICD, detect when changes are happening that could affect the data in some way, and we communicate the impact of those changes bidirectionally before they reach production. So that is ultimately what Gable is about. There's a lot more complexity to it, how we layer on data contracts or data governance as code, how you start to shift data management to the left. Lots of really fun things that hopefully we're gonna get into more today.
\[00:01:25] Benjamin: Awesome. Sounds great. So before we dig into what Abel is actually all about, and I think you're the 1st vendor we have in the show for a while, so I always love these episodes with vendors. Like, maybe give us a bit of background. Right? Like, what did you do before Gable? What motivated you to start the company? What sparked the idea to start this and get this going?
\[00:01:47] Chad: Yeah. I mean, I spent a long time in data infrastructure. Originally, I was a data scientist. I was always more interested in the infrastructure side of things, so I moved into data engineering. But then I realized that as a data engineer, it was pretty challenging to launch the large scale projects that I thought needed to happen in order to change the way that companies thought about data, use data culturally. And then I moved more on to the infrastructure side, so the actual buildings of the systems themselves. And I did that at some large companies, Sephora, Subway, Oracle. I was a tech lead on the AI platform team at Microsoft. And then I I've owned all of the data platform and AI platform at a late stage startup called Convoy. And, really, everywhere that I went, I kept running into the same set of challenges, which is no matter how good our data quality posture was, no matter how thoughtful we were being about data governance and data usage, And no matter how much time we spent building out metric layers and systems to grant certain access rights to particular engineers for certain things, the quality would always degrade over time. Governance would always degrade over time. And the reason was because there's 2 sides to what we have started calling the data supply chain. There's the 1st mile and the last mile. The last mile is everything that's happening after data lands in a warehouse and is transformed and you use something like DDT and spin up data models and build dashboards on top of it. The 1st mile is everything that happens from the source to effectively cloud storage in most cases. And we were applying all of our quality checks, all of our governance, all of our best practices on the last mile, and none of it was being applied to the 1st mile. And because that's where all the sources were, that's where all the changes were happening, obviously, you would incur exponentially more data quality issues over time without those guardrails in place.
\[00:03:51] Benjamin: Okay. Gotcha. And then Andrew Gable, which is all about data contracts, data change management, can you give us basically a 1 minute rundown of what I would use Gable for in my data pipeline?
\[00:04:05] Chad: Sure. So Gable does 3 things very well. The first thing that it does is we identify what are the places in code that data is being produced, who are the owners of those codes? So who's the software engineers actually responsible for generating the data and making changes to it over time? And what does it mean semantically? And meaning can be extracted by looking at the context of the code and documentation and things like that. That's layer 1. Layer 2 is if you can do that in many different technologies and many different repos, you effectively have multiple nodes in a graph. You can string that graph together and understand the directionality of data flow. So how does data move from that source system ultimately into a database, through a Kafka topic, into cloud storage, and then to a data warehouse? And how is it transformed along the way? The big problem or one of the very large problems now is not that a data scientist or a data engineer can't go into a repo and code and subscribe to that GitHub repo and then get alerts if it changes. It's that a data scientist might speak the language of SQL. A software engineer might speak the language of of Java. So there is no translation layer in between helping each side understanding the context of, a, how sources change, and then, b, how data consumers might want to use data in new and in interesting way. So that's the second thing we do. And then the third thing we do, we call data governance as code. So how do you establish the policies about how data should be treated, how it should evolve over time, what the expectations in SLAs are, and then translate those expectations into enforceable integration tasks running in the code base. That's the data contract, and those are the 3 things that that Google does.
\[00:05:49] Benjamin: Okay. Gotcha. So one thing I saw when scrolling through your website is that actually on, like, your about us page, there's, like, kind of Bao from Monte Carlo who's in advisory role, whatever. Like, I'm actually super curious. It's like, how does a tool like Monte Carlo compare, which is also a lot about data quality, data observability. Right? It's like, to me, like, I'm a database internals guy who, like, builds database systems, so this part is always super interesting to me. And I'm not sure I 100% get it yet, to be honest.
\[00:06:18] Chad: So a tool like Monte Carlo let me use a metaphor. If you think of the software engineering governance or quality stack, right, we don't really talk about software in terms of, like, having governance and having quality systems, but they exist. GitHub is a change management system for code. Right? That's really all it is. People are making changes to their code base and all the different functionality of pull requests or merge requests or looking at the the various code diffs or being able to comment on those various PRs or even the act of merging and branching in and of itself is just a mechanism to manage change easier. And you need something like that. Otherwise, you you don't have an audit log of how code has changed over time. You don't have a way for humans to easily insert themselves into the review process while you're trying to manage an an agile, constantly evolving code base. So that's one pillar of this triangulate of quality and and governance. The 2nd pillar is is your monitoring. So these are tools like Datadog. Right? You have checks looking at what are customers actually doing in your application, HTTP requests and errors and things like that? And if you detect something that's wrong, usually, the engineer so the the producer who's ultimately responsible for customer outcomes will go and debug those changes and use the data from Datadog to handle that. And then the 3rd component is you need some mechanism of collaboration and documentation. And what does this data mean and or what does this code mean and where is it found and what does it do? And those are your tools like Jira or Confluent that are effectively catalogs. Right? They're catalogs of information about your software. So you've got the change management system, which is GitHub, the quality and monitoring system, which is Datadog, and the information repository, which is Confluent and and Jira and things like that. And so you can imagine that same sort of tricumference in the data space as well, but the needs of data teams are different than software engineering teams. But, usually, when you think about a a change in code, it's 1 software engineer working with another software engineer on their team. Right? I'm updating my repo. You understand my code base, so I want you to look at my code. But in the data space, that can't be the case. Right? If I'm a data scientist, usually, I'm going to be impacted by someone else making change in a very different part of the organization. So change management has to not be within a team. It has to be across teams. And that's a unique role that Gable fills as we are looking at code, understanding when changes happen to data specifically, and then understanding what the impact of those changes are gonna be on the teams who ultimately use that data for some meaningful purpose. Whereas Monte Carlo fills more of that Datadog role, which is looking at the contents of the data itself and helping engineers debug when things go wrong. And, of course, you have data catalogs, and data catalogs are sort of your Jira confluent equivalent, if that all makes sense.
\[00:09:19] Eldad: So maybe in some sense, a Monte Carlo comes after the change, and it's too late. So, yes, it's a barrier to protect you and tell you your data quality might be at risk, but it's not a data quality issue at all. It's a chain reaction that started with a PR by an engineer writing Java that was propagated down the chain. Eventually, somehow, there was an error coming back from Datadog, and now you need Jira and all the tools that Chet just described, 5 tools, 4, at least, if I remember right, different roles to figure it out that there is a connection between that PR and that error coming back from Datadog. And, yes, the data quality is wrong, but not necessarily. I think I get it now.
\[00:10:09] Chad: This is the problem that I think so many software engineers don't understand because in their world, a bug, it is very deterministic. Right? Like, I have a set of requirements for my code. And if those requirements are no longer met by the code, that is a bug. And now I can run tests on it. I can go root cause. It goes to my incident management system and and so on and so forth. But in data world, you might have a producer of code do something that makes complete sense for their use case. Right? Maybe I have a time stamp field, and I wanna change that from local time to UTC. Makes sense. That's probably what you should be doing. But if somewhere downstream, a data scientist has built a machine learning model, the the time stamp field has some sort of predictive weight to it. Those predictions will be wrong if they're not informed that change is coming and and update their model accordingly. So you can create data quality issues from changes that are good, and that is a very different sort of concept than most software teams are are familiar with.
\[00:11:10] Benjamin: Gotcha. So that was, I think, a really good, like, top down bird's eye view explanation of the space. Right? And I appreciate that. So for me as well, like, it makes a lot more sense now. Maybe another way to approach this is actually like a bit more bottom up, right? Say I'm the software engineer, I own a repository, I publish some messages to Kafka as part of that. And like somewhere, power downstream in the organization, there's a data science team which also builds a predictor on that. I make a code change now, and I make that code change that goes from, let's stick with your example, local time to UTC. Like, will my CI fail? Like, what's actually like, what does Gable do then in order to alert me? Hey. You're gonna cause some damage here. Right? It's like, how can I imagine? Like, how does it impact my life cycle as a developer, basically?
\[00:12:02] Chad: So the way that we think about it is that my intent with Gable is to ultimately light the spark of culture change in a company, which is a very lofty vision, but that is ultimately the goal.
\[00:12:14] Eldad: I love this vision.
\[00:12:15] Chad: Thank you.
\[00:12:15] Eldad: Ambitious, great vision. We've been doing the Jira things for 30 years.
\[00:12:20] Chad: Yes. Exactly. And so the real high level question is, how do you change the culture of how teams think about data, use data, all these types of second order effects? And my view is I believe this very strongly in this concept of of Conway's law, if you're familiar with it. And Conway's law essentially says that the systems people use to collaborate with each other and talk within an organization is ultimately reflected in their architecture and their systems and their processes and the product itself. And in the data space, you think about how people talk to each other. Well, you've got these producers and consumers that are the 2 primary sides of the data equation. You could argue that there's, uh, platform teams and sort of data engineers in the middle. We're not very good at talking to each other, and that's because of all the problems I laid out before. There's this missing translation layer that's not actually helping each team understand the context of what they need to know in order to make a good decision. So what is the right information to share with the right person at the right time depending on the change that's happening? And then what you do with all that information is the the next set of problems to be solved. So now that we understand that, I'll drill down and make it very tactical.
\[00:13:37] Eldad: Okay. So quick question. So the problem exists everywhere. Every company that builds software to deliver value to customer goes through that challenge. And then over many cycles, the only solution that was available for everyone was build a global semantic model. It solves maybe a third of the problem, but let's get everyone those different cultures, those different well deserved opinions because as you said, those are different environments ran by different cultures, different team, different purposes that needs to be combined together to provide customer value and the only thing we could do is just bend reality, which is build a global semantic model so everyone talks the same language. That doesn't work. You just can't enforce that single Latin language on everybody to try to solve that problem.
\[00:14:30] Chad: Yes.
\[00:14:31] Eldad: And my question is, how do we replace those semantic models, which obviously are very challenging? And they do solve many problems, by the way. Vertical, right? If you do drill downs, yes, you can solve a lot of problem with the semantic models, but it doesn't solve what the problem that you are talking about. How is AI coming in to change the equation? Right? Because this is the big change. Going from a semantic human managed project management kind of mindset to a AI mindset. And this is super interesting. So tell us, how is that change happening and what is happening there?
\[00:15:07] Chad: You're hitting on all the points. So, yeah, to follow on the first part of your statement, which is we had this concept of a semantic layer or this large semantic model in the entire business would agree on what certain terms were and how to use them. And then anyone who is writing a query in the future or generating data that sort of corresponds to those terms would need to follow all these. Of course, the challenge there is that works reasonably well if you have an incredibly centralized data production and consumption organization. Meaning, I cannot create a new database without going through some centralized body that has the context of all this semantic information and can make sure I'm doing it the right way. And I cannot build new data pipelines without following a similar process. Right? So it is that steward in the middle that is keeping the semantic map of the world in their head and requiring enforcement in all these different places. That becomes a lot harder to do in this modern era of the cloud that is totally decentralized bidirectionally. Right? You have software engineers that are encouraged to move as fast as possible, build microservices, solve customer problems, right, go fast, fast, ship things. And then on the consumer side, now that you've got tools like DBT that are effectively federating out data modeling, you've got a lot of analytics engineers and data scientists who are building their own data models, and you start to see the warehouse expand out this way as well. Right? So that makes the maintenance of a semantic layer very challenging. Who is even doing that work? Is it the central team? Is it your central governance team? It just becomes very hard. The effort to maintain the system radically exceeds the amount of resources you have to do it. So that's the current state of the world. So what does the world actually need to look like? At least, again, this is my opinion. But what I think is that the people on either end of that process, the producer and the consumer, have all of the context that they need in their head, and they reflect that context through the code that they write. Right? So if if I'm a software engineer and I am building a service that is collecting customer data, which I then use to populate a UI, a reasonable software engineer, like a mid level software engineer, would be able to look at that code with a little bit of context maybe from documentation and infer what's going on. Right? Like, I know this is a website front end. I know that someone is clicking on a button to sign up for my website, and I can see that creates a new record in my database. Right? Like, you can do that. And then on the other side, a reasonable person can look at a SQL query, and they can look at the names of the columns that are being joined together and come up with a reasonable idea of, like, what is the output this is supposed to be generating. And, of course, there's a lot of context around that as well, like the actual dashboard and documentation and yada yada yada. And what my view, where AI fits in all this, is that those types of things, looking at code to infer intent, is exactly what AI is extremely good at. Like, that's one of the main things that it's good at, actually. There's a lot of things it's not that good at, but that like, extracting semantics from code by following convention and patterns, it actually is really good at. So the question is, how do you leverage AI in order to extract that meaning from both sides and then effectively manage the translation in between what is created and and, ultimately, what is used. I think that is the problem to be solved, and there's some sophisticated ways that you can do that, but this is really at the core of what Gable is. Sorry if that was a very long winded explanation.
\[00:18:57] Eldad: I laughed. It was perfect. So quick follow-up question on that. How do you educate your customers and prospects in the market which already spend years, right, building a physical semantic layer of people, ownership, departments, you've mentioned centralized. So many companies, so many teams head to resorts to a centralized semantic or middle gate, like a gate. Right? It's always a cultural thing. Every company goes through that. Right? They start decentralized and then a subset of that ends up centralized because of all of those challenges, and then they give up. So they apply people manual processes on that state. Now we're saying we don't need to continue and invest and keep that state going. We can actually start dissolving it slowly and carefully and figure out how people that are involved today in keeping up that process going, what are they going to do? What's the transition going to look like Yes. For them. That's is an equal of a challenge to solving it technically. Because once you solve it technically, you deal with the human challenge. Right?
\[00:20:14] Chad: Yes. I think that is exactly right. So there's a few interesting things. The first interesting thing is I don't think that the way any of this can work is to take an LLM and to throw it at the problem and say, go do magic. Right? But it doesn't work like that for a lot of different reasons. One reason being, at least the current state of AI, and this will improve over time, certainly. But in the current state of AI, these models do not do a very good job retaining memory. Their state management is quite bad. Especially when it comes to the analysis of code, state management is important. I need to guarantee that anytime that function is used, it's taking the same 2 values effectively. And the AI is effectively guessing at what the function takes and what the values are because it doesn't really have a concept of of state management today. There's a bunch of other reasons where if the context windows are big, then the cost is just insanely massive. Trying to reason over all this stuff all the time becomes exorbitantly expensive. Anyway, point being, that what makes this problem a lot easier for a model to go and solve is when it's directed. Right? When you can give it context and help and hints about how it should be doing interpretation, you're shrinking that context window, and you're making it a much more powerful and scalable system that applies not only to what code should it be looking at in order to infer what is this data and where is it being used. That's a challenge in itself. But, also, you have like you said, you have all this existing context and existing work that's been done that can make these models so much better. And the more context you add over time, the smarter the models actually become and the more effective they become at doing the tasks. So I I think there is this transitionary period where the people that are currently spending all their time managing these massive ontologies will gradually transition to saying, how can I teach the model about everything that I've I've built over the past 10 years or or 15 years or or whatever it is? And that's not gonna happen in a day or in a 2 days. It's gonna take time. I like to think about AI, at least a a relatively basic untrained AI, almost as a new intern to your company. Right? Like, it's not that great at its job, but it can be in many different places at once. It's almost like a universal intern, like an intern that's just everywhere.
\[00:22:41] Eldad: We love interns.
\[00:22:42] Chad: Yes. And so the question is, well, it's probably not useful for that system to remain an intern forever. And it it could end up costing a lot more than it's worth if it's if that's the level. So you you have to raise the knowledge and raise the awareness and the value of that intern. I think that only comes through these more manual human in the loop steps. And I think that what that actually does is that it allows the data team and the data engineering organization to transition to the work that's actually interesting and useful. Right? Instead of constantly trying to keep track of the semantic state of the world, you start thinking about strategy. How should we be keeping track of things? How do we want to orient our business? How do these various teams start to take a semantic concept that's been defined over here and merge it with another semantic concept that's defined over here? This is the types of problems they'll begin working on.
\[00:23:33] Eldad: So for most information workers today, AI, they mostly consume it. Right? They consume a model. It's a black box for them. One of the challenges is, and and I completely relate to what you're saying is, how do we take all of the information workers and their value is their context that they have? This is something no system can take away from them. They applied the same context previously on software to build stuff. Now you're saying we will start transitioning away from that into teaching models, into building interfaces so information workers can provide context back. That will not necessarily be a diagram or an input box, but there needs to be some interface allowing information workers to feed the model back. And this is like absolutely waiting for disruption because the model is a black box. It's being handled by very few people who know how to generate it and everyone else which is 99% of the information workers are basically kept out of it. They don't know how to tell the model how to provide context, background, data points. They do it on legacy tools all the time, so they know how to do it. We know how to provide context. Just have someone build those tools to help us do it. We are all interns in my example.
\[00:24:53] Chad: Sure. Sure. Sure.
\[00:24:54] Eldad: If you give us the tools to provide context to the model, to the AI versus the legacy tools, then actually, I think, like, that's gonna happen. At least to to me, everything connects and it makes perfect sense.
\[00:25:07] Chad: I think you're exactly right. One of the fascinating things that I found is in this space, it's much less about building a technology. Although we do have to do that. Right? And building the technology is very hard. And we've had to hire PhD researchers and folks who've worked in code analysis for many years, and, like, that's very challenging. But it's not really about the technology. It's more a question of incentives. There is some information I need to extract from various people along this workflow. And then there's certain changes in behavior I need to inspire. So how do I create the right incentive so that, a, I'm getting all the data that I need? And, b, how do I initiate a change in the culture? And I think there's a lot of interesting ways to do this. So on the consumer side, one of the best incentives is if you provide context to the system about what does the data mean? What do you care about? What indicates a data quality issue? How should the world actually look? Which to your point, an LLM will never know that. It will never have all that information. A person will. Well, if you give provide that information, then you get high quality data back. It's a feedback loop. So if you want to productionize some data product that you have created, and that could be a machine learning model. It could be I have a report that I sent to the CMO. And if it's wrong, they come and chew me out on a Friday, or it's this is actually public information that goes to the board or whatever it is. Right? If you want that data to be right and you want there to be very clear, explicit ownership and thoughtfulness around your use case, then you need to provide the context of what you need, which sounds like a very basic and obvious idea, but it's not something that has really happened generically. It's certainly not through vendors and platforms in the data space. On the other side, on the producer side, there's also an incentive. Right? This is where I I see most especially data folks in particular will give me this feedback and say, well, my data producers don't care about me. They don't care about data. They are not interested in changing their work or taking on more work that makes my life better. I don't think that's actually true, not based on the the thousands of conversations I've actually had with software engineers. What they usually say is, you need to make it easy for me. Right? If I have to go and own a bunch of technologies and systems that I have no knowledge of and no experience around, I need to own data quality tools. I need to look at a data catalog and figure that whole thing out. I need to own DBT. I need to own testing and validation systems inside of whatever your analytical database is. They're like, yeah. I'm not gonna do that. I have enough work on my plate just dealing with the systems I'm maintaining today. So you you have to make it very easy. The second thing is they need the context like what we've been talking about. They should understand where is the data that I'm producing going? Who is using it? What are they using it for? Is it important? How important is it? Right? And then there's the expectations. So what is expected of me? What is expected of the data? And how would I know if those expectations are are not being met? Right? And then e even after all of that, there are going to be situations where you might have to change the data in a way that does objectively cause problems for a downstream consumer. But there should be a workflow around handling that. If I'm making a breaking change to someone, the problem is not the breaking change. The problem is the time between when I make the breaking change and when it actually affects that consumer's query or dashboard or or whatever it is. And so if you can create the right systems that are incentivizing the producers to say, oh, I know who my consumers are. I know how they're using this data. I know that this thing that I'm doing is gonna affect them. I know how it will affect them. I know who I need to communicate to at what point in time, then it's much easier for them to follow that process than deal with all the fallout of trying to fix an issue they shipped to production 3 weeks ago.
\[00:29:06] Benjamin: Right. So at this point, like, I'd actually love to drill down again and return to this previous conversation because I think it's a natural kind of here it connects. Right? I make that PR. I kind of change from local time to UTC. Like, how do I interface with a tool like Gable? Like, now both on the producer side and on the consumer side, who have to then adapt to it?
\[00:29:28] Chad: So this goes back to the sort of 3 layers of technology that I mentioned in the beginning. Now let's maybe drill down to another layer. The first step, and I think the lowest level of maturity in this big cultural change we're talking about is you just have to know what data is being produced where in your codebase. So what Gable does is we have a command line interface. You can point it at all the repos in your company. We know how to scan that production code, and we can essentially identify where there is data being emitted and and pushed from one system to another. Once we can identify that point in code, we can reverse engineer the payload. So we can say, alright. We know that this is the actual event payload. This is the schema. And we also have a code block that contains all the surrounding context of that particular event, which you can feed into an AI system and ask it, hey. What what exactly does all this stuff do? Once you've defined all that, so that is stored within the Gable UI and the the Gable APIs. We run as a part of CICD. So anytime there's a new PR that gets opened, we understand what the state of the world should be. We take the diff into what the new state of the world is, and we also understand, are those changes going to cause this data object to meaningfully be altered in some way? Right? And those changes alone, as long as we know who is using that data, and there's a whole other explanation on how we know that, We can communicate that information in the language that they understand, English or whatever language it is you speak in your part of the world. Right? Hey. This change is coming. This is how the data will be different from the previous version to the new version, and this is how it affects you. And we know that because we know how it's all connected and and we can trace it. Right? So that's, like, level 1 is you're just tracking how changes occur, and you stick that into an audit log. And so if you wanna do root causing and say, hey. I I wanna understand how my data has changed over the last 6 months. Well, you're you've now tracked everything, and you have all the context.
\[00:31:25] Benjamin: Just to quickly hook in, like, also, like, I guess these layers also connect to then how you can over time gain wider adoption within an organization. Right? Because I'm sorry. I could still, at this point, have my semantic model and all of these things and my higher level concept, and I just use Gable in this, like, layer 1 kind of for, hey. Wow. Okay. Now something went wrong in understanding the root cause, for example, or what upstream change caused that.
\[00:31:51] Chad: Exactly. A very common sort of first use case for Gable that we see is someone will say, hey. I've got some data coming from my production system that is really important to me. It might be data that I sell to a customer. It might be something that feeds into a machine learning model that makes me $100,000,000 a year. It might be data that feeds into our accounting system, and I need that to be right. And I need to know anytime that changes, and I can't be reactive. I have to be proactive in how I deal with that and track it over time. The 2nd layer is feedback to the engineers making the changes of how what they're doing impacts the rest of the system, and you do this with data contracts. So the way that I like to think about the contract is that it is a baseline that represents the expected state of the world for data. So So as long as someone has said, hey, this is the way that I always expect the data to look. If there is some code change that happens that causes that state of the world to be different, you can give that feedback directly to the engineer through a comment in the pull request. You are making a change. Here is how you are changing the state of the world. Here is the people who are dependent on that state of the world being true, and here is all the negative downstream impacts that this change is going to have if you make it without giving them some leeway or whatever it is. Level 3 is, alright, now that we all understand the role that we play in this ecosystem. Right? You as the producer know how you impact me as a consumer. Then we start rolling out this is the governance's code. Right? We're building integration tests so that if the contract is not being followed through a code change, you simply can't make that deployment. And you either have to update your code or update the contract, which has its own approval workflow. So the people who are dependent on the contract will need to say, okay. I understand that this change to the contract is coming. I need to reflect that in all of my queries and so on and
\[00:33:44] Benjamin: so forth. Right? So this is tactically how that change manifests over time. Super cool. This was incredibly interesting. Seriously, it's like such a cool take on organizing data and at scale, seriously. So looking at 2025, right, like, I think once the episode goes live, maybe it's in January already, is like, what are you excited about for Dable? Like, what are your goals for 2025? What do you hope to achieve?
\[00:34:13] Chad: Oh, yeah. That's a good question. So the thing that I think is most exciting to me is to continue to see ways that our customers use a system like this. Right? It is extremely flexible and varied. We have seen folks who have not just started using it for the benefit of their data engineering organization, but are also using it to manage the relationship between front end engineers and back end engineers for changes to APIs. It's all just data at the end of the day. Right? And so if you could say, hey. Back end engineer, you're making a change to the API that I'm consuming to populate all these UI components, that is going to affect me. Now you're starting to make this data management problem something that the engineering organization cares about independent of ML and and data engineering and analytics. And that is really what you want. Right? When the whole company starts to care about data from their perspective, and they start to care about how changes made to data affect other people. And I think the word data mesh, people have done a ton of writing about it. Jamaka is amazing. I think data mesh is a very beautiful idea. But I think that there's actually something that's a little bit more foundational than data mesh, in my opinion, which is just the concept that if you're building software systems, you can never truly decouple yourself. You are always tightly coupled to something. And the previous state of the world has almost been a lie, in my opinion. It's almost like, let's just pretend that we're not tightly coupled to anything, and let's build almost independently of the broader system. What what would be really exciting for me in 2025 is this broader realization that actually everything is a mesh. Any software is actually a mesh. And as a company grows, the mesh gets more and more sophisticated. And the next stage of software management and and data management is how do we exist within the mesh. And there's a lot of other features and functionality and and use cases that you can start to think about and and that Gable will ultimately support in that world.
\[00:36:14] Benjamin: I love that as, like, a 2025 goal. How do we exist in the mesh? That's an awesome closing sentence, Chad. Thank you so much for being part of the show. Really, it was great having you, and all the best for 2025. Thanks, guys. This has been really fun.
\[00:36:29] Eldad: Thank you.
\[00:36:31] Intro/outro: The data engineering show is brought to you by Fireball, the cloud data warehouse for low latency analytics. Get $200 credits and start your free trial at fireball.
# AI and Data Movement: Trends and Best Practices with Estuary’s Daniel Pálma (/blog/ai-and-data-movement-trends-and-best-practices-with-estuarys-daniel-palma)
In this episode of *The Data Engineering Show*, the bros sit with Daniel Pálma, Head of Marketing at Estuary, to delve into the intriguing world of data engineering and marketing. Daniel shares his transition journey into marketing from data engineering and how his technical proficiency has been leveraged to market to engineers. The conversation cuts across the importance of AI in data movement, the future of data engineering, real-time data integration challenges, and the evolution of data integration.
Listen on [Spotify](https://spoti.fi/42Ss5jQ) or [Apple Podcasts](https://apple.co/3ExRJQW)
**Intro/Outro - 00:00:04:**
The Data Engineering Show is brought to you by Firebolt, the cloud data warehouse for low latency analytics. Get $200 credits and start your free trial at firebolt.io.
**Benjamin - 00:00:15:**
Hi, everyone, and welcome back to The Data Engineering Show. Today, we are super happy to have Dani join us from Estuary. Great to have you on the show, Dani.
**Daniel - 00:00:23:**
Yeah. Hi, everyone. Thanks for having me. I'm happy to be here.
**Benjamin - 00:00:25:**
Perfect. So your job title at Estuary says data engineer and marketing. Tell us about how these two things connect and what you spend your days doing.
**Daniel - 00:00:35:**
Yeah, sure. So I actually totally forgot to update my title on LinkedIn. It should say head of marketing as of three months ago.
**Benjamin - 00:00:42:**
Nice. Congratulations.
**Daniel - 00:00:45:**
Thank you. Thank you. Initially, when I joined almost a year ago, it was data engineering and marketing. It started as a content creator focused role with a lot of expertise in data engineering. So I come from the data engineering world. I worked as a data engineer for pretty much a decade before going into marketing, startups, enterprises, and most recently, like five years in consulting. So my job was to know the target customer, the target audience, which for Estuary is KC, our data engineers, know their problems and know, what they look for and provide them with useful content that ideally helps them on board to Estuary. But if not, still solves part of their problem, at least. And still, even as head of marketing, that's still my main goal to help Estuary users and prospects help in all kinds of data engineering issues.
**Benjamin - 00:01:41:**
Awesome. That's super cool. For the listeners who never heard about Estuary, you want to say a few sentences about what you guys actually do?
**Daniel - 00:01:48:**
Yeah, for sure. So Estuary is a, data integration platform, specialize in real time data movement. So if you work in the data world, you are probably familiar with Fivetran or Confluent or Hevo. Like there was a lot of similar vendors working in this space. Estuary is an all-in-one platform, unified in the sense that we do both streaming real-time data integration and batch data integration as well for analytics use cases. So the shortest pitch is that if you have any kind of data movement challenge, like you want to move data from point A to point B on your own cadence in your own budget, then Estuary can do it for you.
**Benjamin - 00:02:30:**
Nice. That's awesome. Cool. So, moving from data engineering to marketing, tell us about that shift. What's difficult? What's harder than you expected? What's easier than you expected? Take us through that journey, basically.
**Daniel - 00:02:45:**
Yeah, it's an interesting journey for sure. Even before going full-time into marketing, I used to do a lot of content creation mainly. I was always writing blog posts for my own blog or for other vendors or other companies' blogs, mainly technical how-tos, tutorials, and yeah, just my thoughts about the data industry in general. So that was my introduction to marketing. And I started liking it more and more. And I started getting a little bit tired of the actual data engineering, like the actual individual contributor work. So I decided, like I saw that this role was open, and I thought that this is a perfect combination of data engineering skills that I can use and providing useful or creating useful content for fellow data engineers. So I decided to give it a go. And it turns out that I really like doing it. So I decided that it would be a good idea to do even more of marketing. And yeah, I've been learning a lot obviously, marketing as an industry or as a vertical is, it was fairly new to me. So there's been a lot of things that I had to catch up on. But yeah, so far, it's been an amazing experience.
**Benjamin - 00:03:56:**
How much do you think your data engineering background actually helps you now in terms of being a marketer, right? Because obviously, these are very technical products, kind of like with a lot of moving pieces, a lot of complexity. Tell us a bit about that.
**Daniel - 00:04:09:**
Yeah, a ton, honestly. I couldn't imagine doing this without my data engineering background. I think the biggest challenges in data or marketing to data engineers or data practitioners is the same challenge that marketers face when marketing for or to software engineers or anybody in tech, pretty much, that people in tech hate being marketed to. And I know that because I used to work as an engineer, and I hated being directly marketed to. And I think it's super easy to see through like, lame marketing attempts. And that's where the biggest challenge is to actually market your product in a way that is not like in your face and you know, if you're running ad, this is the best service that you can ever imagine and solves all your problems. But to actually show how it solves their issue and provide useful content. So it's a very fine line between being obnoxious and being useful. And if you find that sweet spot, then that's what I call successful marketing to engineers.
**Eldad - 00:05:07:**
Maybe another way of putting it is looking at marketing versus product marketing, right? So if you drill into product marketing, you can say, okay, let's start product marketing out of marketing. Or you can say, let's start marketing out of product. If you're into data engineering, and you've been educating mostly, right? If you've been involved, not just in solving problems real time or building projects or delivering projects, but you're actually educating about data engineering, then Being able to market that education, right? Turns that into marketing. And it's a good thing. It comes bottom up. It comes from data engineering up to marketing versus the other way. Therefore, it's authentic. And it's all about solving actual problems at scale. And it's super interesting to see more and more people get into marketing by pulling it into their domain. So you're a data engineer, you pull in marketing, and then that becomes something else, something bigger. Tell us how you apply in your daily work. How do you take that approach to your customers that are always having challenges moving data? How does it work?
**Daniel - 00:06:13:**
Yeah, so at S3, we are focused on real-time data integration. That's where a lot of the technical mode of the platform is. And historically, real-time streaming data movement has been very complex and hard to implement and expensive as well. So that's where the education aspect comes in, that a lot of data engineers don't know much about real-time use cases or how to implement real-time data pipelines or streaming data pipelines because they come from the batch analytics world, like where you have a daily report or a weekly report and you need the data to be ingested every day, and that's good enough. But as you go into those use cases where you need to react on something in as real time as possible, like fraud detection or calculating customer facing metrics, like obviously there's no chance to wait for the next day to calculate those. So we have to do a lot of education on how these real time pipelines work and how they are different compared to the traditional batch pipelines. And a lot of our content is centered around that. So we teach generic pipeline building on our blog and in our videos. But we also have to do a lot of education about the platform itself because it introduces a lot of terminology, a lot of concepts. So it's not built on the most popular real time frameworks like Kafka that most people are familiar with in the space. It's built on a completely new backbone, a new streaming backbone. So that introduces extra complexity. So yeah, it's hard. And it's one of the challenges that, we pretty much face every day in our content on how much we can focus on education, how much we can actually focus on the solution that we are trying to provide. So yeah, it's something that we definitely want to get better at and it will probably never stop.
**Benjamin - 00:08:04:**
So if you actually pitch Estuary, to someone is like, what are you better at as a tool than your competitors? Like what's your unique value proposition basically?
**Daniel - 00:08:14:**
Yeah. So there's multiple things and it really depends on the use case because, as I said, we. We do a lot of things. So we compete with big platforms who are specialized on one use case like real time or specialized on batch data integration. So depending on what's the prospect, let's say is looking for, that's how I usually tailor this value proposition. But in a sense, the biggest difference is that our platform is built for the cloud as opposed to other platforms who utilize a framework that was built for on-premise servers and databases a few decades ago. And because we are built for the cloud, it's pretty much infinitely scalable, way cheaper than alternatives because, we use object storage as our primary backend and we can move way faster than companies or organizations who use Kafka, for example, as their backend. So yeah, the shortest pitch I think is that it's faster, cheaper, and more efficient.
**Benjamin - 00:09:15:**
Gotcha. And how hard is it? We're like Firebolt, like we build data warehouse. So we think a lot about actually kind of people moving from other data warehouses to Firebolt. And migrating from one data warehouse to another is usually a big undertaking for organizations, right? Like different SQL dialects, different ecosystem integrations, and so on. So we're spending a lot of our time thinking out to make that as easy as possible. Like when you build a type of tool like Estuary, how hard is it there basically to make it easy to move, let's say, from 510 to Estuary?
**Daniel - 00:09:46:**
So moving from a competitor like Fivetran to S3 is we're trying to make it as easy as possible. Like the actual data pipelines, you can spin up an S3 data pipeline from, I don't know, a Postgres database to Snowflake in five minutes. And we make sure to get all of the data that is in your source and put it in your data destination without anything missing or without any duplicates. So in that sense, switching is super easy because there's nothing that Fivetran does that we can't do. The complexity comes in like auxiliary things that Fivetran has, for example, a lot of very cool dbt packages that do a lot of transformations in the destination. And that's something that we currently don't have, but we have a lot of requests for something like that for my customers. But as for the platform itself, yeah, it's super easy. And we do a lot of these migrations from all kinds of competitors that are unhappy with price or latency or anything else.
**Benjamin - 00:10:44:**
Gotcha. And then like the longer term goal is basically taking more and more of the actual whole ELT pipeline, basically, that you don't just do the data movement into your sink, but can also orchestrate data transformations there. Is that right? Or like when you talked about the dbt packages, I'm not sure I understood that correctly.
**Daniel - 00:11:02:**
Oh, yeah. So those dbt packages specifically are open sourced by Fivetran and they are available on GitHub. And they are meant to be executed in the destination. But you know, they are tailored to the schema that Fivetran manages from their own connectors. And a lot of Fivetran customers use those as like the first layer in their dbt project. So everything downstream of those expects that structure. And obviously, when we are doing a migration from Fivetran to S3, our schema is a little bit different, which is a good thing in some cases, because that allows us to be like a fraction of the cost as we don't normalize the schema into many, many tables like Fivetran does. But this is one of the drawbacks that we have to reimplement some of those transformations. But for the end goal, as you said, it is to do everything data movement. So you have to move data from source to data warehouse, we can do that. And we can also trigger transformations. And we can also do movement from a data warehouse to SaaS tools or back to a transactional database like reverse ETL style.
**Eldad - 00:12:05:**
I have a question on AI, actually. So if you think databases from an interface perspective, it's very easy, right? You write in English and that generates SQL queries and everyone is happy. But when you go to data movement, the story becomes more complicated because there's so much more context needed as you are actually moving between different systems, right? Data warehouse has the data, has the schema, has the SQL. It has everything baked in. But data movement is much more complicated from a migration or from an integration perspective. And, right, a big part of the challenge with ETL or ELT or any of those tools was always actually understanding context and migrating context and metadata and knobs and configurations. And you need to understand the source and the destination so well, right? So you can have your data transfer being amazing. But if you don't have understanding, about the source or destination, then you're in a challenge. And this is for, right, the tool, having the right tool that knows how to properly connect in and out. Becomes very relevant. So my question is, how is AI affecting the user experience, right? So I'm thinking ELT tools going forward, I would want to say here, listen, I have that source, I have that destination. Actually, sometimes I also have a different product called Fivetran or any other one of the great products there. Can you use those things to generate my, use my product, clone that and just run it? And that would go and really, right? Like fill in. All the workflows, do all the steps, and get you started very easily. Therefore, I'm kind of democratizing or simplifying data movement, enabling it for so many new use cases. It's not a person, just a person anymore, owning data movement by using a tool. It opens up many new use cases. And I was kind of curious to get your thought. As everyone is thinking, rethinking how AI can simplify and get them more productive. What's your take on movement and AI?
**Daniel - 00:14:15:**
Yeah, I think it's a big topic, both movement itself and then how AI can help there. So we are currently working on opening up the platform for AI-driven use cases. And there's a lot of things that we have to change for that. But the end goal is, or one of the goals is, that anyone should be able to spin up a data pipeline by typing or typing in English or saying some words, you know, instead of having to fill out a bunch of configurations. The hard part is that data movement is just super complex. There's so many edge cases that even if we cover all of those, the actual systems that we move data from and to are constantly innovating and changing as well. And that always introduces new edge cases and new features that we have to follow up. So the work for the connectors functionality will never stop. So enabling all of those to be used from like ChatGPT interface is, I think it's possible, but it's probably a little bit further away. Yeah, it's just a super hard problem. As for using AI to actually build data integration pipelines, it can work for up to a certain complexity. But after a while, the source and the destination systems just become so complex that I haven't yet seen an AI-generated data movement script or application cover all of the challenges.
**Benjamin - 00:15:41:**
Are you, like, for these types of people, also now customers shifting towards AI applications, are you seeing shifts in the type of data people are moving around? Or it's basically still the same data people are moving around to Power AI applications? And then it's more about also for you, how you can give human interfaces to, or plain text interfaces for people to spin up data pipelines easily.
**Daniel - 00:16:07:**
Yeah, so most of the AI-based use cases that we see are people capturing data from a source and moving it into a vector-based representation. Usually, it used to be a dedicated vector store, but nowadays people use their usual data warehouses because everybody started supporting vector data types. So vector stores are becoming less of a thing lately. But as for the sources, I think the most common source still is an operational database that people just want to get their data out of. So meaning a backend database for an actual web application, usually Postgres or Mongo or MySQL. So those kinds of guys. And in addition to those, we see more and more use cases of people trying to enable AI for their Salesforce data, for their HubSpot data, for their NetSuite data. So I think more and more enterprises realize that they have these cool, super strong services that store a lot of data, but it's locked inside. And they are trying to... Get more value out of it. And now it is so easy to spin up a chat application for these. But a prerequisite for those is that you have to get your data out of those, create those vector embeddings, and store it in a database that actually can plug into some of these RAG chat applications.
**Benjamin - 00:17:22:**
Right. Like we're obviously as a data platform also thinking a lot about that because we want to power these types of RAG AI applications. And one thing we've been seeing more and more is basically that people are trying to figure out how to make these AI applications work well on structured data. So if you look at like vector embeddings and vector search, like it was like original kind of RAG pipelines. But the type of like you mentioned, right, like Salesforce data, like and so on, like this is inherently structured. And actually building kind of powerful AI applications on top of that is something that we're also working to figure out with our prospects and customers. And I think it will be super interesting to see how that evolves over to like kind of 2025 as people really figure out how to build amazing AI data applications on top of not just unstructured kind of vector databases, but also structured data.
**Daniel - 00:18:13:**
Yeah, for sure. Like does Firebolt support some kind of vector embeddings generation internally? Or is the use case more of like not vectorized data serving as the backend in a chat application or an AI application?
**Benjamin - 00:18:26:**
We support vector search and all of that, right? As you said, all cloud data warehouses are now pushing to support also these more unstructured use cases. But what we're actually thinking the most about and seeing the most with customers and prospects as well is how can we make it work on structured data? In those cases, you don't want to take your 12-byte text and move it into a 20-kilobyte vector just to do nearest neighbor search, right? It's kind of like people still want to build efficient data applications on top of their structured data. And at least we currently don't think that vector search will be what's going to power these types of applications in the future. One thing that we're seeing, which I think is quite interesting, is a lot of text search use cases, right? Because all of a sudden, there's a lot of text-based human interfaces to data. And this is something that we, for example, have been focusing on.
**Daniel - 00:19:17:**
Yeah, I think that's super interesting. And it's weird that somehow the tech scene collectively decided that vector search or similarity search is the de facto way to build chat applications. But in fact, you just have to somehow stop things into a prompt and that's the end anyway. So any kind of search works. So yeah, I think that's definitely a way to go to expand the horizon and look into what actual search solution is the best for those.
**Benjamin - 00:19:42:**
If you look at database history, like this is a very common theme, right? It's like, okay, a new type of workload emerges, systems emerge that kind of try to redefine how to build these applications or power these applications. But in the end, and this is also what we believe now, okay, relational SQL databases, just absorb it.
**Eldad - 00:20:01:**
We've had a guest actually a few episodes ago. He had a very nice angle to it. He said, nobody cares about writing code. Everyone cares about being able to reason about it. So if AI generates the question, we want to reason about it. We want to be able to audit it. We want to be able to understand if it asks the right thing. So SQL is suddenly getting that AI love, not because people like SQL. People hate writing SQL, right? They use BI.
**Benjamin - 00:20:28:**
I love writing SQL, Daniel. Don't just say that.
**Eldad - 00:20:32:**
Of course, Benjamin, you love writing physical plans, but yes, and sometimes SQL on top. And people, but you know, Benji, people have been using all sorts of tools just to avoid writing SQL. And, for all the good reasons. But now that you have new kinds of tools generating SQL, it's just a great way to reason. We see many startups. We see, as Benji mentioned, startups that apply SQL on all sorts of data, not necessarily through a database. But really try to abstract different kinds of data through SQL without talking about the database in the middle, right? And the database becomes a component. So from an AI perspective, if you're an AI engineer, you're not necessarily thinking about the databases you're used to. You're thinking about databases, just access to great data. And the AI takes care of a lot of stuff. So again, this is very forward looking. And that episode was, I don't know, a few months back. We love catching up with our Daniels. So kind of see where those theories got us, see kind of what really happened. And it's always funny to see kind of, right, how we're so bad in predicting the future, yet are always right and, you know, following the right path. So I love it. I love to see how data movement and AI come together. By the way, I don't think it's a prompt saying move from here to there. I think it needs to be a conversation where you're actually trying to tweak. You said it. There is no way for something to just do it automagically. There's iteration and understanding and figuring out the data and the quality and the metadata, right? There needs to be, everything needs to fit. So from a user perspective, instead of going through a workflow, like a Visio kind of experience and learning how to do pop-ups and everything, taking that same thing and applying through a conversation means the user just becomes much smarter, right? So the same user now becomes a genius. They can use the right knobs. They can tune it. They can exploit the system. So if you guys are better at using specific features of a specific driver. And know and spend hours, days, months perfecting it to become better than everyone else, then your biggest challenge is exposing it to users. Because the user will need to drill in so much, so deep. To actually find that advantage. To the point that they will never find that advantage. And what we're seeing is with AI, users become geniuses. They are now able to figure out the knobs, to figure out the weights. And it's really, again, it's untouched territory. No one has figured it out. And that's what makes it interesting. But definitely my biggest takeaway is that it's not replacing users. It's making them super smart. If you serve it in the right way through that, and you have the right product, it has those knobs, then suddenly you see amazing results. And I think data movement, just as any part of the data funnel, they need efficiency. They need optimization. They need specifics. So I hope to see that GAC bridge being specific about solving problems and AI helping you get there. Of course, we're all generics, right? We don't have the time to drill into the details.
**Daniel - 00:23:52:**
Yeah, I think that's a very eloquent way to put it. Yeah, on the other side of the coin, I've seen people feel dumb when they are trying to solve a problem. They are met with this huge list of configuration options and they have no idea what to do, all this new terminology and everything. And that's definitely a gap that AI can fill and actually empower a user by asking the right questions and providing a friendly interface. So yeah, for sure, I can definitely see that happening.
**Benjamin - 00:24:19:**
Great. So connecting industry trends, right, we talked a lot about AI now. I saw one of your recent LinkedIn posts about whether Iceberg is the new Hadoop. Tell us, maybe let's close.
**Eldad - 00:24:29:**
What an insult. What an insult to Iceberg.
**Benjamin - 00:24:32:**
What an insult. No.
**Daniel - 00:24:34:**
To what?
**Benjamin - 00:24:35:**
Let's maybe kind of close out on this. And especially, I think, for you as a data movement tool, right? Like, I guess Iceberg is also changing how companies and organizations and data engineers think about data movement between different platforms. It's like, yeah, how's Estuary? How are you thinking about Iceberg these days?
**Daniel - 00:24:53:**
So we see Iceberg as something that is very popular. I think it's more popular than actually used. A lot of people are talking about it and not that many people are actually implementing it as like a lake house solution. But I think it will get there. Like, it's definitely on the track. There was the table format wars that kind of ended with all the different table formats like Hudi or Delta tables. And it kind of went in the direction of Iceberg, which I think is a good step for the industry to standardize on one solution, which gets more attention, more development work. Because it's open source. Many other contributors are able to put work into it. And because of this, it will become, in my eyes, the de facto solution in the near-term future for data lake houses. I think lake houses themselves are the obvious next step in the evolution of a lot of data stacks. For a lot of use cases. It is the peak of decoupling compute from storage, in my opinion, which is a good trend for flexibility, for cost management. And a lot of use cases are covered by it. So we see that a lot of people want to or are interested in migrating to Iceberg. Like they have a Snowflake database or data warehouse or...
**Eldad - 00:26:14:**
You know why they all want to migrate to Iceberg? Because they want to try Firebolt without having their boss snooping around and asking them questions. Is it secure or did we move data out? Right? Did we move? That's the worst nightmare, moving data out. You're right. Iceberg is an enabler waiting to explode, waiting to enable. And 2025, I am super confident we will see Iceberg going from big promise to being an enabler. And the reason is, again, that those two characters are, the reason is AI, is because... People have been spending so much on building the modern data stack and connecting all the tools and figuring everything out. Now AI comes and AI gravitates the persona. Now engineering is back in the house. Now companies take their best teams and say, you work on the next AI infrastructure stack while all the legacy runs. Some of them call cloud their warehouses like legacy. Okay. And the reason is, again, like focus and shifting. Now the most valuable data sits in those repositories because it is a business data. It is well-cleaned. It is accurate. It's what they call great data. It's fresh, accurate. Now just, you know, flipping it, the Iceberg magic button means that many other products can come in and use that data. If you're into AI and you want to build a new stack, you don't necessarily need to think, oh, I'm using Snowflake. So let's just go and register to Snowflake AI. You can use your own AI. You. You can use any AI. You can use, one of other 20 options but you need the data and you need that contract with the data warehouse, or in that case, Snowflake. So that clones the outcome of those previously modern data warehouses and allows a new stack to grow. And obviously when it comes to AI, efficiency is king, moving data is king, because in many cases don't need all the cleansing. It's AI. So sometimes people say, I'm willing to get it less structured, less cleansed, less modeled, because I can figure it out as part of the AI feature. So a lot of interesting things, kind of what it means about centralized schema management, but Iceberg is just in a neighbor. So it allows engineers and data engineers to kind of expand that experimentation without restarting, without losing their previous investment. So I think it's a great opportunity for the ecosystem to evolve around it. And that's the only thing we see over the last six months. Like a lot of focus on trying to figure it out. And I hope everyone will figure it out in 2025.
**Daniel - 00:28:57:**
Hopefully. I think the only issue is that it's too complicated currently. You have to deal with catalogs, maintenance, like all kinds of stuff that nobody wants to deal with and shouldn't have to deal with. So once those ironed out...
**Eldad - 00:29:08:**
A new hadouk. This is what Benji said.
**Daniel - 00:29:11:**
Exactly.
**Benjamin - 00:29:12:**
Exactly. As an engineering team building Iceberg support into a system, it is very complicated. Yes.
**Daniel - 00:29:20:**
Yeah, it's rough. Yeah. But once those are ironed out, it's going to be really the democratization of data for many, many organizations.
**Benjamin - 00:29:26:**
Awesome. Cool. Dani, it was amazing how you learned so seriously, like such an amazing kind of hearing your perspective from Estuary, connecting it to how we think about data warehousing. So many exciting things happening in tech right now. Any closing words from your side?
**Daniel - 00:29:43:**
I think this is pretty much the golden age for data-related work and data engineering. I think if anybody watching this is thinking about jumping into the data role, then I would say do it. A lot of people ask me if this is the right time. And yes, this is the right time. Like data is only going to get bigger and bigger and we need more and more data engineers and data analysts and other practitioners. So yeah, go for it.
**Benjamin - 00:30:07:**
Awesome. Thanks for being on the show, Dani. And we look forward to catching up later this year to see which of our predictions from today actually turned out.
**Daniel - 00:30:15:**
Yeah. Thank you so much. This was great.
**Eldad - 00:30:17:**
Thank you. Take care. Bye-bye.
**Daniel - 00:30:18:**
Bye-bye.
**Intro/Outro - 00:30:22:**
The Data Engineering Show is brought to you by Firebolt, the cloud data warehouse for low-latency analytics. Get $200 credits and start your free trial at firebolt.io.
# Architecture and Internal Representation of the GEOGRAPHY Data Type (Part II) (/blog/architecture-and-internal-representation-of-the-geography-data-type)
In [part I](https://www.firebolt.io/blog/building-geospatial-support-in-firebolt-part-i) of this series of blog posts, we talked about the differences between GEOMETRY and GEOGRAPHY, as well as the S2 library that powers Firebolt's GEOGRAPHY data type. In this post we'll take a deeper look into how Firebolt processes and stores geospatial data to enable fast and efficient query execution.
This post will cover:
1. How geospatial data is validated and normalized upon ingestion
2. How Firebolt structures and stores GEOGRAPHY objects using S2 cells and shape indexes
3. How Firebolt optimizes query performance with geospatial data pruning and partitioning
By the end, you'll have a comprehensive understanding of how Firebolt prepares and optimizes geospatial data behind the scenes. Let's dive in.
### Validating and normalizing inputs [#validating-and-normalizing-inputs]
Geospatial data is ingested into Firebolt through one of three standard formats: [Well Known Text (WKT)](https://en.wikipedia.org/wiki/Well-known_text_representation_of_geometry), [Well Known Binary (WKB)](https://en.wikipedia.org/wiki/Well-known_text_representation_of_geometry#Well-known_binary), or [GeoJSON](https://datatracker.ietf.org/doc/html/rfc7946). These inputs are immediately converted into Firebolt's native representation based on the S2 geometry library.
A GEOGRAPHY object in Firebolt can contain different types of geometric components, such as points, line strings, and polygons, and in the case of a GEOMETRYCOLLECTION, a combination of these. All components of a single GEOGRAPHY object are stored inside an [S2ShapeIndex](http://s2geometry.io/devguide/s2shapeindex.html). A shape index is S2's most generic abstraction. It can store any combination of points, line strings, or polygons. S2 also provides functions to compute relations between two shape indexes, such as containment, and intersection which we use to implement the functions exposed through Firebolt's SQL interface.
However, many functions working on shape indexes require that within a shape index, no component intersects a polygon's interior. Firebolt ensures this in a normalization step by merging intersecting polygons, removing any component that is completely contained in a polygon, and cutting line strings at polygon boundaries. This ensures that later operations can be as fast as possible because they don't need to validate the input shapes. In some cases, like wrong polygon orientation, or self-intersecting polygons, Firebolt also fixes invalid inputs while converting them into the S2 format. Our [documentation](https://docs.firebolt.io/sql_reference/geography-data-type.html) covers the cases where Firebolt fixes invalid inputs or applies normalization in more detail.
The example below shows an example of a GEOGRAPHY object before and after normalization. Both polygons are merged and the line string is cut into three pieces.

### Storing S2 cells to enable fast functions [#storing-s2-cells-to-enable-fast-functions]
Geospatial queries, especially those involving intersection, containment, and distance calculations, can be computationally expensive. A naïve approach would require scanning every geospatial object in a dataset to determine which ones match a given query—an operation that does not scale well as data volume grows. Efficient indexing and filtering mechanisms are essential to keep query execution fast and scalable.
Many of the performance optimizations Firebolt uses for geospatial data use [S2 cells](http://s2geometry.io/devguide/s2cell_hierarchy) and [coverings](http://s2geometry.io/devguide/examples/coverings). S2 cells are a hierarchical partitioning of the Earth's surface, designed to represent spatial data at various levels of precision. The example below shows two of 6 cells of the highest level of the hierarchy. The right cell is further subdivided 4 times into lower level cells.

A covering, in this context, refers to a set of S2 cells that together encompass a given region. For example, in the image below, a polygon outlining Florida is approximated by the shown covering of 22 S2 cells at different levels of the hierarchy.

One of the key advantages of using S2 cells is the efficiency of intersection and containment tests. Since each cell is represented by a single integer, determining whether one cell intersects with another can be done using simple integer comparisons, which are extremely fast.
After creating a valid shape index, we compute a single S2 cell that fully covers the entire shape index\*. This cell is persisted in our storage format and used to speed up function evaluation on GEOGRAPHY objects. You can learn more about these optimizations in part III of this blog post series.
### Data pruning for geospatial predicates [#data-pruning-for-geospatial-predicates]
Data pruning is one of the most important performance optimizations in a database. By using aggregated information about large parts of the data, Firebolt can often quickly deduce that this data does not need to be read at all. For types like integers, timestamps or even strings, we compute statistics such as the minimum and maximum value of a column within a [tablet](https://docs.firebolt.io/Overview/data-management.html#creating-tables) while writing to Firebolt storage. These statistics can then be used for pruning during query execution, i.e. skipping entire tablets if all the data inside is filtered out by the WHERE condition of a query. Because geospatial objects cannot be ordered, there are no minimums or maximums for GEOGRAPHY columns that could be used for pruning. Instead, we compute a covering of S2 cells that covers the entire column within the tablet.
The coverings stored for each tablet can be used for queries using geospatial functions as filters. Currently, these are ST\_COVERS(A, B), ST\_CONTAINS(A, B), and ST\_INTERSECTS(A, B), where either A or B is a constant, and the pattern ST\_DISTANCE(A, B) \< C, where C is a constant number and either A or B is a constant GEOGRAPHY object. Firebolt's query optimizer recognizes these patterns and adds an additional filter that skips a tablet if its covering does not intersect the covering of the constant input. For the ST\_DISTANCE(A, B) \< C pattern with constant A and C, the covering for A is constructed such that it contains everything within a distance C of A.
The image below illustrates two tablets that contain all of the displayed points. The purple and gray boxes represent the coverings stored for the tablets. When a query is executed to retrieve all points within the polygon shown, the red covering is computed. The purple covering corresponding to the first tablet overlaps the red covering, so the first tablet is scanned. The gray covering corresponding to the second tablet does not overlap the red covering, so the second tablet is completely skipped.

### Improving pruning performance using partitions [#improving-pruning-performance-using-partitions]
Depending on how you ingest data, tablets might have objects scattered across the globe which can severely impact the effectiveness of pruning entire tablets. In order to improve this, an option is to partition the data by a spatial attribute using the [PARTITION BY](https://docs.firebolt.io/sql_reference/commands/data-definition/create-fact-dimension-table.html#partition-by) clause when creating a table. This could be a zip code (or a prefix of the zip code), or the ID of a point's S2 cell at a given level using the [ST\_S2CellIdFromPoint](https://docs.firebolt.io/sql_reference/functions-reference/geospatial/st_s2cellidfrompoint.html) function. By partitioning the data spatially, every tablet only contains objects from a single partition so the tablet pruning conditions can work at full effectiveness.
However, it's important to avoid creating too many partitions, as this can negatively impact performance. Ideally, the partitioning key should have a moderate number of values—around 100 partitions is a reasonable target. For zip codes, this could involve using the first two digits, while for S2 cell IDs, the appropriate cell level can be determined based on the [average area of cells at that level](http://s2geometry.io/resources/s2cell_statistics). For example, when dealing with points distributed across the land mass of the Earth, a cell level of 3 would be suitable, whereas for points within the USA, a cell level of 5 would be more appropriate.
### Outlook: GEOGRAPHY in the Primary Index [#outlook-geography-in-the-primary-index]
Currently, Firebolt does not support GEOGRAPHY columns in the primary index. In the future, when this feature is added, geography data will be sorted using a method called a space-filling curve. Space-filling curves work by assigning every point on Earth a single number in a way that keeps nearby points assigned to similar numbers. This organization helps group related data together, making it faster to perform operations like checking whether areas overlap or whether a point is within a specific region. The same space-filling curve is also what powers S2 cells and enables fast intersection and containment checks using simple integer comparisons.

By using this ordering within a tablet, each [range](https://docs.firebolt.io/Overview/data-management.html#creating-tables) contains only a small region. We can then store an S2 cell covering of all rows within the range. This covering is then used to efficiently prune ranges in the same way as we prune tablets. In the future, we will also use S2 cell coverings to support fast spatial joins where optimized data structures on top of S2 cell coverings can help to quickly find join partners.
Now that we have learned about the fundamentals and data format of Firebolt's GEOGRAPHY data type, we will dive into the details of snap rounding and performance optimizations for spatial operations in part III. In the meantime you can try out Firebolt's geospatial functionality yourself using our [demo project](https://github.com/firebolt-db/firebolt-demo/tree/main/geospatial).
> \*Finding a single cell that fully covers a shape index is not always possible because the largest cells cover only 1/6th of earth. In these rare cases, Firebolt can also store multiple cells for a single shape index.
# Automatic Cache Warmup (/blog/automatic-cache-warmup)
Firebolt relies very heavily on caches to deliver low latency query performance. When engines start up, all caches are empty. We call an engine in this state "cold". As engines serve requests, caches warmup and query performance improves. To speed this process up, Firebolt introduced the AUTO\_WARMUP feature. This enables engines to remember the state of their caches, and proactively retrieve that data when clusters start.
## Implementation [#implementation]
Auto warmup consists of two components — snapshotting SSD cache state and reloading cache from snapshot.
### Making a Snapshot [#making-a-snapshot]
Explaining how this works requires looking under the hood of your data on Firebolt. Let's take a simple table with two columns:
```sql
CREATE TABLE my_table
(
id INTEGER,
name TEXT
) PRIMARY INDEX id;
INSERT INTO my_table (id, name)
VALUES
(1, 'Alice'),
(2, 'Bob'),
(3, 'Charlie'),
(4, 'David'),
(5, 'Eve'),
(6, 'Frank'),
(7, 'Grace'),
(8, 'Hannah');
```
The primary unit of storage on Firebolt is the tablet. As data is inserted, it is combined into tablets, and column data within each tablet is generally stored separately to enable fast aggregations and scans. The tablet also contains sparse indices that allow efficient access for queries with selective filters. The above data could for example be laid out as a couple cartoonishly small tablets:
Now let's make some queries to this table:
```sql
SELECT MAX(id) FROM my_table;
SELECT * FROM my_table WHERE id = 3;
```
To serve these queries, we need to fetch a subset of the data above. We need to fetch all the data for one tablet to serve the SELECT \*, and we need to fetch id data for all tablets to serve the SELECT SUM(id). We'd then expect our SSD cache to look something like this:
```sql
tablet1/
id
name
tablet2/
id
other_old_tablet/
foo
bar
...
```
We continuously track when tablets in the cache are used to ensure we handle eviction properly. We can use the data from our evictor to construct a list of recently used tablets and columns on every node in our cluster. We use this data to push a snapshot for our engine to s3:
We continually push an updated snapshot like this to cloud storage while the engine is running. When it shuts down this data will be persisted for up to two weeks. When we then restart that same engine, it will be able to retrieve that snapshot and know which data to fetch.
### Reloading Cache From Snapshot [#reloading-cache-from-snapshot]
With the snapshot prepared above, reloading our cache is now fairly simple. We simply fetch the snapshot if one exists, then load the tablets in the snapshot from least to most recently used until one of a few conditions is met.
* We've loaded everything: this is our best-case scenario, and means our cache is fully restored
* We hit our timeout: by default we're willing to spend only 30 minutes loading the snapshot, but this can be adjusted as needed.
* Our SSD cache is filling up: if you switch from a large node to a small node, or storage optimized to compute optimized, there might be less disk cache available. We abort the warmup if we see the SSD cache is filling up to ensure there's plenty of room to cache data for user initiated workloads.
Since everything at Firebolt runs on SQL, we use queries under the hood to fetch this data. We restrict the number of threads on these queries to keep the CPU overhead of the cache load around 10%.
## Applications [#applications]
What does this feature mean for workloads in practice? Let's look at a few examples.
### Online Upgrade [#online-upgrade]
One place where auto-warmup is extremely helpful is in [online engine upgrades](https://www.firebolt.io/blog/live-engine-upgrades-zero-downtime-the-firebolt-method). When we upgrade an engine cluster, we create and mirror traffic to a "shadow cluster". Using auto-warmup, shadow clusters can be warmed up much faster, since they proactively fetch a recent snapshot of the main cluster's cache, rather than relying entirely on mirrored requests to greedily warmup.
Let's look at cache on a high traffic engine during our 4.23 release.
The new node starts up at 16:20 (blue line), and as it receives queries it starts caching data. After an hour, the engine is upgraded, but its cache still hasn't fully converged with the old cluster it's replacing (red line).
Let's compare this to the 4.24 release where auto-warmup was enabled.
We see the same engine start an upgrade at 13:50. By 13:58, the snapshot is fully loaded, and all queries seen during the duration of the shadow period hit cache.
There are a couple major benefits of this approach:
* The final cache is much warmer. This minimizes the risk of a performance degradation to the customer workload after the upgrade. This also ensures we warmup data that wasn't touched during the upgrade period.
* The upgrade took less time. The first upgrade timed out after an hour, and probably should have been run for even longer to fully warm the cache. The second upgrade finished after just 35 minutes when we saw performance fully converge, and could have been promoted even sooner. Firebolt pays the bill for shadow clusters, so there's no direct cost implication for customers, but this lowers our cost of operating engines, and lets us keep Firebolt affordable.
### User Workloads [#user-workloads]
Auto-warmup is clearly beneficial in preparing shadows to receive traffic, but what about running this during an intense user workload? To benchmark this, I'm using a modified version of the [Firescale benchmark](https://www.firebolt.io/blog/firescale-benchmarks-a-deeper-dive) dataset. The benchmark consists of 50 concurrent requests, each picking one of 100k random point read queries against a 1 TiB table, repeated indefinitely.With no auto-warmup, we see a logarithmic decrease in data fetched from s3 throughout the benchmark, as a higher and higher percentage of queries hits cache, but even after 30 minutes, we're sometimes pulling data from s3 as we serve certain queries for the first time.After recording a snapshot from a fully warm engine and trying again with auto-warmup enabled, we can see the difference in the s3 read pattern. The engine is still under a heavy query load and is greedily pulling data to answer those queries, but it is also proactively fetching data it had touched in the previous run. This leads to substantially elevated read traffic initially, but means we eliminate cold reads much faster.
Over the course of the 30 minute benchmark, the engine with greedy warmup averaged 91 QPS and the engine with auto-warmup was able to average 194 QPS as faster reads led to increased throughput.
### Auto-Scaling [#auto-scaling]
When engines [auto-scale](https://docs.firebolt.io/guides/operate-engines/understand-autoscaling), new clusters in the engine come up cold. Since the oldest cluster in the engine is maintaining an up-to-date snapshot, the new cluster has a recent snapshot to fetch from when it starts up. Firebolt's load balancer routes to engines based on reported load, so we will already avoid swamping slow clusters, but auto-warmup will much more quickly bring the performance of new nodes to parity with old ones.
## Try It Out! [#try-it-out]
If you think your workload would benefit from auto-warmup, give it a try. You can enable it on an existing engine by running.
```sql
ALTER ENGINE myengine SET AUTO_WARMUP = true;
```
# Beyond Database Optimization with AI (/blog/beyond-database-optimization-with-ai)
In this episode of The Data Engineering Show, the bros welcome the CEO DuckDB Labs and co-creator DuckDB, Hannes Mühleisen. They delve into the groundbreaking journey of DuckDB, an analytical database that processes billions of queries every month. Learn why DuckDB prioritizes broad compatibility over specialized optimizations, how its extension model works and the emerging solutions for database technology in the age of AI.
Listen on [Spotify](https://spoti.fi/41C0wJx) or [Apple Podcasts](https://apple.co/43T0wHA)
### Episode Highlights [#episode-highlights]
##### The Purpose of DuckDB (01:04) [#the-purpose-of-duckdb-0104]
Hannes gives a full description of what DuckDB is as well as what it is designed to do. He describes the tool as one that understands SQL and is specifically designed to simplify complex analytical use cases.
##### SQLite vs DuckDB (02:53) [#sqlite-vs-duckdb-0253]
Hannes compares two different tools stating that SQLite is an amazing system that is not meant for analytical queries but for transactional use cases while DuckDB is specifically designed for that exact purpose - analytical use cases.
##### The Importance of Collaboration (08:14) [#the-importance-of-collaboration-0814]
Hannes states the need for community collaboration as the database engine space seems to have hundreds of brilliant people trying to solve the same problems. He shares his profound admiration for a team in Munich, praising them for their exploits in implementing concepts only described in paper.
##### The Component-Based Architecture of DuckDB (11:25) [#the-component-based-architecture-of-duckdb-1125]
Hannes highlights a special feature in DuckDB, that is, it can be used as a component and he explains that the in-process architecture is a success because of the memory of data sharing that can be achieved.
##### The Parquet Reader Journey (17:51) [#the-parquet-reader-journey-1751]
Hannes explains how he built his Parquet Reader out of necessity, although he would have preferred not to. He shares how a creator named Ove Korn from Germany donated the reader to a project named "The Arrow Project" and managed it to the degree that the entire project depended on the use of the Parquet Reader and it became an issue to use both independently. Hannes adds that a parquet reader that is competent has no choice but to become a database engine which is one of the interesting things about development.
##### The Role of AI in Database Interaction (22:41) [#the-role-of-ai-in-database-interaction-2241]
Hannes states that he doesn't think that AI has a place in a database engine but rather, it is needed for optimization because the researchers who built their careers on optimization are out of jobs. He explains that the role of AI should be for assistance tasks and not for a total execution.
##### SQL - A Defined Interface (29:20) [#sql---a-defined-interface-2920]
Hannes introduces us to a tool that allows us to pro-programmatically build a query called relational API stating that it helps to simplify the tasks of a programmer. Although, Hannes agrees that using a well-defined interface is important for components like databases, he also argues that SQL can provide a relatively defined behavior within a single system.
##### The Golden Age of Database (38:57) [#the-golden-age-of-database-3857]
Hannes concludes the episode by appreciating Firebolt and other engineers for taking on core engine tasks. He shares his excitement for the golden age of databases where there is a showcasing of what is possible.
# Bill Inmon, the Godfather of Data Warehousing (/blog/bill-inmon-the-godfather-of-data-warehousing)
As people in the data industry go, Bill Inmon is among the top, often seen as the godfather of the data warehouse. In this Data Engineering Show episode, Bill Inmon talks about surviving rabbit holes throughout the evolution of data, the data modeling renaissance, and why ChatGPT is not Textual ETL.
Listen on [Spotify](https://open.spotify.com/episode/3IvsspPtiYeU57McqDCB9k) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/bill-inmon-the-godfather-of-data-warehousing/id1561927688?i=1000623790403)
**Benjamin:** All right, cool. Hi, everyone, and welcome back to the Data Engineering Show. It's a super exciting episode because we have a total data celebrity joining in today. Bill Inmon, thanks for being on the show.
**Bill Inmon:** It's my pleasure to be here.
**Benjamin:** Awesome. And we also have a special guest as co-host today. Robert Harmon, kind of veteran data practitioner. He's an SA at Firebolt right now and apparently a recovering race car addict. So good to have you on the show as well, Robert. Cool. Do you want to kind of intro Bill? I mean, he really needs no introduction because I think all the listeners will know him, but let's still do it.
**Robert Harmon:** Yeah, it's almost humbling to try and introduce Bill. As people in this industry go, he's obviously among the top, often seen as the grandfather or the father of the data warehouse. I like to think of him as the data OG. He's really the guy that set the stage for the rest of us to have careers. And obviously for that, I'm grateful.
**Bill Inmon:** Robert, I like to think of myself not as the grandfather of Data Warehouse or father of Data Warehouse, but as the godfather of Data Warehouse.
**Benjamin:** love that. Nah.
**Robert Harmon:** That gives it much more of an almost infamous tinge, and I think it's fitting.
**Bill Inmon:** Yeah.
**Robert Harmon:** So welcome Bill, how you doing?
**Bill Inmon:** I'm fine.
**Robert Harmon:** Well, so in preparing for this, of course, I was completely out of my mind because it's like meeting Taylor Swift. I mean, this is huge for me.
**Benjamin:** I'm going to a Taylor Swift concert next year. So it's like perfect, perfect. You're giving the example
**Robert Harmon:** Oh.
**Benjamin:** because I'm super stoked about the Taylor Swift concert and also about today's episode.
**Robert Harmon:** So I'm preparing for this interview and I'm thinking, well, what the heck do I ask Bill Inmon? Because if I ask any of the normal questions, well, then I look like I don't know what I'm doing for a living because why would I be asking these questions? And then if I ask the questions I wanna ask, well, everybody's gonna be lost because we've been doing this too long. So I'll lose all of the new practitioners. And just serendipitously. So, you know, one of my contacts came up with a question that I saw on social media that was, that kind of triggered every, you know, my mindset on where to go. And he asked, and he stated, being one of them, if you started your data career post 2010, you're at a massive disadvantage. And this is one of those things that I've been thinking about a lot offline is that not so much the technology or the products or any of that. more the social ramifications of our industry. And I think that's one of those statements that kind of brings that out. So, I'm thinking about generational issues and sometimes I'm thinking about inclusivity issues because a lot of the guys in our industry look a lot like you and me, Bill. So, I kind of wanted to explore this idea that people before, that joined the industry prior to 2010 might be missing a few things. And somehow this opened it up to an idea of, well, hey, I've got Bill here. I've been around since longer than then. Let's, let's try and educate some of the younger guys on some of the stuff they've missed prior to them showing up. Does that make sense?
**Bill Inmon:** Yeah, let me kind
**Robert Harmon:** What?
**Bill Inmon:** of back up and give some perspective.
**Robert Harmon:** Sure.
**Bill Inmon:** The way I look at things is, what we are experiencing today and what we
**Robert Harmon:** Mm-hmm.
**Bill Inmon:** experienced yesterday is nothing but a big evolution. The evolution really started about 1960, which is probably before most of you were born. But in 1960, the world became introduced to the computer. And at that point in time, there was no one, not one person that had any background. It was all fresh, fresh material. And so what we've been witnessing for the past 50 years or so is an evolution. And I have to agree with the person that asked the question. Um, have they missed something? Uh, the answer is yes. They've missed many of the evolutionary, uh, rigmaroles we've had to go through. Uh, uh, but, but is there still opportunity out there? There is, we haven't even begun to start to explore the opportunity that's out there. So yes. Uh, you, you who have just joined have missed. a lot of these struggles, a lot of the fairly nasty stuff that has occurred in the past 50 years, but is there opportunity in front of you? The answer is absolutely yes. Now if you were to ask me where is the opportunity, there's really one answer and the answer is business value. the people that go and find true business value for their company, their corporation and themselves are going to be the people that advance the people that are the most valued in the corporation. And is there business value out there to be found? We haven't even started on business value. So I look at it from a larger evolution and yes, I agree that people have missed some things. Quite frankly, a lot of what they have missed, a lot of what they've missed has been worthwhile missing. I can recall some really stupid things that people said years ago and people believed and I remember when we were told that... secretaries are going to be writing code. That's one thing that was said. I remember somebody saying, we need to develop application programs without programmers. Now, I don't know how that person thought that was gonna happen, but that was once what we thought. And there's been a long list of things that we've been told that to be... turned out to be totally false, but our industry's gone down. And so some of the things that people have missed have been very worthwhile missing because you didn't need to go through the pain that we all went through. However, having stated that, Whatever you do, if you want to get ahead, if you want to take advantage of where we're at, two words, business value. Take technology, done. Once upon a time, we looked at technology for technology's sake. We can't do that anymore. We've got to look at technology as a way of enhancing business value. And once you understand that, then, then you've lost nothing by joining our profession when you've joined.
**Benjamin:** I love that as an entry question Rob, because obviously I am someone who joined the data industry post 2010. So Bill, like for someone who's new in this space, right? You said, okay, there's kind of quite a few ideas in the past, which really didn't pan out and technology goes in cycles, right? So protect me from having the same ideas basically again and reviving them. Do you want to tell some maybe war stories from, yeah. things that really didn't pan out, but that people were excited by in the past.
**Bill Inmon:** You don't want me to start to talk about war stories because I've lived through all of them. But yeah, there's been some, one simple thing that happened long ago. When we were learning to program, we were told that programming would all be simple and easy if we didn't use go-to statements. Now, that's humorous today, but- But there was an element of truth there that the code that was being produced was really quite fragmented code. And indeed, not using the go-to statement did improve things. But was it the solution? Was it a silver bullet? And the answer is absolutely not. And that's one of many things. This other bit about... Even secretaries can start the program. That was a ridiculous idea that somebody that said that didn't know what they were talking about. And I have nothing against secretaries, by the way. My family has secretaries in the family. But there's a skill set and a mindset that's needed for coding. that you normally don't find in secretaries. And so that's another one. And every few years, our industry comes up with what they call the silver bullet. And I remember when IBM told us, man, if you just go DB2, that your problems are gonna be solved by using DB2. Well, that... that didn't work out well either. And so, and that's being polite about how poorly that worked out. So our industry goes nuts over these silver bullets that said, gee, if you just do fill in the blank, if you just do whatever, then everything's gonna be okay. And I remember when we were told, man, going big data. You've got to go get hadoop. You've got to go and do big data. And that's going to solve your problem. Well, guess what that didn't. So I think someday I may actually write a book on all of the rabbit holes that our industry's gone down. And it's a wonder we all survived. Now, part of the fallout from that is that the IT organization has lost great amounts of credibility. And if you don't believe me, go into an IT organization and find out who's in charge of the decisions and who's in charge of the budget. Once upon a time, that was the IT organization. Today, you find that it's the end user. It's marketing, it's sales, it's finance. Those are the people. that are because they trusted IT to help them. And IT kept going after these silver bullets and the corporation no longer trust the IT department. I'm sad to say, in fact, it pains me to say that, but it's the truth.
**Robert Harmon:** And, you know, I can only augment that Bill because I've lived a lot of that myself. I was a practitioner working in data teams for what, 25 years. And I've seen all of this. Um, I've also seen what we used to call rogue development where the end users and start running off in their own direction to solve their own problems. And that's all in my eyes. That's always been a strategic threat to an IT organization, because if the end users are busy solving their problems around you, obviously you're not doing your job right.
**Bill Inmon:** Yep.
**Robert Harmon:** which comes back to your first statement. We need to deliver value as an IT organization or we'll just be ignored. And if we're ignored, then why are we here? Now, you did mention something and the quote that I brought you was from Mark Freeman, just so for clarity sake, so everybody knows I'm not stealing his work. I think why he picked 2010 was the big data explosion.
**Bill Inmon:** Yep.
**Robert Harmon:** And... But from my memory about that time, and maybe my memory is going on me, but about that time, we saw a huge employment increase in our data profession. A number of new people came on board. So there was a swell there among a number of companies. So I really think that might be what he's referring to. What I'm seeing lately, though, is more noise about going back to the pre- big data world where we're starting to hear people talking about modeling again and interesting ideas like maybe we should have constraints. When do you think that it's finally coming back?
**Bill Inmon:** Modeling is not something that is associated with any particular discipline. Modeling is something that is useful in lots of places. And let me ask you this, if you were to build a cabin in the mountains, would you need a blueprint? And you either would need a blueprint in your head, or you would need a blueprint that somebody built for you, but you would be dumb to go off into the woods and build the cabin without having a plan, a model. And so a data model has widespread usage. It's not the IT organization that owns data modeling. It's everybody owns data modeling because we need that. And so today, the end user is waking up the whole motivation for doing new systems. Once upon a time, the IT organization built new systems. Today, it's vendors and end users that are building systems. And the advent, the waking up of the data model, the end users are discovering, oh my gosh, I'm building something that's complex here. I need a data model. So I think that the renaissance of data modeling has come from the awakening of the end user that they're the ones that are building something and they need the data model.
**Robert Harmon:** That's really an interesting perspective that I hadn't quite processed yet. I'm not sure I have a response. I'm gonna have to sit and think about that for a day or two to try and work out all the specifics. But no, I think that's a very valuable observation. So my next question, so we talked about the past a little bit. Are there any advancements currently going on that you're... particularly excited about.
**Bill Inmon:** Well, I have to preface this with a disclaimer that I'm involved with this, but yes, I think there, in fact, I've been asked by a number of students graduating college, I wanna start my career, where should I start my career?
**Robert Harmon:** Mm-hmm.
**Bill Inmon:** And let me answer the question that way. I liken the opportunity. for business value in the world of text to be like California in 1848. We are told by the historians that in 1848, you could walk down to the streams of California and pick up gold. You didn't need a shovel, you could just pick it up out of the stream. And it was there waiting to be found. And I think that the... corporation after corporation is letting text go through their hands and doing either nothing or very little with it. And I think that there is tremendous opportunity there. And I could outline some of the opportunities. One of them is in terms of sentiment analysis, of understanding what your customer is saying. Another one is in terms of medical records. That medical records is kind of an interesting case. I don't know if you've ever taken a look at or had the opportunity to look at medical records, but medical records are written in the form of text. Now, the medical records that we have today are designed for and good for one doctor and one patient. So when a patient is being looked at by a doctor, the medical record is there for them. But what, because it's in the form of text, what a medical record isn't designed for is the whole notion of looking at a hundred thousand patients at the same time. So if we have something like COVID come along and we need to look at a hundred thousand patients, you can't do it if the information you have is in the form of text. The only way that I'm aware of that you can do it is if your text has been transformed into a database. And once your text has been transformed into a database, then you can start to ask the question. Interesting questions. Let's take COVID. How does COVID react to people who smoke? How does COVID react to people that take certain medications? How does COVID react to gender? Are men more affected by COVID than women? How does COVID react to age? What role does that play? And answering those questions are very, very important. But as long as your information you have going through your healthcare system is in the form of text, you can't do, not easily, you can't do. that kind of analysis. So whether it's sentiment, I'll tell you another one that is one of my pet peeves. And that one is corporate contracts. When we first started doing what we were doing, I used to talk to groups of executives and I've asked groups of executives, how many people in this room know what's in your corporate contracts? And to date, not one executive has ever raised their hand and said, oh yeah, I know it's in our corporate contracts. And then I say, well, I guess I'm kind of confused because you guys are corporate executives. Aren't you in charge of liability? Oh yeah, Bill, liability, risk management, we've got that down cold. Well, I said, well, I guess I'm really confused then. You tell me that you have risk management and... liability to your corporation taken care of. On the other hand, you tell me you don't have any idea what's in your corporate contracts. And I said, don't you think in your corporate contracts that there's liability? And the truth of the matter is every corporation out there has got liability in their corporate contracts. That's one of the major purpose of a corporate contract. And at this point in time, The executives say, oh, well, we can't do that. I said, well, why can't you do that? And they say, well, Bill, you see, we've got a million contracts. And each one is different. And we can't possibly know what's in our contracts. And I said, oh, no, you can. Let me show you how you can do it. And trying to convince an executive that they should look at their corporate contracts. is like trying to sell caskets to living people. People only buy a casket when they need it when they're being laid in the ground. And so I got tired of talking to executives, but indeed corporations can know what's in their core. And by the way, from a standpoint of business value, do you think there's business value wrapped up in knowing what your corporate contracts is? there is huge amount of business value wrapped up in those corporate contracts and nobody's looking at them. So that's a that's you know you say where's the opportunity? The opportunity is in tech whether it's corporate contracts or medical records or sentiment analysis. By the way there's a lot more than that that's the tip of the iceberg.
**Benjamin:** Like one thing I'm curious, right? Like kind of, it makes kind of perfect sense, uh, kind of in terms, in terms of the problem statement, and I totally buy into that. So I'm actually kind of a guy who builds databases, right? So I kind of build query parsers, kind of execution engines and so on. And one thing I'm curious about now is like, how does a data processing system look? Right? Like kind of how does the data stack look in the world where you have all of those? kind of just text-based records, right, in terms of query language, in terms of the systems, etc.
**Bill Inmon:** Okay, I'm going to warn you, you ask the question, I'm going to give you the answer. This is going to sound to be self-serving, but you're the one that asked the question. There is technology out there called Textual ETL. Textual ETL takes the raw text and turns it into a database. Now, in order to do that, it's taken a long time for myself and my corporation to figure out how to do it. But indeed, it's a very, very difficult task. We can do that today and I'm happy to say there are people out there that they have gotten the message about Text and they are starting to look at it now Confusing matters immensely is this chat GPT stuff People think oh chat GPT. That's text. We now have a handle on text No, you don't and let me tell you why And by the way, I'm a fan of chat GPT. I have nothing against chat GPT, but chat GPT answers a different question. Chat GPT is good for taking language and text and turning language and text into a question, into a query, into your computer. That's what chat GPT does, and it does it well. What ChatGPT doesn't do is look into the language and take the value and data that's in the language that's there. And I know this for a fact because we've looked at it and played around with it. And so, and it sounds the same. It sounds, oh, ChatGPT text, that solves our problems. No, that doesn't solve your problem because ChatGPT does not go into the text and find what's in the text. Then you say, okay, Bill, thanks for telling me about textual ETL. I've never heard of it before, but it's time you hear about it. Textual ETL, the heart of textual ETL is something called an ontology or a taxonomy. and ontologies and taxonomies are how, that's the magic ingredient of how you go into text and start to understand the text to the point where you can turn it into a database. And you probably don't want to get me off on this subject because I've been working the last, oh. 13 years of my life on it. And I'm gonna tell you, it's a very complex subject. Let me give you a couple of examples of why text is so complex. Let's take the word fire. What does the word fire mean? Well, it could mean that your house is burning and you're on fire. That's one meaning of it. It can also mean that your boss doesn't like you anymore and you got fired this morning. Or it could mean that you've got a gun in your hand and you pull the trigger and you fire the gun. So there's lots of meaning. Our language is full of double meanings, triple meanings for all kinds of things. And in order to do a proper analysis on text, you've got to be able to distinguish between what's being said. And that is no, I've been doing this for 13 years now. And I can tell you there's really two components here, text and context. Text is actually fairly easy. It's not, I don't know, it's fairly easy. What's the devil is context. Context is not easy at all. That's where the problems come in. That's where the difficulties come in. So anyway, how do you start to go from text to a database to where the person like yourself can start to do their analytical magic is there is technology out there called Textual ETL that indeed... does what you are asking it to do. And by the way, when we started, okay, when we started on Textual ETL, we looked at something called NLP, Natural Language Processing. And NLP is an academic exercise. It was never designed to be a commercial product. It is complex. It takes a tremendous amount of time. and it requires high price consultants to make it work. When we started off to build Textual ETL, we wanted all of those things to not be. So we've created something that is inexpensive to use, that is not complex, doesn't require an army of technicians and is fast. And so we have a commercialization. really and truly what Textile ETL, you can think of it this way, is as a commercialization of NLP. But the people in NLP, they want to cling to their life raft. They're not about to hear it. But let me tell you, business people. So when we first started, we thought, well, we'll talk to people in the world of NLP. And they don't want to hear from us. I'll tell you who does want to hear from us is people in marketing, people in management, people in sales, people in finance, and they love what they hear.
**Robert Harmon:** It seems to be a recurring theme through this conversation, Bill, because, you know, at the end of the day, it gets all the way back to the beginning, provide value. And if the people out in the, you know, in the business want to hear from you, you're obviously providing value. And that's sadly not always the case with IT teams. The other thing I notice here is if this does catch on, that's going to be a lot of data. So Benjamin and I may have a little. to do what you get going.
**Bill Inmon:** Yeah, I'll be retired or dead by then, one of the two.
**Robert Harmon:** Hahaha
**Bill Inmon:** but you're
**Robert Harmon:** So,
**Bill Inmon:** absolutely right.
**Robert Harmon:** yeah, so swimming
**Benjamin:** Thank you.
**Robert Harmon:** and more data. I'm not sure I appreciate that, but at least I won't be unemployed for a while. So really, Bill, this did not go the direction I expected it to, and that's a really good thing because sometimes I'm a very boring person. I do have some other more personal questions.
**Bill Inmon:** Sure.
**Robert Harmon:** You're a car guy. You're a car
**Bill Inmon:** Yeah.
**Robert Harmon:** guy, yes. How did that happen?
**Bill Inmon:** You don't know it, but you hit a really sore spot. I am not a car guy.
**Robert Harmon:** Oh
**Bill Inmon:** In February, my
**Robert Harmon:** Uh-huh.
**Bill Inmon:** car had its catalytic converter stolen
**Robert Harmon:** Oh no!
**Bill Inmon:** and my car has been in the shop since February. I've been using my wife's car for half a year now. I don't know if you've tried to get a catalytic converter, but it's like gold or platinum or diamonds. Give me a break. And so I don't know when I'm going to get my car back.
**Robert Harmon:** What?
**Bill Inmon:** But I've owned in my lifetime six Porsches, one Ferrari. I can't resist this. The two best days in the life of a person owning a Ferrari. is the day they own the Ferrari and the day they get rid of the Ferrari. And, oh, you, if somebody came to my front door, parked a Ferrari in my, in front of my house, gave me the keys, I would run as fast as I could. I, I'm not about to take it. Now, now having stated that, Porsche, pardon me. Porsche is as well made as Ferrari is a piece of crap. And I love my Porsches and we need to have my company progress a little bit further. But when my company progresses a little bit further, I'm gonna be getting my seventh Porsche.
**Benjamin:** As a German, I appreciate your support of the German automobile industry.
**Bill Inmon:** Oh.
**Robert Harmon:** Honestly, I think it's possible Bill and I paid for the entire German automotive industry over the years.
**Bill Inmon:** You know what car I loved and people didn't think much of it, but I had a Porsche 914 and that was the rear engine and the thing that I loved about it is the way it handles. It handles differently and better than any other car I've had.
**Robert Harmon:** Exactly.
**Bill Inmon:** But when you tell a Porsche fanatic that you like the 914, they automatically think of you as a real wimp. But I love they're my 914.
**Robert Harmon:** wonderful cars.
**Bill Inmon:** It was a wonderful car.
**Robert Harmon:** yeah, I've owned a couple of them and they're absolutely wonderful cars. They're not gonna tear anything up on the, compared to modern cars, but it's a delightful experience. Everyone should do it once.
**Bill Inmon:** You don't know what driving is like until you've driven a 914.
**Robert Harmon:** Awful experience. Well, there are some things to it. The transmission, the shift forks, I swear were designed for a tractor.
**Bill Inmon:** Yep.
**Robert Harmon:** So you do have to get your, you know, you got to know what you're doing to a little bit, to a solid extent. And then they're great little cars once you get them figured out, but they're a little quirky.
**Bill Inmon:** Yep.
**Robert Harmon:** Well, I don't know as I had a whole lot more at this point. Benjamin, you've been very quiet.
**Benjamin:** I've been asking tons of questions.
**Robert Harmon:** Okay. Well, maybe it's time to start winding this down.
**Benjamin:** Sure. So Bill, it was an absolute pleasure having you on the show, getting to know you. Thanks for everything. Thanks again for all of the kind of, yeah, making the data industry what it is today. It was an absolute pleasure having you on. Yeah. I hope you have a great rest of your day.
**Bill Inmon:** Thank you so much, Benjamin. Robert, nice talking with you. We'll talk again.
**Robert Harmon:** All right, we'll talk to you later, Bill.
# Block Bad Data Before the Write with Nike’s Ashok Singamaneni (/blog/block-bad-data-before-the-write-with-nikes-ashok-singamaneni)
Nike's Principal Data Engineer Ashok Singamaneni joins Benjamin and Eldad to discuss his open-source data quality framework, Spark Expectations. Ashok explains how the tool, which was inspired by Databricks DLT Expectations, shifts data quality checks to before the data is written to a final table. This proactive approach uses row-level, aggregation-level, and query data quality checks to fail jobs, drop bad records, or alert teams - ultimately saving huge costs on recompute and engineering effort in mission-critical data pipelines.
Listen on [Spotify](https://tinyurl.com/56jz6tda) or [Apple Podcasts](https://tinyurl.com/bdheke9h)
\[00:00:00] Ashok: DLP expectations gave an idea to the industry that you can do data quality before actually writing the data into your final tables. As the scale of the product increases, it becomes even more difficult for us to find exactly where the issue went wrong, like, even if a production job fails.
\[00:00:18] Benjamin: Ashok is a principal data engineer at Nike, and he worked on a variety of cool, actually, open source projects as part of his work there.
\[00:00:27] Ashok: I think over the time, in my experience, what I learned is this ingestion layer and the transformation layer, you should treat that as a software product, not like a data engineering product.
\[00:00:39] Benjamin: Hi. This is Benjamin. Before we start with today's episode, I wanted to quickly reach out on a personal note. We've just launched Firebolt Core. Firebolt Core is the free self-hosted version of our query engine. You can run Core anywhere you want, from your laptop to your on prem data center to public cloud environments. Core scales out, and you can run it in the multi node configuration. And best of all, it's free forever and has no usage limits. So you can run as many queries as you run and process as much data as you want. Core is great for running either big data ELT jobs on, for example, iceberg tables or powering high concurrency customer facing analytics on big datasets. We'd love for you to give it a spin and send us feedback. You can either join our Discord, enter our GitHub discussions, or you can just shoot me an email at [Benjamin@Firebolt.io](mailto:Benjamin@Firebolt.io). We'd love to hear from you. We added a link to Firebolt course GitHub repository to the show notes. And with that, let's jump straight into today's episode. Hi, everyone. And back on the data engineering show.
Today, Eldad and I are actually in person in Munich together, which is very nice. He's real, same room. We're both real. Exactly. And we're super happy to have Ashok on. Ashok is a principal data engineer at Nike, and he worked on a variety of cool, actually, open source projects as part of his work there. So one is called BrickFlow, and one is called Spark Expectations. Excited to have you on the show today, Ashok. Do you wanna quickly introduce yourself? Tell us about your background, how you got into data.
\[00:02:04] Ashok: Thank you. I am Ashok Singamaneni, and this might be interesting. I'm a mechanical engineering. I've done mechanical engineering basically. And then I switched to data like probably twelve years ago and it's been a long journey and I've done my master's.
\[00:02:21] Eldad: The industrial revolution ends with data, unfortunately.
\[00:02:24] Ashok: Yeah. So I figured probably when I read the MAPRED newspaper, that's when it clicked that I think this is going to be good. And then Cloudera, Hortonworks, and all of that happened, and I write some books. And then I did my masters and did thesis on, big data analytics on vehicle data, and then later moved into banking industry and then worked in health care and then now in retail domain. So it's been interesting journey.
\[00:02:52] Benjamin: And so after twelve years of data, do you ever think back? Are you like, hey. I next kind of, like, in a couple of years, I wanna become a mechanical engineer again, or you're, like, kind of
\[00:03:01] Ashok: No. I think it's a perfect blend right now. Look at Tesla. You can be a mechanical engineer as well as also be in software. And, also, you can be in robotics as well as in software and data.
\[00:03:11] Benjamin: That's awesome. Cool. So at Nike, right, you worked on both these kind of BrickFlow and Spark Expectations framework. Tell us a bit, like, how did those projects start? Kind of why did you decide to actually open source them? What are they all about? We'd love to learn more.
\[00:03:27] Ashok: Sure. I think coming to Spark Expectations, right, being in the industry for so long and have been in lot of production calls and misfires happening in the production because of the data changes, unexpected column changes, etcetera, in upstream or downstream, and having been in lot of recompute of the data. This all happens because of the majority of the times data quality issues, which can be removed upfront while the data is being processed, as well as also some of them are the processing or or planning mistakes that happens regularly, and that's kind of common. But when I looked at Databricks DLT pipelines that was released a while ago, they have introduced a DLT expectations, which kind of is interesting because if you have seen Great Expectations, Great Expectations is a tool that actually gives you data quality report post processing of the data. After you process the data, you will see the data quality report that you can run on the data and then the report is really fantastic. But DLT Expectations gave an idea to the industry that you can do data quality before actually writing the data into your final tables. And then I reached out to Databricks team like if we can work on, something related to Spark as well. Because many, many companies, we know that using Spark for data processing and transformations. If this can be available in Spark, that would be great. But the timelines didn't work out, so I started working on the project called Spark Expectations in the same name of weird expectations, weird expectations as this project is related to Spark. So I named it Spark Expectations. And, it has more functionality than DLD right now, but it is only supported for Spark. I wish that it supports for all the frameworks like Pandas and other libraries as well, but right now, we are supporting Spark. And what this does is before you write the data into your final layers, the data quality checks happen and restricts the data that is not up to the data quality standards what you want so that you'll need to recompute as well as you will have alerts on the data that has gone bad, and then you'll be able to alert your upstream teams that this data is bad. This need to be resend and reprocessed.
\[00:05:44] Benjamin: How do you define these expectations? Is it like a YAML file, a JSON file? Like, take us through that part.
\[00:05:50] Ashok: I think it's up to the teams which are implementing because right now, the rules, the way Spark expectations expertise is part of a data frame. So you can write it as a YAML file, JSON file, or might as well have that in a table as well and load that as a data frame and give it as an input to Spark expectations. And there are also, like, three kinds of rules that you can provide, like, row level data quality checks, table level aggregation level data quality checks, as well as query deque, which we call it as, like, for referential integrity checks or if you want to have, like, some complex rules that you can run on data quality.
\[00:06:27] Benjamin: Okay. And then kind of that runs basically before I do any of my batch ELT jobs? Does an initial pass over to the data, makes through, yep, everything kind of looks great before I start spotting a lot of
\[00:06:39] Eldad: like, is it sampling the data? Is it running the previous snapshot?
\[00:06:44] Ashok: So there are, like, different layers in Spark expectations. Right? So when you load the data let's say, for example, if I'm ingesting the data and transforming it, when I load the data into the data frame, like the initial data frame when I load it from my source table, you can run source level, query dequeue, or aggregation level checks that you want to do. And then you can run data quality or row dequeue checks as in, like, on each column. Let's say the date has to be greater than this level or any other validations that you want to do on a particular column, you'll be able to do. Not a sampling. It runs on the whole dataset. It gathers all the information. And you can also have conditions like ignore, drop, and fail the job. If a rule fails, you can ignore it, but still alert it that this rule is a soft failure. It failed, but you are getting still an alert. Or you can drop the record. Like, let's say a product table doesn't have a product ID or the product ID configuration is wrong, then that is something serious that should not be there in the data. So you drop that record on the whole and put that in an error table and give that alert to the engineering team that there is some error in the error table you can look at. And you can even fail the job. If it's mission critical, then you fail the job, not process the data, and don't put that data into the final table so that you don't need to recompute that again.
\[00:08:04] Benjamin: But then for me to understand this a bit better, like, does this actually run as a hook within the ELT job itself? So that, basically, every time I start scanning a table anyways for my big batch processing, I also run the data quality check. It's not like I have to scan the table twice now. Right? Like, that I will first pass over all the data, which is really expensive, run the kind of spark expectations, throw everything away, and only when that looks good, I run the batch ELT. It basically overlaps. Is that accurate?
\[00:08:35] Ashok: Yes. I think we use decorator pattern in Python for this. So whenever you tag that spark expectations decorator, there is something called with expectations. If you tag that inside your function and your function returns a data frame, on that data frame, whatever rules you define, it runs all of them and generates the report.
\[00:08:54] Benjamin: Very cool. What have you seen in terms of the, like, overhead this actually introduces for, like, bigger batch ELT jobs? Is it noticeable? Is it very fast? Take us through that.
\[00:09:04] Ashok: Yeah. I think the road e q checks that happens are very fast. It should happen as a pretty standard checks that happens on the scale. But, obviously, definitely, there will be an overhead. So you wouldn't want to put this on all the layers and all the jobs. You only want to put it at the final layer or the final right step where you are actually writing the data, and that is, like, the mission critical production data that you want to have it. And for the query DQ, as well as aggregation DQ, you need to be careful and optimize it well so that the time is not heavy. Like, you need to optimize your queries, obviously, like any Spark job, enter it. But that is an overhead for sure. Like, if you're running data quality checks every day on top of, like, huge scale of data, then obviously you'll have. But it's also possible in streaming too. Like, if you're doing, like, a micro batch, then you would be able to do that on micro batch too.
\[00:09:55] Benjamin: Okay. And so let's say I have, like, I don't know, medallion architecture. I go, like, brass layer, silver layer, gold layer, which you're basically saying, like, on the right to the gold layer, I would basically hook in Spark expectations. And at that point, which then connects with the framework and also error out, I could really make sure that, like, no insert into the gold layer even finishes, um, if there's not a certain kind of data quality check passing.
\[00:10:24] Ashok: Yes. Put it. Exactly.
\[00:10:26] Benjamin: Very nice. Now would it
\[00:10:27] Eldad: be fair to kind of compare that to constraints in SQL or kind of that domain assertions constraints?
\[00:10:36] Ashok: Yeah. I think it's something similar to constraints. Right? But it's also the constraints has a limitation that you can only put constraints on that particular column values. Like, this is what this column expects, but spark expectations does something more than that. Like, you can have, like, aggregation DQR. Like, let's say, for example, the whole count or the sale in that table should be greater than 100,000. If, let's say, there is a big sale that is going on and if the sale is, like, less than 10,000, then that doesn't make sense. There is something wrong in the computation, etcetera. Or differential integrity checks, like, if you have, like, multiple tables in the data model and you want to have, like, correlations that are happening properly or not, if the business logic makes sense or not for this data to be there. So those kind of checks also can be done.
\[00:11:25] Eldad: So if someone is you know, for people who own production grade pipelines, specifically kind of at those stages where the data is not fully cleansed. Right? Like, at the end of the pipeline, everything went through cleansing. It's always safe. It's it's there's a value for every product. But as you go closer to where data is being born, it's nastier. Right? It's not clean. There's no formal way to do it. So every data engineer I know of has a set of toolboxes based on intuition, experience. Obviously, depends if they're engineers and they're writing spark or getting deep into the right into how those queries run. What would you recommend us? Like, when should we seriously start looking into those frameworks and kind of what would be your guideline on transitioning to that mindset?
\[00:12:17] Ashok: Yeah. I think over the time, in my experience, what I learned is this ingestion layer and the transformation layer, you should treat that as a software product, not like a data engineering product like from per se. You will have all the data modeling, data governance, data observability, all of those tools that are there. But I think the general mindset should be the ingestion layer or the raw and the bronze layer and the silver layer should be like a software product. You should have all the checks and balances in place, like data quality and unit testing, integration testing. All of those should be in place so that any mishaps that happens, it's easier to find out and debug. Because as the scale of the product increases, it becomes even more difficult for us to find exactly where the issue went wrong. Like, even if a production job fails, it takes time for you to debug and see, like, lot of human effort also involved, not only the recompute that is happening and the compute costs that are happening, but also couple of engineers, a product person has to be involved for a week or week and half to fix that issue that is happening. But the final goal layer where mostly people use SQL, that can be treated as like a pure data engineering. Like, write SQL, do it fast, build dashboards, break them, and then fix them as much as you can. But like the initial two layers are mission critical that has to be treated as like a software product. That's at least my experience, what I have seen.
\[00:13:49] Benjamin: Very nice. Take us through your experience, like open sourcing a data project, right? It's kind of like, did you go to some meetups? Kind of like, do you actually know some of your users? We'd love to learn more about that.
\[00:14:01] Ashok: Yeah. I think it's been a fantastic journey for open source. It's my first project to to do that, but I have really great mentors. A big shout out to Adi Aditya Chaddhivedi, who's distinguished at Nike and also Scott Haines who was an author at O'Reilly and other people who have helped me through the journey. My senior director, Joe Hollow, who helped me through the process. These guys have helped me through the process of, like, open source. And, also, there is an open source community effort that is happening at Nike as well. So getting through the approvals and getting the org set up and and the project set up and publishing it, it has been great. And I also had an opportunity to talk at Databricks AI Summit. Um, so networking happened. Couple of engineers also looked at the project and held through. And I think there is more that can be done from the project side, and I wish there need not be Spark expectations. There should be something native to Spark so that anyone can actually use directly from the native Spark module.
\[00:15:06] Benjamin: Very cool. So how are you now using like Cursor or, like, any of these Gen AI tools to actually accelerate kind of progress within your own open source project? Like, are you vivecoding a good looking UI?
\[00:15:20] Ashok: I think I've been using Cursor and Cloud Code since, like, over one and a half years' time. I mean, time flies. It feels like eternity that I've been using them for a long time now. But wipe coding is really good if you know exactly what you're doing, at least from the data engineering standpoint. Like, for the regular software engineering, building websites, etcetera, that can be very quick. But from the data engineering standpoint, we have to take some guidelines into place. I've seen people using plot code or cursor on production data directly when they're building and drop the datasets. So, so misfires happen a lot if you are using white coding directly on data engineering projects, guidelines that need to be taken. Use service principles or AWS rules, etcetera, so that they have restrictive permissions when you are using plot code or cursor. That is something that I have learned myself as well. I was working on a small pet project, and I was using SQL mesh. And there were some commands which I didn't know that it destroys the whole dataset. And it's a sandbox handler, but it destroyed everything. Right? So having those guidelines in place is really critical from the permission standpoint. And as well as from wipe coding, I think there are different patterns people are using these days on, like, creating a PRD document first, have a GitHub issue, and set up a process for yourself so that there is a journey from what is the problem that you're trying to solve, creating a proper GitHub issue for all the issues that you are working on, and then helping through the journey of, like, breaking that data issue into small tasks as checklists in your plot code or cursor, and then checking through each one of them, writing unit tests as you go, and having integration tests, and human in the loop that you have to review the code that has been written. Ultimately, at the end of the day, you are responsible when you're checking in the code. It's not Claude or Karsar. That will be blamed if something goes wrong.
\[00:17:26] Eldad: Listening to your think, it actually makes perfect sense to have those cleansing and expectation libraries being used for often nowadays when no one actually knows who's behind each and every change. So as it starts, right, forking, modifying, changing, updating production eventually, setting expectations makes more sense. So I think there's a bright future for kind of right away even say, like, focus on the expectations because we can't control the modifications themselves. So at least we can own the expectations. But this will be super interesting to see how this kind of semi automated yet highly expectation pipeline will turn out.
\[00:18:08] Ashok: Yeah. And it's also the trend that I'm seeing these days, is from the data governance teams and data governance standpoint as well across the industry. Data observability and quality is becoming prime because of AI integrations that are happening. There are tools that are coming out where a CEO or CTO from a company can directly ask questions in natural language and hit the production data and get the data for themselves rather than Someone goes,
\[00:18:37] Eldad: codes it, writes it, generates the reports.
\[00:18:40] Ashok: Yeah. So the leadership is directly looking at the data and if there is something wrong in the data, then there can be some serious repercussions happening on the business decisions. So that's one of the reasons why I see the industry moving towards the trend that rather than having bad data in the tables and then recomputing or reclarifying things, let's not put that data first in the first place and then redo whatever needs to be done from the upstream and fix the data.
\[00:19:10] Benjamin: Cool. So if you look ahead, like, year to year, what are you most excited about in the data space?
\[00:19:17] Ashok: I think from the data standpoint, I'm looking at how the white coding improves and the tools that comes for white coding in terms of the data engineering space exclusively. Because right now, it's all generic patterns that are evolving in the industry. Like, if you look at Reddit or any other blogs, there are certain patterns that people are following for UI coding, iOS, Android apps, but there is not specific that I have seen for the data engineering space. This is how you exactly do white coding. I think I'm more excited towards that on, like, how that can be evolved as a tool or a framework. If there can be something as a spec, it, like GitHub has really suspected recently on, like, how to do wipe coding. But if there is something like that for data engineering, that could be awesome. I would like to collab on that. Nice. Awesome.
\[00:20:10] Benjamin: So to all our listeners, if you want to work with Ashok on Vibe coding for data quality and data engineering, reach out to him. It was so great having you on the show. Seriously, kind of thanks for sharing your experiences. It's always exciting to learn about new open source projects in the data space. So thank you for being on.
\[00:20:29] Ashok: Thank you. Thanks for having me. Have a good day.
\[00:20:32] Unknown: The data engineering show is brought to you by Firebolt, the cloud data warehouse for AI apps and low latency analytics. Get your free credits and start your trial at firebolt.io.
# Building Customer Trust: A CISO's Perspective on Security and Privacy at Firebolt (/blog/building-customer-trust-a-cisos-perspective-on-security-and-privacy-at-firebolt)
### Security as a Core Value [#security-as-a-core-value]
At Firebolt, security isn't just a feature of our service. From our co-founders' initial vision to our cutting-edge technology, through tuning business processes to fortifying our cyber defense, every aspect of Firebolt is designed with security in mind, influencing every decision and innovation we make. Firebolt's holistic approach to security ensures that every aspect of our operations is geared towards maintaining the highest standards of protection for our users and their data while underscoring the critical importance of our security practices and the continuous enhancement of our security posture.
With that, our commitment to robust security measures from day one ensures that every customer's data is protected against evolving threats, allowing customers to trust Firebolt as a secure platform for their most sensitive data. We are proud to have achieved key industry certifications, including ISO 27001, ISO 27018, and SOC 2 + HIPAA, which validate our adherence to stringent security and privacy standards.
### Integrating Security Throughout the SDLC [#integrating-security-throughout-the-sdlc]
As a provider of Cloud Data Warehouse as a Service, Firebolt faces unique challenges in securing vast amounts of data, and at the heart of our Security DevOps approach is the seamless integration of security throughout the software development lifecycle. From the outset, secure coding practices are employed, and as development progresses, each phase undergoes various security testing. This proactive scrutiny helps identify and mitigate vulnerabilities early on, ensuring robust defenses from the ground up. We also actively engage in essential practices like penetration testing and [fuzzing](https://www.firebolt.io/blog/fuzzing-firebolt-catching-0-days-as-fast-as-our-query-processor), allowing us to uncover potential weaknesses and fortify our systems against emerging threats.
### Data Protection and Network Security [#data-protection-and-network-security]
Data protection is paramount, and we employ stringent policies coupled with end-to-end encryption to safeguard data at rest and in transit, maintaining confidentiality.
* Data at Rest
* We leverage AWS's built-in features, like KMS, to handle storage encryption at rest for customer data stored within our systems. Our systems are securely integrated with AWS KMS to ensure seamless key management for data encryption through AWS's infrastructure.
* Encryption keys managed by us securely generate, store, rotate, and retire the encryption keys that we manage, particularly those used for encrypting sensitive internal data
* We enforce strict access controls and separation of duties for individuals authorized to manage the encryption keys under our control.
* We safeguard the storage of encryption keys to prevent unauthorized access.
* We enforce robust authentication mechanisms for users and applications accessing the data. This means that every client, whether it's a UI, SDK, JDBC, or another interface, must be associated with a unique login and user. This approach ensures that each client is individually authenticated and authorized to access the data.
* Initial authentication is handled by our Identity Provider, which validates the user's or application's identity.
* Once authenticated, our fine-grained RBAC mechanism further controls access to specific data.
* We've implemented auditing and monitoring mechanisms to track and log access to encrypted data and encryption key management activities.
* Regularly review and analyze the audit logs for any suspicious or unauthorized activities.
* We ensure the secure transmission of encryption keys from the key management system to the encryption/decryption components to prevent interception or tampering.
* We've implemented secure data deletion practices to ensure that data is adequately wiped or destroyed when no longer needed.
* We've implemented data integrity checks to detect unauthorized modifications or tampering with encrypted data.
* Data in Motion Firebolt
* We've implemented TLS to ensure secure communication across all client-server interactions. This applies to external clients, such as a UI accessing a gateway, and internal communications between components within our infrastructure, providing an extra layer of security against man-in-the-middle attacks.
* Our policy exclusively includes strong cipher suites that offer robust encryption and resist known cryptographic attacks. These suites utilize AES (Advanced Encryption Standard) with 128-bit or 256-bit keys and Galois/Counter Mode (GCM) for authenticated encryption, ensuring data confidentiality and integrity. Additionally, the selected cipher suites support Perfect Forward Secrecy (PFS), meaning that even if a server's private key is compromised, past communications remain secure, as session keys are not derived from the server's private key. We intentionally exclude weaker ciphers and deprecated features like RC4, MD5, and SHA-1, which are vulnerable to cryptographic attacks. By default, we negotiate TLS 1.3, with the ability to fall back to TLS 1.2 if necessary.
Our network architecture is carefully designed, with segmentation for production, development, and testing environments. Each segment is fortified with tailored security controls and subjected to regular security audits. In addition a comprehensive multi-layered security strategy that includes Web Application Firewalls (WAF), Denial of Service (DoS) protection, and bot mitigation is implemented.
### Privacy by Design [#privacy-by-design]
Privacy is another core principle in our system design and is embedded to minimize data exposure and proactively mitigate risks with the addition of comprehensive access controls and robust governance frameworks. We protect user and organizational data. Each tenant's data is securely isolated, preventing unauthorized access or data leakage. Our approach ensures that sensitive information is protected through encryption, access controls, and privacy-enhancing technologies.
1. We follow a data minimization strategy, collecting only the essential information for service functionality. Where possible, we apply anonymization and pseudonymization techniques to reduce the exposure of sensitive data during the processing, analytics, and testing phases.
2. We leverage AWS's encryption services, such as AWS Key Management Service (KMS), Secrets Manager, and Vaults, to securely handle encryption keys and sensitive information. Data is encrypted at rest and in transit using industry-standard algorithms like AES-256. Additionally, secrets such as API keys and credentials are securely stored and rotated.
3. Tenant data isolation is enforced using Kubernetes namespaces within our Amazon EKS environment. Each tenant's resources are segregated into dedicated namespaces, ensuring strict boundaries and preventing cross-tenant access. These namespaces are further secured with role-based access controls and network policies to maintain isolation.
4. We implement holistic runtime security using a defense-in-depth approach. All of our binaries are built with Position Independent Executable (PIE) to randomize the address space for code execution, stack protection, and Control Flow Integrity (CFI) to prevent attacks that target both forward as well as backward edge flows by overwriting return addresses on the stack. Next, these hardened binaries are run inside non-root containers with strict namespace isolations and limited capabilities, thus adhering to the Principle of Least Privilege (PoLP). This guarantees that even in the improbable event of an exploit that defeats binary hardening, an attacker still cannot break out of the container and escalate to root. Furthermore, we deploy an advanced runtime protection tool that observes every process for suspicious activities and malware that alerts and blocks anomalous events. Ensuring that our systems are resilient against threats while safeguarding sensitive data.
5. Access to sensitive data is governed by fine-grained RBAC policies, enforcing the principle of least privilege. These controls extend across all system components, ensuring only authorized users and services, including internal employees, can access specific datasets.
6. Logs are explicitly configured for internal use, forensic needs, and compliance, excluding personal identifiers, authentication tokens, and other sensitive attributes. Where logging is essential, sensitive information is hashed, tokenized, or fully redacted. Additionally, we enforce strict retention policies and secure storage for logs containing minimal necessary data.
7. Our platform is designed to support data subject rights, such as data access, rectification, and deletion requests.
### Change Management and Governance [#change-management-and-governance]
Our change management process is designed to ensure the secure delivery of services. Before deployment, changes undergo thorough risk assessment and testing, minimizing disruption and potential security risks. Governance is central to our security strategy, with policies and procedures in place to ensure compliance with industry standards and regulations. Regular audits and assessments verify adherence to these standards, providing transparency and accountability to our customers.
### Securing Customers Trust [#securing-customers-trust]
We believe in a collaborative approach to security and privacy. By aligning our best practices with customer-managed controls like **Single Sign-On (SSO), Multi-Factor Authentication (MFA), Network Policies, and Role-Based Access Control (RBAC)**, we create a synergistic approach to security. This collaborative effort ensures a layered defense that enhances overall robustness and resilience against evolving threats, demonstrating our commitment to our customers' trust and peace of mind. More information is on our [security documentation](https://docs.firebolt.io/Overview/security.html) page.
# Building Data Products For Data Engineers (/blog/building-data-products-for-data-engineers)
How does a tech stack that always needs to be at the forefront of technology look like? [Roy Miara](https://www.linkedin.com/in/roy-miara-73776a56/) from [Explorium](https://www.linkedin.com/in/roy-miara-73776a56/) talks about building data products for the audience that can't be fooled – data engineers.
Listen on [Apple Podcasts](https://podcasts.apple.com/us/podcast/building-data-products-for-data-engineers/id1561927688?i=1000534802627) or [Spotify](https://open.spotify.com/episode/7dAo7XjwjZ44xz68T0J6Ti)
**Boaz:** Hello, everybody. Welcome to another episode of the data engineering show. So happy to see you with us again.
**Eldad:** Yes, I came back from abroad.
**Boaz:** I missed you.
**Eldad:** I missed you so much.
**Boaz:** So, welcome to another episode where we host data practitioners from all around. Today with us Roy Mira. Did I pronounce it correctly, Roy
**Roy:** Yep.
**Boaz:** Great. So, Roy an engineering manager at a super interesting company called Explorium, engineering manager of data and ML who's been in a variety of data engineering and machine learning focused roles, a variety of startups in his career. And typically, we talked to big brand names before, but it's important to also talk to sometimes smaller, interesting companies...
**Eldad:** Emerging brands.
**Boaz:** Like Explorium so companies that actually build stuff for data engineers and data scientists, which makes Explorium interesting. So, Explorium, if you haven't heard about them recently landed 75 million in investment. What they do, they help with data enrichment and help you sort of use external and public data sources to simplify your training data procedures and stuff like that. We'll have Roy expand on that a little bit but before that, actually Roy ran late a little bit. Roy, why did you run late to this webinar? I think you have a story for us.
**Roy:** Changing a flat tire.
**Boaz:** For whom?
**Roy:** For a pregnant woman.
**Boaz:** This is so nice. Sometimes we have people who are also amazing humans outside of the data engineering show. Thank you for helping...
**Eldad:** And your wife is not expecting.
**Roy:** No, it's not my wife. I had a feeling because I was texting Boaz and then I said, oh, it's going to pop. So, I text him I'm going to be late.
**Boaz:** So, take an example from Roy. If you see pregnant people who need to change flat tires in the heat help them out. Okay, so Roy, tell us a bit about Explorium in your own words, please.
**Roy:** Among data scientists today, we hear a lot that it's all about the data. It's no more about the models, this is the era of data centric AI in general. And big data analytics has always been around data, who has the best data, who is the most accurate, the most relevant data to answer some business questions and in Explorium we're looking at this whole world of machine learning and big data analytics and we say it's all about having the right data. So, what Explorium does, it enables organizations; large organizations, small organizations, to have immediate access to the most relevant data for their business problems. And we do it by obviously aggregating a lot of data sources in a lot of fields, modeling the data correctly and enabling this kind of search engine over our dataset with accordance to what the user has and already has as an internal data from his company. But also, in cases where the user doesn't even have data and only has some business question, he wants to know all the companies that have more than 10 employees and more than $5 million in capital in the East Coast, for example.
**Boaz:** So, who typically are the end users? Is it more for data engineers, data scientists? Is it sometimes for business users as well?
**Roy:** Yeah, our users span from data scientists and deep ML engineers on the far hand and on the other hand, you have people from the business and analytics side that use the platform to do all their exploratory data analysis. So, we see everything in between, obviously.
**Boaz:** Okay. So, we're talking about a company who manages a lot of data, makes it easily accessible, both for engineering and for a business. Super interesting. So, tell us a little bit just so we can understand the data challenges you guys have. What data volumes more or less are you guys dealing with?
**Roy:** Data volume here is tricky because we're processing somewhere around the two terabytes a day, but when I'm saying processing, I'm talking about already structured tabular data. So, we're not looking into raw data, we're looking at fine-grained, high quality data from a variety of sources. And I think one of our major challenges is having this variety of sources and different schemas constantly evolving. So, around a couple of terabytes a day this is what we process and in volume, we're talking also a couple of hundreds of terabytes constantly updating.
**Eldad:** And is that one data source? Is that one global schema with many tables in it?
**Roy:** No, there's hundreds of different sources, most of them are structured but we're talking hundreds of data sources, thousands of features or thousands of different schemas because from every source we generate and aggregate data to many, many different use cases.
**Boaz:** Because Explorium is company where the data is the product, essentially. How does the split between engineering, data teams, data engineering sort of look like? So maybe give us an overview of how many people are in engineering, how many people deal with data, how many people have a data related title versus an engineering related title. Do people without a data title still work on data stuff?
**Roy:** That's a good question, actually, because I claimed that everybody here are data engineers, scientists, and analysts, kind of a mix. We have an infrastructure organization, data infrastructure, which is the team that I'm leading. We're about nine data engineers, a data ops engineer, ML engineers so we kind of work in between. We have another team that's building Explorium's feature store to say.
**Eldad:** So, a user can just add a one of many data sources that are completely structured into their schema, into their model, enrich it and query it, basically.
**Roy:** We have these two flows. One flow, we call the auto ML flow, the ML engine. So, in this flow, a user is coming with his data set, internal data that he collects internally inside his organization and he has some target that he wishes to predict like the classical ML flow. What happens at that point; he can connect to our platform and basically the platform will enrich automatically. So, add sources automatically, according to the context and the analysis of the original core data set that the user uploaded. The platform will automatically analyze the data, understand which data sets are the most relevant, connect them, train the model, evaluate the model based on how much gain do we see from those external features. So, it's connecting the data, it's running feature extraction, it's running feature selection, closes the loop with a model and then finds the best features according to the target.
This is what we call the auto ML, ML engine flow. And we have a flow that is more towards general like, you know, analysis. Meaning that you can upload your data, we analyze the data, but we will present our internal catalog of features. So, you'd be able to see the sources, you'll be able to see the coverage, you will be able to see the feature is with all the description and everything, and the user can add features. So, add enrichments, he can run transformations and basically build what we call the recipe. He could build the recipe of the features that he wants to add with all the transformation and then he can query this recipe in production, he can use this recipe to schedule batch jobs that will update this data. And now we're, we're kind of connecting all of these flows together, create one unified flow where a model is just a part of the recipe.
**Boaz:** And what does your data stack look like?
**Roy:** Our data stack is quite varied. We have a couple of internally built tools, our feature stories internally built but what my team is managing in terms of looking at the entire data stack, we're looking at databases, Postgress, DynamoDB, we have Elastic Search. Those are the main databases that we have. We have a cloud data warehouse, we are using Firebolt for a lot of our internal processing and also for exposing some of the data directly to the users. We're using Spark a lot, we're using DBT, we're trying to be very on the edge of the technology because we are facing new use cases all the time. We kind of have this challenge of trying to be good with any data so we going have to be familiar with any type of warehouse, any type of lake architecture you have to like keep track with what's happening. This is a big part of what we do is try to understand where is the data engineering world is going, because this is where also our users and customers are going.
**Eldad:** By the way, did you try the new latest DBT, Spark integration, any feedback on that?
**Roy:** We have tried it. We actually trying it as we speak so this is something we're building.
**Eldad:** Nice
**Roy:** Yeah. Working nicely. We also ran it based on Presto, which worked nicely and I'm hearing integrations are coming also from your end so we will be trying that as well. As feedback, I love this tool, I think this is the way to go. For us it's been very fitting because of the variety and the constant need to create more pipelines and other complexities and orchestrate everything, and to have it in a way that we kind of democratize it among data scientists and analysts both internally and externally in a way. So, you kind of have to have a unified layer where engineers can sleep in their beds quietly at night without having to worry, waking up on pipelines breaking, over schema changing stuff like that.
**Boaz:** Let's do a quick switch to a fun and blitz round, in which we will ask you a variety of questions where you're not supposed to think too much, just answer quickly. There are no wrong answers, only yes and no. There are no wrong answers only except the ones that are wrong.
**Roy:** Okay.
**Boaz:** Okay. So, are you ready?
**Roy:** Yes.
**Boaz:** Commercial or open source?
**Roy:** Open source and commercial.
**Boaz:** Batch or streaming?
**Roy:** Batch. Streaming is mini batching.
**Boaz:** Write your own SQL or use the drag and drop vis tool.
**Roy:** I like the SQL, come on.
**Boaz:** Work from home or from the office?
**Roy:** Office.
**Boaz:** AWS, GCP or Azure.
**Roy:** Oh, wow. Hopefully all of them, but AWS, GCP.
**Boaz:** To DBT or not to DBT? Although, I think you hinted...
**Eldad:** Not to DBT, always DBT.
**Roy:** To DBT.
**Boaz:** To Delta Lake or not to Delta Lake?
**Roy:** Delta Lake.
**Boaz:** Okay. Thank you. I think Roy answer differently than typically.
**Eldad:** Yes.
**Boaz:** I think you're the first one who said from the office, like home or office.
**Eldad:** Yes\*\*.\*\* Everyone was confused.
**Boaz:** People typically say either home or both.
**Eldad:** Yeah.
**Boaz:** Nobody's just says office, except Roy.
**Boaz:** How many kids do you have?
**Roy:** I have one. Right now, we're living in a very small apartment so working from home is super hard, but also, I like the fact that office is where you work and home is where you live.
**Boaz:** They're separate.
**Eldad:** Old school.
**Roy:** Yeah, old school.
**Boaz:** Nice.
**Eldad:** Nice.
**Boaz:** So, what were some of the bigger data challenges or bigger projects you guys had at Explorium in the last year?
**Eldad:** Traumas, huge success, huge surprise.
**Roy:** As I said before, one of the biggest challenges I think that we have is that we kind of have to be good with any data so the platform and the infrastructure that we build has to be kind of generalized from inception which is hard. And most of engineering organizations and I'm guessing that also your engineering organization, is trying to the right thing and not generalize too early, and not try to have the wrong abstractions over things but when we kind of have to, because in a way, if we're not abstracting in the right way, or if we're not generalizing enough, then with any new source that we find, we kind of have to tweak everything around it. So, this is one big challenge I think that we have. One other challenge that I think is interesting is that you have to understand your user when he uploads, if you look at the auto-ML flow, for example. When the user uploads his data, you have the challenge of understanding exactly what he's searching for and exactly what his data means and that's one challenge and also what he's actually searching for.
Because sometimes you'd be surprised it's not the features that bring the most correlation, sometimes it brings more knowledge. Sometimes knowledge is important when you're doing that analysis. We have users coming to us saying, this is an interesting feature. It's not always with the highest correlation to a target, it's not always with the best statistics, but it's interesting because it tells me something about my business that I didn't know. So, this is another challenge and democratizing this data platform that we built. This is also a big challenge because you have to enable data scientists, both in Explorium and outside and data analysts who actually work with high volumes of data, complex pipelines, and be able to build their own processing pipelines and features and you have to enable them and abstract them from the engineering underneath. You asked before about DBT on top of Spark, I think the cool thing is that you're just writing SQL queries and underneath it runs on hundreds of machines on Spark.
**Eldad:** This is amazing. I remember the days it wasn't long ago where people are so excited about Scala and Java and writing those Spark jobs and owning the threads and the machines and the hardware and the wiring. And it was all poof, it was all gone. I think part of that, or maybe a big part of that is cloud native data warehouses like Snowflake and BigQuery that actually kind of taught many of us that it's okay to abstract, to simplify, and it's okay to have that decoupled from your day to day so thank you Snowflake and BigQuery for teaching Databricks that SQL is good. Now we see everyone, many people we talked, they're using SQL over Spark. That's kind of the biggest change we're seeing moving from developing it to actually declaring it using SQL and as you said, you love it and most people do love it, and it's a good change. SQL is back. We were confused for a few years. We had no SQL, we had new SQL, we had side SQL.
**Roy:** And now SQL is back.
**Boaz:** This is the longest secret rant I've heard.
**Roy:** Exactly. And it's important to rant about it.
**Boaz:** Okay, this is all exciting stuff. Let's talk about it from a more, maybe personal perspective. Tell us about something that didn't go well. Tell us about what we call an epic failure that you guys ran into. Maybe an approach that didn't work well, lessons learned and such.
**Roy:** I have a personal failure so I'm going to talk about my personal one, to own it. You asked me about batch versus streaming. When I started Explorium one of the things that was clear to me is that we need to have this kind of stream ingest, complex stream ingests, that we will have to report some events because we're working with such a variety of sources some of them are more dynamic by nature, APIs, for example. We have many of those and we needed some way to dynamically both enrich using those APIs and then propagate the data again to the data lake and re-ingest it and reprocess it. So, I was building this wonderful Kafka based streaming connector for everything with the best abstraction in the world and it flunked essentially.
**Eldad:** Abstraction is slow.
**Roy:** Yeah, but you got to learn, you got to learn the hard way.
**Boaz:** Why didn't it work?
**Roy:** It was a premature obstruction in a way. And it didn't match the way that now we're looking at processing because we have so many processes running offline because we have to do this complex modeling and connect different sources into one knowledge that we do which happens offline in batch jobs. And it's very important to be able to support quality at scale. Data quality is one of our team's priority and challenges that how do you maintain quality over such a variety of data sources. And there isn't really a good way to provide super blazing fast latency along with quality, because there is some processing that you have to do behind the scenes and so that streaming approach was a bit premature. And now when we're looking at streaming, we're looking at streaming as an additional entry point to those periodic jobs that can run or more batch type of jobs that can run, and look at streaming as it's just another way to get data into our lakes, warehouses, and then, from there to processing.
**Eldad:** So, what you're saying, I might stream data in, stream on write, but it's always batch on read because the batchy part of the schema will make all the streaming one's batchy. So streaming is a one-way ticket, and then you start needing to analyze.
**Boaz:** We see it often. Often, we'll see streaming as just the way I put data in my lake or whatever, but it doesn't mean that end users really enjoy that streaming and that low latency of the data coming in. Sometimes we see that, but definitely it's actually rare.
**Roy:** For a lot of cases, we have updated live data. For example, take weather. Whether is something that, first of all, how much history of weather do you keep and how you do it? So, with whether we work with APIs, for example, when users have to have the data for now, or today, or one hour ago, or right now, you have to deliver it through APIs in live in real time. Every recipe and every auto ML pipeline that we build internally also has this real time face where the user wants to consume it because he has his user or his customer waiting in checkout or whatever. So, we have real-time, but the complex processing and understanding exactly what is the data and modeling it, this happens in batch,
**Boaz:** Let's now move to something positive. Tell us about something you're proud of, a big win in your data work.
**Roy:** We have a big win. I think that combining modeling with the right serving infrastructure, it is a tool that we build internally that enables us to have a more quality matching capabilities over variety and complexity of data. I think that the win was when we started onboarding more and more and more data sources and you have this feeling when you take this new product to production and it works.
**Eldad:** A machine, a working machine.
**Roy:** Yeah, so you kind of reflect and you say, okay, so it was worth investing all of this time really understanding modeling. Data is always talking about something in the real world. Always there is something in the real world that generated this data and when you're actually able to model the real world correctly, then somehow the data behind the scenes kind of falls into place when everything kind of makes sense and I think this is what we saw, and this was really exciting.
**Boaz:** This is a project. Initially we started talking about data engineers versus software engineers and we often talk about the boundaries of blurring between the two. So, such a project where you build that super-interesting flow, do you consider this a data engineering challenge, a software engineering challenge, or both? And which skill set did you need internally to deliver that?
**Roy:** Wow. That's a question I actually talk to a lot of people a lot - what is a data engineer and how is it different from being a software engineer or big data developer? So, I look at the developer world, we have software developers and you have software engineers and you have data developers and you have data engineers. I look at development like writing logic versus engineering, which is construction actually, in a way it's construction. Foundations and understanding things in lower levels.
**Eldad:** So, you're saying that in many ways you're dividing engineers into two groups, those who are responsible to generate the data and deliver it and those who build something on top of it.
**Roy:** I'm honestly not sure where the line is because I'm also looking at the line between where is data engineer meeting an ML engineer, or a software engineer, are they the same person? But I think essentially that the problem I was talking about was the data engineering problem, essentially. Because it had this element of data model and schemas, indexing, how do you index data correctly? It has a lot of those elements that I think makes it very data engineering. So having these schemas and model on one hand, but understanding, which is the right infrastructure to kind of hold everything in place, so it could scale natively I think it's a more of a data engineering problem.
**Roy:** You mentioned quality a lot, data quality and the importance of data being at high quality. How do you treat quality internally within your data pipelines and flow and data stack?
**Roy:** When you're thinking about finding the right data, when you kind of try to think, okay, my user wants to use the platform to actually find the right data, sometimes it doesn't even know what is the right data. It has two elements to it, I think. All of them are under quality in some, in some form, but it has kind of the matching problem which is when a user is coming and he's talking about a certain company or a certain place. How do I know to match it? How to find the geography, the place that the user is talking about, or the organization that the user is talking about, or the combination between the two? So matching is one thing that leads, so our approach there was enabling experimentation. Because it's very hard to get like ground truth because almost every aspect of data that you look at has like those amazing challenges and complexities that only when you start getting your hands dirty you understand what it is, what are they.
So, matching really affects quality. We treat matching as an experiment as understanding how do we tune exactly in every case, the system, and this is where enabling a generalized system that works well with metadata and enables the data owners and the domain experts to change and tweak it and play with it to see that it fits with the real world or with their expectations, so relying heavily on domain experts here. So, this is one thing, the other aspects are correctness. Even if I found the right organization, the right place that the user was talking about and I want to get back points of interest, for example, I need to make sure that if I said, there's a coffee shop there, then there is a coffee shop, that the data is correct.
So, I probably found the right place, which is one thing, but now I need to make sure that I retrieved back data that is correct. And then after you did these two, which are super complex, you have to ask yourself, is this relevant? If my user is trying to predict the sales of his products or optimizing routes for shipping or whatever, is having a coffee place something that he needs to know? Maybe if there's a coffee place in Tel Aviv, people are stopping their car to get coffee and there's always you know, cars, tailing, it takes five minutes more or whatever so those small nuances. So, combining all of those, I think this is how we treat data quality.
**Boaz:** Interesting. Thank you for that. What gets on your nerves the most in your daily work with data? What contributed the most to your hair growing whiter?
**Eldad:** Frustration.
**Roy:** Inconsistencies. We work with a lot of data vendors and you can actually see vendors that they own their data and they provide a consistent stream or consistent delivery. Schema is our consistent file format. You know what, I'm writing my answer. File formats. It's 2021, use compressed RK, use something with Schema. It's not that hard. We were working with textual data, delimited with pipes, and weird stuff and when you compressed with weird compression, you ask yourself why?
**Roy:** I think that's a good new corner we need in the podcast.
**Eldad:** Yes, data format.
**Boaz:** Not data formats, just let people let off steam from the stuff they have, because let's face it, engineering has some pretty nasty parts in our day to day.
**Roy:** Data counseling with the Data Bros.
**Boaz:** So, the rant about your stuff you hate corner.
**Eldad:** Do you support XML as a data source?
**Roy:** We have to.
**Eldad:** Wow. That's amazing. Nice.
**Boaz:** Wait, what are you saying? If you hear me then?
**Roy:** Then use a compressed parquet, Snappy or whatever.
**Boaz:** So, data engineers worldwide, let's agree to use only that from this point forward to make our lives better.
**Eldad:** Please use faster decompression with Snappy. That's the ask here.
**Boaz:** Standardized boom.
**Eldad:** Everyone that hears that. No Zip, no nothing else. Just Snappy.
**Boaz:** How do you stay on your toes in terms of being updated in what's going on in the data world? Any sort of tips on who to follow?
**Eldad:** Aside from this podcast.
**Roy:** I'm an avid Googler when it comes to maintaining my knowledge. I'm following several... I'll try to find it, but I'm using Daily Dev, I don't know if you know them. I think they're Israeli as well, I'm not sure. So, I use Daily Dev, which is kind of a Chrome home page that connects me to really interesting sources.
**Boaz:** Daily Dev, okay.
**Roy:** I'm following several. You probably get lots of awesome pages that are awesome data engineering, awesome ML ops awesome ML engineering so I'm following those. And I'm trying to find projects to follow and kind of lead from there in a way. One example, and people make fun of me on that, but I'll say it anyway. Now I have it on record. You know Spark, everybody knows Spark. I was following Spark as a project and then I started following kind of the people behind Spark and the laboratory in Berkeley behind Spark and then they emerged with this new laboratory that brought us eventually Ray which is a project that I've been following for two or three years. No one was talking about it, no one was using it and a while ago, I'm kind of iterating back and I was getting into the project again, and I'm saying they really made progress and they released a general version like the first full production version and that was really exciting and we started talking to the team behind it.
So, I think it's a lot there. So, finding those projects and then trying to find a way to actually talk to the people. Everything is virtual today, but talking to the people is great, using like Slack channels. I think those are the places where you're actually able to, as you said, kind of be on your toes and be on kind of the verge.
**Boaz:** Yeah. I think like the community and people aspect in data engineering is huge.
**Eldad:** Huge. 36
**Boaz:** Much more than even generalized software engineering because its space is growing so fast, changing so quickly and unless you keep track and keep your eyes open to the open-source projects, to people talking about this and that you lose a lot of great things. And I think there's a lot of community power even without that being formalized that impacts this space we're in, which makes it exciting.
**Roy:** I have this WhatsApp group of data engineers and yesterday someone asked a question about how do I process? I want to process this and that data and the data is partitioned this way and so on. And I was tagging Boaz, I think it's a relevant case.
**Boaz:** It's an interesting case let's share with our listeners. Typically, when we say about communities, we talk about all the famous and open and...
**Eldad:** Big broadcasts
**Boaz:** Forums and insights we're on. Here we have something very local to Tel Aviv maybe, but there is sort of a local data engineering group on WhatsApp, which is very popular here for instant messaging, it's a very local, but very effective group or local practitioners get advice on their daily challenges with data. So maybe the tip from us would be localize more communities, it doesn't necessarily have to be worldwide. Sometimes if you're in the valley, if you're on east coast, if you're on the west coast, if you're here or there, talk to people around you, they're always in the same times or to hop on a quick call, it makes a little bit more personal. So, it's an interesting approach, definitely worked well here for us.
**Roy:** I think this makes the difference. I'm not sure it has to be local, but it has to be a group where people feel comfortable enough to ask stupid questions sometimes.
**Boaz:** Exactly.
**Roy:** And get very straightforward StackOverflow, you know, top result answers. Once you have that, you know you have the right group because people are feeling comfortable and they will ask questions and discussion will move from there.
**Boaz:** Yeah. Good point. Thank you. Okay, I think we're almost done any last famous words for our listeners?
**Roy:** Listen to the data. In a way...
**Eldad:** Nice, nice.
**Roy:** Data engineers, it's very obvious with data scientists and data analysts, I came from data. I did a couple of years of data science, pure data science, and I actually got into data engineering by accident twice. So, the beginning of my career, and then I deviated to data science, and I got back to data engineering again, somehow.
**Boaz:** Keeps pulling you in.
**Roy:** Yeah, it found me, I didn't find data engineering, but listen to the data. If you're able to have the understanding of the data scientist as an engineer, I think it your life a lot easier.
**Eldad:** So, listen to the data and if you don't like what you hear cleanse it.
**Boaz:** Listen to the data is great. It's t-shirt bubble.
**Eldad:** Yes.
**Boaz:** Abstract enough to be a debatable for hours, and it's catchy, well-done Roy. Okay, good. So, thank you so much everybody for joining us for another episode, we'll see you next time. Thanks again, Roy. Bye-bye.
**Eldad:** Bye-bye, everyone.
# Building Geospatial Support in Firebolt (Part-I) (/blog/building-geospatial-support-in-firebolt-part-i)
As the demand for location-based insights grows, integrating geospatial capabilities directly into a data warehouse unlocks a range of powerful use cases—from real-time tracking to spatial analysis and mapping. In this series of blog posts, we'll take you behind the scenes of how we developed robust and fast geospatial functionality for our cloud data platform, enabling users to seamlessly store, query, and analyze spatial data alongside their traditional datasets. We'll walk through the key challenges we faced, the technologies we leveraged, and the architectural decisions that made it all possible.
Geospatial support allows Firebolt to store and process spatial data, such as points, lines, and polygons, that represent real-world locations and features. We'll dive into the various components that make this possible, from specialized libraries to performance optimizations and data pruning.
In Part I, we'll begin by exploring the two main types of geospatial support: GEOMETRY and GEOGRAPHY. We'll explain the reasons behind our decision to implement the GEOGRAPHY type and introduce the S2 geometry library, which we use to power geospatial features in Firebolt. We'll also discuss the benefits and challenges of leveraging this library with a brief look at how S2 Cells can help speed up geospatial functions.
In Part II, we'll examine the underlying architecture and storage infrastructure that enable both high-performance geospatial functions and efficient data pruning through spatial indexing. Here, we will discuss the importance of verification and normalization of inputs during ingestion and see how Firebolt persists S2 cell coverings in its storage layer which enables powerful pruning as well as function optimizations.
In Part III, we'll take a closer look at how we achieve fast and reliable geospatial operations with S2. While S2 offers robust tools for geospatial functions, these capabilities can come at a performance cost. We'll dive into how careful engineering helps mitigate these performance trade-offs by making use of fast checks using S2 Cells before doing expensive snap rounding operations that are required for robust functions.
### GEOMETRY vs GEOGRAPHY [#geometry-vs-geography]
When dealing with geospatial data in a database, most systems use either a GEOMETRY or a GEOGRAPHY type. At Firebolt, we decided to implement the GEOGRAPHY type as it is more accurate for geographical data. Let's have a look at some of the differences between the two types:
GEOMETRY models data on a flat plane. Since Earth is, in fact, not flat, this means that distances and other relations between objects, like containment, become inaccurate the further apart and larger they are. You also have to decide on a projection to map your inputs onto a flat plane, adding complexity to your queries.
GEOGRAPHY models data on a sphere, which much more closely resembles the shape of the earth. Distances and other relations between objects remain accurate, even on a global scale. A projection is also unnecessary since coordinates are directly mapped to the sphere according to the WGS 84 coordinate system.
For example, let's have a look at the following query where we test whether the point at longitude -100 and latitude 45.1 is contained in a large area defined as a polygon within the USA.

The corresponding query using the GEOMETRY and GEOGRAPHY types using the PostGIS expansion for PostgreSQL return different results as shown below.
```sql
-- GEOMETRY: Returns false
SELECT ST_Covers(
ST_SetSRID(ST_GeomFromText('POLYGON ((-115 45, -115 35, -90 35, -90 45, -115 45))'), 4326),
ST_SetSRID(ST_GeomFromText('POINT (-100 45.1)'), 4326)
);
-- GEOGRAPHY: Returns true
SELECT ST_Covers(
ST_GeogFromText('POLYGON ((-115 45, -115 35, -90 35, -90 45, -115 45))'),
ST_GeogFromText('POINT (-100 45.1)')
);
```
### Implementing GEOGRAPHY using the S2 geometry library [#implementing-geography-using-the-s2-geometry-library]
Implementing GEOMETRY or GEOGRAPHY brings many challenges. Consider for example the common problem of determining if a point x is contained in a polygon. We can do this by choosing a point y that we know is outside of the polygon, and counting how often a line between x and y crosses a polygon boundary while making sure to account for edge cases like crossing through a vertex of the polygon, where we might falsely determine that we cross two or zero edges of the polygon.

Of course, implementing GEOGRAPHY comes with its own set of problems: We need to implement all geometric primitives in a way that accounts for earth's curvature and make sure that objects crossing the antimeridian (180 degrees east or west) or the poles are handled correctly.
Solving these problems is hard and takes a lot of time. Fortunately, the S2 Geometry library that powers Firebolt's GEOGRAPHY type and functions can do a lot of the heavy lifting. It provides abstractions for shapes like points, polygons, and line strings and can do many useful computations on them like testing whether one shape contains another shape. It even provides the tools to build powerful spatial indexes using space filling curves (more about this in a later blog post). Thank you to Google and Eric Veach for open sourcing and maintaining the [S2 Geometry library.](http://s2geometry.io/)
S2 also provides powerful tools for spatial indexing by dividing the Earth into a hierarchy of approximately square-shaped cells. This recursive division allows us to approximate complex shapes or collections of shapes through what are called coverings—sets of S2 cells that together represent the shape. One of the key advantages of using S2 cells is the efficiency of intersection tests. Since each cell is represented by a single integer, determining whether one cell intersects with another can be done using simple integer comparisons, which are extremely fast.
This efficiency is especially valuable for spatial queries with selective filters on spatial relations. For example, when querying a table to retrieve all points within a specific polygon, we can significantly reduce the number of points we need to examine by testing whether the covering of the polygon intersects with the covering of the point set. This allows us to quickly rule out large portions of the dataset that don't meet the query criteria, improving both the speed and scalability of spatial queries.

We will go into more detail of how spatial indexing and pruning works in Firebolt in a future blog post.
Of course, these are not the only interesting problems we have to solve and optimizations Firebolt can do. Stay tuned for future blog posts where we talk about some of the most interesting problems we had to tackle as well as how we tuned our in-memory and on-disk representation of GEOGRAPHY to enable fast geospatial functions and powerful pruning.
Try Geospatial today by [signing up for Firebolt for free](https://go.firebolt.io/signup) and look into our [GitHub](https://github.com/firebolt-db/firebolt-demo/tree/main/geospatial) for a cool [demo](https://demo.docs.firebolt.io/geospatial/).
# Building Uber's AI Assistant: How Genie Revolutionizes On-Call Support with Paarth Chothani from Uber (/blog/building-ubers-ai-assistant-how-genie-revolutionizes-on-call-support-with-paarth-chothani-from-uber)
In this episode of The Data Engineering Show, the bros speak with Paarth, a Staff Engineer at Uber, about his work on Genie - an innovative AI assistant that revolutionizes on-call support by combining RAG (Retrieval Augmented Generation) with agent-based automation to help engineers find solutions faster.
Listen on [Spotify](https://bit.ly/4nZd6wE) or [Apple Podcasts](https://bit.ly/415Xrll)
**\[00:00:05] Benjamin:** Hi. This is Benjamin. Before we start with today's episode, I wanted to quickly reach out on a personal note. We've just launched Firewall Core. FireVault core is the free self hosted version of our query engine. You can run core anywhere you want, from your laptop to your on prem data center to public cloud environments. Core scales out, and you can run it in a multi-node configuration. And best of all, it's free forever and has no usage limits. So you can run as many queries as you want and process as much data as you want. Core is great for running either big data ELT jobs on, for example, iceberg tables or powering high-concurrency customer-facing analytics on big datasets. We'd love for you to give it a spin and send us feedback. You can either join our Discord, enter our GitHub discussions, or you can just shoot me an email at [Benjamin@Firebolt.io](mailto:Benjamin@Firebolt.io). We'd love to hear from you. We added a link to Firebolt course GitHub repository to the show notes. And with that, let's jump straight into today's episode.
Hi, everyone, and welcome back to the data engineering show. Today, it's our pleasure to have Parth joining from Uber. He's a staff engineer there. Welcome to the show. It's really great to have you. Do you wanna tell us a bit about yourself, about your role at Uber, what you're working on?
**\[00:01:13] Paarth:** Thank you, Benjamin. Thank you, Eldad, for having me. And, it's my pleasure here. Yeah. So I've been at Uber last four years working on Michelangelo, which is our, like, SageMaker, like, ML platform at Uber. I've been working on feature store, you know, online serving at scale. Basically, we have millions of requests that we need to scale for for with with many models. And then last couple of years, I dove into Gen AI, working on rag, vector search, and building apps like Genie that we are going to talk about, where basically we had on-call productivity, you can say, pains that we wanted to solve out on our own. And that's how Genie organically just came out of a hackathon, as you're talking about. And then it just literally just started growing, last couple of years. And, Arnab and team have been one of our partner teams, and we have been working with many partner teams to really uplevel the bot because we realized that just the bot, the plain rack, doesn't work. So that's how we started this journey of accuracy improvements. Yeah.
**\[00:02:17] Benjamin:** Okay. Very cool. So for those of a kind of listeners who never heard of Genie, I hadn't heard of Genie before. Like, you wanna get the bit more context of what role it fits within Uber, kind of what workloads it serves, basically?
**\[00:02:30] Paarth:** Yeah. Yeah. For folks who don't don't know Genie, basically, Genie is, like think of it as, like, your on call assistant. Right? So and Uber is big into Slack usage. So different infra teams have their Slack channels, whatnot. I think different companies use Teams also maybe. But, basically, you go to a Slack channel. Let's say you have a problem with Spark or you have a problem with Fling or any of those open source technologies. You go to a channel where you have your infra team engineers helping you. And because these technologies are widely used, you have to wait a lot. Sometimes, engineers are dealing with, high severity issues, and they don't have time to help you with every small issue that you're running into. And documentation, as you know, engineering is not easy to maintain. So that's where, you know, the pain started coming that, okay, we are waiting on calls to respond, and then you're feeling frustrated. And that's where it felt like almost like a bot assistant would be very helpful, which can search for you from different services, different documentation, and give you the answer so you do not waiting on on calls as much.
**\[00:03:31] Benjamin:** Nice. Super cool. Can you take us through, like, that data pipe that powers Genie? Like, where are you getting data from and to what systems does it feed? How much data are you handling? Tell us more about it.
**\[00:03:41] Paarth:** Yeah. So, I mean, data sources, what we realized is for our engineers to really get help, data sources really should be internal only because we customize lot of these open source engines for, making it work at Uber scale. So we have data sources like Wiki, Jiras, Stack Overflow. We have our own version of Stack Overflow. We have Google Docs, and then we have people even storing data, custom, you can say, policies, custom information in PDFs. So there is variety of these data sources, and, obviously, source code is another one now. So we listen to and ingest all these different data sources in our own hosted vector DB solutions. And then on top of it, we basically want to do, look up semantic search. And then we're obviously what we're going to talk about, like, how we are customizing search and retrieval for each particular use case to really make it better.
**\[00:04:37] Benjamin:** Okay. Are you able to share how much data Genie is handling overall?
**\[00:04:41] Paarth:** Maybe not the specific numbers, but at this point, close to three fifty plus channels. That's quite a lot. And every channel basically has some flavor of their data that is we are ingesting. So what we have tried to do is instead of building because we are a very small team, instead of building a mega scale pipeline that just ingest all data sources and then keeps a central data source solution, we instead are giving users the flexibility to ingest what data sources they want. Right? And then what we found also is that works better in terms of search. So at least for every team, they are looking at, like, let's say, hundreds of thousands of Wiki pages, hundreds of Google Docs, and then we have design docs that we also, you know, ingesting that. So I would say per use case, you're looking at with a compressed ratio, at least few 100 megabytes to gigabyte or several gigabytes depending upon how much the user is wanting to ingest right here.
**\[00:05:37] Benjamin:** Okay. Super interesting. And as the core vector store, is that also open source technology you're customizing at Uber? Is that like an out of the box existing piece of data infrastructure? Tell us more about that maybe.
**\[00:05:50] Paarth:** Our team doesn't manage the core technology itself, but, yeah, we use OpenSearch as one of the vector DB solutions. Um, we also have our own ingrown vector DB solution as well. So which is what our sister team in the search org manages for us. But, yeah. I mean, it's it's one of the those things that has become popular in general. We're talking about Lance TV is another alternative. And you guys obviously do Firebolt as well, but, yeah, OpenSearch seems to at least our team search team seem to feel like OpenSearch would be a good solution for us. Yeah.
**\[00:06:23] Benjamin:** Right. Makes sense. Was that the first time in your life you built these, like, kind of rag style pipelines at scale? Like, it's new tech for everyone. Right? Kind of how did you even onboard into that? Like, how did you figure out how to make good technology choices? Where did you learn these things from?
**\[00:06:40] Eldad:** And does it feel like doing the Hadoop days at the beginning again? So it's kind of like, okay. Basically, blank sheet. We need to come up with the whole stack from scratch. There is some open source spread at some places, engineers going back to building new infrastructure to serve new workloads. That must be, like, three times in a lifetime kind of exciting thing. Right? Doesn't happen every day.
**\[00:07:06] Paarth:** Yeah. Yeah. 100%. And it this funny thing that while I was at Amazon, I was part of this AWS chatbot team, and we were exactly building same same thing there. And that was pre-LLM era 2018, and it almost seemed like I was like, okay. And it feels nostalgic to literally rebuild what I was building at Amazon before.
**\[00:07:24] Eldad:** So you know how everything ends up. Right? Everything ends up as a cloud native data warehouse. That, that's what happened to Hadoop eventually. Right?
**\[00:07:32] Paarth:** Absolutely. Yeah. Yeah.
**\[00:07:34] Eldad:** But, like like, looking at the tech, looking at the workloads and and those new use cases, like, it feels like innovation at scale is re-happening from scratch again. And this is exciting. To us, it's super exciting, and it must be even more exciting to you given the access to the data that you the team has. So tell us more. Like like, how did you stitch it together? What are you specifically proud of? How do you see that evolving in the next few months? Let's not even think twelve months.
**\[00:08:02] Paarth:** Absolutely. Just to even connect the journey. Right? Like, when we started, it it almost everything just happened very organically. I don't think we planned everything. But, so there was, um, your question of how do we stitch everything together? It yeah. It almost felt like they're doing what EMR was doing. , you know, you have your Hadoop and big data technology, and we needed these pipelines to basically process all this data quickly. And then that's where we started betting on Spark to really help us. And Uber, uses Spark a lot. So we were like, okay. Let's let's go with what is proven well at Uber. So we bet on Spark. We use Spark a lot for, you know, data processing, data parallelly, having to chunk it, shuffle it really quickly, and make sure we can create embeddings at scale because that's really the main two bottlenecks. You chunk your data, you have your all data ready, and then parallelly create the embeddings at scale. Right? So we had to basically scale our you can say the, um, whole infrared layer to chunk data faster to be able to create embedding set scale. And so that that meant also we had to scale our, what we call as gateway engine that, really helps us create those embedding set scales. So we had to scale all these layers. And where I see going with this, I think that's maybe, a very hard question given how fast this technology changes and the expectations of customers is growing so much. But I do see, like I mean, I think, in general, what we are seeing definitely is people want more customization, so which means more Uber internal data sources needing to come on the platform. And then, definitely, everybody is going, going with the agentic route now at this point. And, few use cases that that we have really, really nailed down and worked well, including the one that we published in the blog, you know, we definitely see agents as the way to interact, and especially MCP servers have come now. So agents, MCP servers, and your data custom data sources, stitching it all this together is really how you can make, at least in my opinion, very good, Genya app now.
**\[00:10:08] Benjamin:** The innovation, the space has been, like, super crazy. Like, how quickly there is new frameworks popping up, kind of gaining tractions, etcetera. Like, we're working on some agents now that, for example, like, optimize your SQL queries and kind of these types of things and are deeply built into the product. For me as well, like, kind of, like, having, like, this system c plus plus programming background, it's been fun learning about this, like, completely other part of tech ecosystem. And then one thing I'm personally very curious about is, like, the intersection between, like, LLM and kind of SQL query optimizers because I have this, like, more traditional compiler query optimizer background. Right? And, like, now we're kind of fusing it with LLMs and, like, figuring out, okay, like, when are LLMs great, kind of when are more traditional query optimizers that, like, reason about correctness great. Crazy how in, like, so many pieces of technologies, like, you're now, like, infusing these LLMs and kind of realizing, oh, damn. Like, you can build mind blowing things that were impossible before.
**\[00:11:06] Paarth:** Yeah. Yeah. Totally. And that you hit the point really well that this whole new suite of ways to connect with the databases have come. Right? Like, it is not there. I was reading some blogs about how Google has, basically had this agentic thing which can connect you to all Google Cloud technologies. And, I mean, we at Uber also have done I believe there's another blog by a sister team done basically query Copilot with what they call it, which is your query optimization using LLM and, you know, making sure you can write quick queries. Right? And that's a new suite of things that is not even you know, nobody even thought about that. But, yeah, you can pretty much write all your queries using LLM and build that framework.
**\[00:11:47] Eldad:** Let's not get into predictions on how painful that's gonna be on the market. But if we just focus on the optimizer within the database, it being owning that brain for the last forty years, being responsible to make all smart decisions for any user, any query, any architecture. Like, the one thing that hasn't changed is the way the optimizer feels about the users. It needs to make most decisions for that. Now it's changing. Now the optimizer has part of its brain being outsourced through an LLM, gets back the recommendation. It still needs to do optimizations at lightning speed because, right, you're getting tens of thousands of queries in. You need to optimize all queries. So you need to get feedback from the optimizer, and then your database now needs to think differently. It needs to allow to infuse those hints, that context, that insight that the new LLM optimizer went through into any query. And that's changing how database engineers and databases are thinking about optimization, but it what it also does, which is fairly excites Benjamin and us and anyone who deals with efficiency, it changes the way databases are being released and what the focus is. If my LLM can go for thirty minutes and go over 50 dimensions of optimizations and figure things out that it would take me maybe a week or two people in my company on Slack giving an advice how to change the granularity of a block or how to change the threshold of RAM or whatever optimizers do. Now it's all about how much can we expose, how much technology, how much variation, and, like, how specific can we get. Right? There are 100 ways to implement a join algorithm. Now we're gonna expose all of them because now DLLM can actually make sense out of it. Like, that was unheard of. So it's all about obstruction. Right? Like, you always have to be straight off. Every database engineer will tell you. There's a trade-off between being fast and how fast you get to being fast. It's hard to get fat to be fast so that that, right, that snowflake, abstraction that we like to say comes in, but now it's changing. So now everyone is mama's genius with an LLM. Every agent can just spend this thirty minutes on running queries and getting actual results, actual telemetry no human being has ever looked at. There are levels of telemetry that no human being should ever look at, but agents love looking at it.
**\[00:14:16] Benjamin:** Even I have certain levels of telemetry I don't wanna take a look at when building the database where agent will happily do it for me, and he'll thank me after saying kind of for the opportunity to look into it. Exactly.
**\[00:14:29] Eldad:** So, so, yeah, so it's exciting. And then instead of show tables and show statistics or explain syntax that you build for users, you're now, okay. So how do I actually provide granular, summarized, relevant profiling information for agents so they can learn fast sending less data. So this is very exciting. Um, and this is gonna absolutely change how we build software, and we're not even getting into infrastructure. So it's really exciting that you guys like, the way Uber has always been flexible and open-minded. Right? Like, it's for the last, I don't know, more than ten years, going to the Uber blog, you go there, and there's always engineers experimenting at scale. So meeting you now is exciting and really, like, hearing about it. So tell us, how do users react when you're telling them you use Genie and then you feed it with your completely random Slack chatter, and then it gives you documentation you could have drained off? How do they react to that?
**\[00:15:29] Paarth:** Yeah. And I think we've even evolved from, okay, just being like, okay. Hey. We will give you the right documentation to, like, okay. It was starting to evolve into a situation where you're like, okay. We'll also start taking actions on your behalf. Right? And that's really where I see, I think, to our conversation about the databases being smart and, you know, doing so much preprocessing before you even come to the database. I think it's getting similar situation where the bot can do a lot more now. Right? And I almost connected back to my AWS journey where we're literally trying to do the same thing in natural language where you could manage all your AWS resources using natural language. And and that's really where I think we are going with this. Any problem, be it like, okay, I have a permission issue. Okay. I I need to, let's say, update my spark, resources. I need to get more capacity. Pretty much all of that can be done now but behind the scenes not only bought looking at your documentation through MCP servers now coming and pre previously, obviously, we had the tools that we are all the internal tools that we built to have the LLM be able to leverage it. All of them can pretty much now take actions on your behalf. Right? And that's where agents become so powerful that you can have different sub agents which can take specific action for you, and then you have your this intent agent supervisor which does preprocessing of the query and takes and figures out, okay, which particular sub agent is really the right one for me to, you know, resolve this question? And then that sub agent goes, takes actions, look at the documentation, can do a lot more things, because of this whole agentic framework that has come with LangGraph and other technologies, obviously.
**\[00:17:08] Benjamin:** Nice. Super cool. So if you contrast this to your time building similar looking technology for the end user at AWS, right, like, it is a completely different technology stack, right, in the sense that it's like you solve the same problem, but it's like man or humankind figured out, like, a smarter abstraction to actually solve these problems in a better way. Like, maybe take us through that also. Like, do you feel like the things that you did then at Amazon helped you become a better engineer in this new age, or do you actually think, okay. It's like going back to kind of zero and kind of rebuilding all knowledge you have in this space?
**\[00:17:44] Eldad:** Just my instinct that I've build over my career, and I can't explain that. Like, yeah. I think I mean,
**\[00:17:50] Paarth:** I feel like as you're saying, right, Piel, that right? Like, I mean, no instinct that you derive from goes based. Right? So at Amazon, whatever we built, it was pre-LLM era. So, definitely, I I would say the technology was maybe not as robust mature before. , but that intuition that, you know, you comes from, okay, building this kind of bot, I feel like that intuition came again as we were starting to see this technology come, and we're like, hey. This looks like, okay. Where you can pretty much fit all these pieces together. So I almost felt like, that experience in Amazon was like a starting experience of, okay, how the chatbots can really do lot more. And then this was like a stepping stone to say, okay. Now the technology is finally there. Now you can stitch everything together. It almost feels like going from level zero to level one as building the same similar experiences again. Yeah.
**\[00:18:39] Benjamin:** Yeah. It's funny. Like, we have these conversations every now and then with, like, data engineers. Right? Like, similar to what Eldad said. It was like, okay. Like, back in the Hadoop days and then, like, kind of, like, modern cloud data warehouses and then, like, next generation of data warehouses. And it feels like there's these cycles in every kind of area of technology where just you take that leap, but then some things kind of stay the same. So looking ahead, like, what are you guys working on right now? What are you particularly excited about at the moment? Kind of what are new challenges you want to solve that you might not be solving perfectly right now? Kind of take us through the next couple of months.
**\[00:19:16] Paarth:** I think where we have landed last six months, including the blog we published. Right? Like, so I think we found out that how we can do this agentic solutions, and I think what I call is agentic Genie now. Basically, Genie was our you can say traditional plain React rag style board. Now genie has become agentic genie. Right? So what we have seen with several use cases is agentic genie works well when designed well. When you've analyzed the problem of which type of subproblems the bot should resolve per channel, per use case, and then you go go with, solving each subproblem with each sub-agent. If you do that way, it works really well. Right? And that's what we have seen as delivered good success. I think where our challenge right now, which is where I'm still thinking and it's not something we have designed or thought too well, but it is funny that how or funny or good in a way that how cursor and other IDs have up leveled what agentic experience can mean for code. And I I kind of want to take that same cursor like experience for agentic genie, where you just come and without you having to do anything, figures out all the right sub problems for you as a channel owner, as a use case owner. We decide underneath which particular sub agents to call and, which particular MCP server tools to interact with, and it just figures out everything for you. That's what we'll where I would like this North Star to be. That way, it becomes agentic genie is like a perfect on call assistant what Cursor and other IDs have done for coding.
**\[00:20:52] Benjamin:** Nice. Okay. Super cool. So are you guys using Cursor to build Genie?
**\[00:20:57] Paarth:** I mean, every I think every developer is is likely using Cursor and any other of those, forms of ID experience and CLI experience right now. I mean, yeah, everybody uses that. Absolutely. It makes your life better for sure.
**\[00:21:11] Benjamin:** Yeah. Same here at Fireball. Very cool. So maybe zooming out from Uber a bit. Right? Like, what are you excited about in this space in general? Is there, like, kind of, like, how do you actually stay up to date when building these things? How do you learn about other how other companies are kind of tackling similar problems? Are there meetups you're going to in the Bay Area? Are you just reading blogs and watching YouTube videos all day? Like, take us through kind of that.
**\[00:21:36] Paarth:** Yeah. No. Absolutely. Obviously, meetups is a great way to meet. I think I've been to some of those recent, like, um, meetups. One of them was by Meta, which is called Scale at AI. It was really, really very well done meetup, and I was impressed with how Meta has optimized every single layer of infra, and they have done amazing. They have developed something called MetaMate, which is super cool. And when you look at those experiences, you're like, wow, there is so much more to do here. So I definitely deriving inspiration from all of the, you know, smart co peers in, different industries, different companies, and that's definitely one day. And then I think I I do see still, like, you know, when you apply those inspirations to your use case, there is still a quite a bit of steep learning, trying out different things. So nothing just works what work for other companies just like that. , there is, you know, dollar challenges, core cost challenges, scale challenges that you have to go back and redraw and figure out how it works. But, definitely, I think meetups, blogs, podcasts, like, what we are doing together. I mean, yes. I think looking at all of this is definitely one way to learn. Yeah. And this I think the Bay Area is really good in that sense that there's a lot happening here. So people are really excited to build things. So you find those exciting builders from startups to all the way companies that I've opened yet. Just trying so many things. So it's very humbling to be in Bay Area, and, you know, you just know that you're never done learning because there is new things coming out here in the market all the time, every single day right now almost.
**\[00:23:11] Benjamin:** Yeah. I mean, you're at the epicenter of, like, this next generation of software, basically. Very cool. So this was super interesting, Parth. Any closing thoughts from your side? Like, something you wanted to really kind of chat about, something you wanted to bring up on the data engineering show today?
**\[00:23:27] Paarth:** Basically, I would just say that I think people who are trying out, all these technologies, I would recommend everybody try as many things as you can. And, obviously, keep a problem in mind because just trying in itself, you can just keep endlessly experimenting right now, and you will not go anywhere. I think having a problem in mind always helps. That way, this it's the, energy is little bit focused and directed. I would say that is definitely my learnings, and I think keeping eye for experimentation open. I do see one little bit of a challenge building things. Right? That you build something, you take a bet on one technology, you build something, and then underneath, new things have already surfaced that keep you outdated. So there is I would say it's also a little bit of a pressure situation that whatever you're building is not enough because the expectation has already gone to the next level. So, um, I think it is a little bit hard to, you can say the pace is too fast right now. So as a developer, keeping your customer happy is hard, but then also setting good expectations with your customers is very, critical more critical than before. Because of what we shipped before in software, you know, the pace has completely changed now. So having that expectation and, obviously, keeping your, um, eye for getting feedback from customers is equally important. Um, being humble enough that you know that what you developed is not perfect, and you need to keep iterating and keeping making better. I would just say those would be my quick takeaways as we experiment and go with the next flow of technologies that are coming in here.
**\[00:24:59] Eldad:** Amazing. Nice.
**\[00:25:00] Benjamin:** Couldn't agree more. Thank you for having joined us on the data engineering show today, and we look forward to meeting in person when we're in the Bay Area next time.
**\[00:25:08] Paarth:** Absolutely. Thank you, Benjamin. Thank you, Eldad. It was great to to you guys, and I I really looking forward to meeting you guys also in person very soon, hopefully.
**\[00:25:16] Outro:** Sounds great. The data engineering show is brought to you by FireVolt, the cloud data warehouse for AI apps and low-latency analytics. Get your free credits and start your trial at firebolt.io.
# Caching & Reuse of Subresults across Queries (/blog/caching-reuse-of-subresults-across-queries)
### TL;DR [#tldr]
Ever wonder how you can speed up complex queries without additional resources? In this post, we will show you how Firebolt optimizes query performance through caching and reusing results of parts of the query plan (= subresults) across consecutive queries.
### Introduction [#introduction]
At [Firebolt](http://www.firebolt.io) we are building a data warehouse enabling highly concurrent & very low latency analytics. Another way to frame this is that we strive to be faster than the blink of an eye:

Firebolt's main use cases are "data intensive applications", such as interactive analytics across various industries. These workloads typically consist of high volume, sub-second queries stemming from a mix of tens or even hundreds of patterns.
Such repetitive workloads can benefit tremendously from reuse / caching. In analytics systems, caching as a concept is ubiquitous: from buffer pools over full result caching to materialized views. Here, we will present our findings in a surprisingly little-used approach: caching subresults of operators. The idea itself is not new, with first publications appearing in the '80s, e.g., [\[Fin82\]](https://dl.acm.org/doi/10.1145/582353.582400), [\[Sel88\]](https://www.sciencedirect.com/science/article/abs/pii/0306437988900142). More recent results include [\[IKNG10\]](https://dl.acm.org/doi/10.1145/1559845.1559879), [\[Nag10\]](https://homepages.cwi.nl/~boncz/msc/2010-FabianNagel.pdf), [\[HBBGN12\]](https://vldb.org/pvldb/vol5/p1436_alexanderhall_vldb2012.pdf), [\[DBCK17\]](https://cs.brown.edu/~kayhan/papers/hashstash.pdf), and [\[RHPV et al. 24\]](https://www.amazon.science/publications/why-tpc-is-not-enough-an-analysis-of-the-amazon-redshift-fleet).
In this blog post, you will discover how we implemented a cache for subresults of arbitrary operators in the query plan, including hash tables of hash-joins, and how this can significantly enhance the system's performance. By sharing the impact we have seen in our production environment, you will gain insight into how caching can help reduce query times and improve efficiency for your own workloads. Additionally, we will explore a variation of the eviction strategy which we benchmarked and tuned on real-world data, showing how it outperforms traditional LRU, ultimately saving you time and resources.
Since our cache for now is purely in main memory, we ideally store the subresults in compact data structures which have a small RAM footprint. We therefore devised a custom, space-optimized hash table for our novel "FireHashJoin". With it, we achieve more than 5x memory savings (as measured on production data) compared to the hash table implementation we used before. It thus enables us to cache significantly more of such subresults in the same amount of RAM. We will present this in a follow-up post.
### What do we Mean by "Subresult"? [#what-do-we-mean-by-subresult]
Let us have a look at the following made up query based on the [TPC-H](https://www.tpc.org/tpch/) benchmark schema (see, e.g., page 13 [here](https://www.tpc.org/TPC_Documents_Current_Versions/pdf/TPC-H_v3.0.1.pdf)).
Below we show a simplified query plan for this query. This is a graph which has "operators" (such as "*Aggregate*") as nodes. In our figure the data flows from the bottom to the top and each operator takes care of one processing step, such as aggregating the incoming o\_totalprice & row-counts and grouping them by n\_name. For a more detailed explanation, see, e.g., this [Wiki page](https://en.wikipedia.org/wiki/Query_plan).
So here is the query plan for the example query above:
For each operator, we call the entire data streamed out of this operator a **subresult** – it is what the operator and the subplan below it have computed. Example: all combinations of c*custkey, n\_name streamed out of the lower-right \_Join* together form the subresult of that *Join* operator.
We also refer to the subresult streamed out of the *Aggregate* operator at the top as the **full result**.
For the final artifact which we consider a subresult, let us add a **quick side note about our *Join* operator**. We implemented it as a so-called hash join algorithm. This is perhaps the most common join implementation used in databases. Briefly summarized, the algorithm proceeds in these two phases (for more details, see, e.g., this [Wiki page](https://en.wikipedia.org/wiki/Hash_join)):
1. In the *build phase* it computes a **hash table** for the "build side" of the join. In our case this is the right side of each of the joins depicted above. For the lower join, the computed hash table maps n\_nationkey to n\_name as read from the nation table.
2. In the *probe phase* the algorithm iterates over all rows of the "probe side" of the join (in our figure, the left sides of the two joins are the probe sides). For each row it probes the hash table, i.e., looks up the corresponding key coming from the left. In our example for the lower join, the algorithm would lookup n\_name given the c\_nationkey coming from the left side. The operator then outputs the combination c\_custkey, n\_name for each of the probe rows.
We also consider the **hash tables** computed by Join operators as "special" subresults of the query plan. In our example, the lower join builds a hash table n\_nationkey => n\_name and the upper join a hash table c\_custkey => n\_name.
So to summarize, each operator produces a subresult (the data it streams out) and *Join* operators also create hash tables as "special" subresults.
### Why and How to Cache & Reuse Subresults [#why-and-how-to-cache--reuse-subresults]
Let us now have a look at *why* one would cache and then later reuse subresults of consecutive queries. For similar approaches and in-depth motivation & impact, see, e.g., [\[IKNG10\]](https://dl.acm.org/doi/10.1145/1559845.1559879), [\[Nag10\]](https://homepages.cwi.nl/~boncz/msc/2010-FabianNagel.pdf), [\[HBBGN12\]](https://vldb.org/pvldb/vol5/p1436_alexanderhall_vldb2012.pdf), and the more recent papers [\[DBCK17\]](https://cs.brown.edu/~kayhan/papers/hashstash.pdf) and [\[RHPV et al. 24\]](https://www.amazon.science/publications/why-tpc-is-not-enough-an-analysis-of-the-amazon-redshift-fleet). [Below](#caching-hash-table-subresults) we give more details about the actual impact of our implementation in production.
For full results the motivation for caching is pretty obvious: why recompute if the exact same query comes in a second (or third or …) time. If the full result is small enough, we can keep it in a cache and reuse it as long as the data scanned did not change. Workloads with many consecutive, exactly identical queries are actually more common than one might think. The recent VLDB paper analyzing Amazon Redshift's workload [\[RHPV et al. 24\]](https://www.amazon.science/publications/why-tpc-is-not-enough-an-analysis-of-the-amazon-redshift-fleet) impressively shows this somewhat surprising fact. [Below](https://www.firebolt.io/blog/caching-reuse-of-subresults-across-queries#caching-hash-table-subresults) we also share some stats on our production workload.
Many query engines have a full result cache. Often it is simply based on a hashmap of query-strings to full results. We do this in a somewhat different manner which has advantages given later in this section. But we actually go much further, our system supports caching (and reusing) results of any subplan. These subresults are placed in our in-memory **FireCache** which can use up to 20% of the available RAM. Here is how we determine which subresults to cache:
* On the one hand, our optimizer may add so-called *MaybeCache* operators above any node in the plan. As the name suggests, this operator may cache a subresult – if it is not too large. It may later fetch from the cache and reuse a subresult if the exact same subplan (with the same data being scanned) is evaluated again. Currently, the optimizer places a *MaybeCache* operator 1) at the top of the plan, for a full result cache and 2) at nodes where "sideways information passing" is happening to speed up joins (where the probe-side has an index on the key that is being joined on). For 2) let us only point out that this can be very beneficial for the case described in the next bullet. We will not go into more details here for simplicity. The *MaybeCache* operator is versatile, it can be placed anywhere in the plan. In the future, we may investigate adding it above other, selected operators, e.g., pipeline breakers such as *Aggregate*.
* On the other hand, our system places any subresult hash-tables built for *Join* operators in the FireCache if they are not too large. The reasoning behind this is that these hash tables are relatively expensive to compute and it is very advantageous to reuse them in case similar queries come in consecutively.
So this is what the simplified plan for our example query would look like and what would be placed in the FireCache on the first evaluation:
A *MaybeCache* has been placed at the top of the plan and it places the subresult collected on the first run into the FireCache, as do both *Join* operators with their respective hash tables.
On a subsequent run of exactly the same query (over unchanged data), the *MaybeCache* fetches the subresult from the cache and the entire evaluation can be skipped. This means the latency drops to very low milliseconds. In our toy example this corresponds to a **speed boost of more than 100x** (on a single node, medium engine running over TPC-H with scale factor 100).
But what if the user adds a filter which may be changed across each of the subsequent queries? This is where caching *actual subresults* kicks in. Say the user wants to restrict the date range, e.g., by adding ... AND o\_orderdate >= '1998-01-01' ... (for changing dates) to the query:
The following figure shows how the query plan and the corresponding cache interactions would look like if the query without the restriction was run previously:
Notice that the added filter has been pushed down by the optimizer to be right above the scan of orders. The plan as a whole has changed, therefore the subresult stored by the *MaybeCache* previously cannot be reused. On the other hand, the **right** **subplan** below the upper *Join* (surrounded by a dashed box), has not changed! Therefore the previously computed and cached hash table can be reused in that *Join* operator. This saves us from evaluating the subplan and building the hash table again.
In the toy example, this leads to a **speed boost of > 5x** on subsequent queries (even when each query has a different date restriction).
Note that we would still profit from the FireCache in case restrictions on the customer table would be added – here we would be able to reuse the hash table for the lower *Join*. To further extend the utility of the cache, the optimizer could decide to lift up such filters to have larger common subplans across consecutive queries. This is something we plan to look into in the future.
### Avoiding Thrashing [#avoiding-thrashing]
The example also already shows one potential (and common) challenge with such caches: how to handle possibly many useless items being added to the cache. Say hundreds of such queries come in per second, but each with different date restrictions. All of the final results added to the cache may never be used again. There is a danger of the cache being thrashed and useful items being evicted by these useless ones. In [this section](https://www.firebolt.io/blog/caching-reuse-of-subresults-across-queries#better-eviction-strategy) below we show how we improved on the well known LRU eviction strategy to make such thrashing less likely.
### "Semantic" Cache-ids and their Impact on the Full-Result Cache [#semantic-cache-ids-and-their-impact-on-the-full-result-cache]
In order to add and retrieve items from a cache, one needs a way of deriving keys / ids to identify them. In our case, for each *MaybeCache* and *Join* operator the corresponding cache-ids are determined from the subplan below it. This is described in detail in the next section.
Before we dive into this, let us have a look why this approach can be beneficial already for a full-result cache alone – compared to simply deriving cache-ids from the query text as is done commonly for such caches: since we are operating on the query plan, the cache-ids are "semantic" as opposed to merely text-based. Different query strings may be semantically equivalent and result in the same plan and hence same cache-ids. E.g., all comments have been removed. Ignoring comments is important for BI tools such as Looker which add metadata in comments to queries sent out.
Subsequent queries may be exactly the same, except for the metadata added in comments. The query string changes on each run, but for the FireCache this still results in the same cache-ids and therefore cache hits.
Moreover, the cache-ids are computed after the plan was fully optimized. Rewrites such as constant folding may normalize "different looking queries" to the same plan.
We can even go a step further (as of this writing this is work-in-progress) and push the *MaybeCache* operator below any operators at the top of the plan which do not reduce the cardinality. Note that since we store only relatively small subresults, there is very little (or even no) noticeable overhead when evaluating that final operator on a small subresult fetched from the cache. Why is this useful? Let us look once more at our toy example query. After running it once, we may decide that we are actually interested in the average instead of the total price, i.e.:
This results in the following plan. The final *Projection* preserves the subresult cardinality, therefore the *MaybeCache* can be pushed below it.
Even though the query has changed significantly, we still get a cache-hit and the bulk of the plan is not executed. Say, we then notice that we actually wanted to order by the average cost. The *Order* operator is cardinality preserving as well, so the *MaybeCache* is pushed below it (and below the *Projection*). Therefore, we again get a cache hit!
### Caching in Distributed Engines [#caching-in-distributed-engines]
In distributed engines with many servers, each server has its own, independent FireCache. For system simplicity, there is **no cross-server synchronization** on whether or not to use the FireCache for individual subplans. Individual servers may or may not have cached subresults available, thus some servers may answer a suplan from the cache while others may compute from scratch. Such distributed settings bring interesting challenges, described in [this section](https://www.firebolt.io/blog/caching-reuse-of-subresults-across-queries#consistent-subresults-matter-in-distributed-settings) below.
### Cache-ids for Arbitrary Parts of a Query Plan [#cache-ids-for-arbitrary-parts-of-a-query-plan]
As a well known [joke](https://martinfowler.com/bliki/TwoHardThings.html) goes, there are only 2 hard problems in computer science: cache invalidation, naming things, and off-by-1 errors. Let us now look into the first problem in the context of the FireCache. We tackle it by deriving "cache-ids" for subplans which capture the "full details" of a represented subplan, including which data is scanned over. By construction of these cache-ids, we never fetch invalidated data from the cache.
To compute a cache-id which represents a subplan (say the one below the *MaybeCache* in the previous figure above), we start at the leaves (here the three *Scan* nodes) and then propagate "subplan fingerprints" up through the plan.
A subplan fingerprint needs to represent exactly the data streamed out of that subplan. The slightest change to subplan or the underlying data, e.g., a new row added to the orders table, a new or changed filter, a different aggregation etc. must result in a different fingerprint.
For each operator we implemented a specific rule of how to combine the incoming fingerprints from its children with any parameters / configuration the operator may have (e.g., a complex, nested expression to be evaluated for a *Filter*).
We will skip over the details here and will only briefly touch on how this is implemented for the *Scan* operator. In Firebolt's managed storage, data scanned is split into immutable tablets. If data within an existing tablet is modified, a new tablet is created, reflecting the change (note that the change may be a minimal extension, e.g., reusing all files from the existing tablet, simply adding a delete vector – for our purposes here it is important though that each tablet is immutable). The *Scan* operator is thus configured with exactly which tablets it should scan on which server. Usually, this is a subset of the tablets a table is composed of (e.g., narrowed down by primary-key filters or distribution of tablets across servers). So to compute the fingerprint of a *Scan* operator, we simply hash all tablet-ids and combine these hashes with hashes of the scanned variables.
By construction **we therefore avoid ever serving "stale" cached results for a query!** To illustrate this, let us look at the example plan above once more. Say, from one run to the next nothing changed except that a single row has been added to the orders table. This results in a different set of tablets to be scanned in the leftmost *Scan* node – which in turn gives a different fingerprint for that node. Since the fingerprint is then propagated up through the plan, the fingerprint (i.e., cache-id) used for the *MaybeCache* at the top also changes => We do not use the previously cached, now stale subresult. Instead a new one is computed and stored in the cache.
Note that since nothing else in the plan has changed, both *Join*s actually still receive the same cache-ids for their hash-table subresults as before => These can still be reused. We hope this gives an idea of how the propagation of fingerprints nicely leads to cache-ids which enable reuse in unchanged parts of the plan while avoiding to ever serve stale data.
To make hash-collisions of the fingerprints extremely unlikely, we use the 128 bit [SipHash](https://en.wikipedia.org/wiki/SipHash) function. To date, there are no known collisions of such SipHashes.
### Consistent Subresults Matter in Distributed Settings [#consistent-subresults-matter-in-distributed-settings]
After so-called *Shuffle* operators (which distribute data across servers), one needs to be careful to not cache & reuse possibly inconsistent subresults. This is relevant for SQL queries which may give non-deterministic results on subsequent runs, even if the underlying data did not change. As an example, the following query returns 3 nations from the underlying table, but it is not defined which nations.
Similarly, SUM(o\_totalprice) due to floating-point rounding may give (slightly) different results depending on the order in which the underlying double values are summed up.
Let us look at a made up example to give an idea why it can be problematic to cache non-deterministic subresults on different servers. Say we execute the following query on an engine with two servers.
Let us assume that the small 3-nations-subquery is executed on one server and the subresult is sent to both servers with a so-called "broadcast" *Shuffle*. Also, say that we split the evaluation of the join with customer across these two servers (each server taking care of roughly half of the customer table). We may pre-aggregate the count on these servers independently as well. In a final step the pre-aggregated subresults would be shuffled to one server and merged there to compute the final result: **the count per nation for three arbitrarily selected nations.**
If we are not careful, the subresults of the 3-nations-subquery may be cached independently on the two servers (note, as mentioned above we do not have any cross-server synchronization of the FireCaches). It could happen that we end up with two inconsistent copies of that subresult on the two servers – for instance, due to different evictions happening on these servers. E.g., we may end up with the three nations "GERMANY", "ALGERIA", and "CANADA" being cached on one server and "FRANCE", "INDIA", and "BRAZIL" on the other server. If we would use these cached subresults for the two parts of the join being evaluated, **the final result of the query would have 6 nations** instead of only 3 as expected & the reported counts would likely be off by \~50%! This is of course incorrect and needs to be avoided.
We therefore need to avoid caching possibly inconsistent subresults. To this end, we disable caching on nodes in the plan if these two conditions hold: 1) "separate copies" of a subresults may exist, i.e., post-shuffle and 2) these subresults may be non-deterministic even on unchanged data. We compute this via depth first decent in the plan. In parts of the plan where there are parallel paths, we only allow caching of deterministic subresults.
Note that this does not mean we need to disable caching entirely above a *Shuffle*: if the "separate copies" are combined again into a single copy further up in the plan, from then on even caching non-deterministic subresults is OK again. In particular, the important case of caching full results is always OK, since these full results are always on a single server and there are no separate copies.
### Better Eviction Strategy [#better-eviction-strategy]
As always with caches, there is a risk of "thrashing": if many subresults are added to the cache in quick succession, important ones may be evicted by unimportant ones. I.e., subresults that are not reused in the future may cause others that are still useful to be thrown out of the cache if the memory limit has been exceeded. The FireCache is relatively large, we reserve 20% of RAM for it currently and the subresults stored are in most cases rather small (different thresholds, e.g., 1 MB for the full results). Nevertheless it is important to choose a good eviction strategy to **maximize reuse and thus optimize the processing time saved by the cache**.
The most common strategy is the [least-recently used](https://en.wikipedia.org/wiki/Cache_replacement_policies#LRU) (= LRU) eviction policy. It is easy to implement, fast, and intuitive, since the item not accessed for the longest period of time may be the one least likely to be accessed again in the near future.
In our case we have additional information readily available which could be, but is not used by LRU: we know how long it took to compute a subresult and we know the size in RAM of the subresult. Also, it would be nice to track how often an item was used – this might also indicate how useful it is in the future.
### The "NEW" Strategy [#the-new-strategy]
We thus combined these in the following manner to form our NEW strategy. This is essentially the same as the policy described in the section "Benefit-Based Result Selection" in [\[NBV13\]](https://ir.cwi.nl/pub/21352/21352A.pdf).
* As base-weight, we combine the processing time needed to compute a subresult (and to potentially save when reusing it later) with its size as:
time in milliseconds / size in bytes
In a sense this gives us "how much bang we get per buck".
* For the frequency component we essentially sum up the (decayed, see next bullet) base-weights per access. E.g., if there were 3 consecutive accesses close in time, the final weight of an item would be \~3x its base-weight.
* Finally, to keep an LRU-like component, we decay the weight exponentially with respect to the current vs. the last access time. For this we take a "half life time" of 5 minutes. I.e., after 5 minutes the weight of an item is decayed by half. We experimented with different half-times and the differences were small, but 5 minutes seemed like a good tradeoff for different workloads.
When the FireCache runs out of memory and needs to evict a subresult, it chooses the one with the smallest weight. Computing this is relatively expensive, but with some tuning, we got this to run in 500 to 1000 nanoseconds per eviction which is OK for such "coarse grained" caches.
### Performance in Simulations with Different Cache Sizes [#performance-in-simulations-with-different-cache-sizes]
We benchmarked NEW vs. LRU on real-world data of several customers and compared the total processing time saved by the two strategies. Below is the result for the workload of a selected customer over multiple days. In our simulation we varied the cache size and show the percentage of time saved by the cache:

The results for very small and very large caches are not so interesting: for the former nothing can be cached and for the latter essentially everything may be cached and then reused many times by these very regular workloads – very few exact duplicates, but a lot of reuse potential due to \~100 query patterns, each with large subplans occurring many times in consecutive queries.
For the interesting mid-range of cache-sizes (4GB to 18GB) we get up to **2x more processing-time savings with NEW compared to LRU**, see also this figure which shows the ratio:

## Impact [#impact]
### Caching Hash-Table Subresults [#caching-hash-table-subresults]
For the interesting case of customers with mostly unique queries (i.e., few exact duplicates), stemming from tens / hundreds of repetitive query patterns, we often see very nice speed boosts of 10-100x for individual queries. This is, e.g., the case for the customer mentioned [above](https://www.firebolt.io/blog/caching-reuse-of-subresults-across-queries#performance-in-simulations-with-different-cache-sizes) when discussing our NEW eviction strategy. The customer's SQL queries are often > 100 lines long and contain many joins. But, consecutive queries from the same pattern change only in one literal of a crucial WHERE restriction. Note that in this case the savings come from reusing hash tables in *Join* operators previously stored in the FireCache.
The chart below shows the overall hash-table *build-time saved vs. build-time spent*. This is across all customers in production over several weeks this year. The thick blue line gives the ratio of saved / spent. It shows that **for hash-table build times we overall reap 10x savings** on many days with some spikes being even higher\*\*.\*\*

### Full Result Cache [#full-result-cache]
The *MaybeCache* added to the top of the plan has obvious advantages. As mentioned many other systems add a similar cache (although to the best of our knowledge without the full benefits of the "semantic cache-ids", see [above](https://www.firebolt.io/blog/caching-reuse-of-subresults-across-queries#semantic-cache-ids-and-their-impact-on-the-full-result-cache)).
It is interesting to see for individual customers how much they may profit from such a full result cache. The answer is some customers can profit tremendously: here is a chart with the simulated time savings of the top 15 customers (anonymized), ordered by the potential savings across \~5 days of production queries. I.e., customer "A" may save > 95% of their total compute time if queries would be answered by a full result cache. Note: we will redo this measurement to report actual time savings when the full result cache has been in production for a longer period of time (this was launched only recently).

## Conclusion [#conclusion]
This blog post gave a deep dive on how at Firebolt we give our customer's queries a speed boost by enabling caching and reuse of subresults. For some of the interactive use cases we serve, this is of crucial importance. Without caching results of complex subplans certain query patterns would not achieve the low milliseconds latencies required.
As shown in the "Impact" section above, this also in general yields tremendous savings for our customers, e.g., by saving up to 10x in processing time spent on building hash tables for joins.
We hope we were able to give an idea on how with propagation of fingerprints & derivation of cache-ids we can determine in fine-granular manner which parts of a plan can be reused and which should be recomputed. We dove into the challenges posed by distributed settings. Finally, we gave a quick overview of a modified eviction policy to improve the usefulness of the FireCache under heavy loads when many items are being added (and conversely need to be evicted).
Since our cache is purely in memory for now, compact representations of data structures to store subresults are very beneficial. In a follow up blog post, we will describe in depth a custom, space-optimized hash table which we devised for our novel "FireHashJoin". With it, we achieve more than 5x memory savings (as measured on production data) compared to the hash table implementation we used before. It thus enables us to cache significantly more of such subresults in the same amount of RAM.
# Data engineering from the early 2000s till today - BlackRock (/blog/data-engineering-from-the-early-2000s-till-today-blackrock)
When it comes to data management, have we come a long way since the early 2000s? Or has it simply taken us 20 years to finally realize that you can't scale properly without data modeling. With over 20 years of experience in the data space, leading engineering teams at Cisco, Oracle, Greenplum, and now as Sr. Director of Engineering at BlackRock, Krishnan Viswanathan talks about the data engineering challenges that existed two decades ago and still exist today.
Listen on [Spotify](https://open.spotify.com/show/6hMdnrFKlPbia2k6MkFs8U) or [Apple podcasts](https://podcasts.apple.com/us/podcast/the-data-engineering-show/id1561927688?uo=4)
**Benjamin:** and post. Cool. All right. So welcome back everyone, kind of to the data engineering show. Good to have you. So it's the second episode in a row now without the other data bro, Eldan, because he has out today, but he'll be back in the next episode. But anyway, so let's jump right in. It's a total pleasure having Krish Nan here today. He's a senior director of data engineering at BlackRock. And like he's done... He's done everything, really. It's so cool having you on. He started out at Cisco, spent 10 years there, then went more to the vendor side as a principal PM at Oracle and later as a director of product management at Greenplum. And then most recently, he's actually been as a kind of senior director of data engineering at BlackRock. So great to have you on today, Krishnan. I look forward to talking about kind of the data challenges you're facing, kind of your background and everything. Yeah, do you just wanna kind of give a quick intro of yourself?
**Krishnan Viswanathan:** Awesome. Thanks, Benjamin, for having me. This is a pleasure. And like you said, I've been in the industry for almost a quarter of a century. So now that I think back, it's been a long time coming. But I'm really excited to be part of this conversation. I think the data space and the data explosion in the last couple of decades and all the modernization has put a deja vu into the data. organization, uh, every day seems like a new day, but also seems like a groundhog day in some ways. So I would love to share some of the things that I've learned. And back to my background, I have been, I started off as a data engineer. Actually, it was more of a sensitivity. I did my first initial project as a, as an executive information system for Cisco. Uh, for the on the web based application and then realized at the back of it that we actually needed a data warehouse and a data model and a data cleansing and all of the things that goes with data pipelines to get that numbers right. So that was my marijuana moment.
**Benjamin:** What year was that roughly?
**Krishnan Viswanathan:** That was 1996.
**Benjamin:** Wow, so take us through the stack back then. Like, kind of how did you have a lot of options? What types of systems kind of were out there at the time?
**Krishnan Viswanathan:** Yeah.
**Benjamin:** And I guess it was all on-prem as well, right?
**Krishnan Viswanathan:** This was all, well, think about this. 1993 was when internet became kind of prevalent. Cisco was the main child, main vendor pushing that pipeline and the routers for internet. And before, so when I joined in 96 at Cisco, the executives used to get their reporting in Excel. Every week they would get, and it would be a high level number, pretty much that's it. And no other information. And then they would make phone calls to all the different departments, different groups. And then, you know, by the time they get to a problem statement, the problems are either already solved or it has escalated to an extent that it was too late to solve, right? Whichever way it went. So the premise in 96 was, hey, we are the pioneers in internet. We should build an internet technology that can help showcase what we can actually present to the executives. Out of it was born the executive information system, we call it EIS, the first generation of it. And it was homegrown. I was one of the first, one of the four engineers and eventually you are the lead engineer there. We built it with, we got based out of an Oracle ERP extract for, and then we had an Oracle a large scale data warehouse, scale to single instance Oracle data warehouse. And it had its own challenges extending, but it was that. And then the front end was a purely web-based application. And web-based application, it was pre-Java, pre-EJV, it was all pure HTML. And what we did at the back end was we used to run this extraction process and build this data staging layer, for lack of better word. because he was still, I was not a data engineer at that time and I would do so many things different, but we did that. And then we would run a process between that and the web server to create all these different HTML pages because HTML was limited, right? Web content, web configurations were very limited. There wasn't a real dynamic XML parsing. I think XML was just coming up to speed, but not even that. But anyway, the Java EJB was not even part of the picture. No servlets, right? If you think about those things. So it was all pure HTML, PHP-based query process. So that's that. And the expectation, since it's executive reporting, was they were looking at the overall high number, and they would drill down all the way to, and by region or by product, eventually to a sales rep to a transaction. to see where, if there was an issue, if there was a question, they could go all the way down. So the process itself in the beginning took about six hours for us. Interesting, and again, reliability was a whole different problem. I will talk about that later. And then the data accuracy was another level of challenges that we solved in subsequent releases. But the first version of it, for all the good things that we did, it had its own challenges. And part of it was technology and the immature technologies and the kind of band-aiding them together to make it work. And then the other part of it is, of course, we also could have gone and done a better data design process
**Benjamin:** Right?
**Krishnan Viswanathan:** and data alerting process. So that was fun days. But it was a hugely successful endeavor. If I were to just talk about the outcome of that, the appreciation we got out of the data something like that, it was actually showcased by Cisco as one of the pioneering projects that helped. In those days, Cisco used to call it the virtual close. What it basically means is it used to take like a month for companies to close their book of records. Using this application, this was one of the key applications, I won't want to call it, this is the only application. Using this application, Cisco said they could actually close their books within a day. So.
**Benjamin:** Wow.
**Krishnan Viswanathan:** from the vision to what we could accomplish, notwithstanding all the other disclaimers I put in for it, it was a success. And they were, we actually got funding to make it better. And we did do better. And the second iteration of that was a whole lot better, a whole better concept. And I'll talk about that a little bit later, but that was the genesis of my data engineering role. Actually, let me talk about the second version of it, because I led the second version of that project, which we call E-Exec. And part of that project was to find a better capability, a better tool. And this came around in 2001. So if you are all aware, and if you have followed the internet bubble, and the burst of the internet bubble in the early 2000s, I don't know Benjamin, if you were there, but
I was writing the whole Cisco bubble and I still get nightmares about that because one of my roles in this project as the exec was we would wait for the SVP of Cisco to go in and provide a confirmation on forecasting and judgment, which is basically saying, yeah, I agree with this forecast and this number is going to look good. I was the one who would say, OK, it's done. Click the process, right? Data engineer, making sure that everything is done. The financial analyst would tell me that, and then I would make sure that the process gets itself. But I would validate that everything is working. I still remember that day in October of 2000, some early October of 2000, or maybe October of 2000, when there was a huge downward revision of the forecast from what I have seen in the past. The forecast was always going up, up, up. And then it was like, Well, that was like a fairly deep decline. I was like, something doesn't look right. But I was still an engineer. I had to get the job done. I just did it and went off. Three months later, you saw the whole internet crash and Cisco going from its huge bubble. So I lived that downtown. And I lived it very, very vividly, both in my financial terms and at Cisco. So remember I said, we... Assume that this application will help executives give an early warning and give the executives a good indicator of how the business is performing and if you wanted to do a virtual close-up. Within a day they can close. This application kind of didn't give that exact, it gave them good information, not good at all. It didn't give them the exact alertness that they were looking for, which is something is wrong proactively telling them that, hey, something is wrong, you're gonna have to, you're gonna miss this. And there were a lot of other business reasons for that from a technology point of view. It didn't do that. And that was the next version of this. So the next two years from 2001, 2000, 2001 to 2002. two years, I spent a whole lot of time digging into the data to understand how we missed all of these indicators because it was a big impact on Cisco. We had layoffs, our stock prices never recovered from that. Even today, it hasn't recovered to the pre-internet days. So I did a lot of work on that. And we came to a conclusion that, of course, our engine was not a real engine. It was just a reporting engine. It was not an alerting engine. So we kind of thought that was important. And we went in and we needed a more dynamic and scalable model to improve this, and so that we can easily change some of these parameters. Toward that end, we chose a technology which eventually got called a Siebel Analytics. I don't know if you're aware of that term. But Siebel Analytics was one of the products that Siebel had brought in. And it's supposed to provide this automated dynamic dashboard creation, fast and easy, and it had alerting capabilities, it could multi-device. And this is right at the time when mobile phones are not the smartphones, just the mobile phones were getting more prevalent. Pages, mobile phones, dynamic content, PDAs, I think BlackBerry's and some of those things were prevalent at that time. I'm dating myself here. And we did an analysis of that. I brought that product in. we delivered on that application, and that became another huge success. So that was the genesis. So out of the background work and data validation, we had a second version. And that application actually, if I'm not mistaken, Intelligent Enterprise, if I'm not mistaken, gave that application one of the best BI application in that time. I think it was 2004, if I'm dating myself correctly. And so we got that between with Cisco and Cibo. At that time, we got that. Eventually, the stable got acquired by Oracle. And hence, my transition from data engineer at Cisco to product manager at Oracle, because I was one of those guys who understood the ins and outs of that particular product that became Oracle BI.
**Benjamin:** So, I mean, you, you told us about that kind of stack, right? In the very early days. So having like this kind of like RLQ database generating static HTML, loading that in the browser, like take us through data volumes here. Right? Like if you, if you still remember like what was, cause this was big data back in the day, right? Like kind
**Krishnan Viswanathan:** Yeah.
**Benjamin:** of what, what was big data actually then in the late nineties and early two thousands.
**Krishnan Viswanathan:** Cisco was actually on the pioneering of that. So big data in those days. So to think about Oracle applications, SMP infrastructure was, you can only scale so much. I think it was one single box, 64 cores if my memory's, 32 cores or 64 cores, I don't remember now. And it probably had about 256 GB or maybe even, memory, I have no idea. I think it might be less than that. And it was expensive machines. of a million dollars for that one single box, right?
**Benjamin:** in the like just availability of like that type of hardware. Nowadays, of course, like that's not considered a big box, right? But like, sure, like more than 20 years ago, that's absolutely insane.
**Krishnan Viswanathan:** That was insane. Yeah. And, and on top of it, it was, it was shared everything. And so our process from midnight to like 6 AM, we would take up a lot of processing problems and we would ask nobody else on it, but then there was still conflicts. And when the conflict occurs, there wasn't an easy way to troubleshoot, right? Because everything's there. So we had to start killing other applications in sometimes they would kill our application too. But.
**Benjamin:** Gotcha.
**Krishnan Viswanathan:** I was the HOV lane. I was like the ER lane, right? We would get high priority, high visibility. So that was good to have. But it was also bad because if something happens, I would be on call the next day morning with the VPs and the directors in that office just trying to figure out why I did what I did and sleep
**Benjamin:** Okay.
**Krishnan Viswanathan:** at nights to fix all that. So in terms of data volume, going back to your question, it was internal Oracle applications. There wasn't a lot. Again, we have a lot. So if you expand our bill of materials, there was a lot of transactional level details. So I'm probably thinking, I don't have the exact number, but it's probably closer to like 20, 30,000, maybe not. Maybe about 100,000 records per day.
**Benjamin:** Okay?
**Krishnan Viswanathan:** And we used to do this three times, for three different time zones, Asia Pac, EMEA, and the US. And the US was big. So about 100,000 records per day. The challenge was not the number of records. The challenge was trying to get to an accuracy and a lot of things. So people are entering data on multiple different terminals, orders come in and vendors will load in that information. Somebody fat fingers and instead of one zero, they put two zeros or they put a different number. Executives don't wanna see that and because it would throw off their numbers. One thing I realized in this whole conversation was, at the time of BI and even today it's true. people who really know the data, who really are the business, who are engaged with the business, they intuitively know that data better than us engineers. So why I say that? If there was a fat finger, if there was a number that was incorrect, I would get a call from the SVP of sales of the particular region and say, this number is wrong. I know this cannot be the truth because I have not seen this. I would have known this. So any which way, they were almost always correct. So we would have to... Our challenge was to make sure that those kind of numbers in the upstream system, impacting the upstream system, does not automatically flow through. So we had to build some triggers and metrics. So our processing and data accuracy with lack of actual tuning, all these things were just getting started in pre-internet days. How do you build a reliable data pipeline? How do you build data accuracy and at every stage of your data transformation, how do you put in... controls and alerts in place. I think that was the biggest challenge and that was what took us longer. But in the end, I am a big proponent of data controls, data governance and notification and QC for data as a separate entity. But we did a lot of work on that internally, but yeah, that was my biggest learnings in this whole process.
**Benjamin:** Okay, gotcha. It's so crazy to me because like so many of the things you're talking about from kind of back, back then in the early 2000s, right? Like they're just as relevant today. And if you talk to data engineers today, they're struggling with similar things like data quality, data observability, kind of all of those things, just of course, at a different than scale in many cases. Do you think we've come a long way? Like if you look back, or do you think the problems are actually just the same in many cases?
**Krishnan Viswanathan:** I'm still struggling doing the same things that I used to do back in 20, 25 years ago. Now in my new role at BlackRock. So no, what has happened? I think it is, it's interesting to look at the shift that happened in 2010. So before 2010, 2005, 2010, data warehousing was a big thing. Data modeling was a big thing. Bringing data in. and making sure that the data goes to a certain rigorous data modeling data by accuracy data validation was big to the extent that it actually stifled some of the data volume movement. So there was a data explosion happening on the periphery with internet and the logs, but then there was this use case about decision support systems as they were called in the past. They were going through a certain changes, they were immune to that other world. So you had certain things, like people like us who are building this by hand and taking our sweet time. So building a data warehouse would be a two year effort. And it was granted, like, hey, two years effort. If you want to make a change to a dimension or bring in a dimension, that's a six month effort. And everybody lived with it. And then suddenly this whole data explosion came along in 2009, 2010. And tangentially, MPP became a huge deal, massively parallel processing, analytical databases. So back to 2000, 2005, there was one database for all, like the one thing for all. You either choose Oracle or you choose Sybase or you choose Microsoft, they all did this OLTP extremely well. And then as an afterthought, they give you this big hunk of a hardware, single, SMP single instance hardware that you could just Take the data from your OLTP, do your ODS, and do your data warehousing, and scale. That was your data warehouse, and that was your scale. All the magic was in data modeling. MPP came along, and they broke that paradigm. And they said, you don't really need this big infrastructure. You just load the data and let it scale. And so if you are doing trending, and if you're bringing data from logs, or if you're bringing data from machine-generated, And you're doing trending data. Great. That works awesome. And so it became a huge success. So you could do things that you could never do before. You could never get clickstream applications or log analytics. Great applications, but you couldn't do that within a reasonable size and within a small budget, you couldn't do that. But what it also brought in along with that was this whole concept of data lake, which eventually turned into data swamp. for many people, but that I think we are now about 10 years after we are now starting to realize that. We are starting to realize that, no, you can't just scale your data and your infrastructure good as it may be. You also still need data models. You also still need data quality controls. You also need data governance and right. Those, and those are now, there are vendors who are now doing it at scale for cloud which in the old world mostly done by Informatica on a single node machines. And that's why I think we're seeing a lot of differences now. So to your original leading question, going back to that, I think this next few years, you will see, you're already seeing a lot of that hype. You're talking about data observability, which is pretty much this, right? Data observability goes across the board, not just on availability, but also on a lot of notification and processing time and QC, governance, data maturity. I think those will come into play. I think we are ready for that, but we are not, we don't have a good technology today in my mind that solves for all of that. I also think because of the volume of data that are potential AI ML use cases, right? Which can look at historical data. observability being one of those cases where you can actually bring in a good quality data QC or a data governance model that can give you early warning so you can build on top of it. You can actually be much more effective now than you would have been 10 years ago. So I'm very optimistic about the future.
**Benjamin:** Thank you.
**Krishnan Viswanathan:** We are trying to do some of those things in my current application, because if you see my past, I've been doing some of the data pipelines for supply chain and marketing companies. Clorox was a marketing company for the most part, marketing and finance. But now that I am in BlackRock and it's a financial company, that is again, our challenges are back to data accuracy. providing the right data at the right time with good quality controls. So
**Benjamin:** Gotcha.
**Krishnan Viswanathan:** we're seeing
**Benjamin:** Yeah.
**Krishnan Viswanathan:** that again.
**Benjamin:** I mean, kind of take, take, take me a bit through that. Right. So, right. You have to super long history kind of you've been in BI since the kind of very, very early days at Cisco. So now you're moving kind of into like this kind of more like financial space. And at least in my head as an outsider, right. Can you always like, when you think of finance, you think of like legacy systems kind of, and then very, very high bar in terms of like data quality. Right. Cause okay, like kind of like there's a lot of money at stake in this space and so on. So. take us maybe through the types of data challenges you have in that space and how it's unique.
**Krishnan Viswanathan:** I'm still trying to figure out the uniqueness of our scope. I think a lot of times I've heard about applications and legacy applications in finance and the very hard and well-structured processes and both of them are very true. And most of the financial application, at least in where I'm working, the challenges are... that we have very strict and well-built processes that have survived for multiple decades, and those code has not been touched for multiple decades. People have moved on, but the code's not been changed. We are bringing in modernization into this industry. So now I'm straddling between legacy code and newer infrastructure with data modernization. The challenge is we don't quite understand an existing legacy code which has a ton of processes, which has manual input at every stage and move it to a new technology without breaking something. So that's the challenge. I don't want to get too deep into that, but at the same time, yes, what you are saying is absolutely accurate. If we break something, the unintended consequences are very high. How do we make sure that we minimize? I don't think there will be zero unintended consequences, but how do we make sure that we can minimize those to the extent that the risk is well captured, the testing time can be extended. But that's our challenge. That's also our opportunity in terms of how we can scale because technology is not a challenge anymore when it comes to data processing, infrastructure, architecture. These are all pluggable, quickly adjustable components. I think technology has governed. and moved quite fast. It's now bringing the rest of the organization to validate and use this technology absolutely correctly to make this happen. I think that's going to be the challenge for the next generation.
**Benjamin:** But from a quality perspective, that sounds super interesting because we were talking about this data quality aspect. Did someone enter an extra zero here? Did the distribution of my data change? Should I be firing an alert? I guess here you have this second dimension to it. So you're touching legacy systems, which maybe don't have a lot of observability in place and you might not be able to retrofit that onto them. You're building something new and trying to figure out, does it? Does it do the job? Like what starts going wrong when I turn off the old thing and so on. So that's interesting.
**Krishnan Viswanathan:** Yeah, I think so. Again, I think for what we do, a lot of data comes from external third parties. So we don't control anything here. Uh, but we, and that's what I was talking about using AI and ML to, to fingerprint for like a better word, to look at how we used to get the data in the last few years. And if you see a large enough standard deviation, can we predict and at least quarantine those data so that you can actually have somebody look at it. validate it before you accept it. So those are the kind of processes that we are talking about. I think the key for that is you need to have a proper metadata. You need to have a proper governance model because yeah, I can put a process in place today to do AIML and look at my historical trend. I can come up with an area where I can see some exception but to act on that and act on it real time, I think that's a bigger piece of puzzle. Getting the right infrastructure because people who have been there more siloed in their operation. So I think it is definitely a shift in what we are doing, but I think it is exciting in the sense that it's new. And that's the technological evolution that we're going through.
**Benjamin:** Okay, super, super interesting. So how big is your team, if I may ask?
**Krishnan Viswanathan:** So we are a newly formed group called data platforms and services. And my team, we manage reference data, which is kind of a fairly large amount and it's used across the enterprise. And I have about 25 people in this intervention. And this is just engineers. And then we have product managers and data stewards and other operational analysts, operational engineers and things like that. So yeah.
**Benjamin:** Nice, super, super cool. And also cool that you have like the kind of organizational buy-in to tackle these, these types of challenges. So that's, that's awesome. Um, cool. All right. So in, in the intro, you were kind of mentioning this like data deja vu, right? Uh, kind of seeing, seeing things, uh, over, over and over again. And we talked about that data quality aspect of it, like in, in what other areas of like data engineering as a whole, are you kind of having this like data, data deja vu nowadays that you had in the past?
**Krishnan Viswanathan:** Yeah, I think so. So we talked of quite a few things, right? So what I faced in 2000, 2001 for a company financial metric and seeing that again today.
**Benjamin:** Right.
**Krishnan Viswanathan:** So I moved organizations, but I see the same thing happening here. And there's a lot of things that I am noticing that are consistent in those days here. Some of them is because we operate as a startup and so there are some challenges there. But The other thing is it's also a financial industry and there is a very strict review. So they're very conservative in how they work on this platform. Also remember from 2006 approximately to like recently, I've gone towards the vendor side and I've also worked in other companies which were more open to buying vendor products. We are back here to a point where we most of it is. in-built and in-house. So that's another shift. So I know from a vendor perspective, there is a lot of availability in terms of tools and technologies that we could easily incorporate and put in. But that's not how we build at BlackPak. We invent. We do it here. And a lot of reasons is because we eventually end up sharing it with our clients. So if I were to go and bring a data in, I don't have to only look at. how much money I spend in buying that. I also have to license it for other users. So to make it profitable for us, we have to be able to build it and scale it store-wise. So... The engineering world that I'm in recently, it is interesting because what I'm trying to do is given all the experience that I have had, can I take my team? And my team is fairly young. Average age is probably 30. And I'm bringing the age up pretty high, right? But 30, 32, maybe. So how do I take them through this when they have... One, I think part of them is mostly software engineers, not data engineers.
**Benjamin:** Right.
**Krishnan Viswanathan:** They are, I have to transition them. So I have to go back to dig into my old days of how I reacted to this, get them data savvy. So that's that, right? The whole people transformation, because these are great people. So I need to translate. And then how do I also translate that into what we can objectively achieve year over year and trend better. And I've been only here for a very short time. So it's a long process. I'm still learning a lot of things and a lot of challenges, but I believe that is a, like I said, the next three, four years is going to be massive in terms of how the whole industry changes to solve for this, but also how our company is gonna make a big difference. And I see a lot of potential there.
**Benjamin:** Okay, awesome. So another thing, right, we talked about it earlier in terms of the like, okay,.com, right? And kind of the predictions you were seeing and so on. Like now as well, right? That's kind of like at the macro level at the moment, like not a great time and especially for tech. So one thing which is coming up more and more kind of in conversations I'm having with people in this data space is this question of kind of, right, proving that all this data you're collecting is actually worth it, right? That at the end of the day, it's kind of contributing. at 2D Bottom Line of the business. What are your thoughts on that?
**Krishnan Viswanathan:** Uh... I have dealt with two sides of that coin. So remember my pre-green plum days and my post-green plum and post-green plum days. So when it came to, so the pre-green plum days, I was like, why are we collecting data that we cannot even process? What's the value of that? Why are we creating a data swamp? So I have that, that's part of my brainchild. But when I was part of green plum and I was the product manager for green plum, I had to put that on a really bad burner and not bring that up. I had to talk about all the cool things you can do. But reality is, and this is why I think as a technical and technology industry, we didn't do a good job of educating our clients. You can't just continue to collect data and not process them if you don't even know what exists. And the Hadoop days and the data breaks days, those days... are great for technology and great for infrastructure spend and glad that we had easy money at that time. But now I don't think that's going to fly. Because even at a company like Clorox, then I did a couple, I joined Clorox. I didn't talk enough about Clorox, but I joined Clorox to first upgrade, modernize that data platform. First on-prem through Oracle data warehouses and Exadata and stuff like that. And eventually I did it to migrate to. Google Cloud and Azure. So even there, because we are CPG and margins are really small, there was always this underlying question. Why do we need to collect so much data? How can we optimize our recovery process? So what I ended up doing was not only doing the data transformation, doing the data pipelines, and bringing the data in, I actually ended up building a couple of applications on top of it to showcase what data means to the company like Clorox. So this ties in very well with your question because executives are not sending any black check at that company. We use data to try and predict forecasting accuracy for products. Can I predict how much can I sell it? So we couldn't do it really good, so that was shut down. We tried using NLP and voice interaction to see if we can. predict and call out any product concerns. So that kind of worked out OK. I built a new version of the Executive Information System for Clorox. We called it the Daily Briefing, which was a mobile
**Benjamin:** Thanks
**Krishnan Viswanathan:** application,
**Benjamin:** for watching!
**Krishnan Viswanathan:** but pretty much following the same standard. And it was on a mobile phone. And that became a huge success. That was my last Clorox. And those are all built on the cloud. All of these things are built on the cloud. So I always worked in industries where we had a very narrow path between how much data we collect and what's the value of that data. Nobody gave me a blank check except in my Greenplum days. Even at Greenplum, we were telling customers, nobody gave me a blank check to go in and load as much data as you can. We did one project called the, which eventually became called the CDP, the Consumer Data Platform. And we brought in cookie information and cookie data. And I did it for about one year. And we were collecting about a billion records per day when we explored that marketing data. But in a year, we could not find any valuable metric out of that. At least the marketing team couldn't find any valuable metric, too. And we have brick and mortar. We're not really online. So it may have been better
**Krishnan Viswanathan:** online than us. But nevertheless, that was shut down fairly quickly. And it was all on-prem. We had a Hadoop cluster. I set it up. But again, that goes to show there were some people who were not in the technology world who didn't buy into this whole hype on collect as much data as you want and then run your algorithm on top of it and you'll get the results. I think that it worked. So that was always a challenge.
**Benjamin:** super, super interesting. Yeah, especially that like consumer goods, goods perspective is interesting there. Super, super cool. Awesome, cool. So yeah, do you have any kind of closing things you wanna say, right? So maybe one thing for people getting into this space, cause you said, okay, you're taking software engineers now kind of at BlackRock, right? And getting closer to like the, to this data engineering world. If I was starting today, right, kind of out of high school, kind of going to college to get into this space, like any, any advice for...
**Krishnan Viswanathan:** Again, I think when I grew up, there was no Google, there was no cloud, there was no YouTube, there was no much, I think that's what it's called, online learning, online training, call out a course at us and the, so all those are available today. So all I have to do, which makes my job a little bit easier is point my team to the right direction and tell them go and learn and get better. The biggest shift in from software engineering to data engineering in my mind is the data domain knowledge. I mostly intuitively do that because of my past experience, but understanding how data is processed on top of the code is such a big thing, which software engineers don't care. Software engineers think of data as an afterthought, whereas data engineers think of data as the first thought. So how do I scale? How do I make this consistent, secure? So... That would be my first thing, but in terms of learning, I think there are already so many tools and technologies in place. It's easier. It's a whole lot better today than it was in my days. Now, you might think of me as going back to the old day, but that's true. We didn't have the right tools and technologies and I can get the job done much faster today. In some cases, you're stuck with legacy code. Well, too bad. We've got to get that out. But I see a great potential for people who are... coming into the data world. And I wouldn't have said this earlier this year, which I'm saying now, right? The AI winter is possibly near and here, but it's gonna pass. That is so many opportunities with AI, with data, and with the whole data industry that we don't even know what the future applications are gonna look like. Future applications... are going to be driven by data. Making sure that you have the right data in the right place, processed properly, will make your company more effective, more worthwhile. I've seen this again and again and again. And like I said, I did an application in 2000 without understanding what the implications are. I did the same application at Clorox with the intention of nobody gave me the requirement. I came up with that model. I had my, I pitched that idea to my CIO,
**Benjamin:** and that's just awesome.
**Krishnan Viswanathan:** made that success. And the whole reason was I had that background in data. I had the background in business knowledge to say from an executive point of view. This is what they would look at. And of course I had help. I'm not going to say that I didn't get, but trying to fine tune that I could do that on a different technology on a different platform, because I understood the data and the value of data and how people process that data. So
**Benjamin:** Thanks for watching!
**Krishnan Viswanathan:** data itself is abstract data is difficult to comprehend. Making it easier. That's the job of a data engineer, data scientist. However you want to label it, data analyst, data scientist, data engineer. Those are all different flavors of the same thing. Make sure when you are in front of an audience and trying to pitch, you understand the data. Have a good data representation, valid, not biased, valid data representation that can help others see the picture you want them to see. And... That itself is a big hurdle to pass. And then once you have done that, the third aspect of it is all the cool technologies that are in place. You can build AI and you can build ML, but if you don't do these two things and give them what can be the truth, what are the future of the possibility, I think data will just remain unprocessed, understand.
**Benjamin:** Is this something you can kind of learn or you just learn kind of over time as you're doing it, right? Because especially if you're a kind of data team catering to different parts of an organization, you will also have so many different requirements and kind of different teams trying to do different things with the data and so on. That actually catering to your audience might be really, really hard because it's just, there is no one audience.
**Krishnan Viswanathan:** This is probably the biggest learning I have. I come out and say, like, I know all these things. I would tell you how many times I've gone into meetings. So remember, and I can remember very vividly, I was at Clorox. And this was my first day at Clorox, first few days at Clorox. And I was pretty confident because I knew data, I know what it is, and I could talk data, right? I've done some work on it. And I went into this marketing meeting, and I was... kind of referencing certain things with a different terminology than finance or marketing people would reference. And these are pure marketing guys. And they're like, what the heck are you talking about? I have no idea what you're saying. It's just like we're talking two different languages. So I had to go back. I had to spend time with the marketing team to understand, go back and understand that use cases, understand exactly what they're trying to do. So I had to bite the bullet and spend time, bad as it may sound for me now, to acknowledge. I have to relearn data. That is no easy answer. I think that was my earliest mistake. And now I'm a little bit more careful. And every once in a while, I get overconfident, and I try to think like I know everything, till I get knocked down and said, nope, you don't know. And that's always the case. Data will always have a part of it that's different than what everybody will perceive it to be. And that's the challenge. You are closest to the data. But you still don't have a grasp, I believe, and this is all coming from, this is not my opinion. Just to be clear. You'll still have a grasp of the data, but it will be contrary to many of the biases other people have built in. People have looked at, it's like the, I don't know if you know, and it might be an Indian term, but seven blind men touching an elephant. That's the story of data. Somebody's touching the trunk, somebody's touching the tail, somebody's touching the legs. That's your data. The only person who actually has a full view of the data is potentially a data guy, a data engineer, a data scientist. And how do you present it to the people who are primarily biased and biases is the blindfold and have a view of that bias and they're higher up in the organization? So ego-sacrifice. So how do you do that? And there are areas where you don't confront, and there are areas that you go in and present in open ended points of conversation and eventually bring them along. And that's going to be the challenge. I think today the world is in a place where everybody understands that data is big. Data is huge. You know, LLMs are proving that if you have large learning models, you can actually do a lot of fun stuff, but again, that are going to be skeptics. So making sure that you can wrap your use case wrap, whatever you're trying to do within a right. information providing both the views of it right so you don't want to only provide views that are validating your point of view you also want to provide areas where it says oh here are the area that i found that are contrary to my views but these are potentially having explanation of why you think they are Something that builds to your narrative, something that builds to your storyline is important. So end of the day, telling a story is critical. And data engineers need to get better at storytelling. Data scientists and data visualization engineers need to get better at understanding the data. I think it's a two-way street there. So together, I think we have a very strong team.
**Benjamin:** Awesome, perfect. I think those were great closing words, Krishnan. It was great having you on the podcast, really kind of very, very awesome. I hope you have a great day. Yeah.
**Krishnan Viswanathan:** Thank you very much, Benjamin. I appreciate your time. Thanks for giving me this opportunity. I'm excited to be sharing my information. Take care.
# Data Observability with Millions of Users - Barr Moses (/blog/data-observability-with-millions-of-users-barr-moses)
Barr Moses, CEO of Monte Carlo explains the difference between data quality and data observability, and how to make sure your data is accurate in a world where so many different teams are accessing it.
Listen on [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-data-engineering-show/id1561927688?uo=4) or [Spotify](https://open.spotify.com/show/6hMdnrFKlPbia2k6MkFs8U)
Benjamin: Welcome everyone to another episode of the Data Engineering Show. Today, we are joined by, Barr Moses, who is the Co-Founder and CEO of [Monte Carlo](https://www.montecarlodata.com/). For everyone who has not heard about Monte Carlo before, basically it's a data observability company or a data observability platform and I'm sure Barr will go into much more detail. Some maybe high-level facts about Barr, about Monte Carlo. They have recently raised $135 million series D in May of last year, 2022. Barr is the Co-Founder. Before that, kind of spent time at a bunch of different places and I'm sure she'll tell us more about that in just a minute.
Eldad, of course, thanks for joining as well, as always good to have you.
Eldad: Thanks for having me, as always. Perfect introduction. You're getting better with each episode.
Benjamin: It's amazing All the prep work and kind of practicing in front of the mirror paying off.
Barr: Can I get compliments too on my progress throughout podcast? Is that possible?
Eldad: You're amazing.
Barr Moses: Thank you so much.
Benjamin: Our sample size is one, so I'm not sure if we can. Perfect!
Barr, do you want to basically give us the high-level story of what Monte Carlo does? What data observability is all about?
Eldad: Let's start with yourself as well.
Barr: Yeah, for sure. I'm happy to. It's a big question.
My name is Barr. I started Monte Carlo three and a half years ago. Monte Carlo's mission is to help organizations actually use data by reducing what we call "data downtime." We've coined the term data downtime. In brief, data downtime is periods of time when your data is wrong, erroneous, inaccurate, or just unusable or untrustable for some reason.
That is a reality that data teams everywhere experience. So, we're fortunate to work with hundreds of customers ranging from folks Vimeo, Drata, Gusto, CNN, New York Times, Roche, and many others. All of whom have data teams who want to deliver trusted data and that's really hard to do if you're not thinking about data observability.
Take a step back and thinking about my background, I was born and raised in Israel. Started my career in the Israeli Air Force, and throughout my career, worked with data in different formats. It was actually sort of present in what I would call this acceleration of data.
I guess a decade ago, we said we were data-driven, no one was really using data. We maybe collected a little bit data and thought we were cool and moved on with our lives.
Then, maybe I want to say 3-5 years ago, people realize that you can actually make decisions based on data, and in fact, your decisions might be better if you use data. I think that was a big change for companies. I think we're still in the period of time when we're trying to figure out how to do that. We haven't figured out how to do it quite yet.
One of the most interesting trends, I think that's the backdrop to that, is what people call today data products. By that, I basically mean people using data, whether it's for internal reports, for data teams to make decisions based on data, or actually customers who are using data. Maybe that can be a dashboard that your customers use or your website. There's lots of examples for how customers actually use data and once we start putting that data out there, it'd better be accurate.
My personal story - Prior to Monte Carlo, I was at a company called Gainsight. Fortunate to be part of the category creation story at Gainsight. At Gainsight, we helped organizations basically reduce churn, and increase renewal rates and upsell. Basically, increase customer happiness based on data. I was responsible for the team that was using data, both to make decisions internally but also to surface with our customers. And the problem was that data was wrong all the time, literally just all the time.
Fast forward to today, there are some examples of where data being wrong is really painful. One of the most examples from that time, 2016, was that Netflix was down for 45 minutes because of duplicate data. Netflix was down for 45 minutes is really long time.
Eldad: I remember that moment.
Barr: Do you remember? Where were you in that moment?
Eldad: I was too young. I couldn't enter Netflix.
Benjamin: And also where are you like now?
Eldad: You mention Benjamin?
Barr: Eldad, you actually flew out of the picture at some point. I didn't know if…
Eldad: Oh, really?
Benjamin: You became invisible.
Barr: You literally flew out. You exited on the left. It was very dramatic. There you go. Okay, now you're back.
Benajamin: Welcome back.
Barr: Oh, now you flew out again.
Eldad: Okay, I'm going to fix it in a second. Give me a moment.
Barr: No problem.
Eldad: Please go on.
Barr: I thought it was by design. I was like, "oh, that's a cool feature."
Eldad: Okay, we have no electricity or equipment doesn't work. That's the reason.
Barr: Yeah, no problem. Where were you?
Benjamin: Going, right? Netflix being down for you 45 minutes.
Barr: Yeah. Everybody should remember that day. but guess what? They were down because of duplicate data. How crazy is that? In this world where actually data being down is the main culprit or even worse than applications being down and in that world, you need to make sure that your data's accurate, otherwise your applications and infrastructure are going to be down.
And that is a big change that has happened over my career in the last decade or so. I remember this was in 2016, I was responsible for a data team that was living data. And as I mentioned, the data was wrong all the time and I tried to fix it and I was like, this is so freaking hard. I remember going into a room with a whiteboard and trying to draw things and I'm not an engineer, and I'm like, this is so terrible. Why is this so freaking hard? And I remember asking our customers too, and they were like, "yeah, there's just no way to do this. We just have 6 eyes on every report." And I was like, really? We need 3 or 4 different people to read every single report. That's where we were at.
Eldad: That's a BI mindset.
Barr: Yeah, exactly. Is that normal? At some point, you don't know if you're crazy or the world is crazier or both are like, what's actually happening?
Basically, that actually inspired me and my team to try to build something. Me, and the person on my team, his name is Will Robbins. He's actually at Monte Carlo today, leading work on product and customer success. So, it's pretty cool to continue that path.
The bottom line is we hacked something together which was pretty crap, to be honest, but it worked well enough and improved like nothing that we had today. And then we implemented it with some of our customers and it worked well for them too. And I was like, Hey, why don't we get someone who's is actually an engineer to build this and let's see what happens. Can we build something?"
And then when we look at our software counterparts, they have solutions like New Relic and AppDynamics back in the day, and then Datadog now, and you have to be really irrational to build an engineering team without some observability platform or some observability measure. Why are data teams releasing data out in the wild without, something like observability or ways to make sure that data is accurate?
That basically kind of inspired starting Monte Carlo, spoke with hundreds of data teams, asking them what's keeping you up at night? And kept coming back to this problem that folks had, literally people sweating on Monday morning because they're going to be sharing a report and they're not sure if the data is right. I don't know if you have that, but you can hear your heart beating and you're like, "oh man, is this going to be right? I don't know?" Who's going to call out?
Eldad: We talked to the same people. They couldn't sleep at night. We took the data warehouse part.
Barr: Okay!
Eldad: You're saying so many things and resonate so well. Data was serving consensus and BI and it was all about having multiple versions of the truth and we couldn't get out of it. If you look at companies over the last four or five years, they have actually started to sell their data, utilize their data, build products on top of their data, and we're moving from having three people looking at a dashboard in the best case to having tens of thousands and hundreds of thousands of people, and you can't apply the same mindset anymore.
Also, you compare the Monte Carlo to Datadog and others. I don't think so. I think whenever I hear Monte Carlo, first, it's on a whole different ball game, so it's kind of less observability, more really it's all about data observability, data quality. Data quality needs to come first before it gets out to the customer. Observability goes after. So, it serves different needs. This is why I love Monte Carlo because to me, this is the first company that looks at data, looks how users use data, and rethinks observability from scratch. And it's also nice that you brought real engineer eventually to build something serious out of it, which is amazing, and we hear you all the time.
So, tell us, you quit the second and why you realize that nobody understands you when you drew that on the whiteboard, including the customers. How did that feel?
Barr: I'm not sure was that dramatic, but I think the experience of working with customers and the data being wrong all the time was terrible, really bad. I think that was for me personally. I would get emails from people, like why is the data wrong?
As a data professional, you're like, "I had one job, I was just get the data right." When you can't get that right, I think that's very frustrating. I think when talking to other people and they're like, "yeah, that's kind of like what happens?" We sort of came to accept that as sort of normal. Like yeah, it is wrong and that's okay. I remember living in that period and I was like, "there's just no way that I can accept that." And I don't think that makes sense. We have to get over that.
The days where it's fine that the data's wrong, it's just not cool anymore. Because data is used, for example, I'm reporting numbers to the street. So companies that almost reported wrong numbers to the street. That happens or companies that actually lose millions of dollars because the data is wrong. I'll give you an example. The most notable examples for the last few months. Unity is a gaming company, that basically made one mistake with their ad data and that one mistake cost them a hundred million dollars. One mistake! That's like a big freaking deal. It's not one mistake over a long period of time, multiple issues. One issue - a hundred million dollars. That's crazy! And that was made public. Think about all the issues that are not made public. They're way worse.
Then to give you an example that it's more on the social side. Equifax for folks who aren't familiar is a credit score company. So, basically a assigns credit scores to users and allows those users to take loans, take mortgages basically to live their lives. And it was made public that Equifax issued millions of wrong credit scores to users based on wrong data. That means millions of users that have the wrong credit scores today and literally…
Eldad: Inflation.
Barr: Yeah. They can blame inflation exactly or whatever it is. You think about the impact of bad data has on personal life. Even one person that's impacted in that way is terrible. Think about millions of people who have that is horrible.
I think there's like these trends, one is we're using data more and more. To your point, it's not just three people in some team looking at data once a year, maybe once a quarter, it's actually millions of people using the data, and actually, the stakes are way higher. They're not just looking at the dashboard. They're actually making decisions based on it, or they're trying to get a mortgage based on it. In all those instances, the data being wrong is just way, way worse.
I think there's just this point that we've crossed in the last few years. We can't look back. We can't unsee what we've seen with data now.
Eldad: We went all in on data and it's too late to go back.
Barr: Exactly.
Eldad: 20 years? From a product perspective, data is moving around so much, and there are so many hands being involved, apps, scripts, processes, steps, and in each step, data gets changed, gets expanded, and enriched. Where do you fit in, and if I'm a data warehouse or an engineer, can I use you as a data source to make sure that my data is kind of flow, no matter where it's coming from, I want to use you as my formal data source. Does that connect to my jdbc driver? How does it work from a product integration perspective?
Barr: Yeah, for sure. A few things, I would say, first of all, just to use this as an opportunity to explain the difference between data quality and data observability. I think a few years ago when data teams were thinking about data quality, they were really pulling a report from SAP or something where they had all the data in one place, you just had to dump the data once and then you would use it once a quarter or something like that.
In those instances, making sure that I think the concept of "garbage in garbage out" was really popularized at the time and that made it really important. I think that was the rise of data quality. During that time, it was very important for data teams or for analysts to make sure that the data's accurate, single point of time and that's it!
But the world in which we're today, which is a pretty cool world by the way, in terms of how we use data, is one in which there are so many people using data. So there are engineering teams, upstream, influencing, schema changes making adaptations or changes to code that have downstream implications. There are data engineers. There are data scientists, analytics engineers, and machine learning engineers. The list is long and each of these people actually want to use data and each of them needs to be able to use data.
Lots of things that are kind of a common trend or common thing that customers ask us is, how do I democratize data health for data quality? How do I make sure that the ability to know that data is accurate is not just for one person on the data team, but for everyone who's working with data? I think that has to do with your question around where does data live, and the thing is because there are so many people working with data, data also lives in different places. In order to make sure that your data is actually accurate and trusted, it's no longer sufficient that you're just looking at data in one place. You actually need to make sure that the data is accurate wherever it is. That may be your data lake, your data warehouse. You want to use some of your eTL orchestration solutions to help make sure that the changes that you're making there are not influencing data quality as well as your BI. At any given point in time, you have to be aware of the changes made to data and making sure that data is trusted, remains accurate and trusted.
I'll tell you a little bit about the history and evolution from my perspective. Before starting Monte Carlo, when I personally experienced this, noticed that all of my customers experiences, decided to leave Monte Carlo. I decided to start a company and I in fact started 3 companies in parallel. So, worked on two different ideas that are totally different, to kind of see what kind of pull looks.
In those early days, in order to test this idea of data being wrong, I actually reached out to lots of people and asked them does this ever happen to you? It was really hard to explain what does happen and I was looking recently at some of the wording that I used. It was like "reports breaking" or "dashboard is wrong" and that reality has not changed. Those words still hit a cord with data teams. You still get dashboards breaking. You still have reports going wrong. The thing is it's not just because of one source, it's because we now have data coming from your data warehouse, from your data lake, going through so many different transformations, it's actually hard to trace why the report broke or why the dashboard is wrong.
But that core problem actually hasn't changed, and I think we can trace that back to years ago when people are starting to use data and started to ask themselves, can I actually trust this data?
Eldad: God bless, copy and paste. The source of all evil!
Barr: Exactly. That's right. Multiple sources everywhere.
Benjamin: Awesome! Can you give some concrete examples, say I integrate with Monte Carlo today? What are the actual types of insights you can provide as a product?
Barr: Yeah, totally. I'm trying not to pitch Monte Carlo here. But basically, 'll give specific…
Eldad: We're getting great, existing and future product features from Monte Carlo. So, if there are startups out there trying to compete with Monte Carlo, now would be a time to listen carefully.
Barr: Great sound bite. Not going to happen. Just kidding!
Here's how to think about what data observability actually means and I think we touched on a couple of those already in very tactical sentences.
One is it has to have coverage end to end. By that we mean wherever your data is, it has to cover it. For some data teams that might mean, Firebolt or other competitors, not to be named, but others like AWS, Snowflake, Databricks, different types of places to store your data, aggregate, and analyze it, as well as your orchestrator. So you might have dbt as well, just use that an example and let's choose a BI. Let's say you have Looker. For many modern data teams, be a modern data stack. Actually having a solution that can cover all of those three categories is very important, data warehouse, data lake, ETL orchestration and BI. Again, because data is each of those things, and so actually having integrations that work which each of these is very important for your observability platform, whether that be Monte Carlo or not.
Second important thing is what does it actually do? Or how easy is it to get started with?
Oftentimes in data quality solutions in the past, everything was manual, so you had to manually specify the thresholds that you want for your rules, and you had to manually specify what lineage looks like. That doesn't cut it anymore. Not even close. Solutions need to be automated, and so within 24 hours, you need to have a view of your lineage, both upstream and downstream, both table and field level and you need to have a baseline for what healthy tables look like.
For example, if you have particular distribution of a field, let's say there is some specific null rate. It is possible to learn automatically what is the acceptable null rate, and let users know if that's being violated, without actually any end user specifying that as an example. Having this kind of aspect within 24 to 48 hours, getting started with those sort of machine learning out-of-the-box elements is super important.
Third thing that I would say will be called, the "5 Pillars of Data observability."
Those are:
* Freshness.
* Volume.
* Schema.
* Data Quality, and
* Lineage.
I touched on each of them, but just to kind of explain the history and why they've come together. I mentioned we spoke to hundreds of data leaders before we even started the company. By now, we've spoken to thousands of them. We've worked with thousands of data engineering users and hundreds of data teams, and actually ask them, what are the main reasons for why data goes wrong? And what does it look like when you're trying to troubleshoot it and resolve it?
In reality, there's a lot more commonality than people think. I think one of the main objections that folks have for observability is that they think that everyone is a Snowflake and your data is different. And yeah, that's true. You are a Snowflake. You're very special and your data's different, or actually, there are some patterns along with other folks, in how they use data. And so there is a certain amount of help that automation can introduce.
By codifying these 5 pillars, we are able to capture what the best data teams look for when automating observability.
The first round, "freshness," pretty straightforward, but basical is your data up to date? is easiest way to summarize that.
The second is "volume," which is basically like, is the volume of the data that you have in line with historical patterns are what you'd expect it to be. So, literally in terms of the size of the file or number of rows.
The third is "schema changes." Schema changes is a big deal for folks who know. It's also kind of tied to different trend called data contracts. We might touch on that later. Schema changes cause problems. It is the bane of existence for many data teams and so actually tracking schema changes in an automatic way is a huge help. Fun fact - It's one of the very first thing that we did in the product because it was so impactful on data teams to know what are some of the Schema changes that folks have.
The fourth is "data quality," which is basically you can think kind of from the worlds of profiling and actually making sure that the data itself at the field level is accurate. So, values that you expect. I talked about null rates, unique IDs, etc.
Then the fifth, which is "Lineage," which kind of brings it all together. Talked a little bit about lineage, but power of lineage is that when data goes wrong, the first thing that people ask themself is, where and who cares about this?
We talked about everybody hoarding data and using a lot of data. If you're dumping data and nobody cares about, no one's using, maybe you like only Eldad is looking at it at 6:00 AM, that's it! No one other than Eldad. So, maybe it doesn't matter, maybe it doesn't have to be accurate. But if Ben looking at the data or if Bard's looking at the data, then you really want to make sure the data is accurate, or maybe your customers are using it. And so being able to answer that question, Hey, is there anyone downstream who is using this data? Who should care about that? And if so, this should be a high priority. And maybe you can start thinking about actually announcing automatic severity to your data incidents and saying, Hey, there's specific conditions under which this data needs to be, there's a higher standard for it to be accurate and trusted.
Bringing it all together, I think that's sort of what makes a strong data observability platform, and we see amazing data stories of customers. I just use JetBlue for an example, for folks who've been flying a lot recently for the holidays and with the snowstorm and everything, there's a lot going on. For a company JetBlue, they manage tons of data. So, they use data both to drive their operations, whether it is like flight time or where's your luggage and what's your connecting flight? but also to manage their support and so the data team at JetBlue is actually like a great story of a team that's very, very thoughtful about making sure that data's accurate, so your flight is on time, so you get your luggage, etc.
Working with the JetBlue data team has been really cool. They, for example, manage 100% status rate incident every single week. Where literally every single week they go through each and every incident in Monte Carlo in particular and then triage that to make sure that the data quality issue is resolved and we're actually able to reduce significantly the number of data downtime incidents as a result.
Benjamin: Awesome! especially looking at data lineage, how does, especially in big companies, the adoption of a data observability tool look? Is it you need buy-in from everyone and then you need to integrate it across your entire stack because that's how you get the most value? Or is it like here's a single data team in your big company and you can integrate it with your local partner already get value?
Barr: Yeah. Every big journey starts with small steps towards that big journey. For sure I'm definitely in the camp of like, starts small and grow from there. Mostly because in any organization you want to show value, and have a win story really quickly. I think that that's true for anyone, not just observability, but I think in data in particular, if you put yourself in the shoes of the data teams in large organizations, they're in a tricky spot and I'll explain why. They invested a ton in the last few years in the best data infrastructure. World-class data warehouse, data lake, world-class infrastructure and also they invested a ton in hiring lots of people to use that data. We've seen the rise of data scientists. All of those roles have been growing a ton in the last few years.
Here's a tricky spot! Now, you need to deliver. You need to show that you can use the data, and that's new reality, which I don't think many organizations have proved. You're on the hook to see the ROI of all that investment. Oftentimes, the ROI doesn't exist because people don't trust the data or can't use it.
I would say in order to avoid being in that situation, starting to think early as you're rolling out your data infrastructure, how do I actually make sure that people also trust the data is really important. And for larger organizations, yes, I do think that starts with identifying a particular use case or a particular team. For example, it can be a team that's supporting financial data in particular, that's very important. If you're reporting on revenue growth or customer growth, you want to make sure that there's no questions about that data. Maybe starting with that. Sometimes there's a marketing team that you might want to start with. If they have hundreds of millions or tens of millions dollars in budget, you want to make sure that that money is deployed, that resources are deployed.
Another kind of examples of companies in healthcare, for example, potentially you have production data and you want to make sure that clinical data is accurate. In all of those instances, starting small, making sure that folks have that technology, but also have the mindset. Think about it, going back Eldad's point from before, maybe the difference is that for an organization like Datadog, the concept of observability and engineering is something that lots of folks have done before, but for data teams, it's a new motion. And so understanding who's responsible for what and what you actually do is a totally new ballgame for most people.
So, actually most of the work that we do is helping organizations think through what is a great data observability practice looks like? And oftentimes that has nothing to do with the technology. It's more around the practices and the types of culture that we have in the data team.
Eldad: Love it.
Benjamin: Definitely! Take us through the action part. I'm a data engineer now I get kind of my alert in my Slack channel. I don't know how it works. Kind of this upstream data pipeline is looking weird right now. What actually happens? Is it about spotting failure early to make sure you can kind of keep surface area small? What's usually the immediate action you take when you see these types of alerts?
Barr: Yeah, great question.
Yes, oftentimes, teams actually love getting that information to them directly and so that might be in Slack or Teams, or even via email if you'd like. Large majority, I think use Slack today. You get an alert in Slack, and that could be something, "Hey, this table that gets updated typically every hour, stopped updating for the last day and hasn't gotten an update for the last, let's say seven hours."
There are a couple of things that you can do. First of all, is an automatically assigned severity, maybe you see that it's a high sever severity might be a situation where you drop everything and take a look at it immediate. On the contrary, maybe this is a data set that you know is not being used and not looked at, then maybe you can ignore it for now.
Maybe you can snooze it actually and snooze it until tomorrow because today you're working on a production issue and you need to come back to it later. So assuming you didn't snooze it, you determine that this is actually high severity based on the assigned severity, and the impact of assets that it looks at. What you can take a look at is look at impact radius and start seeing who are the folks who are impacted by this type of issue.
For this particular table, there are three reports in your BI that are actually using this and the particular team that using that as a marketing team, maybe what you could do is you could tag the person on that team and say, "Hey there's an issue here, FYI, I'm investigating."
Then what you could start doing is click into that alert and start investigating and say, "what else is happening around that time? What else is happening to that table that might give me clues as to why this issue is happening?" And then maybe you can see, okay, well the job is running, but actually, no data is arriving, so maybe there's a problem with the low that's feeding data, just as an example, or maybe someone upstream, you noticed made a scheme of change and that scheme of change made it such that the implications are that now data's not arriving anymore. Maybe they change the field type and that messed up the scheduling of the data into that particular table.
On the other hand, maybe when you look at the downstream implications, you recognize that you need to reload all the data to fix that and to make the report accurate now. There's actually a lot of work that might be, involved in both understanding why this issue has happened and also making sure that the data is sort of back to normal or back to its schedule.
It's a combination of looking at both the data itself, looking at metadata and oftentimes looking at the code that's driving all of this to make sure that the data pipeline itself is healthy. So, you might be moving between those different systems and where you're at, and Monte Carlo in particular, we try to bring all of those different feeds into one place. So you actually get both your alerts from Monte Carlo as well as your alerts from dbt in one place, for example. If you're using dbt in this example, you might be not notified about that same issue in the same place.
Bringing all of that together, maybe upon investigation, you realized you understand the root cause of the problem. You understand how to resolve it and now the problem is that you're not responsible for the table itself. So, you can't actually fix it. So, you need to find who is the person who owns this. You can see that information as well and then ping them and work with that person to fix the issue.
At a high level, there are three different things that we're trying to help users do.
One is to know about the problems almost in real time. Just to give you insight into this, most teams are not even aware of these issues and so getting that alert in Slack is a big deal.
The second thing that we help with is resolving these faster. Oftentimes, data teams won't know all this information about, like schema changes or freshness or volume, all the examples that I mentioned, and it would be harder for them to pinpoint what exactly is happening and why?
The third thing is by increasing the communication collaboration on this, you're making it easier for everyone who's touching data to know about data downtime issues and eventually actually preventing these from happening.
Eldad: Would you say that in a few years from now, we will use less classic black box mindset observability plus Jira and more of new context-driven, lineage-driven observability, which listening to you, it's so different. It's so, so different from looking at CPU or disc utilization while someone altered the table but that table affects a hundred Looker dashboard and that's not important. What's important is one of those dashboards serves a customer. There's no way to know that without spending three weeks involving six people, at least using tons of tech and companies try to do it on their own. They try to stitch it. They try to build it, and it's really sad.
But listening to you makes me happy because it is a real thing, and as engineering organizations are trying to become data organizations, I think they have a lot to learn from that. From data engineering mindset where it's taking very seriously. So when someone is modifying something for those scheme, that's like code, that's like a product release. And having data observability that has the schema, that has the data warehouse, it's aware, it knows what an information schema is. It's not like just raw JSON pushed into some time series database, which is meaningless for most people. I love it! It's fascinating!
Barr: Yeah, I think that's spot on. I think as a data industry, if we're successful, it's because we adopted what we need from engineering practices and built that into data engineering. It's going to be impossible for us to actually really, truly become data driven if we don't do that. So, yes, and I think our approach is that starting with the observability is the right place because it's the most imminent problem for engineers.
If you look at data engineering today, the most imminent problem that they have is, "hey, the data is wrong all the time." That's the number one thing that they have as an issue and as a result, data consumers can't trust the data and can't use the data.
I think there's a lot more that we can learn from engineering and that we should over time bring in a lot of those concepts into data, but starting with observability is what we think is the thing that will bring the most value to customers today.
Eldad: Nice.
Barr: And I'm curious to hear more about how you all think about it at Firebolt, but so feel free to let me know if you all are upright.
Eldad: We feel the pain and need to make sure that data quality is right. We've seen some really, scary stories. You've mentioned some of them, but data is so actionable today that customers use data-driven products, their business run on those products. If I'm subscribing to a sales optimization product and that sales optimization product has given me a recommendation how to do my ad spending, and I'm going and just doing this ad spend based on that recommendation, and it's wrong. That affects my business, and it it's real. So, it's much more serious than just looking at two different types of sales numbers, one coming from Salesforce and one from Excel, which is also important. We've been dealing with that for 20 years. But the beauty is that data is really driving the business now versus just being used to understand the business.
So, having data quality should be kind of trivial in core stack of any data driven team. They are listening to you and looking at kind of what you're doing there. We really want and wish Monte Carlo to succeed big time. And having said that maybe we can wrap with that.
If you can share some insights, some recommendations, something with our fellow startups, with our fellow users, how do you see this year coming? What would you recommend for startups going forward? Anything valuable would be highly appreciated.
Barr: I can share only non-valuable advice.
Eldad: Perfect!
Barr: Actually, on that note, I will say, for startups or listeners, I will start by saying that you should not listen to advice. That's my biggest takeaway. I think for any question that you have, both in data in startups and otherwise, you will always get, 50% of people will tell you something and 50% of people will tell you the otherwise, the opposite. So, really I think the answer actually lies in the data. There's no one else but you to take a look at the data and listen to your customers and see what it tells you and actually draw conclusions based on that. Folks who are asking themselves like what does this mean for me and for my team and for the industry, you have the data, you are closest to the customer.
Pick up the phone, talk to the customer, ask them what they think. That is the data that will help you get to the answer, whatever that may be. That's my take on advice.
With that in mind, looking ahead to this year, I think this is the year that's sort of a natural progression of the importance of data teams. I think obviously, like it's no secret, that the market has changed difficult times for lots of folks across the economy and yet I continue to see data teams growing stronger and stronger in organizations. And that means companies doubling down on their data strategy, building more and more data products, and investing in data teams because recognizing that data is a foundation to making strong business decisions, gaining competitive advantages, and generally furthering business outcomes for their customers. And so I think when we zoom out and think about where are data teams today? I think they're on the hook to deliver ROI from all of their investments in the last few years and this is the time when folks are going to be more scrutinized. Yes, I do think that folks will be asking themselves, is there actually ROI on this? And it's a great time for data teams to be able to articulate that and say, yes, we have awesome data, we have best in class infrastructure, and we can also trust the data so we can actually use it to your point, to drive the business. And I think that's an inflection point that I'm really excited for data teams to bring. I hope 2023 will be the year that we see that. I don't think that we can see it soon enough, but I do think it's an inflection point that the data industry needs to drive. And I think there's no better time than this year to do this because data teams are more important than ever. Data is more important than ever and time, the onus on us to prove that is today. So, I'm excited to see us do that.
Eldad: Boom!
Benjamin: Awesome! That were amazing closing remarks. Barr, thank you so much for your time. We really appreciate it. Have a great rest of the day and see you next time on the Data Engineering Show.
Barr: Thanks, great podcast
Eldad: Thank you everyone, looking forward.
Benjamin: Thanks.
# Data Rewind: Conversation Highlights from Zach, Matthew, Joe, and Krishnan (/blog/data-rewind-conversation-highlights-from-zach-matthew-joe-and-krishnan)
In this special roundup episode of *The Data Engineering Show*, the Bros revisits some of the best bits from episodes with data thought leaders Zach Wilson, Matthew Housley, Joe Reis, and Krishnan Viswanathan, spotlighting essential trends and lessons learned across the evolving data engineering landscape.
Listen on [Spotify](https://spoti.fi/40qWYLa) or [Apple Podcasts](https://apple.co/4f2I9CX)
The Data Engineering Show is brought to you by Firebolt, the cloud data warehouse for low-latency analytics. Get $200 credits and start your free trial at firebolt.io.
Benjamin - 00:00:10: So on my end, I'm just kind of like C++ database nerd. I care about how to build concurrent index structures. I don't know, like how to build a fast multi-threaded join and so on. And whenever I look at data engineering, right, and kind of like I interface with a lot of data engineers, Firebolt kind of always seems kind of really daunting just because like the kind of breadth of the field in terms of like just amount of different technologies you have just seems so crazy, right? So like how do you decide on kind of what to like focus on in terms of actually kind of, yeah, like teaching people kind of skills that are valuable for their career?
Zach - 00:00:52: Like for my bootcamp, for example, right, there's six weeks. Only two of those weeks are actually tech specific. The other four weeks are not tech specific. They're tech agnostic. So where, because I think there's a couple things and a couple philosophies that are really important in data engineering that are actually like, they apply regardless of if you're using Spark or Snowflake or Databricks or Presto or Flink or like whatever, you know, tech you want to use for it. And there's a couple of them. Like one is like around like data modeling, how to do proper data modeling for dimensions and facts and how to really get those things like compacted down. And there's a lot of trade-offs in that space that is very art, very, it's not as science and it's a lot more art and you have to understand. Like your consumers. Yeah. And they have to have that empathy. And that part is powerful. Like for example, for me, like when I was working at Airbnb, like I would say 80 to 90% of the impact I had was going to be in two buckets, right? It was in the leadership bucket of inspiring other people and helping them grow. And the other one is data modeling and getting, making robust data models that can then be used by a large number of people downstream. What wasn't as important was like, how good I was at Spark. Like, I mean, I look at Spark as more of like a means of accomplishing something or like as like a, it's just one kind of path forward and that like you just, it solves the problem and maybe it can be a little bit faster. Maybe it can solve those things. But really, this is the fundamental thing that I think a lot of data engineers need to remember is your product that you sell is data. It's not a pipeline. A pipeline, it can help, right? In term of maintenance and like pain and suffering. Like if your pipeline sucks, then like the data is going to be annoying. But like generally speaking, the value you're providing is in the data sets that you provide. And that is like, and if those data sets are not modeled properly, then that's where you can have a lot of unnecessary cost, right? And like, in big tech companies, these mistakes actually cost them millions and millions and millions of dollars a year. Because of like, if you don't model things the right way, then downstream, the compression doesn't work the same way. And then it can blow the data up again, right? And there's a lot of interesting, like tricky things that I've noticed with like how data modeling works. So that's one. I'd say another kind of tech agnostic thing is around data quality, right? And understanding, again, there's technologies here like Amazon DQ and great expectations. And there's going to be 10 trillion more coming. And but like, it's more again around like how to test data like of like, is this quality? Is it not quality? How to validate it, right? That's very like agnostic of like the tech that you're using. And you should definitely be able to do that. And then the last bucket of things that are tech agnostic is storytelling, right? You tell a compelling story, can you like make some cool charts? Can you persuade people to give you the time to make this data? And other things like that, because there's like the story, there's like the before data story. And then there's also the after you have data story. And both of those stories matter. And being able to construct those narratives in a compelling way. Very important persuasion. And then the tech is the last one. And like, and like, in some ways, I think the last one, but also not as important and- But it's tricky because and this is the thing I hate about industry in some regards is that like, If you go into an interview, right? And you go into the interview, like 80% of the questions are going to be on like Spark or Flink and be like, oh, do you know this very specific minor detail about Spark? And it's like, dude, this doesn't matter that much actually in the end. But that's how things are tested, right? And I hope that industry changes in that way.
Benjamin - 00:04:49: Gotcha. So, I mean, looking at this bootcamp, because you're framing it in that context, kind of you said in the beginning, it was like the goal is kind of from good to great. So usually like those will be people who already know Spark, right? Kind of know how to maybe like write Scala code for their Spark stuff, like those types of things. On the other end, and it's hard to say from good to great, is like, okay, you have someone, right? Maybe that, okay, not the influencer with 25,000 followers talking about data engineering, but just someone wrapping up high school who wants to get into data engineering. And if you have to pick up some technology, right? And like there, it just seems kind of like daunting to in a sense, like make that choice, right? Kind of what horses do you bet on? Kind of with what do you get started to actually start with that career?
Zach - 00:05:33: I mean, I totally agree. I think it's similar to, like, so I'm a pretty like athletic, sporty guy, right? And like one of the things that I remember as a kid growing up, my parents were always like, you got to do sports, right? And then I was like, okay. And like, then I tried a bunch of, I tried like soccer, I tried basketball, I tried baseball, I tried like all these different sports that like were all interesting and different. And like, one of the things that I learned about it was, especially going through that process was like, yeah, you just got to pick one and be like, I'm going to get good at this one. And for me, that was basketball. And I mean, I got lucky, I'm tall or I'm 60. So basketball was the easy, obvious choice. And that's one of the things that's tricky about tech sometimes is that the choices aren't so obvious, right? They're not like, oh yeah, this one is 6'7'' and this one's 4'2". So we should go with a taller one or whatever, right? It's not that like obvious a lot of the time. And so I think there's kind of a couple pieces there, like on like how to pick like technologies is one is going to be like, okay, what do you see on social media? Like, I know that that's like, I'm not going to say all social media because I still don't trust TikTok here. But on LinkedIn, like if you have enough of a network on LinkedIn, do a poll. Polls are broken on LinkedIn. Like if you do a poll on LinkedIn, like even if you have no followers, it's going to be seen by like 10,000 people because polls are broken and they're very good. And the reach they get is too good. And so you can learn, you can ask, right? And I found that there's like two or three really high fidelity sources of like, where to get like good information. You have LinkedIn. LinkedIn doesn't give you... The thing about LinkedIn though, is it doesn't give you very... It's not as good about negative feedback. Like... So if you're like asking someone to be like, hey, say why my video sucks. Most people aren't going to do that in the comments section on LinkedIn because like, they're like, I don't want to look like an asshole. And so that's one. Reddit. Reddit's better for that. Reddit's almost too good for that. If you want people to tear you down, go to Reddit. Like Reddit or Blinds even better if you really want to get... But those places, you can also get the more like guidance on like, okay, should I learn these texts or these texts, right? And like, I think like really the big things are going to be just like learning the languages first. SQL, Python, right? If you really want to break into DE, SQL, Python, learn the languages first and then you can learn the tech after that actually. Because you can do... If you just do like Postgres to learn SQL, just like a very basic database, then you can build into the more complicated ones like Snowflake or Firebolt or Spark or whatever you want to use, right? And then like, you can kind of go into the cloud. And like, I found that kind of building more locally first and like kind of learning languages that way. I've had more success with some students kind of teaching them that way as opposed to like being like, okay, here's how you set up an EC2 instance on AWS and now you have a computer in the cloud and it's going to crunch the data for you. And like a lot of that feels like a lot more complexity that they don't really need yet until they have more confidence in their own skills.
Benjamin - 00:08:45: Looking at SQL databases, I guess kind of one of the good things is looking at the space right now. It's like so many, like so many systems are actually kind of converging around the Postgres dialect. There's not like you need to kind of learn like seven different kind of flavors of like SQL or like window function syntaxes, whatever. Actually, like a lot of systems tend to be at least more similar to date than maybe they used to some time ago. So that's super cool. So thanks for the insights there. I mean, like I'm sure a lot of listeners will appreciate that. Cool. Zooming out a bit, right, from the kind of specific technology, like, do you see any kind of big kind of trends at the moment? Like, when I look at LinkedIn, for example, like, one thing that keeps coming up is, like, data observability, kind of data monitoring, data quality. Like, those seem to be what some of the things kind of generating a lot of buzz. Kind of, yeah, what else is out there?
Zach - 00:09:39: Yeah, like, data monitoring, MLOps, data versioning. There's all sorts of, like, interesting things. And then there's, like, a couple of them that come back around sometimes, like, data mesh. Like, I hear data mesh, like, once every three months or something like that. And I'm like, hey, it's there. It's a thing, right? But, like, I think a couple of things that I really am seeing, yeah, is definitely data observability of, like, yo, like, how is this data changing over time? And, like, it's very closely linked with data quality. And honestly, they should be closely linked because if you aren't aware of, like... How your data, the shape of your data, what it looks like over time, then you really don't have good data quality checks because you haven't done your due diligence on looking at what is normal and what is abnormal. I mean, there's data quality checks out there that are very easy to know or what is normal and not normal. Is there any data? No data is abnormal, right? Or this column's null when it should never be null. That's abnormal. Very easy check. But then things, what you define as normal versus abnormal, it gets more and more complicated as you look at more and more different data points together. And that's where if you look at week-over-week row counts, that's going to be one that, what is abnormal versus normal? It depends. Because a lot of times those week-over-week row counts, on Christmas Day, they fail because there's not as much data or there's too much data. And it's actually not... You're looking at the wrong period. Instead of week-over-week, you really should be looking year-over-year and looking at it on zooming out to find the actual pattern that matters the most. And that's why people do week-over-week instead of day-over-day because of the Sunday-Monday phenomena. And Sunday-Monday is super annoying as well. That one's very common to trip people up. That's why week-over-week is better. But it still misses the holiday patterns. So I think that those kind of observability, things are super important because it's linked to quality and that is linked to trust. And because it's like without quality, you don't have trust, right? And definitely, I think that that, I would say, is the big thing that I've definitely been seeing. I've also been seeing a little bit more of a push towards streaming and trying to get more people involved. I've been hearing about ClickHouse so much recently. Everyone's like, you got to try ClickHouse. You got to try ClickHouse. I have not tried ClickHouse yet, but I need to. Just because it's been... I've seen it in every single comment section of all my posts. So yeah.
Benjamin - 00:12:19: One thing I'm curious about in general, right, kind of like coming out of this is we're saying, hey, our tools got much better, right? But we still have many of the same problems we used to have. We still keep cycling around using kind of topics. And both you, Joe and Matt, kind of, right, you're teaching a lot. Kind of you're doing thought leadership. Kind of you're writing blogs. You kind of wrote that super well-known book. You're affiliated with the University of Utah. You're consulting. Kind of like arguably you could say, okay, if we have all of those amazing tools now and we're still cycling around the same kind of types of problems, right? Maybe we're just not teaching it well enough. So what does that mean for your approach to kind of delivering these things to A, professionals, students, those types of things?
Rob - 00:12:58: We failed.
Joe - 00:13:00: The kid comes out swinging.
Rob - 00:13:07: That's a good question.
Joe - 00:13:08: I mean, I think part of the problem, and this is not to trash vendors too much, I think vendors build fantastic products.
Rob - 00:13:14: Horrible products.
Joe - 00:13:14: Yeah, yeah, yeah. But I mean, if I'm in sales for a vendor, I'm not necessarily focused on how I use the tool. I just want to get the tool out there and get people using it, right? And that's where there is more need for people on kind of the meta level to come in and say, all right, you've decided on X, Y, and Z tools. How can we actually use these to help the company? And do that training all along. I mean, I think Joe and I have complained a lot about the lack of training for undergraduates and data specifically. And part of that training as we build it out needs to be, obviously, they need to learn data fundamentals like data modeling, thin ops, cost management, but also what it's like to work inside a business and the kinds of things that businesses care about and how they can communicate better. I mean, communications are notoriously difficult to teach, right? Because how do you teach someone out of a textbook how to communicate with someone? But we need to keep thinking about these problems and figure out how to give students practical, concrete experience with communicating with businesses and stakeholders.
Benjamin - 00:14:16: So how do you do that? Because that was also a very abstract answer.
Joe - 00:14:20: Fair, fair. I mean, I think from our perspective, a lot of this comes down to building better collaboration between undergrad and master's programs and businesses. Shockingly, often we see that you've sort of got this MBA world, that operates almost in a vacuum separate from the business world. And that's not ideal, right?
Rob - 00:14:40: Well, the academic world operates separate from the business world too.
Joe - 00:14:42: Yeah, absolutely.
Rob - 00:14:43: In some cases, that's good. In a lot of cases, I think it's pretty bad. It does a disservice to students. So that's one thing I'd like to see change, right? So you talk about concrete stuff. I would also like to see more apprenticeship-type programs. I think the notion of a university being a necessity, I think is absolutely the wrong way to go. So I think more people could be trained on this from practical things apprenticeships... I'm creating a new MOOC class right now, a course for a really big MOOC. One of the things I'm doing is... It's a simulator. It's your first day on the job as a data engineer. You get to go do business requirement gathering. You get to go find out what stakeholders want. And part of it is identifying, okay, so you're given this list of requirements. What are people actually asking for? So I think that's the other part. We spent too much time teaching tools and not enough time teaching the techniques. So I think those are concrete ways that we could address it. Because it's easy to do PySpark tutorials and stuff. But that's the wrong way to teach data. I think the way we teach it is absolutely backwards. Know the techniques and then learn the tools. That's why we wrote the book the way we did it. It's technology agnostic, for example. And pretty much every company in the universe is using it for their data teams right now. Almost every university that we know is increasingly being used as a default textbook for data engineering. So to me, that's part of the process. But it's not going to be an overnight thing. But I think the way we approached our book is similar to how Martin Kleppmann approached his book. It's agnostic. It stands the test of time. And that's kind of where we need to get to. We are making an effort. It is slow, especially universities are slow. They're so slow. And that's part of the problem with them.
Benjamin - 00:16:20: In the intro, you were kind of mentioning this like data deja vu, right? Kind of seeing things over and over again. And we talked about that data quality aspect of it. Like in what other areas of like data engineering as a whole are you kind of having this like data deja vu nowadays that you had in the past?
Krishnan - 00:16:39: So we talked of quite a few things, right? What I faced in 2000, 2001 for a company financial metrics and seeing that again today. So I moved organizations, but I see the same thing happening here. And there's a lot of... Thing that I'm noticing that are consistent in those days here. Some of them is because we operate as a startup. And so there are some challenges there. But the other thing is, it's also a financial industry. And there is a very strict review. So they are very conservative in how they work on this platform. Also remember, from 2006 approximately to like... Recently, I've gone towards the vendor side and I've also worked in other companies which were more open to buying vendor products. We are back here to a point where most of it is in-built and in-house. So that's another shift. So I know from a vendor perspective, there are a lot of... Availability in terms of tools and technologies that we could easily incorporate and put in, but that's not how we built at Blackhawk. We invent, we do it here. And a lot of reasons is because we eventually end up spending, and sharing it with our clients. So if I were to go and bring a data in, I don't have to only look at how much money is spent in buying that. I also have to license it for other users. So to make it profitable for us, we have to be able to build it and scale it to our clients. So the engineering world that I'm in, recently. It is interesting because what I'm trying to do is... Given all the experience that I have had, can I take my team? And my team is fairly young. Average age is probably 30. I'm bringing the age up pretty high, right? But 30, 32. How do I take them through this? When they have one, I think part of them is mostly software engineers, not data engineers. I have to transition them. So I have to go back to dig into my old days of how I reacted to this. Get them data savvy. So that's that, right? The whole people transformation, because these are great people. So I need to translate. And then how do I also translate that into what we can objectively achieve year over year and turn better? And I've been only here for a very short time. So it's a long process. I'm still learning a lot of things and a lot of challenges, but I believe that is, like I said, the next three, four years are going to be massive in terms of how the whole industry changes to solve for this, but also how our companies are going to make a big difference. And I see a lot of potential there.
Benjamin - 00:19:26: Another thing, right? We talked about it earlier in terms of the dot com, right? And kind of the predictions you were seeing and so on. Like now as well, right? That's kind of like at the macro level at the moment, like not a great time and especially for tech. So one thing which is coming up more and more kind of in conversations I'm having with people in this data space is this question of kind of proving that you're... All this data you're collecting is actually worth, right? That at the end of the day, it's kind of contributing to the bottom line of the business. What are your thoughts on that?
Krishnan - 00:19:56: I have... I have... Dealt with two sides of that coin. So remember my pre-greencrumb date and my post-greencrumb date. So when it came to, so the pre-Green Plum days, I was like, why are we collecting data that we cannot even process? What's the value of that? Why are we creating a data swamp, right? So I have that. That's part of my branch, right? But when I was part of Green Plum and I was the product manager for Green Plum, I had to put that on a really back burner and not bring that up. I had to talk about all the cool things you can do. But reality is, and this is why I think as a technical and technology industry, we didn't do a good job of educating our clients. You can't just continue to collect data and not process them if you don't even know what exists. The hard-op days and the data breaks days, those days are great for technology and great for infrastructure spend. I'm glad that we had easy money at that time. But now I don't think that's going to fly. Because even at a company like Clorox, then I did a couple, I joined Clorox. I didn't talk enough about Clorox, but I joined Clorox to... To first upgrade that, modernize that data platform. First on-prem to Oracle Web Data Warehouses and Exadata and stuff like that. And eventually I did it to migrate to Google Cloud and Azure. So even there, because we are CPG and margins are really small, there was always this underlying question. Why do we need to collect so much data? How can we optimize our data process? So what I ended up doing was not only doing the data transformation, doing the data inference and bringing the data in, I actually ended up building a couple of applications on top of it to showcase what data means to the company like Dropbox. So this ties in very well with your question because executives are not sending any black check at that company. We use data to try and predict forecasting accuracy for products. Can I predict how much can I sell it? So we couldn't do it really good. So that's what we did. We tried using NLP and voice interaction to see if we can. Predict and call out any product concerns. So that kind of worked out okay. I built a new version of executive information system for Clorox. We called it the daily briefing, which was a mobile application, but pretty much following the same standard. And it was on a mobile phone. That became a huge success. That was like my last run. And those are all built on the cloud, right? All of these things are built on the cloud. So I always worked in industries where we had a very narrow path between how much data we collect and what's the value of that data. Nobody gave me a blank check, except in the green firm days. Even in the green firm, we were telling customers. Nobody gave me a blank check to go in and load as much data as you can. We did one project which eventually became called the CDP, the Consumer Data Platform. And we brought in cookie information and cookie data. And I did it for about... One year, and we were collecting about a billion records per day, right, when you explore that marketing data. But in a year, we could not find any valuable metric out of that. At least the marketing team couldn't find any valuable metric. And we are brick and mortar. We're not weak. So it may have been better for online than us, but nevertheless, that was shut down fairly quickly. And it was all on-prem. We had a Hadoop cluster. I set it up. That goes to show. There were some people who were not in the technology world who didn't buy into this whole hype on collect as much data as you want and then run your algorithms on top of it and just get the results. I think that did work.
Benjamin - 00:23:49: So, yeah. Do you have any kind of closing things you want to say? So maybe one thing for people getting into this space, because you said, okay, you're taking software engineers now kind of at BlackRock and getting closer to this data engineering world. If I was starting today, out of high school, going to college to get into this space, any advice from your end?
Krishnan - 00:24:14: I think when I grew up, there was no Google. There was no cloud. There was no YouTube. There was no, I think that's what it's called, online learning, online training, course errors. So all those are available today. So all I have to do, which makes my job a little bit easier, is point my team in the right direction and have them go and learn and get better. The biggest shift from software engineering to data engineering, in my mind, is the data domain knowledge. I possibly intuitively do that because of my past experience, but understanding how data is processed on top of the code is such a big thing, which software engineers don't care. Software engineers think of data as an afterthought. Whereas data engineers think of data as this first step, right? So how do I scale? How do I make this consistent, secure? That would be my first thing. It talks of the... Learning, I think there are already so many tools and technologies in place. It's easier. It's a whole lot better today than it was in my days. You might think of me as going back to the old days, but that's true. We didn't have the right tools and technologies. I can get the job done much faster today. In some cases, you're stuck with legacy code. That's too bad. We've got to get that out. I see a great potential for people who are coming into the data world. And I wouldn't have said this. Earlier this year, which I'm saying now, right? The AI winter is possibly near and here, but it's going to pass that is so many. Opportunities with AI, with data, and with the whole data industry that We don't even know what the Future applications are going to look like. Future applications. Are going to be driven by data. Making sure that you have the right data, the right place, processed properly, will make your company more effective, more worthwhile. I've seen this again and again and again. And. Like I said, I did an application in 2000 without understanding what the implications are. I did the same application at Cora. With the intention of, nobody gave me the requirement. I came up with that model. I pitched that idea to my CIO, made that success. And the whole reason was I had that background in data. I had the background in business knowledge to say, from an executive point of view, this is what they would look at. And of course, I had help. I'm not going to say that I didn't get. Trying to fine tune that. I could do that on a different technology, on a different platform, because I understood the data and the value of data and how people process that data. So data itself is abstract. Data is difficult to comprehend. Making it easier, that's the job of a data engineer, data scientist. However you want to label it, data analyst, data scientist, data engineer. Those are all different flavors of the same thing. Make sure when you are in front of an audience and trying to pitch, you understand the data. Have a good data representation, valid, not biased, valid data representation that can help others see the picture you want them to see. That itself is a big hurdle to pass. And then once you have done that, the third aspect of it is... All the cool technologies that are in place. You can build AI and you can build ML, but if you don't do these two things and give them what can be the truth, the future of the possibility, I think data will just remain unprocessed.
Outro - 00:27:52: The Data Engineering Show is brought to you by Firebolt, the cloud data warehouse for low latency analytics. Get $200 credits and start your free trial at firebolt.io.
# Decomposing Firebolt transactions (/blog/decomposing-firebolt-transactions)
This week, Alex Miller wrote a blog post titled [Decomposing Transaction Systems](https://transactional.blog/blog/2025-decomposing-transactional-systems). He approached explaining transactions from a new fresh angle, but one which immediately felt intuitive. Alex decomposed transaction into four steps:
Every transactional system does four things:
* It executes transactions.
* It orders transactions.
* It validates transactions.
* It persists transactions.
Alex gives examples how these steps map to both generic OCC and PCC systems, as well as examples of real-world systems like FoundationDB and Spanner.
Then, Mark Brooker, followed up with a post of his own about how these steps map to [Aurora DSQL](https://brooker.co.za/blog/2025/04/17/decomposing.html).
So I naturally thought to map them to Firebolt. We haven't published in-depth technical articles about Firebolt transaction manager, but there is [10 minute video](https://www.youtube.com/watch?v=i6yz3Xm3eOg), where we describe how it works to CMU's Database class.
So let's go step by step
> *Executing* a transaction means evaluating the body of the transaction to produce the intended reads and writes.
Firebolt execution happens in parallel across the nodes of Firebolt engine cluster. There is no coordination between nodes or between clusters at that stage. We use MVCC, so all the writes create new version of data which is invisible to other transactions.
> *Validating* a transaction means enforcing concurrency control, or more rarely, domain-specific semantics.
Firebolt employs OCC and therefore validation happens after execution, when transaction is ready to commit. Default isolation level in Firebolt is [Snapshot Isolation](https://jepsen.io/consistency/models/snapshot-isolation), which means that transaction manager needs to check for Write/Write conflicts, i.e. to check if the current transaction (the one about to commit) did not write to the same rows as transactions which were active when the current transaction started, but committed before. If such conflicts are detected, the current transaction is aborted. Note, that validation step happens independently without coordination with other transactions trying to commit.
> *Ordering* a transaction means assigning the transaction some notion of a time at which it occurred.
> *Persisting* a transaction makes making it durable, generally to disk.
This is where Firebolt differs from classic OCC sequence of Execute -> Order -> Validate -> Persist. In Firebolt, Order and Validate are swapped, and Order happens together with Persist. After Validation step, Firebolt's transaction manager will atomically pick the Log Sequence Number (LSN) to correspond to current transaction's commit, and persist it in Write Ahead Log (WAL). But it also checks whether some other transaction managed to commit since validation was completed, because if it did - the validation step needs to be repeated. Think about it as [Compare and Swap](https://en.wikipedia.org/wiki/Compare-and-swap) operation - transaction manager atomically assigns and writes new highest LSN, but only if it didn't changed since the validation step.
Firebolt uses [FoundationDB](https://www.foundationdb.org/) to store all of its metadata including the WAL. So the Order + Persist step is implemented as FoundationDB transaction to ensure atomicity. But also, once persisted, the WAL is automatically replicated just like any write into FoundationDB would.
In addition to that replication which is there to ensure durability, Firebolt also replicates WAL's content into Kafka topics. This happens asynchronously, and is not part of transaction protocol, i.e. it is not needed for correctness, but it is a performance optimization technique - all of Firebolt compute engines subscribe to that Kafka topic, and see all changes, which allows them to do some prefetching and prepare for the new state of the system.
#### Additional note [#additional-note]
Alex has the following note in his blog
> MVCC databases may assign two versions: an initial read version, and a final commit version. In this case, we're mainly focused on the specific point at which the commit version is chosen — the time at which the database claims all reads and writes occurred atomically.
This is exactly what's happening in Firebolt's implementation of MVCC. When transaction starts, it obtains from transaction manager LSN to be used for reads - which define the *snapshot* in Snapshot Isolation.
For read-only transactions, there is nothing else, as their commit is no-op. So read-only transactions only get Execute step - there is no Validate, Order or Persist steps.
The read-write transactions get LSN of commit during Order + Persist step described above.
# Diving Into GitHub's Data Stack (/blog/diving-into-githubs-data-stack)
It's the mother of all development projects. You use it daily. And so do 65M developers around the world. This time on the Data Engineering Show – A deep dive into GitHub's data stack. Arfon Smith KimYen (Truong) Ladia shared GitHub's data engineering challenges and solutions and explained why every developer should know and adopt the ADR protocol.
Listen on [Apple Podcasts](https://podcasts.apple.com/us/podcast/diving-into-githubs-data-stack/id1561927688?i=1000539285425) or [Spotify](https://open.spotify.com/episode/0J4HkBtksS6VIXTqRwczVG)
Boaz: Welcome everybody.
Boaz: Hi, Eldad.
Eldad: Hey, Boaz.
Boaz: We are here with another podcast of the Data Engineering Show. That's the name of our show? Right?
Eldad: Yes.
Boaz: This is it. And with us today, we have two great guests. We have Kim Lydia and Arfon Smith, both from GitHub. So, thank you, Kim and Arfon, for joining us.
Boaz: Welcome our guests.
Eldad: Welcome, welcome everyone.
Boaz: So actually, you know, we are very excited to have you on the show when we were sort of discussing in advance, we wonder how GitHub is dealing with data from an engineering perspective, and you know, how all the way we work with software has tickled into the way we work with data. That is an interesting topic in itself; but for starters, how about you two introduce yourselves? Kim, let's go with you first and tell us a little bit about yourself. What do you do at GitHub? And then we will switch to Arfon.
Kim: Hi everybody. I am Kim. I am a software engineer at GitHub, and I have been at GitHub for almost two years now and I am working on the data platform team. Before joining GitHub, I was helping to build a small data warehouse for a nonprofit.
Boaz: Arfon, now it is your turn.
Arfon: Yes. Hi, my name is Arfon Smith. I am a product manager for data at GitHub, and I support data engineering, data science teams, a whole bunch of, sort of, data-focused teams internally. I have a background in astronomy, so actually, previously, I was running a data archive on behalf of NASA for the Hubble space telescope in Baltimore. Before that, I was actually at GitHub. So, I seem to have liked working at GitHub. This is my second time in the company. I was back at GitHub in 2013 for about three years as well.Boaz: Awesome, Awesome. So, we want to dive into, you know, the data world. What do you guys do? So let us start with sort of, just tell us about your data stack and the kind of challenges it has intended to serve.
Kim: Okay. So, our data stack, at a storage layer, we use Azure Dialect storage, and then we use Azure Hadoop and distributors service or Insight. For KeyCloak, you have metastore for transform data. For ingesting data, we use Airflow. We have people using AML for the machine learning pipeline. For compute, we have Presto, Snap, Azure data explorer. And, for BI, we have some custom tools for the individual customer, like Product 360 that you will hear more about later; and we have a more general tool, like Looker and Power BI as well. So, that's a lot.
Boaz: Yeah.
Kim: We have different things going on.
Boaz: Maybe I should have asked what you don't have.
Eldad: S3, no S3 buckets.
Boaz: I think you're the first guests and we have done like, I think, six episodes or so, so far; all have been on AWS. I wonder, you know, if we can talk a little bit about, sort of, Azure. Were you always on Azure? Was there like it was a transition at some point? Was AWS on the table, if you could share some thought processes there.
Kim: Yeah, we used to be on AWS and we made them move to Archer recently. And during that move, we were not just moving from one cloud to another cloud, we were actually moving the kinds of services that we use as well. On AWS, we were operating hundreds of nodes on the Hadoop cluster and it cost us tons of money. When we moved to Archer, we actually used Archer Hadoop and distributors Insights, which manage all of that infrastructure for Hadoop for us and we pay for service rather than paying for the machine first, and so, it is a lot better in terms of how much we pay the bill for it. So, yeah.
Boaz: And was it a complete transition as in Azure only at this point or do you still keep both in, sort of, a multi-cloud strategy?
Kim: The whole data warehouse is fully operations at Azure, right now.
Arfon: That was a big lift for the team.
Eldad: Who was the person who deleted the S3 bucket, kind of when all the tests are passed and then like someone must have pressed the button, right? "drop bucket," which probably still runs.
Boaz: Terrifying moments, probably. How long did the transitional process take?
Kim: There is a lot of transition that carries out. So, we think we have a lot of things like we may take a quarter to transition one thing and then keep everything else the same. And then, every quarter, we move another product over. I was not involved in the beginning, but I was involved in moving Airflow at the very end. So, yeah, we took 3 months at least to move Airflow from operating on AWS to Azure.
Arfon: Almost it took over a year to move. I think it was a long time. It was a big lift.
Boaz: Are all the mental scars healed by now from that transition? You know, people typically dread these things, like it sounds tough. It typically is tough, but you have done it and it worked well, so kudos! Awesome! Let's talk about all the people involved with data. If you could tell us how the data teams are structured? You know, which software teams deal with data? Are they the data engineering teams? Or they are, you know, data science, and everything in between? We love to understand how that works at GitHub.
Arfon: Yeah. So we have, a fairly new structure actually. So, I am going to talk about what we have just reorged to and maybe about why we have made that change. So, we have a data platform team, which encompasses, sort of, a collection of data engineers; but coming the way, we think about that, they sort of run the common fabric of the data warehouse. So if you are a team that wants to provision storage or, you know, some kind of infrastructure for your service, then they sort of build that paved path for provisioning, compute, storage, that kind of thing, role-based access control, that kind of stuff. That team also owns this BI experience. So, the tools that you would use to interrogate the warehouse go around, and you are on self-service kind of questions as a member of staff. And then, we have a collection of what we call "verticals." So, these are teams that are, sort of, close to full stack in terms of the skills as analysts, data scientists, data engineers, and their managers; and they are focused typically on product areas. So, we have GitHub internally, the way that we build the product that sort of strategy around the service that we are building is broken into a few different verticals. And so we have, effectively, data teams aligned with those. And then, there are other data teams, that are not within this core unit, they are sort of a centralized data team for the company. There are other verticals that are outside. So, they are, sort of, more like satellite data teams, maybe in revenue, sales, that kind of thing. So, they are specifically serving their customers, who have sort of different types of questions. So, this is sort of, I think, an idea that has been pretty hot recently, as a sort of data mesh idea. So, building products for a particular set of customers, really owning that relationship and having sort of long-term relationship with those teams that you support. So, we have been doing this without the reorg for about a year now; and it is working pretty well. And so, we sort of fully embraced this model, actually, just this late last week, we have sort of finally kind of crossed the eyes.
Boaz: I mean, you did a lot of changes recently moving to Azure, reorging the data teams. I wonder what is next in terms of change.
Arfon: Wow. It is stability, I hope. So, you know, I think part of what we have been doing as kind of growing up, like reorging the structure to be ready for sort of next level of growth of the business. I think so there's a lot, you know, the demand for data is very high and I think we have sort of found it a little bit hard to support some areas of the business traditionally. And so, you know, especially with the centralized data team, it could be hard to prioritize across the whole business. So now, we are sort of saying, this is a team that is focused on this area of the business or the product; and if it needs more capacity, then it should go and get a headcount and fund it that way, so it sort of gives us as a data team a little bit more of a logical sort of scaling unit.
Boaz: From a headcount perspective, all the teams we now covered, all the people involved with data, how many people are we talking about?
Arfon: So, in terms of sort of data engineering, data science, I think, we are probably at about 35 or 40, not very big. I think compared to other companies that exclude all of the people who store git on file servers and that is all completely separate. So of course, there is a whole sort of data infrastructure team, a set of teams around, running large-scale distributed git, all the application servers, all of that stuff. And in fact, actually, system observability is also separate from that. So. If you think about that without those headcounts, where they are sort of, I think, it's about 35 maybe, maybe 40 now.
Boaz: So let's get back to the two of you. Now, that you have, actually, mapped out the structure of all the people involved with data, back to you guys. How do you fit in with that? You know, both of you have different sort of roles; Kim, you are an engineer, and Arfon you bring sort of a product hat. So, if you could elaborate on what you do within that.
Kim: Yeah, so for me, I have been on the data platform team from the beginning so this reorg doesn't actually affect me.
Boaz: Got it. Okay. Can you share one of, sort of the recent projects that were, sort of, more memorable or exciting for you? Beyond transitioning Airflow to Azure.
Kim: It is more of transitioning Airflow from one version to another version. So, Airflow is an open-source product. It is continually being developed on, and there are new features right now, and we do want to keep on top of all the bug fixes. So, the most recent project for me was just the mixture that we have, AirFlow 2.0 running.
Boaz: Within, you know, the data platform team, what is the distinction between what a software engineer would do and what a data engineer would do?
Kim: I think all of my team members are software engineers and we consider ourselves as data engineers. We solve any problem that our user have or come to us with from operating the platform for them, helping them integrate from one technology to another technology, or even build new features too. For example, Airflow, I have been building the backfill tools that open source do not have, to solve the needs that we do have.
Boaz: And Arfon. Let's get back to your responsibilities.
Arfon: I have sort of, a bit of, a weird product manager role, I would say. So, we have lots of customers for data internally.
Boaz: Arfon! you're not weird. You're super cool. Come on!
Arfon: You're very kind. It is a bit of a mixed bag, honestly. We do have some products that we serve customers with. Some product teams we have this service product called Product 360, which is a thing we built with software engineers and data engineers, and that is sort of a data product that teams use internally to understand common engagement metrics, acquisition, retention, and churn. That kind of thing around how people are using different parts of GitHub, the product? So, sometimes this is a quite traditional sort of product work where it is kind of figuring out what we should do, managing that backlog, understanding what that product should do? what the customer is? But then, there is also a much more sort of general kind of workaround with all these potential customers we could serve internally, where is the lowest hanging fruit? where is the highest return on investment? So, we have quite a lot of autonomy about how we spend our resources. SWe have got 15 data scientists. What should they work on? You know, I think that is a really good question. And so sometimes I spend my time developing ideas with other teams, sort of getting work to the point where it's ready for a data scientist to go invest some time, where you are sort of doing speculative work. We think there is some potential here. We should spend some time doing some R & D to mature an idea to then go to a product team and say, "Hey, look, we think we could, you know, use data to like, you know, supercharge our product in this way or something." And so, it is a mixed bag. It is pretty varied and I enjoy that.
Boaz: It is very interesting. How many pure product manager or product people are there that have a data-specific road?
Arfon: It is just me. So, we are hiring. We just hired a second person. I am delighted to say because there is a lot. One PM to 40 engineers is way off a good ratio from my perspective. So just add two.
Eldad: So now you have high availability at least.
Arfon: Yes! Yes!
Eldad: What about metadata, by the way? Like, what is the most amazing metadata you have been playing with? Like I am thinking GitHub. I am thinking it is best semi-structured. How does that turn into a Looker dashboard? Or if you can share on the metadata side.
Arfon: Yeah. We have a sort of collection of data streams that we put in the warehouse. We have a thing we called hydro, which is, I think just Kafka, you know, that they streamed a collection of services, write events there. This is the sort of paved path if you are a product team. For this particular part of the product, you want to instrument, measure some behavior in the product, then, you write an event in the product. There is a sort of pretty good abstraction as a developer, in Ruby for writing there or in Go, you know, there are mature client libraries that developers use. And, anything that is in hydro in this Kafka stream ends up in the warehouse automatically, and then, we have got, you know, query tools on top of that, Presto or something else. And then, we have a lot of bags that run and turn those into things that are more consumable. So, yeah, I think probably a good example of where we have to invest a lot. If you think about something like security scanning as a product, I have a vulnerability in my project or dependable, that kind of thing, where there is some dependency that I have in my project and I want to be alerted if, you know, there is some vulnerabilities. So, we care a lot about how successful that product is and whether people are responding to those alerts and whether they saw them and whether they did anything, whether they resolve the alert by pushing new code. And so actually, building a set of dashboards that the product team could use to make data-driven insights about how well the product is performing. There was a lot of data engineering work to design those tables, to collect the right data, and build a performance data model that they could query easily. Because there is a lot of dependable alerts and there is a lot of vulnerabilities that get exposed and people get alerted to, so there is actually a lot of work to make a single part of the product understandable for a product team. So, yeah, it probably will be a good example of how uniquely you get a problem. Yeah.
Boaz: What are sort of the key architectural lessons learned? I mean, you know, when you did the shift to Azure, probably huge architecture redo in a lot of places. Can you talk about Hadoop and how you change your approach there? What are the kinds of things that you do not mess with the old architecture and you are smarter about now?
Kim: One of the bad things that I do not mess is, operating ECG instances. I do not like having to check to see if they are dead or alive and they need to restart or how healthy they are. When we moved to Azure, we also adopted Kubernetes and Terraform. So, we use Terraform to deploy our infrastructure, we use Kubernetes to run our containers. So, it is a lot easier to bring up a new container or deployment. It is less time for a user. So yeah, those are the two big things that I am really happy that we are laying on.
Boaz: Which data volumes were we talking about at GitHub?
Kim: About 40 storage account and they range from a couple of megabytes to a couple of petabytes, and they are different data sets. So, it really depends on the topic or the kind of data that is in the storage that determines how big they are; but yeah, we have a wide range of the data sectors.
Arfon: One good example is the hydro topics. I think they are the most voluminous. It is about 4 billion events per day that get written to the warehouse. So, just to give you an idea. I think that is our biggest the hydro topic, pretty sure.
Boaz: What are some of the use cases that are more real time in nature that are running today?
Arfon: This is a satellite data team. So in the sense like this isn't us, but platform health would be a good example. So spam protection. Those teams use these hydro events, you know, so, very low latency sort of system events, and they are running, their own ML algorithms on the fly to detect spammers and quickly remove them. So, I think one of the things that I kind of think is pretty awesome about the GitHub experience is, it is really rare to see spam, and so that's amazing, so that team is incredible and they are just laser-focused on keeping spammers off the platform. So that is a real-time thing. I think also all the stuff that, sort of, other platforms teams do around availability and GitHub receives a lot of denial of service attacks still and stuff like that, but that is not our world, but I think those are some real-time stuff that are pretty important.
Boaz: It does sound like, you know, at GitHub, there is a culture around being data-driven. I mean, there are so many initiatives going around data, data teams are embedded within different services. Is there any standardization of how to approach becoming data-driven within, sort of, new services being launched or a new initiative or is it all very different across different teams with their own analysts and data people?
Arfon: I think this is part of the challenge of the work that we do. I think I am correct in saying that product operations as a concept is quite new at GitHub. This idea of what should you do when you launch a new product, right? Like there is a playbook, like sort of due diligence. I mean, I think we are good at making sure the docs are in a good state, but what sort of instrumentation should you do in the product to measure success? Do you have success metrics? All that kind of stuff I think actually, GitHub is getting much better at that, but that is not a historical strength, I think. So, I think, there are playbooks now. We built services; it is Product 360 service that we built. Really all you have to do as a product team is change a configuration file and then we effectively ingest that as part of the dag, when it runs at Airflow. We ask product teams to define engagement with their product, so if a customer is active in these areas of the product, then that counts as engagement. This counts as a contribution. This counts as acquisition, that kind of thing. And then if they do that, they have got all these sorts of free dashboards that are built every night. And so I think our general tactic is to try and encourage people to do the right thing by building them good tools that if they just follow the playbook, then they are going to get, you know, they get carrots. But more and more, I think the expectation is that a product manager can talk about their product, can talk about how many users they have, can make intelligent comments about, you know, users that are being particularly successful and those less so. And so I think, sort of the culture of being able to reason with data is much, much stronger at GitHub today than it was, let's say, five years ago. So I think you are being very nice when you are saying it sounds like we have a very strong data culture. I think we are growing a stronger data culture. I think traditionally observability has been amazing at GitHub, like the whole kind of ChatOps stuff and the ability to really manage a service from a chat channel was, you know, foundational in GitHubs' engineering culture. And that's a decade old. GitHub has been doing that forever.
Boaz: Let us do spend a couple of minutes on the observability. Even though it's a decade old. Tell us a little bit more about that.
Arfon: I don't know. I can tell you some. We don't own that, but I know we use services like Centrify, Splunk, that kind of thing. Those are services I know about. I know we are moving towards Open Telemetry for logs and metrics. That kind of thing.
Boaz: I actually, yeah, I caught that on your engineering blog.
Arfon: Yeah. There was a nice post about that recently. I think one thing I would say though, is that we have a divide in the way that we think about that data, which is not good, which is observability, is like an engineering problem and like, you know, metrics and instrumentation is a product thing and in some sense, those data, it doesn't make sense to think of those as separate data streams, you know, product managers can build a dashboard from logs if they wanted to, but they traditionally haven't.
Eldad: This is a great point actually. We see that a lot, like, engineering needs data to figure out how to build something and product needs data to figure out what to build; and when they collide, it usually sparks, you know, in there because each side just looks on its own need and it is really hard to understand the other side. So engineering will say, we cannot wait for a full-fledged data warehouse and we cannot wait for this data pipeline and we need it now, now, now. And the product will say, Well! We have figured out most of the stuff on how to do it. Let us figure out what to do. I think we are entering this kind of era where, as you said, those two parts will eventually converge into one big stream where everyone can take their part. But it is interesting to see that. We see it all the time. Even internally at Firebolt, like, there is a race to space around engineering versus product who can build a better system.
Boaz: We managed a big list of things that usually go through the most horrible grunt work. So what is the worst grunt work in your day-to-day job that you wish could magically disappear?
Kim: That would be Python dependency again, not waking up in the morning and figuring out something broke because of a transient, transient, transient dependency that we did not pin devotion on.
Eldad: Say no to Python, just kidding.
Kim: I have been using Python for a while and it has been the single thing that I have been complaining about Python over and over again.
Boaz: It is a love-hate relationship, I guess.
Kim: Yeah.
Boaz: Tell us about, you know, a great, win from a sort of the last year or something that you are proud of or had a good experience with.
Kim: I would say the backfill tools for Airflow. Airflow is very great with managing a workflow, but they don't backfill at an operation that people do day-to-day, and at GitHub, we have so many people writing deck, writing workflow, shipping day data and the data can change like schema change or the metric needs to be updated, the calculation needs to be updated. They are more factor that they learn about and once more accurate model, things like that, that we try backfill where you are running the whole workflow over a long period of time, could be years of data in. Therefore, does not have a good story around that. So we have to build in our own tooling to let user to use backfill or to run backfill on Airflow easily. Reaching out to the open-source projects and hopefully that in the new versions of Airflow, I would engage more in that and help share our learning and share our use-case and hopefully share the new backfill tools in Airflow.
Boaz: Awesome, super interesting, but stop bragging. Now tell us about a failure. What didn't work so well? what can you warn others about or sort of a lesson learned?
Arfon: I was going to say one thing. I will give you success and then I will give you a quick failure because I do want to tell you about one thing. We did miss 4 billion events per day stream that we have. One thing I think we have done that is really good is we have managed to use this common data source. So that 4 billion per day event stream is actually a log of every request to the Rails application that GitHub like monolith, so it is effectively the Rails log, so every request, it was about 4 and a bit billion per day, and we have managed to sort of figure out how to drive the majority of our company core metrics. So things like your monthly engaged users, your top-level business things that attract by the leadership team and Microsoft and using this event stream and the majority of all the individual product areas. So they use a sort of common data source, which felt really, really good to have this common way of approaching metric definitions. That felt like it was very fragmented beforehand. And so, that felt particularly good. It was very satisfying to sort of be able to, sort of, see them rolling up to some sort of total numbers so that I would say that was a big success in terms of trying to take a sort of rational approach to thinking about metrics across the business.
Boaz: Any special tips, I mean, as the experts in everything Git or is there any unique way in which you approach all the data projects from a Git's perspective, something all the regular Git users can learn from?
Arfon: I have one suggestion. It is not really about Git. It is more about the process, which is I think, I appreciate the ADR format. I do not know if you know, that is an Architectural Decision Record. Is that what it stands for?
Kim: Right.
Arfon: Yeah. So one of the things I do enjoy is that GitHub uses the product extensively still. So the idea is you can just go browse any team, what they are working on. But this folder with all the ADRs and now, which is really lovely, so if you are trying to understand, like, why is the service built this way? Like what decision-making process did we go through? There is usually an ADR. If it was a big decision for a team, there is a thing you can go and read and its reasons and what else did we consider? I like that a lot.
Boaz: Who owns the ADR? Like who in engineering?
Arfon: Usually they are sort of principal engineers responsible for building, you know, the decision-maker, but the fact it is in the repository, means it has been through an extensive review. You can go and look at the pull request associated with that review, you know!
Boaz: Awesome!.
Eldad: Is that somehow related to the change in how kind of, you know, how we are working together in a distributed way where time zones are different, people are moving across places. So getting to those basic disciplines suddenly becomes necessary because you just can't hop into a room and just wrap it up quickly. I think, you know, it's huge and small companies, if you want to work with a distributed workforce with engineers, especially, those advisors are gold.
Boaz: I think we should definitely look into it ourselves. I mean, we've been scaling and starting with remote culture and just, you know, collaborative decision-making in a remote way, it is not as easy as it sounds, so super! Yeah! Thanks. Okay! Now. It's, the advice corner, anything sort of inspirational or who to follow, what can you sort of give back to the community in terms of what to look out for, who to follow things like that? And any less famous words from the data people at GitHub.
Arfon: I guess a company that is well in set technologies that really inspires me is the Jupyter project. So this is Jupyter notebooks, you know, one of the things that they do. I think that this is a project that came out of the IPython originally, this was scientific Python ecosystem. I just think they have built a really compelling set of technologies, cross-platform, cross-language, sort of the de facto way to sort of work as a data scientist today, and the way that the project runs is really incredible, as well as they've got really, really sort of enlightened approach to governance and participation and stuff. So yeah, I think Jupyter is, you know, sort of an exemplar for how I think really good open-source projects can be run. And of course, you know, that sort of toolchain is adopted by all the cloud vendors, you know, SageMaker, Azure notebooks, there's Google collaboratory that are all using the same thing, right? That just seems amazing to me. And I think, yeah, that's a great project. I think they're doing great work.
Boaz: Okay. Now for a quick blitz round:
Eldad: Open source or commercial?
Kim: Open source or commercial? Depends. You have me, I'm a little bit skeptical on either one in terms of whether or not I have used it yet. So I want to get my hand on it before I recommend.
Boaz: AWS or Azure.
Kim: Azure.
Arfon: Azure.
Boaz: Work from home or from the office?
Kim: Work from home.
Arfon: A bit of both.
Boaz: Yeah. Okay. That has been awesome. Arfon and Kim, thank you so much for joining us. It has been super interesting and we will see you around.
Eldad: Thank You.
Arfon: That is cool. Thanks for the time.
Kim: Thank you.
Boaz: Our pleasure.
# Eliminating Redundant Joins in Firebolt for Faster SQL (/blog/eliminating-redundant-joins-in-firebolt-for-faster-sql)
## Tl;dr [#tldr]
In this blog post, we provide an explanation of how our query optimizer is able to detect and remove redundant joins. Eliminating redundant joins is critical for real-world SQL queries, especially those generated by query generators which often contain a significant amount of redundancy. It is also essential for Firebolt's ability to optimize complex correlated subqueries. We'll start by explaining what we call a redundant join operation, and then we'll introduce the data structures we use for detecting them in a complex query plan.
## Introduction [#introduction]
In Firebolt, we are obsessed with removing redundancy from complex SQL queries. We recently presented a [paper at BTW 2025](https://www.firebolt.io/content/optimizing-correlated-aggregate-subqueries-in-firebolt) about how we are able to optimize correlated aggregated subqueries in a way no other system can. In the paper, we explain our approach for optimizing the following query pattern with aggregated correlated subqueries in the projection of the query scan the same data:
```sql
SELECT c.*,
(SELECT min(price) FROM orders WHERE status='O' AND c.key=ckey) "min",
(SELECT max(price) FROM orders WHERE status='O' AND c.key=ckey) "max"
FROM customers c;
```
Our optimizer is able to rewrite the query in such a way that the orders table is only scanned once. That is achieved through a combination of query rewrites, all of which are explained in the paper. One such rewrite, and the focus of this blog, is our ability to remove redundant join operations. The usefulness of this query transformation is not limited to the query pattern presented above, but also in other scenarios as well, for example, the result of inlining views accessing the same set of tables may lead to this type of redundancy in the queries.
Eliminating redundant joins is crucial for several reasons in real-world SQL query optimization. Firstly, we have observed that many complex SQL queries, especially those generated by automatic tools, often contain a significant amount of redundancy. Identifying and removing redundant joins in such queries lead to faster query execution and improved overall system efficiency. Secondly, redundant join removal is an essential step in Firebolt's de-correlation process. Our ability to optimize complex correlated aggregate subqueries, as detailed in our BTW 2025 paper, heavily relies on simplifying the query plan by eliminating redundant operations. This not only keeps the final execution plan concise and manageable, but also enables further optimizations that would otherwise be impossible, ensuring highly efficient query processing.
## Redundant inner joins [#redundant-inner-joins]
A redundant join is a join that could be removed from the query plan without affecting the query results. Eliminating redundant joins reduces the amount of data the engine needs to process, which leads to faster query execution. In other words, you get the same correct results with less computational effort – improving overall performance. Additionally, removing redundant operations can simplify the query plan in ways that enable further optimizations that wouldn't have been possible otherwise.
Let's consider the following example:
```sql
-- joining t with itself
SELECT x.*, y.* FROM t AS x INNER JOIN t AS y ON x.a = y.a;
```
Any non-null value of x.a is guaranteed to have at least one matching partner from y, since x.a and y.a have the same provenance, t.a. If a was a unique key in t, then the following query would produce the same results without performing any join:
```sql
-- self join of t is removed
SELECT x.*, x.* FROM t AS x WHERE x.a IS NOT NULL;
```
See how the entire join has been replaced with just a filter removing the null values from a, as the former join condition used to do. Also, notice that all the columns selected from y have been replaced with the corresponding columns from x. In our terminology, we call y the redundant join operand, as it is the join operand being removed, and we call x the replacing operand, since we are replacing all the appearances of columns from y with the corresponding columns from x.
This self-join on a unique key is the most basic example of a removable join. Let's consider now the following query where column a is no longer required to be a unique key of t:
```sql
SELECT x.*,
y.a1
FROM t AS x
INNER JOIN (SELECT DISTINCT a AS a1
FROM t) AS y
ON x.a = y.a1;
```
Similarly, y can be removed from the query because every tuple from x with a non-null value in column a is guaranteed to have a unique matching partner from y. x and y are equi-joined on columns with the same provenance and the joining columns form a unique key in the redundant operand y. The equivalent query without the redundant join is:
```sql
SELECT x.*, x.a AS a1 FROM t AS x WHERE x.a IS NOT NULL;
```
The join needs to be on columns that form a unique key in the redundant join operand. Otherwise, the output cardinality of the rewritten plan is not guaranteed to be the same as with the original plan with the join.
Another required condition is that the replacing join operand should be able to provide all the required columns from the redundant join operand. Consider the following example where y cannot be replaced with x even though they are joined on a column with the same provenance that forms a unique key on y, because x cannot provide the value for y.max after removing the join.
```sql
SELECT x.*,
y.*
FROM t AS x
INNER JOIN (SELECT a AS a1,
Max(b) AS max
FROM t
GROUP BY a) AS y
ON x.a = y.a1;
```
Finally, the common source between the replacing and the redundant join operands must be unfiltered in the redundant side. For example, the replacement in the following case wouldn't be valid, even though the first two requirements are satisfied, because the rows from t1 are filtered in y.
```sql
SELECT x.*
FROM t AS x
INNER JOIN (SELECT DISTINCT a AS a1
FROM t
WHERE b > 10) AS y
ON x.a = y.a1;
```
The query above is effectively a semi-join filtering the rows from x where there is at least one row in t1 with the same value for with a value for b greater than 10. Rewriting the filter to filter on x.b, as in the following query, would not be valid unless b is a functional dependency of a, where the value of column b is fully determined by the value of column a.
```sql
-- if b is functionally dependent on a
SELECT x.*, x.* FROM t AS x WHERE x.a IS NOT NULL AND x.b > 10;
```
## Redundant outer joins [#redundant-outer-joins]
Query optimizer people know that the real fun always starts with outer joins. There are two types of redundant outer joins we can identify and remove. The first type is shown by the following example where a is again required to be a unique key of t:
```sql
SELECT x.*, y.* FROM t AS x LEFT JOIN t AS y ON x.a = y.a;
```
This query could be rewritten as the following:
```sql
-- if a is a unique key of t, remove the left join
SELECT x.*,
CASE WHEN x.a IS NOT NULL THEN a ELSE NULL END,
CASE WHEN x.a IS NOT NULL THEN b ELSE NULL END,
...
FROM t AS x;
```
Note that the filter "x.a IS NOT NULL" is no longer added since left outer joins must always preserve the rows from its left input. Instead, we need to replace all columns from y with columns from x, but making them NULL when column a is NULL, using CASE WHEN operator, since that is the only case where a row from the left-hand side won't find any matching partner in the right-hand side of the left join.
For an equi-join whose join condition does not reject nulls, for example, a join using IS NOT DISTINCT FROM comparison, there is no need to wrap the columns from x within these CASE WHEN expressions when replacing the columns from y. Implementation-wise though, we always add this wrapping CASE WHEN expressions, use the original join condition as their condition, and let the optimizer simplify them. This way, the code has less special paths to maintain. The resulting optimization steps for these two queries are presented below.
```sql
-- example 1
SELECT x.*, y.* FROM t AS x LEFT JOIN t AS y ON x.a = y.a;
-> (redundant join removal)
SELECT x.*,
CASE WHEN x.a = x.a THEN x.a ELSE NULL END,
CASE WHEN x.a = x.a THEN x.b ELSE NULL END,
...
FROM t AS x;
-> (expression reduction)
SELECT x.*,
CASE WHEN x.a IS NOT NULL THEN x.a ELSE NULL END,
CASE WHEN x.a IS NOT NULL THEN x.b ELSE NULL END,
...
FROM t AS x;
-> (expression reduction, if a is not nullable)
SELECT x.*,
x.a,
x.b
...
FROM t AS x;
```
```sql
-- example 2
SELECT x.*, y.* FROM t AS x LEFT JOIN t AS y ON x.a IS NOT DISTINCT FROM y.a;
-> (redundant join removal)
SELECT x.*,
CASE WHEN x.a IS NOT DISTINCT FROM x.a THEN x.a ELSE NULL END,
CASE WHEN x.a IS NOT DISTINCT FROM x.a THEN x.b ELSE NULL END,
...
FROM t AS x;
-> (expression reduction)
SELECT x.*,
x.a,
x.b
...
FROM t AS x;
```
What if there are more conditions on the join? As long as the unique key from the redundant operand is equi-joined with columns from the replacing operand through columns with the same provenance, we can simply use these wrapping CASE expressions, using the entire join condition as their condition. The following example shows this.
```sql
SELECT x.*, y.* FROM t AS x LEFT JOIN t AS y ON x.a = y.a AND x.b > 10;
-> (redundant join removal)
SELECT x.*,
CASE WHEN x.a = x.a AND x.b > 10 THEN x.a ELSE NULL END,
CASE WHEN x.a = x.a AND x.b > 10 THEN x.b ELSE NULL END,
...
FROM t AS x;
-> (expression reduction)
SELECT x.*,
CASE WHEN x.a IS NOT NULL AND x.b > 10 THEN x.a ELSE NULL END,
CASE WHEN x.a IS NOT NULL AND x.b > 10 THEN x.b ELSE NULL END,
...
FROM t AS x;
-> (expression reduction, if a is not nullable)
SELECT x.*,
CASE x.b > 10 THEN x.a ELSE NULL END,
CASE x.b > 10 THEN x.b ELSE NULL END,
...
FROM t AS x;
```
Let's now consider the second type of redundant outer join shown in the query below, which is the one that we need to cover for optimizing the queries in our BTW 2025 paper.
```sql
SELECT t2.*,
x.*,
y.*
FROM t2
LEFT JOIN t1 AS x
ON t2.i = x.a
LEFT JOIN t1 AS y
ON t2.i = y.a;
```
In this case, we can remove y and replace it with x (or vice versa) as long as column a forms a unique key of t1 even though x and y are not directly joined but indirectly through t2. The resulting rewritten query is:
```sql
-- remove redundant join operand y
SELECT t2.*, x.*,
CASE t2.i = x.a THEN x.a ELSE NULL END,
CASE t2.i = x.a THEN x.b ELSE NULL END,
...
FROM t2 LEFT JOIN t1 AS x ON t2.i = x.a
```
Which can be further simplified to:
```sql
SELECT t2.*, x.*,
x.a,
x.b,
...
FROM t2 LEFT JOIN t1 AS x ON t2.i = x.a;
```
Since the columns from x are already guaranteed to be NULL when the predicate doesn't evaluate to true.
In the example above, both the replacing and the redundant join operands are in the non-preserving side of an outer join. In order for the redundant join removal to be valid, the outer join predicates of the replacing operand must be a subset of the outer join predicates of the redundant operand. In the following example, the join with y can still be removed from the query, replacing its referenced columns with columns from x, as long as column a forms a unique key in t1. However, x cannot be replaced with y due to this new requirement.
```sql
-- cannot remove the join to x due to the extra join condition on y
SELECT t2.*,
x.*,
y.*
FROM t2
LEFT JOIN t1 AS x
ON t2.i = x.a
LEFT JOIN t1 AS y
ON t2.i = y.a
AND t2.j > 10; -- extra join condition
```
Again, the predicates of the removed operand become the condition of the wrapping CASE expressions:
```sql
SELECT t2.*, x.*,
CASE t2.i = x.a AND t2.j > 10 THEN x.a ELSE NULL END,
CASE t2.i = x.a AND t2.j > 10 THEN x.b ELSE NULL END,
...
FROM t2 LEFT JOIN t1 AS x ON t2.i = x.a
```
In order for these rewrites to be valid, we've mentioned earlier that there needs to be a join directly or indirectly joining columns with the same provenance from both sides covering a unique key from the redundant join operand. And that alone is not sufficient. We need to make sure the common source between the redundant and the replacing operands must remain unfiltered in the redundant side. Consider the following example:
```sql
SELECT *
FROM t3
LEFT JOIN (t1
INNER JOIN t2
ON t1.a = t2.i) AS x
ON t3.m = t1.a
LEFT JOIN t1 AS y
ON t3.m = y.a
```
In this case, the inner join between t1 and t2 is the non-preserving side of the outer join with t3. That inner join may filter rows from t1. As a result, we cannot remove any outer join as a redundant join.
## Other types of joins [#other-types-of-joins]
Performing redundant join removal for semi joins is trivial since they are just a special type of inner joins. There is no extra consideration needed for them. In fact, for any equi-semi-join, the columns from the semi-side involved in the join conditions always form a unique key. In the following example, removing the semi-join is possible even if a is not a unique key of t, thanks to the semi-join semantics.
```sql
SELECT * FROM t AS x WHERE x.a IN (SELECT y.a FROM t AS y);
->
SELECT * FROM t AS x WHERE x.a IS NOT NULL;
```
On the other hand, anti joins are not simple. If fun starts with outer joins, pain starts with anti joins. One reason is that there are effectively two types of anti-joins with different semantics on nulls. The following example shows how the two most basic types of redundant anti-joins could be rewritten.
```sql
-- NOT IN
SELECT * FROM t AS x WHERE x.a NOT IN (SELECT y.a FROM t AS y);
->
SELECT * FROM t AS x WHERE FALSE;
-- NOT EXISTS
SELECT * FROM t AS x WHERE NOT EXISTS (SELECT * FROM t AS y WHERE x.y = y.a)
->
SELECT * FROM t AS x WHERE x.a IS NULL;
```
Although we all have rewritten some IN predicate to EXISTS or vice versa in SQL queries, the negated version the two predicates have a subtle difference in null semantic, which is often overlooked.
A simple example is shown below. x and y are indirectly joined. It is obvious that one of the joins is redundant and can be removed.
```sql
SELECT *
FROM t2
WHERE t2.i NOT IN (SELECT x.a FROM t1 AS x)
AND t2.i NOT IN (SELECT y.a FROM t1 AS y); -- redundant (either x or y)
->
SELECT *
FROM t2
WHERE t2.i NOT IN (SELECT x.a FROM t1 AS x);
```
In the following example, y can still be removed since it doesn't filter out any row from t2 that x isn't already filtering.
```sql
SELECT *
FROM t2
WHERE t2.i NOT IN (SELECT x.a FROM t1 AS x) -- cannot remove
AND t2.i NOT IN (SELECT y.a FROM t1 AS y WHERE y.b > 10); -- still redundant
->
SELECT *
FROM t2
WHERE t2.i NOT IN (SELECT x.a FROM t1 AS x);
```
For NOT IN anti-joins, the third condition is a bit different as it is the replacing operand, instead of the redundant operand, where the common source of the data must remain unfiltered.
## Implementation details [#implementation-details]
Our redundant join removal implementation searches for redundant join operands by analyzing the join graph. Within a join graph, the process of removing a redundant join corresponds to removing a vertex (which represents a plan) and an associated edge (which represents the join) from the graph.
Below is a trace from our optimizer after removing the redundant join from one of the examples we saw before.
```sql
SELECT x.*,
y.a1
FROM t AS x
INNER JOIN (SELECT DISTINCT a AS a1
FROM t1) AS y
ON x.a = y.a1;
Sub-plan before rule: RedundantJoinRemovalRule
[0] [Projection] t.a, t.b, a1
| [JoinGraph]
| - Projection: class_0, vertex_1.t.b, class_0
| - Vertices:
| Vertex 0 -> Node 3, provided classes: 0
| Vertex 1 -> Node 2, provided classes: 0
| - Edges:
| type: Inner, predicate: [equals(class_0, class_0)]
| - Equivalence classes:
| Class 0: vertex_0.a1, vertex_1.t.a
\_[1] [Join] Mode: Inner [equals(t.a, a1)]
\_[2] [StoredTable] Name: "t"
\_[3] [Aggregate] GroupBy: [a1: t.a] Aggregates: []
| [Unique Keys]: [a1]
\_[4] [StoredTable] Name: "t"
Sub-plan after rule: RedundantJoinRemovalRule
[6] [Projection] t.a, t.b, a1: t.a
\_[7] [Filter] equals(t.a, t.a)
\_Recurring Node --> [2]
```
Our plan representation uses a DAG representation with binary join operators. However, every node that represents the root node of a join tree is annotated on demand with a simple join graph representing the join graph under it. In the example above, the join graph contains two vertices, which correspond to nodes 2 and 3 in the plan, and one edge, joining the two vertices.
This join graph representation is also used for other optimizations, such as join reordering or aggregation push down.
Equality predicates among expressions from different vertices are represented as equivalence classes. This way we can represent the relation among several vertices with a single edge. The following snippet is the result of adding an extra join operand to the previous example, which is joined to the other two through the same equivalence class, leading to a single edge.
```sql
SELECT x.*,
y.a1
FROM t AS x
INNER JOIN (SELECT DISTINCT a AS a1
FROM t) AS y
ON x.a = y.a1
INNER JOIN t2
ON t2.i = x.a;
Sub-plan before rule: RedundantJoinRemovalRule
[0] [Projection] t.a, t.b, a1
| [JoinGraph]
| - Projection: class_0, vertex_0.t.b, class_0
| - Vertices:
| Vertex 0 -> Node 3, provided classes: 0
| Vertex 1 -> Node 4, provided classes: 0
| Vertex 2 -> Node 7, provided classes: 0
| - Edges:
| type: Inner, predicate: [equals(class_0, class_0)]
| - Equivalence classes:
| Class 0: vertex_0.t.a, vertex_1.a1, vertex_2.t2.i
\_[1] [Join] Mode: Inner [equals(t2.i, t.a)]
\_[2] [Join] Mode: Inner [equals(t.a, a1)]
| \_[3] [StoredTable] Name: "t"
| \_[4] [Aggregate] GroupBy: [a1] Aggregates: []
| | [Unique Keys]: [a1]
| \_[5] [Projection] a1: t.a
| \_[6] [StoredTable] Name: "t"
\_[7] [StoredTable] Name: "t2"
Sub-plan after rule: RedundantJoinRemovalRule
[8] [Projection] t.a, t.b, a1: t.a
| [JoinGraph]
| - Projection: class_0, vertex_0.t.b, class_0
| - Vertices:
| Vertex 0 -> Node 10, provided classes: 0
| Vertex 1 -> Node 11, provided classes: 0
| - Edges:
| type: Inner, predicate: [equals(class_0, class_0)]
| - Equivalence classes:
| Class 0: vertex_0.t.a, vertex_1.t2.i
\_[9] [Join] Mode: Inner [equals(t.a, t2.i)]
\_[10] [Projection] t.a, t.b
| \_Recurring Node --> [3]
\_[11] [Projection] t2.i
\_Recurring Node --> [7]
```
If the extra vertex is joined to only one the other two vertices, i.e. through a different column, our simple join graph will contain two inner join edges corresponding to the two join equivalence classes, as shown in the example below.
```sql
SELECT x.*,
y.a1
FROM t AS x
INNER JOIN (SELECT DISTINCT a AS a1
FROM t) AS y
ON x.a = y.a1
INNER JOIN t2
ON t2.i = x.b;
Sub-plan before rule: RedundantJoinRemovalRule
[0] [Projection] t.a, t.b, a1
| [JoinGraph]
| - Projection: class_0, class_1, class_0
| - Vertices:
| Vertex 0 -> Node 3, provided classes: 0, 1
| Vertex 1 -> Node 4, provided classes: 0
| Vertex 2 -> Node 7, provided classes: 1
| - Edges:
| type: Inner, predicate: [equals(class_0, class_0)]
| type: Inner, predicate: [equals(class_1, class_1)]
| - Equivalence classes:
| Class 0: vertex_0.t.a, vertex_1.a1
| Class 1: vertex_0.t.b, vertex_2.t2.i
\_[1] [Join] Mode: Inner [equals(t2.i, t.b)]
\_[2] [Join] Mode: Inner [equals(t.a, a1)]
| \_[3] [StoredTable] Name: "t"
| \_[4] [Aggregate] GroupBy: [a1] Aggregates: []
| | [Unique Keys]: [a1]
| \_[5] [Projection] a1: t.a
| \_[6] [StoredTable] Name: "t"
\_[7] [StoredTable] Name: "t2"
Sub-plan after rule: RedundantJoinRemovalRule
[8] [Projection] t.a, t.b, a1: t.a
| [JoinGraph]
| - Projection: vertex_0.t.a, class_0, vertex_0.t.a
| - Vertices:
| Vertex 0 -> Node 10, provided classes: 0
| Vertex 1 -> Node 12, provided classes: 0
| - Edges:
| type: Inner, predicate: [equals(class_0, class_0)]
| - Equivalence classes:
| Class 0: vertex_0.t.b, vertex_1.t2.i
\_[9] [Join] Mode: Inner [equals(t.b, t2.i)]
\_[10] [Projection] t.a, t.b
| \_[11] [Filter] equals(t.a, t.a)
| \_Recurring Node --> [3]
\_[12] [Projection] t2.i
\_Recurring Node --> [7]
```
As explained before, in order for this optimization to happen we need to check that the redundant vertex is joined to the replacing vertex through one of its unique keys. As shown in the examples above, we keep track of the unique keys of each node in the plan as computed information that gets attached to the query plan.
The last remaining bit is how to check whether two columns joined in the join graph have the same source. For that, we use another computed property that, for every node, keeps track of the steps its projected columns went through. The following snippet shows our first example again with this extra annotation.
```sql
SELECT x.*,
y.a1
FROM t AS x
INNER JOIN (SELECT DISTINCT a AS a1
FROM t) AS y
ON x.a = y.a1;
Sub-plan before rule: RedundantJoinRemovalRule
[0] [Projection] t.a, t.b, a1
| [JoinGraph]
| - Projection: class_0, vertex_1.t.b, class_0
| - Vertices:
| Vertex 0 -> Node 3, provided classes: 0
| Vertex 1 -> Node 2, provided classes: 0
| - Edges:
| type: Inner, predicate: [equals(class_0, class_0)]
| - Equivalence classes:
| Class 0: vertex_0.a1, vertex_1.t.a
\_[1] [Join] Mode: Inner [equals(t.a, a1)]
\_[2] [StoredTable] Name: "t"
| [Column Provenance]:
| - source: [2]
| column expressions: t.a, t.b
| inverse path:
| - source: t1
| column expressions: t.a, t.b
| inverse path:
\_[3] [Aggregate] GroupBy: [a1] Aggregates: []
| [Column Provenance]:
| - source: [3]
| column expressions: a1
| inverse path:
| - source: [4]
| column expressions: a1
| inverse path: 0
| - source: [5]
| column expressions: t.a
| inverse path: 0, 0
| - source: t
| column expressions: t.a
| inverse path: 0, 0
| [Unique Keys]: [0]
\_[4] [Projection] a1: t.a
\_[5] [StoredTable] Name: "t"
```
By using this property, we can quickly check that vertex\_0.a1 and vertex\_1.t.a have a common source, which is the t.a column.
This column provenance property is effectively a flattened version of the subtree under each node, which makes it a memory intensive property which is also costly to compute. However, it is quite useful for several other optimizations.
# Engines: Online Scaling and Upgrades (/blog/engines-online-scaling-and-upgrades)
Firebolt is a next-gen data warehouse platform designed for data-intensive analytics. It provides high efficiency and low latency, allowing data engineers to deliver scalable analytics faster, more cost-effectively, and with greater simplicity. In Firebolt, an Engine is the compute resource that customers will use to ingest data into Firebolt and to run queries on the ingested data. An Engine comprises one or more clusters, where each cluster is a collection of nodes of a certain type that provides a certain amount of CPU, RAM, and storage. Among the critical capabilities that engines provide to meet the needs of modern analytic workloads are: **1/** Multi-dimensional Elasticity, which allows users to scale their engines across multiple dimensions. **2/** Granular Scaling allows users to incrementally add (or remove) nodes to their engines, and **3/** Seamless Online Upgrades, enabling users to get the latest security updates and performance enhancements without interrupting their workloads. This blog post will discuss how Firebolt implemented zero-downtime upgrades under the hood.
## Firebolt Architecture [#firebolt-architecture]
Firebolt implements a "disaggregated *shared-everything architecture*" with "compute and storage separation." This architecture allows multiple compute resources running different workloads to be fully isolated from each other while sharing the same data. The visual below demonstrates the logical architecture of Firebolt.

**The key components of the architecture are:**
* Firebolt Gateway - Responsible for routing customer requests to the correct Engine.
* Engines Compute - Clusters of machines running in Kubernetes and service customer queries.
* Object Storage - Stores customer data
* Engines Control Plane - Manages the lifecycle of engines and is responsible for provisioning, de-provisioning, scaling, and upgrading engines.
Separating compute from storage allows the compute and storage layers to scale independently, providing customers with flexibility in scaling their engines to achieve optimal price/performance for their workloads. In designing the architecture for engines, we made important design choices around the following:
* **Fast Infrastructure Provisioning**: Enable customers to create and scale their engines quickly.
* **Managing Compute:** Running engines on Kubernetes vs. EC2 VMs directly
* **Multi-cluster Management:** Whether to use Service Mesh for managing multiple EKS clusters
### Fast Infrastructure Provisioning [#fast-infrastructure-provisioning]
To enable customers to quickly create and scale their existing engines, Firebolt pre-provisions a fleet of compute nodes called Warm Pools. When a customer requests a new engine or wants to scale their existing engine, Firebolt leverages the Warm Pool to satisfy these requests rather than going to the underlying cloud provider, thus providing fast engine operations.
### Managing Compute [#managing-compute]
To efficiently manage the compute resources needed for engines, we decided to use AWS Elastic Kubernetes Service (EKS) and AWS Karpenter, an open-source high-performance autoscaler for Kubernetes. The built-in scaling and scheduling capabilities provided by Kubernetes and the autoscaling capabilities of Karpenter gave us a great degree of flexibility in how we dynamically manage our compute infrastructure at scale without the need to maintain raw VM fleets.
### Multi-cluster Management [#multi-cluster-management]
As noted earlier, Firebolt is designed to meet the performance demands of today's low-latency, high-concurrency workloads. To meet these performance requirements, which can be in the tens of milliseconds, we needed a way to support multiple EKS clusters communicating with each other in a fast and secure manner. To manage these multiple EKS clusters, we used sidecar-less multi-cluster mesh to enable direct pod-to-pod [OSI L4](https://en.wikipedia.org/wiki/OSI_model#Definitions) communication between engine nodes without introducing extra hops in network communication.
### Engines Control Plane [#engines-control-plane]
Engines Control Plane is a set of micro-services responsible for Engine lifecycle management. The Control plane is responsible for:
* Provisioning, scaling, and upgrading engines
* Providing engine routing information
* Providing consumption & auditing information
The control plane provides a reliable workflow to manage the lifecycle of engines, interacting with Kubernetes clusters in order to provision and deprovision actual hardware resources. We use [Temporal](https://temporal.io) as a backend for workflow orchestration to satisfy our requirements:
* Provide reliable at-least-once execution for every workflow step.
* Provide capabilities to serialize operations on the objects of interest.
### Engine Architecture [#engine-architecture]
When designing Engines, one of our key goals was to ensure we provide the latest software upgrades without any downtime for our customers. We want to deliver these upgrades to our customers at a regular cadence without requiring them to stop their engines. To deliver these online upgrades without interrupting customer workloads, we considered the two options below:
1. Make changes to the engine nodes in-place while they are running (or)
2. Create an entirely new set of nodes that have the latest upgrades
Although making changes in-place to an engine while it is still running helps deliver the upgrades without customers having to stop their engines, any issues arising from the upgrades could impact the performance of the running workloads. Hence, we decided to pursue the second option above, for which we leveraged engine clusters. As noted earlier, an engine cluster is a collection of nodes. We designed clusters to be immutable objects to avoid disrupting customer workloads during the online upgrades. Making clusters immutable allowed us to guarantee that we won't push configuration changes to already running clusters, which could potentially introduce unexpected and undesirable changes to customer workloads. Before we look at how Firebolt uses immutable clusters to deliver online upgrades seamlessly, let's understand how dynamic engine scaling works in Firebolt. The upgrade process is built on top of dynamic scaling.
### Request Routing [#request-routing]
Firebolt Gateway fleet is a multi-tenant service responsible for routing requests to the correct engine and load balancing across engine clusters. Each engine cluster and each node within a cluster are individually routable. To support that, the Gateway fleet maintains an up-to-date state for each engine cluster, including:
* Cluster state (stopped, starting, running, or draining)
* Routing information for each node of the cluster
* Cluster routing strategy
Once a request lands on one of the Gateway instances, the Gateway service performs authentication engine discovery and then routes the request to the corresponding engine. Each request carries information about the engine it is targeting. At a high level, the flow for the request to reach the engine node is as follows:
1. The Gateway node receives the request
2. Gateway performs authentication
3. Discover requested engine state and routing information
4. If the engine is stopped and auto-start is enabled, start the engine
5. Select the running engine cluster
6. Select engine node within a cluster
7. Route request to selected engine node
The logical engine layout as Gateway fleet sees it is represented below:

The gateway fleet is crucial for load distribution across engine clusters and unlocks online scaling capabilities. Since the Gateway fleet is on the critical query execution path, having the minimum possible request routing overhead times is paramount.
Our goal was to have sub-1ms Gateway latency overhead at p50.
To achieve that, the Gateway fleet implements the:
* In-memory engine routing information on each of the Gateway instances.
* Subscribes to notifications event stream delivered from Engines control plane for any engine cluster state changes. This increases cache hit rates by proactively populating the in-memory cache with up-to-date engine states.
In the unlikely case of a cache miss, Gateway will explicitly fetch fresh Engine routing information from the Engine's control plane via a synchronous request.
## Online Scaling [#online-scaling]
The Control Plane for Engines drives the process of online scaling of any engine.
The process to scale an engine:
* Create a new engine cluster and start it
* Switch traffic over to a new cluster
* Drain the old cluster and stop it
* Delete old cluster
To ensure the scaling process does not cause downtime, the Gateway fleet and Engines control plane work together to avoid traffic disruption. Gateway fleet receives up-to-date state from the Engines control plane regarding the state of each engine cluster. Using this information, Gateway selects the appropriate cluster for sending traffic and avoids disrupting the queries running on the old clusters. The Engine's control plane tracks whether an engine cluster has any active queries before shutting it down during the drain phase.
We'll demonstrate the scaling process in more detail by scaling from a 2-node engine cluster to a 3-node engine cluster. The scaling process along other dimensions (changing the node Type or the Number of Clusters) is conceptually the same.
All operations in Firebolt can be performed via SQL or UI. When the user first creates an engine with the SQL command below, Firebolt will start with a 2-node engine cluster receiving customer traffic:
```sql
CREATE ENGINE MyEngine IF NOT EXISTS WITH TYPE = S NODES = 2;
```

As the workload continues running, let us say that the user wants to scale out their engine from a 2-node cluster to a 3-node cluster to meet their performance demands. They can easily do so using the following SQL command:
```sql
ALTER ENGINE MyEngine SET NODES = 3;
```
The below visual shows what this scaling process looks like in Firebolt.

Now, we will go under the look and look into the details of this scaling process:
### Create and start a new engine cluster [#create-and-start-a-new-engine-cluster]
1. A new 3-node engine cluster is created.
2. Control Plane starts the cluster and updates its state to RUNNING once it's up and running.

### Traffic switchover [#traffic-switchover]
1. Gateway fleet receives a notification from the control plane stating that the Engine state has changed.
2. The Gateway fleet starts routing traffic to the new cluster.
3. The Engines control plane switches the old cluster to a draining state
4. The gateway fleet receives a notification stating that the engine state has changed and detects that the old cluster is in a DRAINING state. The Gateway fleet stops sending traffic to the old cluster.

### Drain traffic from the old cluster [#drain-traffic-from-the-old-cluster]
The Engine control plane tracks the number of queries running on the old cluster. Once the number of active queries reaches zero, it invokes engine cluster shutdown.

These operations allow Firebolt to perform "blue-green style" deployment when changing engine shape to the specification requested by the customer.
### Graceful drain [#graceful-drain]
As mentioned above, old clusters are gracefully drained during online scaling before shutting down. This is necessary to ensure queries running on an engine when the engine was scaled to a different configuration can be completed successfully.
Steps to execute graceful drain:
1. The Engines control plane marks the Engine Cluster as DRAINING. It sends notifications about the state change.
2. The gateway fleet receives an engine state notification, detects a draining cluster, and removes it from the list of clusters eligible to receive traffic.
3. Any new query that reaches the Gateway fleet will be routed to engine clusters that are still RUNNING.
4. The control plane watches the number of running queries on the cluster that are being drained. Once no running queries remain, the control plane invokes the engine cluster destroy sequence that removes underlying Kubernetes resources and engines cluster metadata.
We had several options for how graceful drain would check for running queries:
* Waiting for each engine node to notify the control plane when it has no active queries
* Rely on Kubernetes graceful drain mechanism to self-shutdown each node once no running queries remain
* Making control plane poll engine cluster for the number of running queries
Each engine cluster performs distributed query processing. Engine cluster nodes have to be running to execute distributed queries. Because of that, individual nodes could not make a shutdown decision, and the control plane had to coordinate the process. We opted for the third option of the control plane polling the cluster. It coordinates the shutdown sequence and reduces the amount of traffic flowing from each engine cluster to the control plane during normal system operations.
## Online Upgrade [#online-upgrade]
Firebolt leverages both online scaling and graceful drain to perform engine upgrades.
### Single engine upgrade [#single-engine-upgrade]
The Control Plane is driving the engine upgrade process. It takes a sequence of steps:
1. For each engine cluster, the Control plane creates a new shadow engine cluster with the same specifications as the old one, except a new Firebolt database version is used. The control plane starts a shadow cluster.
2. Control Plane changes engine traffic configuration to mirror mode. Any request to an engine cluster is mirrored to the corresponding shadow cluster. Routing is performed to ensure the same load distribution on primary and shadow clusters.
3. Gateway fleet receives new traffic configuration and starts traffic mirroring.
4. The upgrade verification process starts.
5. The upgrade process waits for verification to finish.
6. If the verification result is a success, the upgrade process performs promotion.
7. The promotion process makes the shadow a new primary and starts a graceful drain process for the old cluster.
8. Once the graceful drain is complete, the control plane shuts down an old cluster and removes it from the system.
9. The upgrade process stores rollback information for the engine in case a regression is detected later, and the fleet needs to be rolled back.
### Upgrade verification process [#upgrade-verification-process]
The upgrade verification process tracks various metrics from old and new engine clusters. Tracked metrics include query latencies and success rates.
The verification process uses a sliding window to compute latency distributions for the queries that ran within verification windows. It detects whether there are any latency regressions caused by the rollout of the new database software version.
When tracked metrics converge within an acceptable margin or error, the verification process returns control to the single-engine upgrade process.
If the error rate for the queries running on the new version of the database software increases, the verification process aborts, signaling the engine upgrade process to rollback.
### Fleet wide upgrade [#fleet-wide-upgrade]
At the fleet-level, the engine upgrade process updates engines based on the segment - a group of customers at a time, grouped by the risk profile in order of:
* Firebolt Internal
* Preview
* Stable
For a single segment upgrade process is:
* Upgrade running engines in the segment
* Update the default version for new engines in the segment
* Update the version for the stopped engine in the segment
If a regression is detected during the upgrade, rollback procedures are performed promptly to mitigate the impact on customers.
The above process continues until all engines are running using the new version.
## Summary [#summary]
Firebolt engines provide multi-dimensional elasticity to our customers, allowing them to achieve the desired price-performance without causing downtime for customers.
Learn more about [engines](https://www.firebolt.io/elasticity), or [get started for free](https://go.firebolt.io) with Firebolt now.
For more information about Firebolt Engines, read our whitepaper [here](https://www.firebolt.io/resources/firebolt-elasticity-technical-whitepaper).
# Exploring your data lake in Firebolt using just TVFs (/blog/exploring-your-data-lake-in-firebolt-using-just-tvfs)
**TL;DR:** Firebolt supports `read_parquet()` and `read_csv()` functions that make it easy to explore data in your data lake or object storage. This blog post describes how these table-valued functions (TVFs) work under the hood and how to use them effectively.
With read TVFs, Firebolt takes care of everything: locating your data objects, inferring the schema, querying them efficiently, and returning structured table values. The `read_parquet()` and `read_csv()` functions are straightforward to use:
1. Provide either:
* A URL to your data, with credentials if required, or
* A Firebolt Location object.
2. Firebolt infers the schema from the most recently modified file in the provided location.
3. The function returns structured rows of data.
4. Profit.
The simplicity makes these functions useful for ad-hoc data exploration and incremental data loading. There is no overhead of defining a persistent schema object (as required by external tables) or ingesting the data (as with `COPY FROM`).
### Show Me What To Do [#show-me-what-to-do]
Here are some examples. All of these work on public S3 buckets. You can provision a Firebolt engine in us-east-1 on AWS and run them right away:
**1. Query public data directly. Use glob patterns to query multiple files at once.**
```sql
SELECT * FROM
read_parquet('s3://firebolt-publishing-public/help_center_assets/firebolt_sample_dataset/rankings/*.parquet')
LIMIT 10;
```
It is often helpful to use the `list_objects()` TVF to discover your file name and directory structure.
```sql
SELECT * FROM
list_objects('s3://firebolt-publishing-public/help_center_assets/firebolt_sample_dataset/');
```
**2. Create permanent tables from query results while customizing CSV parsing options:**
```sql
CREATE TABLE levels AS
SELECT * FROM read_csv('s3://firebolt-publishing-public/help_center_assets/firebolt_sample_dataset/levels.csv',
header => TRUE,
empty_field_as_null => FALSE);
```
**3. To access non-public data, specify the credentials using a location object. This approach is recommended for both ease of use and security.**
```sql
CREATE LOCATION my_location WITH
SOURCE = AMAZON_S3
CREDENTIALS = (AWS_ROLE_ARN = '')
URL = 's3://foo/bar.csv';
SELECT * FROM
read_csv(location => 'my_location');
```
Alternatively, you can specify your credentials using named parameters.
```sql
SELECT * FROM
read_csv(url => 's3://foo/bar.csv',
aws_role_arn => 'YOUR_AWS_ROLE');
```
### Working with TVF Data [#working-with-tvf-data]
After retrieving data with a TVF, you can filter, aggregate and transform using standard SQL operations. If desired, you can also ingest the data into managed storage using `CREATE TABLE AS SELECT` or `INSERT INTO SELECT` commands.
**1. Apply Transformations:** TVFs like `read_parquet()` not only access data but also enable immediate transformations. Here, we're analyzing gaming performance data with column-level transformations applied directly within the query:
The file in this example contains capital letters in the column names. For this reason, we use quoted identifiers to select columns.
```sql
-- Which cars correlate with best player performance?
SELECT
p."SelectedCar",
COUNT(DISTINCT p."PlayerID") AS "PlayerCount",
MEDIAN(r."PlaceWon") AS "MedianPlacement",
AVG(r."TotalScore") AS "AvgScore",
AVG(p."CurrentSpeed") AS "AvgSpeed"
FROM read_parquet('s3://firebolt-publishing-public/help_center_assets/firebolt_sample_dataset/playstats/*') p
JOIN read_parquet('s3://firebolt-publishing-public/help_center_assets/firebolt_sample_dataset/rankings/*') r
ON p."GameID" = r."GameID" AND p."PlayerID" = r."PlayerID"
GROUP BY p."SelectedCar"
ORDER BY "MedianPlacement", "AvgScore" DESC;
```
Notice how the TVF allows you to join data from different sources without first materializing it into tables. This query demonstrates joining two separate Parquet file collections, performing aggregations, and formatting the results - all in a single operation.
**2. Filter by metadata:** Firebolt supports filtering by metadata when using TVFs, which you can leverage to implement retry capability for transient network failures. This works by deduplicating files based on their names, making your data pipelines more robust.
Here's how to implement it:
1. Add a source\_file\_name column to your table.
2. Use the following pattern when inserting data.
```sql
INSERT INTO table_name
SELECT *, $source_file_name FROM read_parquet(...)
WHERE $source_file_name NOT IN
(SELECT DISTINCT source_file_name FROM table_name);
```
### Supported File Formats [#supported-file-formats]
Today, the supported formats are Parquet (via `read_parquet()`) and CSV (via `read_csv()`), the two most common data formats in Firebolt.
We are actively working on support for additional data formats, including Apache Iceberg. Check back soon for more.
### Performance Benefits [#performance-benefits]
TVFs offer scan performance similar to standard query operations. Their advantage is immediate, flexible data access – ideal for exploration and analytics where the schema validation and indexing of formal ingestion can be implemented later if needed.
* **Direct Data Access**: TVFs read directly from external storage without ingestion delays.
* Note that direct reads can be slower than querying indexed tables. For repeated analytical queries on static data, consider ingesting into managed storage.
* **Efficient Schema Inference**: We infer schema through a two-step process: First, we download the initial file to detect structure. Then, we convert its columns to Firebolt columns.
* The conversion step adds minimal latency (only \~70 for a 50 MB file with 10 columns in our tests).
* For most exploratory workloads, this additional download is acceptable. If you need to optimize for minimal latency and networking calls, consider ingesting using another ingestion approach with predefined schemas.
### Under the Hood: Schema Inference and Listing [#under-the-hood-schema-inference-and-listing]
To make it easy to explore your data, Firebolt does quite a bit of heavy lifting. This section describes what happens behind the scenes when you call `read_csv()` or `read_parquet()`.
Before working with your data, Firebolt needs to understand its structure. This is important to validate your SQL query and make sure it's actually well formed. Let's look at some examples:
```sql
-- Firebolt needs to be able to detect whether the column names
-- reference valid names in the underlying files.
-- This query throws an exception: Line 1, Column 8: Column
-- 'this_column_does_not_exist' does not exist.
SELECT this_column_does_not_exist
FROM
read_parquet('s3://firebolt-publishing-public/help_center_assets/firebolt_sample_dataset/rankings/*.parquet');
```
```sql
-- Even when you mention valid column names, you might use operations
-- that aren't well defined on the underlying data types. For example,
-- you should not be able to call text functions on integer columns.
-- This query throws an exception: Line 1, Column 8: function signature
-- 'upper(bigint)' not found, supported signatures are upper(text)
SELECT upper("PlayerID")
FROM
read_parquet('s3://firebolt-publishing-public/help_center_assets/firebolt_sample_dataset/rankings/*.parquet')
```
To perform this validation, Firebolt needs to understand the underlying structure of your data. First, we execute a file discovery process at the specified URL. File listing runs on a single coordinator node, so we've created optimizations to make this process fast:
* Firebolt uses recursive prefix iteration to traverse your storage, breaking down large directories into manageable parallel operations.
* When you apply metadata filters, we exclude any unneeded directory paths from the file listing.
Next, Firebolt's data reader examines the most recent file in your dataset for schema inference. This file typically represents your current structure and allows us to quickly infer schema across multi-file queries, while still giving predictable results. Firebolt uses the Apache Arrow library to parse and infer the file's types, then converts these Arrow types into Firebolt's native type system.
What happens if the inferred schema isn't exactly what you need? No problem. You can either create a destination table with your preferred types (allowing automatic casting during insertion) or leverage external tables for more control.
Once the above steps are completed, Firebolt's planner has validated that your SQL query is well-formed. We can now begin crunching data. To do this as efficiently as possible, we want to use all nodes in your engines and all cores on every node.
Once we've cataloged all files, Firebolt distributes them across your engine's nodes. Our algorithm minimizes the maximum data volume any one node must process, ensuring no node becomes a bottleneck. After the files are distributed, every node can work on individual files in isolation. Each node uses multi-threading to keep all CPU cores busy. By default, we assign different files to different cores. For Parquet, we get even smarter: we can have multiple cores work together on reading a single Parquet file. We distribute work at the granularity of row groups within the Parquet files. This way, even a single large Parquet file can properly utilize all available resources.
To demonstrate scale-out performance, we ran the following query on Type S engines with 1 cluster and varying node counts. The query scans about 130 Gigabytes:
```sql
SELECT hash_agg(*) FROM read_parquet('s3://firebolt-publishing-public/help_center_assets/firebolt_sample_dataset/playstats/*.parquet')
```
| Number of Nodes | 1 | 2 | 4 | 8 |
| --------------- | ------------ | ------------ | ------------ | ----------- |
| Duration | 63.0 seconds | 29.8 seconds | 16.0 seconds | 9.2 seconds |
The improvement is near linear, with only a small constant time used for planning, file listing, and coordination.
## Summary [#summary]
The read TVFs allow you to query external data directly. They automatically detect data structure, support Parquet and CSV formats, and integrate seamlessly with Firebolt's SQL functionality. This capability lets you explore external datasets without full ingestion, saving time and resources during data discovery.
Explore data directly from your data lake using Firebolt's `read_parquet()` and `read_csv()` TVFs — no ingestion required.
If you're new to Firebolt, you can [sign up for a free trial](https://go.firebolt.io/signup) here.
# Firebolt ARM Rollout (/blog/firebolt-arm-rollout)
## Introduction [#introduction]
If you're using Firebolt's Compute Optimized compute clusters, you'll see a performance boost. Firebolt has migrated the Compute Optimized compute cluster family from AWS [c6id](https://aws.amazon.com/ec2/instance-types/c6i/) instances to [c7gd](https://aws.amazon.com/ec2/instance-types/c7g/) [Graviton instances](https://aws.amazon.com/ec2/graviton/).
The new instances offer better CPU performance, DDR5 memory ([50% higher memory bandwidth](https://aws.amazon.com/ec2/instance-types/c7g/)), and higher [network bandwidth](https://docs.aws.amazon.com/ec2/latest/instancetypes/co.html#co_network). The migration is transparent to users, your compute clusters will automatically upgrade during normal operation with no downtime (see "Migration process" below).
This post covers how these improvements translate into real-world results, why Firebolt started with Compute Optimized compute clusters, and the technical challenges involved in the migration.
## Graviton Performance [#graviton-performance]
Migrating to Graviton requires careful validation to confirm real-world gains before committing infrastructure resources. Firebolt used two approaches: (1) replays of actual customer workloads and (2) performance benchmarks.
For real workloads, Firebolt uses "customer workload replays." A compute cluster is created with two clusters, one on x86 and one on Graviton, and users' historic queries are run on their actual data (DQL only). The same query ID is assigned to both runs to enable head-to-head performance comparison.
Using historic queries, Firebolt ran tests across the entire fleet and confirmed improvements over the x86 instances currently in use. Tripledot is a great example: they're heavy users of Compute Optimized compute clusters, and their workload showed consistent improvements.
.png)
Firebolt also measured performance using a benchmark. TPC-H with a scaling factor of 100 was run on an M-sized Compute Optimized cluster, comparing x86 and Graviton head-to-head. As the chart below shows, some queries improved significantly while others moved modestly. Even when a query regressed, it was minimal (within 2%), while many saw substantial improvements.

## Why not all Compute Cluster families? [#why-not-all-compute-cluster-families]
Firebolt aims to deliver best-in-class performance across all compute cluster families, but real-world constraints apply. AWS availability varies by instance family and region, and some Graviton types may not be available at all in certain regions. If there is high demand and low availability for specific instance families, there is a risk that AWS may not be able to provide the necessary compute resources.
There are two specific concerns:
**Slower compute cluster starts:** Firebolt maintains warm pools, idle instances kept on hand to immediately assign to workloads when you start compute clusters. If AWS can't supply new instances quickly enough, Firebolt can't replenish the warm pool after engines start.
**Start failures under shortages:** If a warm pool is exhausted, Firebolt falls back to on-demand instances. During shortages, your compute clusters could fail to start entirely if AWS cannot provide them.
To ensure a reliable experience, Firebolt deliberately chose to only migrate families to instance types with sufficient availability and performance. This limits the use of old generation hardware (high availability but low performance) and certain new generation hardware (high performance but low availability).
Compute Optimized is the only compute cluster family where both availability and performance requirements are met for migration. Firebolt will migrate Storage and Memory Optimized compute clusters once availability improves for the newer hardware.
## Challenges Moving to ARM [#challenges-moving-to-arm]
Getting a binary running on Graviton sounds as simple as compiling for ARM, right? For performance-critical applications, the reality is more complex.
### Physical Cores vs vCPU [#physical-cores-vs-vcpu]
Both c6id.2xlarge and c7gd.2xlarge list 8 vCPUs in AWS specifications. However, c6id.2xlarge achieves this by hyperthreading 4 physical cores into 8 vCPUs, while c7gd.2xlarge has 8 physical cores that directly translate to 8 vCPU.
This changes how work contends for resources and can cause increased lock contention.
### SSE2 Intrinsics [#sse2-intrinsics]
```shell
static_assert(CACHE_LINE_SIZE % sizeof(__m128i) == 0, "CACHE_LINE_SIZE must be multiple of __m128i");
```
\_\_m128i is the 128-bit integer vector type used by SSE2 intrinsics, so it won't compile on non-x86 processors at all (SSE2 is an x86/x86-64 SIMD instruction set).
### Floating Point Differences [#floating-point-differences]
Floating point calculations are inexact by definition, and Intel and ARM CPUs can yield slightly different but equally valid results. Any correctness testing has to account for that, while still ensuring that results are valid.
```shell
Casting from real literals to decimals gives different results in ARM and x86_64:
3e28+999999999999999999 becomes either:
- 30000000001000002375798226943.999999160
- 30000000001000002375504989174.919199928
7e28+999999999999999999 becomes either:
- 70000000001000000708277174272.000000430
- 70000000001000000708242972357.100888494
`exp` gives different results in ARM and x86_64:
exp(6.5) becomes either:
- 665.1416330449221
- 665.1416330443618
exp(5.1) becomes either:
- 164.0219073000141
- 164.0219072999017
`cbrt` gives different results in ARM and x86_64:
cbrt(2435.5) becomes either:
- 13.454349693900086
- 13.454349693900085
```
### SIMD Path Fallback [#simd-path-fallback]
In the x86 code, Firebolt sometimes uses special SIMD paths with fallbacks to standard C++. The special SIMD path performs specific instructions very quickly. The fallback should produce the same result but isn't as optimized. However, during this testing we found some non-optimised edge cases on fallback, never previously used, that would have produced different results, which allowed us to improve the code.
## How Did We Migrate [#how-did-we-migrate]
The goal was simple: a non-disruptive, transparent upgrade. If all you notice is that queries start running faster mid-workload, the migration succeeded.
Firebolt uses the same mechanism relied on for [online upgrades](https://www.firebolt.io/blog/engines-online-scaling-and-upgrades#online-upgrade), with one difference: instead of switching to a new binary version, the underlying instance types are switched.
Online upgrades work by creating new clusters to replace old ones. Within a cluster, parameters can be changed, including instance types. The upgrade flow:
1. Provision warm pools for the new instance types (running double capacity during the migration).
2. Create shadow clusters on the new instance types for each running cluster.
3. Mirror all traffic to the shadow clusters.
4. Wait for shadow cluster performance to match the main cluster by warming caches (via mirrored queries and [automatic cache warmup](https://www.firebolt.io/blog/automatic-cache-warmup)).
5. Cut traffic over to the shadow clusters so your workload continues on the new architecture.
Right after the cutover, your compute clusters are upgraded in-flight without disruption, and new queries get an immediate performance boost.
# Firebolt Auror (/blog/firebolt-auror)
## Preamble [#preamble]
In modern cloud-native environments, ensuring only trusted container images run in production is crucial for security, compliance, and reliability. But enforcing strict image verification often comes with challenges like high latency, unpredictable failures, and limited visibility, especially when relying on third-party tools. At Firebolt, we faced these exact challenges. This blog walks you through our journey of building [Auror](https://github.com/firebolt-db/firebolt-auror), an open source high-performance, custom Kubernetes Admission Webhook that validates container images efficiently and transparently.
Until recently, we relied on a third-party tool to allow only signed images to run in our clusters but it presented two key challenges: unpredictable latency especially for cross-region Elastic Container Registry (ECR) calls and lack of visibility in terms of metrics. As it was a third-party tool, we did not have much control over it.
To address these issues, we built a custom image admission controller that allows only images from our ECR with valid signatures to run in production. We will dive deep into how we replaced the third-party image admission controller with our own custom solution which allowed us to run only images from our ECR with a valid signature. Then we'll dive into how our caching mechanism mitigated the high latency during cross-region calls and resulted in sub-millisecond on cache hits, and how it allowed us to see custom metrics such as external images (coming from outside of our ECR) and cache hits/misses.
## What Is a Kubernetes Admission Webhook? [#what-is-a-kubernetes-admission-webhook]
In Kubernetes, an [**admission webhook**](https://kubernetes.io/docs/reference/access-authn-authz/extensible-admission-controllers/#what-are-admission-webhooks) is an HTTPS endpoint that the API Server contacts to determine what to do with objects when they are created or updated. A webhook is registered by creating an *admission webhook configuration (via Kubernetes resources such as MutatingWebhookConfiguration and ValidatingWebhookConfiguration)*, which instructs Kubernetes on which API operations (e.g., CREATE, UPDATE) and which resources (e.g., Pods, Deployments) should trigger a request to the webhook.
One of the important fields to configure is *failurePolicy*. It specifies how to handle unrecognized errors or timeouts from the admission webhook. It can have two values: **Ignore** and **Fail**.
**Ignore** means that if the webhook call results in an error or timeout, the request is allowed to proceed. **Fail** means that in the case of an error or timeout, the request is denied. By default, *failurePolicy* is set to **Fail**.
Those API Requests are wrapped inside the AdmissionReview which is a standardized API object used in Kubernetes to structure requests and responses between the Kubernetes API server and admission webhooks. They are received from the webhook as POST requests. According to the type of the admission webhook, the webhook can inspect or modify the object and return an Admission Review response, denying or approving the object.
There are two main types of admission webhooks:
* **ValidatingWebhook**: Reads the object creation request and either approves or denies it. Ideal for checks like "Is this image signed?"
* **MutatingWebhook:** Can modify or add fields to the object creation request, such as labels or annotations, before denying or approving.
## Why Our Own Image Admission Webhook? [#why-our-own-image-admission-webhook]
While several existing solutions are available, developing an in-house alternative requires substantial time and resources. You need solid reasons to abandon third-party tools and create an internal custom solution. We had our reasons:
### Latency Spikes & Timeouts [#latency-spikes--timeouts]
With our third-party tool, we encountered timeouts especially during cross-region ECR calls. Same-region calls often took hundreds of milliseconds, and cross-region calls could even time out under load up to 8 seconds. We had no way to tune our third-party tool to address these latency issues.
### Lack of Metrics [#lack-of-metrics]
Due to a lack of reported metrics, we couldn't track how many images passed or failed validation. Also we could not track the data regarding external images.
### Desire for Full Control [#desire-for-full-control]
We wanted to own the validation flow, integrate rich metrics, and optimize for our exact use-cases. As our main goal was to verify images based on signatures, our solution was more lightweight.
## Crafting Auror: A Kubernetes Admission Webhook [#crafting-auror-a-kubernetes-admission-webhook]
Let's move beyond the abstract parts and jump into technical details which made Auror thrive. The main purpose of Auror is to allow only signed images hosted in our ECR to run in our clusters.
Auror is implemented as a **ValidatingWebhook.** It performs two checks when an image admission request comes from the Kubernetes API Server:
1. **Registry check:** Verify the image comes from our ECR
2. **Signature check:** Confirms the image contains a valid signature created with [Cosign](https://github.com/sigstore/cosign) (an open-source container signing tool).
If either check fails, the pod is denied; otherwise, it is approved. By checking external images in the first place, we reduce unnecessary calls to ECR and reject those images directly since we know they do not exist in our ECR. Afterwards we add them into the metrics to track them. Only non-external images that do not exist in our local cache trigger ECR calls.
```go
if len(externalImages) > 0 {
if kind := admissionReview.Request.Kind.Kind; kind == "Deployment" || kind == "StatefulSet" || kind == "DaemonSet" || kind == "CronJob"
{
metrics.RecordExternalImage(ctx,
admissionReview.Request.Namespace,
admissionReview.Request.Kind.Kind,
admissionReview.Request.Name)
}
v.logger.Error("Found external images that will not be validated",
"images", externalImages, "namespace", admissionReview.Request.Namespace, "name", admissionReview.Request.Name, "kind", admissionReview.Request.Kind.Kind)
v.handleFailedVerification(w, &admissionReview, "", fmt.Sprintf("%v", externalImages), fmt.Errorf("Found external images that will not be validated"))
return
}
```
To avoid redundant ECR lookups for already-validated images, we adopted a three-tier in-memory cache mechanism utilizing [Ristretto](https://github.com/hypermodeinc/ristretto) (an open-source cache library):
* **Tag Cache**: for short lived tag lookups. Speed verification for a short duration.
* **Digest Cache**: for long-lived, signature verified digests. Best approach, we never doubt if the tag is the same but the digest is different.
* **Owner Cache:** for dependent objects of the owner resource. In Kubernetes, some objects are owners of other objects. These owner objects carry the same *uid* value in their *metadata.ownerReferences.uid* field as their owned objects. As we cache the UID of the owner resource, subsequent requests from the owned objects are validated directly, as they carry the same UID. This effectively speeds up the validation process.
## How to Be Sure We Do Not Break Things? [#how-to-be-sure-we-do-not-break-things]
To ensure that our new Auror admission flow won't interfere with workloads in production, we had to first test our solution. We followed a rollout plan to achieve it. Let's have a look at those steps:
1. After our initial tests in our local kind cluster, we deployed it into our staging cluster without enabling the ValidatingAdmissionWebhook. To make our tests, we utilized our internal CLI tool "Officer" which we used to send mock admission review requests to our webhook. We ensured that it validates images correctly.
2. We enabled the *ValidatingAdmissionPolicy* by adding it as an argument in audit mode through Auror's Helm chart's values.yaml file. External images and images without a signature were caught and printed as warning log messages. Resource creation requests for resources containing a [podSpec](https://kubernetes.io/docs/reference/kubernetes-api/workload-resources/pod-v1/#PodSpec) (e. g. Pods, Deployments) were still allowed.
```go
// If audit mode
if v.mode == "audit" {
warningMessage := "WARNING: Allowing " + review.Request.Kind.Kind + " creation in audit mode: " + message
v.logger.Info(warningMessage)
v.sendResponse(w, review, true, warningMessage)
return
}
// If deny mode
v.logger.Info(message)
v.sendResponse(w, review, false, message)
```
3. After running Auror in Audit mode in production for over two months, we reviewed the findings and communicated with other teams to either sign unsigned images or carry external images into our ECR. With these issues being addressed, we have plans to transition from Audit mode to Deny mode where unsigned or externally hosted images will be strictly blocked.
## Observability [#observability]
As it is our in-house solution, we log details about the external images, the unsigned images, and the signed images. Via our logging system, we catch those logs and create alerts for the following scenarios:
* **External images.** We can check their CVEs, image scan results and source code if possible. Then migrate to our ECR after validation and sign them.
* **Unsigned images in our ECR.** These alerts help us identify internal images that haven't been signed. We sign them and let them run securely in our pods.
* **Signature mismatches.** These may indicate a CI/CD pipeline bug or a potential supply chain attack. This way we can intervene early and investigate the source.
We also emit Prometheus metrics for external, cache hit/miss ratios, and for latency.
As one of the exposed metrics, *cache\_entries* tracks the current number of items in each cache tier (digest, owner, tag). It is used to monitor cache utilization to adjust *cache\_size.*
```text
HELP: cache_entries Number of entries in each cache
TYPE: cache_entries gauge
cache_entries{cache_type="digest"} 2
cache_entries{cache_type="owner"} 17
cache_entries{cache_type="tag"} 17
```
*external\_images\_total* represents the total number of admission requests for images coming from outside our ECR. In Grafana dashboards, we can track the external images with their namespace and name observed.
```text
HELP: external_images_total Total number of external images encountered
TYPE: external_images_total counter
external_images_total{kind="Deployment",name="e2e-mock",namespace="test-1} 1
```
The third-party tool that we relied on lacked the necessary Prometheus metrics which prevented a meaningful performance comparison with our in-house solution. After addressing its core limitations — such as cross-region timeouts and high admission latency — by developing our admission controller, we benchmarked it against the most widely used open-source admission validation webhook [Kyverno](https://kyverno.io/). Let's have a look at the reasons why we chose to go with our in-house solution:
**Performance:** Auror's two-tier caching consistently delivers faster response times than Kyverno, with average lookups typically well under 30 ms and P95 latencies staying below 400 ms. In contrast, Kyverno's P95 response times frequently exceed 600 ms and can reach up to 1 second under load.
**Resource Usage:** Kyverno's broader policy engine uses more than three times **CPU** and memory compared to Auror's image-validation logic.

**Control & Cache Logic:** We maintain full, low-level control over cache behavior (tag vs. digest) and never mutate resources by default. Kyverno, by contrast, is driven by policies that can mutate or annotate objects and only offers partial caching in Audit mode.
**Extensibility:** Kyverno shines when you need mutation, generation, or rich policy types beyond images. Auror is purpose-built for signature enforcement—simpler, but less general.


When we look into P95 dashboards, we see that worst case peaks are much higher for Kyverno. In general, those results made it clear that we better go with our in-house solution.
## Was It Worth The Hustle? [#was-it-worth-the-hustle]
The only way to understand this is to talk about metrics directly. Average latency is shown below:

And the P95 graph which shows the worst case peaks where it needs to make a call to ECR:

The average response times dropped significantly due to our cache. We rarely observe peaks in the P95 graph due to calls to ECR. To mitigate those peaks we came up with two approaches:
1. We plan to warm-up the cache of our webhook through mock requests coming from the officer (CLI client). As they are not API server requests, they will be used to cache the most frequently requested images\*\*.\*\*
2. As Kubernetes nodes are rotated during routine cluster maintenance, the Auror pods are evicted and the cache is lost. This adds additional latency as we need to fill up the cache again. To mitigate this, we plan to write the caches into the disk periodically, so they will be used again when the node is restarted.
## Conclusion: [#conclusion]
In this post, we have shown how we went from a high-latency, prone-to-timeout, black-box third-party Image Admission Webhook to our rich with metrics, highly observable in-house solution. The result is faster image validation, zero cross-region timeouts since we deployed it into production, and full visibility into accepted, rejected, and external images via Prometheus metrics. With Auror in place, our clusters stay fast, reliable, and secure!
# Firebolt’s zero-copy clone (/blog/firebolts-zero-copy-clone)
Imagine you're managing a massive production table and need to test some changes or run experiments on the data. What's your first thought? Copying the entire table might seem like the obvious solution—but it's slow, expensive, and inefficient, especially if you need to do it repeatedly.
That's where Firebolt's new **zero-copy clone** feature comes in. This capability lets you create an instant duplicate of your table without the heavy costs or delays of traditional copying.
In this article, we'll dive into the technical details of zero-copy cloning and the underlying Firebolt's metadata layer and the design considerations that make it possible.
### What is Zero-Copy Cloning? [#what-is-zero-copy-cloning]
Zero-copy cloning is a powerful feature that lets you create an exact, metadata-level copy of a table almost instantly, without incurring into additional storage costs. The cloned table is fully independent from its source, allowing you to query, test, and experiment as if it were a separate table, all while sharing the same underlying data.
Think of it like the concept of "copy-on-write": when you create a clone the data itself is not duplicated. Instead, both the original and cloned tables reference the same physical data files in object storage. This approach ensures lightning-fast table creation and eliminates the need for costly data duplication.
Since zero-copy cloning is purely a metadata operation, it's much more efficient than cloning the actual data —it duplicates only the pointers to the data files, avoiding any heavy lifting typically involved in storage operations.
In Firebolt, cloning a table is as straightforward as running a single command:
```sql
CREATE DATABASE ultra_fast_staging;
CREATE TABLE ultra_fast_staging.public.playstats_new
CLONE ultra_fast_gaming.public.playstats
```
### Common Use Cases for Zero-Copy Cloning [#common-use-cases-for-zero-copy-cloning]
Zero-copy cloning is a practical feature with several applications, including:
* Creating Snapshots: Save a copy of a table as it exists at a specific point in time for record-keeping or compliance purposes.
* Maintaining Historical Snapshots: Take periodic snapshots (e.g., daily, monthly, or yearly) to preserve a historical view of your data.
* Backup and Recovery: Create a backup of production data before making changes, allowing for quick restoration if needed.
* Testing and Staging: Set up testing or staging environments quickly using cloned data, reducing the time and effort required.
### Background on Storing Data and Metadata [#background-on-storing-data-and-metadata]
Managing metadata in a distributed cloud data warehouse is challenging because multiple independent compute resources can read and write simultaneously. Firebolt's metadata layer ensures ACID compliance at the Snapshot Isolation level, providing robust transaction guarantees in such an environment.
Firebolt's metadata layer tracks all objects within its object model, including schemas, compute engines, RBAC privileges, and tablets.
A **tablet** is a collection of rows stored in Firebolt's indexed, columnar file format on S3. The metadata layer references these S3 objects so the storage layer can efficiently locate and read the appropriate tablets at the right time.
Tablets consist of data-only objects without any direct association to logical objects like tables, schemas, or databases. It is the responsibility of the metadata layer to map these tablets to their respective tables and track their evolution over time.
### Why Do We Need a Metadata Indirection Layer? [#why-do-we-need-a-metadata-indirection-layer]
Why can't we just interact with the objects on S3 directly? Couldn't we periodically list objects in a bucket to determine what data belongs to a table?
This approach introduces significant challenges:
1. **Latency**: Listing S3 objects is relatively slow, taking tens to hundreds of milliseconds. This delay would impact query performance if applied to every operation.
2. **Consistency**: Periodic listing would make it difficult to handle concurrent writes. Data modifications by other writers wouldn't be visible immediately, and ensuring proper conflict resolution between concurrent writers would become increasingly complex.
The metadata indirection layer addresses these issues by efficiently managing collections of objects like tablets. It ensures low latency and supports multi-writer consistency, enabling Firebolt to maintain performance and reliability.
Additionally, this metadata architecture supports advanced data warehouse features, such as table cloning, which we will explore further in this post.
### Tablets metadata [#tablets-metadata]
Tables in Firebolt are composed of multiple tablets—potentially tens of thousands. Each tablet is represented in metadata by an identifier and an address, which serves as a pointer to its corresponding objects on S3.
This tablet metadata is lightweight, while tablets on S3 can be as large as 20 gigabytes, the tablet metadata in Firebolt is only about 100 bytes.
This makes metadata operations to be executed very quickly. This efficiency is a fundamental aspect of Firebolt's metadata layer: working with compact metadata handles is significantly faster and more efficient than directly handling huge S3 objects.
Before zero-copy cloning, Firebolt maintained a strict one-to-one mapping between metadata handles and data objects on S3. With the introduction of zero-copy cloning, this mapping has been relaxed, allowing multiple metadata handles to reference the same S3 data objects.
### What Happens During a Clone Operation? [#what-happens-during-a-clone-operation]
When a clone operation is executed, a new metadata entry is created that points to the same tablets as the original source. This means that the physical data itself is not duplicated. The tablet doesn't contain any information about its table or database, so by duplicating only the table metadata, we effectively store the same logical reference multiple times.
Since the actual tablet is shared between the tables, the underlying data isn't replicated, only the metadata is.
As noted earlier, metadata is much lighter than the actual data. This makes copying tablet metadata far more efficient and lightweight than copying the physical data itself.
### Tables are independent after cloning [#tables-are-independent-after-cloning]
After a table is cloned, the original and the clone evolve independently. Any changes, such as updates, inserts, or deletes, to one table do not impact the other. This ensures that each table maintains its own state, even for the data that was initially shared.
When new data is inserted into either table, a new tablet is created in S3 along with a corresponding metadata entry. This metadata entry is exclusively associated with the table where the insert operation was performed.
For deletes or updates, directly modifying the shared tablet would impact other tables that reference it. To prevent this, tablets are designed to be immutable. This also plays nicely with S3 which doesn't allow in-place modification of objects. Instead, Firebolt maintains a change log for each tablet. This log allows each table's metadata to reference a different unique deletion log, ensuring that every table preserves its own distinct state, even when sharing the same underlying tablet.
### Cross-Database Clone [#cross-database-clone]
Tablet data files are not tied to specific logical objects like tables or databases. This allows the references of a tablet to be copied from one database to another, while the metadata still points to the same shared storage.
Zero-copy cloning is Firebolt's first cross-database operation, enabling data transfer across databases without the need for exporting and re-importing. Cross-database cloning also simplifies setting up dedicated databases for testing or staging environments, which we plan to support in an even simpler way through a CLONE DATABASE command.
Implementing cross-database cloning is significantly simpler than supporting full cross-database queries, as it involves only metadata operations without the complexities of query planning.
### When to Delete Objects on S3? [#when-to-delete-objects-on-s3]
When two tables reference the same tablet in storage, what happens if one table is dropped? Rest assured, your data is safe.
Tablets in storage are not immediately deleted when a table is dropped or data is removed. Instead, they are retained for future cleanup by the storage garbage collector.
For zero-copy cloning, while data isn't kept indefinitely, the garbage collector periodically checks for active metadata references against the S3 storage. Any tablet referenced even once is preserved, preventing accidental deletions.
In summary, dropping a table does not trigger storage cleanup that could affect another active clone of the same data.
### Performance [#performance]
We already covered that Zero-copy cloning is extremely fast because only table metadata is copied. The operation's complexity depends not on the table's size but on the number of tablets it contains at the time of cloning.
For perspective, tables can consist of thousands to tens of thousands of tablets, each with metadata of about 100 bytes. [In our FireEdge demo](https://youtu.be/m_PgdHLc65s?si=g37EFRt93hScrMTW), a 60TB table with \~8,000 tablets generated metadata totaling less than one megabyte.
In our first development version, tablet metadata was written one at a time, requiring multiple round trips to our metadata server. While faster than re-ingesting a table, it wasn't ideal. The first version of zero-copy cloning we shipped batches metadata writes, significantly improving latency. Though there's still room for optimization, the current speed allows us to clone huge tables very quickly. In the demo linked above, cloning the 60TB takes about four seconds.
This figure demonstrates that the time required for a clone operation scales linearly with the number of tablets in the table. It also highlights the significant performance improvement achieved by batching metadata writes, reducing the overall cloning time.
### Clone Storage Costs [#clone-storage-costs]
Zero-copy cloning eliminates the need for physical data duplication, allowing tables to share the same underlying storage without consuming additional space. This approach significantly reduces storage costs.
Firebolt's pricing model separates compute and storage costs, meaning multiple tables sharing the same storage directly lowers storage expenses. For instance, as of today, the storage price in the us-east-1 region is $23/TB per month. Cloning a 60TB table instead of re-ingesting it can save approximately $46 per day. Check out [Firebolt's Pricing Calculator](https://www.firebolt.io/pricing#pricing-calc)
Storage billing and usage details can be monitored through Firebolt's information\_schema views, such as [storage\_billing](https://docs.firebolt.io/sql_reference/information-schema/storage-billing.html), [storage\_metering\_history](https://docs.firebolt.io/sql_reference/information-schema/storage-metering-history.html), and [storage\_history](https://docs.firebolt.io/sql_reference/information-schema/storage-history.html).
### Conclusion [#conclusion]
Firebolt's zero-copy clone feature offers a cost-efficient and fast solution for cloning tables or databases without incurring additional storage costs.
This blog explores the technical design of the metadata and storage layers, which form the core infrastructure enabling zero-copy cloning.
Additionally, we discuss performance benchmarks with practical examples and highlight tools designed to enhance visibility into consumption and storage costs.
# Firing Up Firebolt’s Client Ecosystem (/blog/firing-up-firebolts-client-ecosystem)
## Asynchronous query execution [#asynchronous-query-execution]
Asynchronous query execution refers to starting a long running query without holding up the connection in client, whilst receiving a token to reference said query to be able to see it's status in time. This features enables users to not use precious resources on just maintaining a connection when in fact their client is not doing anything.
Applications can't always wait for a submitted query to finish running before moving on. Imagine performing a bulk insert that will take minutes, and your entire logic flow grinds to a halt while waiting for that to complete. Not a good pattern, right? Asynchronous query execution allows you to avoid this problem, submit a query, leave the execution wholly to Firebolt, and move on. This has landed in all of Firebolt's client drivers, ensuring that your client, application, or hardware can get back to tackling tasks while you put Firebolt to work.
How does it work? In the example below, we assume that the Firebolt driver runs in your VPC on an EC2 instance on AWS (it can be on any other cloud or on prem):
Without async execution:

During the time the Firebolt executes the query, in this instance we assume it takes 1 hour to finish, your resource (EC2 instance) will stay up and running, costing you money even though it is not doing anything useful: it is just waiting for the query to finish.
With async execution:

When executing a query asynchronously, the client will get back a tokenId that can be then used to:
* check the status of the initial query
* cancel the initial query (only if the query has not finished by then)
Notice that you can stop your EC2 instance for the duration of the time it takes Firebolt to actually run the query, and then you can bring it up to check the status and continue whatever work you were planning to do.
This approach not only frees up your application's resources, but also introduces real infrastructure cost savings. By decoupling query execution from your client's uptime, you avoid paying for idle compute — which is particularly impactful when running large-scale insertions or transformations that may take tens of minutes or more. You can integrate this flow into an automated lifecycle where your EC2 instance submits the query, shuts down, and then is reactivated via a scheduled job or event trigger to resume downstream processing once Firebolt has completed the work, or once it's time to submit another query.
The diagrams above illustrate both approaches: the first shows how your EC2 instance stays active (and incurring cost) throughout a synchronous query; the second highlights how asynchronous execution lets you pause client-side activity while Firebolt completes the task independently.
Implementing it with our clients isn't very difficult, either:
```java
// Statically typed drivers require the user to cast/unwrap classes to the Firebolt implementation
FireboltConnection connection = DriverManager.getConnection(url, clientId, clientSecret).unwrap(FireboltConnection.class);
try (FireboltStatement statement = connection.createStatement().unwrap(FireboltStatement.class)) {
statement.executeAsync("INSERT INTO table SELECT checksum(*) FROM GENERATE_SERIES(1, 2500000000)"); //long running query
String token = statement.getAsyncToken(); //token from which you can identify the query by
//checking query running status
boolean isRunning = connection.isAsyncQueryRunning(token);
boolean isSuccessful = connection.isAsyncQuerySuccessful(token)
//you can also cancel the query execution if it is still running
connection.cancelAsyncQuery(token); //this should return true if the query was cancelled
}
```
```javascript
// But dynamically typed drivers do not have this issue
const statement = await connection.executeAsync(
"INSERT INTO table SELECT checksum(*) FROM GENERATE_SERIES(1, 2500000000)",
executeQueryOptions,
);
const token = statement.asyncQueryToken; // used to check query status and cancel it and can only be fetched for async query
//checking query running status
const isRunning = await connection.isAsyncQueryRunning(token);
const isSuccessful = await connection.isAsyncQuerySuccessful(token);
//you can also cancel the query execution if it is still running
await connection.cancelAsyncQuery(token);
```
## Streaming query results [#streaming-query-results]
If you're querying massive datasets, sometimes you're also working with large result sets. Not every query has a LIMIT 10 at the end — and if you're reading a million rows or more, this can pose a very real challenge when it comes time to actually *work* with the results. A naive implementation might simply return the entire dataset to the client all at once. That could work for smaller queries, but with large outputs, it's a fast track to running out of memory, triggering crashes, or severely degrading performance.
But with streaming, Firebolt handles things differently. Instead of delivering the full result set in a single payload, it streams rows in manageable chunks. Your application doesn't need to hold the entire dataset in memory — it can start processing as soon as the first rows arrive, consuming them incrementally in a forward-only fashion. This makes it not only safer (no memory overflows), but also faster in practice, since you can begin working with the data before the entire query completes.
This approach is ideal for ETL pipelines, large data exports, or any kind of batch process where the size of the result isn't known ahead of time — and where robustness is just as important as speed. Firebolt drivers and SDKs provide this functionality in a simple manner, keeping the driver-intended way of parsing results, but not actually storing the entire result at once.
Streaming becomes especially important in production environments where reliability and scalability matter. For example, if you're building an analytics API that returns results to end users or pushing query results into a downstream system like S3, Kafka, or a data lake — you're likely dealing with large, unpredictable volumes of data. Without streaming, you'd be forced to manually implement workarounds like pagination or break your queries into smaller chunks, adding complexity and latency. With streaming, Firebolt handles that complexity for you behind the scenes, so your application can simply iterate over rows as they arrive, no special logic required.
What's more, streaming isn't just about protecting against crashes — it's also about *responsiveness*. Because data arrives as it's ready, your application can begin transforming, displaying, or transferring the output without waiting for the full query to complete. That improves time-to-first-byte and makes your system feel faster and more responsive to users — especially valuable in dashboarding tools or real-time reporting interfaces. In short, streaming transforms how your app consumes data: from waiting and hoping, to processing and scaling.
On the topic of error handling, you do lose the ability to know before receiving the results if there were any errors encountered during processing (i.e.: select 1/(i-100000) as a from generate\_series(1,100000) as i), but each driver provides mechanisms for not generating exception vulnerable code even though an error may be encountered in the middle of the result set.
JDBC supports result streaming by default, so the statically typed driver presented here will be .NET:
```csharp
FireboltCommand command = (FireboltCommand)conn.CreateCommand();
command.CommandText = "SELECT * FROM large_table";
// Execute the query without storing the whole result in memory
using var reader = command.ExecuteStreamedQuery();
// or use the asynchronous version
using var reader = await command.ExecuteStreamedQueryAsync();
// Iterate over the streamed results in the same way as with a regular DbDataReader
while (await reader.ReadAsync())
{
//process results with reader.GetValue(x) for example
}
```
```javascript
// JavaScript's implementation leverages its streaming template
const statement = await connection.executeStream(`select 1 from generate_series(1, 2500000)`);
const { data } = await statement.streamResult();
data
.on("meta", (meta) => {
//store/process meta
})
.on("data", (row) => {
//process data
})
.on("error", (err) => {
//capture error
});
```
## Server-side prepared statements [#server-side-prepared-statements]
If your application accepts user input — whether it's a search box, a filter in a dashboard, or a dynamically generated report — chances are you're constructing SQL queries on the fly. And when you're doing that, you're also opening the door to potential vulnerabilities like SQL injection. A common approach to mitigate this is to use parameterized queries, where user values are passed separately from the SQL structure. But even then, *where* and *how* those parameters are handled makes a difference — and that's where server-side prepared statements come in.
With server-side preparation, the Firebolt engine — not your application — is responsible for securely binding parameters. You send the SQL statement with placeholders, and the actual parameter values are passed separately. Firebolt compiles the query and injects the values safely on the server, eliminating the risk of those values being interpreted as part of the SQL logic. This setup adds a layer of security by removing the burden from your application code and protecting against common injection attack vectors.
It's also worth noting that server-side prepared statements don't just improve security — they can also improve performance. Since the query structure is compiled and cached by Firebolt, repeated executions of the same statement with different parameters can reuse the same execution plan. This is especially valuable for multi-tenant apps, dashboards, or reporting tools that generate the same query with slightly different inputs. With fewer parsing and planning steps, your queries run faster and more predictably.
Firebolt also uses this feature to help standardize how parameters are defined across all supported drivers. While most languages and frameworks have their own syntax for prepared statements — like ?, :name, or %s — Firebolt's server-side prepared statements use a numbered placeholder format like $1, $2, and so on. This consistent format helps reduce confusion across languages and environments, making it easier to maintain query templates in a shared codebase. It also improves readability and makes the logic behind parameter ordering explicit.
That said, Firebolt still supports the native client-side style of prepared statements if you prefer — and in fact, it's the default for many drivers. You can continue using the syntax that matches your language of choice, and the driver will substitute parameters before sending the final query to Firebolt. But if you want to take advantage of server-side preparation — for the added security, performance, and standardization — you'll need to update your queries to use the $1, $2, etc. format and enable the appropriate mode in your client.
In short, server-side prepared statements help you build safer, more efficient applications — while encouraging best practices and consistency across teams, drivers, and use cases.
Two different ways of implementing them:
```java
// Besides GO, all languages need one more element in the connection string to make the connection server-side prepared statement ready. In this case we have prepared_statement_param_style=fb_numeric (the native option is the default, but you could specify it in the connection string if you want)
Connection connection = DriverManager.getConnection("jdbc:firebolt:my_db?account=...&prepared_statement_param_style=fb_numeric....", clientId, clientSecret);
try (PreparedStatement statement = connection..prepareStatement("select $1, $2")) {
statement.setInt(1, 2);
statement.setString(2, "foo");
ResultSet rs = statement.getResultSet(); //then process result
}
```
```javascript
//Node.js allows you to make use of both parameters and namedParameters, but not both at the same time
const connection = await firebolt.connect({
...connectionInfo,
preparedStatementParamStyle: "fb_numeric", //important
});
const statement = await connection.execute("select $1, $2", {
parameters: ["foo", 1], //using the parameters executeQueryOptions field
});
const statement = await connection.execute("select $1, $2", {
namedParameters: { $1: "foo", $2: 123 }, //using the namedParameters executeQueryOptions field
});
```
## Conclusion [#conclusion]
Every driver has these features implemented and they are all ready to improve your experience with Firebolt, be it by a resource freeing aspect, a memory holdup one, or just better security. For better explanations please refer to the driver's documentation and if needed, don't hesitate to contact us!
# From Zero to 100M Users: Inside Notion’s Data Stack and AI Strategy with Sumit Gupta (/blog/from-zero-to-100m-users-inside-notions-data-stack-and-ai-strategy-with-sumit-gupta)
In this episode of The Data Engineering Show, the bros talk with Sumit Gupta, Lead BI Engineer at Notion, about his journey through prominent tech companies, modern data stacks, and how AI is revolutionizing data workflows and professional development.
Listen on [Spotify](https://bit.ly/45is6z0) or [Apple Podcasts](https://bit.ly/43Yv4q0)
**Sumit - 00:00:00:**
AI has made me a lot more productive, but at the same time, it has also made me dumber. Whereas Claude is more particularly great at programming tasks. Complexity is amazing at deep research. I'm a lead BI engineer at Notion, and I lead reporting, dashboarding for marketing and sales teams. Right before Notion, I used to work for Snowflake. I worked for them for a couple of years. If you don't jump onto the bandwagon right now, you might be left out in a year or so.
**Intro - 00:00:30:**
The Data Engineering Show is brought to you by Firebolt, the Claude data warehouse for AI apps and low-latency analytics. Get your free credits and start your trial at firebolt.io.
**Benjamin - 00:00:42:**
Hi, everyone, and welcome back to The Data Engineering Show. Today, we're super happy to have Sumit on. Sumit is a lead BI engineer at Notion right now, has been here for a bit more than a year. Great to have you on the show. Do you want to quickly introduce yourself to our listeners?
**Sumit - 00:00:58:**
Yeah, thanks a lot for having me, Benjamin and Benjamin. I've been so excited to jump on this podcast and talk to you guys about everything AI and a bit about me. So as you mentioned, I'm a lead BI engineer at Notion. And I lead basically reporting, dashboarding for marketing and sales team here at Notion. Right before Notion, you probably will see a Snowflake logo on the video.
**Eldad- 00:01:23:**
No, we haven't noticed.
**Sumit - 00:01:26:**
Yes. And then you have like Hello Data Nation t-shirt in here. So right before Notion, I used to work for Snowflake. I worked for them for like a couple of years. And then before that, I was at Dropbox leading their analytics team, marketing analytics team. So overall, like I have great experience being part of like- I like to use this term called Bay Area Darlings. So back in 2017, 18, Dropbox was Bay Area Darling, right? And then Snowflake became Bay Area Darling because of, you know, IPO and etc.. And now if you think about it, Notion is the Bay Area Darling. So I don't know. I have some affections with darlings, I guess.
**Benjamin - 00:02:03:**
Nice. That's awesome. So what got you started in data kind of like in the first place, right? Kind of like tell me, maybe a bit about that journey.
**Sumit - 00:02:13:**
I graduated from Mumbai University in 2014. And then right before my senior year, I was like exploring what my career choices could be. You know, a typical 18 year old where like he's confused, right? And this is before the age of AI. I know, I mean, a lot of folks nowadays have it easy. Where you could just go to ChatGPT and be like, you know what, help me decide my career. But back in 2014 in India, like you were, there wasn't a lot going on in terms of like tech. Like there was obviously like software engineering as a career. But I realized that software engineering isn't for me because I like, I love talking to people. And then I wanted more interaction. And I did not want to be a out and out management consultant either. Because I still wanted, I knew for a fact that tech has to be in my, has to play a very important role in my career. So as I was exploring, I was like, you know what, Information Management is something that I should probably pursue. And then I decided to pursue my master's in that from Syracuse University in the United States. And then the reason United States Information Management is because it has that mix of data management consultant and yet being very close to like tech. So that's, that's how I got introduced to tech. And since then, I've been, I've been in data slash tech for over a decade now.
**Benjamin - 00:03:29:**
Amazing. That's awesome. So tell us maybe about kind of like the, at least at a high level, kind of the work you're doing at Notion right now. Kind of like what, what are you working on? Kind of how's your data stack? Give us an overview.
**Sumit - 00:03:43:**
Yeah. So Notion, Notion is a very, if you think about it, Notion is relatively new. All the Notion was, I think, started in 2014. But the Notion got really popular due to COVID boom, right? I think if, if I, if I recall it correctly, until like 2020, we had close to like about a million or a couple of million users. Now, I don't know if you saw the news or if you saw our co-founder tweet about it. We have over a hundred million users now.
**Benjamin - 00:04:10:**
Wow, congratulations. That's so impressive.
**Sumit - 00:04:13:**
Yeah, yeah.
**Eldad- 00:04:14:**
That's amazing.
**Sumit - 00:04:15:**
It's wild that, you know, the kind of following that we have, B2C as well as both B2B, right? So if you think about it, going from a couple of million users to 100 million users, if your data stack is not nimble or if it's not modern, right? That's the term that we want to use. If it's not modern, you're going to be stuck, right? Your team will not be able to make any relevant decision in time. So that basically means our stack includes like some of the most common modern tools like Snowflake. We use Fivetran for ingesting some of our data. Obviously, we have some custom pipeline set up too. We use Airflow for orchestration. And then we use Tableau as well as Hex for our reporting in Dashboard. I'm an out-and-out Tableau guy. I mean, not to diss on other tools, but Tableau is great, right? One of the reasons I love Tableau is the fact that it has great community.
**Eldad- 00:05:09:**
By the way, you know, if you love Tableau, you know Tableau went through a long history of improving their query engine behind the scenes. And there was an acquisition done by Tableau of a startup, Hyper, and that query processing team is now driving Tableau, Salesforce, Query Stack. It's an amazing team. And the reason I mentioned it is all of them are coming from Munich. And many of- Some of them are at Firebolt. So there's a lot of geeky query processing stuff happening in Munich that drives many, if not most, of today's data warehouses and data solutions. So we have a special place in our heart for Tableau as well.
**Sumit - 00:05:47:**
Yeah, did not know that. But I know for a fact, like, the Hyper Engine that Tableau has, it has changed a lot. Like, you know, nowadays when I import or extract, like, 20, 30 million rows, it goes by really fast, right? Because of Tableau Hyper Engine. So, going back to the topic about modern-year stacks, our BI reporting layer is Hex and Tableau. I don't know if you guys have heard of Hex, but the way I'd like to think of Hex is like Jupyter Notebook, but on steroids. It does everything that you can imagine, right? Python, Spark, SQL, Matplotlib. When you think of Python, Matplotlib, your pivot tables, your charts and everything. So we have that. We also use nowadays, we also have started moving some of our data loads to like Iceberg tables. That's the new thing in the market, right? Like a lot of folks are moving to Iceberg tables because of the benefit that Iceberg offers.
**Benjamin - 00:06:41:**
So are you leveraging kind of the key benefits of Iceberg already in terms of having flexibility between query engines?
**Sumit - 00:06:48:**
Yes. So I primarily don't work with Iceberg, but I know for a fact that our data platform team does. So I'm not an expert to talk about that. But I think the whole point of Iceberg table was like there's a component of saving cost, right? Because with Iceberg, like with the whole metadata layer and how the data moves into the actual like data layer. That's why we decided. Because we use Snowflake as our analytical data warehouse, but we also have like S3 bucket in our data lake sitting in AWS, right? So we wanted to, because if you imagine, as I mentioned, like when you go so fast, your data is, your data needs are growing so fast. Every bit, right, is expensive. Like all the servers are cheap, but when you're dealing with 100 million users and trillions of rows of data a day, you have to find that one person saving, two person saving or those marginal savings. And Iceberg helps us with that.
**Benjamin - 00:07:41:**
Okay, super, super cool. And very modern stack as well. Kind of like using a lot of stuff to cool, kind of like new tech, I mean like Iceberg hacks, kind of all of that. That's all.
**Eldad- 00:07:51:**
At the front of the modern stack, even the modern stack, right? Like you have that starts to evolve as well. We see the modern stack getting a new definitions across. Across each component, it's beautiful to see it. We love Iceberg. Iceberg makes all metadata and data the same for every vendor, for every query engine, every ETL, ELT tool, every user. Yes, there are many challenges. Yes, there are many gaps, and that's why we're all here. And it's amazing to see the evolution happening, Iceberg, within information and analytical teams. Amazing, go on.
**Sumit - 00:08:29:**
Yeah, and then, sorry, I forgot to mention the... Like if you think about modern data stack, there is one tool that brought this into like in the front, right? And that's dbt. We use dbt too. We are a heavy dbt user. You cannot forget dbt because the whole term around analytics engineering and modern data stack, I think they were the one who like promoted this really heavily. Because before dbt, right? So I was at Dropbox and we used to have our own version of like Hive, right? And then there were times when it would take at least like 18 to 20 hours to run a query. Because you can imagine Dropbox is back in 2019, 2020, right? Dropbox is and was huge. We do like a couple of billion dollars of revenue a year. So and then now with dbt, et cetera, right? You could have incremental models and, you know, basically put things in Snowflake, right? Obviously, when you think of Snowflake and Hive, Snowflake obviously will be a lot faster, although there's a cost associated. But sometimes either you invest in time or you invest in dollars. In the Snowflake case, you invest in dollars because you save time.
**Benjamin - 00:09:34:**
In your data stack, like how is AI kind of like shaping it? Like what's changing? Kind of what's your take on all of it? Give us an overview.
**Sumit - 00:09:43:**
Yeah. So we at Notion use AI heavily. We have our own AI for Work tool, which is inbuilt in Notion. I don't know if you guys use Notion or not, but people or companies that use Notion with AI for work, basically everyone in the company is supposed to use AI for Work. It's like deep research. It has all the context about your data. So you can you can connect your Slack. You could connect your Google Drive. You can connect your other like third party apps into Notion now. And then just like let's say if you have a question about when was the last time a specific metric was updated, right? AI for Work will do that for you, like Notion's AI for work will do that for you. And outside of Notion, we are a heavy AI user as a company, like everyone in the EPT department or pretty much in the company has access to Cursor, right? The coding, I wouldn't say agent, but like editor nowadays, right? So we use Cursor a lot nowadays to like help U.S. Speed up our productivity. And then we have been building a couple of like GTM or go-to-market market, I wouldn't say tool, but like use cases. Like one of the use cases that we have built is whenever we are going to talk to our customers, if they are already, especially renewal customers, if they are already, if they are up for renewal, we have this AI agent or AI workflow setup where you select a customer and then it will give you all the details about total number of users, active users and users who have never logged in in XYZ days. Plus like we also ingest our call transcript, our Slack messages that the sales rep have and then create a short summary of what has been discussed in last 30 days, last 90 days. What has, was there any blockers? Was there any pointers that we can use? So that you have one page document where you can go talk to the company and then, you know, try to, because when you go for a renewal conversation, right? If you have data to back it up and be like, you know what, 80% of your company is using Notion. So it's probably prudent or it's probably wise to, you know, renew it with U.S., right? If you were to move to a new tool, there would be a lot of headache, et cetera. So the fact that we could go have that conversation and the reason we could do that now because of all the context is AI is the reason for that.
**Benjamin - 00:11:56:**
Very cool. Okay. That's exciting. How, how do you stay on top of like all of the stuff happening in AI in your personal life as well?
**Sumit - 00:12:03:**
I think my mode of AI consumption has been podcast recently. When I say mode of consumption, I was recently on another podcast called Your Everyday AI Podcast. Jordan does like 20 minutes AI news every day at like every, every morning. So whenever I wake up, I listen to that. And then a lot of keynotes nowadays, right? Like recently Google had a Google I/O, right? And I would watch that and see like what's happening latest and greatest. Claude recently launched, launched Claude 4.0, like, right? So whenever some new model comes along, I am first to jump onto it, right? And basically test it out and see what's working. And you wouldn't believe I have basically premium access to pretty much all the famous LLMs, be it Claude, be it Perplexity, be it ChatGPT, right? And my use case... Every tool is great at something. Like for me, chatGPT is when I want to, let's say, rewrite an email or need some like text written or, you know, want to review some document, etc. Whereas Claude is more particularly great at programming tasks. Right? Perplexity is amazing at deep research.
**Eldad- 00:13:09:**
Who does the cooking? Who is cooking for you?
**Sumit - 00:13:13:**
Yeah. No, I mean, Perplexity, I guess. Because with Perplexity, right? I think Perplexity was the first tool which allowed you to do deep research where you had direct access to the search engine, right? ChatGPT. And Claude, like they recently launched deep research, right? Gemini obviously do, but Perplexity is where you could go and be like, you know what, like one of the use cases, I was recently searching for like universities for my niece that she was applying for, applying to universities in U.S.. And I was like, you know what? This is the score that she has received. This is like, you know, her SAT scores, her other scores. Go help me find 20 universities that has good acceptance rate. When I was applying, that was like one week of effort. Now it's not even seven minutes of effort.
**Benjamin - 00:14:00:**
That's very, very cool. Nice. How do you think the data space as well is going to evolve over the next couple of years? What are you particularly excited about? What problems do you think will start going away? Tell us about that.
**Sumit - 00:14:18:**
I actually like to divide that into two parts. One is like if you are someone who's starting a new in Data Field, right, as an entry-level data scientist, analyst, or engineer, I would say the value of your technical skills that used to be very valuable until 2021 when there was a boom in hiring. Is not as much as, so the value has decreased, right? Your tech skills are valuable, right? But as an entry level, you don't know if something breaks, how to fix it. That's where your transferable skills or your soft skills comes into picture, right? So if you are to grow in data career, make sure that your transferable skills are outshining your technical skills nowadays, right? That's where the entry will part. But if you're a senior, let's say for someone like us, who has like seven, eight, 10 years of experience, right? Your tech skill is important. You know when things break, where to go and how to look and what to fix. But your transferable skills are going to be paramount in future too, right? Especially at Notion and a lot of other companies that I've heard, right? There's a lot of like, a lot of CEOs nowadays are coming out and be like, use AI. If you don't use AI, you're going to be out of job soon, right? So in that case, right? AI, like nowadays I use it whenever I'm building Tableau dashboard, etc. There were times when it took a couple of days to write a calculated field. It was, if it was complicated because Tableau has its own way of calculating, right? Now when I'm stuck, I just go to Claude and like, you know, write me a calculated field and I get the calculated field done in 20 minutes. So the value that I used to bring as a technical guy has diminished. But if that gets diminished, I have to increase the value that I bring as a professional with like great communication skills, right? Great stakeholder management. So you kind of have to build on that. So I don't know if that answer your question because, data is important, but the role of data as a tech skill.
**Eldad- 00:16:11:**
Be technical goes down, be nice goes up. You know, like that's it. It's really fascinating to see kind of here from someone from the inside experiencing the AI like full-blown in your face not from the like not theorizing about it really practicing it turning into your personal career advantage figuring out how to build your strength around it thank you for sharing it with us.
**Sumit - 00:16:39:**
Yeah, no Notion, like, I use AI heavily, uh, in my work too, but in personal life too, as I mentioned like I re- I was recently on Everyday AI Podcast, and the topic that we discussed there was like, AI has made a, made me a lot more productive, but at the same time, it has also made me dumber. Because I have-
**Eldad- 00:17:00:**
Who would have thought? Who would have thought getting dumber gets you more productive? Who would have thought? That defies physics, defies everything we've learned.
**Sumit - 00:17:08:**
Exactly.
**Eldad- 00:17:09:**
30 years of industry experience. I'm too old. I'm like, I've learned, you know, like, Previous generation physics, everything is changing.
**Sumit - 00:17:16:**
Yeah. And then I gave a couple of examples there. And honestly, like after that call, and I've been like trying my best to not like allow AI to drive my life. It's good. AI allows me to research. But I have like a back of hand rule nowadays where like I use AI for repetitive tasks, right? Like I, to give you an example, like I have an Instagram channel by the name Data by Sumit, right? It has around 21,000 followers, et cetera, right? So previously I used to have a social media content strategist and content researcher. We used to like sit down and, you know, research content ideas for me. Now I have a Make.com like workflow where I have identified a list of, let's say, 100 Instagram users, right? Who are, I wouldn't say competitor, but who are in my niche, right? And then I use Appify's API. Also, I've also tested RapidAPI. Appify's API, which basically I feed the users into Appify. Appify goes into, Appify's API basically helps me extracts, let's say, last 30 days of data for each profile. And then I have a GPT-3 model as part of the workflow where it looks at, I have a custom formula where based on number of views, number of likes, number of comments, I create a custom column, like a calculation of high, medium, low, like, should I be writing about this or should I be scripting about this? And once that cutoff is met, the second part of the GPT-3 model goes and transkypes the reel and then learns from it and then comes up with the new script for me.
**Eldad- 00:18:44:**
Wow. This is when information work becomes operations. Like, you're turning, like, your usage of data with those tools, you- This is operations. You're driving a business. You're driving a unit. You're moving crazy. Amazing.
**Sumit - 00:19:02:**
And honestly, that's how AI is supposed to work or supposed to be used. Someone recently asked, do you think AI will take over the world kind of conversation? I'm like, maybe, yes. But for now, I mean, if it does, we don't know. If it does, everyone would be infected, not only me. But right now, AI is helping me improve my workflow. AI is helping improve my, improvise my workflow and get better at it. Like something which took me a week to get like four or five new content ideas. I just have to click a button and get like 50 different new ideas based on what's working right now. The last seven days, last 30 days. And then I get script too. I just have to. Like one of the, so like an inside scoop. One of the things that I'm doing now is like creating another workflow where based on the script, right? I'm using HeyGen and ElevenLabs to basically replicate myself now. My voice and myself, right? So that's the next step in the process where if I'm not there.
**Eldad- 00:20:01:**
Are you real? Are you real?
**Sumit - 00:20:03:**
I am real. I can pinch myself.
**Benjamin - 00:20:06:**
Six fingers. Six fingers.
**Sumit - 00:20:10:**
I mean, I wasn't doing a lot of big motion. And, you know, when you don't do big motion, HeyGen really works. So you guys probably don't know.
**Benjamin - 00:20:22:**
In the beginning, you were like, oh, kind of fiddling with the laptop, fiddling with the phone. But really, you were just preparing your AI self.
**Sumit - 00:20:31:**
Exactly. Exactly. That's it. That's it.
**Eldad- 00:20:34:**
The free trial was over. Enter the credit card, subscribe, activate the agent. Nice. Amazing. Thank you.
**Benjamin - 00:20:43:**
Yeah.
**Sumit - 00:20:44:**
No, I mean, as I said, I think anything which can be automated, especially a repetitive task, AI is the way to go. And honestly, there are a lot of AI content creators nowadays that you can learn from. And a lot of them are actually offering free workflows, Make.com workflows and entertain workflows too. So if you are into AI, if you are jumping into AI, this is the right time. And the scariest part about the whole AI boom is this is the worst AI will ever be. We get impressed when ChatGPT was launched and initially I was a skeptic. New hype. And then I started using it. I'm like, what the hell? Like, how is GPT so good? Like, it doesn't make sense. And now when I think of it, I'm like, this is the worst it will ever be.
**Benjamin - 00:21:31:**
Nice. I also love how like as an AI content creator, you're at like the front of content creation. And we're getting close to AI content creation being done by AI. Like, we're very close to actually like coming full, full circle and just having AI create content about itself.
**Sumit - 00:21:53:**
But there's also morality, moral aspect of this where I'm like, I want to do the HeyGen and ElevenLabs like workflow. But I'm like, if I was following a content creator that I love and if I knew that all he is is an AI nowadays, I would be pissed. So I'm like, I can do it. If I sit for a weekend, I can set the workflow up. But I've been intentionally delaying it because I wrote a book about Tableau, the Tableau workshop, right? It sells really well. But one of the reasons I wrote it because it was like, if I was to read this, I should be happy about it. Like I should learn something, something. And that is how I view my life. If I will only do something that if I was in their shoes, I should be okay with that or happy about it. So if I build my own AI avatar, right, am I going to be happy following that content creator? Probably not.
**Benjamin - 00:22:43:**
Nice. That's, yeah, that's a super interesting take. I think we've had a lot of content creators kind of on the show already, but like this is definitely the most AI forward kind of take I've heard so far.
**Sumit - 00:22:55:**
Yeah, I mean, as I said, like I think AI is great, both in professional as well as personal life. And there's a con for AI too. So my wife's been searching for a software engineering job. She's an entry-level software engineer. It's been hard. Even in Bay Area, getting a software engineering job, it's been hard. I have so many friends who had offers and then their offers got rescinded because company decided, you know what, we don't want entry-level engineers. So as much as I am utilizing AI for the good part, I have the bad part use case also. Because if my wife was searching for a job back in 2021, she would already have the job. So yeah, there's the good part, there's the bad part. It's all about balancing it. And hopefully there's that 51, 49% chance of, you know, you using AI for the good part more than the bad part.
**Benjamin - 00:23:43:**
Cool. Wow. This, yeah, this conversation was really like kind of like twisted my brain, kind of so crazy. Kind of before we wrap up, right? Like anything else you want to share with the audience, with the listeners, kind of anything, yeah, kind of you wanted to chat about?
**Sumit - 00:23:59:**
Like I think... Even at Notion and Postman also, there has been a lot of focus on AI-heavy workflow, right? So if you are someone new, especially in data, scared of AI or still skeptic of AI, I would say jump in. There's a lot of tools, a lot of resources nowadays. And even if you don't want to use it... For your personal or professional work, use it as a research tool or a learning tool where if you notice AI, you should be able to recognize AI. Like my eyes are trained now to see when I look at something, I'm like, okay, this is AI, right? If you are not trained, right? If you're not aware of how AI actually, what are AI's pros and cons? And you know, like when you look at an AI video, it's easy to make out. If you look closely, the eye movement, the hand movement, right? If you're not trained, you might get flummoxed, right? You might get confused, right? So that would be my tip. Like if you are in the outside or if you're outside trying to get in and still like are not happy about AI or skeptic about AI, jump in. You don't have to use AI in your personal or professional life, but... Trust me when I say this, if you don't jump onto the bandwagon right now, you might be left out in a year or so.
**Benjamin - 00:25:11:**
I think that's a great kind of closing remark. Thank you so much for being on the kind of show, Sumit. It was great having you and we look forward to catching up in the future and kind of seeing what fraction of your content starts getting AI generated.
**Sumit - 00:25:27:**
Absolutely. Thanks a lot, Benjamin and Benjamin. If you guys are in the same time, drop me a message and we can grab a coffee or a beer.
**Benjamin - 00:25:34:**
Would be amazing. Thank you. Bye.
**Sumit - 00:25:37:**
Bye.
**Outro - 00:25:39:**
The Data Engineering Show is brought to you by Firebolt, the Claude data warehouse for low-latency analytics. Get $200 credits and start your free trial at firebolt.io.
# Fuzzing Firebolt: Catching 0-days as fast as our query processor (/blog/fuzzing-firebolt-catching-0-days-as-fast-as-our-query-processor)
At [Firebolt](https://www.firebolt.io/), we are committed to delivering great software that's devoid of security bugs.
As a cloud data warehouse company, we want to provide our customers with a secure product that they can use without worrying about the safety of their data workloads. However, in the complex world of software development, maintaining robust security is a relentless challenge. More features mean more code, and therefore potential for more bugs. And to consistently weed them out from a rapidly growing codebase such as ours, relying solely on traditional testing methods like static analysis is insufficient.
**This is where fuzz testing comes in.**
In this part of our security blog series, we'll highlight how we fuzz Firebolt's blazing fast query processor written in modern C++- a fast, yet memory unsafe language.
### What is Fuzzing, and why is it important? [#what-is-fuzzing-and-why-is-it-important]
* Dynamic software testing
* Send semi-random inputs to the target binary (most often compiled with a [sanitizer](https://github.com/google/sanitizers/wiki/AddressSanitizer)), hope for a crash, and then analyze the code where sanitizer reported the crash
* Catches bugs that if exploited, can lead to serious security issues. In our case, this could potentially translate to reading cross-tenant query data from adjacent memory, to infiltration and takeover of Firebolt's customer engines via an RCE
There are tons of great blogs on various fuzzers and techniques, so we will limit the details here to showcase how we employ some of them to hunt for 0-days in our product, and secure our customer workloads.
### Our approach (to Fire-Fuzzing) [#our-approach-to-fire-fuzzing]
Catching bugs is like finding a needle in the haystack, so effective fuzzing starts with picking the right targets. Fuzz targets are typically interfaces that directly receive and parse external user inputs. As a distributed data warehouse, a lot of our code is also heavy on parsing. To establish an effective fuzzing process at Firebolt, we identified those critical parts in our software, and tailored respective strategies around them. The following sections dive into some of those details.
**User input enters Firebolt in two main ways**:
* SQL statements (DQL/DML/DDL)
* Customer data [loaded](https://docs.firebolt.io/godocs/Guides/loading-data/loading-data.html) from S3 (using our COPY FROM / EXTERNAL TABLE features)
The first type of user inputs are processed by the **SQL compiler**, while the second type by various **File formatters** (depending on the file type being ingested, eg. CSV/JSON/Parquet).
### SQL compiler [#sql-compiler]
To fuzz the SQL compiler, we run two different frameworks in parallel:
* [AFL++](https://aflplus.plus/) to primarily fuzz the **frontend SQL parser**
* A custom [SQLsmith](https://github.com/anse1/sqlsmith/tree/master) extension (privately maintained) to **stress the later stages**
For AFL++, we maintain a harness code linked against a LocalExecutor library which encapsulates our in-process SQL query processor, seed the fuzzer with a subset of DQL, DML and DDL statements from our continuous SQL test suites, and use its [persistent mode](https://github.com/AFLplusplus/AFLplusplus/blob/stable/instrumentation/README.persistent_mode.md) to call the target function "**executeQuery()**" with the fuzzer generated strings.
Most of the time the fuzzer generates inputs that the SQL parser rejects due to incorrect grammar- and that's okay. Lexer rules are also written to handle invalid SQL strings, so we still manage to crash the target and find memory bugs from time to time such as **double-free** and **heap OOB reads**, especially in classes that deal with invalid SQL error handling.
With **SQLsmith**, we add another dimension to SQL fuzzing, as it can generate grammatically correct statements that the parser accepts, yet still find bugs deeper inside the query processing stack (**Planner, Runtime**).
The upstream version supports PostgreSQL and other PG compliant databases (e.g. sqlite and MonetDB). Since Firebolt strives to maintain SQL dialect compliant with Postgres, we were able to extend SQLsmith with minimal changes by:
* importing upstream under our private GitHub org
* implementing Firebolt classes to connect and fetch schema (**col names**, **data types**)
* overriding the [dut\_base::test()](https://github.com/anse1/sqlsmith/blob/master/dut.hh#L46) function to send SQLsmith generated queries to Firebolt
* catching the HTTP responses (when not a **200 OK**) with SQLsmith's internal error handler, which it uses internally to track error types to mutate AST's and generate fuzzing stats.
Being PG compliant, we were also quickly able to piggyback on the existing implementations for registering most of our supported [SQL functions](https://docs.firebolt.io/sql_reference/functions-reference/functions-reference.html) with the corresponding args and return types.
SQLsmith works by connecting to the target database, fetching schemas, and recursively creating random ASTs with schema info and registered Firebolt SQL functions/operators while keeping the grammar intact. The AST's are then converted back into their equivalent SQL string representation and sent to the target database for execution. If the response returned is logically inconsistent, or the server crashes, it [records a bug](https://github.com/anse1/sqlsmith/wiki#score-list).
Additionally, in order to test boundary conditions and discover what happens when SQL constructs are pushed to their limits (e.g. longest identifier, maximum number of tables in JOIN), we use the [tensile](https://github.com/themosha/tensilelib) library written by our CTO [Mosha](https://www.linkedin.com/in/mosha/). We have extended the language support in tensile for Firebolt specific SQL constructs such as lambda functions.
### File formatters [#file-formatters]
To fuzz file input formatters, we leverage LLVM's in-process fuzzer (**libFuzzer**) and our existing unit test suite (gtest), as reaching that part of the code by simply hammering the frontend parser won't be possible due to the invalidness of the fuzzer generated strings. For e.g., we took our gtest code for the CSV file formatter class, and re-wrote it into a libFuzzer target. To boost fuzzing efficiency, we employ techniques such as code coverage analysis (gcov), and feed off that information to guide our fuzzing campaigns.

If it's not covering enough code, we study the target further to identify all possible execution paths, and the conditions under which execution flows into each of those edges. Thereafter, we take two important measures:
* **Re-write the harness:** We instantiate different instances of the target class with different configurations (e.g. enabling [ALLOW\_COLUMN\_MISMATCH](https://docs.firebolt.io/sql_reference/commands/data-definition/create-external-table.html#csv-types) in one object and disabling it in the other) to cumulatively cover all the blocks and edges.
* **Refine seed corpus:** We seed libFuzzer with valid CSV data that adheres to each of these configurations, allowing the fuzzer to mutate these inputs more effectively and reach deeper code paths.
We also write the harness in a way that covers all data-types currently supported for that file format and seed the corpus accordingly, so that de-serializers for those data types are also fuzzed in the same campaign.
For complex data-types such as arrays, we use structured fuzzing to make sure that our target only processes inputs that are enclosed within **\[ ]**.
Using the above techniques, we managed to discover a critical **heap OOB write** in our CSV input formatter. It is also worth mentioning that for fuzzing with libFuzzer, we use Debug builds that enable assertion checks ([\_LIBCPP\_DEBUG](https://releases.llvm.org/12.0.0/projects/libcxx/docs/DesignDocs/DebugMode.html)) in the standard library, in addition to instrumenting with address sanitizer. This has helped us catch buffer overflows in code while iterating over STL containers such as vectors and maps that even ASan once failed to detect (we tested this against non-Debug ASan build with the same fuzzer input).
### Reproduction, Triaging and Fix [#reproduction-triaging-and-fix]
When we discover bugs, we reproduce them both locally, and in a secure isolated environment without affecting customer workloads in Production. So far, all our findings have crashed the test engines due to illegal memory operations, thereby re-invigorating our faith in fuzzing as a sure path to shipping secure software.
Afterwards, the security research team conducts a detailed crash analysis to determine the root cause of the bug. Once an RCA is done, we follow a structured process to address the issue with engineering teams through Jira tickets.
Additionally, we publish internal research blogs after each finding to help engineering understand why the bug occurred, how it was debugged, and include suggestions to avoid similar logic pitfalls in future iterations.
In most cases, bugs are fixed within a week of discovery, with patches rolled out in subsequent releases cut-off for the Staging environment, where we re-verify the fixes before finally releasing it to Production.
### Fuzzing at scale, LLMs and beyond [#fuzzing-at-scale-llms-and-beyond]
We continuously improve our processes to discover bugs at a pace that keeps up with our engineering velocity. Some of these efforts include fuzzing the latest master branch in CI **three times a day**, with automated crash detection and uploads to a private S3 bucket at the end of each CI run.
Additionally, we've developed internal tools to automate reproduction with the crash data locally before conducting the RCA, thereby reducing significant time and manual efforts in eliminating false positives.
We are further researching the potential application of LLM in our fuzz tests, such as in the area of corpus generation for fuzzing file formatters.
Every new bug that we discover and eliminate brings us one step closer to our commitment to shipping a secure product. And at Firebolt, we enjoy that challenge every single day.
# Getting Rid of Raw Data with Jens Larsson (/blog/getting-rid-of-raw-data-with-jens-larsson)
Why would you create ugly data? According to Jens Larsson, don't even go near raw data. Jens started off at Google, continued to manage data science at Spotify, caught the startup bug at Tink, and recently joined an exciting new company called Ark Kapital, together with Spotify's former VP Analytics. Jens explains how he and his team killed the notion of raw data at Tink and walks us through the Google, Spotify and Ark Kapital data stacks.
Listen on [Spotify](https://open.spotify.com/episode/1li4BomGpzfXrDytERxGbF?si=0iXPB-uRQ96AOV3Df0hxug) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/getting-rid-of-raw-data-with-jens-larsson/id1561927688?i=1000554823929)
Guest: [Jens Larsson](https://www.linkedin.com/search/results/all/?keywords=jens%20larsson\&origin=RICH_QUERY_SUGGESTION\&position=1\&searchId=72d96b66-e8a4-41bd-a5df-91797dae829c\&sid=r7r), Head of Analytics at [Ark Kapital](https://www.arkkapital.com/)
Hosts: Eldad and Boaz Farkash, AKA The Data Bros, CEO and CPO at Firebolt
Boaz: Hello Everybody! Welcome! Eldad, How are you?
Eldad: I am good.
Boaz: Welcome to another episode of the Data Engineering Show. Before we start, I do not know how much time will pass since this will be published. So, if you are listening to this, probably like a few weeks have passed since we recorded it, but I did want to send out our support to our Ukrainian friends, at Firebolt and to the whole of the Ukrainian people with such a situation that is out there and we just wanted to send out our support. Now, we can dive in.
Today with us is Jens Larsson, Head of Analytics at Ark Kapital. Jens has quite a background in analytics in general. He spent his initial years doing analytics at Google. Google was too big for him. He moved to Spotify, which is a little bit smaller, and spent a few years there doing analytics and then moved to FinTech in Sweden in Stockholm, which is where he stayed and recently moved to a startup, which is called Ark Kapital, doing exciting stuff with analytics. How are you, Jens?
Jens: Hey guys, I am doing really, really good.
Boaz: Did I miss anything about your story?
Jens: If by the story, you mean where I have worked, then, I do not think you missed anything. I have not moved around that much. But, yeah, that is above my story. I am from an industrial town in Western Sweden and studied engineering and business. I got my first job at Google in 2011, moved to Dublin, which was exciting, moved back to Sweden, did Spotify just like you said. Then Tink, who recently was acquired by Visa and now Ark Kapital, and I think I am more or less 10x downwards in size, every time I have transitioned a company. So next time I join another company, it will be like 0.2 employees or something like that.
Boaz: And it is also with a pandemic, you are working from home. It seems like you are getting more secluded from society over time and so started to be worried, what's wrong?
Jens: No, that is very much true, though I have to say, since I joined Ark in November, I have actually spent most of my days in the office. It is just so nice to be around good fun people and I have truly missed that over the last two years.
Boaz: How did you get into analytics, to begin with?
Jens: I think it started out at Google. We were working with customers doing kind of sales, support, blogging, and stuff, to essentially help the ad-words business. And, I guess I just had a knack for analytics and somehow was transferred into a local analytics team doing, analyzing the performance of our sales teams, analyzing the performance of our ads customers, and so on, trying to optimize the way we sold ads at Google. I was, of course, super exciting and not so much maybe the problems we were solving, but the people I got to work with and that data stack, I do not think it occurred to me at the time I was fairly junior, and it was my first job, but what Google had already back then in 2011, 2012, must have taken me six or seven years before I got to experience anything like that again somewhere else.
Boaz: Let us talk about that. How was Google as a school for becoming a data person? Tell us about that data stack a little bit more.
Jens: Yeah, absolutely. So, we are talking about 2012, roughly, right? So Google has MapReduce and all of that is already 10 years behind Google or something. But, basically, I started out with a bunch of these patterns that people are talking about just now in terms of, let us do ELT, let us load all our data into a big data warehouse and do transformations in SQL, and I never realized how revolutionary that was, because that is the way I have always worked with data. Data stack at Google back then was all loaded into some distributor file systems and queried through a tool. I think it was called Tenzing. I think it is equivalent to Hive, which was eventually open source to backup it. While I was there, they were starting to roll off Dremel internally, which eventually would become a big query, I guess, and I just remember being blown away by the speed. Everyone has this story about how you start to query across gigs or terabytes of data, and then you would go and get a cup of coffee and you come back to see if it is done. We did that transition there back in 2012 or something like that. It was fun, transitioning, rewriting queries from the Tenzing dialect to big query, took a lot of time, was a lot of headache. I remember Dremel when it came, it did not support joins. Then, eventually, it started supporting joins, but you can only have one join in each query. It was quite painful, but the reward was the speed you got in return.
Boaz: Interesting to think back then, did data engineering exist? What kind of positions were around analytics that helped the end-to-end data flow.
Jens: I had never heard about anyone calling themselves a data engineer back then. We were part of the Google sales and marketing organization. There were not many engineers on the payroll in that part of the company. To me, raw data magically appeared in a bucket somewhere in a file system. And, then I would write all my queries in SQL. I would schedule it through the SQL UI, and I would power some internal dashboarding tool off of the back of that as well, all end to end essentially.
Boaz: On the other end, there is somebody sitting and saying, how ungrateful are these people? They get the data into a bucket.
Eldad: He had a mustache.
Jens: It just magically appears.
Eldad: Did he have a mustache?
Boaz: Probably.
Boaz: How many years did you spend at Google?
Jens: I think it turned out to be like three and a half or something, four maybe.
Boaz: And then you move to Spotify. Tell us about that. What did you do there?
Jens: Yeah. So moved to Spotify and side note, the VP of analytics at Spotify, when I started, is actually my current manager. So, if we skip ahead a little bit on the story, that is why I am at Ark Kapital today, but going back to 2014, I joined Spotify, into their still relatively small and unproven analytics team. We were responsible for everything from calculating the royalty payouts that went out to artists and labels, all the way to understanding how people were doing product analytics, really understanding how people were interacting with the application, to trying and what I spent most of my time, understanding how people were using the free product and eventually upgrading to the premium tier and see what can we do to get more people to upgrade and what different product offerings do we need to offer in order to have that? Because this is back to 2014, Spotify has just gone live in the US, I believe, a reason to go live in the US and there was only one product really and that was the Classic 999 Spotify Premium, but we fairly soon launched students tier, a family plan and so on. So a lot of workaround, how those different products would cannibalize our own users or will they complement our own subscriber base, essentially will we eat their own lunch or will someone else do it for us?
Eldad: So, you switched the metadata to playlists, songs.
Jens: Yes, absolutely.
Eldad: Stopped listening after x seconds, time becomes like 2 minutes, 3 minutes when the queen is being played, maybe 10 minutes.
Jens: Yeah, we did a lot of those fun little experiments because we had all this metadata about when songs are being streamed. We could basically detect when a local football team had won a game because the "we are the champion" song would spike in that region when that happened. There were a lot of these kinds of fun things you can do. On the other hand, we were quite limited by our capabilities. I think we used to at least tell ourselves that we had the world's largest Hadoop cluster at that time to process all this data. That does not mean it was very fast for the kinds of analysis we tried to do. That is why I said, it was like going back in time, I had just helped our team at Google migrate off of Tenzing and now, we come back to a Hadoop cluster where people are still writing MapReduce jobs in Python.
Eldad: Progress.
Boaz: The years you spent at Spotify were years of crazy growth, right? How did the increase in scale feel like from your end?
Jens: Yeah, it was pretty intense. We must have gone from 600, 700 to 5000 to 6000 employees over those four years or five years, maybe. Scaling was pretty intense and in particular around the analytics team, and I think analytics and data engineering were some of the areas that really had to scale and scale pretty fast. And in this time span, we also moved from this massive Hadoop cluster into the Google cloud, the world of data services, basically every service we are running on GCP eventually.
Boaz: Do you guys move completely to GCP or was it sort of that you have things in AWS side-by-side?
Jens: I know that we were a few things still running on AWS. The Hadoop, it took a while to properly kill off the Hadoop cluster. But yeah, I think the end goal is everything was moved into GCS.
Boaz: So from 1 to 10, how much do you miss the Hadoop days?
Jens: I have this kind of romantic view of these MapReduce jobs we used to write. I love just going really deep and optimizing the combiners, and figuring out how to avoid an unnecessary shuffling step between jobs and stuff like that. I kind of miss that. I also do not miss that because it was taking way too much of my time. But, it gives you this opportunity to feel really smart about the work you are doing.
Boaz: Yeah and now you look at the younglings of today who do not have to worry about these things.
Jens: Exactly. It is all just drag and drop this and trying to click boxes.
Boaz: Okay, you spent all the time Google, Spotify, learning how to work with some of the world's most complex systems and data sets and then you went on to the startup world, right? Move to FinTech.
Jens: Moved to FinTech to Tink. This was back in 2019 or early 2020. They had just done a pivot from being this business to consumer personal finance management application. They had kind of been a driving force in creating the whole world of open banking, forcing banks to open up APIs. Long story short, Tink started fetching data from banks, closed APIs before there was a mandate, whether banks had to open up or not and the banks deemed it was probably legal. Then the courts realized, no, it is not illegal. It is perfectly illegal and we are actually going to create these mandates that force banks to open up APIs. Tink was really early in that journey, but we are building most of their own applications. When I joined, the company completely pivoted into becoming this platform as a service or API as a service. That kind of unbundled the app fee functionality and features and sold them off to other FinTechs and sold them off to other banks. So, we did things like connecting to bank accounts and fetching all the data so that you can do various analyses, risk analyses, or actually building personal finance management. We were categorizing transactions using various AI machine learning models to figure out all these line items on your balance sheet, on your transaction sheet, what are they actually, which is a surprisingly hard problem because the banks do not include much data in those transaction lines. It is usually like just mumbo-jumbo when trying to read it, sometimes you can read like an MCD and figure out it is probably a McDonald's transaction, but that is basically all you get.
Boaz: And so the analytic stack there was part of the product, the services that the clients were consuming, went through, your stack.
Jens: Yes. The entire backend of Tink is a data platform in a way and it is completely homegrown. It is using things like Kafka and S3 and Cassandra and various databases, but it is all focused on productionizing data access to customers. What I was in charge of building was really this data warehouse, where we could learn ourselves, how our products are actually being used? Instrumenting everything from event collection so that we know what our systems are doing, to how our customers are using our systems, and also some batch processing to get the data back out of Cassandra because Cassandra is not a very nice place to run massive queries over.
Boaz: Can you walk us through the different teams in charge of the different data deliverables. You run the analytics department. There is data engineering, engineering around it. How do all of you interact in terms of responsibilities around the data platform and stack?
Jens: Yeah, so when I was at Tink, I was heading up both the data engineering side and the product analytics side. Most of the analytics we did were product-focused. And most of the data engineering we did was focused on getting this metric data, getting data out of the platform that allowed us to do analysis, create KPIs and metrics, and so on. We were, of course, leveraging a lot of the infrastructure that other teams were building at the company. So, the data engineers were heavily dependent on our infrastructure team that managed not one, but I think 11 Kubernetes clusters in various AWS and on-prem and so on, instances to create these environments that we provided out to our customers. It was an extremely complex setup in a complex environment but the data engineering team had to figure out ways to kind of standardize how we source data from all these different systems and how we put them in a unified data warehouse model.
Boaz: And what did the data stack look like? What data warehouse did you guys use and what was around it?
Jens: Yes, for much of the data, we were kind of bound to you saying, like AWS tooling. So, we were using S3 and Athena. We tried a little bit of Redshift and so on, on that. But for the data where we had anonymized it, taken out any sensitive bits of information, we moved a lot of that over to Google cloud, to use that to power things like interactive dashboards, metric collection. We even used that power, the developer console will be feeding some of these metrics back to our developers that we're building stuff on the platform.
Boaz: In retrospect, you spent a fair amount of time there building that. What would you have done differently if you were to restart that entire journey? What lessons were learned? What could have been avoided?
Jens: What could have been avoided? So much pain could have been avoided. I think what took the most time was to figure out which data we are allowed to do work with? Because you have to realize, Tink is a data sub-process around the GDPR. Does not really own the data that flows through the platform. It is processed there on behalf of someone else and at the end of the day, it is on behalf of the end-user, the person is actual financial data wizes. So, we need to make sure that the data we look at and analyze is only metadata about whether or not someone has to aggregate the data, not the actual aggregate data. And, I think if I were to start this over again, we should have created much, much clearer vocabulary around this really early on in the process, because we could have just avoided so much back and forth discussing security and anonymization and legal. If we would have just said, this data is metadata, about how Tink services are being used, we do not even need to bring that into the discussion and we could have limited the scope drastically. And, we could have probably been more proactive in how we set up contracts with customers to allow that. But yeah, at the end of the day, I think we would have made better progress if we would have been more upfront with figuring out what are the different classes of data and the different use cases we have of that data and realized that we do not have to enforce the strictest restrictions on all of it, because what would you do when there is uncertainty and unclarity in this case, what you do is you apply the strictest rules across all the data, and that kind of also limits what you can do with the data. And it creates a lot of headaches for people trying to do stuff that we have designed, that you cannot do for good reason. And, then you try to approximate or you try to not necessarily sidestep, and not necessarily work yourself through limits and barriers, but you try to create the proxies, try to make an estimate of a metric that you actually could have probably just queried if we had clearer, the clearer delineation between what data is sensitive, which data is not sensitive.
Boaz: Yeah. I guess, as you noted, it is not the kind of thing that you thought would be critical to your work when you first started off at Google. You thought one day I will be regretting, not having planned enough, not having thought about what data I am allowed to keep and what not. But it is true at the end of the day, these things bring the entire difference in the project with a lot of headache or less.
Jens: Yeah, absolutely. Then, there were a lot of good things we did there. One thing we did is we more or less killed off the notion of raw data, because with the kind of idea that there is no such thing as raw data. We are not like pumping raw and crude oil out of the ground and then we have to refine it. Data is something that we control and that we create. We could have probably created crude oil and then, we would have created a process to refine it. But, instead, we decided to create nice tabular data with strictly enforced schemers and contracts from the get-go. And a big chunk of our data engineering work we just never had to do because the data that is streamed in from all these other services, was already pretty clean when it came to it.
Boaz: That is now an interesting philosophy that I never heard anybody articulate that well. But saying, data essentially, it is not born out of a vacuum.
Jens: It is not. It is our system. It is our source code that creates the data. Why would we create ugly data? And, I think one of the reasons why you would create the ugly data is because you created it for a different purpose like you created this data because it is a log record that you want to feed into Logstash and do debugging. We had this alert when you set up this system, you are like, either we tried to build something brand new or we go with what we already have and we tried to retrofit it like a data platform on top of these Logstash data that we had available to us. And, thanks to our lead data engineer who had spent many years also at Spotify and other companies in the past. He was heavily advising against trying to retrofit and repurpose that log data and, actually now let us build a service that we speak through with proto buff, schemas, and contractually sound data. Then, we can stay in control of that data throughout the whole chain.
Boaz: So, the message may be - Let us not be victims of our raw data. Let us take control. You can actually change it up.
Jens: Yeah. Do not even go near the raw data. If you have the opportunity to say no, let us build a parallel system that creates nice data. Do that instead, I think.
Boaz: It is time to question your raw data. So, tell us about Ark Kapital. So, then you decided to join, like you said, your former boss at Spotify. What do you guys do at Ark Kapital? What is the mission there?
Jens: Yeah. So, it is really two-fold and I am obviously most excited about the data part of this. So, I might not give the most flamboyant or elaborate description of the other. But I will start with a "boring part" which is we are lending money to companies to finance tech companies and other modern companies. We do this as an alternative to bringing in say, venture capital and the idea is your current options as a company looking for capital for growth is either you have assets or you have some other security that you can use to get the loan or you go to venture capital and you give away part of your company in exchange for that funding that you need to grow and both of those have a place. If you own property or buildings or whatever, go to the bank and get a loan on those. If your business is not proven, you may not have found product-market fit and so on, go to venture capital. They are perfect for handling that risk. That is exactly what they are experts at. But, then we see all these other companies, they might be too new or they might not be established enough to really go to the bank and get money, but they might also have very predictable growth. They know that for every $10 they spend on marketing, they get $25 back in terms of sales. If you are in that position, you do not necessarily want to give away your company to a venture capitalist either. And, that is kind of where we come in. The reason they do not go to the bank is that the bank does not have access to the information really. They can look at your annual reports from a couple of years ago. If you are a fast-growing company, that is not going to work out. It is really hard to model this and excel in a way that allows you to really understand what is going on. That is where our data platform comes in. So, most of these modern companies are building their organizations on SAS platforms. Modern e-commerce companies built on Shopify or Instacart or any other of those services. They do their advertising on Facebook with Facebook ads or Google ad words. They do their bookkeeping and CRO and so on. And, we connect to all of these systems. We are using Fivetran Airbyte and other services to connect straight to the system and get the source data. Then, we apply models who we have years and years of experience in business analysis and it comes from a VC firm before joining Ark. The rest of us are quite seasoned analysts. We could analyze the business performance using all these different data from all these different sources, and we apply machine learning and we do forecasting to try to really get an idea of where this company is heading. And, based on that, we can really tailor the financing option to that company. So, we call it precision financing, where we learned so much about the company from these sources, and hence, we can tailor financing solutions to them. But, in order for us to kind of digest this, all this data and all this information, we are also building really cool dashboards for ourselves that we are also providing back to our customers. So, we are really building a turnkey solution where you connect your data, and we give you a best-in-class business dashboard with KPIs and forecasts and LTV models and all of that nice stuff that all the big players already have. So, I think of it as a bit of an analogy when you buy data tools these days. You buy a data warehouse, you spin it up and it is empty or you buy Looker and you get this dashboarding solution and you open it for the first time and you are greeted with a blank empty page, which I think is a quite boring way to buy data products. Ideally, more and more products in the future, I hope, will be like turnkey. You buy the product, you authenticate to get your data and then you open it and there is an actual dashboard there already.
Boaz: Yeah. Beautiful. This is an essence you are entering Ark Kapital with a blank slate. Please share, how do you go about implementing the analytic stack, with your experience, but now have the complete freedom of choice?
Jens: Yeah, with the complete freedom of choice. There are many of us in this company that have a lot of experience with GCP. We have decided we are going with GCP. Then, we have all this low-hanging fruit. There are all these kinds of companies that help make our lives easier. One of them is DBT. I have built tooling that works pretty much like DBT several times before, but it is quite nice to just install it, get it off the shelf. We get Fivetran, Airbyte which allows us to connect to all these different APIs. So, we are basically collecting this different software that helps us kind of do our job. But we are also taking quite a lot of time to figure out exactly what that platform is going to look like. There are so many things that we still have not decided on. For instance, orchestration, one of our big headaches. I have heard it on this podcast too, people talking about how much time they are spending, just managing the Airflow instance and trying to upgrade it, and so on. Trying to figure out what is the modern way of orchestration of all these different data pipelines, because our complexity does not really lie in the volumes of data. It is the diversity of data. We were fetching it from hundreds of platforms for hundreds, or potentially thousands of customers and we cannot really, as a colleague told me yesterday, who had been working at Spotify with similar problems in the past, we cannot really take a representative of Google analytics and put them in the room with one from Mixpanel and tell them to start aligning their data models, "Hey guys, can you please start defining amu the same way."
Eldad: Let us join everything together.
Jens: Let us join everything together. Can't you guys just agree on what the daily active user is? So, I do not have to kind of show diverging definitions of the same data. There is a lot of complexity that stems from just the fact that all the data is different that comes from various sources and yeah, the complexities are hard to grasp actually.
Boaz: How do you imagine things looking a year from now? What do you want to look back on with pride having built?
Jens: Year from now, I really, really want as much as possible to be automated, and we have this idea that a customer that comes to us and opens their dashboard for the first time and sees their metric, they are kind of blown away by how rich it is and the insights they are getting. We are currently achieving that, but with quite a bit of manual labor in between, and I mean, creating good visualizations of data or standardizing it is a craft and it is quite hard to automate and do that at scale. So, I am really hoping that a year from now, someone just goes and connects their data to our platform, and they are more or less immediately blown away by the insight that they are able to get. And, hopefully recognizing the numbers they see in our platform from what they have in their own spreadsheets and so.
Boaz: Awesome! This is great Jens! I really appreciate it.
Eldad: Yes.
Jens: Thank you, guys!
Boaz: It was great talking to you, a super interesting, amazing journey, amazing challenges in the past, and amazing challenges you are working on right now. Thanks again!
Jens: Yeah, thank you too.
Eldad: It was great to connect. Thanks.
Boaz: All the best Jens.
# GROUPING SETS as a pure planner rewrite ? Yep - it's possible (/blog/grouping-sets-as-a-pure-planner-rewrite-yep---its-possible)
In this blog post you will learn how GROUPING SETS work and how Firebolt's implementation uses smart query planning to execute them efficiently.
## Introduction - What GROUPING SETS do for you [#introduction---what-grouping-sets-do-for-you]
When working with SQL queries, reporting and analytics often require summarizing the same data in multiple different ways. Let's say we have the following table tracking orders of our products:
We might be interested in multiple aggregations here:
1. the total sales amount per region and product
2. the total sales amount per region
3. the total sales amount over all orders.
Usually we would cover any of these points with a GROUP BY statement. For example, to get the total sales amount per region and product we would query the data with
```sql
SELECT
region, product, SUM(amount) AS total_sales
FROM
orders
GROUP BY
region, product;
```
But how do we get all three aggregations in our query? Instead of writing multiple GROUP BY statements, SQL provides a powerful feature called GROUPING SETS, which allows for multiple aggregations in one go:
```sql
SELECT
region, product, SUM(sales_amount) AS total_sales
FROM
orders
GROUP BY GROUPING SETS (
(region, product), -- Sales per region and product
(region), -- Sales per region
() -- Grand total sales
);
```
With this query, the result table will look something like this:
We just did three aggregations in one swoop: The first four rows are the result of the aggregation with both region and product, the next two rows aggregate only by the region, and the last row corresponds to the empty grouping set (also called *grand total*), that evaluates the aggregate function without a GROUP BY key. Columns that are not part of the grouping set of the current aggregation are set to NULL.
## Rewriting GROUPING SETS [#rewriting-grouping-sets]
So we have seen that GROUPING SETS can be used to request the same aggregates for different sets of grouping keys in a concise way. Imagine you were confronted with a SQL database that does not support GROUPING SETS syntax. The first idea you might have is to write a query like this:
```sql
SELECT region, product, SUM(amount) AS total_sales
FROM orders
GROUP BY region, product
UNION ALL
SELECT region, NULL, SUM(amount) AS total_sales
FROM orders
GROUP BY region
UNION ALL
SELECT NULL, NULL, SUM(amount) AS total_sales
FROM orders
```
The output would be the same, but the query is much more verbose. But maybe one could implement GROUPING SETS this way? Just rewrite any statement containing GROUPING SETS into one that the runtime can already execute. The GROUPING SETS would only exist at the query planner level and would internally be transformed.
But using UNION ALL for this is probably not a good idea. We want to have as few scans over the base table as possible so we only have to read it in once. With this approach however we would have one scan per grouping set. That corresponds to the FROM orders part of the query. The amount of scans over our base table therefore increases linearly with the amount of grouping sets we have. This gets even worse when you calculate your grouping set over more than just a base table. Imagine you compute grouping sets on the result of joining terabytes of data. The UNION ALL strategy would mean that you need to evaluate that very expensive join once for each grouping set.
But there is a smarter way of rewriting a query with GROUPING SETS. Let's take a look:
```sql
SELECT
case when get_bit(gs.id, 0)
then region
else NULL end as region,
case when get_bit(gs.id, 1)
then product
else NULL end as product,
SUM(amount)
FROM
orders,
UNNEST(ARRAY[3, 1]) as grouping_sets(g_id)
GROUP BY
grouping_id, 1, 2
UNION ALL
SELECT
NULL as region,
NULL as product,
SUM(amount)
FROM
orders
```
What happens here? Let's focus on the first SELECT for now. In the FROM clause, we not only scan our base table, but also construct a cross join with a static table grouping\_sets that contains a single g\_id column that encodes the columns participating in the requested grouping set. Each input tuple from table orders is paired with each g\_id value. Our example table after this cross join would look like this:
The values of the g\_id column represent a binary encoding of the different combinations of the expressions passed as arguments to the GROUPING SETS clause. In our case, the two columns used in the GROUP BY clause are "region" and "product". Hence, we have two binary positions that can either be set to 1, if the column is used in the grouping set, or to 0 otherwise. This way, we can distinguish the different grouping sets and do a GROUP BY that covers all of them. For instance, in our example a g\_id value of 3 (binary 0b11) denotes the grouping set that includes both columns, while a g\_id value of 1 (binary 0b01) denotes the grouping set containing only the "region" column. Consequently, in the SELECT clause, we can retain the original column value if the corresponding g\_id bit is 1, and map it to NULL if that bit is 0 (we can check this with the get\_bit function). The aggregate input of the SELECT statement from our example query looks like that.
The GROUP BY clause uses the g\_id and the possibly NULL-mapped region and product columns as the key.
The special case of an empty input requires grand totals to be handled separately. This is because a GROUP BY clause with an empty grouping key should return the default NULL value for each aggregate function, rather than an empty result.
What's so great about this rewrite is that we can compute any amount of grouping sets with at most two scans and two aggregations. And by having one larger aggregation instead of multiple smaller ones we can take full advantage of Firebolt's distributed aggregation computation.
## How GROUPING SETS are transformed from start to finish [#how-grouping-sets-are-transformed-from-start-to-finish]
After the parsing stage, a GROUP BY clause that utilizes GROUPING SETS is internally represented as a special aggregation node within the query plan. This node stores all the information pertaining to the grouping sets by using bitsets, which function similarly to the g\_id values in the rewritten SQL query.
This special node only exists at the beginning of the logical query plan. This logical plan is an abstract representation of our query that we can use to perform optimizations using rules. The rules run over the tree and try to find patterns that they can optimize. In our case we use a rule that matches on this special aggregation node and transforms it subtree implementing the rewrite strategy outlined as a SQL-level rewrite above. The resulting plan consists of nodes that can already be executed by our runtime, so no runtime extensions are needed for adding support for GROUPING SETS.
Let's examine what the rule does in detail.
1. Create the grouping\_sets(g\_id) static table and the cross join. We can just interpret the bitsets in our special aggregation node as integers and build a table this way.
2. Build a projection that optionally maps grouping set expressions to NULL depending on the current g\_id value. In addition, the projection retains the g\_id value and all aggregate input columns.
3. Add an aggregation node
4. Do the same thing, but for the grand total (if requested in the SQL query) and combine both paths
5. Remove the g\_id with another projection
And just like this, we transformed our GROUPING SETS aggregate node into an equivalent logical plan that can be further optimized and later executed.
There is one more edge case that we need to cover in the query planner: duplicate GROUPING SETS. A query like
```sql
SELECT
region,
prodcut,
SUM(amount) as total_sales
FROM
orders
GROUP BY
GROUPING SETS ((region,product),(region,product))
```
has the grouping set (region, product) repeated twice. This leads to the aggregation result being duplicated:
We want to compute the aggregation without the duplicates first, and then duplicate the rows later, when we already have the correct aggregation result. That means our cross join would look like this:
Even though we have two grouping sets that correspond to g\_id = 3 we only have one 3 in the static table. We only consider the duplicates later, by adding another join on top with a static table that keeps the multiplicity:
With this setup, we make sure to compute the GROUP BY both correctly and efficiently—no matter how often a grouping set appears in the GROUPING SETS clause, all aggregate functions are only computed once per distinct grouping set.
### The grouping function [#the-grouping-function]
You might have noticed a problem with GROUPING SETS so far: The data might contain NULL values, so how do we differentiate between them and the NULL values added by the aggregation? SQL also offers a function for that, called *grouping*. Similar to our g\_id, the function takes expressions found in any grouping set as arguments and produces an integer that represents each expression as a single bit. In this case, a 0 bit in the integer indicates that the expression is part of the aggregation.
If our initial query had utilized the grouping function:
```sql
SELECT
grouping(region, product) AS grouping,
region, product,
SUM(amount) AS total_sales
FROM
orders
GROUP BY
GROUPING SETS ((region, product), region, ())
```
The result would look like this:
Grouping functions are also evaluated during planning—for each row the correct value is computed based on the current g\_id value and the arguments of the grouping function call. The values can then later be joined and added to the result.
## Conclusion [#conclusion]
**GROUPING SETS** offer a powerful way to perform multiple aggregations in a single SQL query, improving readability and conciseness. Firebolt's implementation leverages smart query planning to efficiently execute these aggregations by internally rewriting them. When transforming **GROUPING SETS** into a series of joins and aggregations, Firebolt minimizes table scans and optimizes distributed computation, ensuring high performance even with intricate grouping requirements.
# How Amplitude Engineers Process 5 Trillion Real-time Events (/blog/how-amplitude-engineers-process-5-trillion-real-time-events)
Weichen Wang, Senior Engineering Manager at Amplitude, came to meet the bros to talk about Amplitude's cutting-edge data stack and how it processes 5 Trillion real-time events while dealing with mutable data and massive scale.
Listen on [Spotify](https://open.spotify.com/episode/3o0GCNoxaq1g5Dlv65PFW3) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/the-data-engineering-show/id1561927688?uo=4)
**Benjamin:** Welcome back everyone! Today, both Eldad and I are in Tel Aviv. We have the pleasure of having, Weichen Wang, here, who's a Senior Engineering Manager at Amplitude.
A few words about Amplitude and Weichen. Amplitude is actually at this point a public company. Before that they raised more than 300 million in private rounds and Weichen is based in Vancouver, British Columbia. He leads the data connections team there and built on scale, their Vancouver site. So, we're super happy to have you. Do you want to give us a high-level view of what Amplitude does and what service you guys deliver to your customers?
**Weichen:** Sure. First of all, I just want to thank you for having me today, super thrilled! My name is Weichen. Thanks for the introduction!
Maybe I can talk about Amplitude, the team that I'm working with.
For Amplitude, we are a digital analytical platform. We started serving the customers for product mix. We're giving this self-service visibility into the entire customer journey. We have all those tools for funnel analysis, behavior graphs, event segmentation, and all of that. The key takeaway here is we try to help the entire spectrum of data collection, instrumentation, governance, and all to get insight and sharing out. So, it's kind of the one-in-all solution, if you may.
For the team, I'm working on - We are the data connection team and as the name suggests, our mission is to send the data in and out to the right place, at the right time. If you think about the entire space out there, all different sorts of services, including data warehouse, ads, marketing, CRM, and CDP. Sometimes, the data has to be sent in real-time, when time is of the essence. Either time, the data will be sent in batches of different natures. So, that's where our team is building this pipeline and also the tools, and the user interface for our customers to handle those data.
**Benjamin:** All right, sounds awesome. We'll have a lot to talk about. How big is your team and data connectors, can you give some examples of...?
**Eldad:** What this means?
**Benjamin:** Exactly. Like, where's the data coming from? Where's it going?
**Weichen:** If you talk about the Amplitude as a whole, we're about 700 people altogether. And for engineering, I would say it's less than 200 at the moment. This is still a fair small team at the moment. Data collection is roughly less than 30 people, I say altogether. So again, it's a fairly small team. You want an example of the connections. I already named just a few of the categories. In the different places people want to bring data in and these years, the data warehouse has been a really prevalent choice where they served as the first step in all the data, in relation to the governance and everything, and this server, the source of choice. Where in other cases, some clients might prefer trying to get data indirectly from instrumentation, using our SD case and sending directly from the devices or websites. So, whether data hit to us directly in the streaming fashion. So, that's already a different type of ingestion.
When it comes to exports, they are different, again, for example, sometimes people use us at a CDP capacity where they do real-time event propagation. As soon as we receive the data, we probably send them to their data warehouse or anything else. And in the other case maybe they generate insights, for example. They might do queries on user segmentations with audiences and this is a new group of users. They want to start a new campaign, ads. Then they actually want to sync that audience to Google Ads, for example. That's another problem. So, we're trying to support as much as we can at the moment, about a hundred sourcing destinations.
**Benjamin:** Okay. Awesome. Can you give us a feeling of the data sizes or the data volume you guys are dealing with?
**Weichen:** Yeah. We measure the data that is ingested in a number of events. Mostly, we process somewhere around 5 trillion events altogether and most of them at moment is real-time streaming. But we also are ingesting lots of batch data as well.
**Eldad:** 5 trillion is nice. Very nice.
**Benjamin:** Yeah, definitely.
**Eldad:** Is it run on a daily basis or a monthly one?
**Weichen:** No, that's a monthly basis.
**Benjamin:** Awesome. You manage to connect your team, and this is where you have the most experience, but do you want to give us a high-level overview of the internal data stack at Amplitude and where the data goes after it made it through your connector?
**Weichen:** Sure. So, I can give you a high-level architectural overview, although I'm trying to not disclose too much information, that we have been talking about.
**Benjamin:** Of course.
**Weichen:** The most important part was when Amplitude started, there were a few principles that when came, and again, this came from one of the early engineers and the founders and I joined at a much later time. So, one of the key points here is we want to make the data query super fast.
Even with the real-time data coming in, how can you update existing queries in a real-time with as lower latency as possible? So, basically, they chose a column, the storage architecture. Cause you think about other queries you could not make, especially for product mix. So, making like say DAUs or other things, users who did this action after that action usually with just one or two properties that you are interested in among all the event property or user properties for the data. So when you index things in the column, things get you professed. We actually have it in-house column storage system. It's proprietary and at the end of the day, it just stored the raw events and then we can index them in a way. Then, we have a query system we call Nova, that actually builds on top of it. So, it's actually working in a kind of a divide and conquer fashion where it tries to utilize a cluster of computing resources based on those indexes so that you can get a query out super fast. So that's kind of the essence of the system.
For the streaming, of course, Kafka and various stuff. We also use Postgres, lots of things. We started with MySQL, but we're trying to move away from them. And then, everything was on AWS. So, we have to use an S3 for storing those raw indexes and events. So that's probably the most important thing on a high level. But of course, the bunch of things down the roof, like we were Redis with a bunch of other stuff.
**Benjamin:** Got you. Okay.
**Eldad:** Who should take credit for inventing columnar storage and databases?
**Benjamin:** Boom. Okay. Asking you database research questions to me.
**Eldad:** Is it the Stanford guys or is it the CWI?
**Benjamin:** That's a good question.
**Eldad:** Is it Store or is it MonetDB?
**Benjamin:** So as someone from Central Europe, let's give it to Amsterdam.
**Eldad:** Boom! Amsterdam wins.
**Benjamin:** Amsterdam wins, easy, but I mean, that's super cool. So, basically, at least at the core of your stack, you then have this in-house query engine in a sense. Like it's not talking SQL, but this domain-specific language that actually powers all of the stacks. That's actually super interesting. Cool.
Now, what's the hardest thing about your internal data stack? Is it lower latency? Is it high concurrency? What are the biggest data challenges in a sense you guys are dealing with and trying to give your users a great experience?
**Weichen:** If I put it in words, it's just a scale. Everything was easy when the scale is small. But as a lot of demand comes in, even the original architecture would not scale, even the actual feasibility when the query span you cannot scale vertically, just a number of open clusters and all that. And now the other part is cost.
For real-time, it all sounds really good, but of course every time there is trade-off. It's an intentional design that we have, but then that also actually has a recurring burden on us to make everything super fast. Sometimes, there are use cases. I don't know if you actually order food, like with DoorDash. So, maybe you put something in your shopping bucket, but you didn't check it out and you move away. Maybe the DoorDash, hey, you want to send the coupon or push right way in 30 seconds before you gave up and try to finish that order that is super time sensitive. So that's why we need that to be really fast responding. In other cases, that itself, maybe in the monthly data of your inventory and stuff, there's no sense of time, urgency, and that cost is kind of actual.
The other example is data mutability. If you think about the event ingested from an instrumentation SD case, most of you think, oh, the data is mutable and which is the assumption that the system was built before. But there are lots of cases where that assumption is not true. The biggest example is just GDPR. The data deletion request where they actually remove every trace of certain users and that is a big pain point where we had. We actually have a dedicated team to handle that.
**Eldad:** Thank you for that.
**Weichen:** Yeah. Then we are of course rethinking our architecture.
**Eldad:** Also European.
**Benjamin:** Exactly. Like a European, you take credit for the second time.
**Eldad:** Sure, the delete feature in the database and data warehouse.
**Weichen:** If you ever doubt what was that, it's, a definite challenge if it was not designed, to begin with.
**Eldad:** I remember that period when delete was dead and everyone was happy with the pend only, immutable. And there were these two years where everyone was happy, and life was presumably becoming simple. And then, GDPR came and everyone woke up.
Tell me, ecosystem-wise, input, output, you said Kafka and Postgres. We get that a lot. What do you see? Do you see more Kafka or more Postgres, one overtaking the other? We see a lot of comeback of Postgres. What's your take on that?
**Weichen:** Kafka is mostly for the streaming process. Postgres, which we use primarily for storing metadata. Because we have already our separate column indexing system, so Postgres is not used for that purpose. I actually was listening to some of your shows, and lots of people are coming from data engineering backgrounds that they're familiar with just the SQL or just relational database kind of thing. For us, we started with the streaming perspective where the data is coming in and we see every single stream, whereas the batch process or is actually coming in a different stage where we have a separate architectural fit where the system was not designed, to begin with to handle that. But, nevertheless, I think there are benefits that we're designing at a later stage. That means we learn the use cases. So that we know how to optimize that. So back to your question, and Postgres is a popular database. Yes! and we use it all the time. There's a reason that we move it away from MySQL, I suppose. At the same time, they are just serving a different purpose.
**Eldad:** What about output? How do you export your data? Do you see a lot of customers using your platform as a way to cleanse and prepare and then share data with others? How does it play with the existing data? Is it the Snowflakes of the world?
**Weichen:** That's an interesting one. There are a few parts to it. The first is just data governance. There are again, multiple, parts of it. But, you know, I think one phrase I really like is "garbage in garbage out." You have to really control your data ingestion. If you think about all the instrumentation or tools out there, there are certain solutions that actually offer basically with no instrumentation, that's basically auto-tracking and that tracks every button click, every event you get from the app. Of course, that's easy to set up, but then there's a time we had to dig really deep. It's like, imagine you're in this pile of garbage and trying to figure out the gold in it. So it's time-consuming. It may work.
Where's the other part, which is basically what Amplitude, at least at the current stance is we try to do tracking plans and then we try to do versioning of that. So, there are lots of tools we build so that you have a really protected data schema that actually has infrastructure shared across the entire organization. All the teams will actually benefit from it and though it might sound a lot of overhead to start with, and we'll try to make it easier for the small but low-maturity customers. But if you are actually the bigger customers out there, it's actually beneficial to start with, especially to basically make peace of mind for the data leaders. That's one part.
The other part, how we enhance the existing raw events with all the user properties. And then there's the part, we actually try to send it out, in order to touch a little bit, there's a real-time propagation we call stream. So whereas the other, batch export and other time, we send out insights you generated.
So how does that reconcile to the source of truth? I think this is a great question. I guess nowadays, because, Snowflake or the other data route is becoming prevalent, our mission is so that people don't think about that source of choice problem.
**Eldad:** People should always think, I'm telling you, I've learned that. You have been trying to build tools, so people don't need to think and those projects always ended up badly. But, first, your passion is amazing about your product.
**Weichen:** Thank you.
**Eldad:** We love that and listening to you over a few minutes, you first learn how complex it gets when you move from theory to real life, and try to practice data to drive your business, and this is why all of these little things are so, so important.
Tell me, how do you deal with data politics internally? Your team, would that be considered an engineering, a data engineering, a data team? How do they work with the rest of the engineering organizations? What can you tell us about remote and the evolution over the last few years? And then some lessons learned, maybe.
**Benjamin:** Tell us everything.
**Weichen:** You've asked multiple questions, I'm hearing too. Wonders just the internal data policy, it's one. The two is going to be working remotely.
I will start with the first one. Amplitude is coming with a unique edge because we are a data analytics company. Drinking our own champagne is actually a motto for us. We hold ourselves accountable whenever we release something, even if a new feature comes to Amplitude, we actually be the first ones trying it out. If you are shipping a custom-facing feature, you better be using, AB testing experimentation. You better use feature flags and all the other instrumentation in place so people know from, let's say you launch a new, export, the destination of data. So, how many actually click on it? What's the funnel from the user clicking and actually finishing setting up the configuration with all credentials and actually getting data through? So, we have all this tracking and we held ourselves accountable. It's never we shipping features out there. So, in the sense, everyone is a data engineer or a gross engineer so to speak.
I think that definitely provides us with a unique edge, that's one.
Two, just on a remote. I am remote in a sense because Amplitude headquarter is in San Fransico. It's a blessing to me because I was really looking for it. I didn't even know what Amplitude is before someone reached out to me and they opened their office here. So I was hired as one of the earliest managers here without an engineer. And the lot of work I was doing last year is just building the team here.
**Eldad:** Drinking champagne and drinking a lot of ...
**Weichen:** Hopefully, yes.
**Benjamin:** Much, nicer than dogfooding drinking. I would much rather drink...
**Eldad:** He talks and drinks water here. But, some companies are more involved.
**Weichen:** It's just a different way of putting it, but you got it what it is. So working remotely, I guess for us, it's lots of internal debate as well. We at Amplitude want to position ourselves because when Covid started, the situation was staying there. So there is a new reality and a new paradigm that would never go back. Certain, people would just prefer going remote 100%, and they are not even considered working, being the hybrid. We actually had an internal vote and we had an internal discussion.
At this moment, Amplitude is operating in the hybrid remote where we require engineers to be in the office two days a week. We'll still cherish that in-person time, especially with collaboration that's a lot of quality time when brainstorming or face-to-face, which cannot be replaced in my opinion. But yes, there's the time when people can definitely get focused on what they don't want to get disrupted with meetings and stuff.
So, that's where we are at right now. Of course, there's also another debate about whether, for example, the teams should be clustered, geographically. For example, maybe all the teams in Vancouver focus on data connections, and, maybe, teams in London work on something else. So, I don't think that the debate is decisive.
My personal opinion is I do not think geographic location should be a parameter in that equation at all. But you know those people who think otherwise because I do believe that having that flexibility to tap into the talent pool globally definitely provides us a unique edge and the challenge coming with managing a team, it's something that can be managed in my opinion.
**Eldad:** You see Benjamin, it can be done.
**Benjamin:** It can be done. In terms of your team, where are people actually located?
**Weichen:** At the moment, maybe I have to also dive into the kind of a bit of structure, so we run the EPD trail. So, Engineering, Product, and Design. We have pillars and parts. So in data connection, we have three parts and I run two of them. So the streaming and the integrations. For integrations, the majority of the engineers is in Vancouver. For streaming, is probably half-half.
**Eldad:** Tell us about the most embarrassing failure you've had and something that was a big lesson for you. And answering, I'm perfect, I've never made mistake in my life is also valuable.
**Benjamin:** I told you that in interviews, Eldad.
**Eldad:** That's what Benjamin answered in his interview, but we proved him wrong many times. That's why we went on this adventure.
But data teams, we talked to many teams, engineering teams, and when it comes to data, people make a lot of mistakes. So I was just wondering any project that went the wrong direction, any technology that had huge promise and turned out to be a huge flop, something smart for our listeners to learn from in terms of past experiences, the good or the bad. Usually, we love the bad ones.
**Weichen:** Okay. However, I like to share lots of things, I think there is an example I think I can share. If you want to have a third-party authentication system, if you work within a SaaS company and is big enough, they would have authentication done without actually the right keys, so, you authorize on their behalf.
Before going there, Amplitude is in a really interesting position where lots of our customers are mutual that we are their customers and that they are our customers. Also, we share common customers that we have other in a hundred customers, which are both customers of us. In the early days, when Amplitude started, we were still a startup back then, there was a lot of thinking, when we set up those data per account, we just have one and there were also legal requirements where you can only have one account per entity. But, later on, when the system expands, we're building more and more integrations, the problem occurs where the same account was used for both purposes, both for internal teams managing their own data and needs and/or for managing the OR SAP, which is used for our common customers. And this becomes a burden. Sometimes, it is a really heavy burden for us. Then, any migration would come at a risk and also both teams are interfering with the internal teams and internal facing teams are interfering with each other where migration might provide a risk as it might break the data connections for hundreds of customers. It's the burden that's have shown the repeated pattern, and I wish, I know that years ago, so we have, all separation responsibilities for those, internal, and external apps and yes, this actually came with some of the cost. That's kind of the lesson I have. So, my advice for you, try to separate your internal and external systems as much as you can.
**Eldad:** And a cold shower.
**Benjamin:** And a cold shower. So Eldad, you are a bit of a bad cop, good cop. You ask about kind of a failure. I'll do the kind of opposite thing. 2022 is coming to an end. Tell us about something you're super proud of, that your team accomplished this year that really is huge, and that you thought before it was, I don't know, impossible.
**Weichen:** Interesting! Two things in mind. First, we launched a developer portal in early the year. So basically, it's an effort to scale the number of connections we support by supporting third-party developers to build integrations towards us. Because whenever a large client comes to us, Hey, can you have these 20 integrations? And there's only so much we can do, that never scaled. So, we have to rely on our partners to build that bridge and then we provide this generic framework to allow them to submit integrations just in a configuration-based manner. We have a system that parses those configurations and basically, sets it up. So, that was quite successful this time. We probably have almost 50 partners submitting their integrations this year. So, I think that's a huge success for us. That's one thing.
The other for streaming. In the streaming, actually a new product we launched in Q3. Real-time data propagation is increasing demand because at the beginning we think about that we're analytical tools. We won't get data in, but we know we don't want that get data out. But that's apparently not true, especially for the real-time cases. It's basically the path through. They're using our governance features exactly back to that the early point, but, it's again a different system, comes with its own challenges, in relation to reliability, the latency cost and the broadcast overall. So, I think there is a huge shout-out for all the engineers if you're listening on this. So thank you for this great achievement.
**Benjamin:** Wow!
**Eldad:** Wow!
**Benjamin:** Awesome!
**Eldad:** First example by the way, was amazing. As a startup, we meet a lot of partners and platforms and we divide them into two those that have built connectivity. So third party can build their own stuff into the system and those that are not, they're black boxes and a long tail of startups cannot connect to those platforms. So it's really important and you never know how and when someone is using your system and you learn a lot by just opening it up. So congrats on that, especially to your engineers. It's an amazing effort.
**Weichen:** At Amplitude, one of the benefits as a public company, because it actually earns the trust of lots of the partners out there. So, they were linked to us, but of course there are times, those bigger players which never think about us. So we have to still in doing our job, but it definitely still makes our life much, much easier.
**Benjamin:** For us as a database vendor, this is in terms of the dialect and wanting to be compliant with Postgres, to make it easy for the ecosystem to exactly adopt. It's interesting that you guys face similar challenges or similar problems in that regard.
In terms of the second thing, you guys provide these very specific data experiences, to analyze certain user interactions and those types of things, how does this work in terms of like schema management? Do you guys basically say, okay, data has to come to our connectors in a certain schema and then we can give these use cases or like these types of analytics products? Or do you actually kind of manage the schema in your connectors in your system and can accept data in whatever form or shape?
**Weichen:** If you see the pattern for all my answers is like "it really depends." Because as Amplitude evolves, we get more and more customers from different spectrums were in those, the really mature customers where they have their own stuff or in the low maturity customers, we don't have anything which actually easier or somewhere in the middle where they have something. So, we try to be flexible. If they want to use our governing feature, they can, and if we can block all the events. Again, we think things are events. If the event doesn't fall into the tracking plan, we just block them, that's the easier way. And we can use transformations, and filters, certain data only goes to certain destinations and with different solutions, all of that. But, again, we will try to accommodate as much as we can. So that's why it involved a lot over time with all the different use cases coming in. Hopefully, that answered your question.
**Eldad:** It's all about surface area.
**Benjamin:** Yeah.
**Eldad:** Perfect. Okay.
**Benjamin:** Awesome.
**Eldad:** I think we've nailed it.
**Benjamin:** Definitely.
**Weichen:** All right.
**Benjamin:** Alll right Weichen. Thank you so much for your time. We super appreciate it. This was super interesting, learning more about what you and your team are doing, at Amplitude.
**Eldad:** We are going to check out hopefully next year, the connector framework.
**Benjamin:** Definitely.
**Weichen:** All right.
**Eldad:** Come work for us.
**Weichen:** Thank you for having me.
**Eldad:** Thank you for coming.
**Benjamin:** Awesome. Take care and have a great day!
**Weichen:** Take care.
**Benjamin:** Bye.
**Eldad:** Bye.
# How AppsFlyer manages scale without sacrificing performance (/blog/how-appsflyer-manages-scale-without-sacrificing-performance)
Listen on [Spotify](https://open.spotify.com/episode/1dmzUBJXP0WJdPo3ng18Vr), [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-appsflyer-delivers-sub-second-bi-to-1000-looker/id1561927688?i=1000521449360) or watch on [Youtube](https://www.youtube.com/watch?v=MTC9TEMqXcU).
Recently on The Data Engineering Show, we had the pleasure of speaking with [Alexandra Sudilovski](https://www.linkedin.com/in/alexandra-sudilovski-b436a83/?originalSubdomain=il), Senior BI Expert & Looker Guild Master at AppsFlyer. For those who don't know, [Appsflyer](https://www.appsflyer.com/) is a leading SaaS mobile marketing analytics and attribution platform.
Whenever you download and interact with an app, data is sent through their servers. As you can imagine, this amounts to an enormous amount of data.
Appsflyer processes 120 billion events daily and saves 90 terabytes of data in AWS and 40 petabytes in BigQuery daily — and they're still growing.
Appsflyer has exploded in size, growing from a small company of 200 people to 1000 people in just three years. They went from nothing to having thousands of users and hundreds of thousands of dashboards in a matter of a couple weeks.
Dealing not only with a huge amount of data on a daily basis but doing so while growing quickly as a company can come with many challenges. Not to mention, the definition of "huge data" seems to only grow year after year.
We asked Alexandra how she and her team handle these challenges:
*We've learned that you need to adapt quickly and be ready to change your mindset and approach. After we went into production, we had all these people who were hungry for more data, and it became hard to keep up with everything from delivery of new projects, system maintenance, and even bringing on the new employees that we needed. But at Appsflyer, every single department and internal process is measured — and not just the R\&D side of things but also functions like HR, marketing, legal, finance, and more. We want to see everyday how we work as a company and if there are any processes we can improve. We constantly measure ourselves and are improving, and we have seen great results from this. We also use every code that is known on the market to be able to adapt and be equipped to always use the best solution in any situation.*
With so much data to manage, it would be easy for mistakes to fall between the cracks. But Appsflyer takes data accuracy to a whole new level — which is one of the reasons we love this company so much. Alexandra is a leading force behind this drive and mindset:
*If I saw data that was not accurate and not showing the truth, this would make me not believe in my own product. If I don't believe in it, why should others? So my priority is always data accuracy and making sure people can trust us and our product because people rely on that data to make decisions. At Appsflyer, we constantly measure ourselves and double-check our work. There are always multiple pairs of eyes on everything, and we have people dedicated to testing and QA in R\&D. We want to catch any errors immediately instead of waiting to get this feedback from stakeholders and users.*
Alexandra is truly an amazing leader in the BI space, and of course, Appslyer is a company that we believe in and love to follow. We talked to Alexandra about so much more, including how she has blended the roles of BI and data engineering in her work and how she uses some of the hottest tools in the industry today, so be sure to check out the full episode with her.
# How are those data intensive customer facing apps engineered at Gong? (/blog/how-are-those-data-intensive-customer-facing-apps-engineered-at-gong)
[Gong](http://gong.io) manages hundreds of thousands of videoconferences and millions of emails PER DAY, which add up to hundreds of TBs. The Data Bros met Yarin Benado, Gong's engineering manager to understand what is required to move to a modern data stack to support all this, what this stack looks like, and why it all comes down to data quality at the end of the day.
Listen on [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-are-those-data-intensive-customer-facing-apps-engineered/id1561927688?i=1000548477518) or [Spotify](https://open.spotify.com/episode/62BDiWR7CTFZC9aQM3uycL?si=aa14aMlpTIekEtBQQ7D_HQ)
DATA ENGINEERING SHOW INTERVIEWERS:
1. Boaz Farkash.
2. Eldad Farkash.
GONG GUEST: Yarin Benado - Engineering Manager.
**TRANSCRIPT**
Boaz: Hello Everybody! Welcome to another episode of the Data Engineering Show. How are you Eldad?
Eldad: I am good. Thanks.
Boaz: Not tested positive for any variant yet?
Eldad: Dodged it. Had it at home.
Boaz: Talk to the mic.
Eldad: Managed to dodge it.
Boaz: Do not dodge the mic. I dodged it too. Maybe you think it runs in the family.
Eldad: On this version, at least.
Boaz: Yeah. Thanks everyone who joined us. We are here with Yarin Benado. Hi Yarin! How are you?
Yarin: Hey guys! How are you? I am great.
Boaz: Did you dodge the virus too?
Yarin: So far so good.
Boaz: So far. Yarin Benado - Engineering Manager from Gong. Gong, for those of you, who do not know is a really, really interesting and great product, a revenue intelligence product. It started with a product that helps you record and analyze all your conversations, typically for sales teams, but not only, and you can go back and analyze who talked, how long, what words were mentioned, and get insights into how to better, sort of, close deals if you manage sales teams, etc. Yarin joined around a year and a half ago after an acquisition. So, Yarin tell us what the Vayo was, which sort of was your home before entering Gong.
Yarin: I founded Vayo around 2017, with a good friend and a partner, and our goal at Vayo was to basically take the entire customer data that a company is creating constantly, and answer a few simple questions about these customers. Is the customer a happy customer? Are they going to leave the churn? Are there any upscale opportunities? and the idea was basically to automate all the data aggregation, collection aggregation, and reporting around these specific customer-related questions. What used to be done as internal BI and data teams, so we offered a simple solution, where Vayo is doing all the hard stuff.
Boaz: How big was the team at Vayo at that time?
Yarin: So we were a relatively small, very small startup. We ran for two and a half years. We were like 10 employees, mostly engineering, here in Tele Aviv and a bit offshore.
Boaz: Tell us a little bit about your background. What did you do up until that and how did you enter the data and engineering world?
Yarin: Before Vayo, I was leading the engineering team for, a bit bigger startup, here in Tel Aviv. Before that, I held the role of principal engineer for several startups. I think Gong is the biggest company I worked for.
Boaz: How many employees are at Gong nowadays.
Yarin: Almost 900.
Boaz: Wow! You are an engineering manager. Tell us a little bit about your current role and how much of it is dedicated to data.
Yarin: I am an engineering manager in what is called the insights group within Gong. The group is in charge of all the data that is customer-facing basically. So it is either in-app analytics or BI for customer use. Basically, everything we do is around data. At Gong, there are many, many teams and groups within Gong that create a lot of data and we have to build insights on top of this data.
Boaz: Yeah, Eldad and I know, we talked about it in past episodes, I think as well. We always loved the intersection of software engineering and deep data projects. So, let us try to untangle that. Tell us a bit about the current data stack and sort of the current active projects, and let us see what is going on under the hood at Gong with data.
Yarin: Cool. So in terms of the stack, basically, we like to keep things very simple, but at the large scale at Gong. Our data stack is, we have the operational databases, which are serving the entire product and pipelines within the product. In terms of data, we started simple. We just PostgreSQL, then moved to Redshift. Now, we have a bigger project basically to take data from multiple sources for this in-app analytics use case and move everything to a single data warehouse, with a lot of pipelines to aggregate and pre-aggregate, data for basically sub-second query times for dashboards and graphs, etc. For operational databases, we still use a lot of Elasticsearch, PostgreSQL. We have some MongoDB, and as mentioned, it is also for the in-app analytics and we are now moving also a lot of parts from PostgreSQL to Redshift and to Snowflake.
Boaz: These are a lot of things. Walk us through the different teams that are in? How do the teams that deal with data look like at Gong? What do you have between data engineering, software engineering, which teams are dedicated to in-app insights in particular? How does that look like? How do all of these interact with each other?
Yarin: It is a very good question because it is pretty complex at the pace that Gong is growing. We have within the insights group, right now, we have 2 tracks, 2 pods, with products and engineering and data engineering. One of them is to create the in-app analytics track. Another team is focused on the infrastructure of basically building the lakehouse if we can call it like that, both for internal use.
We also have the classic data analysis team and internal BI team. They all are utilizing the same data generated from Gong and others, we have data as mentioned product analysts that are doing a lot of internal work with the data. Overall, I would say there are about 30 different people working around these areas of data that is being generated, on top of Gong's platform.
Boaz: What data volumes are you guys dealing with?
Yarin: A lot. So, most of the data that is not being analyzed or being analyzed but not for data analysis is basically the videos and calls - what we call the media. The last time I checked recently, it was almost 5 petabytes, in terms of data ingestion. So we are talking about hundreds of thousands of videoconference calls per day. I would say many millions of emails per day. Overall, all the data that is not the media, I think takes around a few hundreds of terabytes.
Boaz: Wow! As you mentioned a variety of technologies. Walk us through, maybe, you know, if I am a user, from a user experience perspective, what kinds of analytics am I exposed to as a user? And then what parts of the stack are delivering it to me?
Yarin: The largest part of data, which we call team stats is analyzing how the team is performing in terms of sales calls. Gong has identified some very interesting metrics that help salesperson do better, things like the talk ratio, patience. So we have a very specific part of the product that serves this data, and then, we have some more advanced things that we can basically track what is being said in calls and track it over time. It is relatively new area that we are focusing on at Gong. What we call tracking the strategy or the sales organization strategy. Basically, all the data right now is being served either from PostgreSQL, in which all the data is pre-aggregated or from Snowflake where all the data is basically being modeled using DBT and almost no aggregation is done there. So it is query time.
Boaz: How did the data stack evolve in the year and a half that you are there? I mean you mentioned you are looking into a Lakehouse architecture right now. Is that a new initiative or was that always the case? And when did Snowflake and DBT come in and what was the driver for that?
Yarin: So the entire lakehouse solution is being in progress these days. We are trying to make it more of an evolution rather than a revolution, but still, a lot of the data is being served from the classic simple one big table on top of PostgreSQL, and now we form the new architecture basically, looking for, in terms of scale and building this new lakehouse solution.
Boaz: You also mentioned Redshift, right? Where does the Redshift come in?
Yarin: Mostly around product analytics. A lot of data that is being collected from either third-party tools ends up in Redshift anonymized. We do not have any customer data at all there. So, it is only just IDs and raw metrics.
Boaz: What are the typical frustration that you guys run into in the data stack? What about maybe the internal users like to complain about or what the new evolution may resolve?
Yarin: Users always complain obviously about performance, why this query takes so long, etc. A lot of internal users find it a bit difficult to get the real picture of data because we are using so many disparate databases and also when we try to bring everything under the same hood, data modeling still takes some effort because Gong is a company that moves really fast in the engineering, and sometimes we do not always have a cohesive data model, in terms of data engineering and analytics. So, these are two pretty interesting challenges that we are facing.
Boaz: Who is driving or who is involved in the evolution into the new architecture? How from a process perspective or a human perspective are you managing that project?
Yarin: Everyone, not kidding. The interesting part here is that it is not a classic internal data project, because we are also exposing some of this data, obviously model differently, but to customers. So, they will end up having their own access to a warehouse with all their data. So we have two areas pulling the same string, both from the product and from the internal analytics use case. It is almost everywhere in the organization, starting from customer success to product, to engineering, to pretty complex projects.
Boaz: I think you mentioned sharing data with customers. Can you elaborate more a bit about that? So what's the business need or requirement there and how are we going about it?
Yarin: I think I look at it as the next evolution of APIs. Gong has a pretty robust API, which companies can utilize and build their own products, but the biggest benefit of Gong is the data that it holds, and obviously, our customers want to build data products around Gong's data. So it only makes sense instead of consuming APIs and building the pipelines and data transformation is just hand them over all the data cleanly modeled in a way that they can either just plug in a Tableau or Looker and just start using the data, or just build products on top of data and not API.
Boaz: They are essentially exporting the data, copying it into their own environment through the APIs.
Yarin: Yes.
Boaz: Got it. What are your impressions so far you have been using both Snowflake and Redshift, how is that sort of looking for you guys? What's the conclusion>
Yarin: We are relatively new with Snowflake, but overall I think both have their strengths and weaknesses. Obviously Snowflake, now, in terms of a bit cost and Redshift is not as performant and ingestion is a bit more challenging. Overall, I think they are both okay for what they do. Some use cases are better here. Some are there.
Boaz: Is DBT used with both or only with a Snowflake.
Yarin: Only with Snowflake.
Boaz: And how long has it been since you have adopted DBT? How has that been going?
Yarin: Same pretty, pretty new for us as well.
Boaz: Any tips so far for other newbies with DBT - Things to avoid, best practices from the rather short journey for you guys there or not yet.
Yarin: Yeah, not yet. We are still in the part of, like these things, we are more comfortable with, but we are still learning as we go.
Boaz: Awesome! I wonder you being a veteran in software engineering and essentially now delivering data experiences into products, the intersection of software engineering and data, what kind of practices are important for you that are different from sort of the traditional internal analytics world
Yarin: Mostly it is around the quality of data, which we end up serving our customers. So we treat data projects almost as a software development lifecycle. What is the SLA for supports, monitoring? How can we make sure that the data that customers see, that went through so many pipelines and transformation, is accurate? Pipelines cannot break, so we cannot have data that is, "yeah, it's broken for the past 24 hours." So customers see data that is from the last couple of days. We treat it just like any other large-scale customer-facing software project.
Boaz: Let us maybe use this as an opportunity to pick on you and ask, was there any sort of a failure that you remember that we can look back and then maybe help our listeners avoid? What do you remember as a horrible day? whether it is data pipelines or production issues that we can all learn from.
Yarin: So far, nothing major customer-facing. We did have some mega failures trying to be very naive in terms of how we are going to approach this lakehouse project. We were very naive and said, yeah, let's like take something like a CDC solution that consumes all the possible data, throw it in S3 and then just have Athena or Presto or something like that and it is just a one month project and we goal in, but yeah, this one month was a battle, I think, for about 6 months ago and then, we decided to move on to something more robust.
Boaz: What is that more robust thing. Let us just untangle that a little bit. Let us spend some time there.
Yarin: It is let us not cut angles and just build a streaming service from within the products where we can make sure also in terms of data lineage. So we know when the data was modified and what kind of data was modified. So for example, if we are talking about recorded calls, even scheduled calls. So before the call is being processed and analyzed by Gong, we can know exactly when this call was changed and modified by the host, was it rescheduled, etc., and then, we write everything. We almost created our own CDC mechanism. We dump everything to S3, where we have a like staging area, I would say, being rapidly ingested into something like Snowflake, which is the journal for each entity. And from then on, we can create history tables or daily snapshots. We can go back in time and say, yeah, let us see how the actual data was looking a week ago. So, no shortcuts.
Boaz: How did you go about with the lineage challenge? How did you implement that
Yarin: We still have some challenges there, but we keep track of any transaction and data pipeline, which is internal to the product, and have a record that says that we know when was this entity modified and by whom. From then on, we try to keep a sequence or a batch ID and each transformation has its own records. When was it transformed, by which pipeline, etc? Eventually, we should be able to unroll aggregated data and find out the actual raw data that was taken into consideration for disaggregation.
Boaz: Tell us internal analytics, how does that look like? What tools are being used and which of the different databases and warehouses are they hitting?
Yarin: For internal use?
Boaz: Yes.
Yarin: We have Sisense. We now have Tableau. We have our homegrown analytics, which I am not sure at which front it is being used. Most of the data is being ingested to Redshift, either from third parties or being streamed from within the product. Pretty straightforward. We are using Amplitude for analytics. FullStory for UI analytics, like product user journey. Pretty straightforward.
Boaz: You know, we have picked on you for data tragedy. Now, let us talk about something happy. What are you proud of or projects that went very well or something that we can learn from if you can share?
Yarin: I think one example that I have on top of my mind is like we were looking to replace a PostgreSQL for one part of the product, which we served some stats, and before moving off from PostgreSQL, we map where the problem is? Why is it not working? Is moving on from 2 different datastore is the solution? and at the end of the day, we identify that with just a very, very minor optimization to this PostgreSQL, using RDS or Aurora, so we can more easily scale PostgreSQL, not that it is an easy task in itself, but we had a pretty big table in terms of what PostgreSQL can deal with, it was just research for about 2 weeks and then the solution took a couple of days, and we were back from like 10 seconds loading time to sub-second queries, just with using the right column type, the right indexes, and it was pretty interesting how far you can go with the modest tools.
Boaz: Yeah. Sometimes the simplest things work like a charm, but we oftentimes do not spend enough time figuring that out and go from a robust solution. Awesome! Great story! Maybe an interesting angle to talk about would be hiring. How do you go about hiring engineers from a data angle? How do you make sure they will be able to deliver on those big data challenges? What are you looking for?
Yarin: First, I would say it is a very difficult task nowadays, but yeah, we all can agree on that. In terms of data, we are looking for people that dealt with data, not necessarily at a data engineering level and know all the tools and how to build infrastructure, but have some experience around data, what it means to like query 2 terabytes of data? How databases and data stores are modeled and understand technically the strength and weaknesses of each. It is not necessarily the best SQL experience, but more of the understanding of performance, good data modeling. This is also something that we put some emphasis on. Have the good foundations, 3NF, Kimball, things like that.
Boaz: Awesome! Thanks! Are there any parts left in your stack that you consider very legacy and are the ones that are in line to be replaced?
Yarin: Yes. Although Gong is a relatively young company, the things that are implemented 2 years ago are considered legacy. I would say that we still store data, which is not for operational use in PostgreSQL is something definitely want to move on from, either to the lakehouse approach or maybe some other data store which allow fast analytics at a smaller scope, but at the faster scale and performance.
Boaz: Thank you. I think Yarin I am running out of questions. I mean, that has been super, super interesting.
Eldad: Yes.
Boaz: We love Gong ourselves at Firebolt, and it is definitely interesting to hear such a sort of insight-driven and data-centric product as data flowing behind the scenes. Thank you for joining our episode and any final words you want to spread out to the data engineering world?
Yarin: Yeah, I would say what we always do, one of our operating principles at Gong is first enjoy the ride and the second one is just, do not be so naive and say, yeah, put everything on S3, and then it would work. It takes a lot of effort and a lot of time and a lot of people along the way to create a robust data solution.
Boaz: But the first few days of that naive, false optimism, they are so fun. You think only worries are over before it explodes in your face
Eldad: Yes
Yarin: Feel so super-powered. Yeah! it will take us just 2 weeks
Boaz: Within your head, it is all done already. Yarin, Thank you very, very much and see you around.
Eldad: Thank you so much.
Yarin: Thank you. Bye.
Eldad: Bye-bye.
Boaz: Bye. Thank you for joining us.
# How Bolt Engineers Are Designing Its Next-Gen Data Platform (/blog/how-bolt-engineers-are-designing-its-next-gen-data-platform)
Bolt's ride-hailing app serves over 75M users in Europe and Africa and handles 500K queries every day. Erik Heintare along with Bolt's engineering team is in the midst of designing a new next-gen data platform and is sharing how it's going to solve their biggest data challenges.
Guest: Erik Heintare - Senior Analytics Engineer at Bolt
Hosts: Eldad and Boaz Farkash, AKA The Data Bros.
Listen on [Spotify](https://open.spotify.com/episode/73lYRwqhnEsPYzXn7UDTMG?si=w-HyTZFBT4-7Q8U2v7PaKw) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-bolt-engineers-are-designing-its-next-gen-data-platform/id1561927688?i=1000544916326)
## Transcript [#transcript]
**Boaz:** Hello!
**Eldad:** Hello!
**Boaz:** Welcome to another episode of the Data Engineering Show with myself, Boaz and Eldad.
**Eldad:** Hi, everyone!
**Boaz:** How you have been Eldad?
**Eldad:** I am good.
**Boaz:** Are you ready for a Hanukkahs around the corner.
**Eldad:** Yeah, absolutely!
**Boaz:** Yeah. Do you want to come over to light some candles?
**Eldad:** Light some candles!
**Boaz:** Yeah, do it!
**Boaz:** Okay. With us today, getting ready for Christmas time, I guess, is Erik Heintare. How are you?
**Erik:** I am good. How are you? Hi everyone!
**Boaz:** Thank you so much for joining us, Erik. Erik is joining us from Bolt. He is a Senior Analytics Engineer/Lead Data Analyst at Bolt. He has been there for over 4 years. Before we let Erik introduce us a little bit further, a few fun facts about, Estonia in general, where Erik is from. Estonia, a small country, only around 1.3 million, but is actually ranked number one in the world in unicorns per capita. That's very interesting.
**Eldad:** Woohoo!
**Boaz:** Also regarding Bolt. If you have not heard about the Bolt, means that you probably have not been spending enough time in Europe or outside the US. Bolt is actually a big competitor to Uber. Next time you hop off a plane in a European country, try not defaulting to Uber and see if you can get a ride a scooter or bike, even food delivery from Bolt. A Bolt is active in over 300 cities across more than 40 countries, by now. Mostly in Europe, but also Asia Africa, Latin America, and that is exciting stuff. Bolt was founded in 2014 in San Francisco. But in recent years, especially I think saw tremendous growth.
**Eldad:** Amazing intro.
**Boaz:** Am I wrong? Am I wrong?
**Erik:** You were wrong about the San Francisco part. This was actually founded in Tallinn in Estonia.
**Boaz:** Got it.
**Eldad:** Other than that, amazing intro.
**Boaz:** In the PR, once they opened the San Francisco office, it takes over everything in the PR. But thanks for the correction. It makes a lot of sense. So far, the company is still privately held and has raised over $600 million to date serving more than 75 million end-users. So, Erik, thanks. What did I miss anything about Bolt did I leave out?
**Erik:** No. Actually, maybe a couple of things about Bolt to add more is that we are actually not focused purely on Europe and I think we are the leading ride-hailing and delivery app, not only in Europe but also in Africa. So, we should be the biggest player in Africa as well. So, Europe and Africa together are like 2 billion people. Quite a big market actually.
**Boaz:** Tell us, what do you do at Bolt? Let us here it from you.
**Erik:** Just one more thing coming back to the intro part, then, when you mentioned Estonia and the facts you gave, then actually I prepared one of the fun facts about Estonia and it included that also that we have most unicorn per capita, but when I usually do this intro to foreigners, then I also say we have most supermodels per capita also.
**Boaz:** Interesting, we should investigate a correlation between startups and supermodels, Interesting!
**Erik:** But yeah. Great intro! I have been at Bolt for more than 4 years. Initially, I started off as the first data team member ever. Before me, there was no one working purely on data. I started off as a data analyst, basically trying to get started with our data warehouse together with engineers, try to automate and move away some queries from a production or a pre-live environment to the data warehouse. Also wanted to get rid of Google sheets, all of those things, and try to optimize everything we did in a business. Within those 4 years, I gradually, since we hired more analysts, moved to like analytics managers/lead analyst role, but then after some time, still wanted to be more hands-on. So, now, I am mostly focusing on our analytics platform, and still helping everyone in the analytics also, around all of those topics. So, yeah, this is what I do.
**Boaz:** Awesome! Let us talk about your title for a second because it is interesting your course of a senior analytics engineer, we all know data engineers, analytics engineering is something you heard less about. What do you think what that means?
**Erik:** I think it is the around tooling and the things that we actually do in the analyst world. So this analytics engineer, in general, does not mean there is a big difference between like analyst work or data engineer work. It is just I am more focusing on helping to build the infrastructure for our analysts. Meaning, like I think, the biggest user of analytics engineering is a company DBT labs and they are promoting it heavily. For us, we are not using DBT, but still, the work that we need to do is basically build out a platform for analysts, for business users, for product users to do their work efficiently, smoothly with high quality, high speed. So, I am somewhere in between data engineers and data analysts. So this is how we combine those two together.
**Boaz:** How big is Bolt in terms of headcount worldwide?
**Erik:** Worldwide, I think, we are more than 3000 employees and the data team just reached over 100 members.
**Boaz:** Wow!
**Boaz:** Tell us a little bit about that. What kind of data roles exists in Bolt? And how are they spread out in terms of departments, groups, teams, etc? And where do you fit in?
**Erik:** Yeah, we have 3 different roles. I think it is pretty common. We have data scientists, data engineers, and data analysts. Data scientists and data engineers belong to an engineering organization and are like a centralized group. They report to some higher management in engineering and data analysts, they belong to a product organization, mostly, and they report to product managers. So, they are way closer to the business and to have the effect and the knowledge about everything that is going on in the business. So, this is why we went with this hybrid approach. Initially, with, I think, 3 data analysts in the company, we were thought of going with centralized also for analysts, but it did not make any sense and we are quite happy with the setup at the moment.
**Boaz:** You report to the product as well and not to engineering.
**Erik:** No. Actually, I am part of data engineering.
**Boaz:** Okay.
**Erik:** So we have 4 soft teams in data engineering. Firstly it is a data lake, then it is data transformation, thirdly model life cycle and experimentation platform, and fourth is my team, which is analytics engineering.
**Boaz:** Got it. Let us talk about the data stack and definitely, a bit deeper. Before that, in terms of data volumes, what kind of data volumes does Bolt deal with?
**Erik:** If you want to be fancy and fancy starts, I think from petabytes, then we can say that our Redshift Cluster can handle more than 4 petabytes of data, but actually we are not there yet. I think our data lake in total is somewhere between half a petabyte and one petabyte, somewhere around that.
**Boaz:** In terms of sort of a daily number of events or daily data volume?
**Erik:** Yeah. So I am talking about from an analyst perspective, then, we do, I think, a bit less than half a million queries a day in our BI tools. Then, talking about, as we serve all of the models also as a platform then, I think we do 100 million model life cycles daily.
**Boaz:** Wow!
**Erik:** So, quite a lot, and, yeah, I think there are from those 3000 employees more than half of them use our BI tool also on a weekly basis.
**Boaz:** Okay. So, let's break that down. Tell us about the data stack a little bit, from bottom to top, how does it look like?
**Erik:** Yeah. So, we stream our data from live databases and services with Kafka to S3. Then, we have, of course, as I mentioned, Redshift as a data warehouse. In the S3 and Redshift, we do some transformations in Apache Spark or with Apache Airflow. Then together with Redshift, of course, we use a Spectrum, which is basically a layer to get data from S3 directly without storing it in Redshift. Then, we use Looker as our BI tool. We use SageMaker and all of the data team members use Jupiter notebooks. So, this is, just a brief overview and the really, really high level of what are the tools we use. I think, there is nothing, really, really epically different than the other companies are doing. So, I think it is a pretty common stack.
**Boaz:** How long has this stack been active? I mean, if you remind 4 years ago how did it look back then? And how is the journey?
**Erik:** So 4 years ago
**Boaz:** Google Sheet.
**Erik:** There was not almost anything. When I joined the guys, so cool thing, since we are over on AWS already back then, they were like, "Hey, what is this Redshift. Let's spin it up and see, maybe you can use it." Of course, we still use S3. So S3 was there. We did not use Kafka. So it was a bit different then. And then I think in Redshift when I joined, we initially had like top 10 tables, maybe from the live database, just like getting orders and do understand like where our drivers or how active they are just to get the first initial. I really remember when I joined, I needed to get some data and the engineers were talking to me is like, "we do it from the live database." "We acquire it from the live database. It is like, we do not have time to like a data warehouse. Why should I bother?" And now it is like thousands of tables. I don't know even how many together with Spectrum we have in Redshift. It is constantly improving. So yeah, what's going on there?
**Boaz:** And eventually, because there is Looker that the only BI tool or do you have other methods of visualizing data apparently?
**Erik:** For front-end analyses, we also use Mixpanel, but everything that we as a data team wants to have a better control and better structure and the use of backend data also then, Looker is our, by far the biggest BI tool that we use, Yeah!
**Boaz:** I think I saw on your profile, somewhere you mentioned the use of Amundsen as a data catalog, is that is in use?
**Erik:** Yes
**Boaz:** Can you, maybe, tell us how you guys are using it?
**Erik:** Yeah, I would say it is in the alpha stage for us; but basically, the company is growing so fast. There are more people coming in. There are more tables popping up all the time and we wanted to get a better way to scalably share the information that we have around our data and data discovery has been one of the bottlenecks of scaling. So, we started off a trying out actually two different tools. So firstly, Knowledge Repo, which is basically where you can host; it is a repo filled with Jupiter notebooks where you can add some metadata on top of it and also, Amundsen. We currently use it in a quite limited scope, so we have all of the Redshift tables there. We have key tables, key columns, key schema all, like commented and we have attributed owners to them. So at least if you search for something, you want to know what is going on with food delivery couriers, and you are the first time doing it from the marketing perspective, then at least you will be easily able, with one search, to understand where are the tables? How they are structured? Whether they are like key tables, and who is the owner of those tables? from a data team perspective.
**Boaz:** Can you maybe share how does Bolt approach the data engineer versus analyst relationship? How do you guys make sure that, you know, things are delivered quickly, and the analyst and the other business departments can sort of become, stay self-sufficient and move fast. Has that typically been a friction point or any insights you can share on how you guys do that?
**Erik:** Yeah. I think it is also pretty common; but of course, we struggle all the time either to build the platform to be better and scalable for the future or support the current needs for analysts or data scientists, so they could move on with their projects. So it is a constant battle between the prioritization. What we have done is that periodically we have changed our focus. So, for some time, we focus more on the product requests, making sure that everyone gets their stuff quickly and efficiently. And, then we communicated to our users also, for the next two weeks, we are heavily focused on building our data platform. So all the requests that you have, we will only maybe solve the most critical ones. So please mark them accordingly. It is never easy, like products always try to push their own needs first because why would you bother optimizing something in data engineering. For them, it does not give any return on investment, but their feature, which they want to launch, of course, it is super easy to put a number to it. So it is a constant battle. Of course, we communicate a lot together with a team, so this helps to understand the priorities, and this is the only way to go.
**Boaz:** Let us talk about some of the use cases. So, you know, there is the big Redshift at the center, according to what you say. So what are the sort of uses cases that are run on top of it?
**Erik:** Yeah, I think, Redshift's biggest abuser is Looker. As I said we do quite a lot of queries from our business users, from our analysts. They are the main users. Of course, we have our experimentation platform using it. Then, we have our model life cycle which is basically they are also using to re-train their models, like profiling services fetches some data from it. So basically everyone fights for the spot in a Redshift query queue.
**Boaz:** Can that get ugly sometimes?
**Erik:** Oh, yeah! I think I will leave it to the latter question, maybe potential about the glorious failure.
**Boaz:** Okay. So, let's do a quick stop and move to something funnily quick - the Blitz Question Round. Are you ready?
**Erik:** Yes.
**Boaz:** Do not overthink. Let us see.
**Boaz:** Write your own sequel or use a drag and drop visualization tool?
**Erik:** Write your own sequel.
**Boaz:** Looker or Tableau?
**Erik:** Looker.
**Boaz:** Commercial or open-source?
**Erik:** Heart says, open-source, head says commercial.
**Boaz:** Batch or streaming?
**Erik:** Streaming
**Boaz:** Work from home or from the office?
**Erik:** Hybrids
**Boaz:** AWS, GCP or Azure?
**Erik:** AWS.
**Boaz:** To let people self-serve for analytics or not bother?
**Erik:** Let self-serve.
**Boaz:** Bolt or Uber.
**Erik:** Bolt.
**Boaz:** Boom! I like that.
**Erik:** I did not even leave you time to finish your question. I already said, Bolt.
**Boaz:** Very Good!
**Boaz:** What are the biggest challenges though, with the current stacks? What are your top priorities for next year given what you guys have today?
**Erik:** Since the current stack has been around for more than 4 years and actually, we are currently moving in the phase of finalizing POCs for that next-generation data warehouse, so we are potentially either getting rid of Redshift, they are replacing with some other or adding something on top of the current stack or just playing around with all the different things that are available right now to make sure that we are enabling our users for the next 100x growth also. So, this has been a big focus for us in the last couple of months and, we will try to finalize it during this year and next year, will be where we basically work hard on making sure that we have that next-generation data warehouse ready.
**Boaz:** So you are going through a traditional evaluation process?
**Erik:** No, we are using some of the methods that I have been out there and a lot of companies have been doing their POCs. So, of course, would take ideas from there, but what we do is we also apply a lot of our internal knowledge and let's say like we gathered a lot of internal queries, which are heavily used at the moment and we wanted to see how they would be formed. Because there are some tools that are really, really good when you just need to query data from one table and there are some other tools that are really good at joining together the 20 tables, so what is the best for us? We try to cover all of those things, and, yeah, it has been a heavy job to do those POC and kudos to all of the team members who are doing this. It is not an easy decision to either switch out or keep the current state. So you are only responsible for the most used and valuable asset, but then subsequently will become familiar.
**Boaz:** Any particular technology or feature that is out there that you really are upset, you don't have access to today or you would like to have in your next-gen platform?
**Erik:** I think one of the things that pushed us to move maybe a bit faster with this evaluation process was that we are currently hitting some of the limits when we talk about concurrent queries and that was basically peak hours. So this is where we struggle the most at the moment. Then, this is what we want to solve as quickly as possible, because if business users cannot see the data or need to wait for, I do not know, 5 minutes for it, then it is not worth it.
**Eldad:** So you have products blessing.
**Erik:** Yeah! Definitely!
**Boaz:** So I am assuming the coupling storage and compute will be a big deal.
**Erik:** Yes, yes! Most likely.
**Boaz:** You had mentioned the glorious failure prior. I really do not want to pry, but you know, since you brought it up, tell us about some glorious failure you remember.
**Erik:** Maybe, I hyped it up too much.
**Eldad:** Too late.
**Boaz:** Too late, you have to.
**Erik:** But yeah, one of the things we had, so basically it is a combination of multiple things. Looker, quite recently introduced native integration with Google Sheets and Google drive, which means it is easy to set up a spatial from Looker to those bases. We did it. We enabled it and we thought it is like good thing to have. Of course, business users, you will never ever solve all of the cases. Like people always would try to copy your data tools and Google Sheet and add some manual stuff on top of it.
**Eldad:** Sort the data.
**Erik:** So we thought doing this is a great idea. It reduces manual workload and all of those things, but what we did not realize was that due to some limitations from the Google sheets and Google drive APIs, the scheduling takes a lot longer than, like scheduling slack message or email. So what happened was, people were too happy about this and they set up a lot of things, of course, to the Monday morning together with all the other reports that are running on Monday morning. So, basically what happened was, that we had like, I don't know, hundreds of new spatial randomly popping up on Monday morning all of a sudden, none of our overnight schedules or works did not finish. We did not know what is going on. We saw that there is a load. It was really hard to estimate, like, what is causing it? Is Redshift working slower? Or is it because of Looker doing something wrong? So it was like, we could not find it out on the first Monday. It took us 3 Mondays to actually solve it. And it was a combination of many things, as I said, like from concurrent scaling from data warehouse side wasn't performing as we expect them and Looker was not working as expected. So basically for 3 Mondays in a row, our company users could not query the data in the first 4 hours.
**Eldad:** How did it affect your weekends?
**Boaz:** If it highlights on Monday and a little bit on the Tuesday and Wednesday. Thursday, got it right, had a weekend. It was okay and then Monday, all over again.
**Erik:** Yeah! This is exactly like, by the lunch of Monday, we were seeing like, okay, we killed the top queries. We did a couple of adjustments. It looks like it is working. Let's see how it works on next Monday. And, yeah, it didn't work then as well, of course, and then it was easy to get the priorities, from the product also to speed up all the POCs for the next-generation data warehouse.
**Boaz:** This goes back to the blitz question – Do not let people self-serve because they will over-schedule stuff on their own and ruin your Mondays.
**Erik:** Well, I still say self-serve. Yeah.
**Boaz:** But to care for these, care.
**Eldad:** Gradually.
**Boaz:** Another takeaway, you know, do the crazy stuff on Mondays and not on Fridays.
**Eldad:** Exactly, respect a weekend.
**Boaz:** Okay. But, you know, let us not be so negative. What about, positive stories? Tell us about the great win.
**Erik:** So, since I have heard a couple of episodes before, then I immediately started to think about it then. Well, there are a lot of wins that should be mentioned. But one of the things that I thought about was actually when the COVID hit first and then basically, we lost 80% of our company's revenue and, basically, we got a request from top management, that is like, Hey guys, can you scale down a bit, maybe 60% or so with all of your data infrastructures and actually what we did was we managed to reduce all of our infra costs, more than 50%, within like a couple of weeks, and it did not affect the end-users that much, everything remained to work as is, so basically it gave us a huge boost moving forward. We did quite interesting things there and yeah! from that, we learned a lot, and thanks to that we are now a way lower level than we would have been before the COVID hit.
**Boaz:** Can you share some of the details? How did you go about reducing costs?
**Erik:** I do not think there is only one thing to point out, just on a high level, I think that one of the quick switches was that we moved away much of the tables from Redshift to Spectrum that were not used that heavily. So, this gave us, like, we could easily have a smaller Redshift cluster, everything that I needed to be, like, get out from S3 would be used easily through Spectrum. But, I think there were more than 20 items that helped us to get this low.
**Boaz:** Wow! Impressive!
**Boaz:** Goodwin, goodwin!
**Eldad:** This is how you scale. This is so well spent the time
**Erik:** Yeah!
**Boaz:** Okay. So I think, we are reaching the end almost, maybe before we end, we would love to get a few recommendations from you from maybe recent technologies that you have had a chance to play around and got you sort of interested or excited, even if you do not have adopted them fully. So any tips or things that you ran into recently that you want to give advice to our listeners?
**Erik:** I do not want to point out any specific technology or tool specifically but during those POCs what we found was that be ready to change your perspective on some of the tools? So, maybe you checked out something 2 years ago. Maybe you checked out something 6 months ago. Nowadays, everything moves so fast and, it could be that six months makes a huge difference to the product. So, what I would highlight is like, if you, reject a something, like one or two years ago, and you still have this issue in your hand, take a look at those candidates also, which you rejected back then because so many things are happening in this space and huge, huge improvements all across the board. But, yeah, we saw some, like we needed to change our perspective on some tools.
**Boaz:** I do want to also close the loop on the topic we touched on prior, going back to that big Redshift cluster and how ugly it can get when so many people want to get access to the resource. How do you manage that? How are decisions being made in terms of prioritizing, who gets what?
**Erik:** You mean who gets what? like query prioritization or access to that?
**Boaz:** Query prioritization because so many workloads running.
**Erik:** For us, since Looker is the main user and the majority of the business users use it, so we have prioritized quite a bit to the business user side, so it is fine that some of your, I don't know, experimentation platform can run instead of 5 minutes, 7 minutes; this is fine. But for Looker, if it runs like a couple of seconds or 2 minutes, then it is like a huge, huge difference. So, we prioritize our business users first and also like what we have seen is also coming back to the previous point about changing perspective, then we tried to solve like a lot of this prioritization manually for quite a long time. But what we see is that a lot of tools that have, like internal auto prioritization, auto-scaling all of those that are basically built into the product, then, in the long run, I see that all the auto settings might win. Same goes with, I don't know, Google or Facebook ads from different industries. It is like, you can fine-tune it yourself however you want. In the long run, they will have more data on how to optimize all of this. So you can trust them in the long run. This is also what helped us. And also it is like, you do not need to have the manpower to it. You do not need to, I don't know, monitor it so actively that if they are screwing up or not.
**Boaz:** Also regarding the Kafka. Sorry for suddenly reminding myself of so many questions when we are moving about so close to finishing, but Kafka, are any of the reports actually, you know, closer to real-time or is everything more batch-oriented since using Kafka? What is looked at in a more sort of real-timeish fashion?
**Erik:** Together with the next-generation data warehouse that we are planning, and actually there are two side projects to it also, what we call, like next-generation reporting system, for batch and live. So, in here, we are building out infra to serve not only internal use cases but also external use cases. All of the engineers, I do not know, show some numbers in the client apps or restaurant apps or courier apps and also a batch reporting. If you are a restaurant and you want to get their weekly results, so we are also working on it to get it live already, first half next year. We are not doing that much live reporting at the moment, but we have a dedicated team working on it, so we could have it and also, maybe not necessarily a Kafka live, but still one of the next-generation data warehouse goal was to reduce the ingestion lag also to get it more closer to the real-time.
**Boaz:** Got it. Awesome! Thank you! Erik, this has been super, super interesting. Thank you so much and yeah, it has been great having you.
**Erik:** Thank you. It was great.
**Eldad:** Yes, it is.
**Boaz:** To see around the data world and when we visit Estonia, we will make sure to stop by for a coffee.
**Erik:** Sure! You're welcome!
**Boaz:** Take care!
**Eldad:** Take Care!
# How did Agoda scale its data platform to support 1.5T events per day? (/blog/how-did-agoda-scale-its-data-platform-to-support-1-5t-events-per-day)
Scaling a data platform to support 1.5T events per day requires complicated technical migrations and alignment between hundreds of engineers. What to see how Agoda did it.
Listen on [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-did-agoda-scale-its-data-platform-to-support-1/id1561927688?i=1000542833635) or [Spotify](https://open.spotify.com/episode/3XwPJcJElAlU0O8KzIn7tT).
**Boaz:** Okay, so ready to get started. Eldad! Are you ready?
**Eldad:** Yes!
**Boaz:** Okay. So hello, everybody.
**Amir:** Hello, everyone!
**Boaz:** Welcome to another episode of the Data Engineering Show presented by the Eldad Farkash, right here!
**Eldad:** Hi there.
**Boaz:** And Boaz Farkash, that is me! We are related. We have the same parents, which makes us the data bros, woohoo! So, with us today, from Agoda, we are lucky enough to have been joined by Amir Arad - Director of Machine Learning and Shaun. Shaun your last name, please?
**Shaun:** Shaun Sit.
**Boaz:** Shaun Sit.
**Shaun:** Yeah.
**Boaz:** I had a hard time finding it because your name overlaps with a lot of options there, but I did find it eventually. So Shaun Sit, a Senior Dev Manager at Agoda and currently managing the data platform. So, we are having you guys on the show and, you know, I wonder before we start just, at Agoda, because you are in traveling and all that, is it frustrating working around travel data all the time and you guys are at home? I mean, don't you want to travel all the time? Are there any travel-related perks? How can we get some inside information? How it is to, you know, be frustrated, but enjoy travel life at Agoda?
**Amir:** So yeah! You must travel in order to be working in Agoda. It is not allowed to be stationary. You get fired if you stay at the same place. And lots of perks and everybody at Agoda travels all the time, kind of hard during Covid.
**Boaz:** So, yeah, exactly. So bomber for those who joined Agoda during Covid and did not enjoy perks. Any recent, exciting trips that you guys were on?
**Shaun:** I recently went to Phuket. It was lovely. Absolutely lovely. You guys should come whenever travels open up.
**Boaz:** I'm jealous. I am jealous.
**Eldad:** A few moments of silence!
**Boaz:** I recently came to the office after working from home for a long time. That was refreshing and I went back home and then went back to the office, not fair!
**Amir:** Similar, right?
**Boaz:** These guys go to Phuket and all these exciting places. Okay! So, let us start with a short intro about you guys. Tell us what you do at Agoda beyond the fancy titles? Who goes first?
**Shaun:** Sure! I can go first. As you mentioned, I'm Senior Dev Manager at Agoda. I manage the data platform teams here. I am currently managing four teams. They together manage the entire, deal with anything that is data-related in Agoda itself. So, the four teams managed the pipelines, our data lakes, the self-service data applications that we built for anyone in Agoda to use; and then lastly, we have a team as well that manages the UI, like creating UI that make the experience cohesive for all these tools that forms. So that is what I do.
**Boaz:** How many people are all of these teams combined?
**Shaun:** In my area, we have about close to 30 people.
**Boaz:** Awesome! And Amir, how about yourself?
**Amir:** So I am complimenting Shaun's effort by doing the machine learning part. We have the machine learning platform that we do in Agoda and a lot of tools and good stuff that we do to have the data scientists get their pipelines and get their models in production. And other than that, I have a few teams that do actual business applications, like personalization, marketing efforts that use all of the amazing data platform tools that Shaun's builds to improve the experience of the Agoda customers.
**Boaz:** And how many people over there?
**Amir:** It is changing all the time. But, the entire data platform, at Agoda, we have 4000. I think at the data platform, we have 300 people.
**Eldad:** Wow!
**Boaz:** Wow. Wow! Okay! That's great! So, let's talk data. I mean, at Agoda, I would imagine data volumes are through to the roof. What kind of data volumes are we even talking about?
**Shaun:** If you're asking about like, messages on Kafka, which is our main data pipeline, like how data moves around in Agoda itself, we do about a trillion or so messages a day.
**Boaz:** Okay!
**Shaun:** And then if you're talking about data lake, that comprises about like tens of petabytes worth of data.
**Boaz:** So Amir, on the ML front, a typical slice of data challenges you guys look at, how far back do you go? How many events are you looking at? How much data volume?
**Amir:** In terms of, let's say predictions done daily, we just passed the 60 billion predictions we do per day from our models. And all of that is based on historical events and future predictions that we make. So huge pipelines processing, billions of rows every time.
**Boaz:** So let's break down, you know, the data stack for a second. I wonder Shaun, you know, maybe if you could elaborate when Agoda has been around for quite some time, what does the data stack look like? Or how many data stacks do you guys have? And how do you evolve through the years? How does it look like today and how distant is it from how it was in the past?
**Shaun:** I think today, like I mentioned, the data pipeline we are using Kafka for that. We are using Elasticsearch for logging. We have Grafana with a custom time-series database called White Falcon. It is built in-house. Our data lake solution is HDFS. Then, we have Yarn, Livy, Uzi, Spark, a lot of custom ETL tools. Then, we have also some data governance and discoverability tools. We have a Schema Registry that works nicely with Kafka. Data Market is our discoverability tool. And then, we have custom-like data validation, data quality tools as well. In terms of queering, we are using Impala and Vertica. So that is our ad hoc query story. And then, for UI, we are using Hugo and a custom unified data portal. The team that I mentioned earlier makes everything more cohesive. And then visualization tools, we have Metabase and Tableau, some custom dashboarding stuff for funnels that we built in-house as well. And then that's the data stack from my side. Amir, what is the machine learning side look like?
**Amir:** Yeah. First-of-all, we are using a lot of this; and on top of that, we have some cool in-house build stuff, like a notebook platform that will be Python Opstrat tooling. We are using MLflow, which is a very cool model life cycle management tool, and a lot of Spark jobs also that are used for machine learning as well.
**Boaz:** Are you guys on the public cloud or you guys are self-hosted, self-managed?
**Shaun:** We are On-Premise.
**Boaz:** Everything On-Premise.
**Amir:** We live inside that the data center. We connect the cable ourselves...
**Eldad:** So, that's the real background for a second. Yeah!
**Boaz:** That's why you guys travel all the time. The travel perk includes you must stop a data center and do wire.
**Eldad:** Minus 15 in New Jersey data center.
**Boaz:** As I mentioned, a lot of homegrown tools, maybe a time series data. How do you call it? Falcon? Did you say, Shaun?
**Shaun:** White Falcon.
**Boaz:** White Falcon, first beautiful name. The other names are not as impressive
**Eldad:** Almost as good as Falcon.
**Boaz:** That was a good reason to start some products yourself, you can name it yourself. You get straight to white Falcon, so why not go for something off the shelf?
**Shaun:** I think it is for several reasons. We obviously do try out, like other times you use databases we have always take a look at what is out there, benchmark ourselves around those, come do the feature set comparison. It always comes down to a cost, performance ratio. We think that building it in-house, and we have had this for a long while, and it has been very much key enabler for us to store huge amounts of application matrix on it and across multiple data centers that we have around the world, and also, so it's pretty good. It's pretty great. Shout out to the White Falcon team!
**Boaz:** Awesome! The data stack, I would imagine, serves a lot of use cases. What are the most interesting ones or the bigger ones running on top of the platform?
**Shaun:** The interesting ones, obviously I think, would be, I want to say, Data Market. It is kind of similar to, I guess, DataHub or Amundsen from other companies. I think the DataHub is from LinkedIn and Amundsen is from Lyft, I believe. Did I get it right? I don't know. But basically, yeah, Data Market, it is our discoverability tool that has been instrumental in our data democratization story. We send a lot of data. It does not make any sense if nobody uses it. So, we got to make sure that there are tools out there that they make it so that it is easy to find the data that you are looking for, it makes sense, it is of high quality, and it is usable. It has really been one of the main drivers for us for data usage in a company.
**Boaz:** When was the project launched, sort of, how long has it been in the base?
**Shaun:** I want to say, let's see, 2 or 3 years ago, that is when we built it. Yeah!
**Boaz:** In your current stack, how much would you consider sort of moderate? I am happy with versus legacy. We are always in the process of, sort of, trying to modernize.
**Shaun:** I think that is like a tricky question, I guess, for any engineer. You are never fully happy with a solution that you have. You always want to improve.
**Boaz:** I never met a happy engineer. They are never truly happy. I was always almost there, almost there.
**Eldad:** Always forward-looking.
**Boaz:** Yeah!
**Amir:** It is job security, right? We are never done.
**Shaun:** I think for us, areas which we definitely can improve on, I would say is, around the area of decoupled storage and compute. I will be honest we are a little bit behind in that regard. We have not made that shift, and that is something that we are actively working on right now and it is going to give us that next-generation data platform. So that is something that, yeah!
**Boaz:** How do you go about such a project with those data volumes? How long does it take even to evaluate given the massive lift and shift it would involve?
**Shaun:** So that is one of the pain points, I guess, when you are work data. I think, the biggest, pain point is always migration. All the technologies are pretty cool, like everything; but, in order to use it, you have to migrate from what you are using currently to something new, Right? And that is where things become fairly complicated. Migrations always take extremely long times. I guess the key here is to plan, like plan, plan, plan, plan, plan, plan, plan, plan everything, and then, try and move forward as quickly as you can, as fast as you can, figure things out along the way and then, just adapt to the situation. I would say migration is always the biggest pain point, the biggest challenge.
**Amir:** Yeah! And also in Agoda since we have so many different types of data users, like the data scientists that can write their own code and wizards and they do everything on their own. And it could be like BI analysts that only know how to do SQL and then during these migrations or when you do these changes, you need it to be seamless for them.
So that makes every change even much harder because nobody should even know that something changed behind the scene. Usually, it is impossible.
**Eldad:** The users are basically preventing you from making progress and moving forward.
**Amir:** Exactly. We always ask to just the eliminate the user.
**Boaz:** Which teams will be your early adopters for new tech? Do you have a sort of some internal teams that typically champion for going next-gen, even at the expense of painful migration?
**Amir:** Machine learning, Always!
**Boaz:** Machine learning, always.
**Amir:** Like they get the GPU, the SSA to it and they already started like downloading stuff from the internet, run the latest GPU stuff or TensorFlow.
**Boaz:** If you look at your tech stack evolution, and today you are On-Premise, does that involve moving, becoming hybrid? Does that involve moving to something like S3? Because I guess storage for you is a big thing and a big part of the challenge of migrating somewhere, or how do you think about it moving forward?
**Shaun:** It is an interesting question. I think that the thing is we have always explored the cloud, right? We do constantly explore the cloud. We do have some stuff running on the cloud; but for data specifically, we have always done our research and POC, and we have never found the right motivations or the right reasons or the right…, how would I call it? Like the right thing…
**Eldad:** It is the one big thing that helps everyone make that from transition.
**Shaun:** Yeah, Yeah! Exactly, we would never find that big push, right? that pushes us in that direction. So far, we are pretty happy with On-Premise. I think also the hardware game has changed quite a bit. CPU is now the bottleneck, right? Storage is becoming way, way cheaper, and way faster. Similarly, with networks, right? And so being On-Premise, it does give us some advantage, right? to kind of leverage this system as they come along and to explore them, so I think I would not say one is better than the other. Always, I think in anything, it is just do what makes sense for you. It is just that in Agoda, we have the right people, the right expertise, and the right history as well. We came from On-Prem. So, we have a lot of knowledge in that area and so far, it makes sense for us. That is to say in the future if the cloud makes more sense, we would be 100% on-board. At this point of time, we are still On-Premise, but we are constantly exploring though.
**Eldad:** Makes perfect sense.
**Boaz:** Amir, What about you? Which use cases are the ones that are of the highest-profile?
**Amir:** Yeah. So I think for us, one thing that we got maybe a bit late in the train is like, we were mainly a Scala shop, so we are doing a lot of Spark job and huge parallel. We have jobs using like 13,000 cores for five hours; amazing huge jobs. But people had to be like Scala's expert in order to tune them, in order to get used to them, in order to drive them. So I think, we were kind of late, to see that Python is now really the go-to language for machine learning. And, so we kind of regret not building tools for that sooner; but now, we are already on par. I think very soon Python will win over Scala in at least in the use cases of data application. A lot of, let's say, the less advanced user are already write their scripts and their notebooks on Python and it made like the time to market data projects a lot faster and lot sooner. So that's cool and this is one thing that I think we tuned on maybe sooner.
**Boaz:** From a user perspective - I am a user or an Agoda customer or an Agoda visitor what kind of things happen in the background that sort of start with the melting the time, unaware of even? Can you share some cool things there?
**Amir:** Yeah, so we do a lot. If you and I both opened the Agoda website, we will get completely different experiences based on the past. Like, if I like breakfast and I like breakfast, I will see more photos of eggs and other breakfast stuff.
**Eldad:** Last time we got bacon all the time on your website.
**Amir:** Exactly, exactly.
**Amir:** So, we do try to fit the content to what we think you would like and what you care about. If you are more sensitive to price, then you will get the best offers. We anyway have the best offers, right? But then if we know that that is what you care about in your current trip, then the whole experience will be optimized for that. We do, for example, try to cut snippets from reviews that make sense for you. So if you care about cleanliness, then we will take the review that talked about it, Hey! this hotel is very clean and then, we show that one too. So, a lot of personalization effort goes there, but even before you came to Agoda, right? All of the marketing that goes behind the scene, the email, the popup notification, everything is kind of optimized to make sure that you get what you want and the information you need on the Agoda Website.
**Boaz:** What typically does the initiative for a new ML-based project come from? From your team? Is it sometimes product-driven? How does the thinking around new projects for ML look like?
**Amir:** Yeah! So, that's something that I think we do very coolly. The scrums that we have at our machine learning; they are always a mix of ML engineers, data scientists, and the PO together. And then it's like a 300 beast that kind of try to set the direction. Sometimes it comes from the data scientists, they say, Hey! this is something that we can easily optimize. Sometimes it is the PO or PM that can say, Hey! the business should go that way. And sometimes it is the engineering manager that can say, Hey! other teams did this or we see any other products doing that. So kind of a mix of ideas coming from three different directions, and then get swirled together and the best one wins.
**Boaz:** Interesting. Thanks.
**Boaz:** Shaun, I saw a piece you published a few months back on a medium called "How Agoda manages 1.5 Trillion Events per day on Kafka,"
**Shaun:** Yeah!
**Boaz:** Can you share the backstory there a little bit? Was this published to follow some sort of architectural change or just something that was there for so long and you decided to pass the knowledge out there?
**Shaun:** Yeah. Yeah! I think I wanted to write a blog piece and contribute to the Agoda Tech blog. So I thought this would be a good topic. I think it is an interesting thing because if you read the blog post, it is not so much only about the technical stuff, right? A lot of it is about the human process because you have to remember that at Agoda there are thousands of employees, right? So, certain things you need to think about in terms of scale, not only the technologies that they scale well but the human processes scale well as well. And I think that is something that sometimes we do forget as data engineers that you have to ensure that the human process is scale, so that is a lot of the things that drive towards how Agoda manages at 1.5 trillion. There is all stuff around there like cost management, attribution, even simple stuff like an auditing and monitoring and giving developers the confidence that what they send is exactly what they will receive and having them the self-service ability to just check on those kinds of stuff on their own and then the cost attribution as well, I think, that is a major portion that really allows developers to manage on their own their costs, right? You always have to think about, after all these are company resources. You got to make sure that what you are sending has some business use case, right? You do not want to just send stuff and make the data lake into a data swamp, right? So that is not very useful. So Yeah!
**Boaz:** It's super interesting. So how do you foster a culture or a way where developers are minded to that? Is that something that from day one their mentor to think about? or which roadblocks do you put in place to make sure that it does not get avoided?
**Shaun:** I think there are several ways. One way that we figured out that kind of works pretty well, like we just took a page out of the cloud providers who charge you for every single thing, right? So like, if you mess up a query, you end up paying for it. How much data you store, you pay for it. So I think just building that visibility to allow the developers to see, Hey! this is how much you supposedly will cost a company for sending this much of data or, you know, processing like maybe some, unoptimized query to the actual query engines that we have. So that itself has, you know, driven to make sure like, oh! you know, there is a cost allocated to all of these actions and I have to take that into consideration as well.
**Boaz:** How distributed is the engineering team in Agoda? Which locations are you guys spread out through?
**Amir:** So currently it's a lot more, right? Because we are working from home, so a lot of people are spread around the world; but usually, we have three main hubs. Bangkok is the biggest one where most Agoda seats there. Singapore is also big and we have a small office in Israel that is right next to you. You can jump, maybe eat a sandwich there. It is where we have a lot of smart data scientists there. I think these are the main three and now, we also opened another office in India to increase the diversity and the strength of all developers.
**Boaz:** Got it. So, what are your main challenges today? What takes the most sleep from your day-to-day? What do you worry about?
**Amir:** A lot of things.
**Boaz:** We are not talking about your personal life. We will take that offline.
**Amir:** Ah! Okay. Okay! That's right. I think one thing we find hard, I mean the success that Shaun talked about, kind of making data liability and people kind of owning up to understand that they cannot just send as much data as they want. This is one thing but still, when you try to change the way that people work with data, for example, in machine learning, each team is working with raw data sometimes with just sending SQL or building data frames. And we try to shift everyone to move to more like a feature-solid approach where you first model your data as a feature and then you start thinking this abstraction, okay! this is my feature, how it behaved, how it looked like a month ago and making this change one thing by the other, this is usually that. Like, I wish I could just press a button and everybody in Agoda would just move to work in that way, right? But usually, it is a process that takes a lot of time.
**Eldad:** By the way, this is huge and super interesting and actually one of the biggest things that are happening to engineering, to cloud-native companies or data-driven companies actually moving from engineering, building software to having engineers directly connected to the business feature or building and its cost, its value, its new terms, it is kind of broader thinking on how engineering is done and it is fascinating. So yes! It is everywhere. I can say in Firebolt as well. This is a big thing and being transparent and opening the data and giving visibility, as you said, a payload, an engineer generates payload and that payload ends up at the hand of the user in some form and that changes a lot of things. So, thanks for sharing that is super, super interesting! We should actually drill down on that on future shows as well.
**Boaz:** Yeah! Shaun, what about yourself? What keeps you up awake at night?
**Shaun:** It is interesting, for me, I think what keeps me up awake is trying to figure out what the next generation data platform will look like and trying to see if we make the right decisions, if we have made the right bets along the way. Because in the data space it is not like, yes! we are agile, but it is still like things take time to migrate. There is some time element involved. So that's what keeps me up, right? Is object storage the way to go? Is it not. Is distributed file systems coming back? You know, these kinds of stuff, right? Where is Hadoop going? Where is Yarn going? Right? All that kind of stuff.
**Boaz:** By the way I'm going to reference your LinkedIn account for the second time, looking at your LinkedIn account, your top-line message there says, "I'm hiring, we're building our next-gen data platform. Come join that." And I was wondering if that is there for 5 years or 1 year or a few months.
**Eldad:** It's a good tagline.
**Boaz:** It is always true. As you are thinking about your next-gen platform regardless of what that really means because that could mean a lot of things. Well! What is the objective? I mean, what would you like to plan for to be capable of in 4 years than is it now? Is it more about being future-ready or do you have concrete challenges you want to solve in the near term?
**Shaun:** I think one of the things that we want to solve is the agility of the systems, the agility of the architecture. If you think about the Hadoop space, a lot of things in the Hadoop space are very coupled together, right? The Yarn is coupled with HTFS. It is like Uzi only works on Hadoop, right? A lot of things are very coupled together. So, I think for us, like for me, I am less worried about which systems we ended up picking as long, like in the future, we are in a much better position to have that kind of agility to change something out whenever we need to without incurring that huge long timelines of a migration, right? So that's where I want to be.
**Boaz:** Got it. Okay! Now, we are going to do a blitz question round. We are going to ask you a few questions real quick. Don't overthink, just answer and feel free to cut into each other's answers because there are two of you.
**Amir:** I was waiting for that from the beginning.
**Boaz:** We will count counter whoever answers first, you know, but we love more than the other. We are like the parents loving one kid more than the other. Okay! Let's start - commercial or open-source?
**Amir:** Open-source.
**Shaun:** Both. I say both, do what makes sense for your requirement?
**Boaz:** That's like cheating.
**Eldad:** That's because of Vertica that is why both.
**Boaz:** Batch or streaming?
**Amir:** Batch for now.
**Shaun:** I'm going to say both again. No, the company will use one or the other.
**Eldad:** If you had to choose one, if you had to choose one.
**Boaz:** No, it is something like what makes you feel better?
**Eldad:** Exactly. There is no good answer.
**Shaun:** Okay.
**Boaz:** Are you a Batch person or streaming?
**Shaun:** Streaming.
**Amir:** Exactly.
**Shaun:** From Kafka streaming.
**Eldad:** Everyone wants to be a streaming person.
**Boaz:** We need to highlight the instructions a bit more for Shaun. Shaun! Don't overthink the answer!
**Shaun:** Okay.
**Boaz:** Whatever you are feeling is right for you. Write your own SQL or use a drag and drop visualization tool?
**Shaun:** Drag and drop.
**Amir:** Write your own.
**Boaz:** ML team versus the data development team - interesting insights! So far, no answer that was the same.
**Eldad:** No.
**Boaz:** We are not getting on anything.
**Amir:** Everyone wants to make it interesting.
**Boaz:** Work from home or work from the office?
**Amir:** Work from home.
**Shaun:** Home.
**Boaz:** That is the first agreement.
**Amir:** Yeah! But with a little bit of office here and there.
**Boaz:** To Uzi or not to Uzi? You mention the Uzi so often.
**Eldad:** The original question was - AWS, Google Cloud, or Azure? So Boaz changed that.
**Boaz:** That's true.
**Amir:** Yes! Uzi.
**Shaun:** No Uzi.
**Boaz:** No Uzi. Why not?
**Eldad:** To couple.
**Shaun:** I guess to couple. I cannot run Uzi jobs outside of Hadoop.
**Eldad:** Fair enough.
**Eldad:** To DBT or not to DBT?
**Boaz:** Do you guys use DBT?
**Shaun:** DBT.
**Boaz:** When did you guys start using DBT?
**Shaun:** We only started exploring it. We have not done it yet, but there is a lot of concept and approaches, and ideas that we like a lot, and then we are going to use it for the development of our internal tools to match, you know, what DBT can do.
**Boaz:** What is it you to use for ETL, is it mostly Spark or other stuff too?
**Shaun:** It is driven by Spark, but, it's mainly in-house right? Like we built an ETL tool that is based around spark, but it is extremely easy to do as a whole UI and so on and so forth. It is like anyone in the company just goes in, writes a few SQL, and you are done.
**Boaz:** And no commercials, sort of traditional Informatica style, kind of more these kinds of On-Prem, ETL integration tools.
**Shaun:** No.
**Amir:** Not needed.
**Boaz:** Yeah. Nice. Okay! So now after you guys, trying to be so cool for listeners, it is time to get real and tell us about one project that was horrible for you guys. That didn't go well at all. So we can all learn from your mistakes. Who goes first?
**Amir:** Yeah, that is a horrible one.
**Boaz:** What mistake are you not going to repeat again? Tell us about a project like that.
**Shaun:** I think for me is, I think the failure is that we did not get into decouple storage earlier. Like a lot of our systems are still coupled together and that has hurt us. Because obviously with decouple storage, you could always scale, compute, and storage independently whereas like, in the old ways everything has to be uniform. So, being slower on the bandwagon definitely has some impact on us.
**Boaz:** Yeah. Amir, any glorious failures on your end.
**Amir:** It's too many, every day there is a few but again, not about my personal life, right? It is just always feeling that you are moving too slow. I do not think anything specific, but yeah, getting, for example, by Spark, I think, we should have done it a lot sooner, and in terms of GPU, as we invested a lot in building an amazing project, around Uzi that will allow you to try to kind of do use Uzi to get outside of Hadoop and run stuff on some Kubernetes cluster with SSH commands and stuff. We tried to break this coupling that Shaun mentioned and we failed miserably like it ended up being unusable and we threw it away and our way of trying to rebuilt a new solution but we tried to make Uzi do stuff but he didn't like to do so if you pushed back on.
**Boaz:** But do you still, said, voted for Uzi before.
**Amir:** Yeah. But this is why we insisted, right? Because Uzi is a great tool, even though it's XML based and I think it built in the eighties and the UI is as old as like very old.
**Boaz:** As old as the eighties.
**Amir:** Yeah, something like that, but it is very robust and battle-tested, and it allows you for a lot of features and there is a lot of cool stuff you can do with it. So I wish there was an Uzi replica that is not coupled with Hadoop, but we couldn't find one so far.
**Boaz:** Yeah. We seem to have another issue with Shaun, so maybe you're going to have to answer all questions on your own until we get it back.
**Amir:** I will take it from here. Okay!
**Boaz:** Now on a better note, tell us about the project that did go extremely well or something you are proud of, that is exciting that you want to share.
**Amir:** Yeah, recently we built a very cool too for a model monitoring. Usually, I just talk about machine learning, but this one is also very close to data engineering, I guess, hardcore data engineering because what we saw is that, okay, people build a model, they send it to production. They do maybe an AB test to see that there is business value and that it is actually better than the previous approach. But then after the AB test is over, like nobody's watching over it. Okay. So people see that top-line numbers. People see that traffic is flowing. Bookings are made. People are happy with their product in general, but you don't know what is going on within your model. How good are the predictions that it is doing? What we built based on the stuff that Shaun along the great Kafka pipelines that we have with Shaun, we are sending the data from the model monitoring all the way back to Hadoop. So from the model, the inputs and outputs are sent back and we have a spun of the crunches the statistics of all this data. So kind of trying to find a needle in a haystack, we go column by column, calculate all kinds of stuff, similar to what DQ is doing maybe if you know, TensorFlow data validation, calculating the shape of the data that is flowing and trying to come to insights about how we changed compared to a week ago, compared to what we thought we had when we train the model and these turned out to be a very cool tool. So everything is now connected to Grafana with automatic alerts, and we try to find these anomalies and get them back, getting Shaun back, and, so that was a very cool win and the cool part was the human part. The model owners, they just click a few buttons, register here and there and that's it, their model is monitored. Everything goes automatically for them, and they get these amazing alerts for free and that was very cool.
**Boaz:** How did you build a justification tool to go after a project like that? And how many people were involved?
**Amir:** Yeah, that's cool. So we work closely with the data science department and we kind of search with them, what are the pain points in the beginning? For them, none of them actually complained about that because it was kind of falling between the chairs, between the ML engineers and the data scientist, so data scientists care about, Hey, I want tools to work fast. I want to be able to have a lot of resources to crunch a lot of data. And usually, once the model is in production, they cared a little less about that. Well, the ML engineer kind of felt that, okay, this is a machine learning model, like the data science responsibility. So we kind of know that nobody kind of owns that area and maybe it should be owned by the platform and that was a reason for us to go in and do that. And also there was a streak of failures and actual incidents that happened that caused, like a platform degradation. And we said that, okay, it justified to build a tool for that and a platform from that.
**Boaz:** Awesome. Shaun, what you missed is the question because we lost you for a minute.
**Shaun:** Yeah! Sorry, my place got a blackout like it happens in Bangkok.
**Eldad:** It happens.
**Shaun:** Yeah.
**Boaz:** So now it's your turn to share a project that you were proud of or a great win that you're happy with.
**Shaun:** I think that the Data Market, like the stuff that I mentioned previously, really has beyond that discoverability that it gives to everyone in Agoda. It also serves as a central place to get the information about a data piece, so you can get data quality information from there as well. You get to figure out who is sending this, from where, all that kind of cool information, all in one single piece. And that really has been instrumental.
**Boaz:** Awesome! Okay guys, I think, we are reaching the end of the show. You've been great. Absolutely, exciting to see what is happening behind the scenes with data at Agoda. I'm going to think of you guys next time I book a trip or something.
# How do Canva's engineers and analysts scale data platforms to keep up with growth? — with Krishna Naidu (/blog/how-do-canvas-analysts-and-engineers-scale-data-platforms-to-keep-up-with-growth-with-krishna-naidu)
Canva is one of the hottest, if not the hottest, graphic design platforms out there. With 55 million active users and around 500 million dollars in annual revenue, Canva is an unstoppable powerhouse. They were recently valued at 16 billion dollars!
So how do Canva analysts and engineers scale their data platforms to meet the company's insane growth?
To help us find out, we recently invited Krishna Naidu—Data Engineer at Canva and expert in building large data platforms—to speak with us on The Data Engineering Show.
Listen on [Spotify](https://open.spotify.com/episode/1zMGZ39RmRwMel41EyFPNA), [Apple Podcasts](https://podcasts.apple.com/il/podcast/how-canva-fosters-collaboration-between-analysts-engineers/id1561927688?i=1000521449359) or watch on [Youtube](https://www.youtube.com/watch?v=btF3Twk2_q4).
## How much data does Canva deal with? [#how-much-data-does-canva-deal-with]
Canva processes a huge volume of data. At the time of speaking with Krishna, the Canva data warehouse consists of about 400 TB of data, seeing daily volumes of about 2 TB of raw data. Their biggest data set is their event tracking table, which tracks all of their analytics events.
## How many people work on the data team? [#how-many-people-work-on-the-data-team]
At Canva, the software is always changing so the company requires a large team to keep up. There are currently about 20 data engineers, 20 data scientists, and 40 data analysts—and the team is still hiring globally.
## What does the data stack look like at Canva? [#what-does-the-data-stack-look-like-at-canva]
Canva has historically used both a data lake and a data warehouse, and will likely continue to use both in some form.
Streaming is the main source of incoming data. Incoming data goes to the data lake and is stored nicely in Delta format. Delta format is the foundation of the data.
From there, the data is loaded into the data warehouse, which uses Snowflake.
They also keep a raw data lake. Currently, both the warehouse and Delta Lake consume from the raw lake but Krishna and the team are working on transitioning with Snowflake to consume more from the Delta Lake to cut down on repetition.
Krishna was hired at Canva to revamp their data warehouse. As Canva began its explosive growth, they struggled to scale the existing data warehouse. They needed a better way to gain the performance and storage that they needed.
## What is the data team at Canva focusing on now? [#what-is-the-data-team-at-canva-focusing-on-now]
Now that the warehouse has been revamped and contributions are up, Krishna and his team are working on enhancing analyst and engineering productivity.
With a team of 40 people who might be working at any one time in the development environment, things can get tricky. Using Snowflake, Krishna's team is working on a new workflow that will allow for better testing and rebuilding.
They are also prioritizing giving more control and ownership of the massive raw datasets to backend and frontend engineers, to enable contributions from the broader organization.
[Listen to the full episode](https://www.youtube.com/watch?v=btF3Twk2_q4\&t=1s) for more insights from Krishna and subscribe to our YouTube channel to never miss a podcast episode.
# How Eventbrite is Modernizing its Data Stack (/blog/how-eventbrite-is-modernizing-its-data-stack)
Archana Ganapathi, Head of Data & Analytics Engineering at Eventbrite, shares Eventbrite's data stack modernization process, and how you get engineers to adopt new technologies like dbt which may be outside their comfort zone.
Listen on [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-eventbrite-is-modernizing-its-data-stack/id1561927688?i=1000563268516) or [Spotify](https://open.spotify.com/episode/7sVUVmqQcA28H7kjsyomZf?si=V2johcevQ3mLdLfWtUrSEw)
Guest: Archana Ganapathi, Head of Data & Analytics Engineering at Eventbrite
Boaz: A little bit about Archana. Archana is Head of Data Analytics and Data Engineering, Data Science at Eventbrite. She has been there for almost a year, spent a lot of time at Splunk before, knows data in and out, and she is a little bit upset with us for not using Eventbrite for this event. We cannot please everyone.
Thanks for keeping your smile.
Archana: Will you promise me that next time around, you are going to use Eventbrite.
Boaz: I do not know, talk to marketing about that.
Eldad: Disqualified us for being a startup.
Boaz: Archana, tell us about what you came to do at the Eventbrite?
Archana: I joined the Briteland as we call it about 10 months ago to lead all things, data, essentially what this means is traditionally we have data engineering, the data science sitting somewhere else, and analyst sprinkled all across the board in functional business units. Eventbrite came to the realization, that we need to bring all of these different roles and functions under one big umbrella. So, we can think through this end-to-end and share context as much as possible and solve data problems and look for data opportunities more holistically. We can leverage the scale of all the rich treasure troves that we are sitting on.
Boaz: This is why Archana has one of the best of the things, titles, which is Head of Data Analytics/Engineering/Science at Eventbrite. We need to come up with a new name that can umbrella.
Archana: It is all data.
Boaz: It is all data. So, maybe before we go back to Eventbrite, you spend part of the time at Splunk. What did you learn in that career at Splunk, doing a variety of data positions, that how you feel, lead you to this phase of being ready to take on this new challenge.
Archana: That is a very good question. I think if I reflect back, the journey started much before I even joined Splunk. During undergrad and grad school at Berkeley, basically, a lot of the research I was engaged in was constantly centered around data. This was before big data was a thing or data science was even kind of a formally recognized profession, so to speak. Really, it came down to how we take advantage of traditional computer science systems, research, and technology out there to allow us to scale up insights and get people value from everything that we are accumulating, all shapes and sizes and forms of data. And then, scaleup, compute scale-up storage and make it just easier to drive consumption. That was kind of the backstory, straight out of grad school, I thought if I do not join a startup now, I will probably never go back to it. So, I joined this late-stage start that was then around 200 people at Splunk. I was part of the core engineering team, building some of the platform capabilities to leverage the data that was being ingested, stored, and queried in Splunk. A few years then, I realized, that maybe we are not solving the biggest pain points for our customers. So, I moved to the field to understand where the true pain points are for customers that were trying to use Splunk at that time. And a lot of it came down to people and process gaps, but also, some nice to have again, to drive self-service insights from the platform. Long story, short, soon after we realized we should be drinking our own champagne, just like Slack uses Slack a lot, back in the day. Why should not we use Splunk for internal insights, and I built up the data and insights team, advocated for heavy investments to instrument our product, collect even more data, and then triangulate it with instrumenting our processes to enrich the context there to drive business value. After a good chunk of time, 11 years later, it was time for a change. Then pandemic was a forcing function to want and create that next chapter.
Eldad: Oh! it was a great excuse.
Archana: Yes, it was a great excuse. We are also just partly reflecting on where are the gaps in my own journey here? I thought really thinking about the true scale of impact from data, I need exposure to something that is a bit more consumer rather than enterprisey, and hence, Eventbrite and the mission here is to connect the world together through live experiences, and what better north star for really leaning in on data to bring that to reality.
Boaz: I know even at Eventbrite there is a modernization process now having. Tell us a little about what is in place, what you got in, and what is data stack or environments look like.
Archana: Yeah, absolutely. A bit of history here. When Eventbrite started primarily thinking about ticketing. The data stack as well was designed primarily for ticketing and transaction management and reconciling our books and, those kinds of use cases. Over time, as the demand for insights and analytics grew, there was a lot of duct tape that was fit on top. Early stages, everything threw it all into my SQL database. And then slowly you see, okay, maybe we need some nicer pipelines here and there. So, Spark on EMR was primarily used for compute and the query engine is Presto, storage all our data, moving around S3, HDFS, a lot of the metadata over time is sat on Hive. Then, Tableau was used for dashboarding, and Luigi for orchestration. A lot of technology was right at the time, decisions that were right at the time but if we really think about future-proofing, the infrastructure, this is not the stack that will get us to the future state. So, that is where we are right now, thinking about modernizing.
Boaz: What do you think was your tipping point on that moment of realization, Hey, it is time to reconsider the change.
Eldad: Join. That is as simple as that.
Archana: That is a part of it, but frankly, I think a lot of that, demand also came from bringing data science and analytics under the same umbrella as data engineering, and saying, "Hey, here is what we are trying to do. Here is where the current infrastructure is not meeting our needs." If you think about it, there is reporting, there is ad hoc analytics and then, there is really driving some of these insights from the data back into the product, data power product functionality, whether it is even simple heuristics or fancy machine learning models, requirements change for what needs to happen end-to-end to make that a reality and make it a good experience for our customers as well. I think that was the forcing function and as part of that, we just took a clean slate to say, okay, if we were to design this from the ground up, what do we need to solve for it? Some of the pain points are really like stuck in this chicken and egg loop, our platform is kind of killing over so we cannot support some of these use cases and scenarios, so hold off. But then, in the process of ripping off some duct tape and adding more stuff, you are still stuck in legacy and more things that, snowball and cascade, and any incidents or bug on the product side that impacts data quality, for instance, now becomes the data team's problem, also to stop, undo, redo replay, and it is days of cycle time to get back to clean slate state. I think all of these... it is a domino effect that happened together. And to your point, me joining was a good checkpoint, to say, okay, you know what, let us figure out where we need to go.
Boaz: How do you go about that? The organization needs a big stack of support and then there is a huge journey ahead of us. How do you manage that? How many people are assigned to the new project or are the same people working on the data stack, is there migration plans? Is there a deprecation of all things plans?
Archana: Yeah, that is a very good question. I am under no false misconception that I have all the answers, but I can share what I have done. I guess. The first step, really when I joined was to listen. I get a better understanding of what people want to do, where they are getting stuck or where the challenges are, where their blockers are today, and then figure out how much of this is a fundamental infrastructure technology problem. How much of it is a process issue? How much of it is just like knowledge and awareness? Do people even know this is the place to start? And there is a rich set of data and dashboards that they can leverage already and is there an access problem? Really just going through and figuring out where all the current challenges are and then, that just turns into the requirements for what we need to build. Based on that, the first thing we realized was now we need to modernize our data warehouse. That is one decision we need to make. We also need to make it much easier to instrument and collect richer data at the right granularity and this is upstream of data, back to the dev teams to say, "Hey, this is the information that is missing that people need for what they are trying to do with the data." So, almost kind of teasing apart, data producers' requirements and constraints and data consumer's requirements and constraints, and then on my team's plate is how do we build out a platform that simultaneously solves both?
Eldad: Switching from XML to Jason.
Archana: No comments!
Eldad: Long term, long term that is long term.
Boaz: You managed the part of your origin of stories, putting this under one's roof, the various teams. Tell us a bit about that. What happened? I mean, is there still science to data engineering, so how was the change?
Archana: I think the key is really more communication, more shared context, and shared aligned goals, that matters a lot. If you think about it, if incentives do not align, there is really no benefit to solving for something truly end to end. That is the first step, to say, our biggest goal for North Star here internally is to enable all bright links. So everyone internal to Eventbrite, to have access to the data they need to do their day jobs, arm them with the insights they need to run their own part of the business and make sure we are that bridge between the data producers and data consumers. So, getting folks to talk to each other, put themselves in each other's shoes, and empathize with the challenges or kind of what are each function trying to optimize for and building that awareness went a long way because there was a lot of realization that was not happening around. Here is where I am getting stuck. I didn't realize that it is not a platform limitation. I just did not know that this was the best way to do it, encouraging that dialogue.
Boaz: Now that there is more progress that has been made from the time did you start with your stack, what do you know, for sure, once in a while, how are you getting...?
Archana: I think the top of my list is probably Luigi.
Boaz: Does everybody like the name?
Archana: It is a cool name. I must say that it is associated with the Mario brothers and pizza. Who doesn't like pizza?
Boaz: So Luigi is gone, what else?
Archana: I think next would really just be figuring out how we move folks from having to build pipelines in Spark to just up-leveling that. So dbt as much as possible. Some of the bad behaviors, were kind of, workarounds to the old stack where folks are leveraging Presto as a way to kind of shortcut into the pipeline building which is not the right thing to do in the longer term. That is another thing we will be doubling down on dbt is kind of on the radar.
Eldad: How do you do that? How do you like it because you have a team who writes Scala and build some engineers stuff, and then you counter to switch to dbt and it is such a different universe, different skill set, different people? How do you manage that transition between the team so that nobody gets freaked out too early?
Archana: I think I would almost turn that around to say like, I think historically we were trying to force people that were not comfortable doing some of those things, to use those tools that were outside their comfort zone. Now we are just saying we have the foundations to make it easier for you to do the things that you have to do. Just double-clicking on that a bit. We have a data platform, an infrastructure engineering team, and an analytics engineering team. Analytics engineering should just double down, focusing on driving consumption and building those gold layer data large stack, making everything easier downstream. But that was not necessarily where all their time was spent historically just by nature of the tool stack and toolset that we were using. So, I think it is a big welcome to the change and they are eager to modernize and learn the things that will make their lives easier.
Boaz: Also, Archana is under the BI side. The reports and the dashboards will often have to be remade, are people worried about that?
Archana: That is a good question. Well, it is okay for right now, but I know that there are better tools out there that we need to lean into and I know, Apun mentioned, that Looker was one that they use. There are others as well, that are much easier. I almost think that there is not going to be one size fits all. Also, maybe I need to do the same exercise around, like, what is the preferred interface that each of our data consumers has? In some cases they want dashboards, sometimes they just want the handholding, like, this is the thing you should focus on. I will give you that full service, analytics experience to handhold you through insights from the data and how that should drive your decision? The third flavor, I would say, there are tools that people are already using, and they just want data fed into the tools that they are comfortable with. For instance, sales prefer CRM tools that they are already using. So, the more we can push things automatically into the interfaces they are comfortable with, that is going to matter.
Boaz: That is going to be painful.
Archana: Well everything is painful, but everything is also kind of like if we take it apart into a piecemeal thing that we are solving for each pain point, then we will see some things that we can generalize into patterns and solve for and maybe that is bringing in a third-party tool. In other cases, they already have what they need in terms of the interface. Now, it is just how do you route things to the right people in the right way?
Boaz: What is one thing that sort of so far you are super happy with the specific pain that is already solved for, that you are just happy about?
Archana: I think they just love the partnership with the data team and really being able to lean in on, like, tell me more proactively how I should be thinking more data and really working through with end-to-end from. What should I collect, if these are the metrics that I am solving for and optimizing for and bringing in some of that end-to-end thinking? It is almost like being a consultant in many ways. And kind of having more of the analysts, like, channel the inner consultant in the way they approach solving for the internal stakeholders.
Eldad: I miss the days when you just guesstimate the stuff and use intuition, and gut feeling, and now you were freaked out about following through with the wrong KPI. Makes sense. Life is getting more complex when data is...
Boaz: Migration stories, everybody, some people probably think, ah, you know, I would not bother data stack, those shouldn't care about. One thing is for sure if you are on a data stack from the bottom for a while. So migration stores all is more important than we think
Questions from the crowd.
Audience Speaker 1: Consuming your data at Eventbrite for a long and actually also using it but that is pretty interesting. You have descriptive tags. Do you use a tool like Yokozuna and pipeline into their warehouse? And because you can let people rank the term that maps back to the event in lights more or you let people say I follow this and they get notifications. So you have a feedback loop, not just for your internal consumers, but for your external users and event holders where you can weigh what terms really map to the events people want to sign it for, register for, pay for, and to convert him later. That is the area where you just literally let them follow the tags better. They will forget about the really good tool like this, it is super fast, textbase, and then out of that, you can let get more discovery of earlier events. You can have a higher conversion. They will probably pay for the marketing part.
Boaz: Hire him!!
Archana: Yeah, I am eager to hear more, but I also want to plan, I don't know if you've looked into our new marketing tools suite, but highly encouraged that.
Audience Speaker: I have downloaded a new Eventbrite app, but I was like this Oh it was painful enough to use the old app. So, I didn't look at it.
Archana: Okay, we will chat offline, but certainly something that we are solving and moving on.
Audience Speaker: Actually more useful for me. We will definitely talk about it now.
Boaz: Anybody else?
Audience Speaker 2: Okay. What was the final nail in the coffin for Luigi and why do you think that there is any escape?
Archana: I would not say it is the final nail in the coffin. You still have not done that migration successfully.
Audience Speaker 2: Okay.
Eldad: There are multiple coffins here.
Archana: Yes, multiple coffins. Yes, multiple opportunities to be born again. How about that? It really comes down to just the Daisy chaining of pipelines and the capabilities around automating that better and that is where I think airflow frankly has done leaps and bounds better in that user interface.
Boaz: I have another question based on that, how much do you feel, data engineers in general, prefer working on the data, most interesting. Do you feel the teams that they are excited to do that or it is modernized, or do I know what I am more concerned about that keep doing as I do?
Archana: That's a very good question. I have been doing a lot of thinking about that. I would say it is a 50-50 split. There are folks that are just in their comfort zone with the tools they have been using and I know folks who have gone from one role to another, who just want to follow the stack that they feel happiest about or feel like they are most knowledgeable about. Then, there are other folks who at least right now, the data engineering team at Eventbrite are super pumped to just rip off all things, legacy, and just modernize. It actually makes their lives easier day today. And some of the more modern capabilities around like scaling up and scaling down to zero and the cost benefits of that, the performance benefits of that, and checkpoint rollback recovery, those kinds of capabilities that people have lost sleep over not having. So, that is where a lot of that excitement comes from and push for, yeah, let us do this as quickly as possible.
Boaz: Awesome.
# How Klarna Designed a New Data Platform in the Cloud (/blog/how-klarna-designed-a-new-data-platform-in-the-cloud)
Klarna is one of the leading fintech companies in the world, valued at $45B. While many corporations are "stuck" on-prem, Klarna made the move and today is a cloud-only company. Gunnar Tangring, Klarna's Lead Data Engineer tells Boaz what this new modernized stack looks like.
Listen on [Spotify](https://open.spotify.com/episode/1xZvW4lvciSW8qv7pmGLLj) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-klarna-designed-a-new-data-platform-in-the-cloud/id1561927688?i=1000565772037)
**Boaz:** Welcome everybody to another episode of the Data Engineering Show. Welcome, Welcome! With me today is Gunnar Tangring from Klarna in Sweden. Hi, Gunnar! How are you?
**Gunnar:** Hi, I am good. How are you?
**Boaz:** I am very good. Thanks for joining us. For all you out there who do not know Klarna, Klarna is a super, super interesting company. Klarna is in FinTech from Sweden. It is one of the most valuable privately held FinTech companies in the world. Last year, the last round of investment came from SoftBank, which valued the company at $45 billion, unbelievable. Klarna has been run around since 2005 and has shown an amazing journey and started around just making payments online smoother and a team that kept evolving and pushing throughout the years. They are rather famous in recent years for this kind of, I do not know if you have seen this, but shop-now-pay-later kind of experience. Throughout the years, Klarna has also evolved into being licensed as an actual bank, having its own credit card, which is relatively recent with already more than a million consumers using it, I think. So, a really exciting and interesting FinTech company. Gunnar, is there anything I missed? You have been there for six years.
**Gunnar:** Yes.
**Boaz:** Gunnar is a lead data engineer, architect and more. Tell us a little bit about somebody who has been there in the last six years, and then, we will dive into what you do there and beyond.
**Gunnar:** It is an incredibly good description. So thank you for that. But just my personal perspective more from the data angle, is actually that I remember that I walked into a building and I thought it was kind of a growing startup and it turned out to be a slightly larger company than I thought. But today it is something very different that is incredibly noticeable on the data side. I remember when I joined, there was this BI team, which consisted of a handful of people who handled all the different data requests across the company. The scale we are at now means that we have an entire, what we call the domain of around 60 people doing similar tasks where we have been just five people doing. So it is a completely different ball game, of course, with the type of volumes and there has been a lot of growth on the pivoting into new areas all the time. So, it was always good fun.
**Boaz:** Amazing! Tell us about your roles throughout the years?
**Gunnar:** I started with reporting BI work, traditional working in Cognos at the time and the majority was merchant reporting, but anything people would want to have covered like I want to know this, I want to know that, and that was not exactly what I wanted to do. So, I was working a couple of years on our Hadoop infrastructure, basically building fairly large scale data flows at the time in various forms, a lot of Hive SQL and building frameworks around it. We built a tool called the HiveRunner, as an example, to do unit testing of the data, which is kind of cool. And then, we moved to the cloud. Then you know what you are doing for a couple of days.
**Boaz:** Cognus is not around anymore.
**Gunnar:** No, it is not. But the journey of "Hey, let us move everything to the cloud," I think is when I understood what Klarna was all about. Because it was not some kind of let us do this. It will take time. It is more like, "Hey, let us do this now, fast and get it all done." We did that. We pulled it off within reasonable timelines. So, that was nice. But we shifted our entire infrastructure to a cloud-based only, which obviously data is one of the parts where that becomes tricky because you have to migrate the data itself, but you also need to migrate your tooling. We did not go for a Hadoop on cloud type of setup, so we shifted quite a lot.
**Boaz:** Today, Klarna is cloud only when it comes to the data stack.
**Gunnar:** Yes.
**Boaz:** When was the final switch off for an on-prem? We know we have got companies throughout the years that the length of time it takes to complete is between years to indefinitely. So, just reaching the state where it is completely off is an achievement on its own sometimes. So, how long ago was that?
**Gunnar:** I do not remember 100%, but I think it was probably like three years ago or something like that. But the nice thing was that we were quite happy with our work because we were not the last ones out. So, that meant that someone else was running something on some service somewhere. But when we managed to pull the plug there was obviously an announcement video where someone was going to the computer all that, like unplugging the final computer and then, we ran into a new era.
**Boaz:** That is fine. So many people would not be able to ever experience that. People nowadays are born in the cloud. So, you deserve a medal for being there when on-prem was unplugged for Klarna.
**Gunnar:** We are a fairly modern company, but I still think kids these days do not understand why you have things under your desk, "oh, it is to keep your computer." The computer was the server, which was giving people what they needed. It is a different area, but it has been a kind of a fast transition in the industry. The debate of cloud or not, is fairly done to me.
**Boaz:** This debate is definitely over, I think, even though there are a lot of on-prem activities still around. More than we typically think? Every time there is a survey comes out and it turns out there is so much workload still happening on-prem. Some portions of the industry are much slower to move than we, sort of, on the more advanced side of data and tech tend to realize. Because we live in this modern data stack world and sometimes we forget that we were just a segment and that there is so much workload out there on-prem, but yeah, it is all coming to the cloud for sure. Klarna is over 6,000 people worldwide according to what I see on LinkedIn, more or less, but how many people are spread across the different data teams? How many people deal with data and how is the variety of teams structured at Klarna, if you could walk us through that?
**Gunnar:** I think there are two different stories to tell. One would be that we worked in the data platform in Klarna's domain where I work as a domain architect, but our structure internally is basically that roughly half of our domain of around 60 people is infrastructure, people building the platforms. So, providing tooling for other people to do data processing, in a sense building frameworks and making sure the databases are running and making sure that you have the correct setups of the access to data on block storage, and building the data catalog as well and the other half is doing what we call, core models, basically being end-to-end BI teams. So, they work with more central data that would be not possible to own within a specific domain of the company, but that is the next step. Like a lot of our data teams are either something that is working like in the finance department with data. And, be like a team that is explicitly working on data. And, we have teams doing big data processing in various phases, everything from risk positioning to defraud detection to anomaly detection with problems that can occur with our merchants. So, I see more and more of this, like, "Hey, we need to spin up a data team for it," for something somewhere in the organization, less and less, "Hey, can you please do this centrally at the company because we need this to be done from you" and that is very much in line with how we are trying to build the company, that we want to have. We want to have the domain knowledge close to what is going on and, I think for data that is in particular extremely important because I see a lot of cases where someone is being asked to do something but I do not really know why and they are deemed to be the data people who know exactly how to handle the data. But how to handle data will be very much a product of what the data, like, why is this data generated this way? What does this field mean? All of these things are impossible to know if you are working in some kind of decoupled function, central at the company, you might be extremely good at sparse indexing or sort of really good for performance tuning, but it does not really help if you do not know what you are doing.
**Boaz:** You guys are looking into sort of a data mesh implementation?
**Gunnar:** I would say so. We have looked a bit at data mesh but the thing for us with data mesh is actually, it was not something that came from the sky, like, "Hey, let us do data mesh." We looked at what we were doing. And, then we looked at the data mesh and we realized that this is very close to what we are doing and we can learn some things. We can try some things into terminology. But for me, the key takeaway with data mesh is the ownership aspect. I want to have strong ownership.
**Boaz:** You are saying, it is not about us picking up a data mesh guide and implementing it by the book if the book even exists. As you know, the data mesh in itself at the end of the day talks to something that the industry has always gone back and forth. So centralization versus decentralization in essence, and you are saying you are leaning now towards decentralization and domain ownership much more at this stage where Klarna is, makes more sense for you guys?
**Gunnar:** Yes, but I think we have always been a bit towards that angle, but I do not think we have had the vocabulary for it. But, what you are saying, there is a data mesh book now and I have not read all of it but I will read it at some point and I think it is good inspiration. But I also think it is kind of unclear exactly how you would implement it. Then I see a lot of flame wars of, is this actual data mesh. And, then I just realized that that is not what I want to discuss. I want to discuss, what is the thing that drives our company forward in a good way. I think ownership and how to set the boundaries, those types of discussions are needed. But, I think the data mesh in itself does not really give the answers. It just gives you a framework of how to talk about it in a sense.
**Boaz:** Yes, I completely agree. It is a framework, and it looks different in every company because it depends on people as well as practices and the details of the organization. It is like to some extent they do not feel like being agile. How how do you become agile? There was this decade where everybody was talking about becoming agile in software. There is no one way of becoming agile and there are many ways to implement changes from company to company. But that absolutely makes sense. But, what about data engineering though? There is BI, there is data engineering, how are they split? And is not data engineering more centralized compared to the BI teams or is that also spread across the different domains?
**Gunnar:** I think there are two answers. One thing is the terminology matters. I think we have been discussing if we should release the title, analytics engineer is similar. Because the people we have and who are working with, what a lot of people would call BI would potentially be labeled as data engineers. And that could be the wide-scale, you have everything from someone working with building automation of pipelines and someone implementing the use cases. So, I think there is a space there were titles could make a difference. But, I also think the type of data warehousing work or whatever you would label it, that we are doing might also be different in a sense that you fast fall into the scale of things. So, you do not really have maybe a writing SQL all day, but you are still like, if you do not know how to perform this on SQL then you will have problems like this. I think it is a title question in a sense, but I agree the things we are doing that I would label as data engineering and not analytical engineering are more centralized. If I take the clear example of building and implementing and adopting frameworks for making sure that we can build analytical pipelines, that is something that we considered to be something we should centralize and offer as a standardized component. If you talk about data mesh again, I think this is building the sidecars, but you need to run your analytics products in a company. For us, it makes sense to centralize. It makes sense to standardized because you get so many things out of the box and the way we work is if you want to build your own thing, go ahead; but if you want to integrate with 10 different available tools, if you want to be compliant, then it is probably a whole lot easier for you using the framework that you already have, but it is a give and takes, like if you have a super-specific use case, then you can always build, what you need for that basically?
**Boaz:** Saying though, everything above data pipelines to some extent you prefer an analytics engineer, maybe somebody that can do full stack data essentially, a mix of data engineering and analytics. And, I think in general, that is a trend we see in the markets. The term analytics engineering is half picked up, but people are starting to like it, but actually, I would like to call it full stack data developer or something like that. That is a kind of combination of meshing BI and data generated together because you know, you cannot do much today if you are not able to go further down the stack, roll up your sleeves and do some kind of data engineering to some extent. And, I think more and more data professionals are finding that out and that mixture of bridging the tools is picking up. Interesting to hear that you are going about that at the Klarna. But do you guys already use the title analytics engineer or not yet?
**Gunnar:** No, we are discussing it. But I mean to come with them, like the full stack, I think of it very much just like data being a field, it is not unique to data, but I think in general, I think it is more and more important nowadays to be a bit T-shaped than actually having one technology or something that you are good at and then having something else to combine it with, and that might be that you have, maybe a great building analytics pipelines and you are good at domain knowledge of finance or something or it might be that you have like DevOps capabilities to help with other things and I agree that a wider stack is a strength that you need to showcase. I think it is incredibly hard for everyone to have enough to fill up the entire scale. So, from that perspective, I am sometimes thinking, "Hey, it does not matter what type you have, because I will still have to probe and understand what you are doing." I think that is particularly, I realized as being an architect for a while, because when you talk to other architects, it is just such a scale of everything from, I do not know any details to, I know all the details, to hydro houses, like literally of course, but I have that confusion with some neighbors that they thought I was growing abscess, but I just asked, do you have a database? So, then maybe I can help. But, I do think that what are my core things? What would be my two or three things to pick up on, some kind of strengths? That is how I view profiles in general and I think I am more inclined to want to hire someone who is really good at two different things, as opposed to someone who was more like a Jack of all trades type of person. That is where I might post full stack, but I do hear what you are saying.
**Boaz:** Got it. Walk us through the current data stack. Now everything is in the cloud, how does the data stack at Klarna look like?
**Gunnar:** We have what we call a data lakehouse and we thought we branded the term and the print hats and things, but then it turned out that someone else used the same term before us, but we were happily unaware and just thought we had invented something new, but in essence, we are running on AWS analytic stack. So, we were quite heavy use of.
**Boaz:** By the way, you guys should make some noise about it. Go online and shout from the roofs. Hey, we coined the term data lakehouse.
**Gunnar:** I have told Databricks but then I Googled it and I actually found some reference where I think they were using the term before us, but...
**Boaz:** Never mind the facts. Let us rewrite history.
**Gunnar:** But I mean, we did demand, like, in that sense, we did coin the term, but it is not something super complicated in the sense, but if you think of it from more of an API perspective, it is a platform where you can publish data and as a producer, then, you can consume it as a consumer and you do not really have to necessarily worry about the exact location of the data or exactly how things are working within this box. But it is a combination of basically S3 and Redshift and EMR clusters on the spark jump running and we managed that centrally with the configuration possibilities for the users and iterating on it every day, of course.
**Boaz:** In your transition to the cloud, Redshift was selected as sort of the enterprise data warehouse, right?
**Gunnar:** Exactly. And that was like if you looked at where we were coming from, we were coming from a fairly interesting scenario where we had a mix of PostgreSQL and Hadoop and the actual data warehouse was implemented in Microsoft SQL. And, when I say PostgreSQL, that was not just one machine somewhere, it was not like 10 machines having the same structure either, it was just a wild mix of a lot of different databases all over the place. So, we went for Redshift as being, like the main data warehouse engine with capabilities of also processing less refined data. So, that is where the Lakehouse term really refers to being able to access through one logical environment where you do not have to go to a different place because you need data from a different domain. You can go to the lakehouse. You have the lake and the house, both order and disorder in the same place.
**Boaz:** There is also, like, Finna used around it.
**Gunnar:** We are not using Finna heavily. It is one of the things we are looking at how to leverage more potentially, but it is expanding in usage, I would say. But, we used it extensively, when implementing it; it was like the go-to tool because it is extremely powerful to use the data for interactive results on fairly large data volume and I was so surprised when using the Finna, but I would expect things to be slower than they were being used to, but a lot of the things were quite snappy. But, then when you are running into some limitations, of course, some tools. It is kind of built to have the pet tool and the Hive is very much the opposite where you are focused on resilience and jobs just running until they finish, the Finna is more or just did not work, I am not going to tell you, but you do not get a response. So, it is a different experience, but it worked well, like when we were implementing the phase and needed to quickly look at the data and draw some conclusions. It has just been very helpful. Also very good for doing sanity checks, if you have a data gap for whatever reason, you can refer to it easier to just stop that.
**Boaz:** What else is in using the stack sort of Redshift?
**Gunnar:** We use Airflow for orchestration, and we have built our framework surrounding it. We have a team that is focusing on building frameworks. So, instead of exposing Airflow to end-users, we are leveraging our own CI/CD set-up for it. So, you are kind of forced to come into our setup where we guarantee that you have virtual control for transformations and you have some support outside of what you can get from Airflow and we do not get very skilled people, messing up too much because you have to go through our Jenkins to get things running basically. One other thing that I would mention but I did not is, our data catalog as well, which is the thing we built recently, based on a data hub from LinkedIn. So, this is one of the things we realized was a big gap when we rolled out their first citation, but we did not have a good way to just give the user a way to discover all the data in a sensible manner. So, we decided to adopt a solution that would be flexible enough for our needs. So, we rely on being able to ingest metadata from other services by pushing the data to the data hub. So, that was a conscious decision to go for a flexible open-source project that seems to have some traction.
**Boaz:** Got it. What data volumes are you guys dealing with?
**Gunnar:** Well! It is petabytes at least. So, the timing varies depending on the domain. Our growth of data is quite big and a lot of the things that have gone live later. When you talk about how much data we have and how much it is growing, we do have hockey stick curves for a lot of it and that is a challenge, of course, but we always have to. My experience has been about when you go live with something new, you tend to be at the state where you are producing a bit more data than you need because you are going for a fairly naive setup. So ironically, even though the volumes are growing, you tend to be able to manage it more over time and it becomes more predictable because you can determine if this is useful or not. But the overall influx, I do not have the number.
**Boaz:** How much does all the data end up in Redshift? All of it or do you do it year to year?
**Gunnar:** No. Not all of it. We obviously want to keep it sensible. But this is always a friction point because it is tricky to work with a situation where you would need to send people to different places, depending on what data we need. I would say we have more than we would wish for.
**Boaz:** What do you guys do for an ETL or ELT batch processing and Spark or other things?
**Gunnar:** We run Spark and some Hive as well, but we run a combination of Glue and EMR. So, that gives us the main use case for using best Glue is really that we get access to a run time of serverless Spark and it is quite convenient because that means we have an API that is well known. If we would want to run things on a different computer that would be next to impossible. SQL is a big thing in terms of standard languages, but Spark is gaining some traction and it is interesting to see that. I would predict that if Spark is being replaced with something, I would expect them to try to keep the compatibility with the API to be able to help people to migrate if possible because it is becoming standard for running replacement for heavy lifting.
**Boaz:** Tell us about your day-to-day. So, what are you working on now? What does your job look like?
**Gunnar:** My main focus is iterating on our architecture. So basically, more technical, reworking the typology of how we have our sizing of different components in the data stack is my main thing. Obviously, I am doing various other things, but that is the big thing.
**Boaz:** Where do you want to see two years from now? Where do you want to see Klarna in terms of data capabilities?
**Gunnar:** My dream would be to have CRO concurrency concerns basically. I would want to have a situation where everyone who just wants to do some data processing would be able to do it, and just pay for what they need and not have to worry at all about where the data is. I think that is the main point, even though we have a fairly consistent environment to stabilize the situation where you might be sent to a specific place because of having specific data needs. I would want to break that down entirely and just let people choose the capacity they want and ideally the tools they want. But, I think that is a utopia and I do not think you would not be able to offer everything, but at least like, do you want SQL or Spark, that type of decisions and have frameworks that just support you to the work you need.
**Boaz:** With the data organization that is so big, how do you guys even go about making these decisions? Making decisions that would affect the long run? Can you tell us a little bit about the culture of decision-making around data at Klarna?
**Gunnar:** Historically, it has been a bit unclear, but what we are doing now is going full-blown with the RFCs and the ADR processes for everything. So, to make sure that people have an opportunity to raise their voices about the different decisions, but typically these things also take some time, so it is a process of like, "Hey, we are doing this." Maybe I would propose an ADR for it and then I would get some pushback and then we would move forward with it. But, it is becoming more and more formalized and I think some people love that, some people hate that, but it becomes a necessity when you are growing. You just realize that other people have solved this problem of having growth and it is like some kind of administration, but you can no longer tell people like the coffee machine. We are going to make it change and then it is like, "Oh, but I'm in the Toronto office." "I did not hear what you said at the coffee machine."
**Boaz:** Absolutely, there are the challenges of how you need to adapt the way you work for scale. We feel it also at Firebolt but we have grown tremendously, we spread worldwide and we need to invest more time in writing things properly, sharing them properly, and encouraging people to access them.
**Gunnar:** Exactly. But, how big are you at the moment?
**Boaz:** It is like a drop in the sea compared to you guys, 200, doubled within a year. So it feels huge for us.
**Gunnar:** Yeah, but you need to be prepared for the growth.
**Boaz:** Looking at that journey now, imagine you would do everything from scratch. If our listeners go through the same journey, modernizing their Hadoop and leftovers from on-prem and now sort of are all in AWS, what would you have done differently now that would have saved you some time or headache, that you know today?
**Gunnar:** I think it is a boring standard of things, but I think I would not necessarily focus more on testing documentation, but on strictness on the get-go. I think this is something that will always bite you when you end up making decisions. "Hey, we can do this a bit faster if we cut the bit down on the strictness that you want to have them." And I think going forward with more strictness, this is how we expect data to be produced. This is exactly how things should work, as opposed to letting us try to solve this local problem for now and later, we will implement some chemistry. So that is a very tough transition to make. So, I think that is what I would change. So try to be stricter from the start, and maybe the way you would handle the exceptions because you will always have someone breathing down your neck and forcing you to take some shortcuts. I would probably have some kind of exception process and just gather the poor technical decisions that were made with a clear ambition to move faster. Because I think that would be helpful not to solve those cases, but to get an overview of what we need to learn for the good future cases. So, that would be what I would change I think.
**Boaz:** Awesome, thanks. I am thinking, what else we have not covered? For Klarna, you are dealing a lot with architecture today. When was the point in time where architect for the data became a full-time job, became something people decided to do, we need architects, people to do this as a side job?
**Gunnar:** That is a good question. I think it probably must have been like four years ago or something. And, at the time, it was a product manager who stepped into the first architecture role. I thought it was a bit weird then because I again like titles and how we think about it. But now, it makes a lot of sense to me, I think. I come from a data development background entirely. For me, I realized the type of things I need to adapt and pick up. It is a lot of actual product management because you are building a product thing and you need to think of it that way and we have been looking at architecture like the way we work. I access the architect of the domain on the actual technological platform. But then looking at how we drive architecture, on data modeling, that is more of like team responsibility in a way as well and I think that is all scenarios where it is a bit tricky to like find the exact sweet spot for help to do that. Like how much you should, again, like centralization versus Federation. Like, how do you get everyone to build the data models that are consistent with each other, without having someone centrally telling them exactly what to do and I think that is one of the things that is a bit of a challenge. I do not think there are some standard solutions for doing it, but this surely I need for like, having syncs on how to do that type of architecture and we have some things that come centrally in terms of how you should name your fields or the different like counter codes, trivial funny example, but like, it is easy to standardize exactly how you do it. But then it is just a never-ending list.
**Boaz:** Interestingly, the product person moving into the data or like you said more and more thinking about data, as product, which is by the way, natural link back, sort of to the data mesh story, because in there the causes of the data as a product is also heavily included in the story, but it is true. At the end of the day, I think more and more companies are doing that without noticing just as you guys are going to notice we are doing something that feels like it is a data mesh. Many companies that did, have realized that unless we treat data as if it was a product and understand end-users who are, the internal, could be analysts or whatever, it will definitely make our life easier down the road and it just has picked up like crazy, which is true. What are the top data use cases, workloads that run today, that the company is very reliant on and is interesting?
**Gunnar:** I think that is a hard one to answer, but obviously there are processes that are more important and central and they might have a smaller scale of what they are doing, but be extremely valuable, but then you also have the actual scenario where almost all the different product development teams are doing, like AB testing of their features and when you follow up like train models for how we are doing our decisions and every company is doing bookkeeping and the financial reporting, I think as a company, we are more skewed towards both the product development and underwriting aspects. I think that those are what sets us apart a bit, and it is also partly why we have probably a different load profile from a lot of other companies. I would predict that other banks are not doing the same type of large data processing in the same sense. Because if I look at their web pages, it does not quite look like they are doing AB testing of every feature, it looks more like someone thought something very sensible through and then built it. Then it is Okay. But, for us it is more like to the core; if we release this feature in our app, what will the implication be? And, those types of things. I think that is the type of data use cases where an analyst in a team that is working with a specific module, would want to know, does this work better than if we do this little tweak and this thing, will more people realize that they need to look at this thing now.
**Boaz:** Klarna in essence is very modern when it comes to everything, the data, maybe not what we are used to seeing in the traditional finance world, but definitely representative of the new age, the modern FinTech companies, where everything has to be data-driven and data is entrenched in across the departments, in their decisions and how they work.
**Gunnar:** I think, there is never a scenario where we take positions, and it is okay to not have any form of data. We really need to back up the company and a lot of these things to make sure that the people are not having to do cut, paste calls.
**Boaz:** I wonder for our listeners try to look up Klarna has a very cool commercial that was viral with the fish, sort of moving down, how do you call it, a slide and then smoothly sort of moving across the floor and then the message is just smooth or smooth payments or something like that. So, I wonder how many fish were part of the AB test, types of fish, maybe Tuna, Salmon, and, and there was an AB test and the right fish was selected for the commercial.
**Gunnar:** Unfortunately, I was not part of that, but, AB testing is such a central thing when the core of what you are doing is just removing friction and the fish had no friction. That is probably part of the message.
**Boaz:** Very creative piece of commercial. Okay, Gunnar, this has been super, super, super interesting. Thank you so much for sharing those stories with us.
**Gunnar:** Thank you.
**Boaz:** I hope you had a good time as well.
**Gunnar:** I did. Thank you very much!
**Boaz:** Thank you, folks. See you next time. Bye-bye. Have a great day.
# The Creator of Airflow About His Recipe for Smart Data-Driven Companies (/blog/how-preset-built-a-data-driven-organization-from-the-ground-up-podcast)
According to [Maxime Beauchemin](https://www.linkedin.com/company/40719957/admin/#), CEO & Founder at [Preset](https://www.linkedin.com/company/40719957/admin/#) and Creator of Apache Superset and Apache Airflow, building a thriving company is not so straight-forward. So how did he do it?
Choosing the right system and services is key for a successful start, and can help you avoid the chaos of having too many tools spread across multiple teams.
Max walks the Bros through his recipe for a smart data-driven company, and the genesis of Airflow, Superset & Presto (with some great tidbits about Airflow's old school marketing approach and how the open source platform took on a life of its own).
**Listen on** [**Spotify**](https://open.spotify.com/episode/6FptytEqK20n6ZLDX8pIwO?si=gDBNIlVnQ9-ufTvoi0aDAQ) **and** [**Apple Podcasts**](https://podcasts.apple.com/us/podcast/how-preset-built-a-data-driven-organization-from/id1561927688?i=1000574866747)
**Guest:** Maxime Beauchemin - CEO & Founder, BI Platform - Preset
**Hosts:** The Data Bros, Eldad and Boaz Farkash, CEO and CPO at Firebolt
Boaz: Welcome to the Data Engineering Show! We are here again, this time.
Eldad: Different setup.
Boaz: Eldad and I are not next to each other at the office. We're actually at home.
Eldad: I have my own mic.
Boaz: With us is Max Beauchemin? Did I pronounce it correctly?
Max: Yeah, as good as it comes, except for people who actually speak French as their first language. But yeah, you did well. I am in a beautiful South Lake Tahoe. So, not too far from the lake and somewhere surrounded by mountains. It's beautiful out here.
Boaz: Awesome!
Eldad: Lot of people who live in those amazing places, they're polite, and then they put the background like the live background to see how amazing the place is and then I guess you are kind of enjoying the view yourself. So, thank you for that!
Max: You might be able to see glimpses of the lake in the background and a lot of pine trees, but yeah, I'm saving that view for myself here.
Boaz: I cheated by the way, by pronouncing your last name correctly. Do you know what I did? I used the amazing LinkedIn feature where you can click to listen to the pronunciation of your name. So that's how I heard you pronounce it yourself.
Max: Oh, nice. So, I recorded that at some point, I guess.
Boaz: You probably forgot about that feature, but it comes in handy sometimes.
Eldad: I never knew it existed.
Boaz: Yeah. So if you go to Max's LinkedIn page next to his name, you see this audio button, you click it and you hear Max pronounce his own name.
Eldad: Nice. We'll share the link after the podcast.
Max: That's a pretty cool feature. No one has excuses about mispronouncing my name.
Boaz: I won't record myself. I would like to see the struggle, the first attempt.
Eldad: When you put this signature, your LinkedIn signature kind of puts in action in the link that takes you straight to play off your name and the page opens up and you get Linked in telling your name, pronouncing your name properly.
Boaz: Always more innovative. Amazing, amazing!
Eldad: Tons of innovation.
Max: Yeah. Really don't mind it. Like people pronounced my name. I just go by Max. Yeah, just call me Max. And I don't mind it very much.
Boaz: Max is quite the data guru. If you haven't heard about Max. Max has actually started [Apache Airflow](https://www.theseattledataguy.com/what-is-apache-airflow-data-engineering-consulting/#page-content) back in 2014 when he was at Airbnb. Shortly after in 2015, he started Apache Superset. Later on, he moved on to found Preset, a commercial version of Superset. And going backward in time before Airbnb, he spent time at Lyft and Facebook. So, he is an amazing data professional. And you know, the rule for the data engineering show is if you listen so far, we bring in data practitioners, not vendors, nobody to sell anything they're building. So Max, even though he is the founder and CEO of Preset, will actually talk about Max as the data practitioner and...
Eldad: Burnt scars.
Boaz: From the data world and less about what he is selling in the world out.
Max: Yeah, it's interesting. Over time, I started the company a little bit more than 3 years ago and I was coding a whole bunch at the beginning. I was wearing all the different hats. Then I stopped coding very much in the past year or two, but I still do a lot of data engineering. So I'm still in the data pipeline, still building a dashboard, and still analyzing our data. So, I'm holding onto that data analyst and analyst-engineer-type role for maybe like 10% of my time, but I don't code. I don't develop as much anymore. I don't contribute as much to Airflow and Superset as I used to, just because it requires a lot of contexts, a lot of time.
Boaz: Yeah, it seems the listeners quite surely understand that Max has a problem with delegating. He is a CEO who cannot let go of his passion for data engineering, still hands-on coding until this very day.
Okay, cool! Let's get started. Before we go into actually, Preset and it's super-interesting to hear how you've built a modern data stack there when you started a company, let's go back actually to your time maybe at Airbnb. Tell us a little bit about how you got into data throughout your career and the role you landed at Facebook and then in Airbnb and we touch on all the projects that you did there.
Max: So, I started my career as what you would call today [a data engineer](https://seattledataguy.substack.com/p/tips-for-hiring-junior-data-engineers), but I was a data warehouse architect. I did a little bit of web development. Then I became a data warehouse architect/business intelligence engineer. So, I had a good run, almost a decade worth of using the previous generation tools. So things like business objects, Informatica, and a lot of ELT back then too. I would just write a lot of store procedures either in SQL server or Oracle. So, writing a lot of ETL, building a lot of dashboards, and organizing the data for the whole organization. At that time, it was at Ubisoft. So I did that at Ubisoft video game company. It was super fun. Got my foundation in data. So tons of data pipelines, data modeling, dashboard building, that sort of thing.
And then I joined Facebook in I believe 2011 or 2012. Well, so I skipped Yahoo. So I went to Yahoo. It was the birth of a Hadoop. I didn't stay there very long, maybe 2 years or so, but I remember meetings with the people who went on to start Cloudera. So, my manager's manager Amr Awadallah was there and then some of the...
Eldad: He was always there. Every, you know, impactful, bigger than this world event in the data evolution, he was there, somewhere in the background, in the foreground. It's like always seeing the picture, always seeing Max there. Two peers in Yahoo, exactly the right time, and then, yeah, sorry! go on.
Max: Yeah, I was at the right time and then I joined Facebook, which was a few years later. So I was at Yahoo in 2008, then joined Facebook in 2011 or 2012. And then, there's like a big renaissance of data tools there. So people had to rebuild a lot of things from scratch on top of Hadoop and other things. Because the scale of Facebook was just too big for the Teradatas of the world. There's just no commercial database that could scale to petabytes at the time.
And, I think like during my time at Facebook...
Eldad: Hitting Big tech companies, no database company in the world that can build something that is big enough for us and we just have smart people, so let them do it. And I always wanted to ask and since you did it so many times, so successfully, kind of incubating something in big tech, in a big company, how was it back then? How is it today? Is it still happening today? Or do people just live and open a startup? What's your take?
Max: Yeah, there's got to be phases of a new renaissance kind of era. But, I know at that time we had to rebuild everything on top of Hadoop. We built everything on top of MapReduce and there were no better databases. Oracle's very advanced database or Teradata is actually, really good. Database, it just wouldn't be like parallelism, and scaling horizontally, was just not as much of a premise as it needed to be for a company like Facebook. So, Facebook had to rebuild everything on top of MapReduce or rebuild everything with the premise of things, adding to scale to thousands of machines horizontally. So, they created this culture of, "Hey, we're building everything from scratch." A lot of experiments too. So, we had a bunch of different data pipeline tools. One of which is called Dataswarm became the inspiration for Airflow along with other ones. But there was probably like for every project that stayed on Facebook and got used by people, there's probably like a dozen other projects that didn't go anywhere. So, there's probably been like 20-30 different schedulers built at Facebook and then one or two...
Eldad: So, many dead startups. So many startups that could have happened and didn't happen because just Facebook is like let's have 50 projects on data pipelines and one of them will win. And that's you again, by accident again, you were there in that single project, again, go ahead, sorry.
Max: Yes, I was there at that time when these things were happening. Interestingly, I don't know if it's the engine or if the big tech companies like stopping innovation by putting the brains kind of inside their wall gardens, or are they actually stimulating progress with all the wealth that they generate? So, I think it's a mix of the two, but in the case of Facebook, there's a lot of really good open-source help to break the wall gardens of tech. Because people like me, I'm going to join Airbnb if I can work on open source and then my stage for impact is not Airbnb, it's the world. So, there's been a lot of people before me too, that I've had a lot of success with open source. So, I'm just kind of following those footsteps and say, maybe I'm delusional enough to think that I can do the same as people like Jay Kreps or the people behind Hadoop and some of the open-source technologies like LINSTOR I think have been an inspiration to us all too. So, I was like, oh, maybe I'm crazy enough to attempt this thing and other people have done it. How hard could it be?
Boaz: When you joined Airbnb, what did the variety of data teams look like?
Max: Yeah, it was kind of interesting. So, I really grew up while I was there. So it's hard for me to close my eyes and think about what exactly it looked like on my first day. But, I remember, well, there were like a handful of people, about 3 people, Johnson Parks, Aaron Keys, Sid, think all 3 were working on something called core data. So, I was like, people at Airbnb had suffered enough from handling raw data and they're like, we need to do some data engineering. We need to create some core data sets that we can trust and rely on. Later on, these data sets, there's too much pressure and too much pull from the different teams, so I tried to evolve these things centrally. But for a while, people working on core data, we started using early, early Airflow in production, within a few months, to build the core data stuff and some of the core data sets at Airbnb and soon after, we migrated a lot of the pipelines that existed in some previous scheduler called Chronos that was built on top of Meso and then we just migrated that.
At the time, there were like 3 data engineers. They didn't call themself data engineers. It was not a popular term at the time. I think they called themselves ETLeans or Eliens, so that was the name of the team and then there was probably a data platform.
Eldad: Data mart team, back then, the formal names.
Max: Yeah, little fun name. And then, there were maybe 10 data scientists, 10 people on the data platform and then the team went on to become, I think, there were like a hundred data scientists by the time I left.
Eldad: Phew.
Boaz: Wow!
Max: And there's just an army of data scientists. The data engineering was probably like 15-20 people.
Boaz: That was sort of between 2014-2017, right?
Max: Yeah. So the data platform included all the data functions, and was definitely north of 150 people or so
Boaz: Tell us about the context and the role and how it went about with Airflow during those years?
Max: I started a project in between gigs. So, I left Facebook with the premise that I was going to work on something like Airflow, at least as a side project. I didn't know the name at that time. I talked with people at Airbnb and they're like, we need something to manage our dags. We have tons that we're crumbling under the weight of our own pipelines. So, we need something better than what we have today. So, I was like, "oh, that sounds fun. I want to come and work on this." And then in between jobs, I started working on what now became Airflows, I had I think a two weeks sprint. So, instead of taking a vacation, I just decided to start coding this thing and I put it under my personal GitHub. I would join and say, "Hey, it's already open source." Like, what are you going to do about it? So, it was under my personal GitHub and then the moment I got there, I think we got something in production very, very quickly. So it was like within a month or two with me being there. We had some data marts in production very quickly and within 3-4 months, we had all the core data pipelines and we migrated a lot of the Legacy pipeline to Airflow. Also, moved the experimentation framework as a gigantic dag of thousands of tasks to compute all the metrics and all the experiments data. So, there was a big hunger internally. There are a lot of data professionals and people who need to schedule arbitrary workloads. It was also at a time when things like DBT did not exist. So Airflow was preferred for SQL, for scheduling mountains of SQL too. And then there's a bunch of, you know, people doing all sorts of crazy stuff with RPython and IPython and notebooks and Java or whatever it might be. So, there was really a need to schedule thousands and thousands of jobs, big hunger for that. So, that's how it took off.
Boaz: When did you notice that this is being picked up and is extending its reach beyond this project of mine and getting popularity.
Max: Internally, I cannot understate how much I visited or I did some evangelism around Airflow. I was super early. I would visit any company in the valley that would show interest in me coming and talking to them. I would just go and visit them and answer their questions, talk about their data engineering challenges and kind of convince them that Airflow was probably a good solution to that stuff and the Landscape at the time was oozy.
Eldad: That's how you do it, baby.
Max: Yeah. That's it.
Eldad: You go out there and knock on doors and you just do it old school.
Max: Yeah. And I never wanted to start a company or that was not my intention at all. I just wanted to build something relevant and impactful. So, when the VC started approaching me to say, "Hey, why don't you start a company?" It's like, "are you crazy?" Like, that's like, "why would I do that?" And I'm not an MBA. I'm happy. I used to call it doing open source for the right reasons. I'm just here to evangelize and build something useful, for the longest time that was really what was driving me. And then, it was really progressive. The popularity comes like one issue, one PR at a time on GitHub. You're like, "Oh, here's someone that seems to be associated with this company name." Then the mailing list. So, it is very progressive. So it goes from a handful of people showing a little bit of interest to a small crowd and eventually, it's like a mob. And now I heard recently there's I think astronomers did some analysis and there's probably north of like a hundred thousand companies using Airflow today
Boaz: Wow!
Max: It's just insane.
Eldad: Insane.
Boaz: Insane. From Max refusing to start a company.
Eldad: Shut it down, now.
Max: It has a life of its own. So people ask me, when do you know that your open source product is very successful? You know when, if you would try to stop it, you couldn't, right? If I would try with all my might and all the resources in the world that I have to stop Airflow, I could not, at this point.
Eldad: Amazing.
Max: That's when you know, it's successful.
Boaz: It's like the terminator, once machines take over, you can't stop.
Eldad: AI all over again.
Boaz: And then you also started Superset?
Max: Yeah. The genesis kind of story for Superset. Well, first there's a delusion of like I can build a BI tool. I think that came from Facebook as Facebook people had built all sorts of little visualization tools that were very simple to use, very fast time to chart and time to dashboard. You have a data set ready, you can build a chart and dashboard of that data set in no time. It doesn't have...
Eldad: It's so funny, you mention it. One of our engineers joined us from Facebook, maybe 4-6 months ago. And one of the first things he mentioned was how nice and how well data visualization and data consensus is at Facebook. If you send a link, it opens up a page with all the right charts. There's a discussion on the data. So, just now you're mentioning took me back to that conversation, and Boaz has a kind of query history and all the discussions we're having on how to embed data within conversations. So, yeah, it's super interesting.
Max: Yeah, no, totally. I think it's great. Like to talk and you know, I hate to overpraise any big tech giant or whatever, but like we got to give credit where credit is due. I think like a lot of what we call today, the modern data stack, there was like a microcosm of all innovation and modern data stack happened 10 years before at Facebook. They had something like a, Idata which is a data portal, data catalog with a full lineage of everything at Facebook, you can navigate the metadata graph of all the data objects pretty well and do like lineage analysis and impact analysis. There are all sorts of little visualization tools, little schedulers, and then the databases like a database called Scuba. It's a little Druid and Firebolt maybe too in some ways. Like really fast, real-time in memory.
Eldad: They have a really strong theme by the way. Writing a vectorized query engine is really strong.
Max: The Presto team too.
Eldad: Respectful.
Max: Yeah. Very impressed by the quality of the gray matter at Facebook. Like people are empowered to build new things. Things like Scuba, things like HiPal was like a little bit of a SQL, notebook type thing. There's something called Unidash, that's a dashboard building tool that came a little bit after my time and just this culture of like, if there doesn't exist, I'm going to build it. So I think there's like...
Eldad: Now, they migrate everything to the Metacloud.
Max: That I don't know about. Metacloud I guess is their new data center, I don't know.
Eldad: So all of your friends at Facebook, everyone was sitting there in the room and someone heard, it became this huge, huge open source success beyond control. So you had to feel special. So you went on and said...
Max: I'll do it from Airbnb. I looked to open source on the Facebook stuff internally. It was just difficult because everything was tangled up as you know in data, all the different systems, it takes a fair amount of duct tape and chicken wire to kind of hook a data platform together. And it is really hard after the fact to take a piece of that data platform, that's all duct tape and chicken wire, the rest of everything and cut that out to serve it as an open source project. So, I think it's been done in the past, refactoring an internal technology as an open source project. Some companies have done it with some projects, but I'm guessing it's always with the Premise that it might be open source one day. So let's keep this as a microservice that works well by itself, but yeah, I realize..
Eldad: If all the engineers leave us, we can solve it by open sourcing it and kind of getting engineering love back.
Boaz: We need Max to evangelize it.
Max: That happened in the past with Premise. If we open source this, like the community's going to build it, I think in my experience it is not, as if you guys are like, Hey, we're just going to open source Firebolt so that we get hundreds of contributors for free, that typically does not work. It needs to be open source from the get-go and I think it's really often like a handful of core contributors that are bred, that are very close to the core of the project that built the bulk of it.
Boaz: Now, you started Preset in 2019.
Max: Yeah, early 2019.
Boaz: Early 2019. Okay. So, here's the...
Eldad: So what happened?
Boaz: Yeah. What happened?
Max: What changed? Well, I didn't like my life's goal pretty much to push Superset forward since 2015, so it has been like 4 years that I've been working on this thing. I was like, I want open source to come and compete in business intelligence, and data visualization. That was my life goal. I want to build something relevant that's in every other company. So that every tech company, every company who does data, everyone has heard of Superset. So, that was my goal for the 4 years before.
I'd been sponsoring or incubating my project inside Airbnb and inside Lyft for a little while and these companies are super nice and I had this small team of people working with me on these things, pushing this thing forward. But when I started talking with investors, they were telling me that or it became really clear to me that it would be a really great way to take on capital to really be able to push the open-source project forward. So really it's in the vein of like, if I raise my A round with $12.5 million, that allows me to hire a bunch of people that are very dedicated to working on Superset and making Superset great. So, there's always this duality too, like, hey, you need to build an open core. You need to build some crust. You need to build a successful company that makes money too. But I've seen other companies do it successfully, like companies like Databricks, Confluent, and companies of different shapes and sizes. So, I thought I could navigate as well as anyone else on how to give back to the community to grow an open core, but also like to build something that we can sell in and around it
Boaz: Let's get back to Max, the practitioner, for a second. So you're starting a company. Obviously, you're super experienced with data. Walk us through how you thought about building an organization that is data-driven from the ground up and how you went about that.
Max: At first it's like, when you don't have a product out too, I mean, building a company, there's just like six months or a year just like trash. You need to like set up a bunch of stuff, make a few hires, just kind of get going and there's not a whole lot of data at that point in time, but you know, one thing I've been talking about lately is this idea of like data native companies, the same way that there's like cloud-native companies or digital native companies, companies that were born in a certain era act in different ways. And I think like this generation of companies, like Preset, we add just like really easy access to things like BigQuery and Snowflake to things like DBT and Airflow. Like, we didn't have to build our own scheduler. We just like to pick one up. So you assemble your data platform. Now Fivetran. Now, there are open source counterparts, like Airbyte and Meltano. So for us, we just like to sign up for Fivetran. We get our HubSpot data and our segment data all centralized in a data warehouse pretty easily. Like you can kind of assemble these, pick up these pieces off the shelves and they're all like pay as you go. So, they're really cheap or free if you don't have a lot of data. That's how they get you though but...
At Preset, we offer a premium of up to five seats. At Preset, that's a really good offer. So if you're a small startup, if it is like 2, we make sure that works. And then it's like pay as you go.
Eldad: Is it free, full-featured, like any big limited edition or really kind of about 5 people?
Max: It's like 90% of the features are in. I think we block some things like alerts and reports, where alerts are delivered.
Eldad: Held up, held up and stuff like that up, 5 people and you need that to go away. That's how exactly
Max: But if you're 5 people, 2 or so, then the premium tier, so there's premium and then the premium is like $20 per user per month. If you have 5 users that's $100 a month, and then you can have the SSO and LDAP on that tier too. And, we have an enterprise tier too, but what's nice with these tools is you can sign up for them in 10 minutes and so you can start sending data into BigQuery and the bill is going to be like nothing. So, you can assemble a data platform for us. We picked up Fivetran segment, DBT, BigQuery and just started pushing a bunch of data into BigQuery and started reorganizing the data with DBT. We have Airflow now, too. It took a little bit longer for us to really need an orchestrator, but as you assemble more things, so we have things like Hightouch now, too, which is a reverse ETL tool to send maybe product usage data back to your CRM systems. So, now that we have a complex enough data platform, we need something like Airflow, but it took a little bit longer for us to really have a need for that. But super easy to pick up Astronomer just like it's installed in a half a day and you can start jamming on that stuff very, very quickly.
Boaz: How many you...
Eldad: It's scary and crazy, how fast you can kind of play with those building blocks today and how hard and painful and long it was just a few years ago, which is crazy because it was not long enough in our space, you don't appreciate, you said something, you said kind of as new companies get born, they get born differently, and we are getting old, right? We found companies, but your engineers, your people, they are fresh and they are coming with a new mindset. Just, it's crazy. So, it's a lot of ecosystem things.
Max: It's late if you compare, like, say cloud-native companies, like companies like Airbnb that were born, say like 2000. Airbnb is a little early for that, but companies born after 2010 or so are all built on AWS. And that is also hard to appreciate just how hard it is. We used to order servers and machines and we had to wait for the boxes to make it to the data center. That we were renting and we had to wait for this, this and then, usually balding with a ponytail, to like install that server somewhere...
Eldad: I missed that. I missed that. One of my best friends from my previous startup Sisense, Shlomi. Hi Shlomi, if you see us and listen to us, we love him so much and he used to do it old school. He used to fly to New Jersey and he used to install stuff there and sit there in the basement and everything. Now he's running AWS. It's not the same anymore but in a good way
Max: People would make a nice setup for the network wires, like to kind of weave them together in a specific way and use zip ties of different colors that meant different things. But yeah, so the super transformational companies who were born after, EC2 was born, EC2 and S3 as core components of AWS completely changed the game. You can just be like, hey, I need a hundred machines and you get them. But then I think the data space, it's a prolonger for us. It should be like, I need a data warehouse and I need a scheduler and I need a data visualization tool and you don't need to talk to salespeople. It was not expensive either. All you want is a database tool. That's like 50K, now it's like $20 per user per month. Like why? Like just set it up. You can use it today, you know?
Eldad: Now!.
Max: And I think that's now transformative. Like today, yeah, you don't even need to talk to someone.
Eldad: You can though. If you want, you can and that's when you raise money.
Max: Yeah.
Eldad: Because again, 5 people, it's an amazing team, but it grows and from my experience, everyone needs that at one point or another everyone in the company, almost everyone in the company needs that. And as companies get younger, more people need that. So, it's amazing to see, I've been in this space, and Boaz as well, for so many years and it's just amazing to see that the need for data is never kind of...
Boaz: Never stopped.
Eldad: Nobody gets tired of that, so creative.
Max: There is a question though, it's like, why is data 5-10 years behind some of the software engineering practices or like data engineering is like notably behind I think.
Eldad: Expensive.
Max: Because it's expensive. Yeah. It's not a priority too and we've had tools for a long time that was kind of okay. So maybe for a long time you had, I don't know, business objects or Cognos, and you had like Informatica and Teradata and that was your stack and it was okay. And, maybe that prevented the kind of innovation that we see today. And the data team used to be just like a handful of specialized people, like 2000 to 2010 for me it was like if we had like 4 or 5 data specialists at Ubisoft and we were taking care of all the data needs of the company and that was okay back then.
Eldad: You needed IT back then. That's why and now you don't need IT anymore. That's the only difference.
Boaz: But every data that goes by data engineering and software engineering are getting closer together in sort of bi-directional ways. I always say like a decade ago, a software engineer tasked to do a data project or do something with a database would say that's not for me, that's a job for a student and look down on those tasks. And, today you see that more and more software engineers want to do data-related projects and tasks. So much of the more interesting stuff is happening there. And you see also the way data engineers work is so deeply affected by what's happening in the software engineering world and that's the direction everything's headed.
Max: ML and data science have brought a lot of excitement to data too. So, being able to train models and to predict at scale, I think it is something that brought a lot of attention to data and data engineering. If you want to have the algorithms, you want to have some good ranking and feed ranking, and you want to run a lot of AB tests and know what to ship in your product, you need to have good data. So, I guess that pushed more requirements on being data-driven
Boaz: Yeah. So, back at Preset, how many people are in the company these days?
Max: We're about 70 people or so
Boaz: 70. How many people are in-charge or are on top of that data stack?
Max: We have 2 data specialists now, but those are new folks that started less than 6 months ago. There are an analyst engineer/data analyst and a data engineer now, but it took quite a while for us.
Eldad: And Max.
Max: I mean, but I do spend a fair amount of time doing so...
Eldad: Nobody wants to join that team.
Max: There's definitely pros and cons of joining a team with a super experienced data engineer and I'm kind of anal and picky about certain things. But at the same time, I can be a good peer, a good mentor on some things. Yeah, so we're like two and a half people or so on the data team itself. I think it is right, what you said before, like building a company, you don't realize just all the functions that are required. As a tech founder at first, you're like, Hey, I can do a little bit of everything and it's mostly about building a product that sells itself. And then you realize, "oh, no, I need a sales team. I need SDRs. I need like customer success folks that can help my customers be successful in my product." We need education and the whole marketing side. I'm not even going to talk about it because there are a lot of specialties there. There are a lot of data-driven processes there too. So, it's hard to understand, I think, as a tech founder, what really you're getting into and the vastness of the skills that are required to build a successful company.
Boaz: So for other companies that maybe now in the process of building data stacks from scratch or modernizing. We enjoyed talking about how fun it is nowadays with all these building blocks, but still sort of what lessons learned in implementing a modern data stack can you share? I mean, if you had to do the same thing all over now, what would you have done a little bit differently?
Eldad: Faster. Just faster, for three years everything.
Max: I think the recipe that we chose works very well for us. Fivetran works very well to do data sync. We're a small startup, we have more SAS systems than we have employees. We have more than a hundred SAS services that we do, which is kind of insane. There's a whole portfolio of things that you want to use and there are some really core ones. Your CRM is really important and we use HubSpot, would probably pick HubSpot again, and then Fivetran to bring the data from the different systems into and land it into a warehouse. So that's a really easy data acquisition thing. Then, we have to build our own like scraping and analytics events, inside your product. I'm sure inside your database, you have a lot of sensors and you have a lot of logging, every time someone runs a query, every time someone does an action for us. It's like anytime someone creates a chart, saves a dashboard, alters a dashboard, invites someone, all these analytics events we need to bring to the warehouse. So, we use Segment as a little bit of an ingestion layer for that.
Eldad: You are obsessed with observability?
Max: Yeah. To me, there's operational reporting on one side and that's just Datadog for us. So that's more technical like the hardest systems and machines doing and then there's like their analytics stack which is kind of interesting because they're two different worlds and maybe they don't need to be and I think they're less different worlds than they used to be in the past. I'm sure, at Firebolt, you probably have use cases across the chasm of things like operational analytics and more traditional business analytics; but for me, I'm much more on the product analytics side of things.
Eldad: We're very flexible. We let employees use whatever tool they want, as long as it's Firebolt and if they're not happy, if they don't like using Firebolt, it's okay. It's also okay of course. Then, they have Kafka Streams and they should just manage on their own. I've heard Honeycomb, by many people, and I think there's a great project going on. I think, on our end, we use observability now mostly to define success, enable engineers, and have them figure out how to kind of define the success of a feature they're releasing with observability. So it kind of focuses them and removes a lot of noise, it's just going crazy on matrices and events and just choke the system. So kind of, we went through that journey. How is it with your startup? Where are you now on the observability evolution?
Max: On the observability side of the house? So there's a pretty big chasm between what we do for, I'll call it business analytics or product analytics for us. So this is all for us, it's BigQuery, DBT, Fivetran at Preset to analyze this data. On the observability front, we take advantage of everything that Datadog offers like from their traditional logging to metrics logging. They have some sort of time series database behind the scene. And this is a world that I know a little bit less about, but yeah, Datadog got also, APM, like Application Performance Monitoring.
Eldad: That's when you get into that, don't go there.
Max: Yeah
Eldad: If you go into APM, it gets expensive.
Max: Yeah. So, we used that stack for observability and then we used a different stack for business analytics, and then that's much more, the place that I come from. The lines are getting blurred. Databases like Firebolt, I think, now can take either workload, like you don't need a time series database for your observability and a different database for your data warehouse as much anymore. I don't know how you guys position yourself in relation.
Eldad: Consistency is key to getting that done properly, and we can spend the whole podcast on why consistency is so hard on low latency databases, but why it's so needed, but yeah, we use also multiple systems, not just Firebolt obviously, and we've gone through multiple different products and we just realized that sometimes different tools solve the same problem for different people in a better way. And that's also okay. But we do kind of try to manage cost, kind of adding those tools adds up and get this kind of one point overlapping feature sets and you get confused like 5 logging systems. So, we are also kind of trying to go on a diet, and once in a while stop for a second and say, okay, we've been trying things out. Let's stop for a second, and just pick one or two winners, like the fact that you can connect stuff so fast today only gets simpler, makes that...
Max: But that creates the reverse problem, right? If I can spin up both like Firebolt, Imply, PnO, and like three other databases like today, then maybe I will start using all of them. And then you have this accumulation of systems. I talked about that in my Airflow Summit talk, setting up our data platform at our small startup. And yeah, it's so easy to set up systems that maybe that creates a need for more orchestration and more metadata catalog type things. Because all of a sudden, maybe you have BigQuery and you have Snowflake and you have Firebolt for your super hot data sets and you have a bunch of BI tools because I don't know, they're easy to set up and different teams prefer different tools. So, you end up with chaos in some ways. So, you need to keep things somewhat constrained.
Eldad: Spend on marketing. So you need to spend more on marketing to focus and then help people understand.
Boaz: It's easy to set things up, but there are also things that have a lot of information in the community. There's all of the knowledge sharing going on today. It is much easier to talk to colleagues who have tried things hands-on and you know them in person, you trust them. It's not like far out anymore. It's very easy to get information and find somebody who is really talkative, who's experienced.
Max: It's easy to try software too. You can just go on and try it. But yeah, the reverse problem is like maybe you end up with a startup with 70 employees and 500 SAS tools. But here's, I think that's interesting, like talking to the parallel between data engineering and software engineering. On the software engineering side, we accept that you need multiple databases, multiple languages, multiple frameworks, and multiple libraries like there's such a diversity of systems, services, frameworks, libraries, and languages and it's well accepted, right? There's no one that says like, "oh, no, we should have like one language to rule them out" or there should only be JVM. So, there are probably still people that think that. But I think we accept that it's going to be a highly diverse, different solution for different people, a lot of microservices that we accepted in software engineering. In data engineering, it's like picking one data warehouse and sticking to it, picking one BI tool and sticking to it.
Boaz: It's like in software engineering, people accept that there are multiple ways to reach a great outcome, even though you chose a different path to get there and in data, it's like, No, this is the perfect stack he should've chosen and no, any other whatsoever.
Eldad: No, it's just licensing guys. It's just licensing cycles. That's all, that's the whole difference, and people go and build and decide on a stack and they need to close the license and they negotiate and everything takes time and they close the year subscription or buy credits a year in advance, whatever, and then they take the project to production, and then they move on, it works, they move on. They add another project with a new tool because they like or think that is better and they move on again.
Boaz: New license.
Eldad: But I think like and it's a good and bad thing. On one hand, people have much more flexibility and options and on the other hand, they can just pick another tool. They can pick another tech. So you constantly need to justify yourself. You constantly need to get better, and at Firebolt, for example, consumption, it's even more. You constantly need to earn your customer's consumption. If they consume, it means they're finding you valuable. They don't consume, you're not valuable. They're not using you. And it's very out there. There is no way to hide it. There's no way to hide or postpone the contract. So, advantages and disadvantages.
Max: The trend seems to go in a direction of things becoming more or at least at larger or faster work organization, there's more pressure to decentralize things and to democratize things. So every team can decide what they're going to do, and what tool they're going to pick and use. And then of course that creates a different set of problems. But if you want to move fast, centralized structures don't work as well. Because you need consensus and consensus is expensive.
Eldad: Return of business objects, like retro-release, SuperBGA, business objects.
Max: Yes. SuperBGA, business objects, high resolution.
Eldad: Exactly.
Max: Not working on Macintosh, going back to the desktop products.
Eldad: Old pricing. So you get the retro $1 million license pricing as well. It's amazing. A lot has changed since then.
Boaz: Yeah, good. Max, this has been awesome being in the company of two data geeks and founders. I'm sure we can keep talking for too long, but we are running out of time. So thank you, Max so much. It's been super interesting, and good luck with everything at Preset, and we'll keep definitely watching out for what you do next.
Eldad: We'll see you soon.
Max: Yeah, it has been super fun. There's a bunch of things we haven't talked about, so we got to solve a bunch for next time. But I learned from a pretty good conversation with Eldad that he worked on early MDX compilers, which is MDX.
Boaz: Wow!
Max: MDX like multidimensional version.
Boaz: Don't get him started.
Max: Yeah, so we should do it. I wanted to talk about data apps. I wanted to talk about what it takes in a database, and what properties we need from a database engine for the next wave of highly data-centric applications. Like those data apps. So, it'll be material for another show, maybe in a few months, we connect back.
Eldad: Absolutely.
Boaz: Absolutely.
Max: And talk more.
Boaz: Okay, Max! Thank you so much!
Max: Thank you for hosting.
# How Rising Wave Is Redefining Real-Time Data with Postgres Power (/blog/how-rising-wave-is-redefining-real-time-data-with-postgres-power)
In this episode of The Data Engineering Show, the bros sit with Yingjun Wu, founder and CEO of Rising Wave, to explore the innovative world of stream processing systems. Yingjun shares his journey from academic research to creating a Postgres-compatible streaming system that drastically reduces resource usage. They discuss how Rising Wave's S3-based architecture and Postgres compatibility provide advantages over traditional systems like Flink, and explore the increasing role of Apache Iceberg in data pipelines.
**Benjamin - 00:00:00:**
This is interesting, right? Because this seems to be a general trend in data management at the moment, in a sense that when Snowflake came up, they were the ones who invented in many ways, okay, decoupling storage and compute, using object storage as the lowest tier. And now you're seeing these kind of, I would say, new age systems using the same underlying principle to kind of disrupt that space. So RisingWave for streaming, you had WarpStream, which was acquired by Confluent, kind of doing it around Kafka and so on, right? So I think it's super interesting how this dynamic right now is actually playing out throughout the entire data infrastructure stack.
**YingJun Wu - 00:00:46:**
Yeah, we are definitely the first one to build a kind of RR3-based or systems who are in the stream processing domain, even earlier than WarpStream.
**Intro - 00:00:56:**
The Data Engineering Show is brought to you by Firebolt, the cloud data warehouse for AI apps and low-latency analytics. Get your free credits and start your trial at Firebolt.io.
**Benjamin - 00:01:09:**
All right. Hi, everyone, and welcome back to The Data Engineering Show. Today, we're super happy to have on YingJun Wu, who is the founder and CEO of RisingWave. Welcome to the show. For the listeners who haven't heard of RisingWave and maybe Eldad, do you want to tell us what the system is all about, kind of what you're working on? I think it's super cool, so I'm sure everyone else would love to hear more, too.
**YingJun Wu - 00:01:29:**
Sure. First of all, thanks for having me here. And I'm YingJun Wu, and I'm a founder of RisingWave. So, yeah, people always ask me about, okay, what are you currently working on? Like, I would say, okay, definitely, what is RisingWave? RisingWave is a stream processing system. But this day show a lot of buzzwords about the Iceberg. We could probably discuss that later. But essentially, we found a company four plus years ago. Before that, I would say the Amazon Redshift as well as IBM Research Almaden. So, actually, people sometimes will ask me, okay, why Redshift is like a Bash base? So, the data warehouse, right? Well, why you started working on something like stream processing? A little bit of history of myself. I have my PhD in stream processing and database systems.
**Eldad - 00:02:16:**
You had no choice. You had no choice, basically.
**Benjamin - 00:02:19:**
And you spent some time in Andy Pavel's group as well, right? Like who's a previous guest on the podcast. So I love how we're like connecting. Exactly. Former guest of honor. And now we have you as guest of honor.
**YingJun Wu - 00:02:30:**
Yeah, I'm not sure whether we have enough time to dive deep into details. But essentially, when I was with Andy Pavlo, we were doing something like transaction processing. And essentially, at the time, there was a concept called deterministic transaction processing. Basically, it's like, okay, reorder the transactions before execution. And essentially, the concept is the same as stream processing. And that's why I have that. During my PhD, I studied both stream processing and transaction processing. Because at the time, there was just some concept that would link these two things together.
**Benjamin - 00:03:03:**
Nice. Cool. And then fast forward after your PhD, you said time to build a new database system.
**YingJun Wu - 00:03:10:**
Oh, yeah, definitely. I mean, Redshift was, yeah, IBM, I mean, dbt, right? Well, yeah, Redshift actually at that time was probably 10 years old. If you look at Redshift or batch-based system, but well, how about, I mean, stream processing? Because well, essentially, in Redshift, a large percent of the data is actually came, was ingested from Kinesis or Kafka, right? And these are streaming data, but there was no stream processing systems. Well, I mean, there were some stream processing systems like at the time, Spark streaming, Flink, these are still popular these days. At the time, there was also KSQLDB and Apache Storm, but these kinds of systems were not that easy to use. And there were some fundamental issues with the key concept was called state management. We can probably dive deeper into details, but yeah, there were some issues. So I feel that's right. Yeah, it's the right time to build something new. And that's why we find that company.
**Benjamin - 00:04:03:**
Nice. That's awesome. So tell us about those early days. I mean, like, this is one of those, like, building a system from scratch is like... Crazy daunting challenge, right? Like you need to do everything, like parser, planner, runtime, kind of all of it for streaming system, integrating with storage. Like how long did it take you to get to a first version of the product? When did you launch it? How did your first users like it? Take us through the journey of RisingWave, basically.
**YingJun Wu - 00:04:30:**
It's definitely what's appearing. I mean, during the first two or three years, whenever you started building something, let's say a database system, so I have very hard called low-level data systems, beta-infra system, then essentially it's not about your concept. It's not about your, let's say, the design principles, right? It's more about, okay, how old the system is. It's more about the trust. If you say, tell people, okay, my system was just two years old. Okay, then people will say, okay, that's a new system. That's pretty cool. I will probably take a look at that. And I probably will follow you, but nobody will take it seriously, to be honest. And in our case, we did not really have any customers for the first two years. Actually, the first two and a half years. And the first POC user was like tiny startups. And we found startups because, I mean, they actually just having startups, trusted startups. Also, we know each other. And yeah, so they adopted RisingWave. And I definitely also pitched the RisingWave to some big companies. So, but the world would ever say that, okay, like what, RisingWave is like our new system. And you do not need to put it into production. Just to use it for in some of your testing environments. Because it's so easy to use. We have actually have the single binary. We are Postgres compatible. And I mean, RisingWave itself is not single binary, but we actually made a single binary version just for some customers to self-deploy. No Docker, just a single binary, just like DuckDB. But well, anyways, for the first two years was just like tough. And I think we were definitely blessed that we got first a few customers in the year 2.5. We have at least probably three of them. One of them was like a crypto company. At the time, it was also a pretty tiny company. But actually these days, well, they are still all customers and paying us six figures every year. They've become big. And they grew from a startup to a big company. And the other two are both pretty big companies. One of them actually are Fortune 500 companies. So I feel that's why it's like just a blast and we got lucky. But this is why it's like because we have already had a pretty good customer base and we never need to answer questions like, okay, whether your database system is trustworthy or not, or we just have a page showing all customer logos and that's it.
**Benjamin - 00:06:53:**
Nice. Cool. Zooming out a bit, like take us through the current stream processing space, right? Because there's, nowadays, actually a fair amount of systems. I think like the OG is kind of Flink. That's been around for quite a while. There are systems like Materialize that are also Postgres compliant. There are systems like RisingWave now, which are like built in Rust. Okay, Materialize as well. Kind of like, how are these systems different? What makes RisingWave special? And why would I choose RisingWave, I guess?
**YingJun Wu - 00:07:21:**
Okay, so I believe that's for RisingWave is definitely a number one stream processing system in the world today. I mean, if you want to add something, for sure, I mean, go ahead. I probably removed my sentence. But, well, I mean, personally, I feel that's for we are the number one stream processing system at the moment. I mean, definitely Flink. Well, for Flink, well, people use it for, I mean, if people need to have, let's say, a Java API, for sure, yeah, go for Flink. But, well, if you just focus on the SQL, then it's really hard for us to lose to any other systems. If I talk about design principles, well, then we have several things I want to highlight. The first one, actually, I think we have already mentioned we are Rust-based. I mean, Rust-based, well, it's not like a design principle, but, well, Rust always gives the people the feeling that it's cool. So, yeah, we attract a lot of early adopters, probably just because of Rust. And second thing here, that's well, S3 is the primary storage, basically the corporate-companion storage architecture. And some people would say, okay, it's like Snowflake-like architecture, right? So we started from day one, it's like S3-based architecture. And if you check out some other systems, well, I mean, like Flink, well, they probably just recently announced that, well, they were supposed to put that. But we actually have already adopted this architecture from day one and basically baked it for in production for over four, lots of four years. So there are a lot of like tricks I probably could discuss later.
**Benjamin - 00:08:50:**
This is interesting, right? Because this seems to be a general trend in data management at the moment, in a sense that when Snowflake came up, they were the ones who invented in many ways, okay, decoupling storage and compute, using object storage as the lowest tier. And now you're seeing these kind of, I would say, new age systems using the same underlying principle to kind of disrupt that space. So RisingWave for streaming, you had WarpStream, which was acquired by Confluent, kind of doing it around Kafka and so on, right? So I think it's super interesting how this dynamic right now is actually playing out throughout the entire data infrastructure stack. You have like Estuary, I think, who we had on the show, like Danny Palmer before, who were like building ELT Tooling on top of S3. So I think it's like interesting how that's just happening everywhere, basically.
**YingJun Wu - 00:09:35:**
Yeah, we are definitely the first one to build a kind of Raya 3-based system in the stream processing domain, even earlier than WarpStream. We will be funding in 2021 and all the other systems will probably later. But essentially there are a lot of, there are a lot of like tricks and I have to tell you that it's pretty transparent. Essentially over the last two or three years, we were also thinking about, okay, whether we were doing the wrong design decision. And I can hear you more about that because for a lot of, we are stream processing systems and S3 has two characteristics. Well, one is S3 slow. Second one is S3 expensive. So, I mean S3 storage is not expensive. It's just the 23 bucks per terabyte per month. But if you look at where the S3 is access rate, they actually are charged for get input in all these operations. So I mean, that's super expensive and essentially it costs us a lot of money. And that's why we actually revisited whether we were asking us, okay, if we want to rebuild the system again in 2025, should we adopt the exact same architecture? But essentially we did a lot of optimizations. And now our answer is that's for years. If we build a system again in 2025 or even 2006, we'll stick with this architecture and there's no change. But definitely we prioritize some of our roadmaps. But essentially it's super good for several reasons. The first one is that you never need to worry about the state management because everything is in S3. So yeah, we always trust S3. And secondly, it's about the elastic scaling. In Regional, if we can achieve a second level elastic scaling, but for the other systems like Flink, probably, yeah. Whenever you want to do some elastic scaling, then it probably takes hours.
**Benjamin - 00:11:29:**
Which isn't very elastic if it takes hours.
**YingJun Wu - 00:11:33:**
Yeah. So they actually have dynamic scaling to be honest. Well, I mean, I don't really want to hide anything, but they do have dynamic scaling. But the issue here is that people do not really use that swing adoption. Because for the people who think that it's not good for that, which is true. But in our case, well, people do dynamic scaling all the time. Yeah.
**Benjamin - 00:11:54:**
Gotcha. So RisingWave itself is open source, right? Kind of an Apache license. At the same time, you also have a commercial offering. Take us through that because it's such a common model nowadays, right? Like you have kind of ClickHouse doing something similar. A lot of other players in the space doing something similar. What is free? What is paid? When would I become a cloud customer, basically?
**YingJun Wu - 00:12:15:**
There are several things. I mean, we have the cloud offering, we have BLC bringing on cloud, and we also have on-prem. So on-prem, people say auto-powering business. Well, I mean, but we have to say that so like, if you want to enter banks, all these kind of highly regulated industries, you actually have to be on-prem. We actually enter banks. We also enter the aerospace. Nice. Yeah, I mean, this industry is what I mean.
**Benjamin - 00:12:41:**
Is RisingWave running in space? One of our questions, we had Hannes Mühleisen on who said DuckDB is running in space. So is RisingWave running in space?
**YingJun Wu - 00:12:50:**
I mean, it's definitely not in a satellite, but I actually don't know where they deployed.
**Eldad - 00:12:54:**
It might, it might. For all we know, it might. Tell me something. Who would be the typical user? Where do they come from? They decide to build a new streaming system? Are they coming from a Snowflake? Are they coming from a Kafka, Redpanda? Do they come from Flink? Like, how would you categorize that person?
**YingJun Wu - 00:13:14:**
So we definitely do not really compete with Kafka. We do not really compete with Redpanda. We have super good friends with Redpanda. They're awesome. We're definitely a good friend with WarpStream. But we do compete with Flink, directly compete with Flink. And so basically, there are two cases. The first one is that, okay, people probably are new to stream processing and they want to adopt a new system to process their streaming data. And then they do an evaluation with us. Well, they probably always start with the easiest solution. And we are essentially the easiest solution. Why? Because we are Postgres compatible. That's it. We are Postgres compatible. Challenge me. I mean, you can go with Flink, but I mean, you learn the SQL. But in all case, Postgres.
**Benjamin - 00:13:56:**
Don't have a hard tough crowd here. It's like since Firebolt is Postgres compliant, we're all in on your vision of the world. Like everything should just be Postgres compliant.
**YingJun Wu - 00:14:06:**
Definitely. I think Postgres definitely helped us a lot. Well, I mean, people always love Postgres and there's no such kind of headache learning something new. So basically people are new to stream processing system and they want to adopt a new system. And typically, RisingWave is probably the first choice. And second option is as well, okay, they have already adopted, well, the Spark streaming on Flink. And then they want to, they feel some pain and they want to migrate. So a very reasonable case, which I'm super proud of is as well, I cannot shout the company name because we signed on NDA, but it's a super big company. And they actually run a certain system, either Spark streaming or Flink. I could not tell you which one because there are some issues. And they actually run a system with 20k CPUs, 20,000 CPUs. Wow. And then they use RisingWave and they only use 600 CPUs.
**Benjamin - 00:15:02:**
Wow, that's massive.
**YingJun Wu - 00:15:05:**
I have my PhD and I always challenge such kind of number, right? Well, how you can reduce the, I mean, the number of C2 from 20,000 to 600, how that's possible. So I also dot that for it, but I actually talked to them. So the reason here, that's for their like state management, where they do joins, multi-way joins. So for all the other systems that are not scalable, but for in writing wave is super scalable. And we did a ton of optimization. So that's basically migration case. People feel the pain. They feel that's okay. They spend too many resources in the other systems. And sometimes what people say, okay, I used to use Java, but well, it's so too slow. I mean, in terms of development, I want to shift to SQL-based systems. Or sometimes people say, okay, I don't really want to spend four hours waiting for the Flink to get recovered from failures. Well, that was from a bank. And yeah, they got some issues with failure recovery and they decided to pivot and switch to some other systems from Flink. And that's all these kind of opportunities we got. Yeah.
**Benjamin - 00:16:13:**
Nice. That's very cool. So switching gears a bit, like my link and feed recently has been full of actually posts by you talking about Apache Iceberg. So you're very good at kind of amplifying content there. When I think of Iceberg, I mainly think about it in terms of batch processing, right? Because this is kind of, this is what we're building it for. Take us through how Iceberg is changing the game for streaming.
**YingJun Wu - 00:16:34:**
So there are two things I want to mention. The first one, Iceberg is Bash. I mean, they actually compete with Snowflake, Cresthaft, BigQuery, whatever. And in the past, we basically avoid a vendor lock-in because it's open table format. As long as you adopt Iceberg again, then probably you could use DuckDB to use Firebolt to use any other system to query the data from there, right?
**Eldad - 00:16:58:**
Didn't they say that on CSV5 as well?
**YingJun Wu - 00:17:00:**
Yeah. I was in data console yesterday. Yeah. I know that's where you guys were also there.
**Benjamin - 00:17:05:**
Yeah. Some of our folks from the Bay Area.
**YingJun Wu - 00:17:06:**
Yeah. Yeah. They actually didn't meet. And some other guys who I talked to. So basically, yeah, if you adopt Iceberg and there's no vendor lock-in, right? So you can use any query engine. So for stream processing, you asked me about, okay, well, what's the relationship? Well, so basically in the past, well, I mean, our customers were essentially, after processing, they essentially send the data into Snowflake, BigQuery or some other systems, Redshift. But nowadays they think that's okay. I don't really want to be vendor lock-in, right? Well, let's send the data into Iceberg. So essentially that's why Iceberg became, I think, our top three destinations, right? Because of here.
**Benjamin - 00:17:43:**
This is super interesting to me because when we think of Iceberg, we mainly think about it like as a source, right? It's okay. Like use a Snowflake managed table on Polaris Catalog and Snowflake or something. Can we be more efficient in terms of query processing on top of that? What you're saying is that like for a streaming system like RisingWave, the sync story is actually more interesting than the source story. So you're like the fabric that populates the Iceberg cable. Is that fair?
**YingJun Wu - 00:18:09:**
That's fair enough. Yeah. We do ingest the data from, also from Iceberg. We have the capability. By the way, I mean, most people, I mean, sync the data into Iceberg because they are free of and they're locking and they just want to adopt the Iceberg. We see this pattern pretty obviously in all these kinds of like enterprises. There are VP who will say, okay, like, I don't really want to pay Snowflake too much money. I just want to shift. What's my strategy? Okay, let's do Iceberg. That's the thing. That's always a conversation we've heard about. And yeah, basically the vendor lock-in is a big thing that's a big thing. From the stream processing system perspective, definitely we need to think about the destinations.
**Benjamin - 00:18:47:**
So, Staying with this destination thing, because I want to understand how a data pipeline like that looks, right? Like, would it be, I have this event stream, which is like, has like just bunch of data. I don't want to ingest like the raw data, for example, into my Iceberg Table, but use a streaming engine like RisingWave to compute me one second averages over some metric. And then I ingest that into Iceberg. Like, what is the actual data processing part that RisingWave usually does in these data pipelines, basically?
**YingJun Wu - 00:19:17:**
So like, well, there's always debate about ETL or ELT, right? So basically, ELT means that's okay, let's just ingest all the raw data into my data warehouse. And the ETL means that, okay, let's do some preprocessing before sending my data into data warehouse. I think in Iceberg world, there are typically people do ETL instead of ELT. Why? Okay, GDPR. The reason here, that's what, like, I mean, you don't really want to send out PII data, right, but private data into your lake. Because if you send it into lake, then who to delete it, right? Well, I mean, that's always a challenge. And Iceberg is super good at handling a lot of updates. So essentially, people actually need to delete the data or remove all this sensitive data before sending the data into Iceberg, right? In some other case, it will be like, okay, so people probably have data across the different regions. And they actually want to do some drawings and aggregations, unions. Well, before sending your data into Iceberg. And also definitely people have the requirement of doing enrichment. So let's say that's where I have the Clickstream and I want to also have the user info, right? Well, I want to enrich the data. Why people not just do it into Iceberg? I mean, you can definitely do that for using whatever system you want, right? Like Spark, whatever. But I think here that's where it's already like give people the impression that's where it's historical data. And I don't really want to do a lot of like ETL inside of lake. So people do the ETL before ingesting into lake. But I think what things will be changing because as long as people ingest more and more data into Iceberg, I believe that's where they will think, okay, let's just do ETL inside of lake. Yeah, I'm actually following the icebergs back quite closely. And in v3, they actually have a feature called Iceberg materialized view. And that's essentially for the ETL in Iceberg. And I think Verizon will probably become one of the first vendors that's a part of our Iceberg materialized view because we are essentially ready. And we are just waiting for the Iceberg of v3 got merged. Yeah.
**Benjamin - 00:21:25:**
Nice. Okay. That's really cool. In v3 spec, the MV definition is that like, I mean, one problem is like SQL dialect differences between systems. How do they handle that within like the Iceberg Specification? Like if I define an MV in like an Iceberg Table, what SQL dialect is that specified? And like, that seems like a nasty problem.
**YingJun Wu - 00:21:44:**
Look, in our case, it's just like Postgres, right? For we don't really care about all the other dialects, right? For we just care about Postgres. They might
**Eldad - 00:21:51:**
Actually focus on the schema definition, consistency, and kind of how materialist views commit versus on running on the SQL. I don't see any value in Iceberg to own the SQL, the business logic behind the MV. I can see how Iceberg can expand their metadata layer to support materialist views just from a materialize view object kind of entity aspect.
**YingJun Wu - 00:22:15:**
That's right.
**Eldad - 00:22:16:**
Let's not get confused. Iceberg is not a processing engineer. It's a metadata catalog that sits on top of Arcade. One of the things that people get confused with Iceberg is it opens up a lot of possibilities, but at the end of the day, they're still doing the same thing they've done before. But what does change with Iceberg is actually the last mile of many data pipelines. So if you think about it, if you were using Snowflake before, to get data out of Snowflake into a different system, you needed to talk to the Snowflake team. You had to extend the Snowflake data pipeline to now export into Parquet. Let's not get confused again. Nobody does low latency exports from Snowflake. People do batch every hour, every day, a few times a day. They export their data to Iceberg. They export their data to an S3 bucket for the same matter. The difference is now they just write the end of the data pipe to Snowflake. So they kind of avoid having an extra export, right? They remove that extra export. But that extra export is not just about cost. It's about enabling the data to move to new locations. If one team owns, if the team that owns Snowflake is kind of the barrier within the company on getting right data, correct quality data to more systems, this is a problem. But if the Snowflake team, just like any other team, when they're done with an EAP or a data pipe, it doesn't matter, batch, streaming, it doesn't matter. As long as you write it to an Iceberg Table, you save a hop on the next stage of having to export it. So instead of exporting it, you import it. So people pull data from Iceberg Table and that changes the whole dynamics. But the same old problem still exists. Parquet is not designed for necessarily for streaming format, to do on the meta low latency out of the box, just like it's not designed to do low latency reads. It's not designed for many workload types, right? Scale out is a challenge. So there is a compromise here. You can put Iceberg and get consistent metadata across systems. You can say, actually, I need multiple search formats within Iceberg. I actually expect Iceberg to grow in terms of format. Now that we have consistent metadata across systems, having a global format is less important now. If every system can read any format and everything is protected via Iceberg, then why not introduce extensions and improve formats so that actually workloads can run better through Iceberg? So Iceberg doesn't make workloads run faster. It makes different systems see the same data and it makes different systems lose their kind of ownership over exporting the data forward to the next system. And I think that's kind of where we'll see a lot of innovation. I've heard someone telling a joke, which is kind of becoming something that kind of floats in the industry that says, well, Iceberg was that intention of opening up and unifying metadata across all vendors, but it ended up just helping Databricks deal with all their own formats. So now Databricks will work for two or three years on Iceberg just to get their own data. And like the 20 catalogs that they already support and they already sell some service to, now they can get all of them under one Iceberg. But still, Databricks, Data Warehouse Engine, which doesn't have much to do with Iceberg, is the highest selling product for Databricks. I think for startups, it's a huge opportunity. I think Iceberg for Snowflake and Databricks is kind of basically a done game, right? The architecture is obvious. I do a dbt on Snowflake. I make sure it's getting written to an Iceberg so I can read it from other systems and vice versa. But the real innovation around Iceberg will come from startups and will come from the new data ecosystem, from new companies introducing productivity tools on top of Iceberg. There's a lot of discussion on how to merge data behind Iceberg, right? A streaming system will need different compaction than a system that maybe does post-applegation. Who owns that? Who owns the tablet behind Iceberg? So a lot of open questions, a lot of opportunities for kind of refiguring the ecosystem. And it's super exciting to see, I think, in this show that streaming databases are here to stay. In fact, they're getting so much stronger and through Iceberg are actually getting much more attention. A normal person would not think about going through the streaming journey without a PhD. That's obvious. Like, you shouldn't even mention your PhD. Like, everyone would be surprised if you wouldn't have one. But opening up to Postgres using S3, just figuring out consistency on streaming with that. Like, you need a PhD for that as well. As you said, everything is connected. So it's amazing to see that complex systems like streaming, PhD-grade systems, are getting the right treatment from the right startups. And I'm looking forward to see how RisingWave and Firebolt can coordinate with more startups around that ecosystem, around that Iceberg. Everyone wants to move out of Snowflake.
**YingJun Wu - 00:27:21:**
Yeah, definitely. Iceberg is essentially, people say it's called the Single Source of Truth. So, I mean, like we ingest all the data into one Single Source of Truth. So that's where you do not need to think about, okay, where I have my data copied. Do I have my data, copy my data in, let's say, Postgres or have another copy of my data in Snowflake or have a third copy of my data in some other systems, right? They just sort of think, okay, likewise, the Single Source of Truth. Iceberg is a Single Source of Truth. And once you've got a Single Source of Truth, you can have all these kind of like engines, right? Essentially, my hot take, probably not hot take, but for an obvious take, that is for all the databases will become a query engine on top of Iceberg.
**Benjamin - 00:28:07:**
I don't think it's such a hot take anymore. I think you're two years late.
**Eldad - 00:28:10:**
A Single Source of Truth justifies the source. The problem with data is not the fact that you move it around. It's not that you clone it and write it into a CPU register or RAM or SSD or a replica on S3. It's the fact that when it changes, there is no bookkeeping on that. So you need to backtrack the whole data pipeline, figure out in which step of the data change that source of truth changed. Now with Iceberg, you get a single, unified, consistent source of truth, and that can now tickle down anywhere. And any system that runs on Iceberg that supports transaction can support it. So you can time travel. You can do a lot of stuff you couldn't do without copying everything into a proprietary environment and then using it end to end. And having streaming on one end, just not one end, actually through the ELT as well sometimes, because sometimes ELT is streaming ELT. But then having that switch to a different engine, doing something else like a scale-out group buy, that becomes very interesting. And that monopoly over, you use everything, use a single stack by single vendor. The compute, the metadata, the workload, everything needs to be done by single vendor. That claim gets shaken and we'll see how that plays out.
**YingJun Wu - 00:29:25:**
Yeah. Like, well, I mean, Snowflake can just the data probably can read the data. I mean, yeah, it's great for, but sometimes people just have an impression like, well, I don't really want to get under locking. Right. And all problems with people are saying that, well, like, so I have some, let's say specialized data type and that, which can be only be read or probably write by certain systems. And in the past world, you probably need to build an integration between that kind of system with probably Snowflake. But no, we don't need to do that. We're just to build an integration with Iceberg. So it essentially simplifies the development for a lot of vendors, right? So, I mean, in the past, we probably need to think about that. I mean, we do have an integration with Snowflake. Well, yeah, essentially, I mean, if Iceberg came early, we don't need to do that.
**Eldad - 00:30:12:**
You don't need to talk to Snowflake. That's the beauty. If Snowflake rides to an Iceberg Table, there is no locking. The data is not locked behind the Snowflake wall. It is actually completely open to anyone that can scan or write or read from and to Iceberg.
**YingJun Wu - 00:30:28:**
Absolutely. Yeah. Yeah. In the past, it's like, okay, I need to think about the work. Okay. I'm in front of a vendor perspective. Okay. I need to think about the work. Okay. How to maintain the pipeline to Snowflake, how to maintain the pipeline to integration with Redshift, how to maintain the pipeline integration with Cray. Now that it's no, I don't think you're worried about that. That's right. I mean, even though we have already had that for, I mean, but these days for me, mainly focus on Iceberg.
**Benjamin - 00:30:53:**
Perfect. I think that's a beautiful summary to actually end our show today. Thank you so much for being on. I think you guys are running a bunch of meetups in the next couple of months and we'll be kind of participating in some of the European ones. So listeners, kind of take a look, kind of go to some RisingWave meetups. I think it's going to be a blast. It was great having you on the show.
**YingJun Wu - 00:31:12:**
Sure. Yeah. I mean, we do have our being Europe in May and June and hope to see you there. And I know that's where we have a lot of collaborations. And if people are interested in either RisingWave or Firebolt, just join us.
**Benjamin - 00:31:26:**
Lovely. It's going to be fun. Thank you for being on.
**YingJun Wu - 00:31:28:**
Thank you.
**Eldad - 00:31:30:**
The Data Engineering Show is brought to you by Firebolt, the cloud data warehouse for AI apps and low-latency analytics. Get your free credits and start your trial at Firebolt.io.
# How Similarweb Delivers Customer Facing Analytics Over 100s of TBs (/blog/how-similarweb-delivers-customer-facing-analytics-over-100s-of-tbs)
According to Yoav Shmaria, VP R\&D Platform at Similarweb, the best way to manage data warehouse costs is to tag every table, database or ETL running to have good granularity over every feature. Besides handy cost management tips, Yoav walks the bros through the tech stack he implemented to analyze 100s of TBs of web data to serve fast customer-facing analytics. Full disclosure, Similarweb is a Firebolt customer, but the bros kept it objective, and there's no Firebolt talk in this episode. Let's go!
Listen on [Spotify](https://open.spotify.com/episode/5wCvyjR5DgCGUHbR1wM2u1?si=spxiYNkcTcuXrmP1juT2Bg) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-similarweb-delivers-customer-facing-analytics-over/id1561927688?i=1000569868238)
**Guest**: Yoav Shmaria - VP R\&D, SaaS Platform - SimilarWeb
**Hosts**: The Data Bros, Eldad and Boaz Farkash, CEO and CPO at Firebolt
Boaz: Welcome to another episode of the Data Engineering Show.
Eldad: Welcome everyone.
Boaz: Eldad, How are you?
Eldad: I am great. Yeah, it is always good to visit and feel again.
Boaz: Yeah, you did not come over for a shopper's dinner last night.
Eldad: It is the only place we meet, on podcasts.
Boaz: With us today is Yoav Shmaria - VP R\&D, of the SaaS Platform at SimilarWeb. Yoav, how are you?
Yoav: Never better. How are you guys?
Boaz: Thanks for joining. You are from SimilarWeb. SimilarWeb, if you have not heard, it is an amazing data product. SimilarWeb, our market intelligence, and research platform were publicly traded as of last year. Yoav, tell us a little bit about your journey, to where you are at SimilarWeb or how did your career evolve?
Yoav: My personal journey in SimilarWeb?
Boaz: Yes. Even before. Because you have a nice mix of software engineering, data technologies, and now management of R\&D groups.
Yoav: Yes. I started my journey as an entrepreneur doing my studies. I had some touch in doing everything from everything, from full stack development to marketing and sales, and we had some specialty around payments and registration platforms, and I found my career journey starting in SimilarWeb as a Frontend Developer, nothing data. I have a very strong connection with product management, with business management, and I think, these capabilities took me to start managing, within the organization and eventually, two and a half years ago, I am already six years in SimilarWeb, I joined when my big daughter was six months old and now, she is six years old, so that is quite a journey. For the last two and a half years, I am leading the R\&D group of our main B2B product. We call it, pro internally. When I started my current role, we were a bit less than 30 employees. Today, we are over 80, 15 of them are located in Ukraine, the rest in Tel Aviv. I think that my biggest leap was when I started leading the entire group. So, I had to deep dive into the entire data ecosystem in our organization, which was amazing and still an amazing journey. In my group, I have two main positions, separated into backend engineers and frontend engineers. And what is special about our group in SimilarWeb is that the backend developers are doing some kind of full stack backend engineering. We call it a data server. It is a mix of data engineering, focusing on ingesting data into different databases and the backend layer for the API.
Boaz: Before we dive into that super interesting stack which we want you to dive deep into, tell us in your own words about what SimilarWeb does and the role data plays at SimilarWeb?
Yoav: Getting how to understand about SimilarWeb is that we are not giving you some insights above your data. We are giving insights about everyone else's data and the simplest way to describe it is Google analytics for the entire internet, which it is. We are giving analysis across almost any domain and also a mobile app. We divide our use cases into five separate solutions. So, we are starting with general research. Today, most of the marketing budget goes to digital, not to paperwork, and not to radio. The journey usually starts with research when we get some strategy decisions around how my market looks. I want to get into the website, build their market, who are the main players, in which countries they are, and how the audience looks like.
Boaz: If I am, for example, a digital marketer at Zava, a clothing company, what do I do with the SimilarWeb application?
Yoav: First of all, you would like to explore the fashion industry worldwide, maybe to get some trends around rising countries or some rising audiences to find new competitors, understand maybe the market is getting a lot of traffic from, let us say display ads. Let us say that on average, most of the fashion websites get 10% from display ads and I am getting only 2% for my traffic. So, maybe I am doing something wrong. Maybe, the audience for fashion loves banners. This is very, very high-level market insight. And, then, we are going into competitive research. We call it digital marketing to optimize for each channel for the SEO, PPC, maybe by an affiliate manager. These personas are doing tactical actions, like tracking search engine position, tracking the campaigns, understanding competitors, performance, and doing this benchmark with Zava.com against any other brand and understanding what are the differences in terms of traffic performance?
Boaz: Given this is such a data-centric product, let us go back again to your team. How do you call them the data backend team?
Yoav: Data server.
Boaz: Data server.
Yoav: In my group, we are doing everything. I mean, it is data engineering, backend, frontend, and QA, but I do not have a single backend developer, doing only API, coding, and data engineers. It is like a full stack data server engineer and we changed it four years ago. We had a separate group for data engineering and a web group for backend and frontend and we realized that especially for a product like us, it is super important that the person who is modeling the data and the person who is serving the data should be the same person and this is how we started to get data engineering stack for our backend developers.
Boaz: What titles do the data engineers in that group have? Do they call themselves software engineers or do they call themselves data engineers?
Yoav: In our HR system, it is a data server engineer, but most of them are on LinkedIn, I guess there are somewhere between backend engineers, software engineers.
Boaz: So it is a software engineer, building data-centric products.
Yoav: Yeah, we do separate the frontend from the backend, but the backend is not a trivial backend at all.
Boaz: That is super interesting. We love that. I think one of the courses came up, let us say 10-15 years ago. If you asked the software engineer to build a database or data warehouse or query engine kind of project, people used to look down on that as something beneath their payroll. Today, the most interesting software projects, the hottest thing in software is to build data, data-rich applications and it is a huge shift in the market.
Yoav: Yeah, I love to call it that in SimilarWeb we are doing data engineering upside down. In most of the organizations, I and probably you also, find data engineering is some offline processing. It is saved for BI and analytics. It is not part of the real production. Maybe we are hosting some database, a lot of data to serve something, but we do not have an entire data engineering operation for production. In SimilarWeb, it is the opposite. We are analyzing the entire internet, and we have this funnel to give real-time insights to our customers.
Boaz: Let us talk about the data stack. What runs there in the data server?
Yoav: Everything. I want to deep dive into the entire stack of our data collection methods and machine learning and everything, as I said, it is data engineering all over the place here. But in my group, we have shared the Data Lake but we are still on AWS, which is a very nice name for tons of files, we are using the Glue Catalog of AWS and very important detail for SimilarWeb. We are managing a branch system, a data versioning methodology for our data that is because we have multiple algorithms running to calculate our traffic estimations, and we are changing them from time to time. We are always running with some master brand for our data, and then, we are having in production a mechanism to show our customers a better branch. So look, we upgraded our estimations, take a look, we are validating it for some period of time and when we are deciding it is good enough, we need to make a decision whether we are starting a new breaking point with a new out algorithm or we are running back three months or three years of data. Very soon in the next quarter, we are going to run the entire three years of our mobile web data.
Eldad: What you are saying is the customers are moving away from software versions to data versions. To them, that is much more interesting to be able to play with versions and to be able to kind of understand the data that is driving the results, and that is also a very big shift from how engineering looks at the world. I wanted to ask someone with your background, what was the experience in trying to really reorganize an engineering organization that is more traditional, looks at data as something that is decoupled? How do you reorg from 20 engineers to 80, but also not just grow in size, but really rewire? How does that organization operate on data?
Boaz: How was the experience?
Yoav: I think something that helped them is that we are running, in my group, a totally very strong horizontal teams methodology. That means that we have a metrics organization, and there are a lot of methodologies in HR development. What we are doing is that I have a data server team and each team member is or a couple of them taking part in a different squad delivery and this helped me to keep knowledge in this team. So, we are not losing control but in every team, anyone and everybody doing whatever they understand. So we have a very solid layer of data engineering stack. Now, what we did change when I got into the role, I think we had, two ETLs running on a very premature Airflow and what we did is that we started to take, and that is a great lesson, two or three engineers every quarter from all around the squads and dealing with the infrastructure. I think our two main tasks that we took when I entered the world were first, creating solid infrastructure for an ETL abstraction so that every new engineer can pretty easily run a trivial ETL. We are not talking about a very complex stack, but let us say we are using DynamoDB a lot. If I want to just take a very simple collection from S3 buckets and write it to DynamoDB, it would be like two hours of work. And you have a running ETL supporting branching, everything, it was a big deal. We invested a lot in it. And the second was to find more in databases that we can sell as production because we Edgebase base a lot, which is amazing for concurrency and really, really, fast, querying like the data, but the structure is very key value-oriented. It is like you are keeping a very dummy data structure for a specific key. You can scale, but you cannot really run complex analyses over the data. So, these were the main two initiatives that we took with some kind of virtual team of infrastructure.
Boaz: What data volumes do you guys deal with?
Yoav: It depends. It is basically per feature. I think our largest database in production is around 150 terabytes, compressed, of course. This is for a part of our big data research product and in the digital marketing for the pure dataset. We also have a data set that is some dozens of terabytes, but it is crossing the trillion rows of data. Other than that, we have more than 100 ETLs of data ingestions running every day or month. We have, I think, over 30 or 40 tables in DynamoDB, and in general, I think, we are running over a petabyte in production and that is, of course, only what we sell, which is three years of data and we have more in our Data Lake.
Boaz: When you set out on the journey as you mentioned before on the versioning of the data, I think that is something that is becoming more interesting. How would you recommend it for people who want to go in that direction to take it on? Because it feels like there is no clear sort of market standard on how to go about it. People are trying to figure out if there are even startups and companies being built around that area like lakeFS, what is your take? What would you recommend?
Yoav: Wow! Take for data versioning. First of all, let us understand the challenge. Because theoretically, you can say, I have a new version, let us override. The problem is that you would like, first of all, sometimes you would like to show the customers in production like us two different versions. And, of course, before you release the production, you would like to have some staging product that you can test. So, you want to be able to walk seamlessly with those two versions, maybe more. We are doing a lot of infrastructure work right now on it, but I think the main area to plan is for the actual serving layer for the databases, because in Data Lake, as you say, like lakeFS, where you have a lot of methods to manage like data versions. But, eventually, let us take DynamoDB to the table. I have one table and now, I have another version. So, whether I want to manage a table version, and then if I want to take 12 months from one and then from another that is one kind of complexity. We totally use a lot, prefix or suffixes like the branch name. But you will need to design it in advance because otherwise, you will get into a situation that you want to write about, and again, I am talking about the serving itself. You want to write only part of the data, like for the new branch and you are willing to think about how they both merge together and in big data, if you want to do a union data or join, it is not always a trivial task and you can lose performance for that. So, I would say modeling the serving data, I think, that is the most complex part, because as you say, there are many, I think also Data Lake now has some solutions for the versions. But when we are talking about serving, that is not trivial at all. We are still struggling by the way, what is more relevant and when you are running some analytics engine on the data, it is even more complex.
Boaz: What do you mean by struggling? You told me everything was perfect.
Yoav: Nothing is perfect.
Eldad: The versions are related to the models, right? Different model versions require different data versions, which are translated into a feature version, and the challenges across every step. And you have mentioned up to the serving part but then it starts much earlier. This is why versioning in lakes is becoming such a hot topic. It is connected to ML, it is connected to models.
Yoav: And, of course, you need to hold some metadata store for your versions. Now, I have multiple versions all over the place, somebody needs to hold this information.
Boaz: In which case is what metadata store.
Yoav: So, we are using it, it is an in-house solution. It is a simple database that we are holding in relation to the database with some services that we hold. But again, think about it. It is something that is in your critical path, right? Because, unless you have some cash in place, you need to go through this service, always to ask, wait, where do I get the data from? This is your router. So that is also a critical part.
Eldad: Wired into the product versus just kind of looking at different versions to figure out which report.
Yoav: I always want to rewrite my entire data. I would rather only write parts of it. So, I want to hold the changes somewhere that I can serve seamlessly, let us say a perfect graph, but it is combined with five different versions of data.
Boaz: Let us go down that path, testing. So, how do you go about testing your data?
Yoav: Again, testing is all over the place as part of the modeling of cost, data rating, and everything. In our part, we have a very loud stack of automation testing. That is running on the product and seeing the final results. What is important again in our life is that we have a window release and snapshot list. We have a daily release of daily data and then we have a snapshot where all the monthly calculations are taking part. We have nightly tests, making sure that we have first of all a full sync between our data lake and our production databases, and that we are showing new data. It is supposed to happen that if you have, I know hundreds of data sets that somewhere will get a zero. So, we are always testing that we are not missing anything. This is like the critical test I am always talking about with my teams. Please, when you are releasing a feature with new data, make sure that next month we will see numbers. It is like the dumpiest test, but you always lose something. I think this is the main area. We have multiple tests in place for Linux, but I do not think it is the most relevant now.
Eldad: Testing is moving from units to data units and results.
Boaz: At the end of the day, you are delivering an experience to users, they slice and dice, they look at data from different angles, and users' experience is great and fast, but what are you guiding, I would say product principles. What do you think about the user experience? How slow is the query not good enough to be in the UI? How much does engineering need to be pushed to come up with ideas, to make things smooth from an experience perspective?
Yoav: There is general user experience science, and there is also the part of how easy to use the product. We are talking about basically time to insight in SimilarWeb. How many actions and how long does it take to get an insight from those huge data sets? I think in terms of performance, I believe that and Google has a lot of articles about it, that in some point, let us say around five to ten seconds, this is well, you are starting to lose it, especially now that we are so used for instant messaging and the user will wait for this report, but he would not play with the data and in one slice it over and over again and drill down if every action takes now 10 to 15 seconds. It will not happen and we see it in the data. We did several evolutions in the last year or so, taking the dataset from operational databases or starting a PLC with a low engine and then taking it with a more robust cluster or whatever the solution is, and we see the difference. We see more queries per user as we increase the performance. And, this is from the performance side and the other part is how do we take the most relevant piece of data and expose it to the customer as soon as possible. We have not cracked it perfectly yet, but we are having more and more steps into it and the onboarding experience will be that you define who you are, what is your market or competitors and we will give you the most interesting part of the data.
Boaz: Another question I had in mind. How do you guys hire engineers for the data server team, but the skills you require are not a trivial one. So what is the strategy there?
Yoav: So, in general, hiring is not easy these days. Hiring backend developers or data engineers is really tough, especially experienced ones. We are focusing on hiring smart people, good engineers with one of the aspects of our stack and we have an understanding that we will have to have this ramp up and learning curve for the other part. So, if we have a strong backend engineer, we want to believe that he will be able to pretty quickly take over our backend stack. And, then we will give him all he needs to know about data engineering and if we are getting a strong data engineer, we believe that he will be able to track with our backend stack. So, I do not think we ever hired someone that is fully stacked with both data engineering and backend development.
Boaz: So for that team, essentially, you are looking for both, you would get both data engineering-oriented people and backend engineers.
Yoav: I think I will tell my backend team that part of my problem is that once you are in SimilarWeb, we are like a superstar. You can do everything. So, it is a very, very desired material in the market. But yeah, our stack is not trivial and again, I think also, BI developments got more and more complex over the years, but they usually have a different mission because most of the processing is happening offline and we are taking care of a really big stack of the production data stack, which makes the challenge bigger.
Boaz: For all these years running a data stack in production, which were the most sensitive or error-prone, or risky areas that ended up causing on average more errors in production that you had helped for? Like, if you would go back in time, knowing what you know now, which errors would you have tackled to make for less incidents in production?
Yoav: You were asking like, what should we improve?
Boaz: I mean, most at the end of the day, the incidents that you guys happened that affected the experience, were they more the data versioning, the ETS were serving mismatch of Schema issues. I do not know. What were the most sensitive areas in retrospect?
Yoav: I think, above all-important to say we are running full. We have two regions that we are maintaining. We have a hot backup in two regions. Because that happens. We are managing a totally HPS cluster and things happen and we have a very, very structural methodology of shifting from one region to another and that is something that we know how to do and usually happens with machine errors, this happens other times. We realized that data validation, again between the Data Lake and the actual database is a very, very critical thing. Again, our customers are buying data from us. Saying that it is a very bad experience if you are querying data and getting one response, and the next day, we will get a different number. This is also something that we experienced. By the way, in various areas, it can be that we had some historical bug when loading data because we were writing in batches. We have a daily batch or monthly batch. So, writing to DynamoDB, we had some bug in the Spark client, that we realized that we are just losing rows on the way to Dynamo and it was not logged very well. So in that case, you have keys in the Data Lake that do not exist in production. It happened when we just loaded data that was not synced well. So the numbers, we did not have the same rows count or same aggregation count between the query engine in our Data Lake. So, we had many tests around these areas. And, eventually, you would be surprised, even UI can create some issues, because sometimes doing teamwork, so we say, "okay, it's just a simple average. I will do it in JavaScript." But then, the customer gets maybe a different experience because one developer wrote the code for the client, for the graph and one for the Excel in the backend and one for the API and, then you have some tiny number typos that usually will say who cares, but our customer cares. It is important to mention some of our customers, they have many data analysts that get our data in batches, from API and they are doing the math. They do not need the UI. So, in their world, if we have tiny biases all over the place, it is a big deal.
Boaz: Okay. What parts are of the stacks today do you sort of consider legacy and our plan to be phased out in the next year or so?
Yoav: I do not know we phase out. We are doing a lot of exploration around and changing the way we are maintaining the Edgebase cluster today. It is a technical death during the years. We used to have some very nice hacks in the old versions for how we prepare the data files to Edgebase and how we manage the clusters. So, this is something that we are working on strongly. Other than that, we need the ability to be able to sell as much as we can to sell from the same data source for features that are using the same data. Pre-calculation again, it is great in terms of performance, but the developer experience can be very bad. Imagine that now it happens a lot that you need to take one aggregation over an existing dataset. So, you have this dilemma. We calculated it again and, then you have hundreds of transitions all over the place, and it is really how to track or you want to execute the query on the fly and then the performance penalty is always the question.
Eldad: AAP, aggregation analysis paralysis.
Boaz: Awesome. What did we cover, by the way?
Eldad: Did we cover huge failures while I was away.
Boaz: Maybe you could say because we talked about a variety of things that have gone wrong but we can do that again. If there is one single sort of meltdown that you remember, that you would want to warn others to learn from your mistake and to not get into it again, what would that be throughout the years?
Yoav: Meltdown.
Eldad: I was having a quick question given that you are like, aggregating everything that is out there in terms of data. Can you give us or share with us how fast the universe of data is expanding? Meaning that kind of it is constantly growing. How much of a challenge is that? How fast is data growing, given that it is not per user, it is just massive, massive granular data points across so many industries? How fast is it growing?
Yoav: So you are saying if you are storing more and more data. I can tell that we have data sets where we are loading every month, let us say, two terabytes for seeing the database. So, we are seeing an increase also in coverage. So, we want to cover more and more from the internet. I believe we are not seeing everything in all the darknet and areas that do not exist, but we have more and more websites and, also in the internet area, I think that you have new websites every day. I think this is where the main challenge is because you do not have an absolute number of entities. The entities are growing. So, maybe the data is not always like exponential growth, but the number of entities is endless, keywords are something that all three of us want to buy genes, but we will search for different terms.
Eldad: I search for skinny.
Yoav: Yeah. I had Eldad for the skinny, but yeah. So, these data sets are going massively because the entities are being changed. Of course, you have the baseline, but by the way, we are exploring it a lot and we are seeing that keywords especially are changing on a daily basis. Think about our lives. So, today we have the Russia-Ukraine war. Ten years ago, you had something else, and today you have some fashion trend and tomorrow you will have some fashion trend and you have NFT today and then SpaceX and Tesla, and every day it is something else. So, the information is endless. And for that reason, you cannot assume that most of your data is the same. It is the opposite, most of your data will be changed.
Boaz: Cardinality is a nightmare.
Yoav: But, that reminds me, eventually we serve the entire internet, but our customers are not querying the entire internet. So in that matter, something that they want us to be able to do in the next couple of years is to find solutions, where we can optimize what our customers are more interested in and then to have a better, more cost effective solution to give a very robust solution for the data customers, are querying and having some cold storage for whatever just need to be there and can be queried some time, somewhere.
Boaz: Product cost we did not cover and it is interesting for a company like SimilarWeb.
Yoav: Yeah.
Boaz: How predictable is the cost these days. How tough is it to manage?
Yoav: Also from colleagues in SimilarWeb, I know it is tough. This was one of the first actions I took when I got into this role. I am really glad that I did it. Because we did some gross action again for the data server team. And now, we are tagging every Spark cluster, every database, every table, each ETL running. So, I have a very, very good granularity to know for a specific feature, the entire cost from the ETL and ingest and the data serving, the databases, the staging, the development. So, I have very good visibility. But, otherwise, it is impossible. I mean, I think the number of roles and tags that I have today in our stack is something that you cannot manage if you do not have a standard.
Eldad: You need a data warehouse to manage your tags.
Yoav: That is a very good tip. I think, as soon as you do that, you can control it. Again, shit happens, but at least you can know in real-time what happens. It has happened to me every week. I think there is no way that I am not slacking on one of our engineers and what the hell is that? Why is that on? I see some anomaly in a specific table for a specific team. So that is very, very important. Because otherwise, you do not know what happened. You just know that you have a bias of $10,000, now go figure it out. So, imagine your tags.
Eldad: Keep your tags close.
Yoav: Yeah.
Boaz: Yeah close, as close as can be. This is awesome, Yoav. Super inspiring. It is rare to see a product that is so deeply driven by data. And, I have personally played with SimilarWeb products and it is absolutely amazing. The levels of experience you get there, the breadth of analysis, coupled with such a great user experience instead of speed and insight is something definitely to learn from. So thank you for joining.
Yoav: Thank you, Boaz.
Boaz: And see you around.
Eldad: See you around. Thank you, Yoav.
Yoav: Thank you, guys. Take care.
# Making Observability a Key Business Driver (/blog/how-statsig-engineers-do-observability-right)
80% of the code that you write doesn't work on the first try. And that's fine. But knowing which 80% is not working and which 20% is working is the actual challenge. After 10 years at Facebook, managing and scaling the Seattle site to over 6000 engineers(!) Vijaye Raji founded Statsig to make observability automated and real-time. How is the semantic layer managed? How was the Statsig team able to build an observability product that handles real-time ever-changing metadata? What are Vijaye's main takeaways from engineering at Facebook? Tune in.
Listen on [Spotify](https://open.spotify.com/episode/5hcGDxIPtZ7VmEZposgb1z) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-statsig-engineers-are-doing-observability-right/id1561927688?i=1000587915342)
Eldad: Just before we start, I don't know if you've noticed, I've adopted a new brother, his name is Benjamin, as you can see here. I tried to get the hair close to the family's DNA, but it didn't work very well.
I'm joining Benjamin. Benjamin will run Season 2. He is also our query processing tech lead based in Munich, building and scaling the office there. I've met Benjamin 2-3 now, how long ago Benji?
Benjamin: Probably two.
Eldad: Two years ago, we got in touch. He was an intern, thinking about his PhD in Munich University and this one thing led to another. And here we are today with you, talking about data. So I'm really looking forward and with that, Benji, it's all yours.
Benjamin: Awesome! So, you're introducing me and I'll just go on, and introduce our guest, Vijaye.
Thank you so much for joining us today. Vijaye is the CEO and founder of Statsig, and he founded Statsig in February 2021, if I'm not mistaken. They're based in Bellevue, Washington, and their original office was in Kirkland, where Firebolt also has an engineering office and Statsig is building a product observability platform. And, we'll talk more about it later. Basically, it makes it super easy to manage feature flags and experimentation and yeah, we're excited to have you. So, they raised a series A in August 2021, and now in April of this year, 42 million Series B, led by Sequoia and Madrona. So, that's awesome.
Vijay didn't start out, as a startup founder, he started out in Big Tech actually. So he spent his first 10 years, at Microsoft doing a variety of different things, and then joined Facebook in 2011 where he stayed for another decade and there he led Facebook video and gaming, and was also the site lead for Facebook, Seattle. So, he scaled the site there from, just a handful to more than 5,000 engineers, which is pretty impressive.
Eldad: So, one out of those 5,000 used to work for you. He works for Firebolt now, and we got great, great feedback. So, we are happy and excited.
Vijaye: I love it. I love it.
Benjamin: Awesome. Perfect! So Vijaye, do you want to give a two minute high-level overview about yourself, about what Statsig is doing, before we dive into some of the nitty-gritty details.
Vijay: Absolutely! Hey guys! Thanks for having me. I'm honoured to be here on the first episode of Season 2. Excited to be here and chatting with you both. So, Benjamin, you covered the history a little bit. I started Statsig about 20 months ago, and this is after spending 10 years at Facebook. So, I joined Facebook in 2011 when Facebook was still considered a small startup. And then, obviously the company grew, both globally as well as in Seattle. So I spent, my time also growing the Seattle offices, from you said, a handful of folks to I think when I left about 6,500 people.
One of the things that struck me as I moved from Microsoft to Facebook, I was always an engineer, so I grew up and learned programming and wanted to be an engineer and I spent 10 years at Microsoft as an engineer. And, during these 10 years, it's formative when you see the software processes evolve.
And it went from the shrink box software that sits on Best Buy shelves to shipping every single day and then eventually continuously shipping code. Those are the things that I learned first hand at Facebook. And I was fascinated by that whole concept of how you decouple feature launches from code launches, and the tools that were associated with it, the tools that kept this chaos from getting out of control.
And also, I started understanding how you measure the impact of every single feature you ship. And then, throw away code that doesn't work because just because you thought it would work, it doesn't mean it would and then measuring every single change actually gives you a lot of insights, a lot of humility, I think, and then also a lot of cultural changes. And so, those definitely resonated very strongly to the point where I said, Okay, this is important, that we should go build a company around these sets of tools. And that's why in February 2021, I left Facebook and started Statsig. I brought with me a few folks who joined and built the tools and then now the company is going strong. We have 52 people as of yesterday. So, pretty exciting growth.
Benjamin: Awesome! Yeah, that's super cool! So, I mean...
Eldad: How was it switching from this huge, huge organizations, engineering organization to a startup that just now got to 52 people. How does it feel?
Vijay: It's crazy. Going into a startup, I thought I knew some things, but then over time I realized I had absolutely no idea what I was doing. It's really jumping off a cliff without a parachute, but just instructions to assemble a parachute and then you have to figure it all out before you land. And that's kind of like how it feels.
Managing an org of thousands of folks is a very different skill than building a startup from scratch, hiring the first person, the first engineer, the first product manager, first designer and then getting a product out. It's kind of like when you're building in the early days, you're building it in a vacuum. You have this vision, you believe in this thing, and then you go build and then you put it out. You're really vulnerable, right? So you build out your product and you put it out for somebody to come and take a look at it and you hope somebody would like it, and then, will use it, give some feedback, and then for six months you sit there and nobody is touching your product and you're wondering, Oh, what have I done, I've made the worst decision of my life. But then slowly, the traction happens, slowly people start to notice and then they are like, Oh, this is actually a pretty useful product. And, then you go from this one stage to another stage where people are now even trying out your product for free. And then eventually they want to actually use it and then eventually they want to actually pay for the software, and then eventually they love the product and then they actually want to evangelize the product. So you go through these stages, and each of these stages you're constantly questioning yourself and questioning your decisions. It is a fascinating journey that all startup founders go through. And, this is my, like you said, I spend all my life in large companies where you don't have to worry about this kind of stuff, and then when you throw it into a startup, every little thing matters.
It's been a fun journey, very much a learning journey, very much humbling, just to look around and see how many people that have gone through the same journey as me before and been successful.
Eldad: Nice! And you just started, there's a long journey ahead. Yeah, it is.
Vijaye: I know, definitely. And it's only been 20 months.
Eldad: You picked great timing as well, because I think of raising round be that perfect. Perfect, perfect, perfect! Those are crazy times for building startups. It is always crazy building startups but those times are especially crazy.
Benjamin is young, so everything is new, but we've been there through previous recessions and it's always surprising how different each time it is.
Vijaye: Yeah.
Eldad: But I wouldn't replace it for anything. And, really, we look forward to hear more about what you're building.
Vijaye: Yeah. Thanks.
Benjamin: Awesome. So let's dive right into that. So maybe for our listeners who weren't exposed to, I don't know, A/B testing, experimentation, all those things yet, do you want to give the two minute pitch of why this matters and what it actually is?
Vijaye: Yeah. One of the most important things is when you're building a product, you have a hypothesis and you think that's something that you're building is going to be beneficial for your users or your customers or for your business. And that hypothesis is baked into every single feature that you're building. The idea that we build those things out and then ship it to people without really instrumenting or measuring whether there is that particular belief is actually true or not, is no longer valid. It's no longer okay.
And, at Facebook, I remember this, where 80% of the code that you write or the products or the features that you build, don't work on first drive, which means that you have to go back to the drawing board, iterate on the product until it works or sometimes throw away your idea and it's okay because you can't be right all the time.
But knowing which 80% is not working, which 20% is working is the hard part. So, just understanding and measuring the impact of the changes that you're making to your product, to your users, to your customers is extremely important.
The stage number one is always about measuring, measuring things, measuring things that you build and ship to people.
Stage number two is taking the measurement back and then actually addressing some of these causal inferences. Why did the metrics do what they're doing? Can we attribute it to the changes that we made? And then make the decision, Okay, well what do we need to do in order to... whether double down on it because it's doing really well or wind it down because it's not doing well.
Those are the kinds of tools that people need access to in order to be able to make the right decisions on the product on a daily basis. The state of the art in the industry currently today about experimentation is still A/B testing. And A/B testing is the gold standard for causal inference. However, it's very, very time consuming and manually there's a lot of overheads because you have to first come up with a hypothesis, you'd have to build a variance and then ship out the variance and then you have to allocate the samples, isolate the experiment, run it, for however long it takes for statistical significance, maybe two weeks, maybe three weeks. And then you have your data science team go back and look at and analyze all of these results.
This whole process takes a long time. Most people don't run that many experiments. They pick some of these features that are really, really important, and then they run it through the A/B testing, experimentation pipeline, rest of the features just go out untested.
So, this is where I think companies Facebook, Uber, Airbnb, and even LinkedIn, some of these companies have sophisticated tools that basically take every single code change you're making, feature change you're making, and then have tools that'll automatically run these A/B testing. And then will give you back the results of the impact of those changes. Thereby you don't have to do the previously mentioned long process of A/B testing. You get these results on a daily basis and then you make decisions, which will be much better decisions than before. And that's what we are trying to build with Statsig, it's where we're going with this product observability.
Benjamin: Gotcha. I mean, you have a background in gaming at Facebook, for example. Do you have one or two grippy examples of things where this mattered or where these types of experiments led to interesting insights?
Vijaye: A lot of them. I say, one of the key metrics that I share is 80% of the code or 80% of the features that we think are going to be beneficial are not. And so just going back to the drawing board is a regular thing and it's okay. And, that way people don't get attached to the code they write.
So, I'll give you one anecdote. In video specifically, we have run plenty of experiments, and the general belief is that the higher the quality of video that you're able to provide, the more consumption that'll happen because obviously people like higher quality videos and so, at Facebook, we built, this automatic bitrate detection code, which basically constantly analyzes your network stability, last mile, bandwidth and all of that stuff to be able to, Okay, well we're going to give you the best possible bitrate for the connection that you have and so far, that's been experimented and proven and we continue to increase and increase and increase, and coding have gone up from a 480p to 720p to 1080p and then eventually, 4k, and the belief is, yes, that is always good.
And then, in 2020, the pandemic hit and when most people stayed home, the media consumption grew pretty heavily, almost over a course of a month, it doubled and tripled. And some of the backbones, the carriers and the ISPs, they couldn't basically handle the load, and there some of these folks reached out to Facebook and, Hey, could you help? Because you know, a lot of the traffic that is going out is going to Facebook and a lot of it is for video, could you throttle these for us? And we were, Oh my goodness. If we start throttling, that's actually going to reduce the consumption. That's actually going to leave people not feeling great about the quality and the experience, but we also understand that this is not normal terms. It's crazy, crazy times that we're dealing with. So, we went ahead and started tweaking our ABR algorithm to actually give a couple of clicks lower bit rate than what was actually possible to give. What we found...
Eldad: I always knew this 4K wasn't real 4K when they saw it. Something felt fishy there.
Benjamin: Yeah.
Vijaye: No, we actually showed that we are dropping and then if people wanted to pick 4k, they could go ahead and pick the higher bandwidth. It's just the automatic default that we pick.
When we change the default, normally, obviously, we experiment with everything, this is one of those things when we changed the defaults, we were actually looking at the metrics that were coming in. We were shocked because the usage actually went up. People were using video a lot more and consuming video a lot more with the lower bitrate. Obviously, this was surprising, this was new. It's , okay, well we would've only known because we ran these experiments for every single change we make. Then, we went back and we looked and we did some case studies and we talked to folks, and then later on, we understood that there's a segment of population that is very bandwidth sensitive, and whether they're throttled or whether they're limited bandwidth and so on, generally, they get constrained by how much per day that they can use.
Now all of a sudden, a whole batch of people were now able to use or consume more video than they did before. Then, there's a whole segment of the population that were on the phone that were old outdated phones where 4K or 1080p doesn't actually make a difference. And so, trying to push more and more bitrate was actually not necessary. So there's a set of learning that we had that was very interesting. We wouldn't have had those learnings if not for this experimentation, constant experimentation. Anyways, it is pretty eye opening for me as I was running the video org.
Benjamin: That's pretty interesting, because when I looked at this first and I was, oh, maybe this is about, I don't know, the color of you are pay now button on the shopping cart, but these are really deep algorithmic, core parts of your infrastructure to which you then apply those experiments. So that's cool to see.
Vijaye: Yeah, that's absolutely right. I mean, a lot of people, when they think about A/B testing, you're naturally gravitating towards, Oh, well let's change the text of the button or the color of the button, or change the layout of where things go in order to get more conversion. I think those are also valid A/B tests, but they're more marketing centric and more shallow because your metrics that you're trying to move are directly correlated to the product changes. Whereas when you touch, Okay, well I'm going to change the parsing library, that's in the backend, and then, I want to see if that increases any latency on the large segment population, especially at scale.
Those are the kinds of things that you actually care about when you're modifying a product and you want to verify that it's actually doing what you're expected to do. And then secondarily, you want to verify that there are no extreme side effects. Sometimes you don't know what's going to happen when you actually make these changes. So monitoring those things are important.
Benjamin: Right. So let's talk a bit about what powers actually infrastructure this? So the actual experimentation infrastructure.
The first question I would have is, say I want to run this experiment now, right? I'm an IC at a company, I have this type of experimentation software, and now I'm saying, okay, my video and coding algorithm, I want to change that and see if XYZ increases. Do I have to be the one formulating the hypothesis? Do I have to be the one saying, I changed things and I think XYZ might increase? Or is it there's some core metrics, someone else from the product team defined, and then each experiment is actually tested against these same core metrics? How does it work?
Vijaye: Generally, what you want to be keeping track of is there are some business critical metrics. You don't want to drop those business critical metrics. You want to make sure that they're healthy across all of the product changes that you are making. And then there are all these metrics that are the primary metrics that you expect to change, expect to move. Okay, well, if I'm changing my video and coding algorithm, I expect more video watches to happen or more time spent on videos to happen, and then perhaps even lesser lag or lesser interruptions in the video because of bandwidth issues. So those are the kinds of things that you would normally be watching as a direct result of the change you're making.
I think those are two separate sets of metrics that you should be watching. And one of them is probably picked by your company level product or growth leader. You're kind of like, okay, well I don't want to ever drop my DAU or engagement or retention, or obviously, revenue and things like that.
And then the other one is generally determined by the product team that is making product changes very close to those metrics. I think both are extremely important and then generally, more sophisticated tools Statsig. So you can basically specify if this particular metric, say, my DAU metric drops by below 3%, then I want an alert to fire and people to be notified. And so at any time anybody is changing any product you, if that drops the metric by 3%, you want to be aware of that and you want to be able to get right to the problem of what's happening?
Obviously you can also have these trade off conversations whether it's worth it or not. But that's up to you.
Benjamin: What actually powers these things in the background. There must be some pretty sophisticated statistics or something going on and one thing I just thought about, what about correlation and experiments, for example. I, on my team, am changing the video encoding, another team is changing how much video gets buffered and in the end, we are dropping this into production at the same time, somehow you have to figure out which change actually changes the core metrics you're interested in.
Vijaye: Yeah, this is a very good question. So obviously there's a lot of teams that are changing a lot of things. One of the things obviously you would not launch every product that you're making 100% to the user or splitting the audience into smaller portions and rolling it out.
Typically, we use this exponential rollout model where you take a new feature, you roll it out to 1% of people and then see how it affects the metrics, and then 2%, 5%, 10%, and then 20%, 50% and 100%. So that's the sequence that you normally follow through and making sure that at every stage, you're not dropping any critical metric more than necessary.
When multiple people are doing those, salt or the randomization for each of these rollouts is different. And so different sets of population will get these different sets of features. There will be some intersection and the rest of them are getting their unique experiences and that's how our system determines how to attribute metric change back to the actual feature change. And that's why we also have error bars in our confidence intervals. And then you'll be able to once the error bars are within statistical significance, and then we turn them into either green or red, depending on if it's positive or negative. And so that's how our stats engine determines the causal inference back to the features.
Now, obviously, there are lots of things that you could, also get, employ the tools to validate some of these. So, if you have multiple features that are highly interactive in nature, so, okay, I'm changing something in the product that are very closely tied to each other, then what you do is you have isolated experiments. So you create a layer, which is called the universe in other words. And then you actually say, I'm going to allocate 25% of this universe or layer to this particular experiment and another 25% to this experiment. And so you go like that and that way no single user will be getting two different experiences. They'll actually be isolated and getting one particular experiment that they work on. That way you can also make the confidence interval better, just decreasing the confidence interval and at the same time, also attribute it really clearly, Okay, well this experiment is doing these things to my metric. Those are the kinds of tools that are available.
Now, there's also something that you mentioned, which is, okay, what if I launch these two things at the same time? What happens? The interaction affects the cumulative impact of multiple sets of features. So, we have something called holdouts. So long term holdouts are able to actually measure the cumulative impact of multiple sets of features that you're launching over a course of time. We use this extensively. So basically what that means is the beginning of a quarter or the beginning of a half, you specify 1% or 2% of the population that are in the holdout group, and then you launch a whole set of features and then you go back and analyze, compare all of the metrics to these 1% or 2% of the people that have been held out from these features. That gives you the cumulative impact.
So you have all of these tools at your disposal to get all of that information.
Benjamin: All right, super interesting! Sorry, Go ahead, Eldad.
Eldad: How do you manage this semantic layer of all of that? Because it gets so complicated. I wouldn't consider myself an engineer anymore, unfortunately, but observability to me always seemed that it's easy to collect the events. But it's impossible to connect those events to a semantical model where you can actually drive the business based on that and having so many cloud native companies driving their business through their data, through their product, it seems those platforms are really not just about observability anymore, right? It's really about driving your business. Is there a standardized way to do it? In the industry, is there a shift towards a semantic, like a DBT for observability or something that will make an engineer's life easier to communicate on the inside?
Vijaye: Yeah, that's basically what we're trying to build, which is, you know, the sophistication is actually in the stat's engine and that should not make your product building process any more complicated than it already is. And so the idea behind everything that we're building is engineers should still be building features and when you decorate a feature with a feature flag, we then take care of attribution, the analysis and then also try to do it as real time as possible because you want to get to diagnostics as quickly as possible. You do not want to put out something that is broken experience for your users for longer than is necessary, and then, catch that as quickly as possible and then fix those things.
While it seems like, okay, there's lots of stuff that is happening, all of that is encapsulated by the complexity of the stats engine. But there is also a layer of like, okay, well if you make it complicated for people to understand what's really happening, then you've lost. You have not actually achieved anything.
So how do you take all of that data and simplify it? Some of the visualization innovation that we're doing is to like - How do you make engineers, product managers, even designers, be able to understand the insights easily. However, when a data scientist wants to come in and dig into the code, you want to have this progressive disclosure of complexity, want to actually have the ability to dive into the data if somebody is inclined to do so. So that's the challenge, right? The tooling should make it extremely easy for everyone, but at the same time not block you from getting into the details. That's precisely what we're building.
Product observability is like, if you think about it as an extension of, I always say, tools in the data observability space have become so sophisticated, and then if you look in the ops observability, you kind of like Data Dog or so, in real time, you'll know when one server is misbehaving among a forest of thousand servers or 10,000 servers, in real time, and that's pretty amazing. So systems have gotten so much more sophisticated, than what I remember, 10-20 years ago.
And then when you talk to product folks, you still are in the, Oh, yeah, we're going to launch V2, three weeks from now, and then when V2 launches we're going to wait for three weeks for our product adoption to happen and then we're going to have analytics or data scientists go and look at the analytics. And then have some way of correlating, well, are the metrics going up after the feature launch or the product launch or they are going down? Why are they doing those things? You see how slow and...
Eldad: And it's done manually outside of the system, so it takes your alpha feature eight months to go into beta while trying to make observability work.
Vijaye: Yeah.
Eldad: I think it's amazing. I think the next step is really getting observability to be impactful.
Vijaye: Yeah.
Eldad: It's always been a black box for engineers to do their stuff and seeing observability moving away from just helping engineers do their day job to really driving your business. I think it's a huge step forward, and it's also if you think of engineering, the more data they deal with, the more data that's being impactful, using those kinds of tools becomes not just optional anymore. It's really changing how you build products, and we've seen that firsthand at Firebolt, by the way, how it affects. This is not a hacking, it starts at, okay, let's put a feature flag, but it quickly turns into kind of a way to drive your engineering culture. As you've mentioned at the beginning. Amazing stuff!
Now tell us how complicated these are... like building observability is the hardest part, right? I mean, it's data driven, it's real time, metadata is actually changing unlike your product, which makes it even more challenging. Can you share a bit with us, how does it work? How do you do it?
Vijaye: Obviously, the credit goes to the data science and the infrastructure engineers that we have at Statsig. These guys are pretty amazing. I'll give you the overview of how we ingest the data and the infrastructure that we rely on and then the details of that is actually beyond me.
So, let's talk about… We have a set of SDKs that help you get started relatively quickly. Now, obviously building these SDKs is also a pretty interesting problem because as an engineer you're excited, okay, well there's all these technologies. It's like iOS and Android and React, and React Native. And so you got to build SDKs in every single technology because people have all of those out there. And this matrix continues to grow, especially if you have a new feature, you have to test it across. We have 25 SDKs, I think. So, 25 SDKs...
Eldad: The ecosystem team is going crazy.
Vijaye: Oh, Yeah! We have this on this whiteboard, I'm not kidding. There's this huge matrix where people draw, Okay, well have I implemented this feature in this SDK? And then, the matrix continues to grow. So, that's fascinating. That is step number one.
And then once the SDKs are integrated, they're obviously collecting event data and then sending it back to our servers. On the server, we have a queuing system called Event Hub. So basically the Event Hub is the first stage that actually acquires all of these events and cues them, so we don't lose a sequence. And then we actually have a redundant system as well. So once the Event Hub is receiving all of these things, they actually throw it into a file system, a big large file system so that if anything happens to the Event Hub, we can go back to the data that's sitting on the file system and replay it when necessary.
Beyond that it goes all of this...
Benjamin: Vijaye, if I can interrupt, can we take a step back maybe? What is an event actually in this world?
Vijaye: Yeah. No, absolutely. So, an event could be as granular as, somebody purchased a product in an e-commerce setup or someone, added to cart again, in an e-commerce setup, in a gaming setup, somebody reached a new level or achieved an accomplishment or achievement, and then in a FinTech scenario, someone used a credit card, and so those are all events and those are the ones that we be able to gather and turn them into metrics, really. Imagine if you have a purchase event and that purchase event happens to have a value, which is actually the dollar amount, then you can aggregate a sum of all that stuff, and then you can see revenue as a metric. And now you can verify, okay, if I'm making all of these product changes, is it changing my revenue up or down? And that's how we correlate it.
Benjamin: These events are then basically annotated with the specific feature flag configuration of that customer or user at that time?
Vijaye: No. The events only have user ID information and then on the other side where when we actually deliver the features you ask, for this user id, which features should we show?
Benjamin: Ah, got you.
Vijaye: So that's the correlation. The join is actually...
Benjamin: Join then in the background. Okay.
Vijaye: Yeah. That's why...
Benjamin: I love joins as someone building database systems.
Vijaye: Oh, I love it too because that actually..
Eldad: There's always a join Benjamin.
Benjamin: Always a join. Wow!
Eldad: Always.
Vijaye: Everywhere. Join is actually powerful because it simplifies the two responsibilities. So if I'm a feature builder, if I'm an engineer, I only have to care about, okay, which feature flag am I using? I don't have to worry about all the metrics and annotating the metrics and we do that for you. And so it's part of actually simplifying how we use the system.
So, once the events are in our backend, then we throw it into a data lake. And so the data lake is where we do deduping because sometimes when users are, especially in React setups, constant rendering, you have multiple events generated at the same time, and so you want to de-dupe, do some sanitation of the data and stuff that and so we do all of that on the data lake side of things.
Then what happens every hour, we have this gigantic data bricks job. This data bricks job is where all of the statistical correlation, causation, all of the analysis is all encoded, in Python. Jobs pick back up and then they basically spin. Can you guys hear me?
Benjamin: All good.
Vijaye: Yeah, they basically spin off thousands of servers to go orchestrate all the analysis and then outcome is basically a set of analytics. Those analytics go back to our data warehouse and then, where we actually break it down into lots and lots of pre-computed metrics, so that when you actually go to our console, you'll be able to get these things in real time. So you don't have to recompute anything. The UI is extremely responsive. You can click around and actually have things pop up very fast because that experience is important. Otherwise, you're sitting for even one minute for every one of these queries to run, then you lose the whole fluidity.
We take a lot of care about making sure that those things are working as quickly as possible.
Eldad: Nice.
Benjamin: Nice. It's not about querying anymore. It's just a feature. The fact that there is a query running in the background, users don't care.
Benjamin: All good. Cool. You have these hourly batch processes running, which means, say, I launched something, this experiment starts. After an hour or two hours, I'll start seeing some data on my experiment. What if something goes terribly wrong? Did the worst thing ever and it just breaks everything? Is the safety then that it's only with 1% of my customers and so these couple of hours don't hurt a lot, whereas to a fail safe or crap door if things go terribly,
Vijaye: Yeah.
Eldad: Benjamin, you want them to fix your bugs as well? I mean, come on.
Benjamin: I'm sorry.
Eldad: Don't stretch it. Don't stretch it too far. Sorry, go ahead.
Benjamin: No. I think that's important. I think you, you're hit in the nail because, it's important to roll out slowly in stages. So typically we see most customers when they create a feature flag, it's open to only employees, and so there's usually this dog footing process. If something is terribly wrong, you catch it with just employees before it even hits your customers or users. And then what you do is you roll it out to sometimes some customers have these early adopters that are a little bit more okay to get some beta experiences, and they're okay to tolerate some of these things that are broken. And so you turn it on to the next ring and which will be the early adopters. And then obviously you are monitoring all these metrics. Sometimes you get a little bit of scale, which is good. So it's stress testing your product and so if something is broken, you pause the rollout and you go back and fix it.
Then the next stage is, okay, well now I'm going to open it up to 1% of the general public or you can pick a country. Okay, I'm going to open it up only to the United States. Sometimes people do that because they don't want to invest in localization, and they want to test it in English-speaking countries first before and if everything works, then you invest in localization and so you roll it up to 1% and see how things are operating.
The good thing about this is if something is broken, you catch it within an hour. So, which is as real time as we can get. And so then you are going back and fixing things versus waiting and waiting for your customers to tell you or some support channel file a ticket, things like that. So, those are, I think, why it's extremely important to catch these early and fix them.
One interesting thing is because I'll add this here too. A lot of people think that, oh, you need lots of samples in order to even catch these things. But if something is broken, you don't need thousands of samples. You need 10 samples to tell you, Oh my gosh, this particular metric is dropping, so let's go figure out what's going on.
Benjamin: What volumes of data are we actually talking about here? How much data are you processing every day?
Vijaye: Oh my gosh, terabytes. We have about, 15 to 20 billion events every single day, give or take, based on how one of the large customers ran a sale for three days. And then what happened was their volume tripled for a week and our system is able to scale and take care of that. But give or take 15 to 20 billion and it's been growing. Just eight months ago, we had a million or so events a day, and now we're talking 20 billion events. Those things are raw events, right? And so the job that we need to do is to quickly dedupe and reduce that as much as possible before we toss it over to the data bricks jobs.
Eldad: Benjamin, this is a huge insert into the group by, by the way.
Benjamin: I see.
Eldad: Translating to relational algebra, insert into group.
Benjamin: Ah, now I understand...
Eldad: Billions, billions of unique events.
Benjamin: I always need those relational algebra trees to understand what's going on.
Eldad: That's why I'm here. You're welcome.
Vijaye: And then what happens after that is actually store all the raw events for when our customers need and then they want to be able to go download this thing and when you download, you get this a terabyte file, which is pretty crazy. It takes days to even download.
Then, obviously, we offered several ways to slice and dice, filter, so you can get your data as quickly as possible. But we store all of that stuff and so we're not quite at the petabyte yet, but we're quickly approaching that.
Benjamin: Awesome! So what are actually the scalability limits on those pipelines? What's harder basically growing to 10X more customers or a single customer of 10X the size. Is it similar? Is it very different? How is that looking?
Vijaye: Yeah, it is very different. I think 10X customers is actually easier than one customer growing 10X especially if that customer is extremely large. Because what happens is the joins and the analysis that we do are within the same customer's data. We actually have separation of data. These things take much longer and require a lot more orchestration. It's interesting, some of these things that I never thought we would run into, we are running into. For example, one of the things, where in the data bricks job runs, it spins out thousands or thousands or spot instances, spot VMs, then actually go and do the actual heavy lifting. And then to communicate with each of these spot instances you need local IP addresses and you run out of IP address allocation. Really! I never thought that you would actually run out of those kinds of things and apparently those are resources and then you're, Okay, well how are we going to solve this problem? And then you have to start thinking about, okay, well..
Benjamin: IPv6.
Vijaye: Pretty much, right. It is fascinating, when you're running some of these jobs at scale with so much data. Even the choice of how we store the data, what kind of databases we use, and how we run these orchestrations. It's been affected and it's actually evolving too. Obviously, we didn't foresee all of this stuff, so we built something that worked six months ago, and then you throw it all away and then you rebuild based on, well, we know now, and I am pretty sure six months from now, we'll probably throw all of what we have right now and then rebuild something for whatever scale that we will be dealing with. Because those are the times we're actually going to identify some new problem that we haven't even thought about.
Benjamin: Right. So, what's, for example, now on your current stack, an experience where saying, Wow, we would really like to provide this, but it's kind of something we just can't do on the architecture we have?
Vijaye: Well, I mean, wouldn't it be awesome if we're able to give you more real time than one hour? So imagine if we're able to give you in minutes, 5 minutes, 10 minutes, that would be amazing. Currently, our systems or architecture, our expense, the cost of infrastructure is prohibitive enough that we cannot provide that. So we have to work for an hour, as real time as we can get at this point, but I would love to keep pushing that limit as much as we can because I think we've gone from many, many weeks to an hour, which is a pretty huge step forward. But we're not going to stop there.
Benjamin: Awesome!
Eldad: I smell subscriptions coming along soon.
Vijaye: Yeah. I think especially as we go into this Web 3 world, I'm assuming that that's going to become a real time, an important issue. So we will have to keep investing.
Benjamin: Awesome!
To wrap up the conversation about your internal data stack, I assume you guys are dogfooding basically, so that you're also using Statsig internally to try out the things you are launching. Do you want to quickly talk about that maybe?
Vijaye: Oh yeah. Everything we do is behind the feature flag. So we don't ever throw out any code that is even a little bit sensitive or a little bit, we want to validate it. And so we throw everything behind a feature flag.
Obviously that gives us the separation of, when we can roll out a feature. Sometimes we do an announcement to everyone and then when we do the announcement is when we open up the feature flag. Obviously, we are also monitoring how people use these new features and when we build a new feature, is there adoption to this feature? And then we also constantly consider, okay, should we invest more in this feature or not. Those are product decisions that we make on a daily basis based on the data that we know how people are using our own product.
Our marketing side is full, feature flags. If you go to our marketing site, the button text then gets a demo, that text is actually different for every person, and that is driven by a multi-arm banded experiment, which is automatically deciding which text actually yields the most number of conversions. And then we'll start to show more and more of the winning variants. So we have 8 different variants that we throw into that system. So, we use stat's pretty extensively.
Benjamin: Awesome! Cool!
Maybe to close up today's podcast episode, we'll shift gears a bit to more on the advice side for our listeners. On my end, Okay, Eldad said it earlier I have just more academic background, so I'm relatively new to industry and you've seen all of it, right? Over three decades or well starting your third decade at Big Tech now at startups, do you have any advice for aspiring engineers basically?
Vijaye: Yeah, I think, as an early engineer, I always chase technology. So in the beginning I was, Okay, I want to be, for me, the compiler engineers are like Gods. And I was like, Oh, I want to be a compiler engineer. So, I went in... I was in Microsoft, I worked on compiler technologies. I worked on language services, incremental compilation and stuff like that.
Eldad: Benjamin's upcoming paper is on the query compilation, I'm sorry.
Vijaye: Oh, really.
Benjamin: I haven't told anyone yet.
Eldad: We cut it out, but we don't cut anything out. We'll list it after we release the paper. Sorry, go ahead.
Benjamin: Three months.
Vijaye: Yeah, and I went in, worked on compilers for a while and then over time I started realizing technical problems actually ladder up to user problems. Okay. What user problems are you solving? And so you have a view on, okay, this is what the problems are, and those translate into technical problems. And so you just start chasing specific user problems that you're interested in solving. So developer problems and so I worked on lots of developer frameworks in Windows and such.
And then over time, you start to realize a lot of the user problems after you peel the first couple layers are about the same, the technologies that you deal with, the problems that you solve are repeated over and over and over. You start to pattern match and stuff.
Then, it starts to, okay, there are people that I want to follow, go from technology to use problems to people, you start following people because you pattern match, like this person has done things that I want to do. I want to aspire to be these people. And so, pick those folks. And then, because you don't have to chart your own path, some of these things have already been done. So, we follow them. Pick good mentors. Pick good managers. And learn from them because there's so much to learn from people that came before you.
As you grow, then there's an element of turn around and pay back, because just as you grew on the coattails of other great folks, now there are a set of folks that are joining and learning the trade, turn around and mentor them and coach them, creating that followership. I think that's the journey of an engineer as you grow in your scope, as you grow in your craft.
Eldad: Boom! I love it! It doesn't matter where you work. If you are an engineer and if you love being an engineer and if you actually think about the problems that you are solving as an engineer, there's no limit and Benjamin, I hope I've been mentoring you enough, and if not, I will try harder because, yeah, it's future engineering is all about.
Benjamin: What am I supposed to say now, Eldad? Awesome!
Well, I think those were great closing words. Thank you so much Vijaye, for joining us. This was awesome. I think we've all learned a lot of new stuff. Cool!
Vijaye: Well, thanks guys.
Eldad: Thank you Vijaye.
Vijaye: Thanks, Eldad. Thanks Benjamin. Thanks for having me on your show. I'm looking forward to watching this.
Benjamin: Awesome!
Eldad: Absolutely!
Vijaye: All right, guys. Take care.
Benjamin: Take care.
Eldad: Bye-bye. Take care.
# How Substack's Data Platform Supports 500K Paying Subscribers (/blog/how-substacks-data-stack-supports-500k-paying-subscribers)
Substack is an amazing — if not the most amazing — content publishing platform out there. Essentially, it allows anyone to become a journalist or to start their own newsletters and charge subscriptions for them. So how did they build a data stack that can support all of their 500K paying subscribers?
Listen on [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-substacks-data-stack-supports-500k-paying-subscribers/id1561927688?i=1000530877276) or [Spotify](https://open.spotify.com/episode/0sfJrGQXZ2ddTEo93m1GeK?si=rLvA6cyvROWFuYnsEjLmlw)
**Boaz:** We are very proud and lucky to have Mike Cohen with us today. Mike is a super talented data engineer at Substack. Now, for those of you who don't know Substack, you should! Substack is an amazing — if not the most amazing — content publishing platform out there. Essentially, it allows people like you and me to become journalists or to start our own newsletters and charge subscriptions for them. And Substack has been growing. In February, they reported that they have 500,000 paying subscribers.
**Eldad:** That's old news.
**Boaz:** Yeah, that's probably old news by now. Every month that number grows like crazy. They've been getting a lot of media attention. I think that we should consider stopping this podcast and moving to Substack or something. That's the place to be right now. So, Mike has been working in the data space for quite some time prior to Substack, which we'll hear all about. He also spent time at companies like Capax, Venmo, and a lot of other exciting places.
**Boaz:** At Substack, how much data, in terms of data volumes, do you guys deal with?
**Mike:** I don't know what big data means when people say "big," but I think we're in the small to medium data camp still. We're in the tens of terabytes, not in the petabytes or anything like that quite yet. But the data volumes are growing quickly as more and more people come on. We have a lot of event data that we're logging and that's where we're at today.
**Boaz:** I think the rule of thumb to "big data" is that everybody starts with apologizing that they have data, but it could be bigger. Yeah, we only have hundreds of terabytes. We're not petabytes scale yet.
**Mike:** Yeah.
**Boaz:** I think that is considered big data. And what's the headcount at Substack these days?
**Mike:** We just surpassed 40. In the last couple of weeks, we broke through the 40 mark. We're on a big hiring spree. Come check out our jobs page!
We're trying to scale up the team. Fingers crossed, we'd love to be somewhere in the seventies by the end of the year.
**Boaz:** And how many people deal with data?
**Mike:** We are a very small team in an already small company. We're a two-person team at the moment. So, it was just me for the first 15 months or so and then I recently brought on someone so we have two of us now since March.
**Boaz:** So, you've grown really fast. In terms of subscribers and visitors, you've grown dramatically also in a little bit over a year. How does that look like from the retention perspective?
**Mike:** It's been exhilarating. It's been super fun to watch. And the problems to tackle have grown with the growth of the company in general. Exhilarating is the best way to describe it. Constantly thinking about "this thing that we were doing a week ago — now how do we do it at a much bigger rate and faster pace? How do we design our systems for a couple of weeks and months from now?" Stuff like that. So, it's been super fun.
**Boaz:** So, let's talk about your data world. Tell us what it looks like and what kinds of things do you do with it.
**Mike:** We have a Postgres production database and what we call our events pipeline, which is effectively a Kinesis stream of data that gets processed in parts into S3 and subsequently then dumped into Snowflake. Separately, we also have a process that will mirror our data from production into the data warehouse. Other data sources are getting piped in there too. And so, everything ultimately lands there. Then we have our BI tooling set up on top of that. From there we do transformations internal to Snowflake, and then we pipe that back out to places. I think the phrase I've seen a lot of these days is reverse ETL. But we send our derived or transformed data back out to a separate Postgres database such that the data can be accessed in the product with indexes and be super-fast. So that's our high-level structure today.
**Eldad:** What about BI? Which BI tools are you using there?
**Mike:** We use Periscope Data acquired by a company called Sisense.
**Boaz:** So how much of the data stack at Substack do you consider legacy versus modern? How much has it changed since you joined or is it something that's already built to scale into the future?
**Mike:** That's a good question. A lot has changed since when I joined. When I first joined, we had very little in terms of BI tooling or any data warehousing. So, all of that was in the last 12 to 15 months and that's gotten us to where we are today. As a company, today we are thinking about how we start to really put the data to work now that it's much easier to work with and accessible, and we have the systems in place to put it back into the product or to do BI and analytics. And then I think after that, there'll be the next chapter of asking "What do we do next? How do we build towards more real-time? How do we build towards faster insights?" And unfortunately, as a small team, we have to take things in little chapters. That's how I kind of think of it. So, I think we're on chapter two now and then chapter three will be about figuring out how to ramp this up even faster and do even more.
**Boaz:** Who is driving the requirements for BI? With 40 people, is it specific departments or cross-departmental?
**Mike:** It's a mixture of data exploration work, which can be driven by either the product or data team. We also have other internal teams, whether it's our support team or what we call our partnerships team, and they have data questions. As a data team, we will help them answer those questions by giving them reports that they can monitor in Periscope.
**Boaz:** So how much of a bottleneck do you end up with? It sounds like you have a lot of supporting to do.
**Mike:** Yeah. That's one of the reasons for doubling and why we're seeking to double again, hopefully this year. I'd love to end the year with around four people on the team. And so it's definitely a factor. And I would say it was only somewhat recently we found our team chemistry... we used to be more technical users than non-technical users. In other words, more SQL users than non-SQL users. And only very recently has our growth shifted that equation to where now we have more non-SQL users. And so that bottleneck has started to become more apparent than it once was.
**Boaz:** What does your morning routine look like? Which tool do you open the most to check in on things every day?
**Mike:** That's a good question. I check Slack to make sure there's nothing in the data channel that someone has reported or asked about. Then I have about six Periscope dashboards that I look at in order every single day that are basically pinned in one of my Chrome browsers. They're high-level company stuff. And then I have a spam publication detection dashboard where I look for any bad actors and try to handle that, too.
**Boaz:** It's an interesting use case. How do you find the right actors in the spam publications?
**Mike:** I can't answer that. That'll give away my fraud rules and everyone will know how to beat them.
**Boaz:** Good point good point. I was just testing you, but okay. Let's do what we call a lightning round...
**Mike:** Okay.
**Boaz:** So don't overthink. Shoot straight. Let's see what you come up with. Are you ready?
**Mike:** Yes.
**Boaz:** Commercial or open source?
**Mike:** Commercial.
**Boaz:** Batch or streaming?
**Mike:** Streaming.
**Boaz:** Write your own SQL or use a drag-and-drop vis tool?
**Mike:** Write my own.
**Boaz:** Work from home or from the office?
**Mike:** That one is hard.
**Eldad:** Both. You can have both.
**Mike:** Yeah, I think three days at home, two days in the office.
**Eldad:** Yeah, exactly.
**Boaz:** There are Hybrid modules now so it's legit to say both. AWS, GCP or Azure?
**Mike:** AWS.
**Eldad:** So, you can pick one. To DBT or not to DBT?
**Mike:** Controversial, not to DBT.
**Boaz:** To delt delete or not to delt delete.
**Mike:** Not to delt delete.
**Boaz:** Not to DBT, I think is the first time we had somebody said, no.
**Eldad:** This is the first time.
**Mike:** I know.
**Boaz:** Let's talk about that.
**Eldad:** You're probably mistaken. You probably got that answer wrong.
**Mike:** All your listeners just turn this episode off.
**Eldad:** We put it in the trailer.
**Boaz:** So, elaborate on that a little bit. So, what's your take on DBT and why not?
**Mike:** No.
**Eldad:** Big no. It was a big no, that's why.
**Mike:** It's not, it's not. I don't have that strong of an opinion. I've used it, but I have another system that I kind of put together that does similar behavior and I think allows a little bit more control. Ultimately, with DBT, my understanding is you still need to have something that schedules and orchestrates the jobs and so I've just kind of created a system that does a lot of that. I don't want to say it's anywhere near as comprehensive as DB, but does a lot of that. And it's all just based in Airflow, so it's Python-based.
**Boaz:** You said you guys are on Snowflake. How much processing? Did you guys do ELT exclusively in Snowflake? Do you do a lot of processing also outside of Snowflake?
**Mike:** No. A hundred percent of the processing is happening in Snowflake.
**Eldad:** Have you ever considered using spark to do that? Or just started a clean sheet with Snowflake? No need to migrate anything.
**Mike:** That was my thought. Yeah, it was clean sheet. Start from scratch and we'll see when we need to go bigger than snowflake "can handle." I'm sure they don't want to hear that, but I'm sure there's a point where using something like Spark in a more distributed fashion — where you can have a lot more control — might make a lot of sense. But we're not there yet at least.
**Boaz:** Looking at your pie chart of time spent on which activities, how much time do you spend supporting the BI users and the BI tools versus supporting the warehouse or supporting the pipeline, and so forth?
**Mike:** Yeah, not to cop out on the question, but I'm pretty evenly split at the moment. And there's like another administrative chunk, which is hiring and building the team out so that I can be more places all the time.
**Eldad:** So basically, like most high-growth startups, 70% of your time goes to hiring and then the 30% that's left goes on real stuff, which is great.
**Mike:** Yeah, I would say I'm at 35% hiring, 35% support of different people or functions and meetings, and then the remaining 30% is split evenly between either data engineering work or just my own data analytics work.
**Boaz:** Always hiring is also always a good excuse because if you tell people, "I'm sorry, it will be fixed the moment we hire another person, so I'm actively hiring." It's not your fault essentially.
**Eldad:** Boaz loves hiring. He discovered hiring a few months ago.
**Boaz:** Eldad is always complaining, "Why didn't you do this? Why didn't you do that?" And I tell him I'm hiring for it.
Okay. So, tell us about an awesome win at Substack.
**Mike:** I mentioned before we want to get to a place where there's more real-time analytics and more real-time insights in the product. But right now, it's we're living in a batch world where our definition of "real-time" is really every 20 minutes. We're kind of updating data in place. But I'm pretty proud of the system that we have. We're running a bunch of interesting, complex queries that create meaningful tables that are de-normalized and great for analytics, but also great for serving up things in the product. And then the way we're piping that back to Postgres with indexes in a way that is efficient and scalable is pretty neat. So I'm happy with that system that we have in place.
**Eldad:** Connecting data back to the product and feeding the product experience with data is huge. And you're right, it is super satisfying to get there.
**Mike:** Yeah. And when you send a newsletter a lot, some people want to just refresh, refresh, refresh, refresh, and watch the numbers tick. We're not there yet and I want to get there. There'll be other exciting things between now and then, but that will be a really exciting day for me.
**Boaz:** Now enough with this self-gratification, and then the win stories, tell us about an epic failure.
**Mike:** Okay. That's a bigger list.
**Eldad:** Everything that happened before we managed to get data back to the product.
**Mike:** I guess I should say, thankfully, there haven't been catastrophic errors that we can attribute to the data team, but there have been things we've done poorly. For example, we were writing data too aggressively to this Postgres thing I keep talking about and we ended up filling up the write ahead log and knocking over the database. All queries started to time out and then the site went down. So, we've had a number of learning experiences about how to do things that keep the site running and how to test things a little bit better. We also use some of our metrics and monitoring tools like Honeycomb to have a sense of when things might be going wrong and then we try to prevent that from happening in the first place. So, there's been a lot of small disasters, nothing too catastrophic yet. The keyword is "yet" because I'm sure that it's coming.
**Boaz:** What's the top challenge for data engineers or the data team in general at Substack?
**Mike:** I prefer to keep the surface areas small. So, in many organizations, there might be a data warehouse or something like that and then data is sent back out to a lot of different other services. And I am trying to not send it out to too many places because my fear is that ends up leading to a situation where you have one person looking at something in Google sheets, or Excel, or Air Table and saying, "Oh, I see this number here, but over in the BI tool, the number is different." There's a lot of ways to try to control for that, but one way I think to control for that is to try to centralize and keep things in one location.
And so, the thing I'm thinking about a lot recently is how to make our data and our analytics more self-service? And more self-service for not necessarily technical users. Whether it's building canonical data sets that are easy to query and we give everyone a little SQL lesson, or we get some sort of tooling that doesn't require SQL knowledge at all. How do we democratize the data access a bit more? So, I think about that a lot, and that's maybe more on the analytics side than on the engineering side, but I think that they're two sides of the same coin, really.
**Boaz:** How much of your responsibilities are on the analytics side as well?
**Mike:** All of it.
**Boaz:** Building out company dashboards and stuff like that. All of it? Wow. What gets on your nerves the most in your daily work with data?
**Mike:** Reconciling data from different data sources. For example, yesterday, I spent a long time trying to reconcile data that we're seeing from a test that we're running in Optimizely with our own data events logging and it's hard. It's an example of what I was talking about just a moment ago where Optimizely is a bit of a black box. You want to be able to put trust in the tool but sometimes you verify it yourself and in verifying it yourself, you end up in a rabbit hole. And so that can be kind of frustrating but it's important.
**Boaz:** Yeah. I feel you on that one. It's important, it's frustrating.
**Mike:** Yeah.
**Boaz:** For sure. Okay, so we're close to reaching the end. We want to get your advice on which companies, leaders or people to follow that inspire you or that you find interesting online.
**Mike:** Data Engineering Weekly is a cool Substack newsletter. Tristan Handy has a great newsletter, which is not Substack based but is still a big newsletter.
I'm also interested in what you guys are doing at Firebolt. I love Snowflake, but one of the reasons we're sending data back out to Postgres is because you lose the ability to index or to have functional aggregation that are very snappy. And so, I think getting the data warehousing or the OLAP databases to look more like databases where the index is going to be very, very interesting and compelling in the near future. So that's very interesting to me.
**Boaz:** Thank you so much. I think this is it. Eldad, what do you think?
**Eldad:** I think his interests are really spot on.
**Boaz:** Imagine Mike is not here with us. What would you tell me about him behind his back?
**Eldad:** So, I think the stack is perfect and frustration is hard, but it's a daily frustration we all deal with. And really, I think I wish you all the best. I wish you can scale fast. I wish you can hire a team that you love and can work with and accomplish meaningful things together. Always great to see you.
**Boaz:** And congrats for being a part of the success of Substack. So, keep in touch and thank you everybody for hopping on this episode of the Data Engineering Show. See you soon.
**Eldad:** See you soon. Bye bye.
**Mike:** Thank you.
# How Vimeo Keeps Data Intact with 85B Events Per Month (/blog/how-vimeo-keeps-data-intact-with-85b-events-per-month)
How does the Vimeo data team deal with 2 PB of data and 85 billion events per month? What made them recently build a data ops team? What data tool does the team love? And why (the hell) did they call their legacy platform Fatal Attraction?
Listen on [Spotify](https://open.spotify.com/episode/4hQMUFl974ipOfrUHGSfmL) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-vimeo-keeps-data-intact-with-85b-event-per-month/id1561927688?i=1000532388799)
**Boaz: We recently had the pleasure of speaking with [Lior Solomon](https://www.linkedin.com/in/liorsolomon/), VP of Engineering at [Vimeo](https://vimeo.com/).**
**Of course, Vimeo doesn't really need any introductions. We all have watched a video at one point or another on the platform. Lior has been with Vimeo for almost three years. He joined as the Head of Data, then moved his way up to VP of Engineering. He is an industry veteran, having been in data leadership roles and engineering leadership roles for quite some time.**
**Anything I missed?**
**Lior**: The only thing I'll elaborate on is what Vimeo is. It is known brand and video platform for a lot of people. But you know, for the past five or six years we've been focused mostly on helping small to medium businesses scale their business and bring an impact with video.
So, in the context of data, that's where it becomes interesting. If we're trying to provide these businesses insight about potential competitors or other businesses that are driving an impact, the data should be intact before we try and do that. So, I joined Vimeo with that objective.
**Boaz: "Data should be intact." That's a good tagline. We should make t-shirts. That could be a big seller.**
**Okay. Thanks. So, let's get to it. Let's talk about data at Vimeo.**
**Let's get started with the sheer numbers. What kind of data volumes are you dealing with?**
**Lior**: Last time I looked at it, we collect about 85 billion events per month. We have a couple of data warehouses, unfortunately. We have about 1.5 Peta of data on Big Query, and about half a Peta on HBase.
Most of the data is streaming in as viewership — you know, the analytics around who's watching what and how frequently, the quality of experience, if it is buffering, which regions people are watching from, and if the videos are distributing properly. That's the motive of most of the data coming in.
Beyond that, just like any SaaS platform, is just gathering the user experience, analytics, and any events and trying to make sense out of it.
**Boaz: Ok we'll spend more time diving into the stack because it's interesting how you spread all the eggs across more than one bucket with all these technologies, Snowflake and BigQuery.**
**But, people-wise, I think Vimeo is a little bit over 10,000 people strong, more or less. How many people are in data-related roles? How are those teams structured?**
**Lior**: We're about 35 data engineers. We have a data platform team that focuses on the real-time processing and the availability of the pipelines —mostly Kafka stack. We used to have a managed Kafka clusters, but we moved to a manage one of them in Confluent.
The data platform team is also working continuously on building a framework to make it easier for other engineering groups to consume data from Kafka — just like having their own data syncs dropped or whatever.
There's another team called the video analytics team, which focuses on the consumer-facing video analytics, basically serving the consumer on the website and on the mobile apps. They work on getting the data closer to real-time with all the aggregation and all the challenges around that.
There's another team for enterprise analytics, which focuses mostly on the CDN aggregation. We get a lot of data from vendors on a daily basis that we need to aggregate to understand how much bandwidth is being consumed by each one of the accounts. There, the stack is mostly Dataproc using BigQuery, and some of that data set eventually leads into Snowflake for the BI team. The enterprise analytics team is kind of the gatekeeper of the data warehouse and Snowflake. They're the ones that actually do the ETL integration bags, Airflow and so on.
The data ops team focuses on the time to analysis. They help product teams to onboard their event, payloads, work with them to do the event, and data modeling. They're the team that actually provides a framework internally for the big picture. We like to think of it as a structure stream where any PM or engineering team can go and create their own schema. And once that schema has been created, we provide the SDKs to whatever platform they use. And that framework basically validates the schema and lets those teams find where exactly they want to drop later.
**Boaz: Has data ops always been around at Vimeo or is it a newer addition to the team? Tell us about evolution of it.**
**Lior**: It's pretty new. We launched it maybe eight months ago.
**Boaz: And what drove you to launch it?**
**Lior**: That'll take us into the legacy and the history behind Vimeo.
**Boaz: So, let's roll up the sleeves. Let's talk about the history. How has the stack evolved?**
**Lior**: So originally, we had another home group kind of pipeline for collecting data from different names, called Fatal Attraction. Don't ask me why. That's the name.
It's basically just like unstructured data. You can push whatever you want, but you have like, basically three different columns where you can push. The original idea was like, send me the component where the event actually happens. Send me the page, the actual payload and unsurprisingly after many years, all the massaging of data happened on the ETLs downstream by the BI team. You see ETLs of thousands of lines of code that basically extract the data. In some cases, it's JSONs someone decides to send.
**Boaz: Were these Spark-based?**
**Lior**: No, it used to come in from basically backend logs. We'd aggregate all the backend logs, push it into Kafka, drop it into Snowflake. And then on top of the raw schema in Snowflake, we'll build ETLs that make sense out of it.
**Eldad**: One big variant field with all the adjacent in it. Thousands of ELT processes to parse and structure it.
**Lior**: Absolutely. Yeah. So, you know, there's a couple of challenges here. First, it is so easy to break this pipeline. You know, an analyst worked for quite a while to find his way to the data, and then add a couple of lines of code to those thousand lines of code. And then exactly a couple of days later, someone changes it upstream. The analyst is not even aware of it and we have a broken pipeline.
That's the world we're trying to shy away from as much as possible. Today we're focused on improving the analytics efficiency, letting them be as fast as possible and letting them focus on driving insights. We're probably not the first company to do that.
But going back to Fatal Attraction and the unstructured data. We got to the point where we said, okay, let's build a different framework. Let's make sure that the there's a contract with the client sending the data and we can validate the data. That way we can find the ownership to who's the one submitting the data and emitting the data. And basically we're using schema registry to validate on top of all those contracts, all those like schemas.
For each event, we have a valid topic and an invalid topic. So, once it goes away in that topic, we know, oh, something's wrong. We can go and ping the team that actually is responsible for that pipeline and tell them, "Hey guys, something's wrong." They can go fix the issue. And we can rerun the events because we don't lose them. We still have them.
Problem solved, right? It's awesome. Now everything's like intact. If anything breaks, we know immediately.
**Eldad**: There's always someone to blame.
**Lior**: We don't do that.
But, you know, we created a new problem that goes back to your question about data ops.
From the engineer's perspective upstream, like the one building the applications, you know, working on the user experience, they're like, "oh, that's awesome. I can go into like the big picture UI. I can set up a new schema. I can focus on the questions that I have because I don't care about the other teams. I have a question I want to answer. And I personally just want the data in amplitude." Because, you know, that's how my PM is using… That's how they want their analysis and they don't really think about needs of the analyst or the marketing guy and so on and so forth.
So really quickly you end up with like hundreds, if not thousands of different schemas. The life of an analyst is like: they come in. Someone knocks on the door like, "Hey guys, can you tell people this hypothesis? I want you to analyze this user behavior." And they spend weeks to understand which events they should cherry pick from the list of events that exist.
Yes, they are more reliable, but still, it's not quite clear what the context is, what you should use, And so on.
So, data ops came to that for two reasons. It's first of all, to stop that behavior and say "guys, before you go and create a schema, let's sit down. What are you trying to do? What are you trying to achieve?"
Say if you're trying to track if someone paused the video, I think that already exists. Don't go and create another one. Or if you're just missing some data, let's think about how we can extend the existing event or coordinate with the other teams and make sure it's event data modeling.
That's one objective. The other objective is to create the automation that we need for that specific framework because downstream for the analyst, it doesn't make sense for them to be so familiar and savvy with each one of the schemas. At the end of the day, they want to have click stream.
There's a set of properties that they're looking into, they're looking for some sort of standardization. So, the data ops team was working on creating what we call internally "global properties."
Basically, we have roughly three engineering organizations within Vimeo. It's about 500 engineers. Each one is responsible for a certain application. We signed a contract with all those engineers saying they have to provide us with that information. That's like the basic bread and butter for any analysis. That's how we create that.
So, a lot of the automation is done by the data automation team. And that team is actually a cross-functional team. In that team, you have a front-end developer just for the UI and all the "making fancy," and a data engineer that builds the ETLs.
**Boaz: How big is that team?**
**Lior**: Four right now.
**Boaz: But you mentioned that they're essentially in charge of standardization and consolidation of models.**
**Eldad**: Its basically going full self-service and then turning that into kind of a fully managed self-service. Removing tons of friction, trying to find common data sets, common metadata.
This uncontrolled service, pretty much going back to the Excel spreadsheet nightmare. We all grew on just a better scale with data engineers instead of information consumers.
**Boaz: I wonder, putting that team in place and suddenly telling the org "let's stop for a second," like you said, how do you enforce that? How did you manage the process around that? Because you know, an analyst I might guess that they can just go ahead and do what they did before. Sometimes it would be faster to not even ask if we're around the data ops. So how do you manage?**
**Lior**: Yeah, that's a great question. I won't be able to say, "Hey guys, stop for a month, we're moving." We actually moved, we worked with both of the pipelines, so we still have the legacy to Fatal Attraction, which is slowly draining because as the new products are going, everyone's speaking the framework. So, we are slowly draining that pipeline. I'm having active discussions with leadership to be more aggressive. But that's it: we need to be completely deprecated, even from the legacy parts of the website, because we're going towards using the new framework for experimentation and AB testing, which we do internally.
One of the ways that I got the engineering teams to actually bear with me and work with me on the new framework was to find the carrot at the end of the stick of and say, "Hey guys, how about you get some self-serve analytics on aptitude?" And they were like, "That's awesome."
The analytics team is more focused on business. You know, KPIs, correlating, retention with FTS and subscription. More of the advanced analytics data science work.
And when it comes to, "Hey, tell me how many people click the button. Or show me a funnel." Applicant is just amazing at that. So, I say "go create that event." It immediately appears on aptitude that they created a self-serve world for a lot of the PMs and engineering.
That also created the problem because now again, going back to the original issue, there's not multiple schemas out of team and really hard to control this. So how would you get from this mess? Because it seems like you created a great framework where anyone can create whatever they want and they're happy about it, but the collective is not. So, in the coming quarter, we're working on an enrichment layer and we're looking into KSP over confluent.
The idea is that we want to dumb down as much as possible for the client sending the data. I don't need you to tell me about what subscription the customer has because I already know it. And I have those dimension tables. I don't need you to provide you all that information. I don't need to tell you which team ID they belong to.
So, the reason we can do it so far is because a lot of the obligations. And the business logic is maintained on the data warehouse and Snowflake. And when I'm serving the data from Kafka into amplitude or any of the end points, you know, it could be a Prometheus, whatever you're using, I don't have the access to those dimension tables.
So, using KSP, we're basically planning to have those tables closer to Kafka, KSP runs under the hood with persistent data. And then once that happens, we're going to create basically no more than 20 or 30 generic events, but just asking for the facts and whatever needs to be enriched.
That's going to be done through KSP. It's a framework that you can use in order to drop it to whatever data sync you need. So, you will be able to self-serve. The fact that the event's already designed. The data model is all in place.
We're dumbing down the client. It will speed up the deployment of new products.
And we're keeping the data modeling for data engineering and analysts.
**Boaz: Is this already launched in production or is this an active project?**
**Lior**: No, we just started working on it.
**Boaz: You mentioned before transitioning from managed Kafka to Kafka confluent. What was the story for that transition?**
**Lior**: I think overall, when you need to measure where you want to put your efforts versus maintaining your infrastructure versus putting it on managed, even though you pay for it, it's not for free. Right now, we are focusing our putting our CPUs into driving data insights and helping BI, helping data scientists, machine learning teams to actually move forward with their initiatives.
So right now, as a rule of thumb, if there's a service, I can put it as a managed service and offload that work from me.
**Boaz: You mentioned a huge variety of things at the end of the day that you guys do with data, and it's pretty impressive to see Vimeo being so data driven. Can you share what sort of workloads or use cases you enjoyed seeing that brought a lot of value?**
**Lior**: Yeah. I would say almost any department is actually consuming data from the data warehouse for their own initiatives all the way from marketing to return on targeted ad spend.
A lot of the data science projects are also utilized for Snowflake. I would also say the machine learning team uses some of their payloads. For example, we use Kubeflow together with Airflow to ingest and pull data from Snowflake. An important team is also 'Search and Recommendation', which is running Elasticsearch.
They're basically running our personalization and recommendation algorithms. Also they're utilizing data. The data comes from Slope. And the personalization data is actually stored into a big table and Snowflake. As we talked at the beginning, originally the reason we had both was just because we moved to GCP a couple years ago. When I joined, we just moved from Vertica to Snowflake, Snowflake had only the AWS implementation.
So, we kind of run both and I don't see us moving to Snowflake over GCP.
**Boaz: But what's your strategy there? Is the intent to stay on both AWS and GCP? I mean, you're running both Snowflake and big query — Do you consider that sort of a legacy tech debt, or is it sort of a strategy of using the right tool for the right purpose?**
**Lior**: The way I'm thinking about it is Snowflake to me is where I want to make sure there are high-quality data that we use by this different business. BigQuery and the stack on GCP are for data stores. There's something more engineering-oriented.
The easiest way to think about is like, I won't put any report which is coming from BigQuery in front of leadership. I'm not owning it. It's not my problem. Anything that goes into Snowflake, that's what we do.
**Boaz: So what are the big plans for the next year within your data initiatives?**
**Lior**: We are doubling downloads. For me personally, I'm really, really focused on the data availability and building trust with the overall organization.
We are actually expanding our machine learning teams and really trying to go in that direction. We are trying really hard to advocate for hiring more and to take more risks as a business. We spent a lot of last year on the data availability aspect, but you know, creating SLAs for some of the data stack, making sure that teams have the expectation what they don't like, the business, what is supposed to be done, what's done in response to any data outage. The more we can create that world, you set up expectations with the stakeholders because the easiest statement is saying "the data is wrong." It is not easy to draw those lines like where does the data engineering start? In a lot of the cases, I'm more facilitating the sessions that are not necessarily related to data engineering, but for example, we're trying to help the team by being their ambassador to find where the data problem starts. There's a transaction scope and deal, or maybe payments not getting in on time. It could be third party like one of the vendors, it's not an engineering issue. We're still trying to find this engineering language to help them get there. It's the world of human engineering required to be successful in data engineering.
**Boaz: We don't talk about it often enough, the politics of data.**
**Lior**: Yeah, it doesn't matter how good of a technology we build. It's a combination of people and processes and technology. We implemented Monte Carlo like September last year, which was super successful. I'm really, really happy with it.
**Boaz: Yeah. So, Monte Carlo are trending like crazy. Tell us a bit about the value you got out of it.**
**Lior**: We literally jumped to the future because we were actually trying to build some of those things, and I'll give you an example. So, let's talk about a little bit about data validation. So upstream data comes into the Kafka. Now we have a schema registry and fancy framework that know to tell if the data is right or wrong. Well, what we don't know how to do that yet. To tell tell you if there are any unknowns or anomalies in the data. So, you go downstream a little bit in the ETL and Airflow. We put great expectations there, and maybe the expectations are great, but it takes time to implement for the whole pipeline, all the ETLs. You have to prioritize. And unfortunately, usually you're prioritized by who's shouting at you more.
So we experimented with it a lot. We found it is good for specific anomaly detection, specific business logic. What Monte Carlo did, which wasn't the main thing in my mind, was that it hooked up to Snowflake and Looker.
And basically without setting up anything, they started listening and building those metrics for all the tables, the data sets we have. And suddenly, I'm starting to be aware of problems I wasn't aware of at all.
My first reaction was that sounds like pretty risky, because now we're going to spend our time jumping on like a thousand alerts. How's that going to go?
So, what we did was to actually start focusing on specific schemas or specific data and understand who's the owner of those nodes. Because you know, sometimes you cannot leave the issue with yourself. You need to go talk to the engineering team, understand what's going on. So, we started building those relationships where if I know what dataset, I know who's the team actually driving the data set. I can set up a Slack channel where any alert goes there. I'll always be there as a data engineer. So, I'll be aware of those alerts, but also I'll make sure that the stakeholders that also on that channel and the publishers are also in that channel so we have a team informed.
Now in reality, we need to build that relationship. Some teams are more data driven and they're excited about solving those problems. That's not the biggest problem. The biggest problem is we need to just start building those reports.
I'm getting a report and it tells me, from all the challenges you have right now, from all the teams you're working with, that's the responsiveness SLA. And for me, that's kind of like a get out of jail thing, because whenever someone says, hey, data is bad. I'm like, data is bad where?
Oh, it's bad on the CRM. That's the CRM. Well, you guys are not responsive to any of the alerts for the past month, obviously it's bad, you know? So that starts the conversation in the board where you're setting up the expectations with your stakeholders.
**Eldad**: That's the amazing thing about how Monte-Carlo takes your existing modern stack, your older mess, all the confusion and kind of backwards, built up draws a picture, shows you where the problems are.
That's a new way of driving data. A few years ago, we would have to have a dedicated team just for that. It doesn't work, especially when there's so much self-service going on. It's just amazing to see and we're super happy to hear about Monte-Carlo. We hear that a lot. So, we expect great things from this company.
**Boaz: You shared a lot of great things that you do at Vimeo but we cannot let you brag forever. Now it's time for you to share an epic failure with us. Doesn't have to be from Vimeo, can be from your data career. Let's hear about some things that didn't work.**
**Lior**: When we built this big picture framework, we were so focused on the self-serve aspect. We thought that would be amazing. They can go and create their own schemas. And it's so great. If it fails, we'll find it. That was the mindset. But then we said, hold on. Who's going to be the one that sets up the mandate in regards of like, why are we even creating that? And at what point do we bring the customers data? The analysts and the data scientists should be part of the discussion about what we need.
**Eldad**: Will be interesting to talk to you in a year and see how that experiment went. I mean, in many cases, companies just jump between two extremes. So, they either go full self-service and then they kind of get burned out and then they go full managed and four people try to route every request to the company.
I think that the modern data stack is here to help. So, will be amazing to see kind of how that plays out. Always, as you said, manage expectations, manage people. And prioritize.
**Boaz: Ok, maybe a weird question, but what would happen to the business if the budget of the whole data engineering and data-related initiatives were cut in half?**
**Eldad**: No ads on videos, for example.
**Lior**: Great point, because we don't do ads on videos. Do your research.
**Eldad**: Ahh, plugged on purpose. Of course I know that.
**Lior**: So, going back to the question about managed, not managed, I think right now we're focused on insights.
We strongly believe that data is on an exponential mode because we are a company that has existed for 16 years with no space on video. We believe that if we'll be able to land the processes and teams and infrastructure properly, we can start extracting insights about videos and help businesses thrive.
And really we need that. We won't be just the video platform. We'll be providing video insights for customers. So that's like, that's the main focus and how fast can we get there.
We just had the IPO and are in that sweet spot. It's not as money doesn't count, but right now anything that we can just delegate and just not focus on so that we can just focus on building those insights, that's going to promoted right now. But when my CFO looks at the Snowflake budget, he's really worried because we did grow it. I think from the first year, we grew it by 55% and the second year, I think 70%.
**Boaz: Just by the number of times you mentioned Snowflake, I'm sure your CFO is super worried looking at that bill.**
**These are amazing times to be in data because in the past data budget talks were "how can we, with the least amount of spend, get these annoying reports out of our way." Whereas today, modern companies understand that data is an investment. It's value. It is suddenly an X factor for business models for companies.**
**And it's okay to invest more than we used to, not less, to try and drive value out of it. So, it is super interesting to see how Vimeo does that.**
**Okay, good. So, we want to lighten things up a little bit as we near the end. We'll do a quick blitz round of short, quick questions. Don't overthink. Just answer quickly.**
**Commercial or open source?**
**Lior**: Commercial.
**Boaz: Batch or streaming?**
**Lior**: Streaming.
**Boaz: Write your own SQL or use a drag and drop visualization tool?**
**Lior**: Write your own.
**Boaz: Work from home or from the office?**
**Lior**: For me personally, at home, but I think it's more feasible to work from the office.
**Boaz: AWS GCP or Azure?**
**Lior**: Well, right now, absolutely GCP because we're all hands down into it. But it depends.
**Boaz: To DBT or not to DBT?**
**Lior**: To DBT. Absolutely, all the way.
**Boaz: Delta lake or not Delta lake?**
**Lior**: No.
**Boaz: YouTube or Vimeo?**
**Eldad**: YouTube Premium or Vimeo?
**Lior**: Oh that's a good question. It depends. If you're trying to monetize for ads, definitely Youtube. If you're a business caring about your brand mission and making sure you have clean video with no ads, then that's Vimeo!
**Boaz: That's the difference when we have executives on the podcast versus the hands-on people. If we interview an engineer at Vimeo, "of course Vimeo!"**
**Eldad**: Just true or false. There's no in between.
**Boaz: Okay. Very good. I think we're about to end. Anything else you want to dive into?**
**Eldad**: One quick thing. So, you've mentioned that you want to drive insight from Voice. People communicate and then you get tons of value out of kind of analyzing what multiple voices say on the call. Are you planning to make videos more interactive and embed more? Are you planning on making it less about broadcasting and more about getting some feedback from viewers so that it can actually drive some kind of a video analytics insight in the future? Is that planned? Anything you can share?
**Lior**: Short answer, yes. But I can't share too much.
**Eldad**: Nice! So, there's a big thing planned. Okay.
**Boaz: No comment, no comment. Lior, this has been super interesting. Thank you so much for joining our podcast. We that said we need to talk to you in a year. Please do not disappoint us, we want you to have all these milestones achieved with excellence. Thank you so much.**
# How we built Firebolt (/blog/how-we-built-firebolt)
## Introduction [#introduction]
To understand how we built Firebolt, we need to understand first what we built Firebolt for. We describe Firebolt as a "Cloud Data Warehouse for Data-Intensive Applications". Huh? What does it even mean?
Well, let's start with surveying the landscape of how people do analytics today. Today, we have a generation of modern Cloud Data Warehouses (CDW) such as Snowflake, BigQuery, Redshift, Databricks, etc. They are mature, versatile, battle-tested, and support a wide range of analytic workloads. Today's CDWs can handle BI, reporting, ad-hoc query and exploration, ELT, data science, ML, and, let's not forget AI. They deal with different types of data - structured, semi-structured, even unstructured, geo-spatial, time-series, and, more recently, vector embeddings. CDWs can handle data as batch, micro-batch, and streaming; they connect to a variety of systems - from OLTP to LOB, and, of course, integrate well with Data Lakes and Lakehouses. Amazing!
But there is one thing that the current generation of CDWs struggles with. They are not very fit for supporting data-intensive workloads which require low latency (tens of milliseconds) and high (thousands of) QPS. Workloads which are common in user-facing applications - like mobile apps or Web apps. Today, developers don't dare to connect directly to CDW from an iPhone app and have SQL queries against CDW power interactive experience. This simply won't work, or in rare cases where it might work - would be prohibitively expensive to run at scale. Again, today's CDWs excel at many, many things, but being the backend of interactive user-facing data-intensive applications - is not something they are good at.
So, what do application developers do today? They have to introduce another system for data serving that the application will work against. What system - really depends on the application access pattern. It could be a simple cache; if the application needs key lookup - Redis or Valkey is a good choice; if it needs text search - then Elastic fits the bill. For aggregation queries - Druid, Pinot, or Clickhouse will work. This serving layer could even be MySQL if the data volumes to power the app are small.
There are many specialized systems, each one being very good at what it is designed to do, and solutions employing a dedicated serving system generally work well. However, there are multiple problems with this setup. They all stem from the fact that the data now lives in two systems - one is CDW, where all analytics is done, and another one - the serving system which is exposed to the application. Having two systems is not too bad by itself, as enterprises routinely run dozens (if not more) of systems. However, this approach presents its own challenges:
* Data freshness is compromised, since that data needs to be copied from CDW to the service system.
* Each specialized system has its API, programming model, security model etc. It's not too hard to learn a new API - backend developers will pick it up quickly, but it's a different skill set than doing SQL and database development. And there are many more SQL developers in the world than backend developers.
* These systems all have their preferred data models that they work best with - so data needs to be prepared and molded into that shape. Again, this is not the hardest thing, but yet another set of pipelines to maintain and babysit.
So, we at Firebolt saw an opportunity here: Wouldn't it be wonderful if we could build a Cloud Data Warehouse which can handle these demanding data-intensive applications? The data won't need to leave such CDW, and it will handle both common CDW tasks as well as serve as a backend for workloads requiring low latency and high concurrency. It will have to be a real database system, meaning that it has all the properties expected from DBMS - real SQL (and not "SQL-like"), ACID transactions (and not "eventual consistency"), etc.
So here you have it - that's the Firebolt's vision in one sentence:
*Cloud Data Warehouse for Data-Intensive Applications*
Or, if you prefer in the poem form:
*Low latency*
*Data intensity*
*SQL simplicity*
*Cloud elasticity*
Nice vision and nice rhymes, but does the world need Firebolt? Who needs to expose their data to user-facing applications that require low latency and high concurrency? Or to rephrase it - why doesn't every enterprise build such applications on top of their data warehouses? Those are good questions, and we think that if it were easy to build such applications, every business and every enterprise would do it. Today, it's just too hard and too expensive. Well, Firebolt exists to make it easy!
## The challenge [#the-challenge]
Of course, building a **Cloud Data Warehouse for Data Intensive Applications** is not easy. The rest of this blog is going to give an insight into how we did it, where every decision was guided by the two simple principles:
* Our system has to be a real analytical DBMS. No compromises.
* The most important workloads are data-intensive applications. Optimize for them; everything else is secondary.
While building analytical DBMS, which works at petabyte scale with low latency and high concurrency, is difficult, there are some things which make it easier. The most important one is that application workloads, unlike ad-hoc queries, are very predictable and contain a limited number of query patterns. We fully take advantage of it by being able to set up indexes which fit the workload, utilize caching of sub-plan results, and automatically adjust query plans based on history. This will be covered in more detail below.
And we also want to be very clear that like always in engineering - building a system means making tradeoffs. In software (like in life), it's not possible to have a system which excels at everything. Getting excellent at A often comes at the expense of B.
And it is, of course, true about Firebolt. In this blog, we will show multiple examples of tradeoffs we took in order to comply with our two guiding principles. Some of these tradeoffs will even look counterintuitive, but, I promise, it will all make sense in the end.
One more question that might be nudging you - is it really practical to build a real user-facing application with SQL database as a backend? Answering queries in tens of milliseconds. Well - why don't we take as an example one of Firebolt customers:
## Example of real production workload [#example-of-real-production-workload]
To give an example of such a production workload, let's look at the latency chart of one of Firebolt's customers. It shows median and P95 values for latency in milliseconds - e.g. half the queries finish under 120 ms.

What's more interesting is that these are not simple queries. Let's take a look at a couple of them:
```sql
with O as (
select …
from D
where x IN (...) and y < … and z between … and …
union all
select *, …
from M
where x IN (...) and y < … and z between … and …
),
C as (
select …, coalesce(sum(case when … then … end), 0)
/ sum(case when … then … end) - 1
from O
group …
)
select …, sum(...), …
from O left join C on …
where … between … and …
group by …
```
This query has a couple of Common Table Expressions (CTEs), UNION ALL, aggregations and LEFT JOIN. The input tables D and M are tens to hundreds TBs in size (although WHERE clause filters are very selective).
Another query:
```sql
SELECT …,
COUNT(*),
ARRAY_COUNT_GLOBAL(...) / COUNT(*),
ARRAY_SUM_GLOBAL(TRANSFORM(x,y -> (y - x), …)) / ARRAY_COUNT_GLOBAL(...),
ARRAY_COUNT_GLOBAL(ARRAY_FILTER(x,y -> (x = y), …)) / ARRAY_COUNT_GLOBAL(...),
SUM(...) / ARRAY_COUNT_GLOBAL(...),
FROM (
SELECT …,
ARRAY_FILTER(x,i -> (i = 1), …),
FROM (
SELECT …,
ARRAY_FILTER(t,p -> (MATCH_ANY(..., ['(?:^|[^\\p{L}\\p{N}])(?i)(word)(?:[^\\p{L}\\p{N}]|$)']), …),
FROM D
WHERE a = … AND b NOT IN (...) AND …
)
) GROUP BY …
```
Here, we have aggregations and no JOINs, but a lot of array manipulations, including array aggregations and multiple array lambda functions, as well as the complex regular expression matching function.
These queries power the user-facing application, and with 125 ms median latency, and even P95 around 200+ milliseconds - it provides instantaneous answers to **user experience**.
## SQL API [#sql-api]
We set out to build a database system, and databases speak SQL. In some parallel universe, the term "speak SQL" would've been unambiguous, but in ours - "it is complicated". The SQL language has been [accepted by both ANSI and ISO](https://blog.ansi.org/sql-standard-iso-iec-9075-2023-ansi-x3-135/), since 1986, there were nine edition since (SQL-86, SQL-89, SQL-92, SQL:1999, SQL:2003, SQL:2006, SQL:2008, SQL:2011, SQL:2016, and most recently SQL:2023). Despite all this standardization - every single DBMS system has its own unique SQL dialect. Some are more aligned with ANSI SQL, some are less. We had to decide what Firebolt's SQL dialect would look like, and it was the most natural decision for us to align with Postgres's SQL: It is widely used, people love it, it works with virtually all ecosystem tools, and it is also probably the closest one to ANSI SQL. By aligning with Postgres on our SQL dialect, we achieve two goals:
1. **Ease of use:** Postgres is very popular, and many developers used it at some point and are familiar with its SQL dialect. So they would be immediately productive with Firebolt.
2. **Simplifying ecosystem adoption:** The success of the platform depends on the ecosystem around it. And virtually any application working against a database has support for Postgres. And if the SQL that these applications generate for the Postgres connector works without modifications in Firebolt - it becomes easy to have a Firebolt connector.
Choosing to align with the Postgres dialect is a popular choice in the industry - many other systems have done it as well: [CockroachDB](https://www.cockroachlabs.com/docs/stable/postgresql-compatibility), [DuckDB](https://duckdb.org/docs/sql/dialect/overview), and [Umbra](https://umbra-db.com/#features) are good examples for this.
## Both compliant and fast [#both-compliant-and-fast]
Choosing to be Postgres compliant does come with certain costs and tradeoffs. Particularly, being compliant, means that SQL functions should have exactly the same behavior, including error handling. Let's take the simplest example of the addition operator: x + y. This operator needs to throw an error if the result of addition overflows the result data type - this is both Postgres behavior and is prescribed by ANSI SQL. In real-world data, such overflows are unlikely, but the code has to be ready for them. It turns out that there is a non-trivial cost for doing such checks in a query engine which uses vectorized execution. So it seemed like we had a hard conflict between our two principles - we want to be a SQL compliant system, but we also don't want to sacrifice even a little bit of performance, as we need to power low latency workloads. Both requirements were equally important to us, so we invested engineering resources to satisfy both of those. Details about how it was done can be found [in this blog](https://www.firebolt.io/blog/making-a-query-engine-postgres-compliant-part-i-functions).
## Beyond Postgres [#beyond-postgres]
Postgres compliance is great, but modern data applications need many features beyond what it has to offer. Here, we will discuss one of them - support for arrays. Arrays (alongside with structs) are the two most important building blocks for semi-structured data support. And semi-structured data is ubiquitous in today's data world (thanks to JSON, but also protobufs, Thrift, Parquet, Arrow, and many similar formats). Postgres, being born in the world before that, implements arrays very faithfully to the ANSI SQL specification. Particularly, multidimensional arrays are expected to have same number of elements:
```shell
psql=> select '{{1},{2,3}}'::int[][];
ERROR: malformed array literal: "{{1},{2,3}}"
LINE 1: select '{{1},{2,3}}'::int array;
^
DETAIL: Multidimensional arrays must have sub-arrays with matching dimensions.
```
This is ANSI SQL compliant behavior, but it will be breaking pretty much any modern application. So Firebolt adopted a more flexible and relaxed definition for multidimensional arrays:
```shell
fb=> select '{{1},{2,3}}'::int[][];
?column?
-------------
{{1},{2,3}}
```
Moreover, the ANSI SQL's (and therefore Postgres') way of working with arrays - is to convert them into a relation using a lateral join with UNNEST, and rely on correlated subqueries. This is a very general approach, which allows to do any operation with an array that can be done with a relation using the full power of relational algebra (through SQL, but still). However, it is both cumbersome to program, and query optimizers struggle with efficient query plans on many types of correlated subqueries. Let's take a simple example of filtering out array elements. Suppose we only want to leave even numbers. The Postgres's version of SQL would look the following:
```shell
psql=> select a, (select array_agg(e) from unnest(a) e where e % 2 = 0) f from
(select array [1,2,3,4,5,6] a) t;
a | f
---------------+---------
{1,2,3,4,5,6} | {2,4,6}
```
There is a lot going on here for such a simple task - unnest, correlated subquery with aggregation to reconstruct an array after filtering - ouch.
Firebolt recognizes that if arrays are common, then operations with arrays are common too. So Firebolt has a wide arsenal of so-called "array lambda" functions - functions which operate on an array, and one of the arguments is a lambda function telling what operation needs to be performed on each element.
The above query can be rewritten in Firebolt as:
```shell
fb=> select a, array_filter(x -> x % 2 = 0, a) f from
(select array [1,2,3,4,5,6] a) t;
a | f
---------------+---------
{1,2,3,4,5,6} | {2,4,6}
```
The `x -> x % 2 = 0` is a lambda function here. OK, I know what you are thinking - what about [Codd's theorem](https://en.wikipedia.org/wiki/Codd%27s_theorem) which establishes an equivalence between relational calculus and first order logic, and aren't we escaping the boundaries of the relational model by introducing lambda functions which look like second order logic? Don't worry - everything's still fine; we are still within the relational model. There is proof, but the margins of this blog are too small for it.
But jokes aside, it is much simpler for the optimizer to reason about scalar functions (even if they operate on arrays and lambdas) than about correlated subqueries and unnest, and it can generate better query plans. Our customers use these a lot - in fact, query 2 from "Example of real production workload" was very lambda functions heavy. As a bonus, array functions (including lambda ones) are allowed in Aggregating Index definition, allowing Firebolt to accelerate them.
## Planner: Learn from history [#planner-learn-from-history]
We already discussed that data application workloads tend to be very homogeneous and repetitive. Since the same query patterns repeat over and over again, it makes a lot of sense to learn from the past execution in order to build better query plans for similar queries in the future. This is what we refer to as History Based Optimization (aka HBO). Many HBO techniques have been proposed - ISOMER \[SHMMKT06], STHoles \[BCG01], simplified ISOMER \[BMHM07] and more recently QuickSel \[PZM20].
Of course, Firebolt's planner also uses traditional data statistics, but given our primary focus on predictable and repetitive workloads - we have invested most of our efforts in HBO. The diagram below shows the overall HBO architecture.
To make it work, we first normalize query plans. The goal of normalization is to let queries sharing the same semantics but differing syntactically use the same plan. The normalization process includes but is not limited to
* Remove projections so column pruning and ordering do not matter.
* Join graphs are normalized so join ordering does not matter.
* Filters are sorted so that the order of predicates does not matter. But note that concrete literals do matter.
* Order of window functions applied on the same window frame does not matter.
We fingerprint every subtree in the plan (not just the root), and attach observed execution metrics to it, e.g. row counts. When a new query arrives, the planner tries to match its plan subtrees with the ones learned from the history service and uses historical statistics in the estimation process.
For queries where exact matches are not found, but the plan shape matches (for example, if literal values in the new query are different), we use a variation of QuickSel. A simple example would be the filter operator. Suppose we saw queries with filter predicates `x > 10` and `x > 50`, and learned statistics for them. A new query now comes with `x > 20`. The QuickSel algorithm allows us to estimate cardinality for this unseen predicate. To do this estimate, QuickSel models relations as k-dimensional spaces, and it models query filters as hyper-rectangles in these spaces, where k is the number of relation attributes.
Learning from history is very important for Firebolt main scenarios, and investing in more sophisticated HBO is a priority for our planner team. The tradeoff here is that focusing on HBO means less investment in more traditional CBO (cost-based optimization) and fancier data statistics.
## Runtime: Smaller rather than faster [#runtime-smaller-rather-than-faster]
We already discussed that our production workloads tend to be very homogeneous, with 10s or 100s of predictable query patterns. Such repetitive workloads can benefit tremendously from reuse / caching. In analytics systems, caching as a concept is ubiquitous: from buffer pools to full result caching to materialized views. Firebolt employs a surprisingly little-used approach: caching subresults of operators. The idea itself is not new, with first publications appearing in the '80s, e.g., [\[Fin82\]](https://dl.acm.org/doi/10.1145/582353.582400), [\[Sel88\]](https://www.sciencedirect.com/science/article/abs/pii/0306437988900142). More recent results include [\[IKNG10\]](https://dl.acm.org/doi/10.1145/1559845.1559879), [\[Nag10\]](https://homepages.cwi.nl/~boncz/msc/2010-FabianNagel.pdf), [\[HBBGN12\]](https://vldb.org/pvldb/vol5/p1436_alexanderhall_vldb2012.pdf), [\[DBCK17\]](https://cs.brown.edu/~kayhan/papers/hashstash.pdf), and [\[RHPV et al. 24\]](https://www.amazon.science/publications/why-tpc-is-not-enough-an-analysis-of-the-amazon-redshift-fleet). It also would be appealing to use HBO (history based optimization, discussed above) to help detect common subplans across consecutive queries.
This whole subject is fascinating, and we have a dedicated blog [going into depth on it](https://www.firebolt.io/blog/caching-reuse-of-subresults-across-queries), but here we want to focus on something else.
Since caching is so important for Firebolt's main focus - data-intensive applications - obviously the more we are able to cache, the better. The amount of RAM in the engines is fixed, so the only variable we can control is: using less memory for the cached data structures. And this is exactly how we have designed our runtime data structures. Let us take the one used by our Hash Join implementation as an example. The hash tables built for the Hash Join are great candidates for caching, as there could be many queries with different left sides of the join, but the same right side of the join (e.g., when joining against a dimension or lookup table) - if the right side is chosen as **build** side of join, it will be the one for which hash table is built.
Usually, designing for performance means coming up with the algorithm with the lowest run time. However, since our priority was to squeeze "more data structures" into the fixed sized memory, for building JOIN hash tables, we did not optimize to have the necessarily fastest algorithm, but aimed for an algorithm which produces smaller hash tables. The details will be described in a dedicated follow-up blog, but the general idea is to make multiple passes over the right side. In particular, we compute statistics so both join keys and values can be packed into as few bytes as possible. Additionally, we even sort by hash values, in order to be able to do delta compression on offsets. Yes, all this may take more time, but if we build this hash table once and use it many many many times, and it takes less memory, there is less of a chance for it to be evicted - i.e., it is worth it!
## Distributed execution: Shuffle dilemma [#distributed-execution-shuffle-dilemma]
Firebolt supports multiple dimensions for scaling the processing power of a cluster. One is scale up, where the user simply selects a more powerful node type. However, changes in node type are somewhat coarse grained. If more precise control over sizing is needed, or if there is a need to have more processing power than what a single node can offer - there is a scale out dimension. With scale out, the user can increase the number of nodes in the cluster (and, of course, can also scale up or down the node types within that cluster). Logical query plans for single node and multi-node clusters are the same, but physical plans are different, as multi-node clusters need to add dedicated operators for data exchange between nodes. This data-exchange operator is called Shuffle.
Here is an example query, computing a join and an aggregation:
```sql
SELECT r.c, SUM(s.d)
FROM r JOIN s
ON r.a = s.b
GROUP BY r.c
```
The logical query plans are the same for both a single-node and a multi-node cluster:
```shell
[1] [Aggregate] GroupBy: ["c"] Aggregates: [sum("d")]
\_[2] [Projection] "c", "d"
\_[3] [Join] Mode: Inner [("a" = "b")]
\_[4] [StoredTable] Name: "r", used 2/4 column(s): "a", "c"
\_[5] [StoredTable] Name: "s", used 2/4 column(s): "b", "d"
```
The physical plan for a single-node cluster can execute the entire query in a single stage:
```shell
\_[1] [Projection] ref_0, ref_1
\_[2] [Aggregate] GroupBy: ["c"] Aggregates: [sum("d")]
\_[3] [Projection] "c", "d"
\_[4] [Join] Mode: Inner [("a" = "b")]
\_[5] [StoredTable] Name: "r", used 2/4 column(s): "a", "c"
\_[6] [StoredTable] Name: "s", used 2/4 column(s): "b", "d"
```
In contrast, the physical plan for a multi-node cluster splits query execution into multiple stages which are connected by shuffle operators:
```shell
[1] [AggregateMerge] GroupBy: ["c"] Aggregates: [summerge("d")]
\_[2] [Shuffle] Hash by ["c"]
\_[3] [AggregateState partial] GroupBy: ["c"] Aggregates: [sum("d")]
\_[4] [Projection] "c", "d"
\_[5] [Join] Mode: Inner [("a" = "b")]
\_[6] [Shuffle] Hash by ["a"]
\_[7] [StoredTable] Name: "r", used 2/4 column(s): "a", "c"
\_[8] [Shuffle] Hash by ["b"]
\_[9] [StoredTable] Name: "s", used 2/4 column(s): "b", "d"
```
The Shuffle operator is the fundamental building block for distributed query processing, and there are many possible designs for it - e.g. \[HBEAM18], \[SAJHJ19], \[MYC20], having different strengths and tradeoffs. Usual considerations are how to make shuffle very scalable and resilient to failures. This makes sense in the context of huge batch jobs, which process petabytes, use hundreds or thousands of machines, and run for hours. These, however, are not the main scenarios for Firebolt - we are focused on data-intensive applications, which need to process queries at low latency. So, we made different tradeoffs - instead of prioritizing scalability and resilience, we prioritized latency and network utilization.
One example of where this tradeoff affects design is - whether shuffle should be pipelined and streaming across stages, overlapping their execution, or whether it should be checkpointing state. The benefit of checkpointing state is that in case of failure, execution can restart from the latest known checkpoint. The drawback is that such checkpoints introduce additional pipeline breakers, and also require additional disk or network I/O to actually write the checkpoints. Since Firebolt optimizes for fast queries, we were less worried about the possibility of entire node failure in the middle of query execution, and therefore chose a pipelined and streaming shuffle. In the unlikely event of node failure during query execution, we can simply restart the respective query from the beginning (with mechanisms in place to preserve idempotence of such restarts, i.e. we would never have INSERT adding same data twice). Note, that our shuffle is resilient against transient failures such as network connection drops.
By being able to overlap different stages of distributed execution, we achieve much lower query latencies, since we can usually hide network I/O in one stage behind processing work in another stage. Again, this strategy works very well for user facing low latency queries, and is not optimal for long running ELT batch jobs - but user facing low latency queries are what Firebolt optimizes for.
When pipelining different operators, it is important not to make writing and reading to shuffle a bottleneck, or the entire pipeline will stall on it. Since most SQL operators are CPU or memory bound, we also needed to make sure that shuffle doesn't use too much CPU or memory while trying to utilize all available network bandwidth. We were able to achieve this by doing all network I/O fully asynchronously using the [io\_uring kernel interface](https://unixism.net/loti/index.html). But it turns out that even with io\_uring, it is not trivial to saturate network bandwidth, when AWS gives you up to 200 Gbit/second. For example, AWS imposes limits on a single TCP connection flow, throttling it at 5-10 Gbit/second, so in order to get higher throughput, we are using dozens of TCP connections between each pair of nodes. Of course, since we are no longer using single TCP connection, we needed to build flow control to maintain ordering of packets over multiple TCP connections. The shuffle code became more complex, but we were able to get very close to full network bandwidth - so for us this tradeoff was worthy.
Another consideration is that 200 Gbit/second is getting close to the memory bandwidth, so we had to be very careful not to do extra copies of data in kernel- or user-space. Finally, shuffle flow control has to include backpressure mechanism to prevent excessive buffering of data on either sender or receiver side.
## DML: Automatic incremental index maintenance [#dml-automatic-incremental-index-maintenance]
The classic definition of Data Warehouse was given by [Bill Inmon](https://en.wikipedia.org/wiki/Bill_Inmon) in his "Building Data Warehouse" book \[Inmon02]
"*A data warehouse is a subject-oriented, integrated, nonvolatile, and time-variant collection of data in support of management's decisions.*" \[Inmon]
The book goes on to explain that "nonvolatile" means that data is loaded into Data Warehouse, but not updated, keeping the history of changes. So one would think that in the perfect world, the only DML that Data Warehouse needs is INSERT. OK, and maybe some simple form of DELETE to remove outdated data. But certainly not UPDATE, right ? Well, our world is not perfect, and there are many use cases that require doing mutations in Data Warehouse. Upstream systems need to issue corrections, laws such as GDPR and "Right to be forgotten" require deleting individual records, and sometimes people get carried away and build applications which treat Data Warehouse more like Operational Data Store (ODS) or even a system of record.
Well, we are meeting our customers where they are - if they need UPDATE/DELETEs - we got them. And just like everything else in Firebolt, we always look at these SQL statements in the context of data-intensive applications. For one, it means that these statements themselves should be fast, especially if they update a small number of records. The technique for doing such small deletes is well known - soft deletes, i.e. keeping an auxiliary data structure which remembers which rows were deleted, it is sometimes called [deletion vector](https://delta.io/blog/2023-07-05-deletion-vectors/), and apply it during reads. Firebolt uses [Roaring Bitmaps](https://roaringbitmap.org/) as deletion vector implementation. Roaring Bitmaps is very compressible, yet fast to access data structure.
Below is an example of production workload in Firebolt, where application issues DELETE statements, each deleting a single value from a big table. The charts below show QPS of about 6 per second, and average latency for these deletes is 140 milliseconds (which includes uploading deletion vectors to S3 for durability).


But while performance of DML is important, what is much more important to Firebolt - is performance of reads in the presence of writes. Firebolt is there to support low latency high concurrency data intensive applications, and the latency should remain low even in the presence of updates.
For example, one technique which enables low latency query serving in Firebolt are [Aggregating Indexes](https://docs.firebolt.io/godocs/Guides/working-with-indexes/using-aggregating-indexes.html). Conceptually they are similar to Materialized Views, but with the restriction that they only support a single GROUP BY operator, but with arbitrary aggregate functions. Since Aggregating Indexes contain precomputed results of GROUP BY, query planner can automatically rewrite queries matching (exact or subset) of GROUP BY keys and aggregate functions to redirect them to the Aggregating Index instead of base table.
What should happen with precomputed results during DML - especially for UPDATE and DELETE? It is a tradeoff between write and read latency. Some systems don't update Materialized Views for every DML, and instead recompute them lazily in the background. But we cannot afford doing this in production workloads, because then once DML transaction commits, the subsequent reads will suffer performance drop. So we update aggregating indexes in the same transaction as DML, and we do it incrementally. There is extensive research literature about incremental maintenance of materialized views, particularly the ones containing aggregate functions - \[GMS93], \[GM99], \[DQ96], \[MQM97]. These papers are kind of complex to read, but the idea is relatively simple. Incremental update is straightforward to do for "good" functions like SUM. We can compute SUM (optionally per group) for all values in the transaction. For INSERT we simply add it as a new delta, and for DELETE we also add it, but as a negative delta. It is more tricky with more difficult functions, like MAX. It can still be added in INSERT, and we can quickly calculate MAX between the existing values and the delta. But what to do in DELETE - how do we know if we removed the maximum value per group (maybe there are ties), and what the new value would be ? It is possible to mark affected groups as invalid, but it means that they will have to be recalculated for every subsequent SELECT. In the spirit of doing more work at write to optimize reads, instead we selectively recompute values in the Aggregating Index, by natively deleting the old group by entries, and inserting a fresh values from the origin table after filtering the deleted values, but only for groups affected by the DELETE.
## Transaction manager: ACID at no cost [#transaction-manager-acid-at-no-cost]
Finally, to be a DBMS, Firebolt needed to support transactions. This is something we couldn't compromise on. Firebolt had to have real ACID transactions, not eventual consistency, not "almost transactional" - but true ACID. There are many ways to build Transaction Manager, but for it to fit the Firebolt mission, we put together the following list of requirements:
* 100% ACID
* Within distributed environment, across multiple compute clusters
* With storage compute separation, where any compute cluster can access any data
* With full support for multiple writers, any node of any compute cluster can modify any data
* Immediate consistency: Once a transaction commits, changes should be immediately visible to all other nodes in all other compute clusters
* Strong Serializable Isolation Level for all metadata operations (DDL, DCL)
* Snapshot Isolation Level for data operations (DML and DQL)
* At scale of analytics, where a single transaction can modify hundreds of terabytes
* But also allow low latency "trickle" DML, i.e. DML updating only a handful of rows few times a second
* Optimize for workloads dominated by reads, with writes being a small fraction of the queries
* Allow long running multi-statement transactions
* Enable transactions to span across multiple databases
And if that was not difficult enough, here is the kicker:
* All the above without performance impact on low latency high concurrency data intensive application workloads
To illustrate the last point - the fastest query that one of our customers is running in production to power their application takes 5 milliseconds. So if the transaction manager's overhead is only 1 millisecond - which would be fantastic achievement given the above requirements, that would still be 20% slowdown, so we strive to reduce transaction manager overhead to zero.
Designing such a system is very challenging, and we are not aware of precedents neither in industry nor in research. One tradeoff that we decided to do very early was to prioritize the "happy" path of the read-only DQL SELECT query. Given that analytical workloads have many more read queries than write queries, we wanted to perform best if there were no changes since the previous query, and we could use all the cached metadata - definition of the schema objects - tables, views, indexes etc; snapshot versions of the data (i.e. list of tablets per table), statistics etc. All we needed was one RPC call to the transaction manager at the start of a read-only transaction to verify that indeed nothing changed, and that cached snapshots are still valid in the new transaction.
Since we were optimizing for read-mostly workloads, but with writes that had to be consistent across all the compute clusters, we decided to persist transaction log (aka Write Ahead Log or WAL) in [FoundationDB](https://www.foundationdb.org/). FoundationDB is a proven datastore which gave us production readiness, robustness, scalability, fault tolerance and transactional semantics. On top of that, it offers extremely low latency even at the P99 tail. So in a sense, we delegated the most difficult part of transaction management to FoundationDB.
In order to be able to quickly interpret whether changes in WAL impact the current query, we built indexes over WAL which are used by the most common requests to Transaction Manager, including one which checks if anything changed. Not only that, these indexes allow to easily compute a delta of changes, and replay only the relevant portion of the log. Having WAL indexes has the drawback of slowing down transaction commits which need to build these indexes, but it was acceptable to us, since we were optimizing for reads.
Other techniques include having asynchronous notifications about WAL changes post commit, so all compute nodes have a chance to see them and update their snapshot (and possibly prefetch data as well) before a user sends the next query.
We also invested heavily in network topology of our deployment both for Transaction Manager and FoundationDB, to ensure that within a regional deployment, we cross availability zone boundaries at most once.
So coming back to the "happy" path of the query: A compute node that handles a user query needs just one RPC to Transaction Manager to check that nothing has changed, and Transaction Manager uses WAL indexes and reads few FDB keys in parallel to confirm that nothing indeed changed. We target 1 millisecond for that interaction. But even 1 ms could be too much for low latency queries. We plan to optimistically start the query execution in the hope that the cache is up to date. In the unlucky case when it's not - the query execution would need to be canceled and restarted with an updated version of the metadata. This part is still work in progress.
Since Transaction Manager stores all of its state in FoundationDB, it is a stateless service, and scales easily. (Of course, FoundationDB being a stateful service also needs to scale well - but FDB is a very mature technology, and it does scale). We have put our Transaction Manager through stress testing, loading it with tens of thousands of QPS both read and write transactions, and it holds up well.
## Conclusion [#conclusion]
In this blog we started with answering the question - what did we build Firebolt for, and why do we think the world needs Cloud Data Warehouse for Data Intensive Applications. Once we articulated our vision, it became very clear what our focus and engineering priorities should be in realizing this vision. We defined our two guiding principles - building a no-compromise analytical DBMS and making it perform the best for interactive, user-facing, low-latency, high-concurrency applications. These principles were guiding us in designing and implementing all components of our system. Like always in engineering, we had to make tradeoffs, and the above principles helped us stay focused on our vision, and make tradeoffs that fit the purpose of our system.
The blog gave a handful of examples of how we navigated these tradeoffs in different components of Firebolt - from the API frontend, to query planner, runtime, distributed execution, storage, and transaction manager. And while we didn't cover many other parts of our system - cloud infrastructure, compute, client SDKs, connectors, security manager, metadata service, control plane etc - they all followed the same principles and decision making process.
We believe that being laser focused on our vision, helped us to build a great product which excels at what it is designed to do. But ultimately - it is up to you - our customers - to judge!
### References [#references]
\[SHMMKT06] U. Srivastava, P. J. Haas, V. Markl, N. Megiddo, M. Kutsch, and T. M. Tran. ISOMER: Consistent Histogram Construction Using Query Feedback. In ICDE, 2006.
\[BCG01] N. Bruno, S. Chaudhuri, and L. Gravano. STHoles: a multidimensional workload-aware histogram. In SIGMOD, 2001.
\[BMHM07] A. Behm, V. Markl, P. Haas, and K. Murthy. Integrating Query-Feedback Based Statistics into Informix Dynamic Server. In: Datenbanksysteme in Business, Technologie und Web (BTW), 2007.
\[PZM20] Y. Park, S. Zhong, and B. Mozafari. QuickSel: Quick Selectivity Learning with Mixture Models. In SIGMOD, 2020.
\[HBEAM18] Zhang, Haoyu and Cho, Brian and Seyfe, Ergin and Ching, Avery and Freedman, Michael J. Riffle: optimized shuffle service for large-scale data analytics. In ACM 2018. [https://doi.org/10.1145/3190508.3190534](https://doi.org/10.1145/3190508.3190534)
\[SAJHJ19] Qiao, Shi and Nicoara, Adrian and Sun, Jin and Friedman, Marc and Patel, Hiren and Ekanayake, Jaliya. Hyper dimension shuffle: efficient data repartition at petabyte scale in SCOPE, In VLDB 2019, [https://dl.acm.org/doi/abs/10.14778/3339490.3339495](https://dl.acm.org/doi/abs/10.14778/3339490.3339495)
\[MYC20] Shen, Min and Zhou, Ye and Singh, Chandni. Magnet: push-based shuffle service for large-scale data processing. In VLDB 2020. [https://dl.acm.org/doi/abs/10.14778/3415478.3415558](https://dl.acm.org/doi/abs/10.14778/3415478.3415558)
\[Inmon02] Inmon W\.H. Building the Data Warehouse, 3rd edn. Wiley, New York, 2002
\[GMS93] A. Gupta, I. S. Mumick, and V. S. Subrahmanian, "Maintaining views incrementally," In Proceedings of ACM SIGMOD Conference, pp. 157-166, 1993.
\[GM99] H. Gupta and IS. Mumick, "Incremental maintenance of aggregate and outerjoin," Technical Report, Stanford University, 1999. [https://www3.cs.stonybrook.edu/\~hgupta/ps/aggr-is.pdf](https://www3.cs.stonybrook.edu/~hgupta/ps/aggr-is.pdf)
\[DQ96] Maintenance Expressions for Views with Aggregation Dallan Quass Stanford University, [http://ilpubs.stanford.edu:8090/183/1/1996-54.pdf](http://ilpubs.stanford.edu:8090/183/1/1996-54.pdf)
\[MQM97] I. S. Mumick, D. Quass, and B. S. Mumick, "Maintenance of Data Cubes and Summary Tables in a Warehouse," In Proceedings of ACM SIGMOD Conference, 1997. [https://dl.acm.org/doi/pdf/10.1145/253260.253277](https://dl.acm.org/doi/pdf/10.1145/253260.253277)
# How Zendesk engineers manage customer-facing data applications (/blog/how-zendesk-engineers-manage-customer-facing-data-applications)
This time on the data engineering show, Eldad abandoned his brother Boaz but it's ok because Boaz got the full 30 minutes to talk to one of the most interesting people in the data space.
Ananth Packkildurai is Principal Software Engineer at Zendesk and runs one of the strongest newsletters in data – Data Engineering Weekly. He talked about data applications at Zendesk and how they're built, technologies that excite him like data lineage and data catalog, and the best routes for software engineers to get their hands dirty in the data world.
INTERVIEWER: Boaz Farkash.
ZENDESK GUEST: Ananth Packkildura - Principal Software Engineer.
Listen on [Spotify](https://open.spotify.com/episode/3VdCM7JyQWDtkvRr0TagVw?si=K9c4Nd5dRAKD0Im9YRgh7A) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-zendesk-engineers-manage-customer-facing-data-applications/id1561927688?i=1000551399658)
**Boaz:** Welcome everybody! Thank you for joining us in another episode of the Data Engineering Show. Today, sitting next to me is nobody. Eldad, my co-host and brother, disappointed me. He could not make it today, and he trusted me. Ananth is here with us. Help me welcome Ananth Packkildura, who is a Principal Software Engineer at Zendesk, used to work for 4 years prior at Slack. Has a bunch of very interesting stories for us today, from both. Ananth is also somewhat of a veteran in the industry, has been in software for a long time, and has moved to data for many years. I would like to talk about that as well. But, Ananth is also a mini-celebrity in the data space. He runs the data engineering weekly, newsletter. Tell us about that. How many subscribers do you have today?
**Ananth:** Okay. Great! First-of-all, thank you so much for having me. It is really amazing to talk about data all the time.
**Boaz:** Thank you for joining.
**Ananth:** Subscribing, I think, we just crossed over 1700 or so, I think we are around that mark.
**Boaz:** Nice. If you are not registered to Ananth's newsletter, the data engineering weekly, you have to. If you are in data, you have to. It is a must-read. I follow it. Everything you need to know is in there, so look it up. I mean you are by yourself or do you have a lot of people involved there? How does it work?
**Ananth:** Thank you so much for the kind words, first of all. No, it is all on my own. I started, the background of the story, why I started? Towards the end of my time at Slack, I started working in our observability monitoring infrastructure, essentially applying data principles into the monitoring stack. I started to kind of feel of missing out and there was a pretty good data engineering weekly newsletter before that and got stopped and that was my first go-to source of information before everything, and that is how my morning starts. I had a special alert to read that newsletter and then, I also started to feel that I am going to miss out on something, so I started like, okay! I am going to read something and I just started writing. So, it is out of my own learning purpose, and I still consider this as a good way to learn. I think I would encourage everyone to write most of what they learn and that is how you are structurally learning by yourself. So, it is really, really amazing. Anyone listening, please start your own data engineering newsletter.
**Boaz:** But I am sure it has been getting tougher and tougher because there is much more content to create from out there. I think it is time you find help with that newsletter. It is too much on your shoulder.
**Ananth:** Yeah, that is true. Last week, I was going to curate. I had at least 18 articles to curate and shortlisted. Now, I do not have time to write all those things, shortlisted 10 items or 9 items to do that, and if anyone viewing that we already have data engineering weekly, GitHub link, that, anyone read some article, they like it, you can create a pull request and then we can add that as part of the newsletter in this case. So, there is also a way for more community contribution to the data engineering newsletter.
**Boaz:** Nice. Awesome! Ananth, I would love to hear about your personal career. You started from software and over time, by me looking at your history, it seems that data crept in slowly, slowly, until eventually, you could say, took over. So walk us through that. So when did data become such an important part of your career? What was the tipping point?
**Ananth:** Yeah, I think I started my career as backend engineering, mostly writing code on Java and other stuff like that and it is funny that when I was starting a career because data warehouses were not that much kind of winded off. In most organizations, there is very less visibility on those things, and it is often viewed as you will be using some tools, maybe SSIS package or any of the tools, I will never go to the that saving my time, but that also is a thought beginning at the career. I would say I got carried away and pulled into this whole big data wave or the buzzwords kind of thing, so I started looking at developing systems in Hadoop as part of experimentation in one of my works. I started looking at Hadoop 0.2 at that time or something like that. It was kind of a very early stage and that is the first time I am actually reading a lot about not producing better than when I kind of realized that most of the problems can be solved if you started to think not to produce patterns and most of the analytics problems. So I think that is kind of a very pulling point, and at that point still, the work is more of a backend engineering because you still have to write a bunch of Java code in order to kind of build any kind of analytical ecosystem at the time. I think I just jumped onto the bandwagon of that hype and then, slowly travelled and realized, okay! This is what we are going back to, where I just kind of reinvented the wheel, and then I actually went back and learned about the data warehouse concept and I took a course from the Kimble group. I think that was the last training program he took, so it is very fascinating to learn. Then, I go back and then learn the basics, and then you just continue on those cases.
**Boaz:** I think it is interesting, you mentioned that SSIS part and if we are to be completely honest, I think in the past, software engineers used to, to some extent, even look down at data technologists, thinking it is a different position. That is not for us, you deal with your data warehouse. Whereas now, it has become such a big deal, and most of our engineers are actually looking to be more involved with data technologies. It has become one of the most interesting parts of software engineering in general. Thanks for sharing!
Let us start with Zendesk. You have been there for a year and a half. What did you come to do there? or what do you do at Zendesk?
**Ananth:** Yeah, Zendesk is a very interesting company, right now, as a leader of customer experience platforms, and most of the support platforms essentially. What I am trying to do right now, is involve our customer-facing analytical platform, which is, Zendesk as a company has grown by acquiring more companies. It is like a disparate system involved, but what our clients see or want to see is a unified view of how the customer interacts with their sales and then as a support secret system. Also, not only do they want to view what is the customer experience, but they also want to see more deeper level. My product catalog is there. I wanted to send it to them and tell me which product is doing good or not. They want more than a simple support ticket analytics in this case. So my primary goal - I started in Zendesk to focus on building a data platform and analytics solution to address that particular problem in this case, and as with any data infrastructure, when you start to solve the last mile problem, you realize the problem is actually existing ahead of the problem, like, How do we make sure that we instrument data properly? How can we enable scalable analytics systems on top of it? Right now, I am playing a bridge role to make sure that we are gathering sufficient data, more domain ownership around, and then building scalable solutions.
**Boaz:** Zendesk has been around for many years. I, myself, have been a client of Zendesk for years. For a company that has been around for years, I wonder how does the data stack looks like? How much is legacy versus modernized? How do you go about modernizing a stack and how has it evolved throughout the years?
**Ananth:** Yeah, to my surprise. I would say that Zendesk has some kind of a legacy. We have an analytical system on Mongo DB that is actually still so. There is some system there, but surprisingly most of the part, it is pretty much up-to-date technologies. We have some systems using Google BigQuery. All our enterprise analytics is running on Google BigQuery and then cloud storage in this case and we recently started to adopt Apache Hoodie, which is kind of building this use case.
**Boaz:** Are you guys running multi-cloud or is everything on GCP?
**Ananth:** Yeah! So our enterprise analytics is actually running on Google cloud. Our product analytics running on AWS, various different companies Zendesk had products at some point of a time. That we could call it a legacy. We have two different cloud services for sure.
**Boaz:** Walk us through some of the more challenging use cases, workloads that are currently in place that you are involved with? How far long is that new product you described?
**Ananth:** For unifying the system.
**Boaz:** Yes.
**Ananth:** We are just barely scratching the surface. The way I am looking at this problem, not necessarily from unify the cloud services perspective, but how do we do data sharing between these two different disparate services, I think our challenge or compliant from all stakeholders is not necessarily we are running two different cloud services because two systems are atomic in nature, and then they just doing a fine job. I think the problem is at - if I do the data sharing from AWS to Google cloud, Is there any context that we are missing? Is there any lineage that we are missing? How does the consumer on the other side trust whatever you are sending? Right? because the business logic exists in one part and you just send the derived data set to another cloud and the context is missing in the middle range. So we kick-started this whole data lineage, data catalog project trying to possibly build the full bridge, establishing the full context of it, and trying to give more and more understanding to the consumer side of it to figure it out, and I hope at some point of a time, we will be able to merge these cloud services, but that is like a large project one to take.
**Boaz:** That was my next question. Has consolidating into one cloud been on the table?
**Ananth:** No, not now, at least. It works, but I think adding additional context right now will give us much more visibility to what is going on.
**Boaz:** Tell us how the data-related teams are structured at Zendesk? What kind of teams are there? How big are they? How are the responsibilities split?
**Ananth:** Yeah, that's a good question. I think there is a common pattern I started to see even in Slack and then Zendesk, it is like, we have this product analytics team that is focusing on instrumenting data from our product usages, collecting the data, and then building those, mostly they end up using kind of a Lakehouse architecture. And there is a whole bunch of business operation analytics, sales analytics, and marketing analytics, and these analytics teams have self-contained data engineers and data platforms inside to support the business operational aspect of it. Other teams focus, most of the cases, on the SAS application where you have to deliver customer-facing analytics on top of it. So that requires more and more coding kinds of things other than the SQL kind of workload. So they have a separate team. Zendesk data teams are also organized in these 3 different business orientation aspects of it.
**Boaz:** What data volumes are you guys dealing with?
**Ananth:** I do not have a top of mind, but I think we all log everything all the time. So I do not have a very finite number because it is not like one stream of data that we have been consuming.
**Boaz:** Yeah! Let us say customer-facing stuff, for example. How is that managed? Are these dev teams at the end with customized UI running the show, applying it into APIs, to run queries? or Is it more of an embedded analytics solution? What is going on there on the customer-facing workloads?
**Ananth:** The way we look at customer-facing analytics, it is more of an application. It is kind of a product on its own. We have a product manager to see how fine at SLA, which will be applied there. We source information from our customers, and then we can build those data pipelines over that period of time. That is a good question. Largely, the customer-facing analytics encompasses more of a backend engineering rather than a pure data engineering perspective. I am just slowly introducing to them the data pipelining concept, but very backend engineering focus in this case.
**Boaz:** What is the query engine behind the scenes?
**Ananth:** It is interesting, we right now have two query engines, quick databases, like we are running Redshift and then Postgres, for some historical reason that we keep running on those cases. There is a large project that is going on right now, to kind of unify the data store both in real-time and batch infrastructure to make it more customer-facing analytics. So we are looking at that particular solution.
**Boaz:** Yeah. I think trying to combine historical and real-time is something humans cannot stop trying to question. What is the solution? Any selected approaches already?
**Ananth:** I mean, we are looking at various solutions right now, potential contenders like Druid or ClickHouse or Pinot is one of the exciting ones.
**Boaz:** Which is the last one?
**Ananth:** Pinot.
**Boaz:** Pinot, yeah
**Ananth:** It is kind of a really exciting one. It is a question of right now, what we are doing in the real-time, we do join on the streamside a lot. So I am debating back and forth, join in the stream versus join in the database. I like joining database now because that is what the database is supposed to do. We can always do optimization on the stream. Like why do you want it to wheel on your stream processing? enrichment, yes! But do you need to build up a full flown join system? So those are the interesting concepts that we are exploring right now and dignify our system. Hopefully, we will do something.
**Boaz:** As you mentioned you are looking at Pinot as well. I ran into PC road from your Slack days about Pinot, you adopted that over there as well. Was it over Druid or something? Tell us about that project back then.
**Ananth:** Yeah, Druid vs Pinot. These are some of the things that might have been not relevant because every system always improves at any point in time. I think most of the real-time systems coming out of the use case were doing an ad serving engine or having ad serving capabilities. I think we were kind of an ad engine solution, and ClickHouse to an extent the same thing. The nature of the system is mostly immutable in nature. Even though, always immutable in nature, there is no upsert operation that you wanted to do. So you just see the event and you are just continuously running some time series over a period of time. I think that kind of solution is going to work really well, but in companies like Zendesk, companies like Slack where the SAS application fundamentally tries to solve the workflow over a period of time, right? So tickets have been created. Tickets have been deleted and the ticket can go through its full lifecycle, and we are trying to capture an object and we are trying to produce insight for an object lifecycle. So, we wanted to produce the current state of the object and then we also wanted to kind of do a historical view of how the data transformation goes. So when we wanted to preserve the current state of an object, that is where the upsert operation became absolutely crucial, and Pinot does support upsert operation and work reasonably well, but it is still kind of an afterthought, because, Pinot added upsert operations to solve the use case for Uber eats, which is a similar business process engine. I think maybe that is why we are not able to bridge that real-time and batch essentially building some system in real-time from the ground up, that supports ability and immutability. That could be one reason, I do not know.
**Boaz:** Yeah, this is a tough nut to crack.
**Ananth:** Yeah.
**Boaz:** You were in Slack back in 2016 and you were there for 4 years or so. How different was the data stack at Slack when you joined versus when you left? Because I know you, you pretty much helped build it from the ground up.
**Ananth:** At least at that point when I left, it has predominantly remained the same. One good thing, I do not know if it is a good thing or a bad thing, the early people started to involve in the Slack data infrastructure came from very good previous experience building data infrastructure at scale, and from the good-to-go, we choose some tools that are kind of high on the support side. Just Kafka, Presto and storing everything in the pocket format, tried to use Airflow programmatically to author our system and so on, how do you do modelling and then using structured logging, using a script for logging, not click. So, these things are, we got it right. We spend less time on detecting whether this is an integer or a stream. That kind of a problem I see more and more companies trying to do. Anyone approaching me asking, like, how do I build a data platform? I would say firsthand, please do not use JSON as your data format in this case. So, these are the things we got really well, programmatically author or data pipeline, structured even logging and stuff like that.
**Boaz:** Because of this JSON, that was an interesting comment, because all the new data platforms essentially are encouraging people to use JSON. Everybody is releasing JSON first features, JSON capabilities, so you are saying stay away from that! be wary!
**Ananth:** Yeah. Maybe, I do not know. Maybe, they start thinking of a Schema registry to solve this problem. That could be one reason. But again, the compiler is good at type checking and where do you want it to type-check after the fact. When you are doing a compilation itself, it is better to do that, you can reduce a lot of errors in this case. I think most of the Slack challenge at the point of the time was scalability because the rate at the company had grown, every assumption that we have made in every three months we just invalidate at some point of time we made over three months. So, we have to constantly reinvent or try to scale our system to kind of cope with the amount of volume that we are getting. When I started in Slack, we were ingesting maybe 10K events per second and that is what the Kafka cluster looks like. Then rapidly, nine months down the line, we were just ingesting 3 million to 4 million events per second and we barely supported only one type of use case, on the business side of it. We want the operational side of it and all the other aspects of it. So scalability is a bigger challenge in this case.
**Boaz:** So maybe it sounds like if there ever was a scalability challenge, it is this one at Slack.
**Ananth:** Yeah.
**Boaz:** I am sure not everything went smoothly. This is a good part of our beloved epic failure corner. We grow and learn from things that did not work. Any memories to share things that are a good lesson learned?
**Ananth:** Yeah. I think many things, especially, I do not know how the lesson was learned. I think the Airflow page we had one time, has stopped us from running any code for 8 hours and a lot of small accidental problems. I think this is a very interesting one. Airflow at the time had, I do not know whether it is even now it is true, the scheduler has a problem that it is not able to schedule the rate of spinning off. What we have done is we introduced a Cron job that keeps everyone in our background, just restart the scheduler, just in case the scheduler is exhausted, we just restart in an hour but what we did not realize that that Cron job had a bug. Instead of waking up at the top of the hour and then just restarting one time, it is restarting for one minute continuously. We did not realize that. It was just working fine because it is a very minimal in difference, and one day, we started migrating Airflow from one version to another version and we did not realize this bug when we are iterating that and we are continuously monitoring whether the upgrade is successful or not, we keep seeing this weird behavior and we were so confused, what is happening? Read through all the source code, read through all our deployments. Nothing is happening. What is really happening? That took us like one day to figure it out. Oh my God! there is a Cron bug and we almost forgot that we put that Cron job.
**Boaz:** The last suspect that nobody even bothered to think about was the one. Thanks for sharing.
**Ananth:** The good news is that now I am able to read all the Airflow coding in just one hour because we had to read it and understand.
**Boaz:** If you go back in time, we put you back in 2016 at Slack, starting from scratch, knowing what you know now, what have you done differently?
**Ananth:** I think one thing I would have done differently towards the end of my time at Slack. I think people started to lean more towards cloud data warehouses, 2016 cloud data warehouses might not be much more mature or much more sufficient to handle our scale. I think at this point of a time, I feel like systems like Firebird, Snowflake or something like that could have been our first choice to do that and then more using the tools rather than trying to build everything else that would really give us much more velocity in the way we want it to handle that thing, because we had to manage all our EMR cluster and because we had to manage all of our Presto cluster and Airflow cluster, typical out of toll for us to kind of support the growth that our company kind of going through and that reduced the trust level of the velocity more and we put people adopting to.
**Boaz:** In recent years, which technologies or tools or both did you put your hands on, in recent years, or ones that excited you or that you are excited about?
**Ananth:** In recent years, I think I am excited about a lot of development going on, the data lineage and data catalog. I think this is something that we have not thought through when we started our data team at that point of time. And most of the companies, the data catalog, and data lineage will always be an afterthought. I am so excited to see so much literature, so much talk about lineage data, data discovery systems, and then the data quality aspect of it. I think that is one of the things that I am very excited about. I think adopting that we finally acknowledged we came a long way from Hadoop world to be kind of a heavy hacky backend engineering to acknowledge this is a data system. This is a data management system, properties of the data management system and I think we are coming through the full cycle.
**Boaz:** Which tools have you been looking at so far at Zendesk with lineage, logging, and quality?
**Ananth:** There is pretty good tooling available right now. I think it is a good thing a lot of options are available for the consumers. Now I think datahub from LinkedIn is one of the examples of that. Amundsen is another great. I think a Clan is another interesting tool to do that. I think what I am really looking at is how this tool essentially embeds into our workflow of our analysis that the data engineers integrated into their workflow rather than introducing a data catalog, and then just do a checklist and we have a data discovery and then checkmark is done and then nothing. I am sure you are aware of Apache Atlas, which is an open-source data lineage story developed long back in 2012 or something like that. We do have an Apache Atlas from 2017 in Zendesk, except that no one uses it or half of them are not even aware of it. The thing is that the pattern that I noticed why people are not able to use it, is it is incredibly hard
first-of-all to do that. So the analysts, I observed the behaviors, and again, one interesting thing is an analyst came and said, "Hey! What does this table look like? and I wanted to get an insight. Tell me what are the tables that I should query about that?" And someone trying to point out, well, this is a lineage tool that you should use.
And it is a separate effort. They had to go to a different UI. They had to understand that UI and then they are just trying to query, that it can have information, do not have any information, and then if they are not able to find that information, they immediately go and ask, who is the senior of the team? and they will go and ask the same question, right? You will be using an analysis of the data discovery system again and asking it and there is a high chance that a person has the knowledge about it because of an experience and they answer the question. When the next time a similar problem comes in, the analyst will not go and look at the data discovery system. They directly go because their workflow is now changed because this particular tool is not trying to solve this problem. So, I look at this as a workflow problem rather than a lineage problem. So they need to sufficiently solve that problem, which will be really helpful.
**Boaz:** Awesome! What do you recommend if I am in software, but I have not been involved with data so much and I want to start? What is a good route for a software engineer to get his hands dirty in the data world? How do you approach such an educational program?
**Ananth:** My take on that is, SQL by default is the easy language of data, right? The first thing in terms of the skill set, I would like to say, is to pick up SQL, it is a first-class tool for you to start navigating the data across. And if you are a software engineer, I think one of the good things they can do is take a look inward rather than outward. Many people started taking, wanting to get into data engineering and looking at - I am going to solve a business problem or predicting, sales analytics or marketing analytics. If you are a software engineer, there is most likely that you will be working with some kind of a system, some kind of software you will be deploying to a production system. So that the production system will emit some kind of logs. I think the very first step is to take that log and apply the SQL and try to understand your system from a different perspective.
**Boaz:** Interesting.
**Ananth:** We do have the observability tooling right now, metrics and logs and you do search and try to find the needle in the haystack. But if you try to take that logs and try to understand the long-term perspective, it is particularly helpful, A, because you already know the domain, you already know the system, so you have a very good intuitive scale that you can start to work around and improve your skill and then, once you understood that, the techniques and the domain is transferable, right? The domain is more of acquiring knowledge, the tools, and your way of thinking to figure it out, the pattern is going to remain the same. I think that would be my suggestion - Start with the tool, pick SQL and then look inward in your own domain, trying to figure it out.
**Boaz:** It is beautiful. So, a log-driven approach. If you master the logs, you know everything about your platform. Awesome!
**Boaz:** Okay, great, Ananth! It has been amazing, with so many great insights. Thank you so much for joining and again, keep doing what you do with the newsletter. We love it!
**Ananth:** Awesome!
**Boaz:** Stay safe with the virus and all.
**Ananth:** Yeah. Thank you so much.
**Boaz:** My pleasure.
**Ananth:** Bye.
**Boaz:** Bye.
# How ZipRecruiter and Yotpo power self-service data platforms that work (/blog/how-ziprecruiter-and-yotpo-power-self-service-data-platforms-that-work)
Data engineers are not paid to do support. Liran Yogev, Director of Engineering at ZipRecruiter, and Doron Porat, Director of Infrastructure at Yotpo talk about building resilient self-service products that keep customers happy and engineers calm.
They walked the bros through their data stacks and explained how ZipRecruiter is completely rebuilding its data layer from scratch.
Listen on [Spotify](https://open.spotify.com/episode/076y8o8AYFgVl13r4agedU) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-ziprecruiter-and-yotpo-power-self-service-data/id1561927688?i=1000605558364)
Benjamin: Hi, and welcome back, everyone to The Data Engineering Show. It's super cool. Today, we have another set of experienced podcasters joining us, Doron and Liran.
Liran: I love it.
Benjamin: Thank you.
Liran: You tried really hard. It's just...
Benjamin: I speak German, that's my native language, so we have some hours in there as well. So I practiced really hard before the episode.
Doron: Sounds good.
Benjamin: Awesome.
Doron is a director of infrastructure at Yotpo. I hope I pronounced that correctly as well.
Doron: Perfect. You're the only person in the world that can pronounce this.
Benjamin: Awesome. Liran is a director of engineering at ZipRecruiter and also was at Yotpo before. Do you guys just kind of want to give a brief intro tell us what you're up to, tell us about your podcast, and then we'll dive right in?
Doron: Yeah. Go ahead. I'll start. Yeah. I'll start with me, and then we'll go to the podcast.
Liran: Born in the Yavne.
Doron: Yeah. I was born in Yavne.
Liran: 47 years ago.
Doron: 1923. So, I worked at the Yotpo for 25 years, just a lot longer time.
Eldad: Was born at floor four.
Doron: Yeah. We started at floor number one and we reached up to floor number 26, right above a Firebolt. But, I worked at the Yotpo for a long time. I started as a team leader for the data engineering team. Later on, we became an infrastructure group and the team grew. I became the group leader for the data infra. We built an amazing data platform here together, Liran and I.
He was actually my predecessor. He was the infrastructure group manager, and I replaced him recently.
Liran: Doing a much better job.
Doron: Yes. So I do everything better. Now, everything is better here.
Liran: Less people, right?
Doron: With less people. Yeah. We're running very lean.
Eldad: Benjamin, don't get any ideas in your head.
Doron: No. It works only if you're a woman. We have a joined child. We don't have a child.
Eldad: Oh my God.
Doron: We have a podcast that we've been podcasting for about a year and a half. Also about, data engineering and all the surrounding world.
Liran: But it's in Hebrew.
Doron: So it's in Hebrew. So we're not really competing. It's not the same.
Benjamin: I tried preparing for the show and Tamar said your podcast, and then I was shit.
Liran: So, there is one episode in English, which...
Eldad: So he learned Hebrew.
Liran: There you go.
Eldad: Hebrew. Then you didn't hear the second half. Yeah, he's still...
Liran: If you even started, it's so fast,
Eldad: It's too modest.
Liran: How fast we talk in the episode, so I don't think it's good for you as, as the beginner, wave to learn Hebrew.
Benjamin: Okay.
Liran: Sorry. With friends or something in Hebrew. I do not if there is a version.
Now me. I'm kind of always around the one, so all the stories intersect. I was at Yotpo before that, I did something else. I've created the data platform there, but also I did a lot of other different roles, in my end position I was running, all of the platform engineerings at Yotpo, so backend data, and front end which Doron is doing right now.
Then about eight months ago, I moved to ZipRecruiter. I'm doing the same there, but a bit different teams. So I also run the platform engineering, or you call it enablement and around the area of ML experience, we call it, or ML tools, and all around ML data. And also experimentation, which is a team that is, building an experimentation platform for internally for our organization. So doing heavy testing and measuring everything. So, it's really been fun. And, the podcast.
Eldad: Data sets. Data sets, data sets. More output. More output.
Benjamin: Awesome. For listeners who are kind of have never heard about Yotpo or ZipRecruiter before, do you want to give us a quick, high-level overview of what product the companies actually building or different companies?
Doron: Yeah, sure. So, I think it's funny because I've been in the Yotpo for a while, so we kind of started off as a branding ourselves, as a marketing e-commerce platform and we really became a real platform only in the past few years where we offer a set of different products under the same platform to help e-commerce businesses just, just do better, bit bigger and stuff that.
But recently, I think given the latest changes in the ecosystem and the financial macro environment, Yotpo is trending towards being this retention platform for e-commerce businesses and we do this with the same set of tools, but we're really focused on how we can help e-commerce businesses, online businesses preserve their customers and enlarge their customer base through our products, which are review solutions, loyalty programs, referral programs, communication channels and customer data platform, and more and more products this.
Liran: ZipRecruiter is a Hiring marketplace, I think that's what you call it, basically helps both job seekers and employers find matches and we do that for customers from really small mom-and-pop shops up until customer Amazon. We do have different approaches for each of those customers, from enterprises to small businesses and we heavily rely on AI to do the matching and other things in science systems. So, that's our forte. Helping them really find good matches for both sides. And also balancing the market itself just completely. So, that's the gist of it. And we're based both in Israel and in the US.
Benjamin: Awesome. Cool. Nice. This is the data engineering show, and you guys obviously have a bunch of experience in data at a variety of companies. Take us through the types of data challenges you guys have in your day-to-day and maybe tell us a bit about your stack.
Doron: Okay, cool. So, maybe I'll start with the stack and then I'll go on to the challenges. I think that might make more sense.
Stack, I'll start from the bottom up. That's comfortable for me. So we're ingesting data into the data platform from all sorts of different data sources, whether it's operational databases and third parties event data, whatever and it's basically all streaming into the data lake. So, the whole solution or platform is built around the data lake and it's very data lake centric. Based on AWS. That's all the workloads running there.
Then in the dat lake storing data, different formats, and using different techniques for transforming the data and different engines, mostly Spark these days. I think we're running a long time with Spark, what Liran and I did together in Yotpo for many years is making Spark...
Eldad: Buy a perpetual license, gives you the ability to use it for free on unlimited resources forever.
Liran: Yes.
Eldad: Sorry, go ahead.
Doron: Our challenge years ago was how to make big data tooling available for the generalist developer and later on also for BI developers and stuff that, and that was a big challenge and how to democratize data sets as well as data tooling and we used to do this... I think we had a different approach for this a few years ago, which changed and evolved over time as the platform grew older, and the company also evolved and had different needs and requirements.
Orchestration, we're using Airflow, most of the past pipelines are running using this framework that we built internally in Yotpo. It's also open source. It's called Metorikku where we write YAML files with SQL statements to describe the data pipelines.
Also, with streaming, mostly Spark structure streaming, also Flink pipelines recently. Then, the whole analytics area things.
Plus we have DBT that we started using at Yotpo for the past year or so, well over a year. But we built this whole framework around DBT to visualize the new way of thinking about how we should manage data in the organization. It's also an internal tool that we built and it really connects data producers and data consumers on the other end where we have Looker which we use for internal analytics or external B2B analytics, either embedded dashboards or API, which is also, something we are leaning towards more and more as time goes by. Maybe add kind of analytics is also a big part of the thing, making data available for everyone. So everyone are using Databricks clusters to work on top of the data lake. And it means engineers, BI, analysts, support engineers, solution engineers, and everyone working on top of the data lake.
Yeah, I think that that's the big picture.
And if you ask about challenges, I think that in the podcast, we talk a lot with people and I think a lot of the times we go and talk to people that have big data challenges, it's still a thing. I mean, you think that it's solved already, but people have big, big data challenges.
I think at Yotpo, it's more what keeps me busy. It's more data manageability and how to architect this thing to work well at scale, serving a lot of people, here we have an R and D of 250-260 people and more and more outer circles using the data, and how we optimize this huge machine of money into something that's much more coherent, robust, scalable, and resilient over time.
Data manageability, I think, a wider term, because I can talk for hours about this. I would say that's where I am focused at the moment.
Eldad: Quick question, seems data and everything you do with data is part of the feedback loop within engineering, within product building, , everything you do as engineering and product is using data, how is that translated into user experience, the data? and a lot of what you've mentioned is internal, right? So, it's for building products, how does that get translated into Yotpo's business, for example?
Doron: I think Yotpo is not per se a data product. I think that ZipRecruiter maybe is, is more an example of how is data, centralized and within the product. Data is a big thing, and when I talk about experience, I mostly talk about, and that's what my group is focused on, we talked about front and backend and data, but we are very, very focused on developer experience. I think that where our customers, B2B customers, make the data is, the places where we make it .
Well, I'm not talking about all the machine learning, data science part of things, it's not really under my responsibility. But, analytics is becoming bigger and bigger and I think it's also a part of what's going on, in the world that we are really focused on observability and demonstrating ROI and helping them make the right decisions on. Yopto is a complicated machine and part of it is, I think it's almost actional BI, but external. It's the way that we organize the data in a way that helps them understand how to navigate through the different products to bring more value. So, I think that's where they touch on data the most.
Liran: So I want to add that. I think that what is happening is what we went through in last couple of years is more and more and more use cases were added to the world of the data lake. First, what type of consumers we have for data? So we are mostly being focused on internal, but also, external as well. The data meet the customers in both companies. You don't have to be a data company to have the data reached somehow into production systems.
So, there are more and more use cases just being added. More and more types of consumers. They require different things and their experience or even their capabilities are different. How can they access data? How can they produce data that is high quality? So I think that's what I'm focused on and what is always changing in our ecosystem.
I can but I do not know if we care about that, but our ZipRecruiter stack or not or we just moved on. It's okay. Move on.
Eldad: So, what you're saying is if your data platform shuts down, then internally engineering product business won't be able to operate. It's so much embedded. It's so much interwined. It's as you're saying, self-serve. What is self-serve? Is it opening the data for as many as possible? From your experience, is that real? Is that kind of something...?
Liran: Yeah. So we try to build a decoupled system, right? I don't think it's great if the big data platform, which again, both of our companies is very data lake eccentric. I don't think that if this drops, as soon as it's down, then everything goes down.
I don't want to be at that position. So we need to have some kind of differentiation between all of these backend processes, batch processing, even stream processing that happens, managing both for ML or Funnelytics in the production systems, which actually needs to serve something to the customer, they actually see.
So again, the business will suffer, maybe some late data will arrive to the system. It may be visible to the users, but in the end, we don't want to be, where something actually is not working anymore in production. I think in my opinion.
Eldad: I've seen many peoples saving data sets eventually after they use all the stack, all the tools, they save it in Excel. So, it's a backup, but it's also a failover mechanism and eventually, it all ends up in some report. So, we hear it all the time and I think companies went all in on data, you'll be surprised how dependent they are on internal data.
I'm not talking about external. External is easy to justify, but justifying internal, asking how do I optimize internally? What do I optimize for? Those are new questions we're hearing more about.
Liran: It's even more than that. We need to ask a question, do we need all this data? And it's a question I don't think a lot of companies are asking.
Doron: No, no, we don't need it.
Liran: We don't need all this. Yeah.
Liran: Because Doron and I have been to a lot of discussions about retention periods, for example, for our big data. At some point, just delete it. What will happen? Well, we'll be fine, and I think in a lot of... I'm not actually saying that.
The discussion needs to be made about because at some point it reaches such a high complexity in cost and so many moving parts that you need to ask yourself, do I really need all this? And I think that's something that each company needs to always reiterate on, and ask these questions. I think that's a good culture.
Doron: I would to add something you were talking about questions that we ask ourselves, and I can say that after many years working in data, it sounds dorky.
But I find myself asking myself different questions as time goes by, my concerns shift and I ask myself, recently we started talking about how we should better structure the infrastructure group. And then I started asking myself, where do the lines cross? where does data start and ends when there's backend infra starts and ends, and we have all those interfaces between them. And I think it's a really fascinating question to ask. And it's also in the way that how data stack connects to the APIs or event-based architecture and where do the lines cross. I think it's very, very interesting.
Liran: Yeah. I think we're actually seeing the world move a lot. It's becoming more an engineering world than it was before. It's less and less about just being data. So data is coming from somewhere. It has been produced by someone. It's even been managed by someone that's probably in the product or engineering world.
And I think before we used to have these silos where we had analytics teams just hand-managed data and I think that we're more mature companies. Question is, can they really manage it? Do they really have enough information? Are really taking the responsibility off of the engineering teams?
I think it's because we just talked about boundaries, I think those are also changing all the time.
Eldad: Crazy times, huh?
Liran: Oh yeah. We love it. It's great.
Benjamin: Nice. Maybe take us at how these boundaries look at, the specific companies you are working for? In a sense, you're providing core data infrastructure for your company. Say I'm a neighboring engineering team, I'm trying to build this new, I don't know, data application or internal data experience, whatever, where do I interface with your teams and what are the types of services in a sense you provide them?
Doron: Do you want to start?
Liran: Yeah. We are trying to build the methodologies, processes, and tools around producing and consuming data by all those types of customers you just talked about. That's where we are. And that means that in most cases, when you build the data application or some data experience, then you'll be interacting one of the tools that we have. So it's either going to be something that we bought and we basically implemented and integrated or something that we built that's very specific for just an organization. And what we like to do is optimize that all the time. So that's what we do
So we figure out, okay, we have this customer and we want them to have, to create the best data set or the best data application, so it's going to be the highest quality. It's going to be really fast to create, it's going to be really easy to consume for consumers. So how do we get there? And that's where we add all the different layers of the tools that we either build or buy. So, I think that's kind of where I see the interaction, specifically around the data. Do you want to add some?
Doron: Well for us, first of all, because we talked about teams being really lean, but we're not kidding the teams are very, very small and we support a lot of people, given services. So, we are really focused on building a self-service data experience. I'm going to borrow your words because I like it, but it's all about the experience and how we make this experience better, and help the developers and engineers be more self-sufficient and free to operate within their domains.
It's always maintaining this balance between allowing them to run freely. I'm trying to find the nice word, destroying our vision for how the...
Liran: And also, taking care of our people. We don't want to be just giving out support every day. We are product teams, in a way that we are actually creating internal products, just the Firebolt, for example, does for its customers. But in the end, we both are lean. So, we cannot do support. I mean, no one pays us to do support all the time, and I think it will be really a waste of our time. Our engineers want to build things and not have to support them. So self-service is a really big thing and providing with our customers, with the tool to debug the systems, to run it, to own it, to have...
Doron: And to enforce best practices as well in a way that will not ruin their lives and experience, because it can be a real drag to being blocked in CI with every step that you try. So, it's a matter of how to enforce these in a smart, elegant way, and through this, create the things that you believe in and you think are instrumental for the data platform.
Benjamin: Got you. So going back to this previous example, of retention periods. The day I'm an engineer, I want to use the awesome data utilities you and your teams provide and I said, wow, I really need a 10-year retention period here. At what point does someone question this choice? So it's this part of the core education you guys are doing inside of the companies or is this ultimately up to the consumers?
Liran: I think it's both. We always have a choice between creating validation in CI/CD or creating whatever it is you do. Adding some rules on top of it. We're getting some alerts and monitoring, so we can do everything and we can also do education. We can also go make sure that we have really good guides, and we do sessions with everyone and talk to them about what is actually happening.
Eldad: And it depends on the weather. It depends on how they feel on that certain day. So for example, if they're angry on that, there will be more about consolidation, and measuring best practices applied. When they're happy, then it's about self-serve, pushing more data tools, it's a never-ending cycle.
But the truth is even myself, you constantly ask, does it need to be centralized? Does it need to be decoupled? all of those complicated terms. The truth is you need a team that knows better and that team knows better most of the time, not all of the time. And that team learns faster because they're domain experts. And if they're actually capable of translating that into "best practices", which are amazing, if they work, then they make everyone better. Because the truth is most teams don't have a DNA for data. And the truth is that if you look at data, how it's being applied today versus a year ago, then most teams are not even close to having a DNA for data.
So, I think, The IT Crowd, I do not know, maybe some of you know the show the British, that is how it all started, right? It was all about conflict and they think teams that win with data, they get addicted and they start depending on experts and domain experts like you.
First, we salute you because I think the reason we asked it, it's because it's hard for those teams and in those times, it's even harder because they're now inbound, outbound, they do a lot of stuff.
So, first, we salute you, second, you are important, and third, if you don't deliver, then yeah, the business is shut down in so many ways that you can even imagine.
Liran: I think it close down and there are issues going on. We just help protect and make sure that people do good things and not cause...
Doron: No, I like Eldad's version. I feel very important.
Liran: Okay, we save the world.
Doron: I feel very important.
Liran: Yeah. We're on the critical path. Get that.
Doron: Yeah. I want the critical path.
I wanted to add something. I wanted to add two things, but I forgot one.
Liran: But I'll give you to add your things first.
Doron: No, but I think that another conflict that we have is following what you said, we do data all the time. We practice data. We breathe this. We eat this, and, and we, that's all we do. We do more stuff now, but basically.
Liran: We also have hobbies.
Doron: That's our DNA. We have that DNA, but the problem is working with generalist developers. They have an epic on some data pipeline or something that they have to do with data infrastructure. But the specific developer might not encounter anything related to data for a long time. And by the time they get back to that feature, to fix something, to add something, two years can go by and a lot of the time this stack has completely changed within those two years. And they're like, what the hell just happened?
And I think it's different between the culture and the DNA between the ZipRecruiter and Yotpo, but at Yotpo, the full stack developers, they do pipeline building if they need. All the teams, all the product lines, all the domains, they all have features and stuff that are related to data and data infrastructure. But it's not a day-to-day thing as it is. I don't know, talking about infrastructure related to Java, for example. Because it is not something they do and practices every day. And it's also a battle that we keep trying to solve and get better at how to engage our users to understand what are the pain points.
I remember what, the other thing I wanted to say is that...
Liran: I want to react to what you just said,
Doron: But forget. I'll just say it and then you.
Liran: Okay. Alright. Very fast, I'll forget mine, so.
Doron: One of the things that we use, more and more now, is creating the right observability around data and it started mostly around cost because it's the big thing, right? In the past six months at least. but it's not only cost. It's cost and it's performance and creating this observability in a way that is actionable for the teams. First of all, it's engaging and also really it helps to bring them closer to the material in a way that they can comprehend and understand and relate to their actions. So I also think it's something very stronger we invest in.
Liran: So, I want to add to...
Doron: On the first thing or is the second...
Liran: I don't know. Not about a cost difference, but you now confuse me.
Anyways, I just want to add, that Doron mentioned that the organizations are different and I think when I got to zip, I was under, like, we are the same person, basically, just different.
Doron: I'm a woman.
Liran: Yeah, she's a woman, I'm a man. But the same person.
But in general, we were trying to have our generalist developers use big data, which really, they didn't really know how to do it. So we helped them and created a lot of structures around and interfaces. So, they can do it with really easily.
Doron: In ZipRecruiter, you mean?
Liran: No, in Yotpo. But when I got to ZipRecruiter, I saw something else. They have a new persona that we didn't have before called the data engineer. And I think that's related to what you said before Eldad is actually in both my teams, actually maybe in three of our teams that I have, we're not the experts actually. We are experts in something. We know to how to build infrastructure. We're good at that. We understand the product. We understand our users. We can collect feedback. We understand technologies.
Doron: The platform, architecture.
Liran: The platform, we're good at that. But how to actually, and what kind of data pipelines to write or ML even. We're not data scientists. We're building an ML platform. So that's really different. We actually try not to be the expert in that, in their day-to-day, but on our day-to-day and figure out, again, like Firebolt does for its customer. You don't have to be the best in the data or understand all the different types of data pipelines people use, but more understand what your customers need.
Eldad: Oh, no, our work is harder. Okay, so I apologize.
Liran: I get it
Doron: You need to do both. Yeah.
Eldad: No, it's nothing to say, right? Benjamin? building a database. It's the hardest thing in the universe.
Benjamin: By far.
Liran: Thank you. So yeah, I think for me it's very challenging. I think in data it's a bit different because we do understand data pretty well, but in ML, we have to understand our data scientists where they're not just one person. They are many different people with many different cultures and needs that they need and create a platform that will help all of them at the same time. And that's very difficult, which we are not and really understand what they do. So I think that's also in the data world, for us as well. So our experts are ever in the company and we know how to gather feedback and how to build a really good product. So, but in Yotpo again was different. So, I'm just giving two perspectives.
Benjamin: Right. So, that sounds super interesting, right? Then, coming to ZipRecruiter and kind of adjusting to kind of this different company and different team structure. How's that going basically?
Liran: I need to go on basically, but that's what happens.
Doron: But we do podcasts together.
Liran: But we do it together. But it's neither on a day-to-day basis and it's really missing.
Eldad: So basically, you went from data engineering to data science and it's hard. And, there's a lot of science there, but it's the same data, so at least that, right?
Liran: Yeah. I have one word left from everything I did before. No, I think, whatever I did at Yotpo, it still helps me with the challenges, but it's a different organization, it's a different size.
Doron: The culture is a big part of what we do. It's really important to understand what culture you're stepping into and what's the culture you want to push.
Liran: And I think what I did... I'm going to be really blunt here. What I did when I entered ZipRecruiter, I was like, oh, I know how to do data. We're experts we know what to do. And I tried to push my agenda and I figured out really quickly that this is not what this organization needs. This needs something really different.
Maybe even if I am right, I cannot just push it. Culture is something and culture change takes a lot of time and maybe in some cases you can't even completely, you have to live with something else. And I think that's what I experienced when I joined the ZipRecruiter.
Eldad: Do you have kids?
Liran: Oh, I have too many kids.
Doron: Too many kids. Too many.
Eldad: Okay.
Liran: Three for me, two for her.
Eldad: Your same experience. So, we trust you'll manage and you'll figure the kids out and because it's a new family, but you have lessons learned from the previous family. But it's amazing. It's actually, we kind of never got, at least on our show, to hear that version that those challenges from that perspective. Thank you for that.
Okay. Benji, go to the... We need something formal. So, pick the next formal bullet that Tamar gave you.
Doron: Can we do something formal?
Eldad: We haven't even really opened the formal questions and an agenda yet.
Doron: We're not formal people.
Eldad: That was all intro up until now.
Benjamin: Then that to get more formal, tell us about your goals for the next year because you've already hinted it at six past months, things change, right? Things are focused much more maybe on cost efficiency now and those types of things. So, what are the things you're aiming for with your data teams over the next year?
Liran: Can I start? So I think, Zip is in a strange position. We basically are rebuilding our entire data layer from scratch. It's quite a big company to do that in this time, so it's very challenging. So we're moving away from previous architectures and even technologies and just removing everything to just behave differently. So, right now our challenge is the quality of all of this migration and in general, just quality
Up until now, a lot of the different use cases created, data is not at the top quality. There have been testing before and ways to measure the quality of data and monitoring, but in the end, there were many, many unstructured things along the way. So, we are now trying to create, this culture by technology. So, we are building the different tools to help make that into a structured process that's repeatable. It's easy to use. and it's just no-brainers just give you high-quality data. So, we started with Schema.
So we built a way to document Schema, write Schema in Protobuf for each of your data sets, be able to kind of document everything about that dataset in a way that their consumers can actually, you know, understand what it is they're consuming.
We're creating new ways to create data pipelines based on events or in different ways, either by SQL or in Scala, and we're creating infrastructure around that. So to help the producers create high-quality data as easily as possible without having to actually deal with a lot of the complexities of doing that by yourself. So, adding think a lot for automation on top of that. We are going to build a new semantic layer.
We have a data set that's great but how do we actually use it? How do we aggregate on top of it? How do we join it to different other data sets? So creating that extra layer that explains the consumption patterns of the data. That's something big that we're going to build. We need to switch a query engine. We are right now using Athena, so we need to think of something else.
Doron: Firebolt.
Liran: Firebolt, right? That's an option.
Eldad: It's always the first and last option, but freeze for a second with that thought, and question. You talked about how you wire, you're using Protobuf, are those things that you embed within your legacy or existing, let's not call it legacy, with your production system, assuming that over the next year, step by step, you will kind of be able to unplug and plug and play new stuff or are those practices and principles that you apply on your new projects? Because you're saying you're moving the business to a new platform?
Liran: Moving everything, the end goal is basically deprecating everything that is, old and we have deadlines for it. It's quite aggressive. It's good in a way. We're actually moving. We're not just, doing everything new is going to be like greenfield and everything old is going to remain and crappy. We have to move everything. So that's going to be something that in the end, we'll just have something new and everything else that's not new will just be destroyed.
I think that's why I talked before about the question of what we actually need. So these are questions that I'm not being asked because I'm more on the technological side, but our BI developers and analysts are asking those questions all the time. Do we really need this report? Who actually uses it? Why is it so complicated? Do we really need all the data set behind? So building that data layer is not just about the infrastructure behind it, it's also about the data itself.
And I think all of those layers are being rebuilt, so it's interesting times.
Eldad: Nice.
Doron: I guess it's an entirely different story at Yotpo, but I think it can sound very, really similar if I say it, but I can say looking for the year from now, I think the challenge divides into two parts. One is in the world of analytics, it's actually combined, you know, and again, we're also restructuring the data lake, and also I think the whole escapade with dealing with cost, very, very seriously, has also got us to think of are we doing stuff, efficiently enough?
Are we doing a lot of things not in the right way or suboptimal and I think that the rearchitecting of the data platform, I talked about before about dbt and the infrastructure that we build around it, and it really embodies all the agenda that we have towards how data should look in our organization? We also working on migrating everything. This is very difficult and it starts from the actual S3 buckets and where data lies, it goes on to data contracts and data owners and how we preserve this.
It goes on to data quality and all the way to exactly like everyone talking about semantic layers and data catalogs and how we push this thing forward and I think that, we could have worked on this stuff a year ago, maybe not semantically because something is really happening now, these times it's really, really hot.
But I think even talking about the data catalog, we could have started working on this two years ago or even three years ago. But I think that we needed to get to this, level of understanding and have the organization understand what they need from, from data and for us to reach the same point where we see things eye to eye and the importance of things, and that's returning to what I said before about data manageability and the other part of it that goes together is that, currently a huge part of the data platform or all the raw data?
Most of it, not the big data by the way, but a lot of the data in terms of a number of data sets comes from CDC where we stream data from the operational databases with the BDM into the data lake and we understand that this method of replicating normalized databases and tables into a data lake and having everyone or not everyone, it depends how you layer the data transformations, but having people need to reconstruct the logic and need to understand what the team that built this architecture of how that data is modeled for an application and build this into something that makes sense for analytics, for example, or even for the product when you want to push data into a moderation view for B2B customers. It doesn't really make sense and it doesn't make sense in so many ways.
Liran: It's also about the coupling. You have someone build an application that is used in production to serve some data or to serve something for the customers that is built in a way for MySQL, which is you have a couple of tables, you have some JOIN, that doesn't really make sense for analytics. It wasn't built for that and when you couple that means that a production team or the team in charge of those tables are stuck with those like the Schema forever because someone in analytics actually utilizes it, but it's their tables. Why is anyone using those tables? So I think it's also about that.
Doron: Yeah. It's really understandable that we need to publish these data contracts or these data facades and have the operational teams in charge and be the owners of these data facades and for the analytics to rely on that, moving onwards. And, this goes on to, for us, it's probably going to go to the direction of using the output box pattern, and it's a big thing because it means that we need to restructure the whole bronze layer, and it's just something quite big, more technical than what I talked about before, but I think it's a big part of it when you're talking about re-architecting the platform.
Eldad: So, Benjamin to someone who is just an engineer, not a data engineer, even though you're building a database.
Benjamin: Which is the hardest thing in the world?
Eldad: Think about a lot of that stuff that's being discussed is when we used to write and read design pattern books in software engineering, right? So you'd go, you open a book and you had 25 patterns, and you know that there's so much experience being embedded in each pattern that you even learn them by heart because you just assume that's how it should be done, and data doesn't have it. So, a lot of what that team is doing is trying to figure it out. A lot of the staff is common and you could say, okay, we could start treating that as a data design pattern. A lot of that is unknown, right? I've heard about semantic layers since I was 16 when the first time started dealing with data. I'm 44. We invented the semantic layers. It was 20 years ago, but now\...
Doron: But it belonged to a BI tool, and now it's also decoupled, everything getting decoupled.
Benjamin: Eldad goes all game, for this podcast was just to show how visionary he is.
Eldad: Exactly.
Benjamin: Set out, in the beginning, to just drop this.
Eldad: But the truth is we called it semantic and it was just because it was so nice...
Doron: Rebranding to make it cool.
Eldad:... to have a diagram, the JOINs are just, you drag and drop the boxes and the JOIN line follows you, etc, etc. But you're right. I mean back then data was so easy in terms of metadata, in terms of who's using it. It was nothing compared to now. So just reinforcing everything you were saying, we need it, most of it is not solved yet.
Liran: Yeah.
Eldad: But it's nice to see that there are pattern emerging.
Liran: Just to add, one of the things that we do actually is actually write playbooks for different data scenarios. With the Data Guild, which is another, entity that we have at ZipRecruiter. But that's part of the job to actually create those, and I think each company probably has its own, and that's probably why there's no single book for how to write. Data Mesh tried to, I think create some kind of standards around it. But I mean, it does not dive really deep or into each technology and each kind of a pattern.
Eldad: It's because it's culture driven and it's kind of product-driven. What's the product right now? What's the culture right now? What did they go through? Did they go from Oracle three months ago? Where they're born? Did they go through the \[00:42:49.08] **\_** experience, cloud-native startup, day zero data-driven? It's how we win, so many ways to win or lose with data, so many practices to follow or not, but amazing to hear it. Benjamin, you see, there is a reason to wake up every morning. There is a reason to bend the data. Benjamin asks me sometimes, why do we need to build the database.
Doron: Why are we doing this?
Liran: Yeah.
Eldad: Do you see the pain?
Benjamin: It's so hard
Doron: It is the hardest thing in the world. It's too hard. What's it for?
Liran: Exactly, exactly.
Doron: Say the world. Pinky.
Benjamin: Yeah. I appreciate you giving me back my sense of purpose.
Eldad: That was the purpose of this podcast.
Liran: What are you doing? Why are you talking to us?
Doron: Just leave everything and go build something.
Liran: Yeah. Build some nicer, neat index.
Eldad: Guys, we're frozen in time and space right now. We don't want to disconnect and go back to life.
Benjamin: All right. Liran and Doron any closing remarks from your end? Tips, tricks, anything else you wanted to say?
Liran: I don't know.
Doron: Yeah. I don't know if you have any tips or tricks, but, I just think I said it so many times, but I can repeat myself, but again I want to say that for a very long time, Liran used to laugh at me always that I'm like data infrastructure engineer that doesn't like data.
Liran: It's a cool thing. It's special.
Doron: That's my mojo. It's good for nothing though, but I think that recently and again, the motivation was cost, but I think that we really doubled down on analytics, on what we do. And I think it's fascinating and it's also about self-measurement, also product measurement for the product that we build and also for observability, which I talked about before. I think it's fascinating. It makes our job much more efficient and, and it really helps to bridge with our internal customers.
That's my tip, I think.
Liran: I think if I also going to add to what you just said, those interfaces are what interests me right now. So the analysts called decision scientists, and we have BI and we have data engineers, we have data scientists, we have ML engineers, all of those need to work together and they're all using the same platform. By the way, in some companies, don't use the same platform. There is the Snowflake area, that's only for the analyst and then you have a data platform based on it, that's only for engineers.
I would like to have a world and I think that's what happens in both of our companies where we work really well together and we are using the same platform, the same tools, and we are all enjoying because I think that's how it should work. And I think having those organizations disconnect, and by the way, that's another thing, organization, why are they disconnected and sitting under the same roof? They all need data somehow.
Eldad: Different database licenses are bought by different managers, mostly.
Liran: I guess that's good for you. Not good for the world. But yeah, I think I would like to see us help each other better, all those different departments and I think that's what fascinates me now and I hope to I have a better future.
Eldad: Boom! Amazing. Love it!
Benjamin: Awesome. Thank you so much for joining today. It was a total pleasure and kind good luck with the super ambitious projects you guys have over next year.
Liran: Thank you.
Doron: You too. You have the biggest ambition.
Eldad: Thank you so much.
Doron: Thank you.
Liran: Thank you.
Eldad: Thank you for joining us.
Benjamin: Bye.
Liran: Bye-bye.
# How ZoomInfo transitioned from data graveyards to ROI-driven data projects (/blog/how-zoominfo-transitioned-from-data-graveyards-to-roi-driven-data-projects)
Too often expensive resources and manhours are spent on dashboards no one uses, resulting in zero ROI. Philip Zelitchenko, VP of Data & Analytics at ZoomInfo met the bros to talk about adopting product management principles to ensure data projects have value, and provide an unfiltered peak into ZoomInfo's data stack and unique tech culture.
Listen on [Spotify](https://open.spotify.com/episode/2b56ZcsJioF1FR9hXlEo3b) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/how-zoominfo-transitioned-from-data-graveyards-to-roi/id1561927688?i=1000652603546)
Transcript:
Benjamin (00:01.518) All right. Hi, everyone, and welcome back to the Data Engineering Show. We're super fortunate to have Filip Telychenko join us today on the show. He's the Vice President of Data Analytics and Zoom Info. Welcome. So great to have you.
Eldad (00:17.698) So, yeah.
Philip Zelitchenko (00:19.067) Yeah, thanks for having you Benjamin and Nildad. It's nice to meet you both and happy to be here.
Benjamin (00:23.566) Awesome. Cool. Do you want to tell us a bit about yourself, right? Kind of your background, how you got into data, your current role, and just introduce yourself to the audience.
Eldad (00:24.13) Thanks for being.
Philip Zelitchenko (00:35.515) Sure, yeah. So my name is Philip Zelitchenko. I'm the VP of Data and Analytics here at ZoomInfo. My career journey started in the world of stats. Most of my career is in the data science and machine learning world. That's where I built most of my career. And then in the last, I would say, six, seven years, I've started to expand to other areas, so look into the data platform side, which I understood that is a... big dependency for data scientists on the data platform. So it started from MLOps, then expanded to other areas. Then it started looking into data governance and data quality is important. How can we build things on the data science side with understanding the quality of the data? How do we ensure it's at a high quality? From there, it went to the data engineering fronts, because at the end of the day, as a data scientist or an ML engineer, you need to do things in batch or in stream. So it goes into those areas.
And basically, I started exploring. And the last seven years, I'm in the exploration mode. I'm still exploring and learning about platform, product management, data science, and so on. And how do you apply? And I talk about it a lot because I think it's a very good mental model, which I'm a big fan of mental models. How do you apply Marty Cagan's Inspire and Empower frameworks onto data product management and building data products, which I think is crucial because it's
If you look at the world, I think we spent a lot of resources and time on things that go into the graveyard, the Tableau graveyard, the Snowflake graveyard, you define the graveyard you want to look at. And there's a lot of time that was spent on things that in a lot of cases were not needed. And we paid a lot of money on it. So that's the world I live in and what I try to focus on in the last few years.
Benjamin (02:25.966) Nice. We really look forward to hearing more about this today. So at Zoom Info, do you maybe want to give us a high level overview of your data stack? What's the key technology you guys use? What are the data sizes? Just a one minute rundown before we all in the grave.
Eldad (02:42.434) It's all in the graveyard. It's all in the graveyard. It's all in the graveyard. Sorry.
Philip Zelitchenko (02:47.515) Exactly. No, no, yeah, it's a good question. So Zoom Info is a data company. I feel very fortunate to work at a company like this for many reasons. Everyone does data. Our CHRO writes Python. So just to give you an understanding of how technical the teams are, and it's embedded everywhere. So you're not the SME. SMEs are across the board. We have different data teams. And
Philip Zelitchenko (03:16.731) Generally, we split the data teams into two groups. There is my side of the house, which I deal with everything related to corporate. So if you think about how do we make our business more efficient, gathering information from different data sources, both on the go -to -market side, on the product side, and how do we make sure that we're building the right thing moving forward. And then we have all the data teams that support our products that are built on top of data.
And there's a set of leaders in the org that basically support those initiatives. Within my org, we have the enterprise data platform. That data platform is built out of multiple components. We have things that we build. We have things that we buy. And we have all kinds of hybrid solutions that we encounter. And that platform goes from warehouse. We're a warehouse -first company. We utilize Snowflake to.
Philip Zelitchenko (04:13.243) serverless ML ops solutions, so Databricks and some other solutions that we use in parallel, and many things in between. So MWA on AWS, how do we do orchestration using Airflow. We have something similar on GCP. The stack is pretty wide and has a lot of components to it. From the data observability side, we look at Monte Carlo. We use Atlan for data cataloging.
It's a pretty mature, I would say, serverless stack that is more modern than, I would say, most of the companies that I worked at or consulted for, where they're still in different areas, are still trying to figure out what they're doing. So we can go deeply into the different areas here, if you'd like, but I try to give you a quick overview of what we do.
Benjamin (05:01.838) Yeah, it was a great overview. So one thing I'm kind of particularly curious about is you said in the beginning, you're like your data company, you offer at the end kind of data analytics to your customers at a super wide scale, but you also have internal BI and all of these things, right? Are these like two largely separate data stacks or is it all served on the same underlying data infrastructure?
Philip Zelitchenko (05:13.627) Mm -hmm.
Philip Zelitchenko (05:24.443) So I think it's a perspective question. The way I see it is it's a unified stack. At the end of the day, the fact we are, I think we're sort of being a data company and being leaders as a data company, we have the privilege of dealing with a lot of data challenges across the board. And we are working on big initiatives. ZDP is one of them, basically the Zoom in for Data platform that's being led by a few very talented folks that come with a lot of experience of doing these things. And then the question becomes, OK, so we have ZDP. We have our EDP, which is the Enterprise Data Platform. How do we marry those together? And I think the vision is that we take this EDP as being our main platform that we're going to be utilizing and expanding it with EDP capabilities. So for example, if you need this type of tool for this type of use case, for this type of latency, this type of freshness, there's no reason why.
Philip Zelitchenko (06:23.419) EDP shouldn't be just embedded into ZDP. So in the long term, what EDP will become as part of ZDP and all the great work that the team has done to build the different components will be embedded into that great offering that will be utilized for many purposes, both external and internal.
Eldad (06:39.426) The lines are blurred and it's a good thing. It's a good thing.
Philip Zelitchenko (06:47.931) Yeah, I'm a big fan of adversity in general and being Israeli and Russian probably bad marketing, but that's who I am. And I think adversity grows interesting things and we don't have adversity here. And I think we partner very well with all these teams and we work very closely and we help each other out. But I think these different views on how to build things correctly give us the ability to bring to birth outcomes that are much better than most organizations that I worked for in the past. In the past, I've been in other organizations where there wasn't a lot of communication or a lot of adversity, if you'd like, about how to build things correctly. And the outcomes were usually suboptimal versus here. I feel like the fact that we have so many data experts, back to my point about the CHRO writing Python scripts or other people across the org doing data.
It allows us to make sure that the path we're going is the right path and it's a democratic state so people can contribute and make sure that we're not missing anything.
Benjamin (07:51.79) Nice. So let's talk a bit about data as a product, right? And you mentioned it earlier, like this kind of graveyard of components. Tell us a bit more about your thoughts on that, right? Kind of things being built or never being used and so on.
Philip Zelitchenko (08:05.435) Yeah, the way I see the world, at work at least, is that there's people, process, product, and tech stack. People, and there's a reason why I say those in that order, is because I think that's the way I think about the priority around the company and how I think about how we do things. So the people have, if you invest 10 % in the people, you're going to have a return of 50%, 80 % on that investment. You put 10 % in process. you'll probably have 30 % to 40 % return. You put invest in the product, you'll probably have 10 % to 20 % return. You put it in the tech stack, depends on the tech stack. Usually it has almost no return unless you fix the things up top with the people in the process. Because all products have a great prop or value prop, but it almost always fails on the people in the process side. That said, if you think about.
Data products in general. And we can define data products in a sec. Let's call that for now. Data products are basically a presentation of time and effort that was being put by humans in an organization. So if you sit for, and I'll give you the simplest example. Let's talk about a dashboard that you've built. And I see the dashboard as a data product. If it took three data analysts and two data engineers to build something
Philip Zelitchenko (09:33.755) for someone to consume and the Dow or the Mao or the Wow of these data products is zero, that means that you take the salaries of the people that we talked about, multiply it by the amount of hours and you get the amount of dollars that were invested and ROI of zero.
Eldad (09:54.754) There's a lot of effort put into PowerPoint presentations coordination meetings project management Decision -making leadership everything to get that dashboard up and running and be successful Just to get to what you say. So
Philip Zelitchenko (09:54.811) So if you take that.
Philip Zelitchenko (10:10.139) Exactly. So take that phenomena, multiply it by the amount of dashboards in the Tableau graveyard into the tables in the warehouse graveyard, into the ML models that are running or bad jobs that are running. Take that and multiply it and you get a dollar value. That dollar value is going to be very high. And that's something that we ignore because usually we have the fortune to work.
Philip Zelitchenko (10:39.323) at tech companies, which I'm so happy that I was born at this time in history and not 300 years ago where I'm sure I'd be unemployed or not sure what I would be doing. We have that benefit and because we move fast and we try to move forward, things, we don't have a back -view view. We don't do a lot of postmortems. So we continue throwing things behind the back and continue moving forward. But I think what's interesting about it is if you step for a second and look at the back, you understand that something isn't working well in general. And what's not working well is how do you invest the time in the right things and make sure that what you're building is right and is going to work. The same way you don't build products moving away from the data product world to the real product world, you go through a set of validations, you evaluate. There's more rigor about how you invest your time. But for some reason, when we talk about data products, it's sort of treated as free.
And I think that paradigm shift will happen in the next few years, where this will become something of the past. This would be an absurd, the fact that we invest so much time in things that are not being used.
Eldad (11:48.514) I've heard that about PowerPoint presentations and PDF files and I'm still getting those. I saved dashboards as PDFs by the way. So you can actually read it. But I completely, I'm with you on a lot of stuff you're saying. But sometimes, you know, people get religious over dashboards. But in reality, it's just people translating decisions into a nice presentation. Sometimes the purpose of the dashboard is really to be viewed once because of the effort that was put to behind the scenes to get the dashboard up and running, the cleansing, the modeling, the picking of the right stack. So the dashboard is really the least important thing. As you say, it says you move on, you just evolve and dashboards are just, right? Like those points in time where you did a snapshot of your data, you know, your data capabilities.
So, you know, it's okay, create wrong dashboards. As long as your model evolves, as long as your team gets smarter with the data, make faster decisions. And it's hard, it's really hard. So you're right, in many cases, the dashboard becomes the project. And I think that's when you know something is really wrong.
As you said, it's not the dashboard fault, it's the people behind it and the intentions that drive that dashboard. And I've seen some nasty things in my career, but always good intentions. So yes, I contradict myself and our theme here to be really nasty on dashboard and to go to the graveyard, pick them out, kill them again. But we don't want that. We are good with them.
Philip Zelitchenko (13:24.699) Yeah, yeah.
Eldad (13:39.842) Benji, what's your thought?
Benjamin (13:42.35) So one thing I'm curious about Philip is like, when you're talking about data product management, and you're saying, right, like for internal dashboards, for the CEO or whoever kind of, we consistently create things that have no ROI. Now you're at a kind of data product company. So in a sense, for all of the, many of the data projects you're doing, the ROI is easy to measure, right? Like kind of Zoom info is a public company, kind of you earn money, kind of off. at the dashboards and data experiences you create. Like, why is that going better? Cause in a sense, I guess you're advocating for all of the processes you have around taking data experiences to your own customers, also mattering then internally for any organization that kind of have large amounts of data.
Philip Zelitchenko (14:31.643) Yeah, and I think it's a good question. But again, I think our company is unique in many ways. Most companies in the world are not data companies. And most companies, the data function that they have is a corporate data function that serves the internal product. In some cases, you'll have some sprouts of data that is ingested in different areas of a product. But most companies in the world are in the world where the data function mainly serves internal stakeholders. And the reason like any other company, that function is something that I think needs to be stabilized and built in the right way to make sure that you invest the time of these resources in the right place. And so back to your question of why is it, I think your question was why is it so important to me? Because at the end of the day, the fact that our product team works and thinks about things as products is very dependent on the fact that
Philip Zelitchenko (15:29.435) It is a product and engineering team versus data analytics teams today are not managed as product teams. Most data analytics teams in companies that I've seen that are internal data analytics teams, they don't have a product manager that thinks about it. Mostly it's business stakeholder to engineer. That's the first relationship that you get. And that's, I think, one of the main reasons that that graveyard that we've talked about is so big is because we don't think about what we build as products in those use cases.
Philip Zelitchenko (15:59.003) We think about them as stopgap solutions. And if you look at the most of these stopgap solutions, what happens a lot of time is that we recycle things that are in the graveyard, but we are not even aware of it. And the reason why is because the requirements, thinking about how it's going to be utilized, how it's going to drive value, all these questions are usually being run by the business. And unfortunately, business people, they are amazing at a lot of things, but they're not product managers. They're not thinking about, OK, how do I build an experience? For my internal team to ensure that they can utilize data in a smart way that will help to drive outcomes, that will help drive ROI, that will help us upsell or cross -sell our customers or reduce turn on our customers or save costs around the org. These things are not part of their day -to -day thinking. It's mostly a transactional thing that they do as a side job. And usually it looks like a side job. That's why the graveyard is so big.
Benjamin (16:56.27) But calculating like your data ROI does seem much easier in a company like data product company, right? Cause like you sell the product for X amounts of dollars. Like it's easy to calculate this. If I give a dashboard to some executive and kind of she makes a better decision because of that becomes much, much harder to measure this. So how do you think that kind of transfers, right? Kind of from the world you have at zoom info to maybe some company where it's mainly internal.
Philip Zelitchenko (17:26.427) And I think that's a great question. That's part of what you do when you write a DPRD, when you write data product requirements document, you think about how do you measure the impact of what you do. And in a lot of cases, you don't. What happens in most cases, in most companies that I've worked for, have you heard of the HIPPO?
Eldad (17:38.498) Thank you.
Benjamin (17:42.958) I haven't heard of the hippo, no.
Philip Zelitchenko (17:44.603) Okay, so the HIPPO stands for the Highest Pending Person Opinion. That's HIPPO. And usually the HIPPO rules. In those companies, the HIPPO is the way things are decided on. Don't measure the HIPPO because the HIPPO is the HIPPO. And what happens is, if you don't put any measurements around the HIPPO, and you don't, and forget about the HIPPO itself, the HIPPO is just a representation of it's, yeah.
Benjamin (18:07.47) Measured a hippo, measured a hippo.
Eldad (18:09.73) Hippo!
Philip Zelitchenko (18:12.283) So the symptom, this is just a symptom of what I'm trying to say is when you work on a product, we're trying to release a product, you're going to think about the metrics, how you're going to measure the impact. You're going to think about maybe I'll run an A -B test, right? Maybe I'll see if I'm building a way to improve my sellers, I'm going to run an A -B test between sellers that can use my insights or not and evaluate that. These practices have never been something that the go -to -market teams or the business teams in companies thought about. Because they try to move fast. And moving fast meaning creating a larger graveyard with less thoughts around, what is the framework to do this? Instead of doing 50 projects that are trying to micro -optimize each step here, is there a framework we can build that will help run these more efficiently, but also will be able to measure them, measure the impact of them, and will help us make sure that we're going the right direction?
Benjamin (19:03.918) So this role, you would basically embrace the kind of flows and processes you have at these data product companies kind of much more widely internally as well in organizations that aren't first and foremost data product companies. Like to me, that makes sense.
Eldad (19:14.082) By the way, there are many organizations that will never be data first because they're just selling something else and they will always have a different exposure to data. And most importantly, from a cultural perspective, those companies learn to make decisions.
Eldad (19:42.658) almost completely with humans. So when they go and sell a car, it's the agent that decides the discount and how based on the, I don't know how much the customer blinks. And that's hard. And that's kind of goes way back. But when you're building a data company, you have, you remove so many limitations and you can rethink a lot of the go -to market.
Philip Zelitchenko (19:58.235) Okay.
Eldad (20:11.362) that you would apply, right? Like so many things, little things that you've mentioned when you talked, like, yes, you just change the product you're selling. And of course, anyone, you know, anyone in the services business would say, you can't just have everyone being in the service business. Someone needs to drive those people to work. And, but I think, yes, looking forward and how information workers transition.
Benjamin (20:33.39) Okay.
Eldad (20:41.314) from being consumers to making decisions which they're not with dashboard and this is what you're saying having a dashboard is it's not a trophy okay it's like a spreadsheet it's nothing more it's like whatever um but transitioning information worker it's basically your same information worker got it wrong and this all the stack we've did to serve decision makers right like what is bi what is self -serve?
Eldad (21:09.826) It's like a thousand different ways of doing the same thing, making less mistakes. And we're getting to square one. So data companies and data is a product, is the future across many industries and we're seeing it. So you better work or plan or be somewhere in that ecosystem. What else do you think will happen?
Eldad (21:36.002) As we move forward as we'll have less interfaces to do with right? I mean what you're saying is there is no looker interface in a few years from now Nobody's doing that. So what do you think will happen? How will you sell your data products? Dashboards last for 20 years. Believe me. I know
Benjamin (21:36.686) Remind, remind me, three years is look at that.
Philip Zelitchenko (21:56.827) I think, so I agree with a lot of what you said. I think the big thing that I think will happen is dashboards currently are transactional products that are not meant to stay in the form that we think about them. So that takes me back to the Dow, the Mao, the Wow. You look at these things on these metrics, on these assets, and you see that they're not meant to stay.
Eldad (22:06.178) But what if your AI call pilot generates that dashboard which
Philip Zelitchenko (22:25.115) because it always declines with time. Well, it...
Eldad (22:32.29) between us, right? Like this is kind of first feature for a co -pilot. Isn't that just yet an animated PowerPoint with the right data? So the dashboard itself, as we said, is not the point. It's the journey, the model, the cleansing, everything that happened to get there, but then generating the dashboard is free. It's like a PowerPoint. You get it for free, but it's everything else that matters. So.
Philip Zelitchenko (22:52.571) That's a good question. And I think that what is the time horizon we're looking at? Are we looking at 50 years from today or five years from today?
Eldad (22:59.426) Will you change your perspective on dashboards? And that's my question. Will you change your perspective on dashboards when they become yet another file format? So on Wired, unfortunately, it's 12 months in reality, I guess 15 years.
Philip Zelitchenko (23:22.331) Yeah. Yeah. And if you think Elon Musk would say it's going to be ready by end of year, right? So I think at the end of the day, I think the timeline is much longer than that. And I think in the meantime, what will happen is instead of creating tens or hundreds of dashboards that are going to be consumed by people at the organization.
Philip Zelitchenko (23:51.707) Yeah, no worries. So I think what will happen is that in the short, medium term, we'll move from dashboards to data apps, which I think is going to be the new world. Because when you think of building a data app, you're going to start thinking about the experience of the people who are going to access the data app. And now it's getting closer to a product rather than a transactional object that is there to serve you for 30 minutes for a meeting. And when you start thinking about it, you will start thinking of frameworks.
Eldad (24:18.562) and I'm going to be talking about that in a minute. Thank you.
Philip Zelitchenko (24:38.524) So that will be a centralized place, one central place that people can consume. But that's not the only one, because from an activation perspective, some of the data will flow into a data app. But some of the information will flow into your systems of engagement, where you are trying to drive a certain behavior, where there it's going to be not just another field in Salesforce or another field in whatever system you'd like, but it's going to be driving a workflow in that experience that will drive that workflow for the
Philip Zelitchenko (25:06.748) for online practitioners using those systems of engagement. And I think the activity.
Benjamin (25:08.174) What's the, I never heard the term system of engagement. Like what's the actual like textbook definition on that for someone who doesn't know.
Eldad (25:10.882) Absolutely. This is.
Philip Zelitchenko (25:18.588) So there's two systems that usually people talk to about, system of record and system of engagement. System of record is basically a place where you maintain information. So if you think about Salesforce today, removal of the Salesforce marketplace and all the things that connect to it, it is a system of record. People go there and put in notes or put in details and so forth. A system of engagement is a system where you log in and you engage with it.
Philip Zelitchenko (25:47.324) to make actions. So it's not just a place where you drop data in, it's a place where you activate different initiatives or you send outreach as an email approach that SDRs use to send out emails in bulk, define workflow. It basically helps automate a lot of the work. Slack is a system of engagement exactly in some sense. Some things can be both, right? So it depends on what basically the product is.
Eldad (26:05.282) Slack.
Philip Zelitchenko (26:16.06) but usually refers to some product, not data product, product, that has qualities of either writing, or writing and reading, or writing, reading, and acting upon whatever is happening in that system.
Eldad (26:29.666) And it's smart, you know, Salesforce, being Salesforce, being smart and being a great company. Like you just described like outladed for everyone, right? Like build a system of record and then we translated it into a system of engagement. So they acquired Slack, they acquired Tableau. They give any way to engage data through the dashboard. Um, that didn't go as well as planned and they got Slack.
Eldad (26:58.594) to have the other kind of engagement, which could they got really well, way beyond what they expected up to the point where this is kind of becoming their operating system of engagement. Thanks for playing us like those two things are crispy for us.
Philip Zelitchenko (27:16.892) Yeah, and I would also add that basically if you think about products in general, some of them are a pool, some of them are a pool and like reactive, proactive, pull and push, whatever you want to call it. I think the data apps will become a pool method where you go and you consume. And the activation layer or the system engagement layer would be the push where you set up automations that are going to be assisting you to help productivity across the different functions in your org.
Philip Zelitchenko (27:45.596) to make sure that you're doing your job in a good way. Because at the end of the day, with all of the things that we can do as a single front -line practitioner at a company, from a seller to a marketing person to a CSM, you can't really do your job in an amazing way without the assistance of some automated processing in the background. Because currently what happens is the method is you call someone, you write notes, you call someone, you write notes, you call someone, you write notes. That is not an effective operating model.
Philip Zelitchenko (28:15.228) And it's probably the productivity, if we had a measure on productivity of human beings and companies, that productivity metric would be pretty low because of of. Yeah, because low interest money is free, so you can throw bodies at any problem and then that's how you solve it in a world where interest is high. You need to find ways to be more creative in some sense.
Eldad (28:26.114) But it worked well with low interest. It worked well with low interest.
Eldad (28:42.594) Exactly. Good to be living in high interest rate times.
Philip Zelitchenko (28:48.572) Yeah, rise efficiency.
Philip Zelitchenko (28:56.316) The Industrial Revolution would probably come much faster in a world where we had a lot of high interest times.
Eldad (29:11.458) Benjamin, I can see you're circulating some thoughts with yourself. So feel free to share it with us.
Benjamin (29:15.822) And I'm still like...
Benjamin (29:20.494) No, so I'm still kind of trying to, to wrap my head around kind of your overall perspective on, on kind of the, these like data pipelines and so on, Philip, right. So first of all, I need to understand what system of engagements are and figure out how this fits in my mental model. I think my question kind of in the bigger scheme is right. Like when you're talking about, okay, people writing notes into Salesforce, kind of keeping track of all of their interactions with customers and so on, and then needing more efficient data products and data pipelines in order to become more effective.
Does this to you also then fall under projects actually owned by the data team within the company where you want to embrace all of this kind of then product workflow around figuring out ROI, kind of figuring out kind of requirements and so on, or is this more like, okay, there's going to be specific products emerging around this that actually make this easier? Like, do you feel like this kind of business function will be built more in -house in the future by data teams? Because it's very specific to the business or people will just buy it from like Salesforce 2 .0.
Philip Zelitchenko (30:31.388) It's a good question. I think it'll be probably a hybrid. There's going to be more tools. And we'll have a new infographic coming out every month showing how populated the whole industry is. And these tools will solve some use cases. But the biggest problem, or one of the big problems, is that as these set of tools continue to grow, the more discrepancy you have in your operating model, for personas that are working on a certain problem. At the end of the day, the number of tools per FLP, per frontline practitioner, has grown, I don't know, 300x over the past 10 years. So when I use, not me, but on an average seller used to sell 20 years ago, they had one system. Today they have 40. And that's true not just for a seller. It's true for an SDR, an MDR, an AE, marketing manager, a data engineer, a software engineer. You choose. Yeah. Yeah.
Eldad (31:35.746) phones, your phone suddenly, you can download as many apps as you want to them. So yeah, I mean, it's interesting to see how that will end and how far that will actually go because it was driving us really nicely, right? All of us kind of having that tool productivity mindset going from using one SAP with 50 apps in it with negative productivity to having that interconnected, well -behaved ecosystem with humans to connect the dots, which I think is much better than a glued single tag stack. But I don't know, right? Times change, things look at snowflake now, right? Like I remember people used to laugh at Oracle for trying to do more than just building a database.
So now it's okay to have a CRM. It's back okay to sell CRM with your database. So it's fashion, you know, it goes in cycles. You never know. You just need to focus, as you said, on great products, finesse and have teams that love the products they build and iterate fast, especially if you're in data. No egos. It's not needed because mistakes are almost practically for free.
Philip Zelitchenko (32:53.276) Yeah, I agree. Yeah.
Eldad (33:01.794) And so you don't need that ego in the room to protect the project, to have defense balance, nothing. No, it's like a query. It's like a model. It's a right. I think listening to you at the beginning, that is what I took. It's like, it gives me flexibility and freedom. So I can approach my team. I can lead my team differently. Maybe we can all build better stuff. Boom!
Benjamin (33:31.822) So maybe pivoting a bit and then kind of also generative AI, right? Like obviously kind of that's on everyone's mind. And I did want to kind of bring it up today because I think you have a really interesting perspective on that because one, you're using data infrastructure so heavily, right? And like systems like Snowflake have Cortex now and all of these generative AI features to deal better with semi -structured data and so on. At the same time, you're also thinking a lot about.
How do paying customers interact with data? Because this is what your business does in the end. Where do you see generative AI fit in? Both in terms of the internal data stack as a company, but then also in terms of how customers and users interact with data as a data product.
Philip Zelitchenko (34:19.868) It's a good question. So I'll split it into two. I think let's talk about ZoomInfo as a company, because I think there is some exciting work that is being done. We're working on releasing something we call Co -Pilot. ZoomInfo Co -Pilot is going to be basically a way to interact with your data and act on it in a more of a natural language processing way. And I think there's some exciting work that is being done currently in the R\&D and product org that will be exciting for a lot of our customers in a few months. And in that world, basically, the approach of write versus writing or reading, all that will be a bit easier, or I think much more easier from a productivity standpoint. If you think about Microsoft Copilot, I think there is research showing 15 % to 20 % to 25 % improvement in productivity of writing code.
I think a similar thing is coming from ZoomInfo for our ICP, our ideal customer profile. So our AEs, our SDRs, MDRs, they're going to have a very high benefit from interacting with a co -pilot that is going to basically help automate a lot of the things that currently they spend a lot of time doing. And there's some interesting research coming both from Forster and Gardner about how much time on average people spend on maintaining tasks that can be basically automated using GenNI. So I think there's great work that's coming from there. On the other side, I think for the.
Eldad (35:55.586) Oh man, what's gonna happen when everyone stops moving stuff around and start thinking? Because you don't need to move stuff around anymore.
Benjamin (36:02.67) So now that Eldad interrupted you already, I might as well follow up with another question. Having experience now taking these products, at least on their way to production, what's actually the hardest part? What's the biggest engineering challenges you have there? Because I think a year ago, everyone was just talking about these LLM companies and, oh, which LLM will win. That debate is kind of.
Philip Zelitchenko (36:04.316) Yeah, I think it's exciting. I agree with you.
Eldad (36:08.418) Mmm!
Benjamin (36:32.174) Over like as someone taking such a data product to production now, like what's the hard engineering about it?
Philip Zelitchenko (36:39.708) I think there's a few challenges there. I'm not sure I'm the right person to answer that question here, just because that's not my realm of responsibility. But I think, in general, that if I had to classify it into categories, I would say A, which is something that we're all probably aware of, the hallucination problems that are happening, how do you solve them, and B, I think, scale. So I'd say these are the two things. And the third one, which is hidden under scale, is cost. So combination of those three factors, but again, I wouldn't call myself the expert here on those areas.
Benjamin (37:11.022) Sounds good. So then kind of backtracking to my original question around kind of these, like you wanted to split your answer in two and talk about the second dimension now as well.
Philip Zelitchenko (37:21.98) Yeah, I think the other piece is how do we not make Gen .AI the new graveyard pool for similar to Snowflake and Tableau and other areas, right? Because at the end of the day,
Benjamin (37:31.854) The data engineering show season 55.
Eldad (37:31.874) You know what's more important? How do we make sure dashboards don't become the best thing that everyone does because everything else is just too expensive and just people go and say, let's just create a dashboard and run it for a few dollars and that's it.
Benjamin (37:47.982) Yeah. 20 years from now, we'll be sitting here again and Philip, you'll tell us about the graveyard of generative AI kind of products you've seen.
Eldad (37:55.938) I'm sorry.
Philip Zelitchenko (37:58.556) It's already being populated by bodies of things that are trying to be built. Every time I see an announcement of Microsoft or OpenAI, I basically see a whole industry just getting wiped. It's like a Marvel movie where basically someone...
Eldad (38:00.546) Who's?
Benjamin (38:03.598) It's filling up.
Eldad (38:05.634) of lost souls.
Eldad (38:18.306) It's crazy.
Philip Zelitchenko (38:25.5) snaps the fingers, and a third of the population disappears. But joking aside, I think in a similar sense, what will happen is that now there's this hype around what can generative AI do. And I think the Kruger effect, which I don't know if you guys heard of it, but you probably have, basically a very famous curve where there's the first hype, and then it goes down, and then it continues growing in a more, I would say, in a less steep way. And what happens in those areas, I think we're still on the first hype cycle where people are saying, what is your AI strategy? How are you going to solve world peace with AI? And all of the buzzwords around that. And I think soon enough, the tide will go down and we'll see who's swimming without any underwear.
These people walk out of the room and we'll stay with things that are really beneficial and help us do things better. So I think that would be the... Exactly. So and we're also...
Eldad (39:29.698) I don't want to compete. I don't want to compete with you guys on the data product on your domain. But I love, I love the, I love the attitude. Like really I love everyone who, every engineer who love what they did and love their peers and wake up every morning and love doing it. So thank you for.
Being an inspiration on that for sure. And there's the graveyard. Yes, there is that slide. And, and in our domain, it's changing actually much faster because there are no constraints. So yes, we might wake up in a few days, weeks or months. And that will be really a restart for, for, for data stack. And when that happened, it's good. That's what we do. But you're still in the data warehouse always, always we need a data warehouse.
Philip Zelitchenko (40:24.156) Yeah, no, I agree.
Eldad (40:29.346) Other than that, I don't know. Other than that, I don't know.
Philip Zelitchenko (40:32.54) I agree 100%. And again, what I'm trying to say is that I'm not trying to aim for a world where there's no graveyard. The graveyard is important. I'm just saying that I think we currently don't measure how big the graveyard is or what is the ratio between people who are alive versus people who are dead. Like humanity is trying to improve life expectancy, lifespan, health span, and how we're doing things better. I think in the same way, we need to be more thoughtful about what we build.
And how do we ensure that the ratio of things that go to the graveyard versus stay in live are being thought through? And the way to do it is to start understanding what exactly, what is in that graveyard and why is it there? How do we ensure to reduce the amount of things that go into the graveyard?
Benjamin (41:15.534) I think that was the most inspirational outro we ever had on this show, Philip. That was great. It was so good having you on today. Really, like I learned a lot. It was so good hearing your perspectives. Thanks a bunch. Any closing words from you other than your perfect words just now.
Philip Zelitchenko (41:38.268) No, just thank you again for hosting me. And again, I think as Eldad mentioned, which I want to echo, no one really knows what's right or wrong. We can be, we're almost around us. There's a lot of smart people. And I think remembering that and just coming open to conversation to hear other thoughts and argue and make sure the right decision are being made, I think that's the key for success of all of us. So yeah.
Benjamin (42:06.862) Thank you so much, Philip. Awesome. Cool.
# Implementing Explicit Multi-Statement Transactions in a Stateless, Cloud-Native Architecture (/blog/implementing-explicit-multi-statement-transactions-in-a-stateless-cloud-native-architecture)
### TL;DR [#tldr]
Firebolt supports explicit, multi-statement transactions using the familiar BEGIN, COMMIT, and ROLLBACK syntax. You can now group DML, DDL, and DCL statements into a single, atomic operation that maintains full ACID compliance while preserving the value of Firebolt's stateless architecture.
These transactions can span different compute clusters, provide snapshot isolation, and continue running even during compute cluster replacements or upgrades, with no downtime. This article explains the architectural decisions that make this possible, focusing on how Firebolt's centralized metadata service enables this capability
Syntax:
```sql
BEGIN TRANSACTION;
-- Read data to inform next steps
-- SELECT ...
-- Perform data modifications (DML)
-- INSERT, UPDATE, DELETE, MERGE
-- TRUNCATE, DROP PARTITION
-- Modify the database schema (DDL)
-- CREATE / ALTER / DROP TABLE/DATABASE
-- Manage permissions (DCL)
-- GRANT / REVOKE ...
COMMIT TRANSACTION;
```
### Why You Need Multi-Statement Transactions [#why-you-need-multi-statement-transactions]
Grouping multiple INSERT, UPDATE, and DELETE statements into a single atomic unit is fundamental to data integrity. Without this capability, partial failures can leave your data in an inconsistent state.
Consider a common ETL/ELT pattern where you update a fact table and its related dimension tables simultaneously. All updates must succeed or fail together to maintain consistency across your data model.
An even more demanding scenario is blue-green deployment: replacing a production database or table with a new version from staging. This involves a sequence of DROP and ALTER RENAME TO commands that must execute atomically. Partial success could leave your production environment broken.
In production-critical data analytics, atomicity isn't optional, it's essential. The challenge is providing these guarantees within a stateless, distributed architecture.
### Decoupling transactions from compute clusters [#decoupling-transactions-from-compute-clusters]
Traditional data platforms often tie ACID transactions to a single, stateful compute node. While familiar, this approach creates rigidity that can't handle modern, distributed workloads. You sacrifice elasticity and resilience because a single point of failure can disrupt your entire workflow.
Firebolt decouples transactions from the compute layer. Transaction state is managed centrally by the metadata service, allowing transactions to be long-running and resilient to infrastructure changes.
A single transaction can seamlessly span compute cluster replacements, scaling events, or [online upgrades](https://www.firebolt.io/blog/engines-online-scaling-and-upgrades) without interruption or downtime. You can perform complex, multi-step transactions without worrying about the underlying compute cluster infrastructure.
This is a unique capability in cloud data platforms, providing integrity and control that works with, rather than against, a stateless, cloud-native architecture.
### Why Firebolt uses centralized metadata architecture [#why-firebolt-uses-centralized-metadata-architecture]
This table compares different approaches to transaction management in distributed systems:

### Stateless execution with metadata-coordinated transactions [#stateless-execution-with-metadata-coordinated-transactions]
Firebolt's approach to explicit transactions leverages a centralized metadata service that is decoupled from compute clusters. This service acts as the single source of truth for all transaction states and content.
When you issue a BEGIN command, the metadata service provides a unique transaction ID. This ID serves as the identifier for an isolated workspace in the metadata service where all metadata changes, followed by INSERT or UPDATE operations, are staged.
These changes are initially available only to your session using that specific transaction ID. Once you commit the transaction, the metadata service atomically validates and marks the transaction as complete, making the staged data visible to all future requests.
This design ensures that compute clusters remain stateless and interchangeable, allowing transactions to maintain full ACID compliance even as they span different compuete clusters and resources. You never need sticky sessions or compute cluster affinity.

### Metadata service architecture and design [#metadata-service-architecture-and-design]
The metadata service stores all system information that isn't user data. This includes table and database structure definitions, access control permissions, and other system-level information. Your actual data is stored separately in a data lake. The metadata service only stores metadata pointers called "tablets" that serve as a map to where your data resides.
The metadata service is built on FoundationDB, a distributed, transactional key-value store renowned for its high performance, strong consistency, and ability to handle large-scale distributed workloads with millisecond latencies.
This service maintains all metadata transactionally in the order it was committed, allowing Firebolt to construct a complete picture of the database state at any given point. To ensure strict, chronological ordering between transactions, Firebolt uses a Log Sequence Number (LSN) that acts as a global, monotonically increasing timestamp for the entire database.
### Understanding the Log Sequence Number (LSN) [#understanding-the-log-sequence-number-lsn]
Think of the LSN as a precise timestamp that marks when each transaction begins and commits. It ensures transactions execute in a strict, predictable order.
Firebolt generates this LSN using FoundationDB's Versionstamp feature, which embeds the unique transaction's commit version into a key or value. This produces a textually sortable 20-hex-character LSN (e.g., 00000000f03ac8b10000) used extensively for transaction identification and for storing cache entries on the database compute cluster.
When a transaction begins, it receives an LSN (called begin\_lsn) that defines the "snapshot" of the database the transaction will see. This ensures your transaction starts with the most up-to-date data, including all data from transactions that committed before it.
When the transaction commits, it receives a new LSN (called commit\_lsn), which is guaranteed to be greater than any previous LSN. This commit\_lsn marks the point at which changes become visible to other transactions.
This process guarantees that commit\_lsn > begin\_lsn for any given transaction. More importantly, the begin\_lsn of the next transaction to start will be greater than the commit\_lsn of the previous one. This strict, chronological ordering prevents transactions from seeing or interfering with each other's in-progress changes, maintaining the isolation and consistency of the database.
**Building the transaction log**
To ensure atomicity and isolation, Firebolt creates a dedicated transactional workspace for every transaction, whether explicit or implicit. All changes, from a single row update to a new table creation, are logged to a Write-Ahead Log (WAL) identified by the unique begin\_lsn.
This log-based approach ensures data durability and acts as the building block for the entire transactional system. Changes written in a transaction are only accessible via the unique begin\_lsn, guaranteeing they remain completely isolated from other transactions until they are officially committed.
The commit process makes these changes visible to the rest of the system. Instead of rewriting or moving the data, the transaction simply receives a new, unique commit\_lsn. This commit\_lsn serves as a pointer to the begin\_lsn, making the changes permanently visible and part of the official transaction history.
This allows Firebolt to replay the transaction log in the correct commit order to reconstruct a snapshot of the database at any given point in time across all compute clusters.
To read the transaction log in the correct order, Firebolt first reads the list of ordered commit\_lsn values. Each commit\_lsn acts as a redirection to the corresponding begin\_lsn that holds the data, allowing the system to accumulate the content needed to construct a consistent snapshot.
**Key concepts:**
* The begin\_lsn defines a transaction's consistent snapshot, and its order reflects when transactions started.
* The commit\_lsn is a globally ordered and sequential number that defines the final, true order of transactions, used to replay the transaction log and reconstruct a consistent system state.
**Example:** A transaction with begin\_lsn 10010 can see and delete a table because it began after the table's creation was committed at commit\_lsn 10070.
**Another example**: A transaction with begin\_lsn 10005 can insert data into a table because it started after the table was committed. However, if that transaction is not committed, its data remains invisible to others.

### Indexing WAL for Optimized Read Patterns [#indexing-wal-for-optimized-read-patterns]
To avoid scanning the entire transaction log for every query, Firebolt maintains a crucial indexing layer on top of the WAL. These indexes quickly differentiate which committed changes are relevant to a given query, dramatically improving read performance.
The indexes are created by storing metadata changes, such as new or modified tablets and objects, in a separate, structured way. They are organized based on common access patterns (e.g., per-database or per-table) and sorted by their commit\_lsn.
This approach provides two main benefits:
**Single range read:** The underlying key-value store is highly optimized for reading a range of ordered data instead of many small reads. By using these indexes, a query can access all relevant log entries in one operation. This avoids the inefficiency of many small, scattered reads and reduces I/O and latency.
**Targeted access:** The indexes enable reading only what you need, allowing a query to request only the portion of the log relevant for specific tablets or objects, within a specific time frame. This means you can read only the incremental changes from a recent snapshot the compute cluster already holds.
These indexes transform the WAL from a simple append-only log into a highly performant and queryable data structure.
**WAL index structure**:

The WAL index sorts data first by the read pattern (e.g., database for reading schema, table for reading table data), then by commit\_lsn.
The index allows for a single, efficient read of a specific table, even within a given commit\_lsn range. For example, if an compute cluster already knows the table's state up to commit\_lsn 10112, it can request only the incremental changes in the range (10112, 10122] and apply those changes.
### Compacting the Transaction Log into Snapshots [#compacting-the-transaction-log-into-snapshots]
To manage the growth of the transaction log, Firebolt periodically compacts the recent transaction tail into a single, concise snapshot. This process creates a deterministic picture of the system at a given point in time, permanently applying all committed deletions and modifications and eliminating them from the log.
This compaction serves a crucial purpose: it allows Firebolt to safely remove old records from the transaction log and its indexes. After a snapshot is created and has passed a defined retention period, the log entries it replaced can be safely deleted.
This retention period enables features like time-travel, allowing you to reconstruct a system snapshot from a requested time in the past. The oldest snapshot you can use to reconstruct the system's state is determined by the retention period (also called the time horizon).
Since Firebolt regularly compacts older transaction logs into snapshots and then deletes the logs, you cannot construct a view of the system before this time horizon. This is why the time-travel feature is limited to the defined retention period.
A long-running explicit transaction is not a problem even if it runs for hours and tries to construct an "old" system state, because the data retention period is much larger than the transaction's lifetime.
## Isolation levels and conflict handling [#isolation-levels-and-conflict-handling]
In a system with parallel transactions, a conflict occurs when two or more operations logically contradict each other, potentially leaving the database in an inconsistent state. To handle this, Firebolt employs Optimistic Concurrency Control (OCC).
This approach assumes conflicts are rare, allowing transactions to proceed without locks. Conflicts are detected during a final validation step before a transaction commits. The key difference between isolation levels is the specific types of conflicts they prevent using this OCC mechanism.
### Isolation levels [#isolation-levels]
**Strict serializability:** For critical metadata operations (e.g., creating or dropping a table), Firebolt applies strict serializability, the strongest isolation level. The OCC system checks for both read-write and write-write conflicts to ensure all structural changes are globally ordered.
For example, a transaction attempting an INSERT will fail if the target table was concurrently dropped, as this is a read-write conflict on the metadata.
**Snapshot isolation:** For data operations (e.g., SELECT, INSERT, UPDATE, DELETE), Firebolt uses snapshot isolation. A transaction is given a consistent view of the data from when it began, regardless of concurrent changes, which eliminates read-write conflicts.
Therefore, the OCC mechanism only needs to check for write-write conflicts, aborting the transaction only if it tries to modify data that another concurrent transaction has already committed.
### How Firebolt enforces isolation [#how-firebolt-enforces-isolation]
To enforce the specific isolation levels described above, Firebolt tracks every transaction's activity using a system of objects called dependency tags. These tags are the key to detecting conflicts during the commit phase.
Here's how it works:
**Tracking during the transaction:** Following an operation such as INSERT, when a transaction performs a WRITE or READ of metadata, the compute cluster creates a simple record, a dependency tag. This tag notes the type of operation (read or write) and the specific data involved.
**Validation check:** When a transaction is ready to commit, the metadata service examines all the tags created by other transactions that have committed since this transaction began.
**Conflict detection:** The metadata service then checks for conflicts. For example, to prevent a read-write conflict, it checks if any of the data this transaction READ has been written to by another committed transaction. If it finds a match, the transaction is immediately prevented from committing.
This process provides a flexible and precise way to enforce isolation. Firebolt simply configures the system to check for the specific types of conflicts required by each isolation level, preventing write-write conflicts for snapshot isolation and both read-write and write-write conflicts for strict serializability.
### Example of read-write conflict validation [#example-of-read-write-conflict-validation]
Consider two concurrent transactions, T1 and T2, interacting with a table named sales\_data.
T1 begins an UPDATE operation on a row. It first reads the sales\_data metadata to confirm the table's structure. The compute cluster registers a dependency tag for a READ on the sales\_data metadata key.
T2 begins a DROP TABLE operation on sales\_data, which is a structural change. The compute cluster registers a WRITE dependency tag on the sales\_data metadata key.
T2 commits successfully. T2's WRITE dependency tag on the metadata key is recorded in the committed log.
T1 attempts to commit. The system checks T1's READ dependency tag against committed tags since T1 began. It finds a conflict with T2's WRITE tag on the same key.
T1 is aborted. Because T1's operation depended on a state that was concurrently changed, it cannot commit. This mechanism ensures the system's structural integrity by enforcing strict serializability for metadata operations.
Example:

In an explicit multi-statement transaction, you don't need to wait until COMMIT to realize there is a conflict. To improve your experience, Firebolt performs this validation after every statement rather than only at commit. This "fail-fast" behavior gives you earlier feedback on conflicts, saving time and resources.
### Controlling Sequential Execution Within a Transaction [#controlling-sequential-execution-within-a-transaction]
In a distributed system where any statement can be routed to any compute resource, ensuring sequential execution within a single transaction is crucial. A transaction depends on a consistent snapshot to operate, so allowing multiple write operations to run in parallel within the same isolated transactional space could introduce inconsistencies before the transaction is verified and committed.
For this reason, Firebolt explicitly restricts parallelism within a single transaction, rather than allowing undefined behavior like some other database systems.
Firebolt achieves this by generating a sequence\_id along with the transaction\_id. On every statement, the sequence\_id is atomically checked and updated in the centralized metadata service and then sent back to the client.
The client, when submitting a subsequent statement in the same transaction, must attach the correct sequence\_id in the HTTP headers. If the provided sequence\_id does not match the expected value, the transaction is immediately aborted.
This mechanism guarantees that statements within a transaction execute sequentially, maintaining integrity even in a highly distributed environment.
### Security Considerations [#security-considerations]
Security enforcement is real-time: Every statement in a transaction is individually authorized based on role permissions.
RBAC (Role-Based Access Control) is always checked against the most recent state of permissions, regardless of when the transaction began. This ensures that if your privileges are revoked mid-transaction, your next statement is immediately blocked. Access control is never stale.
#### Real-time Security Enforcement [#real-time-security-enforcement]
The crucial point is that permission checks use the current, up-to-date state rather than the snapshot state at the beginning of the transaction. While data operations within a transaction see a consistent snapshot defined by the begin\_lsn, security checks must always query the most up-to-date, globally committed state of the system's permissions.
This means that a separate read is performed on the permissions metadata for every statement to ensure up-to-date state. Even if a transaction has been running for a long time, its access is validated against the very latest security policies, preventing any staleness.
#### Atomic DCL Within Transactions [#atomic-dcl-within-transactions]
A powerful and unique feature of Firebolt's security model is its full support for Data Control Language (DCL) operations within transactions. You can create a new table and atomically GRANT permissions on it in the same transaction. This guarantees that security policies and schema changes are applied simultaneously, preventing any inconsistencies between access controls and database structure.
Just like other metadata changes (e.g., CREATE TABLE), DCL operations like GRANT and REVOKE are also written to the isolated transactional space. They are represented as specific objects in the transaction's WAL (Write-Ahead Log).
When the transaction successfully commits, both the new table and the corresponding GRANT are simultaneously made visible to the rest of the system using the commit\_lsn. This prevents a race condition where a table is created and visible, but its intended permissions haven't been applied yet, ensuring the system is never in an inconsistent state.
## Conclusion [#conclusion]
Firebolt delivers robust multi-statement transactions without introducing a stateful compute layer. The centralized metadata service, powered by an append-only Write-Ahead Log, ensures data durability and consistency.
This architecture allows for the atomic execution of DML, DDL, and DCL operations within a single transaction, a capability that sets Firebolt apart from platforms like Snowflake and Databricks. As a result, you can:
* Run transactional statements across different compute clusters
* Replace compute nodes mid-transaction without losing state
* Atomically update both data and schema in a single transaction
This approach preserves the core benefits of Firebolt's cloud-native design that optimises elasticity, resilience, and statelessness, while delivering one of the most powerful and highly requested database features.
# Implementing Firebolt MERGE Statement (/blog/implementing-firebolt-merge-statement)
The MERGE statement is a powerful SQL command that allows you to simultaneously perform multiple INSERT, UPDATE, and DELETE operations on a single table. It's a classic PostgreSQL feature that simplifies data synchronization tasks by
1. supporting complex conditional business logic, and
2. moving the entire data mutation operation into a single ACID transaction, thereby natively protecting against stale reads from frequently updated data sources
This blog goes deep on what it takes to build a scale-out merge operation that can run over terabytes of data.
## MERGE Statement Overview [#merge-statement-overview]
As an example, if you had a large table of students that you wanted to sync after the course registration period completed, thus making recent enrollment changes visible publicly (e.g. students that transferred in, that dropped out, or that changed their planned graduation year, etc), you could use MERGE to do so. Here's a toy example.
```sql
-- dated 990 A.D. upon school founding
CREATE TABLE hogwarts_students (
id INT,
full_name TEXT,
dob DATE,
graduation_date DATE,
owl_exams ARRAY(TEXT)
);
-- dated 1017 A.D. alongside new privacy and ethical recruitement reforms
CREATE TABLE public_hogwarts_students (
id INT,
full_name TEXT,
graduation_date DATE
);
-- running the sync from the staged source data to the public target table
MERGE INTO public_hogwarts_students AS t USING hogwarts_students AS s
ON t.id = s.id
WHEN NOT MATCHED THEN INSERT (id, full_name, graduation_date) VALUES (s.id, s.full_name, s.graduation_date)
WHEN MATCHED THEN UPDATE SET full_name = s.full_name, graduation_date = s.graduation_date
WHEN NOT MATCHED BY SOURCE THEN DELETE;
```
The MERGE statement works by comparing rows from a source (which can be a table, view, or subquery) with rows in a target table. Based on whether rows match or not, you can specify different actions:
* **WHEN MATCHED**: This clause defines actions to take when a row in the source matches a row in the target. You can specify UPDATE or DELETE operations here.
* **WHEN NOT MATCHED (e.g. WHEN NOT MATCHED BY TARGET)**: This clause defines actions to take when a row in the source does *not* have a corresponding match in the target. It is used for INSERT operations, in order to backfill the target table with rows from the source.
* **WHEN NOT MATCHED BY SOURCE**: This clause defines actions to take when a row in the target does *not* have a corresponding match in the source. This is used to UPDATE or DELETE rows from the target table, to clean up rows that no longer exist in the source.
Row "matching" is evaluated by an explicit JOIN condition. Traditionally, this is done over a set of primary key columns. However, any subset of columns from the source and target tables can be used, or any transformations therein.
You can find common use cases and examples of the syntax in our [documentation](https://docs.firebolt.io/reference-sql/commands/data-management/merge). For the remainder of this blog, however, we proceed with the low-level details of how we implemented MERGE.
## What MERGE Requires Beyond Existing DML [#what-merge-requires-beyond-existing-dml]
At a high level, MERGE needed to wrap together functionality that we already had. We already knew how to modify tablets using the logical representations of INSERT, UPDATE, and DELETE requests (physical tablet management is explained below). We already knew how to transparently modify aggregating indexes, alongside any affected table (these are the consistent materialized views that the planner uses to optimize aggregation queries: [documentation](https://docs.firebolt.io/overview/indexes/aggregating-index#aggregating-index)). We already knew, even, how to run multiple DML operations under a single transaction, since the UPDATE command reduces to a DELETE and an INSERT op. And, quite significantly, we already had battle tested distributed executions of massive joins.
So what was missing? Mainly, adding the parsing/planning layer to translate a user's MERGE query into our proven, scalable building blocks. While not breaking any of the existing code paths 😂. Let's dig into that translation.
## Ordering User-Provided Clauses [#ordering-user-provided-clauses]
While these terms are personal colloquialisms, they are useful for describing nuances of the MERGE syntax that have to make it into the implementation.
**Match category**: A set of clauses within which order matters. E.g. the *category* of MATCHED clauses, or the category of NOT MATCHED clauses, or the category of NOT MATCHED BY SOURCE clauses.
**Match clause**: A tuple of match category, optional conditional, and action. Several unrelated examples:
```sql
-- Delete action with the (default) catch-all condition
WHEN MATCHED THEN DELETE
-- Update action with a data transformation
-- Note the nontrivial condition
WHEN MATCHED AND t.cond > 5 THEN UPDATE foo = s.foo + 1
-- A different match category, with a no-op action
-- Note that conditions don't have to be meaningful 🤔
WHEN NOT MATCHED and FALSE THEN DO NOTHING
```
As defined by the syntax, ordering within each category provides **short-circuiting** (e.g. a row matching an earlier MATCHED clause condition will be sent to that earlier action, and will not be considered for any further MATCHED clauses). So within a single category, the order of clauses that the user defined is important.
But what if we look across different categories? Conceptually, it cannot be, for example, that a target and source row pair both MATCH and do NOT MATCH BY SOURCE. **Match categories act on non-overlapping sets of target-source row pairs**. Put another way, given a well-defined JOIN condition, every target-source row pair will fall into exactly one of the match categories, or into zero of them (in case the JOIN condition evaluates to false or NULL). Therefore, clauses from different categories will have no effect on each other's row sets. This means that it is irrelevant in which order the user defines the different categories, and makes no difference whether the user intermixed clauses across categories.
An aside: If we add to these observations the thoughtfully-limited number of category/action pairings that the MERGE syntax allows, and the errors that are thrown when attempting to update the same row multiple times, we get the following lemma: The MERGE syntax **guarantees that each row from the target table will be modified (updated/deleted) at most once and each row from the source will be inserted into the target at most once**. It is a nice cap on the maximal number of operations.
## Degenerate No-Op Scenario [#degenerate-no-op-scenario]
Let's discuss the base case before the inductive case: In the most basic scenario, a MERGE statement represents a DML no-op. It must pass parsing and be registered as a successfully run DML, however it adds or affects exactly zero rows. Whereas this may be discovered at runtime due to data distribution (e.g. the source and target were already in sync and there was nothing new to handle), there are certain degenerate MERGE queries that represent global no-op's and actually require special handling by the query planner.
These consist of all DO NOTHING clause actions. Consider this example, where the syntax is used in a valid though futile way. The user is literally not asking for any actions to be taken:
```sql
MERGE INTO target t
USING source AS s
ON t.a = s.a
-- single DO NOTHING branch
WHEN MATCHED THEN DO NOTHING;
```
So how do you translate a DML statement with no requests for action? Instead of defining a runtime primitive for a "non-action", we chose to catch such cases at the earliest convenience and provide a well defined DML plan. For simplicity sake, it is: INSERT INTO target SELECT \* from target LIMIT 0. And here's what the planner simplifies a LIMIT 0 down to: a local read from an empty static table.
The query plan generated:
```shell
[0] [TableModify]
\_[1] [Insert] target table: "target"
\_[2] [Projection] c_0: NULL, c_1: NULL
\_[3] [StaticTable] Columns: [][]
```
Neat! Note that distributing DO NOTHING actions across multiple categories or conditions still hits this degenerate scenario.
```sql
MERGE INTO target t
USING source AS s
ON t.a = s.a
-- Match category I:
WHEN NOT MATCHED BY SOURCE and s.name in ('Harry', 'Ron', 'Hermione') THEN DO NOTHING;
-- Match category II:
WHEN NOT MATCHED and s.color = 'red' THEN DO NOTHING
WHEN NOT MATCHED and s.iteration > 10 THEN DO NOTHING;
```
As an aside, you can also get "zero effect" queries via always-false conditionals. However, it is not the parser's job to catch those. So the following plan will still have a DELETE node in it, but note that the planner has collapsed any real scan of the source and target tables into, once again, a local "read" from an empty static table.
The query:
```sql
MERGE INTO target t USING source s ON t.a = s.a
WHEN NOT MATCHED BY SOURCE AND 7 < 5 THEN DELETE;
```
The plan:
```shell
[0] [TableModify]
\_[1] [Delete] target table: "target", tablet_row_number_column: 0, tablet_txid_column: 1, tablet_id_column: 2, source_node_id_column: 3
\_[2] [Projection] c_0: 0, c_1: '', c_2: '', c_3: 0
\_[3] [StaticTable] Columns: [][]
```
## Single Node Query Plan [#single-node-query-plan]
So what does the query plan look like for a MERGE statement with at least one interesting action? Let's try an INSERT action, where two local tables are simply generated and filled with 10 and 20 rows, respectively.
\[EXAMPLE 1] The script:
```sql
CREATE TABLE target (a int, b int);
CREATE TABLE source (a int, b int);
-- Populate 10 entries with keys 1,11,21,etc
INSERT INTO target SELECT i, i%17 FROM generate_series(1,100,10) i;
-- populate 20 entries with keys 1,6,11,16,etc
INSERT INTO source SELECT i, i%17 FROM generate_series(1,100,5) i;
-- Note that the source contains all the keys already in the target, and then some (specifically, 10 other rows)
-- Merge in new entries from the source, with an extra conditional based on the mod of the key space, and a transformation of some of the source values.
-- The condition is picked purely for demonstrative purposes. It happens to capture 6 of the 10 possible rows in the target-source row diff.
MERGE INTO target t USING source s ON t.a = s.a
WHEN NOT MATCHED AND s.b < 11 THEN
INSERT VALUES (s.a, 100 * s.b);
```
The executed plan:
```shell
[0] [TableModify]
\_[1] [Insert] target table: "target"
\_[2] [Projection] source.a, multiply_checked_0: (100 * source.b)
\_[3] [Filter] ((CASE WHEN (c_0 IS NULL and (source.b < 11)) THEN 0 ELSE NULL END) = 0)
\_[4] [Projection] c_0, source.a, source.b
\_[5] [Join] Mode: Left [(source.a = target.a)]
\_[6] [StoredTable] Name: "source"
\_[7] [Projection] target.a, c_0: TRUE
\_[8] [StoredTable] Name: "target"
```
What are we looking at? Going bottom up:
\[6, 8] At the bottom are the scans of the source and target tables.
\[7] Before we get to joining rows from the source and target tables, we introduce a constant boolean projection onto the target side, mimicking a **mark join**. Mark joins are an optimization technique that propagate, alongside every source row, the result of some condition in a temporary "mark" column. Sometimes, this row-wise evaluation, passed along to a JOIN, is actually enough to compute some subquery like EXISTS or IN, without continuing to propagate the full data set itself.
Firebolt doesn't have native support for mark joins, but we can utilize the technique by manually inserting this 'magic' constant column. The point is that we are about to compute the set of source rows that do not have matches in the target table, i.e. we really want a LEFT ANTI JOIN. However our join infra will return the full set of source rows, extended by either the matching target rows or by all NULL's. To distinguish a NULL value set because "there was no matching target row" from a NULL value set from "a target row consisting of all NULL's", we need an extra marker that will be TRUE alongside all target rows (even a completely NULL one), and NULL anytime there is no matching target row.
This constant (dubbed c\_0 in the plan) will be filtered against further up.
\[5] A left side outer join, which will return all rows in source, matched either with a row from target or with all NULL's.
\[4] A pass-through of all the source columns required for insertion, plus the target mark
\[3] A filter that encapsulates the NOT MATCHED BY TARGET clause and the additional condition that we passed. Note the use of c\_0 from before, to isolate only source rows without matches. It may look odd that the CASE statement in the filter is returning 0 and NULL, rather than true and false, but given more merge clauses, it will be classifying each row as belonging to clause index 0, 1, 2, etc, up to the number of clauses. More on that in the CASE WHEN section below.
\[2] The computation of the transform we requested on the data to-be-inserted.
\[0,1] The top level goal: we're modifying the table target by inserting some rows into it.
Ready to dive in deeper?
### Choosing the JOIN Type [#choosing-the-join-type]
Plain WHEN MATCHED clauses require simply an INNER JOIN between the source and target tables. Target rows can be updated or deleted based on conditions on the source row values, and only row-pairs with direct matches are candidates for consideration.
In the prior example, the NOT MATCHED BY TARGET clause requests a LEFT OUTER JOIN because it's looking to operate over source rows (and the JOIN inputs happen to be ordered as source table scan on the left, target table scan on the right)
More generally, NOT MATCHED BY SOURCE clauses will require the flipped RIGHT OUTER JOIN in order to be able to update/delete target rows.
If both the NOT MATCHED categories are requested - the JOIN type will be expanded into a FULL OUTER JOIN.
The JOIN type is dynamically chosen to the strictest one possible, given which subset of merge categories are requested.
### CASE WHEN: Stacking Conditions [#case-when-stacking-conditions]
Let's look at the CASE statement from that example's filter step \[3].
The MERGE syntax mandates short circuiting between different match clauses of a single category, and a CASE statement is exactly the primitive to describe that behavior. This is a PostgreSQL operator that enables describing a series of if-then-else statements. Once a row matches one of the CASE filter conditions, it will not be evaluated again. So in the previous example, we assign the value zero to rows matching the INSERT clause condition, and then filter for exactly those rows that were evaluated to zero. Not too complex. However, consider the case where there are multiple match clauses within a category, with different filters and even potentially different actions.
\[EXAMPLE 2] Given this query:
```sql
MERGE INTO target t USING source s ON t.a = s.a
WHEN NOT MATCHED AND s.b % 2 == 0 THEN DO NOTHING
WHEN NOT MATCHED AND s.b < 11 THEN INSERT VALUES (s.a, 100 * s.b)
WHEN NOT MATCHED AND s.b >= 11 THEN INSERT VALUES (s.a, 1 + s.b);
```
We want to route row pairs to four possible outcomes (fine for rows to fall through and land in the fourth and default "ELSE do nothing" outcome. At this point, the filter step will be elevated up (to sit below each of the INSERT operators), and the CASE statement will be clear within its own projection step. Note the assignment of a unique index to each of the possible outcomes, including the first DO NOTHING group:
```shell
[Projection]
source.a, source.b, multiIf_0:
(CASE
WHEN (c_0 IS NULL and ((source.b % 2) = 0)) THEN 0
WHEN (c_0 IS NULL and (source.b < 11)) THEN 1
WHEN (c_0 IS NULL and (source.b >= 11)) THEN 2
ELSE NULL
END)
```
Although previously mentioned as the recipe for "degenerate" merge queries, DO NOTHING clauses are typically quite useful. They are used to take out a subset of rows from consideration. Since match clauses are evaluated in order, having a DO NOTHING clause lifts a conditional from having to be repeated in every later clause.
This overall projection step shows that the result of the CASE is forwarded alongside the source data (dubbed multiIf\_0 in the plan). This allows DML nodes further up in the plan to filter on precisely the subset of rows that they are meant to handle.
### Loopback Shuffle [#loopback-shuffle]
So what does the plan look like for \[EXAMPLE 2]? With multiple different row subsets getting inserted (each with its own custom filter and transformation), it's getting a bit thornier to read. I have greyed out the sections that have not been changed - namely the topmost INSERT node and the base JOIN.
Observe that we are now working with a DAG-shaped plan.
```shell
[0] [TableModify]
\_[1] [Insert] target table: "target"
\_[2] [Union]
\_[3] [Projection] source.a, multiply_checked_0: (100 * source.b)
| \_[4] [Filter] (multiIf_0 = 1)
| \_[5] [Shuffle] Loopback with disjoint readers
| \_[6] [Projection] source.a, source.b, multiIf_0:
| \ (CASE
| | WHEN (c_0 IS NULL and ((source.b % 2) = 0)) THEN 0
| | WHEN (c_0 IS NULL and (source.b < 11)) THEN 1
| | WHEN (c_0 IS NULL and (source.b >= 11)) THEN 2
| | ELSE NULL
| | END)
| \_[7] [Join] Mode: Left [(source.a = target.a)]
| \_[8] [StoredTable] Name: "source"
| \_[9] [Projection] target.a, c_0: TRUE
| \_[10] [StoredTable] Name: "target"
\_[11] [Projection] source.a, add_checked_0: (1 + source.b)
\_[12] [Filter] (multiIf_0 = 2)
\_Recurring Node --> [5]
```
Going bottom up once again, beyond the base join of source and target tables:
\[6] is the projection discussed in the CASE WHEN section above, which is classifying every row-pair in terms of which merge clause it should be handled by, and forwarding that classification in the new column dubbed multiIf\_0. All the c\_0 IS NULL checks are coming from the fact that all three clauses are of the NOT MATCHED BY TARGET category.
\[5] Introduces what Firebolt calls a **loopback shuffle**, a building block that enables multiple operators to reuse the same input stream. Conceptually, this can be viewed as a form of "fork" operator which allows both \[4] and \[14] to consume the data produced by \[6]. In this case specifically, there are two Filters that are going to split the classified stream of row-pairs into sections, and pass only subsets of the stream up further to the Insert node.
Note that the "shuffle" keyword is usually reserved for distributed join or aggregation processing in which nodes need to send data to each other and hence data is "shuffled" across the network. In contrast, the loopback shuffle does not move any data off the node and simply broadcasts data from one local input stream to multiple local output streams. It keeps track of which blocks have been consumed by which output stream, and applies backpressure onto the input stream so that no consumer gets "too far ahead" or falls "too far behind". (It would lead to a memory usage explosion if one consumer was reading very slowly, whereas another was speeding ahead and asking to load the whole input stream into memory.)
At Firebolt, we use loopback shuffle for other DAG shaped plans, such as join pruning using a filtered primary index key set that is passed sideways from one side of the join to the other. More on this in a separate blog.
So now you can see why and how branches \[3,4] and \[13,14] are reusing node \[5].
\[3,4] Is filtering out the subset of row-pairs that match the second WHEN condition, and then preparing the data for insertion: WHEN NOT MATCHED AND s.b \< 11 THEN INSERT VALUES (s.a, 100 \* s.b). These are the rows that were assigned classification 1 by the CASE statement \[6].
\[11,12] Is filtering out the subset of row-pairs that match the third WHEN condition, and then preparing the data for insertion: WHEN NOT MATCHED AND s.b >= 11 THEN INSERT VALUES (s.a, 1 + s.b). These are the rows that were assigned classification 2 by the CASE statement \[6].
Note that there is no action branch to handle rows matching (multiIf\_0 = 0). These would be the rows caught by the first match clause of the query. As a DO NOTHING action, its purpose was precisely to classify some subset of rows for non-action. And so its execution is satisfied by the CASE statement assigning those row-pairs a unique classification which will not get picked up by any of the downstream filters. There is "nothing more" to do for those row-pairs.
\[2] The Union operator joins the two transformed streams into a single input for the Insert step. Handing the Insert operator the full data stream will let it best distribute the data across all newly generated tablets, thereby improving tablet quality.
\[0,1] Are the same top-level operators as before.
In a nutshell, while performing a table synchronization or a data deduplication task, the introduction of the loopback shuffle is the *key performance boost* that MERGE offers over running multiple separate DML statements. The fact that the base join can be done once and then reused by all the data mutation branches is crucial.
Now, let's see what the reuse looks like when there's even more action types requested. But before that, let's step to the side for a moment to refresh what it actually means to INSERT, UPDATE, or DELETE from a managed (internally persisted) Firebolt table.
### Tablets: What INSERT, UPDATE & DELETE effect [#tablets-what-insert-update--delete-effect]
At Firebolt, tables are persisted in units of "tablets", and this is what all DML operations ultimately operate over. A tablet is a collection of data and metadata files that are (for the most part) immutable. As data is ingested into a table, one or more new tablets are written and stored somewhere in cloud. At the conclusion of the insert transaction, our metadata service attaches those new tablets to the tablet list that defines the table.
The sole caveat for tablet immutability comes in the form of deletion vectors. Either through UPDATE or DELETE, a user may ask to delete some subset of rows from a pre-existing tablet (say all rows where column foo is divisible by 5). While we could immediately construct a new tablet that has only the rows to-be-kept-alive, that is often overkill, and we punt on that heavy copy operation until a VACUUM op is run. Instead, a deletion vector is attached to the tablet files (conceptually a bitset of which rows are still alive; implementation wise it is a RoaringBitmap). This bitset is then referenced prior to any future reads of tablet data, ensuring old rows are ignored.
If multiple mutation statements are applied to a single tablet, the previous deletion vector will be read each time, and a new one will be generated to reflect the net sum of the changes. The fact that there is only one trusted deletion version at any one time actually explains why it's important to have at most one Deletion operator per table per transaction. Just planting the seed here, but this turns out to be the root cause behind adding the UNION operators to the DML nodes in the MERGE plan in the first place.
To reiterate: other than this mutable deletion bitset concept, all files in a tablet are immutable. And DML operations operate over tablets. INSERT creates new tablets; DELETE marks subsections of a tablet as deleted, or drops tablets entirely; UPDATE applies a DELETE on all rows that are to be changed, and then inserts new tablets with precisely the affected rows and their new values.
### Multiple Action Types in MERGE [#multiple-action-types-in-merge]
Let's go all in. What if a user requests the classic "please sync my tables" MERGE statement? Insert the new rows, delete the old rows, and overwrite any duplicate rows. In other words, the source table is more up-to-date and we would like to trust it.
In the general MERGE query, of course, there could be multiple match clauses per category, and every clause could have further conditionals. We present the simplest case of one clause per category and zero extra conditionals, purely to demonstrate how the Insert/Delete nodes roll up.
\[EXAMPLE 3] Query:
```sql
MERGE INTO target t USING source s ON t.a = s.a
WHEN NOT MATCHED BY TARGET THEN INSERT VALUES (s.a, s.b)
WHEN NOT MATCHED BY SOURCE THEN DELETE
WHEN MATCHED THEN UPDATE SET b = s.b;
```
The executed plan, for readability having extracted the CASE statement like so:
```sql
multiIf :=
(CASE
WHEN (c_1 IS NOT NULL and c_0 IS NOT NULL) THEN 0 -- when matched
WHEN c_1 IS NULL THEN 1 -- when not matched by target
WHEN c_0 IS NULL THEN 2 -- when not matched by source
ELSE NULL
END)
```
```shell
[0] [TableModify]
\_[1] [Delete] target table: "target"
| \_[2] [Projection] target.$tablet_row_number, target.$tablet_txid, target.$tablet_id, target.$source_node_id
| \_[3] [Union]
| \_[4] [Projection]
| | \_[5] [Filter] (multiIf = 0)
| | \_[6] [Projection]
| | \_[7] [Shuffle] Loopback with disjoint readers
| | \_[8] [Join] Mode: Full [(source.a = target.a)]
| | \_[9] [Projection] source.a, source.b, c_0: TRUE
| | \_[10] [StoredTable] Name: "source"
| | \_[11] [Projection] target.a, target.$source_node_id, target.$tablet_id, target.$tablet_txid, target.$tablet_row_number, c_1: TRUE
| | \_[12] [StoredTable] Name: "target"
| \_[13] [Projection]
| \_[14] [Filter] (multiIf = 2)
| \_[15] [Projection]
| \_Recurring Node --> [7]
\_[16] [Insert] target table: "target"
\_[17] [Union]
\_[18] [Projection] target.a, source.b
| \_[19] [Filter] (multiIf = 0)
| \_[20] [Projection] source.b, c_0, target.a, c_1
| \_Recurring Node --> [7]
\_[21] [Projection] source.a, source.b
\_[22] [Filter] (multiIf = 1)
\_[23] [Projection] source.a, source.b, c_0, c_1
\_Recurring Node --> [7]
```
Having again greyed (\[8] - \[12] above) out the bottom most Join, as well as the topmost TableModify operator, I'd like to point out the high level hierarchy.
At the top, multiple types of DML operations (Delete and Insert), are now inputs to a single TableModify node (\[1], \[16]). This is what is enabling multiple types of data mutation to be run under the same transaction. The final query will succeed only if all children of the TableModify node succeed.
Both the Delete and Insert nodes happen to have multiple row-pair classifications sent to them, and hence are receiving Unions of substreams (\[3], \[17]).
Why are they both receiving multiple classifications, given the three original match clause actions? Because the subset multiIf = 0 , expressing the WHEN MATCHED clause, is an UPDATE action, which splits into a Delete and an Insert component. The subset multiIf = 1 has only an Insert component. And the subset multiIf = 2 has only a Delete component.
All these filters and projections are reading from a single base loopback shuffle \[7], and the base join is now a FULL OUTER (marked) join, rather than a LEFT OUTER (marked) join, because both target and source row values are required in order to execute all later branches.
Note also that there are now markers c\_0, c\_1 applied to both source and target rows, respectively, such that we can tell apart a source row of all null's, from a target row without a match in the source table, and vice versa.
Phew. There you have the composability of this query plan:
* Complex conditionals will make it into the branch filters.
* Complex data transformations (upon Insert or Update) will make it into the projections before Insert.
* As many clauses as requested will get unioned and piped up to the DML processing nodes.
What are the differences introduced by making this plan distributed?
## Distributed (Multi Node) Query Plan [#distributed-multi-node-query-plan]
We've mentioned a type of shuffle already: the loopback shuffle. Now we introduce the Hash and Key-Identity (or Homecoming) shuffles that will actually move data between the nodes of a distributed engine, at two key steps.
Let's continue using the same full-sync query \[EXAMPLE 3].
### Difference 1: Hash Shuffle of JOIN inputs [#difference-1-hash-shuffle-of-join-inputs]
```shell
\_[Join] Mode: Full [(source.a = target.a)]
\_[Shuffle] Hash by [source.a]
| \_[Projection] source.a, source.b, c_0: TRUE
| \_[StoredTable] Name: "source"
\_[Shuffle] Hash by [target.a]
\_[Projection] target.a, target.$source_node_id, target.$tablet_id, target.$tablet_txid, target.$tablet_row_number, c_1: TRUE
\_[StoredTable] Name: "target"
```
The source and target tables (being default FACT, rather than DIMENSION tables) are now distributed, and each node can read only a portion of them locally. On a multi-node engine then, given the join condition of target.a = source.a, each table's data will get shuffled across the engine, partitioned by the hash of each of those columns. More complex join conditions will hash all relevant columns from each base table.
Note that once the full data partitions land on every node, the join can proceed node-locally. Nothing about the node-local reuse of this base join (via loopback shuffle), nor the filter branches higher up, is affected.
Fundamentally speaking, the scalability of this basemost join is what enabled MERGE to be scalable out-of-the-box. Convenient!
### Difference 2: Key-Identity (or Homecoming) Shuffle of Delete inputs [#difference-2-key-identity-or-homecoming-shuffle-of-delete-inputs]
```shell
[0] [TableModify]
\_[1] [Delete] target table: "target"
| \_[2] [Shuffle] KeyIdentity by target.$source_node_id
| \_[3] [Projection] target.$tablet_row_number, target.$tablet_txid, target.$tablet_id, target.$source_node_id
| \_[4] [Union]
| \_ ... multiple branches
```
For this difference, I have to more deeply explain how delete works. For starters, a Key-Identity Shuffle sends rows to a specific node ordinal (rather than hashing row values to decide the receiver node at runtime). But why would you need to shuffle the stream of (tablet, row\_number) pairs sent to a Delete node? And it looks like we're sending each pair back to its $source\_node\_id?
Here's why. On a multi-node engine, the entire query pipeline is duplicated on every node. So every node will have its own active Delete operator that can receive row sets, write and serialize RoaringBitmap deletion vectors, etc. But a limitation mentioned earlier is that at the end of any given transaction, a persisted tablet must have exactly one deletion vector specified. Which means you cannot allow multiple nodes to delete rows from the same tablet - in this case they would upload multiple independent (and partial) deletion vectors, and all but one of those would fail to be committed to the tablet's metadata. It is like some deletions would have never taken place.
So it's great to have an active Delete operator per table per node. However each of them must handle rows from distinct sets of tablets.
The workaround is to assign to every tablet a "parent" node, here called the "source". That's the node responsible for initially scanning the tablet and then for uploading a single and complete deletion vector for that tablet at the end of every DML query. (Note that it behooves us that this node will already have the previous version of the deletion vector cached.) Applying the Key-Identity Shuffle using this "source" node id lets us guarantee the correctness of the distributed Delete step. I've always thought it a bit sweeter to call the mechanism a "Homecoming Shuffle" instead 🐑, for all the wandering lil' rows to be sent back to where they came from.
Note that this shuffle is introduced only if at least one prior layer of the pipeline ever shuffled data away from its source node. (In our case, it is the distributed join at the bottom of the MERGE plan which added the initial hash shuffles.)
And now a neat aside - to be fully transparent, implementing MERGE did not require building out or even newly instantiating this Key-Identity Shuffle layer. It was an existing building block, already used to support distributed UPDATE and DELETE queries. What MERGE introduced was the first ever time that more than one data stream could contribute to a distributed DELETE of the same table, all within a single query plan. (Consider multiple match clauses, with different conditionals, that all request the DO DELETE action.) So whereas the shuffle layer was already in place - this is actually the true reason that we started adding UNION operators to all the branches funneling into DML nodes. I mentioned earlier that unioning INSERT streams improves tablet quality. But unioning DELETE streams is truly a basic correctness requirement to ensure that all rows-to-be-deleted make it through the same shuffle infrastructure, and get marked exactly once per tablet.
## INSERT ON CONFLICT: Syntactic Sugar for MERGE [#insert-on-conflict-syntactic-sugar-for-merge]
The final cherry on top of building out this new syntax was enabling a syntactic "candy wrapper" around the merge plan to support INSERT ON CONFLICT syntax. See the [documentation](https://docs.firebolt.io/reference-sql/commands/data-management/insert#insert-on-conflict) for the full syntax spec and its current limitations.
The MERGE spec provides full flexibility on the number and order of merge clauses that you choose to interleave. However many common "upsert" patterns can be expressed more concisely. Specifically - say that you want some tuples to be inserted into a target table without creating duplicates. Once a duplicate is encountered, the only decision is whether to keep the original row or to overwrite it. PostgreSQL defines this as an ON CONFLICT DO \ clause on top of the classic INSERT statement. With options to DO NOTHING (e.g. INSERT IGNORE) or to DO UPDATE (e.g. INSERT UPDATE).
And the key observation for implementing this translation was that the ON CONFLICT syntax can be rewritten as a MERGE. Observe the following.
**Original INSERT query:**
```sql
INSERT INTO target (id, value, last_updated) VALUES (1, 'retro', GETDATE())
ON CONFLICT (id) DO UPDATE SET value = EXCLUDED.value, last_updated = EXCLUDED.last_updated;
```
**Translates to the following MERGE statement:**
```sql
MERGE INTO target AS T
USING VALUES (1, 'retro', GETDATE()) AS EXCLUDED (id, value, last_updated)
ON T.id = EXCLUDED.id
WHEN NOT MATCHED BY TARGET THEN
INSERT (id, value, last_updated)
VALUES (EXCLUDED.id, EXCLUDED.value, EXCLUDED.last_updated)
WHEN MATCHED THEN
UPDATE SET
value = EXCLUDED.value,
last_updated = EXCLUDED.last_updated;
```
Therefore, some parsing magic was all that was required to support folks' favorite UPSERT commands. Now you know what's happening under the hood when you invoke them.
## Looking Forward [#looking-forward]
MERGE is currently in public preview. Thanks to our powerful runtime building blocks, crafting the right MERGE plan made the execution "just work". Use the EXPLAIN command ([documentation](https://docs.firebolt.io/performance-and-observability/query-planning/inspecting-query-plans)) to see the query plans for yourself.
Several of our customers are already using MERGE for table sync or ingest, repeating the query pattern across many of their tables every couple minutes or hours.
Word to the wise:
```plaintext
Anywhere from 0-100% of your table can be updated on sync. To reduce churn, avoid empty rewrites (i.e. deleting and inserting a row that is exactly equivalent).
```
We're looking forward to hearing how MERGE makes a difference in your day-to-day data-filled lives. Please reach out with any performance or scalability issues that you uncover.
Thanks for learning about the underlying query plan and implementation. Query on everyone.
# Introducing Firebolt Core - Self-Hosted Firebolt, For Free, Forever (/blog/introducing-firebolt-core)
Today, we're excited to announce [**Firebolt Core**](https://github.com/firebolt-db/firebolt-core). Firebolt Core is a forever free, self-hosted edition of Firebolt's distributed query engine. It provides high-performance data warehousing capabilities that can be deployed anywhere you'd like to deploy it. Core is designed for engineers who need low-latency, high-concurrency analytics. You can use it for one-off ad-hoc analysis, to power 24/7 production workloads, or anything in between. The only thing it cannot be used for is to build a hosted SaaS solution which competes with Firebolt's managed service.
Firebolt Core contains most of the features present in Firebolt's managed service. Most importantly, the query engine in Firebolt Core is exactly the same as in the managed version of Firebolt. Without any handicaps or limitations. And it is free. Forever.
## What Is Core? [#what-is-core]
Firebolt Core is a production-grade, distributed query engine you can run locally on a single laptop, in the cloud, in a massive datacenter, or anywhere in between. It comes with no cost and no lock-in. It is incredibly easy to get started with Firebolt Core. Just type the following to start a local Core instance on your machine and drop into the CLI.:
```bash
curl -fsSL https://get.firebolt.io/ | bash
```

You can find Helm charts to run a multi-node setup in our [GitHub repository](https://github.com/firebolt-db/firebolt-core).
What makes Firebolt Core particularly exciting to us is that it's not a stripped-down or limited edition of Firebolt. It's the same engine that powers our cloud platform: a fully distributed, vectorized query engine with aggressive indexing and modern SQL support. It's built to serve large analytical workloads and can handle the same scale whether you're running on a local dev machine or across a Kubernetes cluster. We believe that there is no real alternative in the market today. With Firebolt Core, you're not just getting a taste of the engine, you're getting the full experience.
## What Can You Do With Core? [#what-can-you-do-with-core]
Firebolt Core excels at handling a wide variety of analytics workloads and it is built to handle production workloads: the query engine is the same as the one used by Firebolt's production customers today.
Firebolt Core can power data and AI applications and take on highly-concurrent analytical workloads. With Firebolt's optimizations and features like [aggregating indexes](https://docs.firebolt.io/overview/indexes/aggregating-index#aggregating-index) and [subresult reuse](https://www.firebolt.io/blog/caching-reuse-of-subresults-across-queries), you can achieve millisecond latency for your queries on Core. We use Firebolt Core for our [Clickbench](https://benchmark.clickhouse.com/) submission, where as of the time of publishing this blog, it holds the #1 spot.

You can also power your ELT workloads at scale. Core [integrates natively with Iceberg](https://docs.firebolt.io/reference-sql/functions-reference/table-valued/read_iceberg#read-iceberg), and is easy to [set up for scale-out processing](https://docs.firebolt.io/firebolt-core/firebolt-core-operation/firebolt-core-deployment-k8s). This empowers you to build a wide range of high-performance analytics applications.
## Why Did We Build Core? [#why-did-we-build-core]
Why give away the same high-performance engine we've spent years building? Because this move actually makes business sense for us. We believe it will mutually benefit both the engineering community and Firebolt.
Data architectures are becoming more modular. More teams are opting for open table formats, decoupled storage and compute, and hybrid environments that span public cloud, private cloud, and on-prem. Firebolt Core fits into this vision cleanly. It supports Iceberg out of the box and plays nicely with object storage. It's a flexible foundation that you can adapt to your stack, rather than the other way around.
This makes it easy to adopt Firebolt for modern software projects. No matter whether you start your journey on a laptop, a VM, or a local cluster, you can leverage Firebolt from day zero. For projects that later decide they'd rather not manage infrastructure or want seamless scale and support, Firebolt's fully managed cloud offering will be there. We want you to start where you are, and grow with you at your own pace.
We're living through a shift in how infrastructure is consumed and trusted. More than ever, developers want to try before they buy on their terms. With Firebolt Core, they can. It's fast, free, and self-hosted. If you've been curious about what Firebolt can do, there's never been a better time to try it.
Firebolt Core is built to respect the "try it before you buy it" paradigm and the shift towards increasingly technical and knowledgeable users who have no issues with self-managed solutions. Too often, powerful tools are locked behind paywalls, complicated procurement processes, or opaque platforms. We wanted to do something different: give you something you can use today, without needing to talk to sales or commit to a contract. We trust that once you experience the performance and simplicity of Firebolt, you'll understand why we built it.
We're also hoping Firebolt Core becomes a platform for learning and innovation. Whether you're building an internal analytics stack, testing new data modeling approaches, or simply exploring what's possible with Iceberg and SQL, Core is there for you. We're excited to see what use cases and tools the community builds around it, and we plan to support and learn from those efforts.
## Why Isn't Core Open Source? [#why-isnt-core-open-source]
We want to address the question that many of you have: why isn't Firebolt Core released as an open-source project? This is something we've debated at length internally as well.
We want Firebolt to be a great business. With GenAI accelerating software development, defensibility is more difficult than ever. Our query engine is built on years of innovation in database performance by Firebolt's engineering teams. GenAI can't replicate that work yet, but open-sourcing it would make it easier to copy and commoditize. To protect that edge, we chose to keep Core closed-source.
We also want Core to be a full-blown version of our query engine that can power your production workloads. If we open-sourced Core, it would either have to be a limited version of the product, or we would hurt our ability to build a great business.
There's another reason this decision makes sense: database lock-in is becoming less of a concern. Open table formats like Iceberg are democratizing data access, and we're betting on that future. Core can [directly read from Iceberg tables](https://docs.firebolt.io/reference-sql/functions-reference/table-valued/read_iceberg). With a single [COPY TO command](https://docs.firebolt.io/reference-sql/commands/data-management/copy-to), you can export data to open file formats—and soon, directly to Iceberg. Firebolt aligns with Postgres for its SQL dialect, making it easy to move queries between systems.
This means that using Core doesn't require an unreasonable leap of faith. We mean what we're saying, but you don't know us. But you're also not betting your career on us, as moving away from Core is easy.
Engineers powering production workloads want a database that's efficient, easy to use, gives them control, and doesn't lock them in. We hope that for many of you, this will be Firebolt Core.
## Conclusion [#conclusion]
We're releasing Core because we believe that giving away our query engine for free and achieving commercial success are not in conflict, but rather will be a virtuous cycle. Free doesn't mean unsupported: it means inclusive. Firebolt Core is our way of giving to the community, helping to democratize high-performance analytics. We want to earn your trust, and look forward to building together.
If you want to chat with us about Core, [join us on our Discord](https://discord.gg/UpMPDHActM).
# Is Self-Service BI a False Promise? Lei Tang of Fabi.ai Thinks So (/blog/is-self-service-bi-a-false-promise-lei-tang-of-fabi-ai-thinks-so)
In this episode of The Data Engineering Show, host Benjamin interviews Lei, Co-founder and CTO of Fabi.ai, to explore how AI-native BI platforms are reshaping data analytics and empowering non-technical users to derive meaningful insights from complex datasets.
Listen on [Spotify](https://bit.ly/4n2s03T) or [Apple Podcasts](https://bit.ly/4oSCd4R)
\[00:00:00] Lei: For the past decade, it's really difficult to make sure the self-service BI can work. And then now with AI, the worst part is that it can run properly, but the numbers are wrong. Right? So you you immediately, like, might lose some trust with the person, with the user.
\[00:00:18] Benjamin: Hi. This is Benjamin. Before we start with today's episode, I wanted to quickly reach out on a personal note. We just launched Firewall Core. Firebolt Core is the free self hosted version of our query engine. You can run Core anywhere you want, from your laptop to your on prem data center to public cloud environments. Core scales out, and you can run it in a multi node configuration. And best of all, it's free forever and has no usage limits. So you can run as many queries as you want and process as much data as you want. Core is great for running either big data ELT jobs on, for example, iceberg tables or powering high concurrency customer facing analytics on big datasets. We'd love for you to give it a spin and send us feedback. You can either join our Discord, enter our GitHub discussions, or you can just shoot me an email at [Benjamin@Fireball.io](mailto:Benjamin@Fireball.io). We'd love to hear from you. We added a link to Fireball course GitHub repository to the show notes. And with that, let's jump straight into today's episode. Hi, everyone, and welcome back to the data engineering show. Great to have you on. Today, I'm really happy to have Lei joining us, who's the cofounder and CTO of FABI AI. Elad is on vacation today, so he can't join us, but he'll, I think, be back ten to four the next show. But, yeah, great to chat with you today, Lei. Tell us what you're doing at FABI, and tell us how you actually got to starting an AI native BI company.
\[00:01:35] Lei: Sure. Yeah. So thanks for having me here, Benjamin. So right now, I'm building FABI AI. So it is a AI native BI platform that essentially combines SQL, Python, AI altogether so that it allows anyone to do vibranetics. So no matter where your data lives, it could be in a database data warehouse, but it could also be a Excel file, a Google Sheet, or even just a API from application. You could be able to pull the data through Fabi and then be able to join this data from different sources together to do some analysis. The cool part is that now within FABI, you can really leverage AI to supercharge your data analysis, like, 10 x faster. So our users include not just data engineer, data scientist, data analyst, but also, like, say, other data practitioners such as product managers, founders, gross marketers, operation team. So anybody who want to embrace data to make decisions, you can use Fabi to do the analysis. Right? So that's about Fabi. So about myself, I've been working in this data domain for quite a while, over a decade. So I was trained as a computer scientist, and then I got my PhD in machine learning. So while in graduate school, mostly, I use a method loading CSV files, doing some research. And then later, like, I joined Yahoo. At that moment, like, big data, like, Hadoop, Spark, that has been the main thing. And later moved on to work on sales forecasting analytics and help growth as well. So that's where I get exposed to NoSQL, SQL database, and Net XDB. Right? So the one common theme I have been experiencing is that normally would work with other business stakeholders, could be marketing, could be operations, could be sales. And then on one hand, they have lot of questions, like, say, coming to the data team and then ask some questions. But then our data team normally is always underwater. So that so many things need to do. So never can satisfy the requirement. So that's why when I was at Lyft, I actually tried a few different attempts really trying to empower my business partners. We tried to do some SQL training for our marketing team. So I would say it was a very limited success, probably only quite a few. After training, be able to really, like, use SQL by themselves to do some analysis. But on the other hand, my team, actually, we build all these dashboards like reports, but we'll check the log. Most time, we'll build a dashboard for somebody for certain type of questions, but then it would never be used anymore. So whenever there's another question, they would still come to us. So that's almost, like, always a pain for my team, but also for my business partners. That's why in 2022, when AI really kinda chat to be was out, I said, man, this is the time, this is the moment to tackle problem. Like, my team, myself has been suffer for a long while. So that's how we started Fabi. So it has been a fantastic journey so far.
\[00:04:40] Benjamin: Okay. Nice. That's amazing. So I like this term you said in the beginning. Like, I was like, VIBBI, right, which kind of is makes perfect sense. And, like, you have such a prolific career in data across organizations. Right? Like, at Walmart, at being chief data scientist at Clari, being director of data science at Lyft. It's like but you were always in these, I think, very traditional out of world use to run. Right? Okay. Someone needs a specific dashboard. Someone needs a specific type of, I don't know, CXO report that they look at every morning to influence certain decision making. It's like, all of a sudden, you wake up one morning and you at least in principle, any person in your organization should have access to be able to answer any question about the data that's collected. It's like, how do you even build a data platform like that? What are you doing different now with Fabi than you would have in a traditional BI platform? Like, I don't know, Looker, Tableau, Sisense, etcetera.
\[00:05:32] Lei: Yeah. Very good question. So if you talk to anybody working in the BI space, like self-service BI, that has been termed for maybe for the past decade. But I have to say that is a false promise.
\[00:05:43] Benjamin: Right.
\[00:05:43] Lei: Yeah. Is special for data engineers. The one key point is that there's a lot of upfront cost in order to set up this service BI to work. Right? So you have to define all the data semantics and potentially centralized data and make sure the data is in a good quality, define the semantic in certain, like, using LookML or maybe some other type of language. And then you want to encourage your business stakeholders, like less technical partners, to use a product in order to do some self-service analysis. But then on one hand, most organization either documentation or semantics is quite missing. And then on the other hand, most organization, the data, like the schema, the kind of business, like metrics and logic, has been constantly evolving. So it's really difficult to catch up with what is going on within the business. And then on the other hand, you will see the more senior level the leader is, the more likely they would just refuse to go to a dashboard and check out the metrics. They would prefer to okay. Just send me the numbers or send me a Excel file so I will come up with that. Right? So and, moreover, most of the time, these dashboard are very static. Of course, you can put some input filters, but then people always have some customized, like, follow-up questions. Oh, there's some metrics going on, and I can zoom in into certain product lines or certain regions. But moreover is that the BI dashboards is static. It focus on one type of analytics that's like describe what is going on. But the more likely, especially from the company, from the organization, they care more about those who have a why question. What can we do about it? Like, for example, last week, the metric went up by 10%, like, for new user onboarding. What what's going on? Is it we did something right or, like, is something else? So there's a lot of, like, hypothesis testing. That's why you have all the data scientists need to zoom into, kinda say, oh, potentially, because we launch a new product or maybe sometimes it could be like the ETL pipeline has a bug. We need to kinda make sure the numbers actually make sense. So for the past decade, it's really difficult to make sure the self-service BI can work. And then now with AI, one attempt is to let people talk directly to the database, which I would say that's, like, highly not recommended because enterprise data tends to be very noisy, very complex. And then on one hand, the data semantics, of course, like, you define all the database schema, you can get some information there. But more importantly, like, lots of this business logic is spreading our course different teams, different individuals. So if you just let anybody to talk direct to your database, most likely, you will get kind of a wrong result. K? Using all these AI and ARM, they have no problem writing syntax crack SQL queries, but the worst part is that it can run properly, but the numbers are wrong. Right? So you immediately, like, might lose some trust with the person, with the user. Moreover, as I mentioned, like, SQL is only one part of the analysis. It's almost like the full app to pull the data in. But then afterwards, you might run some statistic, like, correlation analysis and then potentially even, like, pulling some machine learning model. Like, I want to forecast look based on the trend, forecast what would be the numbers, like, next month, next week. So AI actually today can make them very easy to use. So that's what, like, Fabio has been really focused on. We're saying that we really want those data team to be able to, like, say, what type of data is exposed to, like, say, less technical folks. And then you can your customer instructions and the configurations, like, say, semantics and business logic. And one difference compared with the past BI platform is that right now, you can define the data semantics or business context in a very fluent way. Like, you can just upload a context, and then the AI can really understand well. Rather than in the past, you have to go through all these tables, like, configured one by one. But right now, you can do the customization. And one more step further is that the AI can learn from the interactions. Right? So depending on how people interact with Fabi, with the product, and also what type of queries, what type of metric has been using the dashboard, and then the AI can learn from the past interactions and then be able to answer, like, relevant or similar questions in a very accurate way. So that's what we are super excited about, like this viburnetics. And now you can have data team to really focus on, make sure the data is in a good form for the AI to use, But then the business stakeholders, data practitioners, they can ask question, but they have some boundary, like, which is set up by the data team so they can trust, like, the model output. And most of the time, when you have a sense about the business by just looking at the numbers or just by looking at the chart, you immediately get a sense saying that, okay. This doesn't look right. And then the AI can figure out, okay, maybe it's doing some top accounting, doing the drawings, or maybe something else. Right? So I think it's really powerful to allow anybody to bounce ideas, like, go through multiple iterations quickly to see, like, get the numbers, get the analysis as one needs.
\[00:11:07] Benjamin: But then, basically, like, in that world, for me to also understand where you think that is going, okay, we'll have a data team that basically procures datasets and important business metrics, and that might tell a tool like Fabi, hey. This view contains, I don't know, our, like, annualized revenue broken down by month. Whatever. Right? And, like, then Fabi basically gets that context provided by the data team. But when a CXO just asks random questions about their business, the goal of the Vine BI tool basically becomes to match that human question to the entities or underlying kind of things provided by the data team. So you're still ultimately in the business of building, like, a semantic model, data model, all of these things. But all of a sudden, it's not just this political thing within the organization. It really becomes that tool that empowers all of your AI vibe interactions with data. Is that accurate?
\[00:12:09] Lei: Yeah. So in order to build AI, native BI, I would say the focus should be how human interact with AI. That's why, like, many of these efforts, like, for example, you wanna define semantic layer. But right now with AI, you probably want to define in a markdown text, which is more appealing to the AI to consume. And then combined with some of these MCP servers, you can pull, like, external documentations or, like, say, Confluence page or maybe tickets. All these can be combined to provide the context. So the semantics of itself becomes part of the context for the AI engine. And then on the other hand, as I mentioned, is that the AI itself should learn from past interactions. Of course, like, as the data team, you can configure, like, custom type, like, instructions, like, say, or really emphasize, like, how the metric is defined or, like, say, only focus on certain type tables. Take myself, like, I used to like, when I was working on left, I searched for some revenue, kinda, like, a number. And then immediately, I got, like, 200 tables containing the revenue. Right? So in the end, what I did is, like, in the past is to talk to the person to really understand which table, which column I should use. But right now, you can just use free form text to config for that and the pass to AI as a context. On the other hand, as I mentioned is that the AI should be able to learn. It's like you consider the AI should have some form of memory. By learning from these interactions, it could somehow remember, okay, these are the context. It's almost like you have a new hire in your team gradually as the indirect other coworker.
\[00:13:47] Benjamin: We're using Devon, for example, to make, like, documentation changes. That's similar. Right? Like, it learns over time, formalizes knowledge to do better code changes in the future. So it makes perfect sense to have a similar way of generating knowledge or context through interactions in Vibe BI tools. Nice. Super interesting. So okay. Now I understand. Fabi connects to a bunch of data sources, maybe even, like, Google Sheets, Postgres, Snowflake, etcetera. It can produce data in multiple shapes as well. It can produce your Google Sheet. It can produce your Slack message, etcetera. Tell us about your internal data stack. What connects these two things? What do you run internally to connect the different pieces of technology? How's your basically data stack that powers Fabi as a product?
\[00:14:36] Lei: Yeah. So the underlying tech structure is that we build this agentic flow for our AI system. So one part is, like, this context engineering. So it would have the rack system to retrieve from all these, like, different contacts. Like, could be database schema metadata, could be, like, a past query examples. And, also, like, you can upload any documentations. So all these would be part of the context to be retrieved for in general, the answer. But on top of that, we are also working on this MCP server so that you can essentially connect with your own MCP server depending on your organization, like, where your documentation lives. Right? So you can just connect with, like, a Notion doc, like a MCP server, then you can just, like, pass all this, like, semantic information or, like, some business context probably. So that's one. Then the other part is that in order to make a agentic flow, we have all these dry runs to ensure that the whenever this AI generates some code, it could be SQL, it could be Python, it will run and we verify this, like, first of all, there's no syntax error. And then it will dry run-in a way so that it'll verify with the database or, like, table or API schema, and then be able to present you the result in the chat interface. Behind that, we have this containerized environment. So it's like a kernel environment, like, to allow you to run SQL, allow you to run Python. So that I think this is, like, especially when talking about many of these organizations, security is actually a key issue. And then we want to ensure all this code are running in a containerized environment so that you don't need to worry about, oh, this accidentally, like, have some malicious code, like, swipe out swipe out kind of your database. I think there's many interesting story about that.
\[00:16:24] Benjamin: When I run SQL in the Fabi environment, in this kind of container environment, what query engine powers that? Or it's like the SQL in your source system. So it might be Snowflake SQL. It might be Databricks SQL. It might be PostgreSQL.
\[00:16:40] Lei: So it really depends on the SQL dialect because our text tag has primary in kind of Python. So all these BI platform, we kinda we strongly believe that BI should be stored as code. So we can do versioning. You can manage the workflow, how to run it properly. Right? So in the back, like, everything's converted into some form like a Python code. And then to connect to to the database data warehouse, we essentially use SQL Acme, and then you can use a different to pull in. Once you pull in, it becomes, like, a pandas or, like, a pullers, like, data frame, and then you can do subsequent analysis or visualization as you want.
\[00:17:18] Benjamin: Ah, okay. Interesting. So everything is basically pulled in from sources through data frames. And then if I wanted to join Google Sheet onto my Postgres table or something, that would actually not happen within the SQL ecosystem. It would happen more, like, within the Python ecosystem.
\[00:17:34] Lei: Yes. And the one thing to add is that our kernels actually is stateful, so the keep all these variables in memory. On top of that, we also dynamically cache the data. So you don't want to constantly send all the query to a Snowflake or, like, a BigQuery instance. So it's all cached in certain way so that we based on all these code blocks, we determine the dependency and then see whether or not this code accurate needs to refresh. And then we select it, like, say, oh, this code blocks need to run. This one actor can just reuse the cache on this or in memory so that you don't need to worry about that, like, blowing up your, like, say, data warehouse bills. And then it also, like, keep the latency really fast. Like, especially when you're doing this Vibe Analytics, it's like you're constantly on the go. It's like, try out different ideas and see the result. Right? So this is actually really critical as, like, you have a short latency and be able to try out different ideas, run the analysis, especially when you are working on, like, connecting with different data sources together to run the analysis.
\[00:18:37] Benjamin: Right. Nice. So super cool. Looking ahead now, right, like a year from now, like, Vibelytics and so on, what are you excited about? What's gonna happen in the space? What are you guys working on at Fabi that you're excited to launch soon? Look a bit into the future. I know it's super hard in the space that's moving that quickly, but would love to get your take on that.
\[00:18:57] Lei: Sure. Yeah. So we believe that, first of all, within a couple years, Vibe Analytics, like Vibe building, that will be the norm for anybody to interact with the data, run some analysis. And then we believe that, essentially, this BI system or, like, AI BI system would be more like a agent, and then it'll actually looking for, like, business opportunities and insight and surface to you. Right now, you can see, like, all these BI dashboards is, like, static. You have to go to the dashboard in order to see what's going on. But in the future, it would say, look at all these dashboards, look at all these, like, metrics. And the potential surface is like, oh, there's something going on. I want to run some analysis for you. These are certain things, like, you need to pay attention. So you almost have somebody to regular review your, like, product house, business house, and then be able to surface the insights and the opportunity for you. And then one more thing is that the interface between AI and also human beings, We believe that will be some new invention, like, a new form will be coming. Like, today, we are still getting used to this chat interface and, like, some form of static BI dashboards, but it could be way more interactive and dynamic. One idea is that you may kinda have some debrief, like summarize everything into a podcast and deliver to you, or maybe just like you generate all this exact summary in a few slides. And then so we believe that a lot of things will be coming. We I would say, reshape how people interact with data and be able to deliver the insights as they want.
\[00:20:31] Benjamin: Nice. Yeah. It's a super interesting space, Lei. I look forward to continuing to just watching what you build at Fabi. Excited to see how you also help build the future around Vybelytics. Would love to stay in touch. Thank you so much for being on the data engineering show. It was a pleasure having you.
\[00:20:46] Lei: Thank you for having me. Yeah. It's a really nice conversation. Yeah.
\[00:20:50] Benjamin: Awesome. I feel the same.
\[00:20:52] Outro: The data engineering show is brought to you by Firebolt, the cloud data warehouse for AI apps and low latency analytics. Get your free credits and start your trial at firebolt.io.
# Joe Reis and Matt Housley on the fundamentals of data engineering (/blog/joe-reis-and-matt-housley-on-the-fundamentals-of-data-engineering)
After co-writing the best-selling book 'Fundamentals of Data Engineering', Joe Reis and Matt Housley joined the bros for some much-needed ranting, priceless data advice, and good laughs. So why are we still talking about providing business value and dashboards, even though we don't really have anything new to say? If there are so many great tools in the data stack, why are we still so troubled? How can we focus more on things like data governance and data quality that'll actually push the industry forward?
Listen on [Spotify](https://open.spotify.com/episode/0rEzqKOvhOYVUQxji5DJlg) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/joe-reis-and-matt-housley-on-the-fundamentals/id1561927688?i=1000626948271)
**Benjamin:** Hi, and welcome back, everyone, on the Data Engineering Show. Today, we again have two amazing guests, kind of Joe and Matt. You probably know them from the book Fundamentals of Data Engineering, which is really well known. And we again have our awesome co-host, Robert, joining in today, because I know you are super excited about getting to talk to Joe and Matt again.
**Joe Reis:** I'm just getting off a good start. All right.
**Benjamin:** So yeah, welcome everyone to the show. Perfect. Rob, take it away.
**Robert Harmon:** Oh, straight to me, huh? Well, it's great to see you guys. You again, Joe, and it's great to meet you for the first time, Matt. Really, we try and keep things fairly low key in these and try and stay out of trouble if we can, but that doesn't always happen. But I do have some questions, and I figured, I've got you guys here and you really do work the entire industry rather than a single product or a single content, you know.
**Matt Housley:** Great to meet you.
**Joe Reis:** Hehehe
**Robert Harmon:** consulting firm, etc. So you've got a different visibility than many of the people that we get to talk to. The other thing I noticed is you guys both have a very long background in data analytics, but you decided to jump over the dark side with, you know, guys like me and get into data engineering and data architecture.
**Joe Reis:** Yeah, that's... we can talk about that.
**Robert Harmon:** So that doesn't sound like a subject you want to talk about.
**Joe Reis:** No, it's fine.
**Robert Harmon:** So how did this occur?
**Matt Housley:** Joe, do you want to go first? Do you want to take this one this time?
**Joe Reis:** I mean, I think we both jokingly call ourselves recovering data scientists. I mean, I think Matt and I had also, I mean, we could, we could write code before. I mean, we kind of grew up around computers and all that fun stuff. But what we realized was, especially with data science and analytics and so forth and the popularity of data, especially the 2010s, you know, you'd see a lot of companies hiring data scientists and I think forgetting to build the foundation that would help data scientists succeed. And so that's, I think it was a lot of our.
**Robert Harmon:** Mm-hmm.
**Joe Reis:** I suppose. So we joined the dark side, right? We found data engineering is actually, I guess it wasn't really called that back then, but it's a sudden necessity. You have to, you know, you could basically not do your job or you could figure out ways to make yourself successful. And I suppose that's at least how I envision it. I don't know. What about you, Matt?
**Matt Housley:** Yeah, yeah, I think it was a combination of factors. So when I was doing a lot of analytics and data science work a few years ago, we had Teradata and Hadoop both on-prem. And both of those systems at some point started becoming a bottleneck for me. And this sort of coincided with this era when a lot of data processes were moving to the cloud. And so I got to participate on big data POCs in both AWS and GCP and see the possibilities of those technologies in terms of scalability.
So now I didn't have to buy all the nodes that I needed to run a really big job, or I could dynamically scale up capacity to support more analytics instead of having just a static system that I pay for upfront. And so that's kind of what pulled me over to the engineering side of seeing the possibilities of those tools and how those could support analytics and data science. And so...
Like both Joe and I are familiar with on-prem systems and we do trainings around those as needed where companies have a need for those, but I think our real emphasis too is the possibilities of the cloud and scalability, et cetera. And how those really, when we say we're recovering data scientists, it's not out of disdain for data science, it's rather out of the possibilities that data engineering can bring to machine learning, data science and analytics.
**Benjamin:** Thank you.
**Joe Reis:** Speak for yourself now, just kidding. Ha ha ha. I'm joking, geez. Sort of.
**Matt Housley:** I mean, disdain is fun, right? It can be a good brand online.
**Robert Harmon:** If it helps any recently, I got to use the phrase recovering architect for the first time in 20 years. I'm not the architect of anything. I'm not an integration architect. I'm not a solution architect. I'm not even a data architect anymore. So I'm recovering as well. I make, you know, touch base with you later on notes on how to get through this transition.
**Joe Reis:** Mmm. What's up? Yeah, yeah. I mean, we could, there's a support group out there. You know, we can all get together and, you know, really, just help each other out here in our, in our recovery. But no, it's recovering architect. Did it feel cathartic when you started calling yourself that?
**Matt Housley:** Several, yeah. Grind to our beer, yeah. Yeah.
**Robert Harmon:** Uh, it's new, man. I'm still trying to grow into the idea. So it's, you know, it's like buying a new car. It works kind of like the last car. It's just really, really different. We'll figure it out.
**Joe Reis:** The recovering title sort of has a new car smell to it too, for a bit. And it wears off. Then you're, then you're sick. Then you're just a car mudgeon. Yeah. So.
**Matt Housley:** Yeah, yeah.
**Robert Harmon:** Bye.
**Matt Housley:** Yeah, give it time. Just like modern anything. Yep, yep.
**Robert Harmon:** Hey, okay, so I do have some of those tendencies. You got me there already. Maybe that's part of why I'm looking to walk away from that title. Speaking of, you bring up the term curmudgeon and this may touch a little bit on that. We've been through some stuff in the data industry. You and I, we're of similar vintage. We've seen some things. I'd like to think we're kind of coming off that whole big data thing and we're...
you know, strategically shifting as an industry. So, you know, phrases are popping around, things like we need to deliver value, for instance. We've seen this a lot lately. Now, I like to think I've been delivering value for a lot of years in different ways, different methods, because my job as a data architect is to get as close to the customer's experience as possible and truly try and influence, reach out from my little data pit to make that happen.
Where do you guys kind of see this going in the upcoming year to five?
**Joe Reis:** I don't know, Matt, you're writing a book on data value. You should speak on this.
**Matt Housley:** Yeah, I like what you say about getting really close to the customer. I think in terms of delivering value, that's one of the main channels is like really connecting with the customer. And I think that's where we've seen a lot of really exciting work in data recently. It's also getting close to supply chain. It's getting close to the sea levels. It's getting close to all the goals of your business, right. And actually try and deliver value in those areas. I think where we saw data science, big data engineering, data engineering go astray over time.
**Benjamin:** Thank you.
**Matt Housley:** is the kind of gee whiz aspect of technology and new domains, right? So data science craze a few years ago, everyone wanted to be a data scientist, companies wanted to hire data scientists, but that customer focus was often lacking. So it's just like, oh, what cool thing can I do with Kaggle data? Now can I do that with my data? But wait, what am I actually trying to do for the business or what am I trying to do for the customer? That often got left out of the conversation.
So I hope that now as the market is tightening, as the job market is a bit more tight in the tech industry, we're actually thinking about those questions both for data science and data engineering and trying to do things that the business actually needs. I don't know, what are your thoughts on this, Joe?
**Joe Reis:** I don't know, I'm tired of the word value, right? And sorry that you're writing a whole book on this, but it's talking to, I think, Malcolm Hawker yesterday and he asked me about business value. And I said, if I keep hearing this or reading this on LinkedIn, I'm probably gonna jump off a bridge pretty soon. It's just, we've been talking about the same stuff for decades now, right, Rob? And it's like, it's the same, I mean, I feel like I'm just like an old folks home where you just need to keep talking about the same old war stories over and over again. And it's...
**Robert Harmon:** Yeah.
**Matt Housley:** Good old days. Yeah.
**Joe Reis:** Yeah, yeah. And I just hope that we can move past this. I mean, my dream actually is that we just stopped talking about data. Right. And that's, I think, when you, when you finally delivered values, when you don't have to acknowledge it, because you're just delivering it. You know, you don't have to like rant about it like, hey, I'm delivering value. Because it's like, if you have to scream that far, you know, scream that loud from the rooftops, and you know, I think that there's an inverse correlation to the amount of data or value that you're probably delivering, in my opinion. You know, I mean, you don't see your, your
**Robert Harmon:** Right, right.
**Joe Reis:** accountant saying, oh, I deliver value, right. And they just do books and then your books are done. That's pretty easy. And I, I just hope, you know, over the next few years, I just hope that, you know, data becomes, um, a lot more silent. Right. And I wrote about this in one of my blog posts where I, I think I said it's, um, you know, stop using the word data when you're talking to the business. I think that's, if we can reach that goal, you know, or we said that ideal or data sort of, um, it just happens and we're just delivering whatever value we say we're delivering. Then I think that's a win.
**Robert Harmon:** Yeah.
**Joe Reis:** If that happens the next five years, super duper. Now I do feel, um, we are at sort of an interesting moment as an industry, right? There's a lot of attention being paid to things like AI now, you know, AI sort of jumped the shark and I think this is, if we're ever going to get anything right in this industry, now's the time to do it. We don't have this opportunity that often or things like data quality, governance, management, all these enterprisey things are suddenly like, you know, you need to get these things right for AI to work. Um, then my concern is if we can't get this right.
**Robert Harmon:** Mm-hmm.
**Joe Reis:** When are we gonna get it right?
**Matt Housley:** Yeah, and that's the question, right? Can we do something different this time or are we just gonna repeat history again and again as we have with every new data fact?
**Joe Reis:** We'll see. I mean, I was joking with somebody, I think it's Malcolm's post too. He posted something about like large language models and data governance. And at this point, I'm kind of like, whatever Hail Mary you need to do to make data governance work, go for it. Cause I think at this point, we keep trying to do the same stuff over and over again. And that's the definition of insanity really. The success rate on these kinds of projects is not that great. So I mean, I'm hopeful that we can finally figure things out, but we'll see. We'll see.
**Matt Housley:** Yeah, I guess we will.
**Robert Harmon:** You know, you raise an interesting issue there, because when I jumped into the data warehouse world way back in the day, the success rate on data warehouse projects was dismal, 10%, 15% tops. It just didn't happen. And I'm looking at the world today, and you do bring up something interesting, Joe, is are we that much better today?
**Joe Reis:** Well, 10, 15%, that's 85, 90% failure rate, right? According to whatever criteria, and that's sort of the stats that Gardner always throws out and all the other pundits, you know, so, well, you just need to move the goalpost for what success looks like. So, you're suddenly just killing it. So, I don't know, but you're absolutely right. I mean, but you know, walk me through this over, Robert. I mean, back in the day when you got into the industry,
**Robert Harmon:** Well that's not hopeful.
**Joe Reis:** whenever that was data warehousing, you know, what was, what were some of the contributing factors to success and failure back then.
**Robert Harmon:** I can't speak for the entire industry because I was so busy in my own project that that's all I could see. I put blinders on because it was such a pain. Um, very, you know, we did really well. It was just a really big project. I got through in because I had somewhat of a background in process management. And that married really well with the business that we were, we were managing. So I could take that process management background, marry it with the wonderful things I learned from.
**Joe Reis:** Yeah.
**Joe Reis:** Mm.
**Robert Harmon:** Bill Inmon and company through reading and start to build out structured strong data warehouses that met the business's needs. So when I started really, what I was looking at was the outcomes. You know, everybody talks about all kinds of things like models and structures and streaming and whatnot. And the customer doesn't care. All the customer cares about is that they get what they want on time. With high quality.
**Joe Reis:** Mm-mm.
**Robert Harmon:** and that we can do it repeatedly. So that's where I set my sights with that project is, okay, I'm gonna map out everything that touches a customer in this company. And that's all sorts of things, whether that's customer service calls or product delivery or any of these things. So I map out these processes and then I apply those same four basic metrics to it. Did it happen on time? What's the backlog? What's the volume? What's the rate? The very boring mathematic process.
But since it was so business process oriented rather than technology oriented, it gave me a lot of latitude within the warehouse itself to just things like software and hardware, nobody cares, we just need to hit these, you know, these things. So that was the goal way back then. And it continued to be that way, at least in my world for at least another decade.
But then things went a little silly for a while. I like to think we're slowly trying to scratch that back. Because if I can deliver on those things, then the customer's happy. The business has no choice but to succeed. Now it's not gonna happen overnight. They'll improve, the company will improve performance 3% per month forever. But if I can get them on that path, I've got at least a path to success. Does that answer your question, Joe?
**Joe Reis:** Yeah, it's a good perspective, I think. Yeah, it's interesting. I wrote a blog post a couple of weeks ago too about, you know, we have no shortage of great tools at this point in time. I think we sort of have too many great tools actually. It's almost a paradox, right? But even amidst this, you know, embarrassment of riches, as Matt's always fond of saying, why is it that we're still troubled?
So I wrote that, you know, in a lot of cases, I feel like from a practitioner standpoint, I really feel like we need to do better using the tools we have. This comes through like upskilling, learning best practices and so forth. I feel like that's largely ignored. You know, we focus too much on learning to work with, then quote the business and all this stuff. I think if we can start focusing on that, I think that this is actually one of the biggest issues in the industry right now is really the gap.
between the capability that we have with the tools and our ability to properly execute on using these tools. So.
**Matt Housley:** Yeah, the tools are almost too easy, right? And as part of the problem, they've almost become toys rather than professional tools that we use toward achieving a goal. It's like, uh, yeah. Like, uh, what are we trying to do with these tools? Sometimes we don't have a good answer to that question. It's just like snowflake is cool or EMR is cool or whatever happened thing. Yeah, well, Firebolt is awesome, but sometimes we don't answer the question of like, what we're actually trying to accomplish. And like you said, Joe fundamentals, like data fundamentals, like
**Joe Reis:** Professional Tools.
**Joe Reis:** Firebolt is awesome. Yeah, it's from the podcast. Yeah.
**Matt Housley:** quality, for example, and how do we ingest data properly? What do data contracts look like? Those are often missing.
**Joe Reis:** Mm-hmm.
**Robert Harmon:** That is a subject that I've been thinking about and I've been studying on, and I haven't quite wrapped my head around where exactly they came from, but data contracts. Again, I'm from a previous generation. Quality to me is guaranteed by correct schema. Those were the rules from 1990.
Obviously, that's a little more challenging today, especially when we have these massive cloud data warehouses and constraining a thing on one node when it's happening on another node in this giant warehouse is a mathematics trick that, well, I'll leave for Benjamin in another day. It's hard. So these things aren't available. So I can see how an idea like data contracts would come in. I just haven't quite wrapped my head around it. Any help?
**Joe Reis:** I mean, software engineers have been doing this forever though, right? With schema registries and stuff. And, you know, and I think that that's, I think if you want to know where data is going, like with contracts and whatever else, just look at where software engineering has been for the past 10, 20 years and just adopt those practices there. It's, you know, so I think it was Andrew Jones. I think he, he claims he was the first person to come up with the idea of data contracts or the term. Um, so we'll trace it back to him. Um, he's publishing a new book unpacked, which I still need to read. I think we.
**Robert Harmon:** Right.
**Joe Reis:** actually were editing it for reviewing it for a bit there about, but maybe he did more of it. A really cool guy and then obviously Chad Sanderson, you know, is just sort of taking the idea of data contracts. I think really popularized it. You know, he's built a big community around that whole idea and data quality. But it was interesting, you know, one of my software engineering friends, you know, she was at a conference listening to talk on data contracts and she was laughing the entire time.
She's like, this is sad. Like we've been doing, what's new about this? We've been doing this stuff for forever. Like this is nothing, this is nothing new. And so I think she was just kind of shocked at how far behind the data world really is compared to software engineering. So I thought that was really interesting.
**Robert Harmon:** Right.
**Robert Harmon:** Well, and maybe this is another artifact of a previous subject we were discussing. If I'm not reaching out to the customer, if I don't live in the customer's world, how am I going to understand what some dev team is doing with some application?
**Joe Reis:** Bingo. You won't, right?
**Robert Harmon:** So here we are.
**Joe Reis:** Here we are. I mean, the last blog post I had last week, it's a Dev and Data Divide. And I have this picture that I like to show. It's Dev on the left, Data on the right. And it's crap flows downhill, to put it more euphemistically. But that's kind of how it is. It's a one-way street typically right now. Data's on the receiving end of a lot of stuff. And hopefully that changes, especially as we start developing more data products and the feedback loop goes back to.
**Robert Harmon:** Yeah.
**Joe Reis:** the dev side now, to me it's an artificial divide. It's a divide that had to happen back in the day. Cause it was typically a data was an IT function, but that's disappearing as data becomes more front and center for everything. So I think that's, that's what's going to change. If you kind of rewind to your first question of where things going to go over the next few years. I think that's definitely one of them. I, it has to this, this artificial divide between dev and data. It's really crippling and it's a, I think it's a stupid. So
**Robert Harmon:** I agree, Joe. And honestly, for me, it was quite shocking because I spent way too many years at a single organization, well, later than the previous warehouse I spoke of, where I was the data guy for both sides. So, I control the entire BI environment and I control the entire operational environment. So I live with the development staff all day. This is what we do. So there couldn't be a divide.
So I had to play both sides of that game. And when I came out of that environment, out into the real world, it was honestly quite shocking. I hadn't seen this before. I didn't imagine it existed. And I got a lot of lessons in a hurry. But I do think, yes, we need to work more toward that homogenous type environment where data people are embedded everywhere.
They don't have to have a sign on their head that says, I'm the data person, but at least if we've got data people everywhere, then we've got a chance.
**Matt Housley:** Yeah, just moving toward the assumption that a lot of data is going to be customer facing rather than just appearing in reports and quite often, frankly, stale reports traditionally where you get a report after 24 hours or after 48 hours when maybe there are actions you could have taken and it's too late to take those actions now. I think the idea that your data can show up directly in an application, that the customer can get an idea of what's going on with their account or other places and that's all tied into analytics has really taken off in Silicon Valley in the last...
**Robert Harmon:** Mm.
**Matt Housley:** 20 years, but we're still kind of behind in certain areas.
**Matt Housley:** Oh Joe, you're muted.
**Joe Reis:** I was just gonna keep talking like that the whole time. You should let it go for a while, it'll become funny. No, I mean, the world's moved beyond reports at this point, BIA reports and stuff. It's just like, if we're still struggling with that, I don't know, I'll go do something else with my time, go become a veterinarian or something, it's more fun. But no, I mean, that's kind of where we are. I mean, we're still talking about dashboards and stuff. I'm like, seriously? Like, this is, so it's interesting.
**Matt Housley:** There we go.
**Robert Harmon:** Yeah.
**Joe Reis:** We can talk about solutions though, right? I mean, I'm good at being a curmudgeon at this point and cranky and, you know, uh, irritated and stuff. And I don't know it's, uh, but you know, solutions are interesting. And, um, you know, I think that's where the conversation needs to go. Cause it's again, it's just the same old tropes you see on LinkedIn all the time, especially where it's like, you know, deliver business value and all this stuff. And you need to have a data strategy in place and all the, all the stuff. And I'm like, dude, like we've been talking about this for, for ages. Like let's, uh,
**Robert Harmon:** Yeah, and the other half is, consider your audience. This kid just came out of college, it's his first year in a data role and you're gonna tell him he needs to deliver value. How's he gonna do that?
**Joe Reis:** Hmm. Good point. And not that this data value matter. I know, again, I know you're working on a book and stuff, but it's one of these things where I hope you can nail the topic too, in a way that, you know, pushes the industry forward, the stuff I've seen so far, it's like, it's good, but it's like, yeah, it's a tricky subject to tackle. Right. So.
**Matt Housley:** It is, and I've seen way too much vague consultant speak that I want to avoid. I mean, I think there are, I think if you've worked in data, you've seen very concrete ways of serving customer needs, for example, and that's what we need to talk more about. We need to talk about things that are frustrating in the customer experience and how people working in data can help with other IT teams to improve those experiences. That's what we're talking about when we're talking about business value, like things that make the customer happier, that make the business happier,
**Joe Reis:** Yeah, that's my problem with it.
**Matt Housley:** Yeah, that's value very concretely. And the problem is that on the one hand, it's very concrete for data people. On the other hand, it can be a little bit vague from the accounting side, right? Like what is the value of a customer who is happier because they can see what's going on with their account very quickly? It's a little hard to measure, but if we have a strong customer service focus, then there's definite value attached to that.
**Benjamin:** So one thing I'm curious about in general, right? Kind of like coming out of this is we're saying, hey, our tools got much better, right? But we still have many of the same problems we used to have. We still keep cycling around using kind of topics and both you Joe and Matt kind of, right? You're teaching a lot, kind of you're doing thought leadership, kind of you're writing blogs, you kind of wrote that super well-known book, you're affiliated with the University of Utah, you're consulting kind of like.
Arguably you could say, okay, if we have all of those amazing tools now and we're still cycling around the same kind of types of problems, right? Maybe we're just not teaching it well enough. So what does that mean for your approach to kind of delivering these things to a professional students, those types of things.
**Joe Reis:** We failed.
**Matt Housley:** Hahaha
**Robert Harmon:** The kid comes out swinging.
**Benjamin:** Hahaha
**Joe Reis:** It's a good question.
**Matt Housley:** I mean, I think part of the problem, and this is not to trash vendors too much. I think vendors build fantastic products. Yeah, yeah, yeah. But, but I mean, if, if I'm in sales for a vendor, I'm not necessarily focused on how, how I use the tool. I just want to get the tool out there and get people using it. Right. And that's where there is more need for people on kind of the meta level to come in and say, all right, you've decided on X, Y, and Z tools, how can we actually use these to help the company?
**Joe Reis:** Firebolt's awesome, yeah, for example. Um.
**Matt Housley:** And do that training all along. I mean, I think Joe and I have complained a lot about the lack of training for undergraduates and data specifically. And part of that training as we build it out needs to be, obviously they need to learn data fundamentals like data modeling, fin ops, cost management, but also what it's like to work inside of business and the kinds of things that businesses care about and how they can communicate better. I mean, communications are notoriously difficult to teach, right? Because how do you teach someone out of a textbook?
how to communicate with someone. But we need to keep thinking about these problems and figure out how to give students practical concrete experience with communicating with businesses and stakeholders.
**Joe Reis:** I completely agree. Yep.
**Benjamin:** So how do you do that? Because that was also a very abstract answer.
**Matt Housley:** Fair, fair. I mean, I think from our perspective, a lot of this comes down to building better collaboration between undergrad and master's programs and businesses. You know, it's shockingly often we see that you've sort of got this MBA world that operates almost in a vacuum separate from the business world. And that's not ideal, right? You want... Yeah.
**Joe Reis:** Well, the academic world operates separate from the business world too. I mean, in some cases that's good. In a lot of cases, I think it's, um, it's pretty bad. It does a disservice to, to students. So that's one thing I'd like to see change. Right. So you talk about concrete stuff. I would also like to see more apprenticeship type programs. I think that the notion of a university being a necessity, I think is absolutely the wrong way to go. Um, so I think more people could be trained on this, uh, from practical things like apprenticeships. Um, you know, I'm, uh, creating a new MOOC.
**Robert Harmon:** Absolutely.
**Joe Reis:** class right now, of course, for a really big MOOC. One of the things I'm doing is, it's a simulator. It's your first day on the job as a data engineer. You get to go do business requirement gathering. You get to go find out what stakeholders want, and part of it is identifying, okay, so you're given this list of requirements. What are people actually asking for? So I think that's the other topic. We spent too much time teaching tools and not enough time teaching the techniques, right? So I think those are concrete ways that we could address it.
Um, because, uh, it's easy to do like the, you know, PI spark tutorials and stuff. I think, but that's the wrong way to teach data. I think the way we teach it is absolutely, uh, it's backwards, right? Know the techniques and then learn the tools. That's why we wrote the book the way we did. It's, it's, um, technology agnostic, for example, right. And, um, pretty much every company in the universe is using it for their data teams right now, right. Almost every university that we know it's increasingly, uh, being used as a default textbook for data engineering. So
To me that's part of the process, right? But it's not gonna be an overnight thing. But I think the way we approached our book, it's similar to how Martin Kleppman approached his book. It's agnostic, it stands the test of time, and that's kind of where we need to get to. Yeah, so hopefully I answered your question. We are making an effort. It is slow, especially universities are slow. They're so slow. And that's part of the problem with them.
**Robert Harmon:** Yeah.
**Robert Harmon:** Mm-hmm. And, you know, I've experienced this myself and I've had the awesome opportunity to work with some great kids that came straight out of college. They were bright and worked with me for a couple of years and that builds that mentorship type relationship. And then of course I achieved my goal. They get all full of themselves and they quit on me and go somewhere else.
which is absolutely awesome. This is the best day in anyone's life as a data monkey when you send another one off to open his own shop. The problem is, he's not on my team.
**Joe Reis:** Benjamin, when are you quitting? Just kidding. Oh, OK. Sorry.
**Matt Housley:** Yeah, yeah. He's announcing it right now.
**Robert Harmon:** So no, those days do happen. The problem that we run into is how do you do that at scale? And that I haven't figured out yet.
**Joe Reis:** It's an interesting one. Mentorship is that's the other key component of it. I think is mentorship, right? And, um, you know, that that's, it's a huge, huge thing. So I think all the above really, how do you scale it though? I don't know. Right. It's hard because mentorship is inherently kind of a one-to-one type thing. So it's like, I don't know. And a lot of people don't want to be mentors, for example, like it's work.
**Benjamin:** Very encouraging.
**Robert Harmon:** Um, well, from, yeah, it's work. So there's that.
**Joe Reis:** Yeah, it's an interesting one though, but it is something that Matt and I think a lot about. I mean, you know, we're both educators and I think, um, you know, we did our small contribution to the universe with our book and we're writing new books, but you know, books, books will only take you so far, right? This is very much a practice oriented field. So.
**Robert Harmon:** right. You do have to be out in the pits doing the job to get it. And that's hard to explain. There's a million things you'll get hit with on any given day as a data guy that are not in the books anywhere. And the other thing that I'm really, you've heard me or you've seen me write it, Joe, data is not a technical game.
**Joe Reis:** It helps.
**Joe Reis:** Yep.
**Robert Harmon:** It's a social club. It really is. I need to know everything that's going on in my team, in other teams. I need to keep that socialization going or I can't achieve the technical solutions. And that I don't think is coming out of college.
**Joe Reis:** It truly is.
**Joe Reis:** Nope, not at all. So it'll be interesting, but you know, we were talking yesterday, uh, with Hall and Elson about this was a, the rare opportunity, uh, we're doing a podcast with her and we're opportunity of a three math nerds, um, three professors, uh, three O'Reilly authors in one podcast. So we were talking about the, you know, the, the idea of tenure really. And it's a double edged sword for that. Cause like tenured professors, you know, on one hand it's, it's great. It allows you to the, um, academic freedom and the
**Robert Harmon:** Wow.
**Joe Reis:** psychological safety to pursue your work. On the other hand, I think it also provides incentives not to do things in the student's interest. I've seen this happen where some professors and data programs especially won't update their stuff because it's too much work. Or maybe they have other things going on. So these students are learning outdated stuff like this Hadoop and all this other crap. I'm just like, why are you teaching this? This is nothing to do with reality at this point. So, but that's what it is.
**Matt Housley:** Yeah. And it's tough to stay current and you know, it's tough to find people from, from business who want to teach as well. I mean, they, if you have a job in data, you're very busy and you're probably well compensated for your time. And so teaching is almost a charity exercise. Maybe isn't so feeling. Yeah. You know about this.
**Joe Reis:** especially adjuncting dear god that's i mean it's like the worst job in the universe it's the best and worst job you've done it before
**Robert Harmon:** I have not. Yeah. Though I have been involved in a number of K-12 tech institutions and it's the same game.
**Joe Reis:** Yep. Yeah. I will say the best talk I gave this year though, is at my sixth grade, or my kids sixth grade class, we talked about AI, and that was really fun. You know, so I think like, you know, teaching younger kids is almost easier than teaching college in some ways, because they're just more fun to talk to, for one. But yeah, I think, you know, but it's, it's an interesting one. I think there's a lot of anti-patterns established at this point though, and how you could probably improve on with respect to teaching.
But to me, this is the crux in the industry right now. Like again, we have all the tools in the universe. That's not the consideration at this point. You can solve any problem you want to basically, but it's like, you don't even know what problem to solve because you don't know how to think through problem solving. That's a fundamentally different thing. So.
**Robert Harmon:** Well, and I do believe that part of that is just the youth of our industry. Now, sure, we've been collecting data on stone tablets since Mesopotamia, but not like this. This is a new
**Joe Reis:** That's when you got started in data warehousing, right? Just kidding. Um, okay. Okay.
**Robert Harmon:** I may have been there when stone tablets were invented. But not like this. This is a new industry. It's not like I'm dealing with architecture or finance or manufacturing where they've got ideas on how to do their job. We're still working out the kinks on this thing. So, you know, I think some of that is to be expected.
**Joe Reis:** Yeah.
**Joe Reis:** We talked to Bill Inmon too, you know, and I always ask him like, geez, Bill, what was it like back in the day? Um, he's like, oh, it's, we're, we're a very immature industry. I say, I remember, I remember calling him and kind of just. Anx one day. I was just like, why? I'm kind of tired of this industry, Bill, you know, like why, why is it that we keep repeating ourselves over and over? He's like, well, Joe, we're very immature as an industry. We haven't been around that long. I'm like, okay. Like.
**Robert Harmon:** Well, and I think intrinsically, we as data professionals have extremely short memories. Because if we could remember anything, we wouldn't run systems to memorize stuff for us. That's.
**Joe Reis:** But easy.
**Benjamin:** Thank you.
**Matt Housley:** Maybe it's the same thing about software engineers basically being working very hard to be very lazy or what they say about mathematicians, right? It's like you do all this work to write code so you don't have to do the work day to day.
**Joe Reis:** Yeah. But I don't know. It, you know, when you talk to people like Bill though, who's been around, I mean, he, in my opinion, he is the industry. He kind of helped, you know, he's the godfather of the data industry. And so, you know, um, but he'd been programming since 1960, I think. So that's, that's a long time. That's basically stone tablets at that point or a punch cards are basically the same thing. So, you know, but it's, it's. So he convinced me to stick around.
**Robert Harmon:** Yeah.
**Joe Reis:** So I was I was really gonna leave. I was like, I'm tired of this. This is just the dumbest industry I've ever seen. Who knows, I still might, but it's just, you know, I think it depends. And like, if, if we can like make tangible efforts to move the industry forward, you know, I think that's a good thing. But if we're still here talking about the same crap in like five years, like I'm gonna go find something else to do, I don't have time for this, like, you know what I'm saying? So
**Robert Harmon:** I actually did try and exit the data warehouse space after my second data warehouse, I got a new job. They had a whole lot of, they had a whole lot of operational database issues. I figured I could go work on that for a while. So when they, when they hired me, one of my contingencies was I'm not touching a data warehouse as long as I'm here. And they agreed. And then two years later, I rebuilt the data warehouse. So.
**Joe Reis:** It was that bad, huh?
**Matt Housley:** And were you coerced to or you just like, no, I've got to fix this mess.
**Robert Harmon:** No, I can't deal with this chaos anymore. I have to fix this. Yeah, it was pretty much it. Yeah, I tried to quit. I couldn't get out. So here I am. It's just can't finish. So yeah, we've talked about a lot of somewhat curmudgeonly topics and woe was me in the industry, but the upside here is we're all still here and we're all still fighting.
**Matt Housley:** So it's your own damn fault basically. Okay.
**Joe Reis:** That's pretty funny actually.
**Matt Housley:** So it's like smoking, essentially.
**Joe Reis:** Yeah, he's recovering now.
**Matt Housley:** quits thousands of times.
**Joe Reis:** Hahaha.
**Robert Harmon:** Um, I'm not sure why, but here we are. I do want to kind of move to some lighter subjects. If you don't mind. I know Joe and Matt, you guys are everywhere lately. What's coming up next.
**Joe Reis:** Yeah, that's fine.
**Joe Reis:** What about you Matt?
**Matt Housley:** Let's see, I have a couple of things upcoming over the next couple of months. So I'll be on a couple of panels at Big Data London in September. And then there's a conference in Budapest at the beginning of October called CrunchConf. And so I'll be speaking there as well. So I can, if you guys want to put that in the show notes, I can put a couple of links out there. And I'll be at your barbecue on Friday as well. That's right.
**Robert Harmon:** outstanding.
**Joe Reis:** It'll be at my barbecue on Friday. So yeah, that's like the highlight of the year. Um, yeah. And then it gets you right in a book too. Right. So that's, uh, ongoing, but, um, yeah, for me, what I'm starting a world tour this Saturday, actually, small to Australia, and then, um, I do the dbt and Joe Reese road show too. So they, uh, dbt and I have a.
Kind of a traveling circus goes to different cities actually would be in your neck of the woods in Seattle. And September. So you better Yeah, yeah, I'd be like that. Where's Robert? I can't have a party about Robert
**Robert Harmon:** Well, then I'll have to stop by. I have no choice, right?
I could be totally, yeah, I could be unsociable. I mean, that's in my nature, right? Everybody knows I'm shy.
**Joe Reis:** Yeah, you'll just send your kid. His kids keep stealing his fun rails of data engineering book. So it's almost a constant prank.
**Robert Harmon:** Oh, that was no, it wasn't just him. So there is a story here. I wasn't going to bring it up, but since you did Joe, before, before we ever talked to each other, I ordered a copy of the book and then, uh, my sister-in-law and nephew came to hang out for a week over Christmas. They leave. I go looking for the book. It's gone.
**Matt Housley:** Here we go.
**Joe Reis:** That's a funny story.
**Robert Harmon:** So, and then a month later, they come back to visit, book reappears for a day, and then it's gone. Cause the son now stole it. And I finally, I mentioned this to Joe, and he was so kind, he sent me an autograph copy and gave the kids instructions not to steal it. It says so right here. So I finally got through the book the hard way.
**Joe Reis:** No! That's awesome. That made my day though. And it made my day too that your kid I think was interested in data engineering too. Like that was one of the coolest things I'd heard.
**Robert Harmon:** Yeah, you kind of fired. I can't explain it. He's seen me go through an entire career of pain and misery. And then he's decided he wants to do the same. I don't suggest it. But what am I going to do? He's soon an adult. I can't stop him. So in a country where I'm not going to be able to do anything
**Matt Housley:** Yeah, tell us about that. Okay.
**Joe Reis:** Yeah. It's like that old 80s anti drug, you know, drug ad, the parents use drugs have children that use drugs. That was my childhood.
**Robert Harmon:** Oh, yeah. Yeah, I learned it from you, dad. I learned it from you. And you know, he's, there's some, there's some really attractive things about this industry. So I'm not going to tell him no. But I am going to be there, you know, the day when he falls into the industry trap, the consulting trap until I told you so. That's a guarantee. So he's off to school here in a couple weeks. He goes off to
**Matt Housley:** say no. Yeah.
**Robert Harmon:** be on his own in a computer science program, and we'll deal with the rest later.
**Joe Reis:** That's cool. That's pretty awesome. Yeah. So I think, uh, you know, the book is Impacted Kids. It's, it's awesome. Um, you know, uh, other stuff I'm working on, I got a, um, course for Riley coming out, it's a mini course, it's just basically the greatest hits of the book. It's like a, not that long of a course. It's, um, and then, uh, I can't announce it yet, but it'll be a big announcement in November for something I'm working on, so, uh, this is a ways away. But, uh,
can wait. So then I got a new book too, on data modeling that it should have been done by now, but I think a few things, one courses, two speaking, and then three just large language models and chat GPT kind of threw me for a loop with respect to data modeling. So I was like, okay, I don't know what this does. Does it do anything? Is it a nothing burger? But you know, it'd be, it'd be kind of weird to like write a book and not acknowledge that. So you have to really think through the consequences. So I spent the last
I think I'm finally coming to a resolution on it, but it was one of these events that I think it was too big to ignore. You know, so just, yeah.
**Robert Harmon:** Mm-hmm. So even I started playing with chat GPT and data modeling because I've written on this before. I'm a big nerd when it comes to business rules. Business rules drive data models. I write them almost compilable, kind of like Chris Date taught us back in the day. What I've found over the years though is that people write them poorly, so they're not exact. And what I found with...
LLMs is I can take my business rules, feed them to the LLM and say, generate me a schema and I can see immediately where I screwed up on my business rules. And then I can go back, tune up my business rules. And once I've got that, I can support a strong data dictionary, master data management, all that fun stuff. But the reality is chat GPT is an idiot. And if I can convince an idiot to get it right, then I got my business rules right.
**Robert Harmon:** So that's my little foray into that world.
**Joe Reis:** That's pretty cool. It's really cool. I've been spending the week nerding out on a vector databases too. I think that's a really fascinating technology. Because I think, again, with data modeling, I feel it's one of these topics where you mentioned Chris Date and relational model and stuff. And as an industry, we're still talking about things like relational models and Kimball, all this stuff. And I'm like, why are you?
People are just arguing about star schemas. And it's like, they came out in like the nineties, like let's move on. Like use it or not use it. I don't really care at this point. Use one big table. I don't really care, you know, but what's happened in the meantime, right? You got streaming, NoSQL, machine learning, all these new things. And you know, the modeling practices around that still need to congeal. And so that's, you know,
**Robert Harmon:** Yeah.
I, for the life of me, I'm not figured out how to pull off integration and time variance with a streaming source. This seems like a Doctor Who moment. You're going to have to fold time. Not that anyone cares just yet. But
**Joe Reis:** Well, walk us through, walk the audience through what you're thinking, because I know you and I have talked about this, I think, but you know, it's an interesting topic.
**Robert Harmon:** It's just a lot of workload to pull that off. Integration itself, any consulting firms run from integration projects for good reason. They're hard. They're just hard and they're work. You can't buy a product to make this happen. So if I've got like a salesperson table in one application, I've got an employee table in another application.
And really, this individual in both applications is the same individual in a subject-mat-oriented world. I need to put him together before I can make sense of all the data that applies to him, right? Now, try this in a streaming environment. I dare you. It's hard enough in a batch environment, because if there's an interference, at least in a batch environment, I've got an hour to fix this before the next batch comes through. In a streaming environment, I'm lost.
And then you start thinking about things like time variance, where we're looking at temporalities of that employee. I want to see what that employee state was over time and quickly grab the activity data over time for that employee. So he was in Montana, now he's in Wyoming. Did his sales numbers change? This is a tough question to start with because now we've got all this temporality to pull together.
**Joe Reis:** potentially.
**Robert Harmon:** Okay, so now I've got to integrate the employee, pull all the temporality off and do that while that Kafka queue is busy chewing my butt. I don't know how to do this yet. It's a problem I think the industry will have to solve. Probably not today.
**Joe Reis:** We're talking to the guys from SGR yesterday about streaming a kind of a deep dive.
**Matt Housley:** Yeah, you might check out, yeah, they had an interesting idea about this. So you might check out our Monday morning data chat from yesterday, actually. Yeah. So for the audience listening after the fact, that was a August 14th episode with Estuary and the idea there was to, you keep all the incoming data and sort of a raw JSON schema, and then when you have problems with your schema, for example, you actually pause the kind of more refined side of your data.
**Robert Harmon:** can do.
**Joe Reis:** Yeah.
**Matt Housley:** So you have an incoming pipeline, you're ingesting, collecting. You have a transformation pipeline. You basically pause the transformation pipeline and say, hey, what's going on here? You set off an alarm, you have engineers check the schema changes and say, okay, here's what happened. They can make a schema change and then replay from the point where the failure happened. Ideally it happens within like a day or something like this. And so that way as schema and data are changing, you do kind of have pauses on the output side.
The real time side, you want to focus more on the raw data basically, but you're able to go back in time using replay capabilities to fix issues as they arise fairly quickly. And the idea is to make this as agile as possible. So you don't accumulate several problems before you fix them. You're very operational with your fixes.
**Robert Harmon:** Right.
**Joe Reis:** So really too, I mean, if you take about a bounded and unbounded time as well, right? I mean, that should solve a temporality problem. If you're keeping an append-only log, for example, right? It's just, I mean, inherently it should be there unless you have late arriving data, which case, you know, you said need to know like when did the event happen and when did he get it, right? So by temporal, tritemporal would be like, you know, what did I do with it? But that's a different subject. But yeah, it's, in theory, it should be easier with streaming. But
You know, in practice, um, I don't know. We'll see.
**Robert Harmon:** That's not where I'm not seeing it as easier yet. I'm sure if we mature a little bit again as an industry, we'll figure it out. But right now that one's, yeah, it absolutely does.
**Joe Reis:** It'll have to happen though. It has to happen. I mean, Matt and I predicted this in our book, the, uh, we call it the live data stack and it's really a feedback loop between applications and events and, um, analytics and data and machine learning. That feedback loop shrinks. Right. And this actually yesterday, we're talking about how this might actually bridge the, uh, dev and data divide too, because if you can shorten the time, uh, between working with devs, right. Uh, the feedback loop is there in which case you do, you have to collaborate.
**Robert Harmon:** Yeah, yeah, right, right. And exactly. And one of the things that may help there is isolation of some of these ideas. Historically, we had operational data stores as part of our BI environment, where we can grab operational data from operational applications, bring it into that operational data store to co-locate it, and sometimes apply temporality there, but we wouldn't do integration in that place.
**Joe Reis:** It's a forcing function. So.
**Robert Harmon:** And that can really accelerate that communication between Dev and the BI space because this becomes our DMZ. What you guys do over there, you do over there. What we do over here, we do over here. We're going to discuss what's going on in this operational data store as a negotiation. And that helped a lot. The problem is making that jump from an operational data store to a real data warehouse. That's hard to do in streaming. I don't know quite how to do that yet. I'm working on it.
**Joe Reis:** Yeah, I think Druid does a pretty good job at this, but it kind of changes, I think, the nature of what, you know, we might consider data warehousing as well. It's still analytics, but maybe it doesn't, you know, in some ways it fits Bill's original definition, in other ways, you're probably going to have to morph it a bit too, just because of outer necessity.
**Robert Harmon:** Yeah.
Right, right. Or break it out into smaller structures that, like I said, make more sense or leverage the physical stuff underneath. But I think there's opportunity there as opposed to our previous conversations where we were all grumpy. I think there's definitely, yeah. Well, that was kind of random.
**Matt Housley:** We think we can actually make progress here.
Great. I like random conversations.
**Joe Reis:** But no, it's, it's fun. Um, yeah, that's, that's all of Matt's and Matt and his conversations are just, uh, super random, but, uh, happens. I don't know. When we get on podcasts, for example, right? We don't ever have scripts. And I know, and if people want scripts, I usually, I usually tell them I don't want a script, like do not have a script. I don't know, cause it's a, yeah.
But this is what you get. You get, you get the kind of our meandering, uh, you know, um, you know, crank cranky guy fast here. And, um, you know, just except for Benjamin, he's, uh, smiles too much. Um, just kidding.
**Benjamin:** Sorry, maybe in 10, 20 years, I'll smile less. Yeah. Perfect. So yeah, I think that's an awesome conclusion to the episode. Just be more bitter. Joe, Matt, I'll work on that. Thank you so much for joining today. Yeah, it was great having you on the show. See you around.
# Joseph Machado, Senior Data Engineer at LinkedIn talks best practices (/blog/joseph-machado-senior-data-engineer-at-linkedin-talks-best-practices)
Data engineering should be less about the stack and more about best practices. While tools may change, foundational principles will remain constant. Joseph Machado, Senior Data Engineer at LinkedIn, is on The Data Engineering Show to talk about principles that are key to success, leveraging AI for automation, and adopting software engineering methods.
Listen on [Spotify](https://open.spotify.com/episode/3sFQ0Q12dfjTLXx0AhBslM) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/joseph-machado-senior-data-engineer-linkedin-talks/id1561927688?i=1000647509481)
**Benjamin (00:07.422)**
Hi, everyone, and welcome back to the Data Engineering Show for another awesome episode. We have Joseph Machado joining in today, who did his master's at Columbia, then spent 10 years as a data engineer, data scientist in the industry. He's a senior data engineer at LinkedIn right now. And in parallel, he also has an awesome blog called Start Data Engineering, and is teaching data engineering to, I guess,
aspiring data engineer. So good to have you on the show, Joe. Awesome. Cool. So like, where should we start basically, right? There's, there's so much stuff to talk about here. Yeah. Maybe do you also want to kind of say a few sentences about yourself, your background? Exactly.
**Joseph (00:38.008)**
Thank you for having me.
**Joseph (00:54.376)**
Yeah, sure. As you said, I went to Columbia here in New York City with electrical engineering, although most of what I did was like network analysis. So like K-means clustering, that sort of thing. And then I started as a software engineer but quickly got interested in the database side of things. So like automizing indexes, making sure people are writing good queries, things of that nature. So I was in software engineering, moved into data engineering, but at that time my title was data scientist, but...
basically I was just doing data engineering, writing some SQL queries. And this was back when there's a Java MapReduce and HDFS. So I started there and then slowly along with the industry mode with like Spark, Snowflake, Artflow. Yeah, I've seen a lot of tools. Yeah.
**Benjamin (01:39.402)**
You've been through it all. You've seen it all. Hehehehe
**Eldad (01:41.25)**
You know, but it's amazing if you follow each step, it started with implementing an algorithm. So like an expertise with an algorithm, right? Like K means, you said K means, that was the domain. And then it expanded into a micro process and then it became bigger and bigger and then moved into a data warehouse and then you ended up with Snowflake.
**Joseph (01:54.043)**
Yeah.
**Joseph (02:08.533)**
Yep.
**Eldad (02:10.702)**
But I think it kind of tells a story where, and then of course I'm glad that the database always prevails, but that's a side story. The real story is that databases have grown way beyond anything that anyone predicted. And I think today kind of if we'll get your share of the pie.
and your experience around that and how you started there and ended utilizing data stacks. I think that's kind of what I'm personally looking for in today.
**Joseph (02:50.152)**
Yeah, I think that's a good point. I like it started off, I started off with IBM DB2. I don't know if you guys worked with it. It was like back in the day, IBM data warehouse stuff. Yeah, honestly, it wasn't too bad. I mean, the main thing was the data was modeled properly. So it was very easy to use. I worked with data warehouses where data warehouse is great. The technology is great, but data isn't modeled. So it's hard to work with.
**Eldad (03:01.598)**
or Swarp, that was my first IBM product.
**Joseph (03:18.144)**
But when I started off, luckily, the data was modeled very well that made it super easy to work with. But I started off with like just writing Python scripts directly accessing DB2 and Hive. That's pretty much it. And orchestration scheduling was just Python and was that Windows task scheduler. And it continues to work. I think it's been running for like eight years now. Yeah, it works. If you like... Mm-hmm.
**Eldad (03:41.062)**
It worked, right?
Just a second, Benjamin, for Benjamin and the rest of the young audience, there were a lot of keywords that most of you don't understand or know. This were kind of at the beginning, right? I love it. I've been there. But for Benjamin and the rest, this is how it all started. Sorry, go ahead.
**Benjamin (04:01.646)**
Thanks for watching!
**Benjamin (04:06.264)**
Thanks for watching!
**Joseph (04:09.94)**
Oh no, that's a good context.
**Benjamin (04:10.007)**
I love this has become a recurring segment on the podcast. Aladad.
**Eldad (04:11.882)**
I'm trying, I'm trying to tell Benjamin that databases and warehouses were not born in the cloud.
**Benjamin (04:17.802)**
Eldar explaining things to me before that happened before 2010, because I wasn't alive back then. That's awesome. Well, so having gone through that journey, I just like kind of from Hadoop, MapReduce, those types of things to now modern cloud data warehouses, like what changed, right? That kind of like, what's the same? Like when you look at the space today, what are the main challenges you're seeing also in your job at LinkedIn?
**Joseph (04:25.186)**
Sorry.
**Joseph (04:33.234)**
Mm-hmm.
**Joseph (04:45.82)**
Yeah, I think from like a conceptual standpoint, the technology has gotten much better. However, the fundamentals still remain the same. Good software driven practices, proper testing, that's still the same and a lot of places, software engineering based concepts like testing, making sure you have proper CI, CD set up, it's not really followed in the data team. So I feel like data engineering teams are kind of lagging, although that's changing these days,
That's one thing I've seen. The technology hasn't gotten so much better. It also makes it easy to build things quickly without following good practices, which leads to like long-term pain and having to migrate or do things like that. So technology is growing super fast, a lot of features, but fundamentals haven't changed much in my opinion.
**Benjamin (05:37.474)**
So this is actually something I'm curious about is, right? Like when you're talking about testing, so when you started out, like this was, there wasn't like big data engineering teams like at those companies back then, right? Like these were purely software engineering teams probably kind of working with these big data technologies in many cases. So would you say already then kind of people didn't do enough testing on these big data things or would you actually say, well, so 10 years ago we were in a better state?
**Joseph (05:48.152)**
Mm-hmm.
**Benjamin (06:06.39)**
kind of in terms of testing our data pipelines, kind of our infrastructure, maybe because...
**Eldad (06:11.766)**
everything was consistent it worked db2 transactions committed you know yeah
**Joseph (06:16.344)**
It worked, yeah. I wouldn't say better or worse. Like it just depends on the team, but it's a pattern I have seen. Like if you focus on fundamentals, if your team has solid fundamentals and good data platform, it helps a lot. And with the team size, right? Like the data engineering team that you mentioned, yes, there are a lot more data engineers now, but I also feel like there's so much more complexity in most cases unnecessary.
that adds to like the toil, the developer experience toil, if you will, of getting something into production. I used to be able to, yeah.
**Benjamin (06:53.134)**
It's funny you're saying that because the tools should have gotten simpler, right? Like it's been kind of 10 years and like, okay, they are simpler in the sense that you can write something in like 10 lines of Snowflake sequel that would have been hundreds of lines kind of like complex map-produced tasks back in the day. But this explosion of complexity to get things into production, it's like, it's a bit crazy to me. Where do you think that's coming from?
**Joseph (07:13.916)**
It is. I have a hypothesis. I'm not sure how accurate that is, but my hypothesis is that when the SaaS companies start building tools, they make it really hard to test locally. So if you have Postgres or something like that, you could easily test it locally. On Snowflake, it's hard to test. Databricks, it's not simple to test. You can do it. And one of the reasons why DBT is so popular, it makes testing super easy. So with...
with a lot of new features that we were also giving up a lot of these core software principles, like having a virtual environment locally or Docker or whatever you wanna run it and being able to quickly run tests with data teams, that's a hard ask, but it is what it is. And then focusing fully on SQL, while SQL is great, sometimes it's hard to test specifically. So I wouldn't say it has gotten
worse or better. I think it has been messy and it will always be messy and clearing up the mess is up to the individual team. But the technology is far superior now. I don't know how to like, as you said, multiple lines of Java code, compile it, push it and I could just write a SQL query. So yeah, that's what I've observed.
**Joseph (08:39.872)**
Yeah.
**Benjamin (08:42.466)**
Awesome. So what can you maybe take us through some of the challenges you're seeing today, like in your job at LinkedIn or other industry exposure you had around these types of things you're thinking about nowadays?
**Joseph (09:00.308)**
Yeah, I won't say specifically about LinkedIn, but I could say it like as a general, kind of generalized idea, what I've seen is the developer experience is really lacking, especially in the data space. Like I worked on software engineering teams where we can deploy like in an hour, if you put up a hour, someone reviews it, that's it. But that's not always the case with data teams. Sometimes you spend like a month, you spend a certain amount, a week.
**Benjamin (09:09.546)**
Sounds great.
**Joseph (09:28.46)**
a few days to actually validate your data. So I think in that aspect, the data teams could do better.
**Eldad (09:34.042)**
Don't be shy, a month is good, a month is great. Really, 30 days or 31 days.
**Joseph (09:42.69)**
31 days, yeah. But yeah, the developer experience, I wish it were better. And it also partly comes from the whole, I feel like the domain itself, right? Like from like backend engineering or application development perspective, you have clear definitions, you have clear scope, you have clear, let's say UI or clear behavior. But from a data perspective, it's hard to quantify what right data is. So that's the kind of...
difficulty that I'm seeing. Because if you quantify what right data is, what right data means to you, it should be pretty straightforward. But when the data grows in complexity, and there are so many product teams that you have to coordinate with, defining what right data, it's in of itself a huge task. And it's never ending. That's like new edge cases, and then you modify a code, yeah.
**Benjamin (10:30.189)**
Yeah. So what's your take on like data observability tools like Monte Carlo or something like that? Like where do you see their place?
**Joseph (10:39.444)**
I think they do definitely help. LinkedIn has its own system. We were using something like that at my previous place with DBT. It definitely has its place, but at the end of the day, it's just a tool. It cannot define what good data is. It can give you guidelines or freshness, check for these qualities, sure, but the business rules.
For example, like variation of a threshold over time, how should it vary? What is the seasonal? It's hard to automate that with a tool. You need to kind of dig into the data to manually figure that out. But those tools do make it easy to kind of set it up, if you will, super simple. Yeah.
**Benjamin (11:12.408)**
Right.
**Benjamin (11:25.826)**
makes perfect sense. One thing you just mentioned is like internal tools, right? And I feel like, okay, if you're working in the big company, like also like as a software engineer, right? Like you're going to Google, you're going to use a lot of their internal tech around how you deploy things kind of on Google data centers around like their internal version of your PC, all of those things. If you go to Facebook, same thing, internal tools, if you go to Microsoft, same thing, right? It's like, as a data engineer, I assume that's not so different. Like if you go to kind of big technology companies, which have
like exabytes of data they're managing, there will be in-house tools for specific problems. How do you think about that in terms of staying then relevant and kind of up to date with technology, right? Because it seems like that actually makes it harder. And especially data engineering, I feel like it's even more about the tools you know how to use compared to software engineering, which is already about that. So what's your take on that? Data engineering kind of had big tech then.
**Joseph (12:24.52)**
Um, I, I have like a opposing opinion to that. I don't think tools matter. I think the, the principles matter, like test data before you publish it to your stakeholders, um, how do you quantify test? Those sort of things matter. The design principle, if you will. I don't think tools matter as much. For example, you can have spark, you could have snowflake at the end of the day. They're both distributed systems. You, if you know how to.
look at distributed query plans, optimize it, you're good, that's my opinion. But I do know when you apply for jobs, you need to take certain boxes. So the way I think about it is if I have experience in like, let's say Snowflake, I would just try out Spark on my own, see how it works. So yeah, sometimes companies just want, yeah.
**Eldad (13:13.082)**
This is super interesting, you know, because in many ways, what you're saying is many steps within the data pipeline are commoditized. You can pick, you know, like each step you have a choice of 10 tools. And each one of them is unique on its own, et cetera, advantages, disadvantages.
by the end of the day you look at the whole data pipe, the whole funnel, right, that's your product kind of the input and eventual output, right, you're talking unstructured data coming in defining raw metadata on top of it, like this is very delicate stuff completely owned, I think dominated by the human factor, right
Of course, once it gets semi-structured and of course, structured, it's easy, right? Like the universe becomes much easier. But I think in many ways, it's really some of those big steps are mostly about efficiency, getting the job done. How fast, right? How robust can you plug and play each part and have the human part own it?
interesting to get your opinion especially where you're at right where in-house dev happens how you apply AI on those delicate parts of the process right is that applied is there ELT.AI being yeah one
**Benjamin (14:59.646)**
Let's start with one question, Elda. It's been a million questions, I can't keep track. Ha ha ha.
**Eldad (15:03.858)**
And so one of those each one of those is great. I mean, but yeah.
**Joseph (15:09.504)**
I think as for AI, there is an internal tool to convert text to SQL queries. It's like the low hanging fruit type product. I think everyone is doing it. Convert text to SQL queries based on the metadata information we have. But for designs, we do have like a template one can use because at the end of the day, most pipelines are kind of similar. So there are templates we can use to...
kind of quickly spin something up. However, within those templates, we still have to write a Spark job. We still have to see how joints are done. We still have to figure out what is the best way to design that code, how to structure it, how to organize it. That has not been automated yet. So I guess that's why we have a job. But yeah, I don't know if it will ever be automated because there are so many constraints, especially in a big company with so many teams, so many formats, so many...
data nuances, which kind of brings me to another point, which is like the separation of product teams and data teams. I think it's pretty not great. It's bad, I think it's because product team operates on its own, data team operates on its own. The way there is like a huge disconnect. So if there were more kind of connection like embedded data engineers within the product team that might...
kind of enable a more AI driven development, but so far I haven't seen any. No, no, adding, trying more people, I don't think is always the solution is it's just, yeah. This adds a lot of confusion.
**Eldad (16:41.794)**
More people, more people.
**Eldad (16:50.318)**
You know, it's no matter how you place, you know, like you can centralize, you can decentralize, you can do all sorts of options. And I think all of the work for certain use cases at certain times, for certain sizes, if you have a good team, they can utilize any stack and deliver value.
**Eldad (17:22.039)**
And yes, so as you're saying, data is becoming very boring and products don't matter anymore. It's just squeeze out more, you know, more efficiency, more value, no appreciation for the little things.
**Benjamin (17:28.418)**
That's it.
**Joseph (17:28.724)**
I did not say that.
**Benjamin (17:40.302)**
That's it.
**Eldad (17:40.533)**
Great outlook.
**Joseph (17:42.748)**
I do think that eventually, backend engineers will become data engineers or resource as well. That would be my ideal scenario where you have one team that builds the backend system and front-end, if you will, and kind of owns the data as well. That way you don't have these two separate teams with two different roadmaps with two, everything is separate. If it were a single team, I understand it's...
**Benjamin (18:07.542)**
Only engineers. Amazing.
**Joseph (18:09.817)**
Not only engineers, but I understand like bigger teams, you know, with bigger companies.
**Eldad (18:12.954)**
Give everyone an engineer title, everyone, and everything will sort out, no matter how you structure the things. Put them far away, right? Remote, Benjamin? Ha ha ha.
**Joseph (18:18.744)**
There you go.
**Joseph (18:22.303)**
Yeah.
Get Stuff.
**Benjamin (18:28.315)**
Good stuff. So Joseph, one thing you're also actively doing is right, kind of like teaching and thinking about how to like teach data engineering, like you have a really successful kind of blog and newsletter called Start Data Engineering. Take us a bit through that journey, kind of like what, how did you start that? What are you talking about there kind of now? What's top of your mind?
**Joseph (18:50.848)**
Mm-hmm. Yeah, so I started during COVID. I had like the extra commute time of us, like, okay, what do I do with this commute time? Just...
**Benjamin (18:56.758)**
This is like such a consistent thing. Like I love it, right? It's like kind of for everyone we have on the show, we're like in kind of thought leadership and kind of education is always, how did you start? Oh, like during COVID, like we had so much spare time. Maybe COVID should.
**Eldad (18:58.234)**
Ha ha
**Joseph (19:07.989)**
Yeah.
**Eldad (19:09.706)**
Everyone remembers COVID as a very positive memory experience.
**Benjamin (19:14.41)**
Exactly. It got me into data engineering thought leadership. Nice. So sorry for interrupting.
**Joseph (19:14.673)**
No.
I guess so. No, no, please. That's funny. No, yeah, I just started during COVID. I try to write about what people are looking for, not what I like to write about. I would like to write about more low-level operating system. Oh, is there a way we could use systemd for argument? More low-level type stuff, but that's not what people are looking for. People want some...
projects or DBT explanation, things of that nature. So I try to write about what people are looking for.
**Benjamin (19:50.862)**
this. Like this is different because so far always the answer has been, oh, just write about whatever you're passionate with, start with something that even no one cares about and you'll build an audience over time. This is the opposite. Pick what's popular, pick what gets the most likes and then kind of commit to that. I love it.
**Joseph (19:57.904)**
Yeah, that doesn't.
**Joseph (20:08.792)**
Yeah, I try to make it actionable. So always code, not just text. Because that's what I prefer. As you said, I don't like just text. There is some sort of actionable. So yeah, that's pretty much it.
**Benjamin (20:26.622)**
this.
**Eldad (20:26.874)**
So if you had to build the perfect dbt benchmark.
Right? Like run the craziest stuff, the hardest stuff. What would be the kind of the hardest stuff to run on dbt? Have a global open benchmark, right? Just switch, plug and play and run. How would you do that? What the queries would be there? Is there something going on there? Like, right, there's so many things you can solve with dbt. It is such a powerful obstruction. Tell us more like what's out there?
**Joseph (21:04.984)**
I, well, what's out there, basically, I think everyone is just mostly doing the same, 95% of the companies are doing the same, using the dbt kind of project structure to build their own. I do think there is a need where there will be like a business vertical type product. So let's say advertising, right? I can see like someone building an advertising stack. So right from segment to data warehouse.
that you build it with DBT, you can apply it to different advertising companies. Same with like finance stuff, from like pulling data from Experian or whatever it might be, getting some dashboards out to analysts. Because I feel like I've worked in multiple verticals, marketing and advertisements, little bit of finance. So they all have the same sort of input. Like if you look at marketing, click stream, click orders, blah, blah.
**Eldad (21:32.998)**
Thanks.
**Joseph (22:00.068)**
Sales, it's the same thing, opportunities, things of that nature. So if you model it right, and if you make that pipeline specific to a vertical, you can use dbt and just deploy to different people or different companies in the same vertical. It's at the end of the day, it's the same data they're collecting. It's just that everyone does their own implementation. So I think that, that we might see more of.
**Eldad (22:25.39)**
You know, they said the same about SQL. They said, oh, you just write it once and you can run it everywhere on any database. And look, look where we are today. What, what a mess. But SQL is still very consistent. If you play by the book.
**Joseph (22:32.376)**
Haha, wow. So, yeah. I guess that is the... yeah. Mm-hmm.
**Eldad (22:46.182)**
plays well, really well. And DBT is same, it's very similar, like hearing you out, you're really treating DBT as a standard. Super interesting to see where this grows as an ecosystem for verticals.
**Joseph (22:48.772)**
Yeah, there's the, uh, mm-hmm.
**Joseph (23:03.784)**
Yeah, I do see that dbt, you know, there's like a lot of community support as well. So if, if they officially don't support a database, there's like community drivers to enable dbt to run on like different databases. So I do think a lot of companies are moving there, especially, um, new, new startups, LinkedIn is also starting with dbt. Um, so yeah.
quite popular. Hopefully they get to profitability soon and don't change the license but we'll see.
**Benjamin (23:40.062)**
One other thing you talked about recently in your blog post was open table formats, right? And I think this kind of also ties into the kind of modern data landscape. So things like Apache Iceberg. What's your take on those? Like where do they fit? Kind of do you think they're basically eating the world? What are your thoughts here?
**Joseph (23:43.948)**
Mm-hmm.
**Joseph (23:51.98)**
Mm-hmm.
**Joseph (24:01.552)**
It depends on the company size. I do not think they're going to eat the world anytime soon. Just because Snowflake has its own internal format, if you will. I forget the name. It has its own thing, which is very similar kind of to Apache Iceberg. Spark has its Delta Lake format. Iceberg, I think it will be helpful for companies, bigger companies, specifically working cross clouds and cross systems. So...
In LinkedIn, we use Apache Iceberg. I mean, we have a wrapper on top of it, but it's Apache Iceberg. We use it to shift, move data between our on-prem and cloud resources, and we can use Spark and our Trino, whatever it may be. So at a bigger company, it makes a lot of sense because there are so many different stacks, but I do not see smaller companies using Iceberg just because the impact would not be as high. You could just get the same with Snowflake.
And smaller companies are not usually not going to have like two or more data processing systems usually. So that's my opinion on that, but I do, I do think it's growing fast. People, there's a lot of interest, but mostly from bigger size companies.
**Benjamin (25:16.546)**
Yeah, makes sense. I mean, yeah, it's like, to me, I mean, one of the big questions there is, and we'll also have to see what the verdict is around performance, right? It's like kind of one thing that vendors who have kind of first-class managed storage, like Snowflake, for example, would claim, is that they can build a file format and kind of like manage storage is going to be faster than Apache Iceberg over Parquet. Of course, then kind of vendors who are into the open formats, like kind of Databricks.
would disagree on that. And also Snowflake is now moving heavily in the iceberg direction. So I think this will be like very interesting to see how it plays out over the next couple of years.
**Joseph (25:54.12)**
Yeah, they just opened public preview, I think two months ago or last month, something like that for a iceberg interaction. Um, yeah, with Spark, it's super easy. I do think that like, because it's open source, that'll be a lot of adaptions specifically in like features parity, specifically from Snowflake and Databricks side.
**Benjamin (26:16.79)**
The spec is actually huge, like kind of iceberg as a specification, like with all the features it has, it's massive. It's actually very hard to implement into a data warehouse.
But okay, as a user, you of course kind of don't care about whether it's hard or not. It's just nice if it works.
**Benjamin (27:06.926)**
Awesome. Cool, Joseph. So anything else that kind of is on your mind in terms of data at the moment that you wanted to chat about today, kind of wanted to bring up.
**Benjamin (27:23.382)**
You don't have to say anything.
**Eldad (27:27.163)**
Ha ha.
**Benjamin (27:43.126)**
I see. I think those are great kind of closing words. It also really resonated with me, like kind of like this idea around learning like concepts rather than focusing on tools. I totally buy into that. I mean, I come from a software engineering background and like there, this is 100% the case. So of course, like probably the same with data engineering as well. So thank you so much for being on the show. It was really great having you Joseph. Yeah. And see you around.
**Eldad (28:10.822)**
Thank you, Joseph.
# Large scale data engineering at Momentive.ai - Meenal Iyer (/blog/large-scale-data-engineering-at-momentive-ai-meenal-iyer)
As companies scale, data can get messy. The data team says one thing, the business team says something else. Meenal Iyer, VP Data at Momentive.ai, met the Data Bros to talk about enforcing collaboration in large organizations to ensure what she considers the three most important factors in data: Adoption, Trust, and Value.
Listen on [Spotify](https://open.spotify.com/episode/15CdXgmYIWYsV21xkcJ1hw) or [Apple Podcast](https://podcasts.apple.com/us/podcast/large-scale-data-engineering-at-momentive-ai-meenal-iyer/id1561927688?i=1000620875451)
Benjamin: Hi everyone. Welcome back to the Data Engineering Show. Welcome Meenal our guest today. Welcome Eldad, kind of back from vacation after missing out on the last two episodes.
Eldad: Glad to be back.
Benjamin: Good to have you. So we have a great episode planned today. We have Meenal Iyer joining us from Momentive.ai. So, anyone who hasn't heard about that before, maybe you've used SurveyMonkey, basically Momentive is the parent company. And, I'm sure Meenal will tell us just in a minute kind of what other products they're kind of working on? What is the company doing and so on? Meenal is the VP of data there. So, we have a great conversation plan, kind of talking about the data challenges there, kind of data leadership, those types of things.
Meenal, do you just kind of want to quickly introduce yourself, tell our listeners about Momentive and then we can jump right in.
Meenal: Awesome. Thank you again for the opportunity. Hi, Benjamin and Eldad. Hi everyone. I'm Meenal, I head the data team here at Momentive and now going forward, it's going to be called SurveyMonkey again. We are still in the process of renaming our company. So just...
Benjamin: So it's like a flip flop, basically going from SurveyMonkey to Momentive and then back to SurveyMonkey.
Eldad: Because everyone knows SurveyMonkey.
Meenal: Exactly, yeah. Everyone knows SurveyMonkey.
Eldad: Yeah, it works. Sorry, go on.
Meenal: Yeah, I've had an exciting 11 months over here, looking to build a data platform that can allow the organization to kind of make, data-driven decisions and then produce value for the organization itself. We just were simply acquired by a private equity firm and, as you know, it just becomes more imminent that the data team kind of starts producing more value than it typically does in an organization and such scenarios. It's going to be an interesting journey, starting now or going forward. So, I'm super excited, and yeah, very excited to be on the show.
Benjamin: Awesome. Super cool. Do you want to kind of give us a quick recap basically of what got you into kind of big data, data engineering, those types of things? Because you have a bunch of experience across many different companies, so I'm sure our listeners would love to learn a bit more about that.
Meenal: Oh, absolutely. So, no big story there. I have a long story, but I got into data by chance. Realized I really enjoyed working in it and figured that I want to make my career here. So, I kind of dabbled in different industries. The reason being that I wanted to learn new businesses, wherever I went, and then looked to see if I can solve different kinds of data problems everywhere. So, it's now gotten to a point where I have an understanding of the industry, so all I have to learn is the business and then I have a playbook and I essentially use and employ that playbook and make organizations successful in their data journey itself.
Benjamin: Gotcha. Sounds cool. So this kind of leads us right into, I think, the first interesting thing to talk about. So, tell us a bit more about that playbook, basically, and also as times change, like especially in data engineering, things are moving so rapidly, how much do you have to adjust the playbook, basically, or is it actually quite constant?
Meenal: Well, I think I would say my playbook is now about six years old till before then it was kind of evolving and just for the reason that space in itself was evolving till that point we had a concept of where we were very heavily into data warehousing, having warehouses. But now, the nature of data has changed, the usage of data has changed, how organizations perceive data has changed? The value that these teams provide has become very, very different. The value generation is now coming out of data teams, the ideation, the monetization. And so for that reason, the playbook had to evolve in a way that we start looking at data democratization in a much broader sense. Data privacy and governance became large components within the whole model itself.
So, I would say there was a shift or a change over the years as to from where we started. From there it was very simple and not simple in terms of the build out itself, but simple in terms of what the requirements of the data team were. It was that you have or produce the data, you have a model essentially that services it to different teams within the organization. And then organizations had their own decentralized analysts. And people who were really good with data itself on their teams, who would like to pull the data and do work with it. But then over time it changed to where the data team itself needed to be that center of excellence and that value generator rather than just being the holder of data or just having governance over the data itself.
In order for that to happen, the education or the training that the team had to undergo or the way the team has to partner essentially with business rather than just being someone who just takes orders from the business itself.
We become partners because since we hold the data, we have a full understanding of all the data that exists, how the data all ties in together, and the value that can come out of that data that business may or may not be able to see. And so, how can we make that our motto, is how my playbook has evolved into. So, of course we still believe fully in self-serve analytics. So where, we push the data out and we make data available, so that business can make the decisions. But we also hold that center of excellence hub where we can start ideating in terms of what we can produce out of this data? What value can we bring out of it? Because that is the real ROI of what the team can actually provide.
So, yes, the playbook kind of has all of these things, but I would say, data democratization is like the word that would kind of encompass the end-to-end of what I have within my playbook itself.
Benjamin: Gotcha.
Eldad: Okay, so Benjamin, let me explain to you in a nutshell, in 40 seconds kind of the evolution of data. So, at first we served engineering teams and they were using data to build products. It was amazing. But then we got kind of used to it. So, the business actually got on and they started to use data to run the business instead of just building products. So, we switched to serve the business instead, which was also amazing. But then again, we reached the point where we just democratize data and we open it for everyone. So, they can serve it themselves and we just serve the data. So, I think it's kind of getting back to the roots, we only serve the data and if we model it right, we open it right. You've mentioned compliance and we've mentioned all of those things, gatekeeping, an excellence center. I think we're entering an era where it's all about controlling metadata and opening the data, so people can use it themselves. Tell us how it evolved? Actually, tell us when was the first time in your career that you considered data to be a strategic part of your team's ability to serve?
Meenal: I would say, we were naturally doing it throughout my career. Like, but it got to, I worked in companies where I truly was able to see how we could push the value for the organization itself. And let me give you an example. So, when we started off we again started off as a regular organization. We had our warehouse. We were just kind of pushing data out and then as we began conversations with the business, and I got to learn a lot more about what the business teams were doing and the challenges that they were facing, and I realized it is so much more simple if I can help them build this out themselves rather than them doing it. So, it started slowly with automation. There was an automation of an exercise that used to take one person on their team, 76% of their week to do, and that is all they did. And even then, the value that it provided was not complete. And I said I can very easily build this out and automate it for you, create a whole simulation for you and you just have to press a button and like input values and you get an output. And not only are you able to project this for three years, but if you want to project it across five years and see what that looks like, you should be able to do that. And that was kind of the first foray and then I was like, Wow. And, I know it's like a moment, but, it was...
Eldad: Wow, they're all lazy here. I just managed to kind of replace them over a weekend. I need to, I'm leaving.
Meenal: No, but sometimes it's like, there are folks on different teams who are just there for that one specific task, and that very repeatable task, which provides them some level of security, I guess. And I said, we can very easily convert. So we just did. We did go and deliver it to them, and they were so happy. And then they kind of championed us, going forward and soon, we had a couple of teams coming in and saying, oh, can you do this for us? This is what we are struggling with. So, we built fraud models, we built like market basket analysis models, and that's how it kind of started. And I was like, huh. So, that kind of essentially became my thing to do. So, as I started moving into other organizations, conversations with the business became a regular thing, communication with stakeholders, understanding what their challenges are, became a very, very regular thing. And I realized that they all are very willing and ready to share what their concerns or their challenges are. And then you just have to find a way in which you can help and/or assist them. And you would be very surprised but there is always a way for data to assist and people don't just say data is an asset, just like that.
It is truly an asset and if used well within organizations, we have an ability to assist with every business function that exists over there. So, that kind of became my motto. I started identifying who my champion team would be because those champion teams became the sellers of my team, and they became the sellers of our abilities and capabilities as well. And so, going through this champion methodology essentially assisted there as well. So, it's something that I have put in my playbook and I have employed.
Eldad: Awesome.
Benjamin: This playbook, as you move between different organizations to make them more data driven, to kind of champion data teams and so on, like how much does it generalize? So like, let's be maybe very specific, like now that you're at SurveyMonkey how much tuning does it need? How much time do you have to spend just learning a lot about your organization before you can kind of start implementing that?
Meenal: So, the business changes. I have moved across multiple industries. So the business changes. Every business operates a little differently than the other ones. So, of course, there are those nuances. The metrics are the KPIs of the organization measures are different. The business functions are different. Sales looks a little different here versus how it looked as retail sales. So, there are a lot of those nuances that you have to take into account. But then if you look at the overall strategy, you still have a data team. We still have a data science, BI analytics team. We have business functions. We have teams that need to be upscaled. We still have a maturity model, that you understand that at what level of maturity you are and where you have to progress.
So, there are certain things that are common and you just have to realize, okay, there in the journey is this organization. And then, you kind of take that journey across. So, for example, as I get into organizations, I look to see how self-service an organization is
In some cases, you know, you enter an organization, it's pretty mature. Self-serve analytics has already been built out. So then, you think, okay, what next? Like, how do I take them to that next level? In some cases, no, they haven't even got self-serve. The data team is functioning almost as an IT team. And how do you shift that from an IT mindset to basically a data team mindset. So, there is that evolution that happens, but if you look from a playbook standpoint, you just define or understand, okay, as to which and where in that journey they are. And then you basically start from that journey. But the playbook already has it in such a way that I know where, at which point I need to start and then move ahead on this. And yes, there are some small tweaks here and there that we have to do for the organization itself, just based on certain changes that may happen or in the way that you may have to operate. There are some changes, but for the most part that that playbook has been good.
Benjamin: Gotcha. So, how do you kind of along this journey measure the success because you're coming in as a data leader and if you have this specific vision of where the data team moves or how a kind of high functioning data team operates? How do you get buy-in into that and then kind of show, hey, look we're operating better now than a year ago or two years ago?
Meenal: So I think that is, how do I put it? Okay. So say if you're at the beginning of your journey, in the beginning of the journey is where you have to make a case as to that I have to, this is where I have to kind of drive the organization towards. This is where you are. So the first state is where you kind of go and speak to individuals within the organization and get a feel from them. Because you hear a very different story from the data team typically. I have been a developer and I know how I was. I used to always say my code is the best. I can never make mistakes and I produce the best thing. So, it's exactly that way.
Eldad: No, it's not your fault. You're using different excel versions of the same schema definition.
Meenal: Exactly.
Eldad: So, they end up building multiple versions of the truth on the centralized data warehouse.
Eldad: Tell us a bit about, do you use products to enforce collaboration on shared metadata, consistent metadata? How do you do it?
Meenal: So again, the centralized team, so let's talk about these large organizations. So, you have a centralized team. The function of that centralized team is one to produce that golden state of the data and to produce a semantic layer. So what is a semantic layer, is basically a business view of the data, where the business is able to kind of come in and essentially question the data, do anything with the data. So that may be, whether they want to use it to build their own data science, whether to do their own analytics, there are a whole bunch of functions that they can do. In some cases, they have their own sandboxes where they kind of play with the data and see, okay, what could be, and then they push it to the central team to kind of productionalize. So, there are a lot of these functions. So you produce, so you have the semantic layer ready. Your semantic layer has all of your key enterprise metrics and KPIs already predefined. So, irrespective of where it is being used going forward, it is always going to stay consistent. So, tomorrow, it is not gonna be that one team took it out and they have a very different value of what financial sales needs to look like. And then, marketing says that, oh, this is the value of financial sales. No, it all comes from the single layer. So, the very key and important part is that first get this groundwork set in and then what you do is then you have a business glossary, of course, which has the definitions of the metrics and then who the keyword for that metric definition is. So, if there are changes to that, then it's a communicated and published document. Then everyone has, and then they know what it exactly means. So, if there needs to be a change, then we need to go and follow protocol to essentially make that change within the system, put it on the semantic layer, and then go out for that other team itself. Yes, it's a little bit of a process, but it can be optimized. Once you have that, you have your data dictionary, you have your business glossary published, you have your semantic layer. Then, the third thing you do is basically you define the tools that can be used within the organization. So again, that is something that should be managed through the central team. And say for example, you say Tableau is the only tool that we have from a reporting analytics, dashboarding standpoint. And Tableau should be the only one that should be utilized within the organization. Now, if a team comes up with a specific use case and says, oh, you know what, Tableau doesn't work for our needs and we are going to need to use Power BI or we need to use Looker because...
Eldad: Sisense, of course, only Sisense.
Meenal: Yes, Sisense. And, if that's the one that is going to provide for our function. Then, you know, as a central data team, again, it's your responsibility to ensure that that tool is really required because sometimes it is just a matter of preference and not need. And so, it's your responsibility to ensure that tool is really the tool that you need to take forward or go forward with. But my statement here most is around the fact that it's your responsibility to also standardize the tools in use across the organizations and in cases where these functions are going to now be decentralized to the other team, so one thing they're going to require is access to your data warehouse or your data. And they're going to need a playground for themselves where they are going to start building their artifacts.
Eldad: They can actually change everything they want.
Meenal: Exactly. So they can change their stuff and push that stuff. And, then you have to tell, again, tools. So you have to provide them with the tools that they use so that they don't purchase their own tools. And so the total cost of ownership still stays constant. And then once you have that, you build out templates essentially as much as possible so that they can work within that same standards and guidelines. And then of course, continuous education. So, data literacy is a big part of what I do, and continually educating them in terms of what's right, what's wrong, how to use the data, what data exists, and how to use the data in a governed sense and in a private sense? Because in some cases the data that they may be taking in may be sensitive data and usage of sensitive data education is very, very key and important, as to how to do that. So, some of that becomes a little repetitive, but it's very, very essential for organizations itself. And, then, of course, a lot of training on the tools so that they are using the tools in the right way. And they're using it optimally for their needs. Now, this I talk about in very, very large organizations.
Now, you come to like smaller and mid-size organizations. In that case, you should minimize the amount of decentralization that actually happens. So from a data standpoint, you don't expect them to go and be running their own ETLs and doing anything beyond.
Eldad: It's the same stack, but just a free new addition.
Meenal: Exactly. I think that's a perfect analogy. Yes.
Eldad: Remove all the enterprise features, no auditing, no compliance, no security.
Meenal: No. So, you still have all of that, but it's not a central team. You still have education because they still have self-serve, so the ability to do self-serve, but your self-serve is now limited to where they're more dashboarding and then doing stuff from that point onwards. So, they are not responsible for bringing the data in or ETLs and you want to try to minimize that as much as possible, because it's a smaller organization, yours is a smaller team as well, so you want to kind of keep that management of it much, much more centralized. So, education still exists here because they still need to understand the importance of the data, what data exists, and how sensitive and private data should actually be utilized? So, that still exists from a literary standpoint.
Eldad: You picked the right computer so they don't burn the monthly budget.
Meenal: Exactly. So, you know that part of it still continues. But, I would say that's how I typically prefer that we organize data in, because again, large organizations there are just too much to manage. And, it's not essential that every time your data team doesn't have to be for 40 people, like a 40-people team to serve a larger organization. You can actually manage with a smaller team. It's just that you...
Eldad: We're going to do a sister show, that's called Data Politics unlike Data Engineering, very similar to Data Engineering.
Benjamin: The Data Politics show nice.
Eldad: The Data Politics Show and it is interesting though to see how data gravity affects data politics. And, it does.
Meenal: Yeah. I'm sure. I'm sure it does. Again, with the importance that data has across the organization, I'm sure there are politics associated with it. But, yeah, that's kind of how I look at self-serve, and that's how I prefer to manage it within organizations where I go and lead such efforts.
Eldad: Thank you.
Benjamin: In terms of, I just lost track of my train of thought. So, Tamar, when you listen to this, please cut it out.
Eldad: No, please. Tamar, please keep it. We never cut out. We've never cut anything out. We're not going to start now. Now I will ask, while Benjamin is getting his threads in order, what's kind of the most exciting thing coming to your team this year? What are you working on that's big and risky and supposed to make a big impact?
Meenal: Well, there are a couple of initiatives, I obviously can't go into details of them, due to privacy, but there are a couple of very interesting things the team has been working on. So, one really awesome thing that I was able to do for my team earlier this year was doing a data hackathon. So, you typically have software engineering hackathons, but data hackathons are very cool. They're much cooler than software engineers.
Benjamin: As a software engineer, I feel offended, but tell us how to host an amazing data hackathon?
Meenal: So, what we did, sorry, Benjamin. I was…
Benjamin: It is okay. I can handle it.
Eldad: They're doing three hackathons a week in Munich there. His team is doing like, oh, let's do a hackathon. Hackathons are unique, Benjamin, you do it once in a while and you eat pizza. So, but yeah. Sorry, go ahead.
Meenal: No, no. So, we did a hackathon. What we did though, Benjamin, we kind of set the topics previously, and what the topics where is these were longstanding problems within the organization and challenges that the organization was facing. And we were looking to see either to come out with a solve for it or with a prototype for it. So, either the outcome would be a plan, as to how we would tackle the problem or it would come out with a full solve, with a prototype. I'm happy to say like three out of the four projects, I don't have a large team. So, three out of the four projects that we did are like going live already. So, that was like, and those like, Eldad to your question is, the super exciting stuff that we actually did. All of them revenue generating ones again, and things that we were super excited about. So, we are taking some of that learning and essentially we have taken it a step ahead and we are looking for other similar revenue generating opportunities and we have found some similar such ideas and we are moving forward with that as well. So, again, that is like the fun and exciting stuff that's coming up. Again, can't go into much details, but yes so far.
Benjamin: For these open business problems initially, you worked on specific things, given how many different kinds of business functions there are that you guys are helping with as a data team? How did you decide basically which problems were worth tackling?
Meenal: So we have, as SaaS has, like you have a Freemium and the smaller version of your platform, and then you have the much more enterprise version of the platform. Our focus was to essentially see how we can increase the acquisition of customers here, like on the premium side of it, and then retain them and then get them to move to more paid and so what we did is that from the problems that we had. So, we had 12 problems that came to us and we had to choose four out of them and the four we selected were all related to. Eventually, that's how we kind of focused on it because that was very important for the organization and we wanted to make sure that we were helping with that, specifically given the time. So, that's kind of how we prioritized it. That's not to say that the other ones have been ignored. The other ones we have also taken on, we do intend to have like a virtual hackathon. This one we were fortunate to be able to do in person.
Eldad: But it's virtual, I'm just saying.
Meenal: Yeah. So, the intent is to kind of tackle those in the subsequent virtual hackathon. And then, of course, I'll let you all know how that goes because I have never done a virtual hackathon before.
Benjamin: Awesome. That sounds like a kind of huge challenge to get the energy. So, gotcha. So, in a data hackathon you start out with a specific business objective. You pick projects around that as you close out the hackathon. So this is the final part, like, how did demo day look or how did every team close out their data hackathon project?
Meenal: So, this is where our coolness factor ends a little bit in comparison to software engineering hackathon because you all can actually show, like…
Benjamin: I knew it.
Meenal: I made him happy again. Of course, our demo essentially was data, our demo was the solution. So, it was more PowerPoint slides than anything else. But, the outputs itself were super, super exciting. We had data science models built out for all of them. And so, we could showcase those models and show the output of those models itself. So, for our judges, it was so exciting to see that data live and see the answers to these questions. That was like the fun part of it.
Benjamin: But that sounds super cool. Like I don't see what the kind of missing coolness factors like we built kind of the query engine. So, our demos usually are, you click something, it's slow, then after the hackathon, you click something, now it's much faster. Because our query engine, like we improved some algorithm or something. So, okay. Maybe it's levels. We can agree the data hackathon is just as cool as the software hackathon.
Meenal: I agree. I changed my statement.
Benjamin: Perfect. So, on this kind of lovely note, let's wrap up today's episode. Meenal, any kind of closing words to our audience that you wanted to talk about in terms of data teams, etc?
Meenal: One thing I want to close out with, and I'm sorry we didn't get a chance to talk much about is the ROI of data teams. And I think, if you've heard my Montecarlo post as well, I talk about ATV, which is Adoption, Trust and Value. And, if you follow these three methodologies as you go and build out a strategy for your data team itself, you will realize that the value that your team provides far outweighs the cost of the investment that you have made in building the team out itself. So, that's one thing I would love to leave you all with.
So, adoption essentially talks about the fact, data democratization. So, as you build your semantic layer out and as you have the organization essentially adopt, your data platform itself, it's very essential that the adoption piece of it occurs, for the other pieces of it to happen. Now, the second part is trust. Of course, if there is no quality of data or if there's no trust and transparency within your data, the adoption is not going to be complete. And, so you have to ensure that that is the other pillar that you have to take care of and then the value generation starts happening is that once you have adoption and trust in, then you start producing value out of your platform itself.
The second thing I would like to leave you all with is the fact that nothing is possible without the team itself. And so it's very essential that your team has the ability to cross train, upscale and continuous learning has to be provided to the team so that they are kind of growing and provide them the opportunity to become as productive as possible in whatever it is that they do because the job that the data team does is always very underappreciated. And, so in order for the teams to be appreciated and for the teams to be more productive, you have to provide them the environment to actually be productive and really provide a satisfying outcome for everything.
So, two very, very key and very important things to take care of. But that's the message that I would probably leave this.
Benjamin: Awesome. I think those were great closing words. Thank you so much for joining in. So, this was an awesome kind of learning about how you think about building high performing data organizations. We look forward to hearing how the virtual hackathon, data hackathon, goes in the end. Awesome! Thanks for joining in, Meenal.
Eldad: Thank you for joining.
Meenal: Thank you so much.
Eldad: Bye-bye. Take care.
Meenal: Take care. Bye.
# Live Engine Upgrades, Zero Downtime: The Firebolt Method (/blog/live-engine-upgrades-zero-downtime-the-firebolt-method)
At Firebolt continuous improvement is a core principle: it doesn't matter whether we roll out a new feature, make a tiny improvement or address a bug - we aim for an excellent service for our customers. Starting with Firebolt version 4.0 all releases of our database core go through an online upgrade process. In this post we shed light on the internals and explain our decision making.
What bar do we set to ourselves when offering an online upgrade to our customers? It is surfaced on three major pillars:
* Zero-downtime – the service is not interrupted during the upgrade.
* Seamless functioning – the switchover moment requires no actions from customers: no restarts, no client reconnects.
* Unnoticeable – without performance impact on running queries no matter what the current load is.
## Engine Upgrade [#engine-upgrade]
A given Firebolt engine can contain one or more clusters and each cluster may consist of multiple engine nodes. Cluster configuration is homogeneous, meaning that clusters have the same number of nodes of the same type. From the client perspective an engine is exposed by a gateway service that hides from a user all actual running engine nodes. These nodes may belong to different clusters, but all have an exact same version. When we talk about the online upgrade we mean the change of version of engine nodes. All other Firebolt components: the gateway, control plane, UI can be upgraded by a standard Kubernetes gradual rollout process. But what makes engine nodes so special?
A graceful rollout process usually means launching a new instance of a service, rerouting new connections to the new instance, waiting for old connections to drain and then to shut the old instance down. However the process relies on one important fact - both old and new instances do not have any state and if they would have some then that shared state would be a separate service. Database processes of course have state - SSD storage, but besides this they also have local caches and having those caches hot is crucial for a high performance.
Firebolt engines have two main types of caches: data cache to avoid unnecessary S3 reads and subresults cache to optimize query performance. Both caches are local to an engine node. When we do an online upgrade we need to ensure that both caches are in a warm state before making a switch. Thus engine nodes with the new version need to run in parallel to the old one and also process queries. To protect user data from modification these new nodes can only execute read-only DQLs. We call such engine nodes Shadow.

Overall the upgrade process consists of the following steps:
1. Shadow creation – for each engine cluster we create a copy with identical instance type and number of nodes, but a new version.
2. Traffic mirroring – the gateway is configured to mirror all DQL queries to both Main (old version) and Shadow (new version) clusters.
3. Warm-up & verification – we keep running both clusters in parallel warming up Shadow's cache.
4. Switch-over – the gateway is switched to use Shadow cluster; all new queries are processed by the new version.
5. Drain – upgrade process waits for completion of all queries running in the old version.
6. Cleanup – the old cluster is removed.
The warm-up and verification step allows us to compare two versions of the product side by side on the same set of queries. The version comparison also gives an opportunity to try new versions out and do this without actual upgrade. This allows us to catch failures and performance regressions earlier in the release cycle and ensure that the release is in good shape before starting the rollout.
## Warm-up & Verification [#warm-up--verification]
So we have two almost identical clusters executing the same DQL queries and would like to ensure that at the end the new one can show the same performance as the old one. Besides a pure performance goal it also gives a good opportunity to detect binary regressions such as runtime errors. In particular:
* database process crashes
* queries that succeed on the old version but fail on the new one e.g. caused by changed validation rules
* and the opposite - queries that fail on the old version but succeed on the new one e.g. caused by a parser regression
The comparison logic is implemented in Firebolt's control plane, with the data retrieved from the local engine history. In particular it may give us query type, query duration, exception code if the query failed and other execution insights such as number of reads from S3. Once the gateway receives a user query it forwards the query to Main and if the query is a read-only DQL then to Shadow too. The gateway also tags a query with a timestamp watermark to ensure that the control plane can select the same set of queries disregarding the time when they are actually processed by a cluster.
How do we compare the performance of queries executed by Main and Shadow clusters? First logical idea would be to make a pairwise comparison, but in practice it doesn't produce stable results because:
* Main cluster also runs DML queries and they influence performance of DQL queries, thus giving Shadow cluster an advantage.
* Comparison of absolute duration values is significantly influenced by slowest queries contributing to the tail of the distribution.
* Comparison of relative values is impacted by randomness of the fastest queries.


The solution instead is to compare percentiles of distribution of query durations. So instead of making a pair-wise comparison we need to collect a set of queries and then calculate the percentiles. But how many of the queries do we need? A common answer is for moderately skewed data and a 95% confidence interval; sample sizes of a few hundred (e.g., 200-300) are often considered a reasonable starting point. However, for more heavily skewed data or when estimating extreme percentiles, you might need samples in the range of 500 to 1000 or even more. So it's time to look at the actual distribution:

If all queries were of the same type we could expect the distribution to be a sort of Gamma distribution with a longer tail. But we have different types of queries and each of the types follow some Gamma distribution, so the final distribution is a multimodal with multiple "hills" overlapping each other (one can see around 3 hills in the chart above).
Now we need to estimate the number of queries to put in a single sample set to calculate their percentiles. Basically we not only need the percentile value itself but also a confidence interval of the value. The narrower an interval the stricter an estimated value is. Ideally if we plan to accept 10% variation between Shadow and Main percentiles we need to require the confidence interval to be even lower. Going deeper into statistics, a confidence interval is also not defined strictly, but with e.g. 95% of certainty. Luckily all this can be calculated by using the bootstrap approach and then by iterating the sample size and checking when the confidence interval stabilizes.


While calculation of 50th percentile stabilizes at approximately 400 samples, the 95th only at 600. In practice these estimations are very dependent on the distribution and for some even a 1000 items is not enough.
From the implementation perspective the verification process is a loop, retrieving query performance metrics at each iteration. If the number of observations is enough for a confident analysis then we calculate the percentiles and compare the values from Shadow and Main. The process repeats until:
* Metrics successfully converge meaning there is no significant difference between Main and Shadow – a clear success!
* There are not enough observations to make any statistically significant decision, still the cache is warmed up, also a success!
Internally the verification process also looks into S3 reads operations and state of a subresults cache. If the engine is fully dedicated to analytics queries (DQLs only) then both caches warm up quite quickly. In the chart below X axis is local time and Y axis shows number of operations within a time window (defaults to 5 time intervals):


And finally this is how performance convergence looks as percentiles:

But what happens in case of a performance regression? If query durations on the upgraded version are worse than on the Main, then all percentiles shift upwards. Below is a chart corresponding to a performance regression detected in our end-to-end jobs. When such a condition is detected, the upgrade process is aborted and further root cause analysis performed.

In conclusion, Firebolt's online upgrade process prioritizes zero-downtime, seamless functioning, and minimal performance impact. By utilizing shadow clusters and traffic mirroring, the system ensures cache warm-up and runtime verification. This approach allows us to detect binary regressions and performance issues earlier delivering a robust and continuously improving service.
# Making a Query Engine Postgres Compliant Part I - Functions (/blog/making-a-query-engine-postgres-compliant-part-i-functions)
### TL;DR [#tldr]
This blog post gives a technical deep dive on how the Firebolt team forked off the ClickHouse runtime, and evolved it into a PostgreSQL compliant database system without compromising on performance.
### Introduction [#introduction]
Firebolt is a modern cloud data warehouse for data-intensive applications. This means that Firebolt is capable of both running large-scale ELT queries on terabytes of data, and serving homogeneous low-latency analytics at high concurrency (i.e. hundreds of queries per second).
At the heart of any cloud data warehouse lies the query engine. This is the part of the system that computes the query results: parsing and optimizing a query, scheduling it across the nodes in the cluster, and performing the actual computation such as projections, filters, joins, or aggregations.
As Firebolt is all about performant, cost-efficient analytics, the query engine needs to be as well. Since building a query engine is a massive undertaking, we originally decided to fork our runtime off [ClickHouse](https://clickhouse.com/) \[1]. ClickHouse is exceptionally fast and proved to be a great foundation in terms of performance \[2].
Beyond performance, we also wanted to make sure that we get Firebolt's SQL dialect right. Firebolt is SQL-only, meaning all operations are performed using SQL: provisioning compute resources, setting up network policies, interacting with RBAC, and - of course - ingesting, modifying, and querying data.
For us, getting the dialect right meant aligning with PostgreSQL (PG). The Postgres dialect is widely known and close to ANSI SQL. Almost every data engineer has used PostgreSQL or a Postgres compliant system before. This means that someone trying out Firebolt for the first time can feel right at home and be productive right away. As a database startup, aligning with Postgres also makes ecosystem integration much easier: making tools that work well with Postgres (i.e. all of them) also work with Firebolt becomes much simpler.
While ClickHouse is exceptionally fast, its behavior is very different from that of PostgreSQL. For the Firebolt query processing teams, this meant rebuilding large parts of the system. This was both a huge opportunity, as well as a huge challenge. We needed to re-architect the system while making sure that we remain extremely fast.
This blog post gives an in-depth view on how we evolved Firebolt's OLAP runtime to be PostgreSQL compliant, without sacrificing system performance. We give a deep dive into how modern vectorized database runtimes work, and the engineering challenges we faced moving away from the original ClickHouse runtime to the new PostgreSQL compliant Firebolt runtime. We also give a deep dive on how we bootstrapped our Firebolt SQL testing framework along the way, by both writing and generating Firebolt specific tests, as well as porting tests from other systems using custom automation.
Throughout the blog post, we focus on the changes we've made to scalar functions. Note that this is only a very small part of our larger PostgreSQL compliance story: we've built new aggregate functions, new data types for dates and times, as well as a completely new query optimizer from scratch. We'll talk about that in future posts.
***Disclaimer***: *throughout this blog post, we use example benchmarks using
[QuickBench](https://quick-bench.com). This makes it easy for you to play around with the examples
and try out custom performance optimizations you might come up with. When you run these benchmarks
on your own machine the results might be different: different compiler flags, target
architectures, etc. When you build a high performance query engine, you should always benchmark on
the same hardware with the same toolchain that your users will run on.*
### Getting the dialect right - postgres compliance [#getting-the-dialect-right---postgres-compliance]
Before we dig deep into modern OLAP runtime internals, let's talk a bit more about modern SQL dialects. If you've used different database systems in the past, you've most likely run into this: not all relational SQL databases are made equal.
### Why is the dialect so important? [#why-is-the-dialect-so-important]
SQL dialects can be different in many ways. At the surface level, they might differ in how data types or functions are called. When you dig deeper, even functions with the same name might behave differently. Beyond that, advanced language features around type inference, implicit casting, or subquery support (e.g. correlated subqueries) can be completely different.
In principle, there's a standard for all of this: [ANSI SQL](https://blog.ansi.org/sql-standard-iso-iec-9075-2023-ansi-x3-135/). The standard has thousands of pages, and we don't know of a system that implements everything that's outlined in the standard. Many systems also have custom dialect extensions that are not standardized. Out of all widely used database systems, [PostgreSQL is probably the closest to ANSI SQL](https://www.postgresql.org/docs/current/features.html).
For Firebolt's SQL dialect, we wanted to stay close to ANSI SQL. However, we also wanted our SQL dialect to be based on the dialect of a widely adopted real-world system. There are two core reasons why this mattered for us:
1. **Ease of use:** by aligning with an existing system, we can make it easy for new users to quickly be productive with Firebolt. While we also support custom dialect extensions, getting your first workload running on Firebolt becomes much easier by aligning with a widely used dialect. People don't have to spend a bunch of time reading your documentation, or getting used to weird idiosyncrasies of your dialect.
2. **Simplifying ecosystem adoption:** nobody uses databases in isolation. People use tools such as Fivetran for ELT, DBT for transformations, MonteCarlo for data observability, and Looker/Tableau for BI. When people adopt a new database system, they expect their existing tools to continue working. Virtually all of these tools generate SQL to interact with the database. By aligning with a widely used existing dialect, building an integration that can generate Firebolt SQL becomes much, much easier.
It's clear that to get the most value from both of the above points, it's not enough to just align with an abstract standard such as ANSI SQL. You need to align with a real system. For us, PostgreSQL was by far the most natural choice: it's widely used, people love it, it works with virtually all ecosystem tools, and it has the plus of being very close to ANSI SQL.
If you look at the wider database space, a lot of other systems have decided to align with PostgreSQL as well: [CockroachDB](https://www.cockroachlabs.com/docs/stable/postgresql-compatibility), [DuckDB](https://duckdb.org/docs/sql/dialect/overview), and [Umbra](https://umbra-db.com/#features) are good examples for this.
### Protecting users by throwing errors early [#protecting-users-by-throwing-errors-early]
PostgreSQL is quite strict about functions throwing errors. Functions can throw errors for many reasons. Examples might be overflow checking, or arguments being out of bounds (e.g., trying to take the logarithm of a negative number).
Throwing errors in these cases makes it easier to use the system: instead of returning potentially meaningless results, the query engine lets the user know that they aren't using the functions in a way that's safe. This matters a lot in the case of Postgres: if you use an OLTP system like Postgres to power core parts of your business (e.g., your shopping backend), you don't want to receive wrong results because no overflow checking was performed.
In a similar way as for an OLTP engine, throwing errors early also matters for a cloud data warehouse: people use Firebolt for heavy ELT jobs that ultimately serve their internal or customer-facing analytics. Wrong results propagating through your data pipelines can be hard to catch, and the engine failing silently and propagating wrong/meaningless results makes this much worse.
However, extra safeguards such as overflow checking often come at a cost: there are usually extra checks required to sanitize the arguments and check for errors. If you look at the Postgres implementation of adding two four byte integers for example, [you'll find the following code](https://github.com/postgres/postgres/blob/2488058dc356a43455b21a099ea879fff9266634/src/include/common/int.h#L104):
If there is no builtin intrinsic for addition with overflow checking, Postgres casts the four byte integers to 8 bytes, and then adds the two four byte values. This can't overflow. It then checks if the result is outside of the four byte integer range. If it is, it returns true indicating that an overflow occurred. Otherwise it casts back to a four byte integer.
Let's write a C++ microbenchmark to figure out how much more expensive this actually is. We implemented a quick benchmark [here](https://quick-bench.com/q/s1rtHEsryaNgyUIF9Pi5xTpxamY) using QuickBench that compares (1) addition without overflow checking, (2) addition using the manual Postgres overflow checks, and (3) addition with overflow checking using \_\_builtin\_add\_overflow.
If you look at the code, you'll see that we perform the addition on dense vectors of 1024 rows. This is how modern high-performance OLAP engines work internally, you'll learn more about this in the next sections.
The primitives for addition with overflow checking can most likely be tuned further: we could try to manually vectorize them, be smarter about avoiding overflow checks in some cases, or not throw exceptions in the hot loop. But the naive implementation nicely shows that we're paying a hefty price for the extra safeguards: both using \_\_builtin\_add\_overflow() and the manual PG-style check is about 4x slower than the version without overflow checking. The microbenchmark also nicely shows why it makes sense to go for the builtin when available: it's about 1,7 times faster than the manual check.
Because of this, avoiding these checks might make sense for some systems: if you're focused on serving already cleaned-up data, optimizing to squeeze out the last bit of performance can be a good choice. This moves more responsibility to the user who needs to use the system in a safe way, but implementing a very fast system becomes easier.
We believe that for a data warehouse, having defensive function implementations that raise clear errors to users is a must. Aligning with Postgres here was not negotiable for us, as we believe that it's the right thing to do for our users. The good news is that it's still possible to build an extremely fast engine with defensive function implementations: for most queries only very little time is spent in scalar functions as other operators such as joins or aggregations are computationally more expensive. Even for a query where a lot of time is spent in the scalar functions, it's possible to tune the function implementations to provide very good performance.
We decided to spend a lot of time performance tuning our runtime to provide a customer experience that's both safe and extremely fast. We'll drill into this in the next sections.
### Firebolt's roots in clickhouse [#firebolts-roots-in-clickhouse]
Alright, with the background on modern SQL dialects behind us, let's start diving into technical details of how we built Firebolt. This section gives an overview of why and how we forked our runtime off ClickHouse.
Building a database system is a daunting task that requires a very large engineering investment. This runs counter to many of the goals you have as a startup: going to market early, finding first design partners and customers, and iterating with them to build a really great product.
To go to market quickly and iterate with real companies, we wanted to fork off an existing open-source system. For us, ClickHouse was the only natural choice. We had three core criteria when it came to the runtime:
1. **High-performance vectorized engine:** we wanted the query engine to be a state-of-the art, low-latency OLAP runtime. If you've never thought about what makes such a runtime special, the next section will tell you more about it.
2. **Battle-tested:** we wanted the runtime to be battle tested and widely used for production use-cases.
3. **Scale-out engine:** we wanted the runtime to have basic support for distributed query processing. As a data warehouse we need to be able to handle ELT queries on massive data volume within a fixed time budget, and scale-out support is a must for this. Nowadays, we've completely rebuilt our distributed query processing layer, but the ClickHouse runtime was great for us in the early days as it allowed for basic support for scale-out processing.
If you want to dig deeper, take a look at our CDMS\@VLDB'22 paper on just this topic \[1].
At this point, we want to give a huge shoutout to all ClickHouse contributors: **ClickHouse has been an exceptional system to innovate upon**. It's incredibly fast, layered in a clean way, and elegantly makes use of the building blocks required for modern high-performance OLAP engines such as vectorized processing, a multithreaded and push-based query engine, and efficient columnar storage.
While the rest of the blog post will talk about things we've changed in significant ways, this has only been possible because ClickHouse provided a great architecture that allowed us to make these changes. We made these changes because we believe that it's the right thing to do for the workloads that matter to Firebolt customers, but different choices make sense in the context of different systems and workloads.
By forking off ClickHouse, we managed to achieve our goal of quickly having a functioning, high-performance system. However, we also realized that we had a very long way to go to actually make our query engine Postgres compliant.
Historically, ClickHouse was built for massive-scale serving & reporting workloads at Yandex. Because of this, they've made the conscious choice that they want to optimize their dialect for allowing peak performance for these specific workloads.
As we've seen in the previous section, we decided that this wasn't our desired path for Firebolt's dialect: as we want users to run large data pipelines with complex transformations, raising errors early was extremely important for us. This meant that we had to rebuild large parts of the runtime in order to become Postgres compliant. The next sections will discuss this journey in detail.
### Modern OLAP Query engines [#modern-olap-query-engines]
Why are some query engines faster than others? While two systems might speak a similar SQL dialect, they can be built in completely different ways when you take a look "under the hood".
Historically, most systems implemented a row-at-a-time interpreter for relational algebra \[3]. This is also called "Volcano"-style query execution. Many OLTP systems such as Postgres are still built that way. In a Volcano engine, individual rows are moved from one operator in the query plan to the next. They flow "upwards" from scan operators through filters, projections, aggregations, etc. Once rows arrive at the root of the query plan they can be returned to the user.
At a technical level, these systems implement relational operators such as filters, aggregations, or joins using a simple getNextRow() interface. When calling that function on a filter for example, the filter repeatedly calls getNextRow() on its child operator until a row matches the filter. This row is then passed to the parent. Your query optimizer builds the relational algebra tree for the query, and the result can be retrieved by calling getNextRow() on the root operator until the tuple stream is exhausted.
For OLTP systems that don't process massive datasets in the context of a single query, such an approach performs well. Other things like the buffer manager or indices are much more important for high performance in these systems. For an OLAP engine that does analytics on terabytes of data however, a row-at-a-time interpreter leads to poor performance: it has bad code locality, lots of virtual function calls, and isn't tailored to modern CPU architectures.
The MonetDB and MonetDB/X100 projects from CWI redefined how to build a high-performance analytical query engine \[4, 5]. Instead of implementing a row-based interpreter, MonetDB performs in-memory columnar query execution with specialized primitives. X100 went beyond that by passing batches of around one thousand rows through the query engine. Such query engines are called "vectorized" query engines.
The data flow in vectorized engines is similar to the one in a Volcano engine. Rows also flow "upwards" the query plan from scan operators all the way to the root of the query plan. Compared to the Volcano model where you just pass a single row from one operator to the next however, you now pass a whole batch of rows between them. These rows are usually stored in a columnar layout, meaning that a single column in a batch occupies a dense memory region.
Modern vectorized engines are much more efficient than traditional Volcano engines: the working set fits into the CPU caches, code locality is improved and virtual function calls are amortized over multiple rows. Most importantly however: CPUs love operating on batches of data. The code in the engine is much more friendly towards modern CPU architectures. It can leverage prefetchers, branch predictors, and superscalar pipelined CPUs much more efficiently.
The code in such a query engine looks similar to the microbenchmark we showed earlier. A scalar function for example receives a batch containing a few thousand input rows, and computes the function result for the entire batch:
These primitives are the basic building blocks of your vectorized query engine. You need them for all your supported scalar and aggregate expressions. More complex relational operators such as joins and aggregations should also always operate on batches of tuples. This can be especially beneficial when operating on large hash tables: by working on a batch of rows you can issue lots of independent memory loads and saturate your memory bandwidth.
Most modern OLAP engines implement vectorized query execution. ClickHouse, the system Firebolt forked from, is built in exactly this way. This is one of the core reasons why ClickHouse is so fast. There's a great paper about their engine at this year's VLDB \[2]!
Many other modern OLAP systems such as DuckDB \[6] (also from CWI, the group where vectorized execution was invented), Snowflake \[7], and Databricks' Photon \[8] follow this approach.
Note that there is a second approach to build modern high-performance query engines: some systems such as Hyper \[9] or Umbra \[10] just-in-time (JIT) compile a declarative SQL query into machine code \[11]. While this is super cool technology, it doesn't really matter for this blog. In practice, vectorized and compiling engines tend to perform similarly \[12].
### A PG compliant runtime - scalar functions [#a-pg-compliant-runtime---scalar-functions]
You now know how modern vectorized query engines work: they have simple primitives operating on batches of a few thousand rows, allowing them to efficiently leverage modern CPUs. You also know that Firebolt forked from ClickHouse, which is a modern OLAP engine that implements vectorized execution. In this section, we will dig into how we evolved our runtime to become PostgreSQL compliant.
### Keeping users safe - more exceptions [#keeping-users-safe---more-exceptions]
We mentioned earlier that we decided to align our dialect with PostgreSQL because it's widely used, close to ANSI SQL, and protects the users by throwing errors early. This meant that we had to take the parts of ClickHouse that aren't PostgreSQL compliant, and rewrite them to align with our target dialect.
ClickHouse decided to optimize their dialect for pure performance. In the context of ClickHouse workloads, this makes sense. This means that many functions do not throw exceptions when PostgreSQL does. Let's look at a few examples. You can easily try these out yourself using tools like [ClickHouse Fiddle](https://fiddle.clickhouse.com) or [OneCompiler](https://onecompiler.com/postgresql) for Postgres.
All of these examples are quite simple, but imagine subtle conversion errors happening deep inside of an ELT job. This makes debugging hard or – even worse – you might not even notice that things went wrong. This is why we believe that a data warehouse should throw errors in all of the above examples.
## Postgres compliance - one function at a time [#postgres-compliance---one-function-at-a-time]
This meant we had to carefully evaluate [for every function in our dialect](https://docs.firebolt.io/godocs/sql_reference/functions-reference/functions-glossary.html) whether it's compliant with PostgreSQL. For functions that are not supported by PostgreSQL, life is a bit easier, but we still wanted to make sure that our implementation matches PostgreSQL "in spirit" and raises errors early.
It's hard to come up with all the subtle ways functions might behave differently. As a result, for any function that we added to our SQL dialect, we followed a checklist to ensure that we align with Postgres:
1. Port SQL correctness tests from the PostgreSQL test suite
2. Port SQL correctness tests from other open-source SQL test collections. In our case we always checked at least DuckDB, Sqlite, MySQL, SparkSQL, and ZetaSQL
3. Write custom Firebolt tests, explicitly checking for NULLs, corner cases, etc
This effort ultimately made us completely overhaul the scalar and aggregate functions in the Firebolt runtime. For most functions in our dialect, we either (1) completely reimplemented them, or (2) took the original ClickHouse implementation and changed it significantly to change error handling.
This was a lot of work, and we built some tools to make it easier for us to port tests and ensure that Firebolt really matches Postgres behaviour.
#### SQL Test driver [#sql-test-driver]
We have a simple SQL-based test driver that's hooked up in our CI. Using this test driver, you can very easily define SQL tests. You can write arbitrary SQL queries (DDLs, DMLs, DQLs), and then check that the query's result set matches the expected result. Here's an example test that our CTO Mosha wrote that I especially love:
Our SQL tests can run in different environments:
* A "local" mode that mocks away dependencies such as S3/MinIO and our metadata services. This mode allows for extremely fast iteration times when working on the runtime.
* A "remote" mode that can run against a version of "Firebolt light" running on a local developer machine. This makes full use of MinIO and our metadata services. Compared to the local mode, this allows us to test much more of our system: interacting with object storage, SSD caching, and transaction processing.
* A "cloud" mode that runs against a [Firebolt engine](https://www.firebolt.io/resources/firebolt-elasticity-technical-whitepaper) running on AWS. This also tests our cloud infrastructure, routing layers, and authentication flows.
The "local" and "remote" modes run in the CI on every PR that gets merged to main. We nowadays have more than 100.000 SQL queries stressing every part of our dialect that run on every commit. The "cloud" version of the SQL tests runs nightly and before every release.
### Test Transpilers [#test-transpilers]
While almost all database systems implement text-based SQL tests as the ones mentioned above, many systems use slightly different formats. DuckDB uses a similar format to Sqlite. Here's an example of the [DuckDB tests for length](https://github.com/duckdb/duckdb/blob/main/test/sql/function/string/test_length.test):
To make it easy to port tests from other systems, we wrote custom transpilers that take text-based tests from ZetaSQL, DuckDB, etc. and automatically turn them into Firebolt's SQL test format. This made it much easier to quickly cover Firebolt's SQL dialect with tests from other systems.
### The Postgres executor [#the-postgres-executor]
When porting tests from other systems, or writing our own tests, we wanted to make it easy to ensure that Firebolt actually matches the Postgres behaviour. For this, we wrote an additional execution mode for our SQL tests: the Postgres executor.
This mode uses the Firebolt SQL tests, but doesn't run against Firebolt at all. Instead, it sends the queries to a locally running Postgres server. We can then take the Postgres output and render it in the same way as our text-based format requires.
When adding a function, this makes it super easy to actually ensure that Postgres returns the same results at Firebolt. Of course, this execution mode doesn't help with custom Firebolt dialect extensions such as our [array functions](https://docs.firebolt.io/godocs/sql_reference/functions-reference/array/).
### Overhauling our runtime primitives [#overhauling-our-runtime-primitives]
In addition to writing a lot of tools for testing, we also overhauled how we write scalar functions in our C++ runtime. As you can see in the SQL examples above, ClickHouse doesn't raise exceptions for many functions. This leads to some interesting runtime implementation details when handling NULL values.
In the section on "Modern OLAP query Engines" we've seen that vectorized engines such as ClickHouse always operate on batches of rows. These batches are stored in a columnar format. For primitive data types such as integers, the columns are represented as contiguous memory regions (vectors) packing one integer after the next. More complex types such as TEXTs or ARRAYs usually store an additional vector for TEXT/ARRAY lengths.
For columns that can be nullable, vectorized engines usually store which values are NULL in an additional vector. There are different ways to represent these NULLs \[13], Firebolt implements a byte mask indicating which indices in a column of a batch are NULL.
Let's understand how the following data would be represented in an in-memory batch:

The five integer values for Id are stored in an integer column that has four bytes for every entry. The balance values are stored in a double column that has eight bytes per entry. Separate one byte null maps store one byte for every value in the source column. If the null map entry is set to one, this indicates that the value is NULL.
Note that for any value that's NULL, there's still an actual value stored in the data column "behind" the NULL mask. In the above example, this value is set to 0 for the two NULLs in the balance column.
Most scalar functions propagate NULLs: an input NULL also results in an output NULL. At first glance, this seems to yield a very simple implementation for vectorized primitives that contain NULLs. For some scalar function\_implementation, the primitive can just look as follows:
You take the logical OR of the input NULLs, and call the function implementation on all values. The above code snippet is great from a performance standpoint: working with NULLs does not require any extra branches, just the extremely efficient logical OR of two byte maps. Splitting the primitive into separate for loops for the null map and the function implementation improves performance. This is how ClickHouse implements many of their vectorized primitives. **However, this only works if function\_implementation does not throw errors.**
In the example above, imagine we wanted to run a query such as SELECT log(in\_game\_balance) as res FROM players. A primitive that looks like the one above is only safe in the context of the ClickHouse log implementation. As the log of 0 returns -inf rather than raising an error, we can safely perform the computation on all nested values. The output block simply looks as follows:

We keep the original NULL map (we only need OR for n-ary functions), and evaluate the logarithm on all nested arguments. The output column now has -inf stored "behind" the NULL values.
This breaks down completely once you start throwing frequent errors in your runtime. **We cannot evaluate the function on values that are "behind" NULL, as they might cause errors.** In Firebolt, log(0) throws an error, and if our primitives were implemented in the same way, we would throw an error in the above query computing log(balance), even though all balances are positive or NULL.
When implementing the above primitive for a function that can throw errors, you need to be more careful:
This code is less friendly for modern CPUs, as there's a branch (!out\_nulls) in the second nested for loop. However, it's safe if function\_implementation can throw errors. We could also write the primitive in a single nested for loop. However, the above version is more efficient. [Let's write another Microbenchmark](https://quick-bench.com/q/F8cbQk8BnX_21ifzNZowRk1RbEA) on QuickBench to compare the performance. Here, we compare (1) an unsafe primitive that performs addition without overflow checking on all values, irrespective of if they are null, (2) a primitive that changes the control flow to not perform operations behind null values, but still does not perform overflow checking, and (3) a primitive with changed control flow and overflow checking.
Comparing (1) to (2) allows us to see how expensive just the change in control flow is. Comparing (1) to (3) allows us to compare end-to-end performance of the primitives. We run the benchmark with a 10% null ratio.
We can see that the change in control flow with the extra branches creates a \~4.3x overhead. Overflow checking causes another 20% slowdown. When you look at the generated assembly, it becomes quite clear why the version without overflow checking is so much faster: it can unroll the loop more effectively and leverage SIMD registers. We've added some comments to show you what is happening where. Note you won't see the null map computation here. That happens in a separate loop before:
Meanwhile, the generated assembly for the safe primitive with overflow checking doesn't do any unrolling:
This shows that when implementing scalar functions that throw exceptions, it's quite easy to end up in a state where the performance of your vectorized primitives suffers.
It might be comfortable to ignore this problem: when executing a query in a system, there's usually a lot more going on: table scans might read data from SSD, you might perform more expensive relational operations such as aggregations, and you need to serialize query results. In many cases, even a 4x slowdown in your addition primitive won't have a large effect on end-to-end query runtimes.
However, the performance difference of the primitives can become noticeable for some query patterns: complex expression trees, selective filters early on in your query, or small scans where data is stored in main-memory. Firebolt is all about high-performance, cost-effective SQL processing. As a result, just living with a 4x slowdown for some of our core primitives was not acceptable. The next section will provide a deep-dive into how we tuned our scalar functions further to make them as fast as possible.
### Keeping firebolt fast - function dispatching [#keeping-firebolt-fast---function-dispatching]
The previous section taught us that it's harder than it might seem to keep a query engine very fast while still throwing exceptions early. The extra argument and result checks in the inner hot loop of the vectorized primitive make it harder to generate very efficient code for modern CPUs. This section shows how we tuned our runtime primitives to still provide the fastest possible performance in as many cases as possible.
#### Dispatching for numeric functions [#dispatching-for-numeric-functions]
Usually, people run SQL queries that don't throw exceptions. Let's stick with the example of an ELT job. Throwing exceptions matters a lot when getting an initial version of an ETL query running. But once you have that query running, you want it to run every day/hour/minute and ingest fresh data. Exceptions might still happen here if e.g. your data distribution changed, in this case they are helpful and alert you that you should take a look rather than serving wrong results. But exceptions should be the exception, not the norm.
Can we use this property to actually improve the performance of our system? If we know that it's safe to do addition without overflow checking on a block of data, we can save the cost of doing overflow checks. We can then execute a ClickHouse-style, super efficient primitive, and be sure that we don't miss over- and underflows.
One thing that's great about vectorized engines is that since they operate on at most a few thousand rows at a time, the working set of an expression fits fully into cache. Maybe we can do a first pass over the function, do a cheap check with one-sided error that ensures that the addition can definitely not overflow, and then run addition without overflow checking whenever possible.
We know that adding two four byte integers can definitely not overflow if they are in the range of -1 billion and 1 billion. We can use this property to rewrite our vectorized primitive as follows:
Time for [another microbenchmark](https://quick-bench.com/q/EAV5XgdTCfasKVUjP7nF8guy0TU). We're now comparing (1) the super fast unsafe primitive, (2) the "safe" primitive that always does overflow checking, and (3) the attempted optimization above. There are no rows that overflow, and 10% NULLs.
This is a pretty good improvement. We went from being \~5x slower than the original ClickHouse primitive, to being just \~2.7x slower.
What happens if we can't choose the fast specialization? We can also test this in a simple [microbenchmark](https://quick-bench.com/q/GtWkI7at0ylYyKpsqYezbssCw1o).
We can see that in those cases, the extra checks cost about 30%. As these cases should be rare, being \~2x faster in the common case seems worth it. There are ways to improve the situation further: if we see that we often choose the specialization with overflow checks, we could at query runtime decide to just always call the safe specialization and not do the extra range checks. We currently don't do that in Firebolt, but plan on doing so in the future.
Addition is one of the cheapest possible functions you can calculate. So getting within \~2.7x of the original primitive is already quite amazing. For more expensive functions, the performance of dispatching like this will become even closer to the performance of the original primitive.
For different functions, you can use different tricks to maximize performance. In many cases, first doing an efficient loop over the arguments to check for invalid values is a good choice. Then you can implement a very efficient loop afterwards that doesn't require nested control flow for argument checks.
#### Dispatching for text functions [#dispatching-for-text-functions]
Smart dispatching such as above can help a lot for squeezing even more performance out of primitives for numeric functions. However, numeric functions tend to be quite cheap. Text functions in a database system are often much more expensive. Firebolt uses a similar dispatching concept as above to speed up common operations on text columns.
We've seen before how we encode simple columns in our vectorized engine. TEXT columns are stored in two dense memory regions: one containing the text byte sequence, and one containing the offsets of where the individual TEXT values start. There is one more offset than there are TEXT values in the column. The following example column contains three TEXT values: "HELLO, WORLD", "FIREBOLT", and "BLOG POST":

While the text data is variable-sized, the offsets are fixed-size. The offsets act as pointers into the byte sequence and make it easy to work with a string at a specific index.
In Firebolt, TEXT data is always represented as a [UTF-8](https://en.wikipedia.org/wiki/UTF-8) byte sequence. This means that a single text character can have multiple bytes in the binary representation. The [🔥](https://apps.timwhitlock.info/emoji/tables/unicode#emoji-modal) emoji for example becomes 0xF09F94A5, which is a four byte code point.
The offsets act as byte indexes and not character indexes in the column. This means if we replace "Firebolt" with the 🔥 emoji the column would be encoded as-follows:

Even though 🔥 is just a single UTF-8 character, the byte-based offsets show that four bytes storage are required.
Most TEXT functions operate on the UTF-8 characters. When calling length(\) you care about the number of characters, not the number of bytes in the underlying encoding. This means length('🔥') needs to return one, not four.
Because of that, length can't just subtract the neighboring offsets in the text columns. It actually needs to look at the byte sequence and understand the UTF-8 encoding. It's [still quite easy](https://github.com/ClickHouse/ClickHouse/blob/c0b36c946d9b069bfd10687ed6e6e61ab7e7a908/src/Common/UTF8Helpers.h#L61) to build an efficient length function that way, but it needs to iterate over the whole byte sequence.
The good news is that for many use-cases, TEXT columns just contain ASCII characters. These are only single-byte code points. If we know that a TEXT column only contains ASCII we can implement length by simply iterating over the offsets.
Luckily, checking whether there are only single-byte code points is incredibly cheap. For many functions, it thus pays off to first check whether there are only such code points, and then dispatch into a more efficient specialization. This is a very similar trick to the one we've used in the section before for numeric functions.
Let's [write a microbenchmark](https://quick-bench.com/q/fsqTAEZ8Q-vlw0vcsr7adZs96VQ) for the example of length. We compare (1) a version that checks the actual byte sequence for the number of code points, and (2) a version that checks if all code points are single-byte and in that case uses the offsets to compute the length. We run on ASCII strings of size 10.
The improvement here is super impressive. By checking whether there are only single-byte UTF-8 characters and then calling into the faster specialization, length() becomes 11x faster than before! When it turns out that there are multi-byte code points, the above numbers also show that the overhead will be \<10%. This is because the fast version is an upper bound for the overhead of the code-point checking.
We use this trick for other text functions as well. Further nice examples where this trick can be applied are [lpad(\, \\[, \\])](https://docs.firebolt.io/godocs/sql_reference/functions-reference/string/lpad.html) and [rpad(\, \\[, \\])](https://docs.firebolt.io/godocs/sql_reference/functions-reference/string/rpad.html).
Our implementation here is actually even cooler than what we've shown above. We implement these UTF-8 checks as lazily computed column properties. These properties are attached to blocks we pass through our query engine. This means that the primitives don't always recompute whether a text column only has single-byte UTF-8 code points. Instead, the property is computed the first time it is needed and then cached. This reduces the overhead even further.
#### Testing for performance [#testing-for-performance]
As you've seen above, microbenchmarks are awesome. They allow you to test a specific code fragment for performance in isolation. Performance tuning isn't easy, and in many cases choices that seem good at first glance can make performance worse.
For all scalar functions we're building at Firebolt, we're writing our own microbenchmarks to validate that they are efficient. In the same way as QuickBench above, we're also using Google's [benchmark](https://github.com/google/benchmark) library. The microbenchmarks need to cover all the different ways the functions might be called: different argument distributions, different NULL ratios.
One thing that really helped us is that after forking, we had all of the original ClickHouse primitives that don't do argument checking. In the microbenchmarks for our Firebolt functions, we thus always had an efficient baseline to quantify whether we're doing well.
#### Looking ahead - dispatching based on storage statistics [#looking-ahead---dispatching-based-on-storage-statistics]
The more efficient function dispatching helps us in many cases to make our primitives faster. However, there's usually still some overhead associated to check whether we can call a more efficient specialization.
Our goal is to get the overhead to zero for as many cases as possible. And we believe we can get to zero quite often. For this, we plan on integrating metadata from our storage layer to allow for more efficient dispatching. Firebolt maintains statistics such as tablet-level min/max for every column. This means that when scanning a vectorized block from storage, we can attach the tablet's statistics to the block. We can then use those statistics to directly dispatch into more efficient specializations.
In the example for addition without overflow checking, this would mean that any integer block we read from storage with absolute values smaller than 1 billion can be safely added without overflow checking.
Propagating this "up" the query engine through complex relational operators isn't easy. But as the most expensive expressions are usually evaluated right after a scan, even just having the statistics then can already help a lot. We look forward to talking more about this in a future engineering blog once we've implemented it.
### Conclusion [#conclusion]
This blog post gave a deep dive of how we made scalar functions in Firebolt's runtime PostgreSQL compliant. Firebolt decided to align with PostgreSQL's dialect as it's widely used, loved, and provides protections for users by throwing errors early.
To get to market quickly, we decided to fork our runtime off ClickHouse. ClickHouse is an exceptional, modern, vectorized OLAP system. However, our push for Postgres compliance led us to rebuild most scalar functions. We've shown that great testing infrastructure was the backbone of this effort. We have extensive SQL-based tests that can run both on developer machines and on the cloud. And we've invested into test transpilers that make it easy to port tests from open-source systems into our testing environments.
We've seen that making functions that throw errors extremely fast is hard. Sanitizing arguments introduces branches in the hot loops of vectorized primitives, which makes it harder to maximize performance on modern hardware.
To tackle this challenge, we've invested a lot of time into dispatching to efficient implementations of the primitives whenever possible. When we know that addition cannot overflow for example, we call primitives that don't perform overflow checking. When we know that a text only consists of single-byte UTF-8 code points, we call specialized text functions. We're continuing to invest into this by, e.g., enriching our query engine with statistics provided by the storage engine.
This is just a very small glimpse of the work required to build a very fast OLAP query engine that's PostgreSQL compliant. In future blog posts we look forward to diving in depth on how we implemented new PostgreSQL compliant data types, and rebuilt our query optimizer from scratch.
##### References [#references]
\[1] Pasumansky, Mosha, and Benjamin Wagner. "Assembling a Query Engine From Spare Parts." *CDMS\@VLDB*. 2022.
\[2] Schulze, Robert, eta al. "ClickHouse - Lightning Fast Analytics for Everyone". *Proceedings of the VLDB Endowment 17.12* (2024).
\[3] Graefe, Goetz. "Volcano - an extensible and parallel query evaluation system." *IEEE Transactions on Knowledge and Data Engineering* 6.1 (1994): 120-135.
\[4] Nes, Stratos Idreos Fabian Groffen Niels, and Stefan Manegold Sjoerd Mullender Martin Kersten. "MonetDB: Two decades of research in column-oriented database architectures." *Data Engineering* 40 (2012).
\[5] Boncz, Peter A., Marcin Zukowski, and Niels Nes. "MonetDB/X100: Hyper-Pipelining Query Execution." *CIDR*. Vol. 5. 2005.
\[6] Raasveldt, Mark, and Hannes Mühleisen. "Duckdb: an embeddable analytical database." *Proceedings of the 2019 International Conference on Management of Data*. 2019.
\[7] Dageville, Benoit, et al. "The snowflake elastic data warehouse." *Proceedings of the 2016 International Conference on Management of Data*. 2016.
\[8] Behm, Alexander, et al. "Photon: A fast query engine for lakehouse systems." *Proceedings of the 2022 International Conference on Management of Data*. 2022.
\[9] Kemper, Alfons, and Thomas Neumann. "HyPer: A hybrid OLTP\&OLAP main memory database system based on virtual memory snapshots." *2011 IEEE 27th International Conference on Data Engineering*. IEEE, 2011.
\[10] Neumann, Thomas, and Michael J. Freitag. "Umbra: A Disk-Based System with In-Memory Performance." *CIDR*. Vol. 20. 2020.
\[11] Neumann, Thomas. "Efficiently compiling efficient query plans for modern hardware." *Proceedings of the VLDB Endowment* 4.9 (2011): 539-550.
\[12] Kersten, Timo, et al. "Everything you always wanted to know about compiled and vectorized queries but were afraid to ask." *Proceedings of the VLDB Endowment* 11.13 (2018): 2209-2222.
\[13] Zeng, Xinyu, et al. "NULLS!: Revisiting Null Representation in Modern Columnar Formats." *Proceedings of the 20th International Workshop on Data Management on New Hardware*. 2024.
# Making Firebolt Fast By Doing Practically Nothing (/blog/making-firebolt-fast-with-pruning)
## TL;DR [#tldr]
This post describes different methods for reducing the number of scanned rows in a query (a.k.a. pruning), which improves query performance. From these methods, the post chooses to focus on sketch-based row pruning, its benefits, and how it's represented in Firebolt.
## Introduction [#introduction]
A famous saying in the world of low-level development is "The key to making programs fast is to make them do practically nothing". Unsurprisingly, this is also true for SQL databases. During the planning phase, Firebolt tries different optimizations in order to make a query plan more efficient. An inefficient plan usually involves doing the same work many more times then we have to. For example, the plan may run a scalar operation on 1B values, even though we're going to discard them later. One such optimization to avoid doing unnecessary work is row pruning.
To discuss why it matters, let's say that for an online game, we have a table `players` with 30 columns, on which we run the following query:
```sql
SELECT id
FROM players
WHERE hours > 10000 AND country = 'Laputa';
```
Since Firebolt is already a columnar database, we only read the data for the columns `id`, `country`, and `hours` from storage, avoiding all other columns. However, let's say that only 0.1% of our playerbase are Laputans. Furthermore, let's say that due to legal issues, the game was only introduced to the people of Laputa 450 days ago. Since there's only 24 hours in a day, a player that has more than 10,000 hours had to play it for at least 416 days, meaning that only the players who bought the game in the first 34 days after it was published, about 20% of all Laputan players, are eligible, reducing the 0.1% into 0.02%.
If we don't use any of this knowledge, then we are going to end up reading the data for all players and then use only a 5000th of it. For queries as computationally simple as this, where most of the time spent by the query is spent on I/O, reading all that data would be a major bottleneck that we'd like to avoid. It is for this reason that row pruning is crucial.
## What is row pruning? [#what-is-row-pruning]
Just as column pruning allows the database to avoid reading columns that are not required for evaluating the query, row pruning refers to every optimization that allows the database to avoid reading rows that are not required for evaluating the query. Some commonly used forms of row pruning are:
* Secondary indexes - For a specific expression, a data structure is constructed during insertion that allows the database to quickly retrieve, for every X, the list of rows for which the expression is less/more/equal to X. [B-trees](https://en.wikipedia.org/wiki/B-tree) allow retrieving this information with at most `O(log(n) + k)` reads from disk, where `n` is the total number of rows in the table, and `k` is the number of rows satisfying the search criteria.

* Partitions - For a specific expression, the table is internally divided into subtables where, for each row in a subtable, all values of the expression are the same.

* Clustering indexes - The table itself is sorted by the index expression, and the index data structure now indexes ranges (named "granules") within the table instead of individual rows. As the index data structure now indexes fewer things, it can become small enough to fit entirely in RAM.

Firebolt supports both partitions and clustering indexes. Yet, both of these methods present two problems:
* You must explicitly decide which column(s) will be used for partitioning/indexing.
* Unless we're willing to create several copies of the table, only one set of column(s) can be used for partitioning/indexing (note that B-tree based secondary indexes do not suffer from this limitation).
If we look at our original example, the query's `WHERE` clause checked two fields: `hours` and `country`. Unless usage of queries on these fields is very frequent, it is very unlikely that you would use them as an index/partition, since using these fields prevents them from picking other fields, which might be used in much more frequent queries. As we've mentioned above, this specific query is also very correlated with the date when the player registered in the game. However, taking advantage of this correlation requires a planner smart enough to deduce it, which isn't trivial.
In order to overcome these problems, Firebolt uses a fourth type of pruning: sketch-based row pruning.
## What is sketch-based row pruning? [#what-is-sketch-based-row-pruning]
All pruning methods described above require special preparation during `INSERT`: splitting the data into partitions, sorting the data, or building an index structure. Sketch-based row pruning is similar to indexing, but instead of a full index, we only build small, per-tablet sketches. Unlike an index, a sketch can be built separately and cheaply for every column. A sketch contains statistical properties of the column values. Some possible properties that a sketch can contain are:
* A boolean signifying whether the column contains NULL values
* A set of all occurring values for the column.
* A [Bloom filter](https://en.wikipedia.org/wiki/Bloom_filter)
* A binned histogram
* Lower and upper bounds
Firebolt sketches only support lower and upper bounds at the moment. So, as an example:
Since sketches are created for every column, there's no need for you to choose specific columns for sketches. Now let's go back to the players table.
```sql
SELECT id
FROM players
WHERE hours > 10000 AND country = 'Laputa';
```
As mentioned earlier, the players we are looking for registered in the game during the first 34 days after it launched. Under the reasonable assumption that players are added to the database roughly at the same time they registered, it is likely that these players would be stored within a small set of tablets that was added around those 34 days. This means that for all tablets created later, the hours field will always be less than 10000, which would show in a lower-upper bounds sketch, allowing us to discard them without reading their contents.
Furthermore, if our sketches support Bloom filters (or even just a set of all possible values, as there are only 195 countries in the world), we can also get rid of all the tablets that were added before our game launched in the country of Laputa.
This example shows how sketches can implicitly capture correlations between table columns and their insertion time, without these having to be specified explicitly by you. Furthermore, sketches can also help us with equality predicates against very rare predicates. If we store a sketch of unique country values for every 1000 rows, then since only 0.02% of players are Laputans (every 5000th player), we can avoid scanning almost 80% of rows, as most blocks of 1000 rows won't contain a Laputan and can be pruned.
## A real life use case [#a-real-life-use-case]
One of the tables we commonly use in internal Firebolt processes is called `query_history_f.` This table, along with several others, keeps a history of the queries that ran on our Firebolt cloud. Let's say that we are interested in getting information about queries that ran on a specific version of Firebolt that had some unique behavior. In this case, we would use the `query_history_f.packdb_version` field to filter on these queries. For simplicity's sake, let's say we only want to count them, which means that our query would be the following:
```sql
SELECT count(*)
FROM query_history_f
WHERE packdb_version = '4.23.0';
```
Below, we can see what happens when we run this query with and without using sketch based row pruning.
Without sketch-based row pruning:

With sketch-based row pruning:
Let us now also take a look at the `EXPLAIN (ANALYZE)` output for the query with and without sketch-based row pruning. Both `EXPLAIN` outputs have the `read_tablets` table function, and we can see that without sketch-based row pruning, it had to read 2,126,392,391 rows, and with sketch-based row pruning it only read 522,026,101 rows. If we look at the `list_tablets` table function, we can similarly see that both versions read a nearly identical number of tablets (\~1110). The difference between both versions happens between `list_tablets` and `read_tablets`. In the version without sketch-based row pruning, they appear on top of each other, while in the version with sketch-based row pruning, there's a `[Filter]` between them which cuts the \~1110 tablets into only 92. In the next two sections, we're going to explain why we choose this specific representation for sketch-based row pruning.
Without sketch-based row pruning:
```shell
[0] [Projection] ref_0
| [RowType]: bigint not null
| [Execution Metrics]: Optimized out
\_[1] [Aggregate] GroupBy: [] Aggregates: [count(*)]
| [RowType]: bigint not null
| [Execution Metrics]: output cardinality = 1, thread time = 1ms, cpu time = 1ms
\_[2] [Projection]
| [RowType]:
| [Execution Metrics]: output cardinality = 5512650, thread time = 0ms, cpu time = 0ms
\_[3] [Filter] (ref_0 = '4.23.0')
| [RowType]: text null
| [Execution Metrics]: output cardinality = 5512650, thread time = 6236ms, cpu time = 6175ms
\_[4] [TableFuncScan] $0.packdb_version
| $0 = read_tablets(table_name => query_history_f, ref_0)
| [RowType]: text null
| [Execution Metrics]: output cardinality = 2126392391, thread time = 29451ms, cpu time = 17721ms
\_[5] [TableFuncScan] $0.tablet
| $0 = list_tablets(table_name => query_history_f)
| [RowType]: tablet not null
| [Execution Metrics]: output cardinality = 1113, thread time = 8ms, cpu time = 7ms
\_[6] [Projection]
| [RowType]:
| [Execution Metrics]: output cardinality = 0, thread time = 0ms, cpu time = 0ms
\_[7] [SystemOneTable]
[RowType]: integer not null
[Execution Metrics]: Nothing was executed
```
With sketch-based row pruning:
```shell
[0] [Projection] ref_0
| [RowType]: bigint not null
| [Execution Metrics]: Optimized out
\_[1] [Aggregate] GroupBy: [] Aggregates: [count(*)]
| [RowType]: bigint not null
| [Execution Metrics]: output cardinality = 1, thread time = 1ms, cpu time = 1ms
\_[2] [Projection]
| [RowType]:
| [Execution Metrics]: output cardinality = 5512650, thread time = 0ms, cpu time = 0ms
\_[3] [Filter] (ref_0 = '4.23.0')
| [RowType]: text null
| [Execution Metrics]: output cardinality = 5512650, thread time = 1256ms, cpu time = 1240ms
\_[4] [TableFuncScan] $0.packdb_version
| $0 = read_tablets(table_name => query_history_f, ref_0)
| [RowType]: text null
| [Execution Metrics]: output cardinality = 522026101, thread time = 5891ms, cpu time = 4296ms
\_[5] [Projection] ref_0
| [RowType]: tablet not null
| [Execution Metrics]: Optimized out
\_[6] [Filter] (((ref_1 > '4.23.0') or (ref_2 < '4.23.0')) IS DISTINCT FROM TRUE)
| [RowType]: tablet not null, text null, text null
| [Execution Metrics]: output cardinality = 92, thread time = 42ms, cpu time = 39ms
\_[7] [TableFuncScan] $0.tablet, $0.min_packdb_version, $0.max_packdb_version
| $0 = list_tablets(table_name => query_history_f)
| [RowType]: tablet not null, text null, text null
| [Execution Metrics]: output cardinality = 1110, thread time = 47ms, cpu time = 14ms
\_[8] [Projection]
| [RowType]:
| [Execution Metrics]: output cardinality = 0, thread time = 0ms, cpu time = 0ms
\_[9] [SystemOneTable]
[RowType]: integer not null
[Execution Metrics]: Nothing was executed
```
## The classical representation of sketch-based row pruning [#the-classical-representation-of-sketch-based-row-pruning]
These two last sections will be dedicated to explaining how we represent sketch based row pruning in our system. Let's take the original example of our players table, but simplify the query to also include non-Laputans. The new, simpler query is now:
```sql
SELECT id
FROM players
WHERE hours > 10000;
```
What this query needs to do is, for every tablet, look at its bounds and determine whether its upper bound is less than or equal to 10000. If it is, then we can skip that tablet when running the query. Now that we know how we want to execute the query, one question remains: how can we express this behavior in the query plan?
A naive query plan without sketch-based row pruning would look roughly like this (sadly we can't bring an actual Firebolt example here, as Firebolt no longer uses this representation):
```shell
Project id
Filter hours > 10000
ReadFromStorage players.hours, players.id
```
Query plans are trees executed from bottom-to-top, meaning that this query plan represents the following steps: read the columns from storage, then filter them, then return the `id` to the user. Now, let's represent our sketch based row pruning in the most obvious way:
```shell
Project f2
Filter f1 > 5
ReadFromStorage t.f1, t.f2, WithPruning t.f1 > 5
```
While this looks very easy to read, it also has some problems.
* First of all, we have made the semantics of `ReadFromStorage` more complicated since it can now accept an optional `WithPruning` clause.
* More than that, `WithPruning` accepts an expression. Is any arbitrary boolean returning expression allowed in `WithPruning`? If that's not the case, then we need to define the exact semantics of what expressions `WithPruning` accepts.
* `WithPruning` hides the actual runtime behavior from us. `Filter` or `Project` or `ReadFromStorage` can be easily explained, but what does `WithPruning` actually cause the runtime to do? The runtime has to interpret the expression and then decide how to use the sketch to prune, creating a "planner-within-runtime" situation.
* If we ever support reading from different types of storage, or from table functions, do we also want these to support `WithPruning`? In that case, we would have to ensure each of these table source nodes allows such a clause.
The last point is especially important for us, as we want to support reading from different data sources and not just from our managed tables. Firebolt also supports reading from an Iceberg lakehouse, where we'd like all our optimizations to work just as well. The good news is that Iceberg can also provide a lower-upper-bound sketch for its tablets. But does this mean that we should reimplement `WithPruning` for a `ReadFromIceberg` node? Should we have a hidden layer within Firebolt code that generalizes `ReadFromStorage` and `ReadFromIceberg` so that both can handle `WithPruning` the same?
## Composable pruning [#composable-pruning]
What we really want is to be able to express `WithPruning` as its own node in the query plan, to decouple it from ReadFromStorage. This would make it less opaque, as well as allow it to easily compose together with `ReadFromIceberg` or any other `ReadFromXXX`. Such a composable approach is described in the paper ["Big Metadata: When Metadata is Big Data."](https://vldb.org/pvldb/vol14/p3083-edara.pdf) The main idea here is to separate `ReadFromStorage` into two primitives: `ListTablets`, `ReadTablets`. `ReadTablets` will not be a table source node, as that role will be reserved for `ListTablets`. The old `ReadFromStorage` can now be replaced with a `ReadTablets` on top of a `ListTablets`. Here is an `EXPLAIN (PHYSICAL)` for such a query (note that there's still no pruning here):
```shell
[0] [Projection] players.id
\_[1] [Filter] (players.hours > 10000)
\_[2] [TableFuncScan] players.id: $0.id, players.hours: $0.hours
| $0 = read_tablets(table_name => players, tablet)
| [Types]: players.id: bigint null, players.hours: integer null
\_[3] [TableFuncScan] tablet: $0.tablet
| $0 = list_tablets(table_name => players)
| [Types]: tablet: tablet not null
\_[4] [Projection]
\_[5] [SystemOneTable]
[Types]: $0: integer not null
```
Note that in our implementation, `ReadTablets` and `ListTablets` have become table-valued functions (and are also named using `snake_case` instead). `list_tablets` returns a "metatable" containing the list of tablet "paths" for the `players` table and metadata for each tablet. `read_tablets` can no longer decide what it reads on its own, so it has to instead accept input from `list_tablets` which contains the "paths" for the tablets it will read. By using this representation, we have managed to formalize the separation between "read from storage" and "decide what to read" into the query plan itself.
The next step would be to express the pruning of blocks using metadata. Since `list_tablets` returns a list of tablet "paths" and their metadata, we can use said metadata, which would include the sketch, to filter out only the tablets that we want. Doing this will get us the following query plan, where in addition to the obvious `[Filter]` on hours, we also have a `[Filter]` on tablet metadata that runs before we even start reading the tablets:
```shell
[0] [Projection] players.id
\_[1] [Filter] (players.hours > 10000)
\_[2] [TableFuncScan] players.id: $0.id, players.hours: $0.hours
| $0 = read_tablets(table_name => players, tablet)
| [Types]: players.id: bigint null, players.hours: integer null
\_[3] [Projection] tablet
\_[4] [Filter] ((max_hours <= 10000) IS DISTINCT FROM TRUE)
\_[5] [TableFuncScan] tablet: $0.tablet, max_hours: $0.max_hours
| $0 = list_tablets(table_name => players)
| [Types]: tablet: tablet not null, max_hours: integer null
\_[6] [Projection]
\_[7] [SystemOneTable]
[Types]: $0: integer not null
```
We no longer have a special `WithPruning` clause or even a new node representing pruning. Instead, the pruning operation is simply a standard SQL filter done on standard SQL fields, which just happen to be the same fields we will later use to read `players`. Note that `IS DISTINCT FROM TRUE` is used in order to allow a sketch to be `NULL` for backward compatibility with storage where it was not yet calculated.
Finally, as mentioned earlier, using this we can also easily represent such pruning on Iceberg. For example, the query
```javascript
SELECT l_linenumber
FROM read_iceberg('s3://firebolt-core-us-east-1/test_data/tpch/iceberg/tpch.db/lineitem')
WHERE l_receiptdate > '2001-01-01';
```
gets the following `EXPLAIN (PHYSICAL)` (where some field names have been replaced with ... to make the plan more easily readable):
```shell
[0] [Projection] l_linenumber
\_[1] [Filter] (l_receiptdate > DATE '2001-01-01')
\_[2] [TableFuncScan] l_linenumber: $0.l_linenumber, l_receiptdate: $0.l_receiptdate
| $0 = read_from_s3(url='s3://firebolt-core-us-east-1/test_data/tpch/iceberg/tpch.db/', format='PARQUET', object_pattern='*', type=Iceberg, file_format, file_size, ...)
| [Types]: l_linenumber: bigint null, l_receiptdate: date null
\_[3] [Filter] ((max_l_receiptdate <= DATE '2001-01-01') IS DISTINCT FROM TRUE)
\_[4] [MaybeCache]
\_[5] [TableFuncScan] ...
| $0 = list_iceberg_files(url => 's3://firebolt-core-us-east-1/test_data/tpch/iceberg/tpch.db/lineitem/metadata/00001-8e0eaefb-ab69-4db0-99ea-8fe78f974ab7.metadata.json', metadata_json_content => '****', snapshot_id => '...', snapshot_timestamp => NULL)
| [Types]: file_format: ...
\_[6] [Projection]
\_[7] [SystemOneTable]
[Types]: $0: integer not null
```
Using this approach, we've managed to add support for sketch-based row pruning on Iceberg without having to write any specialized pruning code, relying only on our already-existing pruning infrastructure and on reading Iceberg metadata.
## Conclusion [#conclusion]
In this post, we have shown the benefits and implementation of lower-upper-bound sketch-based row pruning. We can also extend what a sketch can contain and use composable pruning to allow for some more useful optimizations down the line:
* Firebolt's join pruning currently uses only the clustering index. It can be extended to take advantage of sketches as well by adding matching `Filters` on top of `ListTablets`.
* Pruning on `WHERE x > (SELECT max(y) FROM t)` becomes much easier to implement. This can be done by adding the subquery into our existing `Filter` node on top of `ListTablets`.
* Queries with a `LIMIT` on top of a `Filter` can use a `Sort` on top of `ListTablets` to sort tablets in descending order according to how many rows are estimated to survive filtering, allowing the query to finish earlier.
# Matthew Weingarten from Disney Streaming about Data Quality Best Practices (/blog/matthew-weingarten-from-disney-streaming-about-data-quality-best-practices)
Matthew Weingarten, Lead Data Engineer at Disney Streaming, talks about principles essential for data quality, cost optimization, debugging, and data modeling, as adopted by the world's leading companies.
Listen on [Spotify](https://open.spotify.com/episode/5viuHVkUHwZ9HKwDJKpSYW) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/matthew-weingarten-from-disney-streaming-about-data/id1561927688?i=1000650462145)
Transcript:
Benjamin (00:01.282): Hi, and welcome back everyone to the Data Engineering Show. It's my total pleasure to have Matt Weingarten on the show today. He's a lead data engineer at Disney Streaming right now, based out of Seattle. Yeah, spent quite a bit of time in industry already. Was at Facebook before, did his master's at the University of Florida. Really great to have you on the show today, Matt. Um, so yeah, do you quickly want to introduce yourself?
Matt Weingarten (00:32.616): Yeah, well, thank you for having me, of course. So yeah, Matt Weingarten here. I've been in the data engineering space for roughly seven years. I technically started as a software engineer, like probably a lot of people in the data world, and then kind of just learned the tools of the trade of data engineering from there, because a lot of the things I was working on were data-centric. And so I felt that's what best applied to me, and I've stayed there ever since, and I really enjoyed it. So I think the data space as a whole, it's really exciting, what's going on in the last few years, you know, seeing how it's evolved from, you know, when I started, which even was a more evolved form from a course a few years before that when, you know, data engineering kind of really started to take off. So it's a great time to be in this space and I'm glad to be on this program.
Benjamin (01:18.326): Nice. Yeah. Uh, that's definitely super exciting. Uh, and I couldn't, I couldn't agree more on your comments on the data space. Um, so what's kind of on your mind these days, right? So kind of, yeah, like what, what are the big topics basically that you're thinking about at the moment in terms of data?
Matt Weingarten (01:38.592): Yeah, so to give some more context, I work on a team that's working with clickstream analytics within Disney. So on an average day, we're processing around a billion records, representing different user actions within the Disney streaming ecosystem. So when you're logging into an application like Disney Plus or working with any of the different ESPN applications, we're tracking all of that data.
Benjamin (01:53.206): Wow, that's a lot.
Matt Weingarten (02:06.596): And so it is a true big data application that we have in place due with the scale of what we're handling. So there's various challenges involved with that, although I think we've gotten it into a pretty good form at this point. So really, what we're focusing on is just business projects, anything that our data can leverage in order to help contribute to business. That's a big aspect of that. But then really, where I pay some attention, is just trying to see how we can do various things better. You know, for example, like one, I think that everybody's kind of trying to approach right now is data quality. Lots of data, of course, can be a cause of a lot of issues. So we wanna be able to see how we can kind of really, you know, look more into what we're actually producing and making sure that what we have in place is like of the highest quality. So that along with some of the other standard challenges, I feel is like where we spend most of our time these days.
Benjamin (03:06.614): So tell me a bit more about that, right? In terms of kind of data quality, cause this is something that comes up very consistently in this show and there's so many aspects to this, right? Proper testing, having a staging environment, writing like unit tests for data, having data observability tools. What's your take on this space, right? What have you seen really provides a lot of value? Have you seen something that maybe doesn't work so well? Give us your thoughts on data quality.
Matt Weingarten (03:34.976): Yeah, so it's kind of interesting, but kind of like a lot of the principles that I like within data quality, I actually got from Facebook, well, what's now Meta, during the time I was there. So one thing that they did really well, and with a lot of the pipelines they had, what you would first do is you would write data to kind of like an intermediate table, and then you would run various DQ checks on top of that data. Once all those DQ checks passed, data was pushed to its final location. So essentially, you weren't pushing data to its final location. Yeah. You had an intermediate layer to basically be like, all right, are we sure everything's good here? And I think that's a great practice. I wish that would, and I feel like that's becoming more standard, but I wish that was like the standard because what I've seen in a lot of places, for the places that even have data quality checks implemented.
Benjamin (04:09.93): Like a bit like write, audit, publish. Okay.
Matt Weingarten (04:32.088): But they'll be doing it on top of like the final layer of the data itself, which is good that you have it at least, but you want to make sure that you have a single source of truth. And if you're running those checks on top of that and the data turns out to be wrong, well, now you have to go back and fix that. You know, teams who are running reports off of that data already have to go back and reprocess their things. I like to keep the final layer as clean as possible, or at least that's my philosophy on it. So... I think that's one good practice that I've seen. And to go back to some of the other points, yes, of course, having lower environments where you can do testing, building in unit tests, which I think is something we're gonna actually start to really try to focus on a little more, at least within my team. I think some of the projects that we've worked on in the last few months have kind of revealed the need to have some more of that in place. So we've realized that we need to have a good practice there because... If we don't have that proper testing in place beforehand, we go and deploy something, we start running, other people point us out to issues, and then we have to start all over again. We have to make the fixes, we have to backfill data, which is an expensive exercise. We wanna make sure we're doing things as cleanly as possible. So I think we really wanna take some lessons there and make it a stronger practice for us.
Benjamin (05:50.094): Gotcha. In terms of this auditing step, right? Of like having an intermediate stage where you put your data before you actually then publish it to your downstream consumers. What are the big things there, right? Is okay, there's some easy stuff you can do. Looking at the distribution of certain values, making sure, hey, are all of the timestamps I have in there actually from the last couple of days? Do we have certain correlations, certain distinct counts, all of that stuff? Anything fancier? Kind of you're doing or kind of you advocate doing. Um, when you then look at whether the data you have in that like staging area basically meets your, your quality bars.
Matt Weingarten (06:30.824): Yeah, so I mean, this is something we have kind of just started approaching. I have a lot of thoughts in this space, but really it's going to be kind of like a continual approach, or an incremental approach rather, to get to that, you know, just taking it one step at a time. So I think first, of course, is defining, you know, the core checks that you would expect with data. Like, hey, this is a primary key, are there any duplicates? Because if there's duplicates, there's an issue. Like, we know that right away. So having that... You know, checking for nulls where they shouldn't apply. Some of those just like, you know, basic checks we should have in place. Then we want to have checks against like our, you know, core metrics, our KPIs, making sure those all look good. And then we kind of want to dive into the area of trend checks. So like, you know, a lot of our data, what we see is a lot of seasonality involved with it, especially because we're working with sports data. We know exactly when, you know, we're going to see spikes in our data. Like for example, the Super Bowl is this Sunday. It's one of the biggest sporting events in the world. So we know we're gonna see a spike this Sunday of traffic compared to other weeks. You know, during the week, things are usually pretty quiet. During the weekend when a lot of the sporting events happen, that's when we see those jumps. So we wanna make sure we're accounting for that. Like over a month or over a week or whatever time period makes sense, we wanna make sure that, hey, we've seen a drop of like 20% of logged in sessions.
Benjamin (07:26.875): Just coming up with a timely
Matt Weingarten (07:51.38): You know, if that might be normal behavior, if we were expecting that, but of course that can totally reveal something different. So we really wanna make sure we have that in place. And then I think another point about data quality that sometimes isn't necessarily discussed enough, but one thing with this is we wanna make sure that, you know, our approach to data quality is a conversation between us and our business stakeholders as well. You know, we have an idea of what our data quality checks should look like, but our stakeholders might have some other ideas based on what they think is critical. So we should take that into account as well, because it shouldn't just be something that applies to us. It's their data at the end of the day as well. So we want to make sure that we kind of go in both direction in terms of how we're incorporating that.
Benjamin (08:32.398): That makes perfect sense. So one thing kind of coming myself from a software engineering background, right? That I'm seeing is okay. Kind of writing tests, checking for the data integrity. Like that makes perfect sense. The next step then is you have something failing, right? Okay. Like you have a test in your CI, uh, that doesn't succeed as a software engineer, or you have some data quality checks kind of that, that fail in your, before you publish that data. Um. How do you approach debugging that? Right? Cause like the actual like root cause for why your data isn't in a good shape might be very far kind of upstream in your data pipeline. And what, what are your thoughts on that?
Matt Weingarten (09:11.884): Yeah, we've definitely had to do a fair amount of that. Sometimes we have to go all the way up to the source of our data and then actually check there to see what's going on. It could just be some things like bad data there or some issue with timestamps, which can come into play all the time. It seems like it's always timestamps, it feels like, more often than not, in some way, shape, or form. So we just have to, you know.
Benjamin (09:33.175): Yeah.
Matt Weingarten (09:37.216): We definitely make sure we do the proper analysis. And of course, I think part of that is making sure that we're enabled to do that analysis quickly, because sometimes it can become a process to really try to actually, you know, uh, dive down into that data, but we want to make sure that we can speed that up so that it's easier to do that because yes, there's a lot of comparisons that we need to do when it comes to that. So we want to make sure everybody's enabled to do that, uh, properly.
Benjamin (10:01.422): Are there any tools at the moment that you're particular excited by in this space? So one company we hadn't shown, for example, in the past was Monte Carlo, who were doing a lot of like kind of data observability and lineage and these types of things. This feels like such a quickly changing space, right? And kind of like there's new tools coming to the market all the time. Anything that maybe you've worked in the past doesn't have to be at your current employer that you find really exciting.
Matt Weingarten (10:29.28): Yeah, so I think Monte Carlo is great. I've been really impressed by Barr as a thought leader in the data space, especially. And I know that we have kind of looked at their product in the past. But the thing with, and this is one thing I always say with companies like the size of Disney, is that what we do in one area can be completely different than what's going on in another place. It's huge.
Benjamin (10:38.679): Definitely.
Matt Weingarten (10:59.324): It's kind of hard to get that full landscape. So, Monte Carlo has been one. I know that we've also just tried smaller things like Deque, the library that Amazon, I think, built a long time ago. Great Expectations is one that a lot of teams use as well. So, we've looked at a few of those things. I think as we start to turn more towards both having a quality, observability, monitoring, all those aspects, you'll need some type of platform for that. So whether it's built in-house or whether it's using one of those tools, I guess that remains to be seen. But yeah, I've been really impressed by some of the developments that have been going on in that area.
Benjamin (11:43.346): That's awesome. Um, cool. So one other thing that also always comes up when, when we talk to people on the show is cost, right? So data quality, of course, is on everyone's mind. We're serving the business in the end and to make good business decisions, we need to make sure we look at the right data. And the other part is we need to provide that insight in a cost-effective way. Um, what, what are, what's your, like, what are your thoughts on that basically?
Matt Weingarten (12:12.12): Well, I think it's kind of funny that you asked that because that's something I'm very passionate about actually is that whole area. And I think a lot of companies have been over the last year or so if not before then, because you know, 2023, if you were going to kind of summarize that, you know, there was layoffs that were happening almost everywhere. Even right now, they're still happening in a few places. But you know, 2023 was really when a lot of that took hold. And we kind of realized that, you know, especially during the beginning of the pandemic, you know, stock market was in a great place. Everybody was hiring like crazy. And now things have to be scaled back a little bit. And one of the things that we can look at first is data products, because data applications, and this is something you kind of overlook when you're talking about costs and optimization and that whole area, which is kind of referred to as FITOPs for those who are familiar with the term. But big data applications are one of the most expensive things usually, because you're working with big data. You set up big servers. You have a lot of data storage. You need to make sure you're doing that in a cost-effective manner, or you'll see a very big bill at the end of the month from whatever your cloud provider is. So we've definitely done a lot of work in that space. And I would say we have, and we still have to continue to do so because it's just something you have to keep working on. You can't just work on this in one month and then say you're done. This whole thing is a continuous effort, and we've made a lot of good progress in that space through some various practices, which I'm happy to dive into more. But yeah, there's still a lot of work that needs to be done.
Benjamin (13:43.21): So let's talk a bit about practices, right? Say I'm a data engineer. My manager is coming to me saying, Hey, like we're spending 500K a year on this data pipeline. Like let's figure out how we can reduce costs here while retaining certain level of data freshness or, or whatever. Uh, like what are tools in my toolbox now as a data engineer? And where can I learn more about that?
Matt Weingarten (14:09.428): Yeah, so the way I approach it is there's two different aspects that I think are critical when you're looking at how to fine-tune these data applications. First side is the storage side. Effective almost always, you have to store data somewhere because it needs to be accessible. If you're using, for example, some file system like S3 coming from an AWS background, that's the first one that comes to mind. So with that, you can see some really expensive costs in S3 because by default, all files are being stored in standard storage, which is great because you can retrieve those things really quickly, but it's also the most expensive form of storage. Now, if I'm storing data there from five years ago that I'm only gonna refer to once maybe every few years for auditing purposes, then you don't wanna keep that in standard storage and you wanna put that in some glacier or something like that, which is much cheaper. And when you start to apply that to tens, hundreds of terabytes for your data, that cost can drop really quickly. We saw a lot of improvements just from looking at that. The other side, of course, is compute. Now, compute, there's a lot of progress that's gone on this space. It just feels like every other week, you're seeing like Databricks or Snowflake or any of the internal tools within any of the cloud companies. They're bragging about how they've kind of helped in this aspect, whether it's having serverless technologies or optimizing compute from some other means. There's been a lot of great development in that space. There's a bunch of different fine points that we've kind of just followed by looking at some of their overall best practices. And we've made a lot of good progress there as well. Just making sure we're using only what we need. That's a big one because sometimes by default, you'll just like copy, clone something and it'll end up being like a 100 node cluster that you could be doing with 10 nodes. So just that type of making sure we're doing things correctly, optimized as much as possible, that's really helped us get to where we want to. Although like I said earlier, still a lot more to come in that space, but a lot of progress to this point.
Benjamin (16:06.348): Right?
Benjamin (16:25.622): Interesting. How about things like modeling, right? So tools are the one thing. Sure. We want to optimize our compute. We want to optimize our storage bill, right? These types of things. But then you also have things like, okay, on the serving side, right? Like, do I denormalize or have like a highly normalized schema to optimize for query performance, which should in the end kind of go into spend, like, do you feel like also in the community as a whole, um, people think about these cost aspects enough when they're approaching how to model certain scenarios, for example, or do you think modeling is actually not that important there?
Matt Weingarten (17:02.452): I feel like that one does get overlooked, but that's certainly a very relevant point. One thing that I think the primary consideration when it comes to modeling, and sometimes what happens first, is you try to think about your consumers, your stakeholders, how you're modeling that data for them. We often see sometimes that you try to stuff a lot of information into one table or like some super big view because... That just makes it easier for consumers to query. They're not necessarily going to be as experienced with joining a bunch of tables together, exploding, nested elements, all those things. So you try to keep it very simple. But of course, that can have implications for the cost. But I feel like that definitely does get overlooked. So I think when designing models, if you want to really do it the effective way, you've got to balance your consumers' needs.
Benjamin (17:38.316): Right.
Matt Weingarten (17:58.92): Along with what you think is the proper architecture. Because of course, yes, you could go for some super normalized form, which if you were in a database classes in college, it would mark off all the checkboxes. You would get an A+. But that's not always the most business effective solution. So you kind of have to weigh those considerations and see what makes the most sense for everything.
Benjamin (18:15.423): Yes.
Benjamin (18:21.326): Nice. Awesome. One thing in these cost debates that then often comes up is right. This idea of data ROI, because it might be perfectly fine for a certain pipeline to cost a whole lot of money. If it's providing that value again, with a certain markup, of course, kind of 40 organization, um, do you have any thoughts on that? Right. It's like, cause as a data engine, data engineering team, you should be able to kind of say, hey, our data pipelines are kind of the dashboards we're providing the analytics we're providing and so on is contributing at the end of the day to the bottom line of the company. Um, what's your take here?
Matt Weingarten (19:01.725): Yeah. Yeah, I think we do need to get to a better definition on that. And I know that's something we're trying to enable. One thing that I feel like has been a struggle, and this is a primary question that sometimes we don't even have all the answers to, is who is using our data? We have this available, and we know some of our stakeholders, but do we know about every single one of them? So you can't really go to the data nail down an ROI unless you can trace that back to everybody who's using your data for what purposes and what areas they were contributing to. Our team, we're more of like a middle man, or a middle person rather for data within the company in that we provide this for other teams and then they're using it for a variety of analytics related use cases. But we don't necessarily see to it too much beyond that. So we can't really get a reasonable calculation on what type of value and ROI our data provides without seeing more details along that. Obviously there's a lot that comes out of it, but we kind of really need to draw that all together through proper lineage and kind of just like proper ownership to really figure out like what our value is overall to the company.
Benjamin (20:13.79): Yeah. My feeling is that in many cases, or at least in, well, many is maybe too much, but in some cases in software, that's easier, right? It's like you're a company, you have a prospect, you want to sell your software to, and then there's okay, like here's like three features that are missing to unlock this much revenue, right? And then it becomes very clear of, Hey, okay. Do we deploy the engineering resources to be able to then win that deal in the end for data, which is mostly about improving decisions, right?
Matt Weingarten (20:25.654): Yeah.
Benjamin (20:40.654): Like this lineage aspect just becomes so much harder because a lot of the decisions in the company that in the end you're powering, uh, like you can't directly trace it back to, Oh, this person like looked at that dashboard, uh, during that day and then made it like decision, which saved the company. I don't know, like $10 million or, or whatever.
Matt Weingarten (21:02.996): Yeah, yeah, no, I definitely agree. Like, of course, if you're working at, like, your typical SaaS company where you're shipping something to clients, then of course, you can really trace that pretty easily. But when you're working within a bigger organization and you serve different teams from the data scan point, it's a little harder to put all those details together. So that's why you really need to have that type of collaboration in place with stakeholders, which we've certainly done a much better job of over the last few years, I feel. But yeah, there still needs to be some work to really get to a proper number of, you know, what is the ROI of our data? Because I'd be curious to know that ourselves. I couldn't give you that answer if you asked me.
Benjamin (21:39.618): Yeah. So we talked a lot about kind of tools now, right? Testing data quality, kind of ROI on data, data cost, all of those things. What do you feel is something that's missing in the ecosystem? Right? Kind of, if you started a company today to work on some tool to make your life as a data engineer easier. Do you feel like there's something where there's a lot of opportunity or kind of not a very mature tool yet?
Matt Weingarten (22:09.3): Um, I feel like we've made a lot of great progress in this space within the last few years. Um, You know one one, uh area that I would think of that first came to mind and then my questions were kind of answered There is the um is the area of cost actually So, you know if you're a data engineer, of course You can always think of you know the days when you would have to you know, debug some, you know slowly performing pipeline um where you'd have to look into the spark plan, look at all the details, is there any data skew or spill, ask yourself all those questions, and then you would have to put together some approach and then just kind of see what the overall improvement of that is. Then when it comes to costs, you kind of have to do that debugging as well and see if you can optimize your pipelines in any way. But, and this was kind of something that caught me by surprise at the beginning of, I think 2022 was when I first heard about it. But there's this company by the name of Sync Computing. I'm not sure if you've ever heard of them before or not, but one thing that they do, sync, S-Y-N-T. Yeah, well, not like sync the engine appliance, but like sync the syncing process. Right. So, yeah, so, you know, I first heard about this tool, I think actually through Reddit,
Benjamin (23:13.034): What was the first word? Like what computing? Sync, like data sync. Okay. Haven't heard of the... Yes, data sync like the data. Yeah.
Matt Weingarten (23:34.476): I use Reddit mainly for work purposes at this point. I'm a loser like that. So, I saw about this tool, but essentially, they could actually look at the logs for your EMR application or your Databricks job, and they could automatically recommend an optimized configuration for it. That didn't really exist before. This was a lot of fine tuning you would have to do. And then of course, you might see some costs that made sense and then you would just leave it be.
Benjamin (23:50.114): Nice.
Matt Weingarten (24:04.128): But of course, perhaps it could even be better. So to have something like that can actually apply, you know, like, you know, that proper modeling and artificial intelligence that didn't really exist in that area to kind of tell you, you know, how you could be doing things better. I think that's a great use case for data engineers because that was one of the ones that, you know, we had kind of struggled on in the past. Yes, data quality and a lot of those other aspects you can have, tools already existed to answer those questions or it wasn't as difficult to put those tools together. But especially with cost becoming such a big concern for companies in the last few years, it's great to see that there's some emergence in the space of those sorts of tools. So I guess to kind of answer your question, I'm not sure, I mean, yes, of course, if you really think about it, I'm sure there's a lot of unanswered questions in the data world, but it just feels like whenever you have an unanswered question, you just wait a few weeks and then you find out, oh, this thing is out there or it's coming out. That's how fast this space evolves. So... I think you just kind of like take it week by week and evolve with the times. What you're thinking of, what you think are big things at the beginning of the year are going to be completely different at the end of the year. I think going into 2023, for example, did we really know what to expect with chat GPT, generative AI, LLMs? I hadn't heard of any of those terms before, and now they're everywhere. And all the companies are going to probably emerge into that space this year if they haven't already.
Benjamin (25:22.068): Right.
Matt Weingarten (25:32.224): Yeah, I think it's really exciting. And that's why I said being a part of the data world is so interesting right now. So I guess we'll see what's in store.
Benjamin (25:42.73): Yeah, definitely. I think you mentioning these like cost optimization tools. That's, that's super interesting. Sync, which you mentioned seems to be mainly focused on Databricks. I know that there's also Kibo on Snowflake, for example, who offer these types of things as well. Uh, one thing I'm always curious about is like,
Matt Weingarten (25:55.161): Yeah, yeah, well... Yes. And I think it's great to have all those aspects covered, of course, because there are companies who go with Databricks, there's companies who go with Snowflake, there's companies who even use both, just depending on what works best in various use cases. So having those tools in place to kind of figure out how you can be doing things smarter is great because Databricks and Snowflake are great tools, but if you don't know how to use them properly, you're going to get hit with a very big bill. They're very good about that. Just as any... You know, cloud application would be, because that's how it's designed. So, got to make sure you're doing it the right way.
Benjamin (26:33.314): Yeah, definitely. Um, awesome. Cool, Matt. So all of this was awesome. Hey, I learned a lot about how you think about data engineering. I think all of the things you mentioned around data quality cost are on everyone's mind at the moment, right? Like everyone is struggling with these things, working towards solutions, kind of trying to implement things, um, in their own organization there. So hearing your thoughts and how you think about that, uh, was, was super interesting. Any closing words from your end and things you wanted to mention or bring up.
Matt Weingarten (27:04.54): Uh, no, I mean, thank you. Once again, thank you for having me. Um, one thing that hopefully everybody who, you know, listens to this program knows that the data community on LinkedIn and kind of just like that whole social media form is very, very strong. Um, so there's a lot of good thought leaders in the space who put out some thoughts, you know, we've touched on some names there and some companies. But, you know, if anybody has any suggestion or any names, um, you know, or any, you know, would like to know about any of those names, you know, feel free to reach out to me, I'd be happy to connect you to the right people. Cause. I've definitely spent a good portion of like the last two years really trying to expand my network. And that's how I end up here and in various venues like this. So we're all a very tight-knit community and it's great to see the evolution that this space has had during that time. But yeah, thank you, Ben, for having me. Thank you to the Data Engineering Show. It was great to be on this program today.
Benjamin (27:54.41): Awesome, it was great having you Matt and see you around!
Matt Weingarten (27:57.301): All right, thank you so much.
# Megan Lieu on powerful notebooks that enable collaboration (/blog/megan-lieu-on-powerful-notebooks-that-enable-collaboration)
There are two types of data influencers on LinkedIn:
1. Those who talk directly about the products and companies they work for
2. Those that provide more general guidance, tips and opinions
Can influencers actually be passionate about the products they're developing and straightforwardly talk about them without sounding salesly?
Megan is one of those influencers that combine the two approaches, and with almost 100K followers, her content seems to be resonating with many data folks. She talked to the bros about her approach to data advocacy as well as the power of notebooks, especially when they become broader and enable collaboration.
Listen on [Spotify](https://open.spotify.com/episode/4I6tmMEQnBJVjM1LZG6Ptb) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/megan-lieu-on-powerful-notebooks-that-enable-collaboration/id1561927688?i=1000640237748)
Data bros (00:00.946) All right, perfect. Hi everyone, and welcome back to another episode of the Data Engineering Show. Great to have you with us. Megan, great to have you as our guest today. Megan is a data advocate at Deep Note. We'll learn a lot about Deep Note, I'm sure. She was a data scientist before, kind of then moved into data advocacy, which is awesome. I actually never met a data advocate before, so it's going to be exciting to hear from you today what that's all about. Do you quickly wanna intro yourself, Megan?
Megan (00:30.047) Yeah, great. I am so excited to be here with you guys. My name is Megan, as you guys said. I got started in the data industry, I guess, as a data analyst. After that, moved to being a data scientist and now am a data advocate or another name would be like a developer advocate or DevRel. And so, yeah, I've held a couple different roles in my short time in data and I'm very passionate about talking about my learnings those different roles and also just soaking up knowledge from experts in the field because I acknowledge that I am still very new to this wild and wonderful world of data. So having conversations like these and listening in on podcasts is one of my favorite ways to learn from other professionals in the field. So very excited that I have the opportunity to like sort of pay back the learnings that gotten from podcasts.
Data bros (01:32.234) Sounds awesome. So take us through that journey, right? Like data analyst, to data scientist, to data advocate now. Like what got you into data in the first place?
Megan (01:42.711) Yeah, after I graduated from school, I was, I had studied finance. I had held a couple of internships in finance, and so I jumped into the finance world. I was working at a Big Four consulting firm as a financial analyst, doing mergers and acquisition, valuation and advisory. Sounds fancy, but realized two years in that was not the path for me. And so had to go back to the drawing board What's the... the many years of my future as a working adult would hold for me. And so I thought back to a couple of courses I took in college related to data analytics and data science. Um, at the time, I think those courses were still, uh, very much in the infancy of like a lot of the data science, um, graduate programs that you see offered at big universities now. So it wasn't anything super rigorous or super hardcore, but it was one of the very few, um, skills or courses that I had taken in my time at the University of Virginia. So I was like, you know, I was like, I don't have any other technical skills that I'm interested in developing really. So thought back to those courses and I was like, I guess I know stuff about data. In reality, I was like, I took one SQL course and I thought I knew what it took to become a data professional. But... Regardless, dove headfirst into that and haven't looked back since. I think, well, actually I think in the process looking backward and like questioning whether moving into data was the right move or not would have been even scarier than just like...
Megan (03:26.883) blindly going in and just being like, yeah, like this seems cool. Data science, AI is like what everybody's talking about these days. So like for all intents and purposes, why would I look backward? Right. But I'm glad I didn't. And I'm glad I've kind of forged the path that I have had in data, which is like I've held a lot of different roles in the space and so have been exposed and worked with a lot of different types of personas.
that you normally would when you are working with data and learned a lot in the process and am still very much learning.
Data bros (04:02.594) Nice. That's awesome. And you're also huge in terms of just like thought leadership, right? Like you have almost 100,000, I checked earlier, like followers on LinkedIn. I think you'll crack the like six figures there soon, kind of where we're rooting for you.
Megan (04:15.443) It's literally all I want right now. It's like, ah, end of the year, please. But it's not, it's not going to happen by the end of the year. That's okay. It's okay.
Data bros (04:24.371) What got you into that? So at some point you got deeper and deeper into data science, became an expert, you said okay, I want to share knowledge now. What happened there?
Megan (04:34.508) Yeah. I, yeah, first and foremost wanted to share some of the ups and downs of the journey. And it wasn't even like, oh, like I am now an expert and I want to impart my knowledge on people. It was more me finding some situations were funny. Some situations were very relatable. And I've always enjoyed writing. So like what better way to kind of react to those funny situations than just like put writing it down into words and releasing it. And some people were like, oh my gosh, like I experienced that too. And so over time I realized that like, I had this knack of picking up situations that happened to me that are like, I don't know, like maybe I get a feedback that I'm not like, that I'm not. doing as well in a certain area of development. And I was like, I can either take that or I can be like, well, I have some more thoughts around it. And I feel like these thoughts could be very relatable to others. And so that process of me just picking out these takeaways from my career and adding more thought around it has allowed me to relate to a lot of other people who go through those moments too. And so that's how I really started. Just like, I treated it like my daily diary. I do have a physical diary that I write into every day and then I think of my LinkedIn posts as just like a more...
Megan (06:07.959) marketing version of what I write in my daily diary. And yeah, it's just led me to this point where I don't take anything that I write in there super seriously. It's just my musings. But then with my current role as a data advocate at Deep Note, I do kind of have to take that more seriously. So it's been interesting to treat that side hobby of me writing thought leadership just on the side and turning that into a full-time job.
Data bros (06:39.146) Nice. So turning it into a full-time job, right, is like... Last time on the episode we had like Chao Kshu and she also kind of is a kind of thought leader in the space. And she was like, hey, I started writing and kind of suddenly some of my posts went viral. Right. And like, it's actually like, it was very unintentional in a sense. Like she didn't plan out like the three year journey to like become a kind of influential, like kind of thought leader. And like, how was it with you? Like, was this kind of something where in the beginning you were just writing and then some things kind of gained traction or were you like from the beginning very focused on figuring out your time? our good audience and so on.
Megan (07:16.931) Ooh, no, it was definitely not. Solid no, it was not the latter. Like, I don't think anybody, I don't know, I feel like if you go into it thinking, hey, I wanna become an influencer, hey, I wanna get all these followers, it's just not.
Data bros (07:18.358) Hell no!
Megan (07:32.575) gonna work out because you're going to start writing in a way that's not authentic to yourself. You're gonna write in a way that you see other influencers or thought leaders have already been doing and they've already perfected their craft. And so if you were to try to inject yourself into that voice, people can really easily pick up that is not how you normally talk. And also when you constantly write in a way that is not... true or authentic to you, it's just gonna make it so much harder to like consistently be consistent with your content. If you are writing about things that you truly enjoy, that's when it flows and that's when people can pick up that you are doing it because you enjoy it rather than like chasing some external goal. So I started with it, like I think my first post was like, oh, I finished this like data
Data bros (08:14.734) That's good.
Megan (08:30.465) and it got like five likes or something. And then I one day wrote a post about, I don't know, like debugging code at midnight or something. And I released that post almost exactly two years ago to this date. And then I woke up the next morning and it had like thousands of likes. And I was like, oh, I didn't know that this could happen. And so...
Data bros (08:52.348) Mmm.
Megan (08:55.015) I, for a while after that, I thought that like all my posts would pop off just like that one. And so I was just setting myself up for disappointment, obviously, because that's just not how the LinkedIn algorithm works. And so for a while in the beginning of my writing journey, it was very much me chasing that feeling of hitting virality. But that's just such an unsustainable model motivator. So over time, as I became more consistent, it was like less of those huge spikes, but more like the tiny ones that were enough to keep me going. And then, um... you write enough and you like hit another one of those big spikes. But what I realized in retrospect is that like most of the followers or like most of the traction I gained is not from those huge spikes, but it's from the cumulative growth of like those smaller spikes, but you will never hit those smaller spikes unless you are actually consistent in the first place. So yeah, long story short, I don't. I don't think that anybody can really go and be like, I'm going to, I have this like picture perfect plan of what my influorship will look like. It's just not gonna turn out that way. And so I think my advice for anybody is just like, just take it day by day and like see where it takes you because it could open a lot of doors that you never expect. But if you had your sights set on like one very specific goal, those opportunities could totally pass you by.
Data bros (10:27.882) Yeah, so that makes perfect sense. And I think it's a consistent theme we had across the podcast. Like I think Zach said very similar things about how he got started and kind of, yeah, just finding your voice. That's great. So what's on your mind these days, right? So what are you using your voice for? What are the big things on your mind at the moment?
Megan (10:41.099) Yeah. Love, Zach.
Megan (10:50.435) Yeah, obviously to do my day job as a data advocate, it's a very different way of writing and creating content. From when I was doing it as a side hobby. When I was doing it as a side hobby, there was no rhyme or reason to whatever I put out on any given day. But now doing it as part of a company and being the only advocate at my company, it has to be a lot more structured because essentially the voice that I'm putting out there is the voice that will be one in the same as Deep Note's voice. And so have to be a little bit more curated there. And so what I have been working a lot more on these days specifically for Deep Note is providing value to my audience in the form of content. So not just like sharing stories, which is what I, like I share a lot of stories about like impactful career moments in my life, right? But getting the tangible value out of those stories and relating it to our product is the hardest part. And so I do that a lot via sharing projects that I'm doing within Deep Note and showing people what I can build in Deep Note. And that has the added benefit of like... me continuing to upskill in this journey, right? Like I am building in public and hoping that people get the same value that I did out of building that project and also learn about Deep Note along the way. I think building in public is so important and it's how I first got started writing on LinkedIn because I was like, hey, here's a Tableau dashboard that I shared into Tableau public like.
Megan (12:37.311) Check this out, right? And so that's how I got started becoming a data analyst. And over the years, kind of lost sight of that. But now working for a company that creates a product where people can build their projects, it's kind of a no-brainer for me to dive into that, to one, help my audience learn about the product, two, help them learn about how to build similar projects. And so.. tying all that together into my content is, it's tricky, but it's definitely something that I'm excited to dive deeper into because at the end of the day, really just wanna help my audience grow alongside me because my audience is very supportive. And so I just kinda treat them as my learning buddies along the way.
Data bros (13:28.962) Nice. Yeah. So it also feels like there's some tension there, though, right? So you started out not being a data advocate for Deep Note. How now actually also advocating for a product, a specific one, has changed also. how you approach kind of your like role then as a thought leader because I'm sure there's a lot of people who say, oh, like another post about deep note, whatever, like, are you trying to keep these things separate? Kind of does it mix together well? Like what are your thoughts there?
Megan (14:00.439) Yeah, that is such a good question. I think about that tension. Tension is the right word here. Or like maybe balance, I don't know. That's something I think about every day. I remember once I made a post about Deep Note and somebody commented, they were like, you are selling out and is this part of your daily job to like promote? Deepnote and I was like, yes, it's literally in my job description. But the way that you do it, um, it can be very, the way that you do it has to be like on an individual by individual basis. Like I know some other developer advocates, um, in this space, like all of their content is related to their company. Um, and I, you know, I sometimes envy them. It's like, they have created a consistent, um, like they've created a reputation where they are known for their product and their audience knows what they're going to get out of that, uh, that individual's content. Um, it's going to be materials about their company. And I don't think that's a bad thing. Um, developer advocates, like their job is to educate, um, people about, or it's like to provide resources for these people to be able to use their product better. Um, and so if you are somebody who is, who, relies on the channel of like social media to distribute those resources, then that is absolutely a way to do it. The problem with me is that like I built my platform off of not doing that in the first place. And so if I were to switch into being like 100% all deep note and all my note, all my contents, my audience would be very alarmed. And it goes back to the whole authentic voice thing, right? Like they would know that is not what I used to write about. And like, let's be honest, a lot of people don't like to be.
Megan (15:56.339) sold to or like have marketing with like marketing in their face at all times. And so that tension, um, is something I've had to figure out for myself. And I, what I end up, like right now what works for me is one or two posts a week that I tie back to deep note. It doesn't necessarily have to be, um, it doesn't necessarily have to mention deep note or I have to tag them, but rather talking about the concepts. and the principles that deep note espouses. So that is first and foremost, like notebooks and how we believe that notebooks are the perfect medium for data scientists to do all of their work. And so my content can center around that without being too like deep note in your face. And so that is kind of the balance or like the answer to that tension question that I have arrived at. It's like to... What our company likes to say is like to wave the notebook flag. And like we literally have a flag in our office that says notebooks. But it doesn't say like deep note anywhere. Right. And so I really like that. At the end of the day, our company is like, not just pushing for our own specific company, but like notebooks in general. And that is something that I can easily get behind without rubbing my. my audience the wrong way, especially because I do believe in notebooks and like how powerful they can be so it's easy for me to create content around it.
Data bros (17:25.826) There's another thing, I think like, cause people think sometimes that advocacy is consulting. And I think there's something very beautiful in advocates that have strong passion on specific products. Not only that, they go and work for those products. Think about it. Most people go work somewhere. I hope they're passionate about it. I hope they pick that place because they could. And
Megan (17:33.519) Hmm. You would hope.
Data bros (17:54.422) you're going to work somewhere, you should be proud. And I think your followers should be proud of you as long as you're authentic and you are. And that whole, right, you need to manage and there's this confusion and tension, but there's also one bigger thing, which is passion. On the subject, like, and the fact that you're there allows you to be also passionate about the company you work for, it's okay. And being authentic doing that is perfect. So I don't see anything wrong here. And trying to keep a balance between everything will just make you not authentic and confuse everyone. And that's how people lose it. And yes, there are advocates that are very strong at being neutral. Like they've taken neutrality, product neutrality to a new level, which is its own niche. But I personally like interacting with people that do invest. Time and effort in specific companies, on specific products, take the time to do it. So yeah, so actually well done to be able to combine those two worlds together.
Megan (18:57.859) Thank you. Yeah. Yeah, no, I think that there are ways for me to go even deeper on that opinionated approach and it's just that I have not been at the company for long enough for me to. exactly like be the picture-perfect spokesperson for it. Yeah, and I think my company knows that. And my company appreciates that, you know, I talk about other things besides notebooks, but when I do talk about it, I hopefully am bringing a level of genuineness that others cannot. And so I think that, yeah, like you said, that authenticity is a strength that people should lean into. than shying away from it.
Data bros (19:46.898) Exactly. Awesome. So let's talk about notebooks for a bit. Right? Like, I for once, I'm not a huge notebook expert. Like, sure, I know like Jupyter notebooks, right? All of that stuff. Like, what's new in the world of notebooks, basically?
Megan (20:01.164) Yeah. What's new is precisely what we're trying to push out. It's not just notebooks, but a medium where data practitioners can come together to tackle the hardest data problems together. And so with traditional Jupyter notebooks, what you're used to is something that's hosted on your local machine where if you were to do an analysis and you were to have the environment set up for your workflows, that is all limited to your local machine. And so if you were to try to, A, replicate the environment that was required to run that analysis in your notebook, and B, also replicate the actual contents of your notebook, it's really hard because that was done in isolation. And so with Deepnote and more modern notebooks in general, it's all cloud-based now, which Jupyter , Jupyter lab so they understand that and so but We wanted to take it to another level, especially the collaboration aspect, by allowing for synchronous collaboration where multiple people can be in the same notebook at the same time. And what we have found is that the people who benefit from that are not only data scientists who have to work together, but also data scientists who have to work with people outside of their team to be able to collaborate with those citizen data scientists and business personas.
Megan (21:35.589) because what people are normally used to doing is sending screenshots or PDFs of their Python file. And nobody wants to see that stagnant document. Right, so, right, or yeah, or a snapshot of a dashboard, right? So collaborating not just within the team, but also.
Data bros (21:50.527) snapshot of a dashboard.
Megan (21:58.579) Intra team is super huge. But also what we found is another party that really benefits from this are educators and students. And not just like educators and students in like academic settings per se, but also like people who are more junior developers trying to, who need to like. code with their superiors in the same environment, or even like, um, people who are administering like live coding tests, right? There's this educational component to notebooks, um, that with notebooks, um, that kind of linear, um, format is, is really conducive to helping like develop those building blocks and knowledge. But also when you add that collaboration component, it just takes it to a whole other level. And so we have found that by bringing the collaboration aspect to notebooks, you can really elevate it and expose it to parties that normally, you know, would not have thought about using notebooks in the first place.
Data bros (23:03.522) So base Notebooks is outgrowing its initial purpose of serving data science teams in a very isolated environment, on the laptop in many cases. And now it's becoming a much broader thing and data gets involved, data warehouses, SQL gets involved, formats change. I love it. I've never used the Notebook by the way, exactly because of that. Like...
Megan (23:09.423) Sure, yeah.
Data bros (23:32.13) towards this data science thing and it runs on a laptop so...
Megan (23:34.379) Yeah. And I, you know, you, we can't blame you because notebooks were originally developed for a very niche, like mathematics, um, very like maybe mathematics, um, and like scientific, uh, personas. And so when we introduced this collaboration aspect, it was not only supporting features that allow you to like bring in other people into your workflows, but also supporting other languages and other formats of coding, like no code visualization blocks, SQL blocks, things that at the end of the day, break down the barriers so that it's not just data scientists who are working with notebooks. It's what we like to call those citizen data scientists who may not have been trained to be. Pure data scientists by education, but they have to interact with data science concepts in their day-to-day job. And those are the people that we're trying to bring into the fold. Because if we were to only focus on the super niche, siloed off data scientists, there's an upper limit to that.
Data bros (24:43.254) So how does this relate to BI, right? So you're saying, okay, like you want to get more people in the organization involved. What does it mean the person where I as a data engineer am now like building a dashboard for us, suddenly I would be building a notebook for that. Like, is that something where you would actually replace dashboards with notebooks or those are two completely kind of different things that they coexist? How do they coexist?
Megan (25:06.263) Yeah, so very great question. We have what we call apps or applications that are a layer of curation and polishing on top of the notebook. And it's supposed to be a one-to-one reflection of the contents of your notebook where you can also configure like, whether you wanna show the code or whether you only wanna show the outputs. And that app... Since it's a one-to-one reflection of the notebooks, it's going to be automatically updating with any changes that you make to the notebook, but the app feature is similar to BI tools and that it is supposed to interface or interact with those business personas who may not care about the underlying code and the inner workings of What gets them that view? But the difference between how we think of apps and dashboards is that with dashboards, there tends to be this problem where like, you spend all of this work putting all the pretty plots and charts in there only for the dashboard to be viewed once. Maybe you make a couple rounds of edits and then like that dashboard is no longer used.
Data bros (26:23.21) If someone is using that dashboard more than once, something is wrong with that dashboard. We're in the recurring theme of there has never been a dashboard with positive ROI, which keeps coming up in every podcast episode.
Megan (26:27.337) Exactly, right?
Megan (26:35.155) Okay, this sounds like you guys are onto something here. We'd love to hear your guys' thoughts on it, but like-
Data bros (26:39.875) I will take a defense here. I will defend the dashboard for a second. It's when you really need it. That's like, sometimes dashboards are there when you really need them. So you shouldn't just count how many refreshes a day you get for that word. You should count when that dashboard was meaningful. And we've had, like I've had a few occasions in my career where.
Megan (26:44.072) Okay. True. Yeah.
Data bros (27:05.494) Like without that useless specific dashboard that nobody cares about, we would be in a shitty situation, in a shitty spot. So have respect to the dashboards, but I think that as you said, the dashboards like are kind of serving themselves now and it's becoming, you know, yeah, we need a, it's like fashion. We need to put it back and move to something better.
Megan (27:09.006) Ha! You're selling it. You're really selling it.
Megan (27:19.728) Hahaha! Fashion.
Data bros (27:37.598) And if notebooks is that platform and if we can kind of communicate, present, because dashboarding is all about presenting, it's all about, like engineers that build dashboards, they spend the energy on that and they format them. It's their assets, it's their careers. It's like their project. So being able to package that in a better way that reaches more people in the company.
Megan (27:44.503) Yeah. Right.
Data bros (28:05.186) that makes it part of a bigger story.
Megan (28:08.119) Yeah. Right. And so
Data bros (28:10.395) What can I say? Highly... it's needed. Definitely.
Megan (28:13.739) Yeah, and so I mean, our answer to that is not just notebooks, but also that app feature on top. And for all intents and purposes, like some people call our apps a dashboard as well. And for all intents and purposes, they are interchangeable. But. We like to think of apps as like, rather than being a standalone thing, it is an extension of all the work that's been put into the notebook. And the app is like the cherry on top that is meant for that presentation layer. And we are putting our bets on apps being what we transform or where we go from dashboards. But I think yeah, there's definitely a place and a time for dashboards and it's probably not going to go away anytime soon. But we do start, we do have to start thinking about alternatives.
Data bros (29:08.994) So this app layer then sounds really like kind of your gateway to less technical users, right? Or less technical stakeholders in the organization. One thing that we're thinking a lot about is customer-facing analytics, right? And kind of really taking the data you have in your organization and showing that back to your customers.
Megan (29:15.359) Exactly. Yep.
Data bros (29:26.686) Is this also, I assume notebooks don't really cater to this at the moment, but is there a path there, right? Where actually like this app layer then goes to your customers and you start generating kind of really that customer value as well from these notebooks.
Megan (29:41.299) Um, sorry, can you ask that question again?
Data bros (29:44.05) Sure, so say... So I put it in a simple way. We have basically two ways to open data to outside users. Most cases like when we go from internal BI, which is, as you said, rarely refreshed and you switch to... customer-facing data apps that's frequently accessed and that's kind of the whole different story, the quality and the challenge is at a completely different level. The thing is you still serve that either as a dashboard or you have a bunch of engineers writing code and building the UI. I think Data bros's question is. Is there a future where kind of that notebook evolution can go from being internal only, right? We collaborate internally within the company across different roles to collaborate externally. So that dashboard journey goes beyond the boundaries of the company and goes into the customers. And, and again, it's all about, uh, dynamic format that fits the, fits the audience versus the other way around. And, uh,
Megan (30:53.567) Hmm. Yeah. I okay. Yes. I understand your question now. So I think there's also another component that is a limitation in bridging those two sides, which is the level of data literacy and data skills that those customers have. You're not always going to be serving your dashboards to people who know the fundamentals of data. And so that's going to be a factor that limits. That... connection or like the ability to interact one-to-one on both sides and so I think our answer that we're betting on here is AI to be able to boost the data literacy and skills of those customers who currently are not at the same level on the internal side. And so using the power of AI to enhance coding abilities or to get insights on a plain English level, that is kind of like the answer that we're going towards to not only be able to bridge the gap between internal and customers, but also people across all. skill levels on the data science spectrum.
Data bros (32:18.786) Nice, I love that. Awesome. Hey, I learned a bunch about notebooks today. This was super interesting. Awesome, Megan. It was great having you on the show. Are there any closing words from your end that you wanted to share with our audience?
Megan (32:24.217) Bye! No, I really enjoyed our conversation about the dashboards and I definitely need to tune into some more episodes to see what other people are saying because I think that's a much broader conversation and I'm sad that we didn't get to dive deeper on that. But no, I'm just really glad that we did get to talk about the topics that we did and hopefully you guys learned a couple or you guys in the audience as well learned a thing or two about notebooks.
Data bros (33:02.262) Sounds awesome. I recommend you start with the episode with Vin Vashishta. He had probably the most controversial opinions on Dashboard. So that's a good one to start. Exactly. That should be a good one. Awesome. Thank you so much, Megan. It was great having you on the show. Thank you. See you around. Take care.
# FuzzBerg: Hunting Bugs in Iceberg and file-format readers (/blog/open-sourcing-fuzzberg)
## Shifting (Fuzzing) Strategies [#shifting-fuzzing-strategies]
At Firebolt, we prioritize software security, as detailed in our [previous blog](https://www.firebolt.io/blog/fuzzing-firebolt-catching-0-days-as-fast-as-our-query-processor). We employ fuzzing at the class level for our query processing and data ingestion stack, utilizing popular fuzzers such as AFL++ and libFuzzer.
Beyond dynamic testing, our security measures include static code analysis and hardening release binaries with stack buffer overflow and Control-flow Integrity ([CFI](https://en.wikipedia.org/wiki/Control-flow_integrity)) checks, Position-Independent Executables ([PIE](https://en.wikipedia.org/wiki/Position-independent_code)), and [FULL RELRO](https://ctf101.org/binary-exploitation/relocation-read-only/) to impede runtime exploits.
Given Firebolt's rapid innovation and expanding adoption, our security strategies must constantly evolve. The launch of [Firebolt Core](https://www.firebolt.io/blog/introducing-firebolt-core) in 2025—a free downloadable query processor that can be hosted anywhere, and our [READ\_ICEBERG TVF](https://docs.firebolt.io/reference-sql/functions-reference/table-valued/read_iceberg)—which allows querying data from S3 Iceberg catalogs, represent significant advancements. These developments also required us to pivot our security-testing strategies, and this blog outlines how we addressed those needs.
## Existing Challenges [#existing-challenges]
Interfaces such as COPY FROM and Table Valued Functions are a gateway for untrusted user inputs. And since they constitute such a big part of a data warehouse's attack surface, we specifically wanted to fuzz them ahead of the critical releases.
But before we talk about **how** or **what** we did, let's understand **why** we got there.
Popular fuzzers like AFL++/libFuzzer are coverage guided, gray-box fuzzers. They instrument their target at compile time to collect in-process coverage feedback, and use that feedback to drive mutations and explore newer code paths.
**However, we found these fuzzers unsuitable for our targets due to:**
* **Simpler Interfaces**: They typically feed mutations to a target's STDIN, or write to a file that they expect the target to open/read automatically.
* **Class level harness**: These fuzzers mostly require a glue code that can call the target interface over a simple API–quite like unit tests. Even though AFL++ supports binary mode fuzzing (if source code isn't available), it still expects the target to accept inputs as mentioned earlier.
* **Blind Mutations**: They mutate inputs solely piggybacking on coverage feedback without being aware of the target's internal structure, or any knowledge of the format.
**But that's not saying we didn't look under the hood:**
AFL++ has a feature called [Dictionaries](https://github.com/AFLplusplus/AFLplusplus/tree/stable/dictionaries), that relies on a user-provided list of tokens inserted at random positions during the mutations. It then observes if or how re-positioning/changing those tokens affect the coverage, long before it spews a correct input.
Therefore, such discrete tokens still wouldn't generate a syntactically correct (yet corrupt input) **100% of the time.** The original author of AFL++ discussed how Dictionaries work and addressed the limitations in one of his [blogs](https://lcamtuf.blogspot.com/2015/01/afl-fuzz-making-up-grammar-with.html).
In recent years, however, a lot of security research have been done on how to make gray-box fuzzers structure aware using context-free grammar (thanks to [Chomsky hierarchy](https://en.wikipedia.org/wiki/Chomsky_hierarchy))
Academic projects like [AFLSmart](https://thuanpv.github.io/publications/TSE19_aflsmart.pdf) have built smart mutators that claim to preserve complex file formats. Similarly, [syzkaller](https://github.com/google/syzkaller)— a popular Linux syscall fuzzer provides its own [formal grammar](https://github.com/google/syzkaller/blob/master/docs/syscall_descriptions_syntax.md) definitions to fuzz syscalls.
But grammar fuzzers too rely on accurate data modelling. AFLSmart for e.g., uses the [Peach modeling](https://peachtech.gitlab.io/peach-fuzzer-community/v3/DataModeling.html) for [file formats](https://github.com/aflsmart/aflsmart/tree/master/input_models) it needs to fuzz. And translating complex file-formats— especially with tightly coupled, binary structures to custom grammar requires upfront investment of time and effort, with future maintenance overheads.
The closest we could resonate with was LLVM libFuzzer's [structured fuzzing approach](https://github.com/google/fuzzing/blob/master/docs/structure-aware-fuzzing.md). It exposes an interface LLVMFuzzerCustomMutator, which allows writing custom mutation logic.
For e.g., say you wanted to preserve a starting magic byte ('**H**') in a corpus of string values ("**H**i", "**H**ello", "**H**aaaai"..) and forward the remaining to libFuzzer's generic mutator from that interface.
```cpp
#include
#include
#include
void target_func(char* Data, size_t Size){
std::string str(Data, Size);
std::cout << "Post libFuzzer mutations: " << str << " \n" << std::endl;
}
// Forward declaration
extern "C" size_t LLVMFuzzerMutate(uint8_t *Data, size_t Size, size_t MaxSize);
// LLVMFuzzerCustomMutator mutates bytes in-place
extern "C" size_t LLVMFuzzerCustomMutator(uint8_t *Data, size_t Size,
size_t MaxSize, unsigned int Seed){
std::string s(reinterpret_cast(Data), Size);
std::cout << "Prior to libFuzzer mutations: " << s << " \n" << std::endl;
// We retain the first byte ('H' in this e.g.),
// and call libFuzzer's generic mutator
return LLVMFuzzerMutate(Data + 1, Size, Size + 100);
}
extern "C" int LLVMFuzzerTestOneInput(uint8_t *Data, size_t Size) {
// Additional logic can be added here,
// to check for format correctness
// before calling target API
target_func(reinterpret_cast(Data), Size);
return 0;
}
```
When libFuzzer calls LLVMFuzzerTestOneInput to pass mutated data to the fuzz target, we can additionally check if the magic byte is preserved.
Such an approach is used to fuzz [libPNG](https://github.com/google/fuzzing/blob/master/docs/structure-aware-fuzzing.md#example-png).Compiling and running the above code, we see that the mutations always start with an '**H**'.
```shell
// Compile with libFuzzer
:~$ clang++-18 libfuzzer_custom_mutate.cpp -o libfuzz -fsanitize=address,fuzzer
// Input seed
:~$ cat libcorp/*.txt
Hi
Hello
Haaaai
// Run
:~$ ./libfuzz libcorp -runs=2000 -seed=10
INFO: found LLVMFuzzerCustomMutator (0x6152e56f09c0). Disabling -len_control by default.
INFO: Running with entropic power schedule (0xFF, 100).
INFO: Seed: 10
INFO: Loaded 1 modules (20 inline 8-bit counters): 20 [0x6152e5737f00, 0x6152e5737f14),
INFO: Loaded 1 PC tables (20 PCs): 20 [0x6152e5737f18,0x6152e5738058),
INFO: 3 files found in libcorp
INFO: -max_len is not provided; libFuzzer will not generate inputs larger than 4096 bytes
.......
.......
Post libFuzzer mutations: H�EEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEE
Prior to libFuzzer mutations: H�EEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEEE
Prior to libFuzzer mutations: H�H
Post libFuzzer mutations: H�HH
Prior to libFuzzer mutations: Hw
Post libFuzzer mutations: Hw
Prior to libFuzzer mutations: H:@@
Post libFuzzer mutations: H:@@@
```
**But our challenges were still, quite different:**
* **Query based API**: Unlike traditional fuzz targets that read from STDIN or open/read a file automatically, our target reads the mutated file *only* after it receives a query over HTTP. To integrate existing fuzzers with such a pipeline would mean fiddling with their existing child process communication interfaces, such as a pipe in the case of AFL++.
* **More Structure than just "Magic Bytes"**: Unlike other binary formats like ELF or PNG, complex DWH file formats like Parquet or Avro go beyond header magic bytes.
Parquet— for instance, contains variable length Footer metadata. Corrupting this will lead to parsers rejecting most inputs, and also won't allow us to stress paths that read the pages based on this metadata (**we actually found a bug using this approach**).
And to preserve this metadata, you'd first need to read the actual size of the metadata from an adjacent 4-byte field!
Given all of the above considerations, we decided to write a custom fuzzer that does what we really need:
* *create partially-corrupt, yet structurally correct file formats*
* *write them to a path the engine can read from*
* *issue HTTP queries to trigger ingestion.*
**This approach yielded immediate values:**
* We **fuzzed the product holistically** without having to build or maintain multiple class-level harnesses, reducing future engineering overheads.
* We built the world's **first Iceberg fuzzer** (to our best knowledge).
* It **discovered 5 critical bugs** in all of our file-format TVFs– including READ\_ICEBERG!
And most importantly– we are now [**open-sourcing**](https://github.com/firebolt-db/FuzzBerg) it, for the larger Iceberg and database community to test their own reader implementations.
## Introducing FuzzBerg [#introducing-fuzzberg]

[FuzzBerg](https://github.com/firebolt-db/FuzzBerg) is a **hybrid, file-format fuzzer:** it combines structured and black-box mutations (using the classic [Radamsa](https://gitlab.com/akihe/radamsa)), to fuzz database interfaces that read complex file-formats.
A [Mersenne Twister](https://cplusplus.com/reference/random/mt19937/) seeds the Radamsa mutator, and randomises the corpus selection.
Radamsa is typically used as a CLI tool, but after a little bit of tinkering we discovered that they do support a [static library](https://gitlab.com/akihe/radamsa/-/blob/develop/Makefile?ref_type=heads#L108) mode, and incorporated that into our fuzzer's build process. It isn't structure aware, but it serves our purpose because we provide custom mutation logic (wherever necessary) on top of Radamsa.
The fuzzer is designed with modularity to allow for extensibility and ease of maintenance:
* **Custom Format Mutators**: Each file-format is fuzzed through a corresponding class, encapsulating mutation logic tailored to that format.
* **Fuzz Target**: To integrate with the fuzzer, a target implements its class that overrides two virtual methods:
* ForkTarget(): Manages the child process creation.
* Fuzz(): Calls the corresponding mutator based on the file format to be fuzzed.

This modular approach allows for the easy addition of new file formats, mutation strategies and new targets that ingest data over HTTP based queries.
And if you own the target code like us, you can increase the odds of bug detection by:
* compiling target with AddressSanitizer
* registering a SIGHANDLER in the target's entrypoint, that calls **\_\_gcov\_dump()** when it receives a SIGUSR1 from the fuzzer. We did this to analyze coverage after a fuzz campaign and write better mutation logic, but it's also a great way to visualize post-fuzzing coverage statistics.
## Alright, let's talk some file-format fuzzing now [#alright-lets-talk-some-file-format-fuzzing-now]
## Fuzzing Iceberg [#fuzzing-iceberg]
[Apache Iceberg](https://iceberg.apache.org/) is a popular open table format for efficient querying data without storing them in a data warehouse. For brevity— it comprises three Metadata layers, followed by the actual Data layer. When a specific version of data is queried, the reader parses each metadata layer sequentially to return the results from Data layer in the most efficient way:
* **Layer 1:** JSON metadata
* **Layer 2 & 3**: Avro metadata (manifest list and manifest files)
* **Layer 4**: Parquet data

Iceberg's official [spec](https://iceberg.apache.org/terms/#decoupling-using-the-rest-catalog) says every layer has some *mandatory* fields (that we need to preserve in the mutations), while some are *optional* (which we probably can remove).
Since our Iceberg seed corpus is generated on S3, we do some in-process modifications before handing it to the format fuzzer. We de-serialize the original seed into a [nlohmann JSON](https://github.com/nlohmann/json) object, and then modify the following:
* Update "**location**" field to a local Minio path
* Remove "**metadata-log**" (if exists) as we are not interested in metadata time-travel
* Remove "**snapshot-log**" (if exists) as we are not interested in snapshot time-travel
* If a "**snapshots**" list exists, we update every "**manifest-list**" field inside (to a local file that the fuzzer writes the Avro mutations to in Step 3 below).
Due to such complex inter-layer dependencies, and varying file formats at each layer, we take a three step, sequential approach to fuzzing Iceberg:
#### Step 1: Blind Metadata fuzzing [#step-1-blind-metadata-fuzzing]
We pick a random seed from the updated metadata corpus, and let Radamsa mutate it as is. At this step, no custom mutation logic is applied.
As Radamsa is not structure aware, it flips random bits and duplicates bytes, which allows us to validate our JSON parsers.
**Original Metadata (trimmed)**:
"table-uuid":"371589cb-27dd-4622-bc18-acc423d1e1be"}
**Mutated Metadata**:
"Table-uuid**z**":"371589cb-27dd-4622-bc18-acc423d1e1-bc18-acc423d1e1be"**}be"}be"}be"}be"}**
But interestingly, Radamsa isn't completely dumb either! It sometimes goes deep into the nested fields to selectively alter values (maintaining type correctness too), such as in the case below:
**Original Metadata**:
\{"current-schema-id":0,"current-snapshot-id":8715189102138002866,....,"schemas":\[\{"fields":\[\{"id":1,"name":"field\_int","required":true,"type":"int"},\{"id":2,"name":"field\_int\_null","required":false,"type":"int"}],"identifier-field-ids":\[],"schema-id":0,"type":"struct"}],"snapshots":\[\{"manifest-list":"s3://iceberg-fuzzing/metadata/manifest\_list.avro","schema-id":**0**,",..,"table-uuid":"d83c214b-5292-4c5f-827a-9d851f1a6137"}
**Mutated Metadata**:
\{"current-schema-id":0,"current-snapshot-id":8715189102138002866,....,"schemas":\[\{"fields":\[\{"id":1,"name":"field\_int","required":true,"type":"int"},\{"id":2,"name":"field\_int\_null","required":false,"type":"int"}],"identifier-field-ids":\[],"schema-id":0,"type":"struct"}],"snapshots":\[\{"manifest-list":"s3://iceberg-fuzzing/metadata/manifest\_list.avro","schema-id":**65535**,",..,"table-uuid":"d83c214b-5292-4c5f-827a-9d851f1a6137"}
#### Step 2: Structured Metadata fuzzing [#step-2-structured-metadata-fuzzing]
We re-use the same seed from Step 1, but now de-serialize the JSON to mutate every field value with Radamsa. If the field is a nested object or array, we descend further (**depth=1** for now), and randomly pick a key/index to mutate, as shown below:
**Field is an array, traversing further..**
**Field Value**: \[\{"fields":\[\{"id":1,"name":"field\_int","required":true,"type":"int"},\{"id":2,"name":"field\_int\_null","required":false,"type":"int"}],"identifier-field-ids":\[],"schema-id":0,"type":"struct"}] ,
**Key**: "type", **Original Value**: "struct", **Mutated Value**: "surstrutrvdct"
This step ensures that field names do not get mutated, so we remain compliant with Iceberg specs, and only change their values.
We now encounter a wider gamut of errors beyond invalid JSON formats, such as type (e.g. "**Not a valid signed integer**") and schema confusions, and we start gaining ground on the Iceberg reader.
#### Step 3: Structured Avro manifest list fuzzing [#step-3-structured-avro-manifest-list-fuzzing]
This is where things get a bit more interesting. The manifest list corpus is an immutable [Avro container file format](https://avro.apache.org/docs/1.11.1/specification/#object-container-files), which is part JSON and part binary encoded. It has a rich data-structure that includes magic bytes, and sync-markers to delineate blocks and mark EOF's.
We initially preserve the 4 magic-bytes ('**O**', '**b**', '**j**', '**1**') in the header, and mutate the remaining bytes with Radamsa.
Then, we apply a set of probabilistic custom mutations:
* **Sync Marker Corruption** targets Avro's 16-byte block delimiters, corrupting both initial and final markers to test parser synchronization recovery.
* **Block Metadata Corruption** attacks count/length fields that define data organization, potentially triggering buffer overflows or infinite loops.
* **Schema Injection** inserts fake JSON schema fragments to exploit schema confusion bugs where parsers might apply conflicting schemas to the same data.
* **File Size Manipulation** tests boundary conditions through random truncation and null-byte padding at the cost of format corruption.
* **Bit-Level Corruption** simulates storage corruption by flipping individual bits to catch subtle integer overflow and off-by-one errors.
* **Block Duplication** creates structurally valid but logically inconsistent files by reordering data blocks, exposing assumptions about block sequencing.
Such a multi-step approach allows us to target different parser components—from low-level file format handling, to high-level schema validation.
## Fuzzing Parquet [#fuzzing-parquet]
[Apache Parquet](https://parquet.apache.org/) is an open-source, columnar file format designed for efficient data storage and retrieval. It serves as the data layer for Apache Iceberg tables, but is also used as a standalone format for efficient query processing in various big data analytics frameworks.

To fuzz Parquet, we preserve the following sections in a legitimate seed:
* 4 x 2 magic bytes in header and footer ("**PAR1**"),
* 4-byte footer (contains the size of Footer metadata)
* Variable sized footer metadata (File + Row group metadata)
We also perform some additional checks on the validity of the Footer metadata size, so as to ensure that the fuzzer itself does not crash from bad allocations.
The remaining bytes (page header metadata and data pages) are then fed to Radamsa. Post mutation, we recreate the Parquet format by stitching up the header, mutated pages and footers for ingestion.
Fuzzing our Parquet interfaces ([READ\_PARQUET](https://docs.firebolt.io/reference-sql/functions-reference/table-valued/read_parquet) and COPY FROM) for about an hour yields **>60% median** BB coverage, and \~**40%** of branch coverage in some of the most critical reader code-paths.

Our custom logic for mutating data pages, while maintaining essential structures trigger a diverse range of errors\*\*:\*\*
* **Schema corruptions**: (e.g. Cannot infer internal type from "map", Failed to read schema from 'fuzz.parquet': Couldn't deserialize thrift)
* **Page Header Metadata vs Data conflicts**: (e.g. *IOError: Number of decoded rep / def levels do not match num\_values in page header*)
* **Compressions errors**: (e.g. *corrupt Lz4/Snappy compressed data*)
* **Corrupted data pages**: (e.g. *Received invalid number of bytes (corrupt data page?) )*
## Discoveries [#discoveries]
During a **total runtime** of less than **a week**, the fuzzer discovered **5 critical bugs** impacting all our TVFs:
* **READ\_ICEBERG** (also affected our **READ\_AVRO**): An OOB index read while parsing Avro manifest-list.
* **READ\_CSV**: an [assertion failure in Arrow's CSV parser](https://github.com/apache/arrow/issues/45497), leading to a buffer overflow. The bug was caught and fixed earlier by the upstream maintainers, but was still present in Firebolt's Arrow version.
* **READ\_PARQUET**:
* A logic bug in our Parquet reader, triggered via a SIGABRT on a code path that should otherwise have been unreachable.
* A logic bug on parsing empty chunks led to a [buffer overflow](https://github.com/apache/arrow/blob/main/cpp/src/arrow/array/array_binary.h#L111-L113) in Arrow's binary array parser.
* An OOB read due to an incorrect logic that attempted to read more rows in a Page column than was specified in Parquet file metadata.
## Further Improvements [#further-improvements]
As a prototype, this fuzzer exceeded our expectations in terms of quickly discovering some deep nested bugs. But as with all fuzzers, there is room for improvements:
1. **Better Execs/sec:** With all the overheads of structured fuzzing and a long ingestion pipeline, the fuzzer was expected to run slow. While speed was not the main focus during the prototype, we still optimized things wherever possible: be it re-using cURL handles to send queries, or passing complex structures by-reference.
We also tested writing mutations to a memory backed file instead of direct writes to the filesystem, but this turned out to be slower due to higher # of expensive syscalls, such as **two ftruncate() calls** (before and after generating a mutation), and **msync()** to flush the writes to disk immediately.
2. **Power Scheduling**: For this prototype we don't differentiate seeds, but future research can be done on gathering runtime heuristics (e.g., unique response signatures) to assign more energy to promising seeds.
3. **Iceberg fuzzer enhancements:**
* **Manifest layer fuzzing**: We currently stop fuzzing at Manifest list layer, as we dedicatedly fuzz Parquet interfaces. Therefore, some Iceberg TVF implementations might stop reading (for optimized reads) when it cannot find the manifest file based on the "**manifest-path**" field value in manifest list.
Overall, the bugs we found with FuzzBerg encourage us to explore it beyond TVF fuzzing— such as stressing the SQL parser. And so FWIW, coverage guided gray-box fuzzing is not the only way to find a needle in the haystack.
Don't forget to take FuzzBerg for a spin, and we hope you enjoyed reading this.
# Postgres vs. Elasticsearch: The Unexpected Winner in High-Stakes Search for Instacart (/blog/postgres-vs-elasticsearch-the-unexpected-winner-in-high-stakes-search-for-instacart)
In this episode of The Data Engineering Show, host Benjamin speaks with Ankit, former senior engineer at Instacart, about the company's innovative approach to modernizing their search infrastructure by transitioning from Elasticsearch to PostgreSQL for single-retailer search functionality.
Listen on [Spotify](https://bit.ly/3VTmgxR) or [Apple Podcasts](https://bit.ly/3Kdfmkp)
**\[00:00:00] Ankit:** Now what different InstaCart was that our items was so fast moving. It's a grocery store. Things go in and out of stock very frequently. Almost everything that we got retrieved had to be filtered out.
**\[00:00:12] Benjamin:** Hi. This is Benjamin. Before we start with today's episode, I wanted to quickly reach out on a personal note. We've just launched FireVault core. FireVault core is the free self-hosted version of our query engine. You can run core anywhere you want, from your laptop to your on prem data center to public cloud environments. Core scales out, and you can run it in a multi node configuration. And best of all, it's free forever and has no usage limits. So you can run as many queries as you run and process as much data as you want. Core is great for running either big data ELT jobs on, for example, iceberg tables or powering high concurrency customer facing analytics on big datasets. We'd love for you to give it a spin and send us feedback. You can either join our Discord, enter our GitHub discussions, I o. We'd love to hear from you. We added a link to Fireball course GitHub repository to the show notes. And with that, let's jump straight into today's episode. Alright. Hi, everyone, and welcome back to the data engineering show. Today, we're super happy to have Ankit on. Ankit is, well, now a software engineer for RateDB, I think. But before that, I was a senior engineer at Instacart, did a lot of work on Postgres, their data stack. Excited to have you on the show. Do you quickly wanna introduce yourself? Yeah. Kind of share your background.
**\[00:01:27] Ankit:** Yeah. Thanks. Honored to be here. So, yeah, as of two weeks ago, I was an engineer at Instacart. Now I've just recently moved. I mainly worked on the Postgres, so storage infra search infra team, and there's a series of blog posts on how the search was modernized that will more talk in detail today. Search was modernized with hybrid retrieval, and there's also talk of how we moved from Elasticsearch to Postgres, and we'll talk about the reasons later. Before that, I had worked in a bunch of different start ups in Canada and India as well, and I've always been interested in databases. And special love for Postgres because in my experience, choice of database boosted dev productivity in general. Think about it. If there's a lot of things that you can get the database to do, then the applications become simpler. And my non Instacart experience has largely been in, like, think it's a pre PMF startups where the approach of abuse your database to its absolute limits, it works wonders. Now you can argue that it can be done at scale too, but that's for later.
**\[00:02:27] Benjamin:** Nice. So take us through search at Instacart. Right? And I think one thing that also is interesting to our listeners is, like and was similar for me as well for a long time. You have in your mind, like, okay. Here's my traditional relational SQL database, then there's search things. And I think traditionally, there would actually be little overlap between these things. And I think what you're seeing nowadays, and I think Gen AI is actually also driving that even harder, is like you see this kind of merging of these concepts into individual system that can do everything. So, yeah, kind of maybe take us through how do these problems actually pop up in the Instacart product.
**\[00:03:05] Ankit:** Yeah. That makes sense. Let's go one step back. Let's talk about retrieval in general. What is the lay of the land of the retrieval? What was the lay of the land? So there is, like, if you go to Instagart app as a consumer now, again, there's different apps for shoppers, there's consumer, and then there's ads where brands do their thing as well. So I'm sure there are retrievals over there as well. But as a consumer, the first thing that you get in touch is auto suggest retrieval. So you type b a n and you expect banana to be there and bunch of banana smoothie and whatever is there in grocery store. So that's one retrieval. Then there is we call it home search, which means you just say, I want eggs, and the app would show you a list of retailers and bunch of items in each retailer. That's one retrieval. And then there's another retrieval, which is you know which grocery store you care about because you're very passionate about I want my eggs from this specific. And that's me. Like, my eggs come from one grocery store.
**\[00:04:00] Unknown:** I know the chicken. I know the mother of the chicken who raised her, and she has great eggs.
**\[00:04:05] Ankit:** Right? So that's the single retailer search. And that's the biggest, you can say, search retrieval by volume and by server strategic importance as well. Like, Instacart is a retailer first marketplace compared to the peers. And this blog only talks about this line. We don't talk about anything anyone, auto suggest. And you can imagine bunch of other retrievals exist that I don't even know. I'm sure some teams are still using Elasticsearch for their underlying retrieval. Like, as manager, I don't I have never even seen that. So it would have some boxes for search. Someone team was also using ClickHouse via a foreign data wrapper in Postgres because it works for them, so why not? And one flag that I get after the blog was released by Elasticsearch fans is that they felt that the blog claims that we kicked Elasticsearch out of Instacart. No. That's not true. We only removed it for the third kind, which is singular tailored search, and there were good reasons for it. So let's deep dive into singular tailored search. So previously, we had an Elasticsearch that was doing the search retrieval, and the other things was application. So what are the other things? Filtering, the retrieval for some use cases, ranking, and hydrating whatever was retrieved with more attributes, and then ranking after hydrating as well. It's a bunch of different steps. And it's a standard model, if you imagine. And that that's what everyone does if you have multiple kinds of filtering and and ranking. Now what different Instagart was that our items was so fast moving. It's a grocery store. Things go in and out of stock very frequently. And we also have people who care differently about different things. Like, I might care very specific about where is this brand I care about. Someone would care. I just care about getting eggs at the point. I just care about ice cream. I just want an ice cream right now. So you can imagine the solution we were moving towards was an ensemble of models and the weights of those models being controlled by the context. The context can be the search query, understanding of search query. Now these days, that some of the context is enriched by LMs as well. Because of the fast moving nature of the grocery store like these going in and out of availability, in pathological cases, almost everything that we got retrieved had to be filtered out. So we go back to Elasticsearch again. Hey. Give us 200 more results or whatever x number of more results. You do that five times. Now this could happen because we get updates from retailer that, hey. This item is out of stock right now, or our machine learning models tell us that the item is is gone. So for an application engineer, think about it. This situation is exactly what you would see if you, by mistake, do an n plus one query. Right? You you have to fire multiple queries. So what's the solution? Well, the solution is to avoid going back and forth across the network where the network delay eats up all your latency. So we created a cluster that had everything in it. It had search index. It had ranking and boosting tables. It had availability, machine learning availability, and all other filter tables that the query may care about. And also, we could add more ranking and boosting as the team wanted. And that's what the Postgres was too. In a sense, we traded off the quality of retrieval, hardcore core retrieval, with the whole system reducing the network calls. So I guess another way is to say that we push down the compute to the data layer, closer to the data layer, which is a, I guess, an approach opposite to what you guys are more familiar about.
**\[00:07:32] Benjamin:** We love pushing computer to the data layer.
**\[00:07:34] Unknown:** By the way, the beauty about our space, it's confused. There's a lot of confusion. You know? Like, you said something now pushing down the computer, the storage. We can even go further and say, to get the speeds we need, we need to offload a lot of the compute from reading to writing. Right. And it goes into pruning, even into the codex you apply on your data. Like, what we say you push compute down to the storage is actually, in many cases, 80% of what gets you from a to b and gets you to production. And, yes, this is all hidden done by engineers. Go on. Right. So you you came to the realization where, okay, we need to do that. So what did you do?
**\[00:08:16] Ankit:** Yeah. So we traded off the Elasticsearch hype, like, the m 25 with the TS vector, and some modifications had to be made to how the TS vector is queried. But, essentially, this gave us avoiding of the pathological cases. And there were side benefits also. Like, you know, data stored in normalized tables was cheaper than, you know, full blown denormalized. And, also, there is operational benefits and simplification of stack. Like, the in front engineers, which is, like, sort of my team, we were not owning, operating, and managing two different systems. We're just the application would call single query that does a lot, and we'd get done with it. Also, one thing which we missed out writing in the blog is that the cluster is not just a search cluster. So search is one of the workloads it's offering right now. It is a place to go to find what item is available, in what store, what item is available, at what price, including full product taxonomy graph and product and ontology. So you could query products, filter, rank them in any sort of flexible way you want in a single query. Now does that sound like a silver bullet? Maybe.
**\[00:09:23] Benjamin:** Yeah. So maybe one or two questions just for kind of, like, me to also understand better the flow. So there's this pre Postgres cluster kind of running. It can do a million different things. It can also do search and ranking. Now I go into my Instacart app, and I start looking for okay. Let's stick with your x example. It's like, how does the request flow actually look here? Like, do you really send a query for every character I type or you predict what character I get how like, a single interaction with the app of, like, cert looking up a certain term, how many queries does this actually kind of cause on the back end?
**\[00:09:59] Ankit:** Right. Right. That that's a very good question. So the single character goes to the search auto suggest service, which is a completely independent thing. It it's still using Elasticsearch. It is always used, and we didn't change it. It was working beautifully. It does what it's supposed to do.
**\[00:10:15] Benjamin:** You got a lot of hate from Elastic fans. Like, whenever you're you're sprinkling a lot of love for Elastic. Like
**\[00:10:22] Ankit:** I mean, it it's good. It's it's very stable, and it does what it's supposed to do. So my intention was never to say that, hey. This is bad. This is good. He was like, hey. This data model and this whole system doesn't work for, say, that use case.
**\[00:10:35] Unknown:** So, basically, you're saying if it worked for spam checking, it shouldn't necessarily work for anything else. And it's good that it still does spam checking and searching. That aspect of searching really well, So all good, and, , no disrespect to anyone. Yeah.
**\[00:10:50] Ankit:** Yes. And we are doing more than that. We're doing some some, like, if it's a brand search. Because I do remember there was a time when we were running a Red Bull campaign, and, you know, someone would type red, the red bell pepper would be at the top. Because guess what? The CDR boosting, the CDR, the click through rates of red bell pepper, there are so many people who buy day to day that they were making at the top. So all that auto suggest, it still would give you the knobs.
**\[00:11:15] Unknown:** Benjamin, top k problem.
**\[00:11:17] Benjamin:** Yeah. Auto suggestion doesn't go into your service. So at what point, right, like, again, like, maybe take me just through the user flow, would my interaction start sending queries against your system?
**\[00:11:29] Ankit:** So as soon as you press enter, that you choose your search term, at that time, the application would construct a SQL query for Postgres, and it directly hits Postgres. There's no memcache. There's no other caching systems. It directly hits the cluster. The application would do certain things like query we call it query understanding to build more parameters for the search query to the Postgres engine. And then when we get the response from the Postgres query, it is ranked again by the application. So we have, like, different passes of ranking. The first pass ranking is done by the Postgres itself. The second one and the division is just because whatever we want to move fast, we whatever application engineers feel that, hey. This is where we want to iterate. We it's better to keep that in application and not push down to the database. And that's it. Now it sounds simple from the architecture side, but each component of it has its own, like, complexity. I don't think I'm even aware of what that complexity happens. So I can talk more from the architecture, from the data flow side.
**\[00:12:30] Unknown:** Hey. It's all because you're using Postgres. This is why it's complicated. You know? Yeah. Go on.
**\[00:12:35] Ankit:** Right. But the flow is very simple. The application prepares query, sends it to Postgres, gets the result, and then there is some hydration that final clients also do by clients. I mean, the actual mobile apps or your or the web application that also is hitting the same post release.
**\[00:12:54] Benjamin:** Gotcha. And so, I mean, this wasn't vanilla Postgres in the sense that you had, like, use Postgres extensions there as well. Right? Like, you use something like PG vector, kind of take us through your modding Postgres kind of journey. Like, what extensions did you use? Did you custom build any extensions?
**\[00:13:12] Ankit:** That's a good question. So the shape of the Postgres changed throughout the years. So the the large team here is having control over our own Postgres. So that's the only reason why this self hosted. It's a self hosted Postgres. It uses a very new proxy called PG Cat. PG Cat has moved on, and now there's a PG dog as well, the same author. It's pretty good. It's stable. It works.
**\[00:13:35] Benjamin:** Would you recommend PG cat or PG dog Dizdiz?
**\[00:13:38] Ankit:** So I have not used PG dog at scale, but I've played it on with it, and I like it. Conditions apply. Like, try it at scale. There could be bugs, but the devs are like, I know, like, Lev is is a friend, and he's very responsive. So what it allows is, , work with a sharded database in a way as if it was not sharded. So it adds the sharding layer to the PD bouncer. So coming back to the question, the grand theme of here is that we wanted more control over the cluster, how to spin it off, what kind of disks it would have. So that cluster is purely NVMe based. We don't use any network attached. , we are okay if the primary goes down for an hour. We would rebuild it back from the backup instead of failing over. Like, I didn't want to be on call for that if the Postgres had to be failover. So all the intelligence here that handles the failover in a way is in application. Like, even if a replica goes down, you need to blacklist that that replica. All that is happening on the application. So there's a you can imagine that the client for the application is a thick client. It does a lot more than just the query. It does all the failover, high availability. So that allows us to be simple on our infra side.
**\[00:14:50] Benjamin:** How does data flow into this Postgres database in the first place? Was it attached to Kafka? Is it just plain updates and deletes kind of being run all of the time? Take us through the entire journey, basically.
**\[00:15:04] Ankit:** It's a first class citizen of how this cluster operates. So all the rights are pipelined. We don't use Kafka. The interface is s three. So we tell teams who want to have their data in this cluster, create an s three home, create either a bucket or a home, whatever they want to do, and tell us that we would sync ourselves. So we'd pull s three for changes and get the data in bulk. And all the rights to the Postgres would happen like a bulk write, so, like, either 16,000 rows or 32,000 rows, whatever that is.
**\[00:15:35] Benjamin:** So it's append only, or there were updates and deletes as well?
**\[00:15:38] Ankit:** It is merge. The Postgres has merge command. So we would tell teams to tell us what are the primary keys or what are the keys that would help us identify if it is a insert or an update. And there was no real time rights or update allowed. It's that we also supported deletes through a very special mechanism. It was a, like, a you just configure that these rows are supposed to be deleted, and we would do it as at a cleanup at a later stage. A big part of the system was also PGD pack. So PGD pack is something that is like a super vacuum or something. It does it very fast, although it uses a lot of IOPS, and it uses a lot of system resources to do that. It essentially creates a copy of your table, real time syncs it, and at the opportunity it finds, it swaps hot swaps it, deletes the old. So the new table that you get is is defragged. It's slightly packed. It's clustered again. Suppose this has this cluster table. Cluster is in a overloaded time. I hate that. It's causes so much confusion. But it what we get is a nicely clustered table that has its rows tied in a way that helps query at the read time, essentially. We were running, like, tons of free pack on like, almost always, there is some repack running on. And it was hard like, it was tuned heavily because what we found is that the read throughput, we can throw more data if the tables are repacked nicely. These are fragmented and they're not clustered, the rights would slow down. The more of a right throughput optimization, and it's probably an order of magnitude more by the repack tuning. So that's how the rights are going in. And from the customization, we had in the past a custom extension for a dot product. One of the rankings we are doing in the Postgres was personalization ranking as well. So you would get your products, and at the end, the final ranking would be personalized. So this was happening in the Postgres to a very like, it's a 50 line extension in c, just to a dot product and give it back to Postgres. Now that was all deprecated because PG vector came somewhere in 2023, and we moved our retrieval to PG vector, but also PG vector could do your own default dot product.
**\[00:17:44] Benjamin:** Nice. Okay. Very cool. Yeah. It's always interesting to hear, like, these production stories from Postgres. Really, like, the amount of moving pieces and extensions being used is really kind of cool.
**\[00:17:56] Unknown:** The interesting thing is you hear about more and more workloads use cases where when engineers say simplicity, they mean, like, we want the Postgres interface. We want to be able to normalize. We want to be able to have more complex queries running. There are many moving parts. It's not just a on the metal time series, real time, single denormalized, five second query span. Like, this is something completely different up to the point where some of the data is being batch processed on the side and hot swap because you need it to be super hot. The data as it's being served to those users, it needs to be super hot. SSDs were mentioned, like, you can't break your way around it, but then you need to deal with Postgres back end, which is at a whole different place. So you need to have modes, extensions, which gives you exactly the flexibility you need as an engineer to take the concept that the kind of the strength of Postgres and evolve it over time. As you said, like, types change, search patterns change. Absolutely. I love it. Thank you for sharing with us kind of the whole thing end to end and, specifically, kind of the whole effort effort you guys are putting on the back end.
**\[00:19:06] Ankit:** It was almost a silver bullet, but not quite. There were certain trade offs as well. Like, eventually, what we found out was that the DevX, the developer experience in how engineers use it, especially search engineers, it was less than ideal because you realize that most engineers who want to work on search, they are more used to the Elasticsearch shape of the query. And even those who knew Postgres among those, they found the final query to be hard to work with. And scaling, provisioning, sharding, operating this self hosted Postgres in a company that relies on managed cloud for everything else. Like, everything else is on RDS. It's amazing. Right? That was, again, a hard thing. So you have this small team that does its own thing that no one understands. And we also noticed that replicating production behavior in a nonproduction, it was impossible. So if someone is making a change to really know what's the impact, it had to be tested in production. And that makes everyone hard. Like, you know, it slower is the velocity, puts more questions in the queue, like, have you made sure everything is right for this? And this sort of thing steered me to what I'm doing now, like, to sort of as of last two weeks, I joined Pro eight DBN. The promise here is that it's a transactional Elasticsearch alternative. So we ship a Postgres extension that is built on Tan TV, which is a rust fork of Lucene, and it uses Postgres storage. And that's my personal sort of observation is that if I have to rebuild the Instacart like search, it will be like a prey DB. Lot of things would be especially the Demex would be, like, much simple. And taking a stating that conversation in different ways. Yesterday, I saw a blog post by Target. They moved their search workload to AlloyDB in a sort of a Postgres compatible way, and it's not clear what were their motivations. They could be similar motivations, but their outcomes are very similar. Like, the relevance is better because they could join more things in in the database. They also saw the cost of the normalized data reduced. So that's pretty much matching with what we saw at Instacart.
**\[00:21:06] Benjamin:** Nice. Very, very cool. Well, Ankit, it was amazing have you on the show. We appreciate the total Postgres Instacart deep dive. Yeah. We'll follow your work at Parete d b. We're excited about all of that. Thank you for being on the show.
**\[00:21:20] Ankit:** It is an honor. Thank you.
**\[00:21:21] Unknown:** Thank you.
**\[00:21:23] Unknown:** The data engineering show is brought to you by Firebolt, the cloud data warehouse for AI apps and low latency analytics. Get your free credits and start your trial at firebolt.io.
# Professors Joe Hellerstein and Joseph Gonzalez on LLMs (/blog/professors-joe-hellerstein-and-joseph-gonzalez-on-llms)
You should've seen Benjamin's face when we told him that we managed book Joe Hellerstein and Joseph Gonzalez for the Data Engineering Show.
Joe Hellerstein is the Jim Gray Professor of Computer Science at Berkeley and Joseph Gonzalez is an Associate Professor in the Electrical Engineering and Computer Science department. They've inspired generations of database enthusiasts (including Benji and Eldad) and have come on the show to talk about all things LLM and RunLLM which they co-founded. If you consider yourself a hard core engineer, this episode is for you.
Listen on [Spotify](https://open.spotify.com/episode/62UFnosPeuUUDH7i3J1Osl) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/professors-joe-hellerstein-and-joseph-gonzalez-on-llms/id1561927688?i=1000642780090)
Benjamin (00:16.882) Hi everyone. And welcome back to another episode of the data engineering show. Today we have two super exciting guests, Joe Hellerstein and Joey Gonzalez. They're both professors at UC Berkeley and are now heavily involved as co-founders in a startup called Run LLM. I'm sure we'll hear a lot about that kind of going forward. Joey, Joe, do you want to quickly introduce yourself to the audience?
Joey Gonzalez (00:43.849) Sure. I guess I can start. So hi, I'm Joey Gonzalez, faculty at UC Berkeley. I do research broadly in machine learning systems, work on everything from crazy neural networks to really fast systems for serving models and even training them. I've been doing a lot of work recently in large language models, everything from Vicuña. to the bigger LM sys, chatbot arena, and even VLM, so the serving infrastructure. And then of course, I'm also co-founder at RunLM, doing really cool stuff there as well.
Joe Hellerstein (01:11.459) Hey, I'm Joe Hellerstein. I'm also on the faculty here at Berkeley, where I joined really long time ago, back in the previous millennium. My background's in database systems, so I was brought up at IBM Research, and then I was a student on the Postgres project. More recently, you know, over... decades here at Berkeley have done research in a wide range of data centric systems from database engines to machine learning systems with Joey and others, to things that were more visualization related. That work led to the Data Wrangler project, which we commercialized as Trifacta. I was also involved in the Green Plum database effort. So I've been in both industry and academia for a bunch of decades now in and around data.
Benjamin (01:55.53) Awesome. I love previous millennium. Like it's not that long ago, but it sounds a really long time ago. Nice.
Eldad (02:03.208) When they used to build databases that just can't be ripped and replaced, like Postgres, impossible. It's just too good. And no matter how much you try, how hard you try, it stays. Just upgrade to a newer version. Thanks for having us. I was pitching Benjamin for a long time. Can we get some professors? Maybe we can kind of expand the data engineering show beyond the usual stuff.
Joe Hellerstein (02:09.584) That's right.
Eldad (02:30.98) I hope that we'll be, today we'll kind of try to do that. Thanks for joining us. Really excited to hear what you're all about.
Joey Gonzalez (02:39.651) Thanks for having us.
Benjamin (02:41.37) Awesome. So, and Joe, like I'm a database nerd, right? Like I've read your papers, kind of like know you from my database education. Now you're thinking a lot about like LLMs. Like tell us like how did that happen? Right? Like how, yeah, how did you get into that?
Joe Hellerstein (02:56.495) You know, um, LLMs happens to me like they happens to everybody else, right? Uh, it's just phenomenal what, um, what we've absorbed in the last year and gotten used to in some sense, at least on my end, I feel much more, um. matter of fact about the whole situation than I was maybe 12 months ago. But really, you know, like my interest in machine learning systems, which goes back to days when Joey was a grad student and I was collaborating with his team at CMU, Carlos Gestrins team there, is around the algorithmics and the scaling of the data that happens in the, and the, you know, computation that happens in these large scale data driven machine learning models. So Joey and I go
Joe Hellerstein (03:41.189) in machine learning. It was actually quite algorithmically interesting, arguably more so than LLMs, if you're a computer science nerd. And then, more recently, I actually sort of backed off of AI work because at one level, Joey was covering the systems end of AI at Berkeley. And at another level, the community had jumped in with four feet. And it didn't seem like I had to be working on it. Lots of smart people were working on scaling AI systems, and I figured I could work on stuff where fewer people were working. Anyway, I'm happy to leave the fray to those who wanted to fight. So Joey has kind of been more involved over the last decade in scaling those systems. And more recently, he and I have been teaming up to build product around it.
Benjamin (04:25.73) So building product around that, right? Like you also said, Joe, like last year was this explosion kind of in this space. It seems to be changing every day. Like how do you actually build a company in this space, right? Kind of when things are moving that quickly, like how, how do you make sure that things you came up with, I don't know, half a year ago in terms of vision are still relevant today.
Joey Gonzalez (04:46.262) Yeah, maybe I can take that one because I was head of product and still am at a company that's constantly changing.
Benjamin (04:47.648) Sure.
Joey Gonzalez (04:52.826) in a world that's constantly changing. So when we launched the company, it's actually remarkable. The broader vision of what we were doing at what was originally called Aqueduct was to really think about how to bring machine learning into the world of data systems, into the world where people would use it every day. And in our early interaction with customers, this was pre-2023, which is shockingly actually more or less pre-LM, pre-mass adoption of large language models. most people are thinking actually not about how to do really cool deep learning stuff in production, but often just how to plumb machine learning into the workflows that they had. And a lot of that meant connecting machine learning to different data backends, to different data systems. A lot of that meant just basic processing of metadata, things that the data engineering community is actually quite good at. And that was sort of where machine learning was, what was And still is, but that's evolved very quickly with large language models. And actually, I think what's most interesting about large language models is they made it more, even more about the data. It was previously like really cool algorithms you had to deploy and they were special. They're special algorithms for each model, special tasks, special systems. Now there is sort of one set of special systems that we need. And remarkably, it's a text in, text out kind of system. So the sort of APIs would have simplified.
And the real challenge is getting the right data in and getting that out, you know, the outputs to where it needs to go. So in some sense, the space of products as a whole actually hasn't changed that much. It's always been about how to connect machine learning to things. The kinds of machine learning has changed. And throughout my career, it's changed many times. As Joe pointed out, when I started my PhD, neural networks were silly. No one did that stuff. We did probabilistic models that were easy to reason about, or at least, you know, analytically reasonable.
And a lot of work went into algorithms and cool techniques and scaling it. And Joe and I worked on some of that stuff. And then it all changed. And deep learning really jumped in very quickly over quickly by today's standards. It took like four years, five years for people to really adopt deep learning across the entire field, at least in research. And then, you know, we had to transition, but again, the story's always been about connecting the data and systems to the, to the algorithms. And then of course, in one year.
Joey Gonzalez (07:11.93) it went to a whole other class of techniques.
Eldad (07:14.602) You know, in databases, we take 20 years to evolve, like, kind of figure things out here. Like it shrinks, all right? 10 years, five years, one year. It is confusing. Please help us.
Joe Hellerstein (07:14.715) So maybe I can.
Benjamin (07:30.397) Help, help Eldad, he's very...
Eldad (07:32.872) HUEH!
Joey Gonzalez (07:33.446) kind of excited to see what this means. Like what is it all the changes in the technology of machine learning mean for the data engineering database community as a whole? And I have some thoughts, but yeah, it's a, you know, an exciting area to exciting time to be working in this space.
Benjamin (07:47.266) Yeah. I mean, looking from the outside, right. It's like many of the things you're seeing now in the like very mature data warehousing space are also kind of popping up in this LLM world. Like some tools doing ELT, kind of creating your knowledge graphs, text embeddings, kind of whatever. Then you have like your vector database, your knowledge graph to feed context into the LLM, like take us through that stack, right? That's kind of popping up around LLMs and your thoughts on that.
Joey Gonzalez (08:16.974) So I can take a quick stab at it. So at a very high level, the basic adoption, the basic pattern that people are using with large language models, LMS, is I want to ask a question and get an answer. I want to have a support request. And so to get to that, I have a model that's essentially a good reader. You can give it meaningful text that'll read it and produce a result. To do that with your company's data, with your company's product processes, with your company's tone. it takes a couple of steps. So the first step is getting that data into the request. And often when I ask a question, it's not feasible to dump all of my company's manuals information in with the question that was asked. So we'd like to say, I'll read everything about my company, answer this one question. That doesn't work for reasons we can get into later. So the first big innovation that people are excited about today and the kind of deployment of large language models is something like RAC, where I have a retrieval process. Using technology that's in some sense antiquated and also brand new for going from a question to the right piece of documentation, the snippets of texts that are most helpful. And then you ask the LM with these snippets of texts that we think are helpful and the question answer the question. And so collecting that, that documentation into a data system that you can quickly retrieve has been kind of the first big engineering pipeline that sort of sits around the LM's to make them more adapted to my companies or my, you know, domain.
And so that is the vector store. It's kind of neat to see that field evolve because looking up stuff by vectors, something we've been doing for a long time, there's a whole field of information retrieval that's been doing this for many, many years. And they have very cool techniques that we haven't seemed to rediscover yet. So things like, perhaps you wanna find documentation that's frequently clicked on. I've had my grad students go, keep returning papers that are correct, but... not well cited, like, yeah, maybe we should include relevance in how we choose the things we return. So yeah, so a lot of machinery and taking text, breaking into little chunks, and then retrieving those chunks. And we're rediscovering a lot of old techniques. And I think there's a lot of opportunity to bring in more disciplined ideas around, hey, maybe I want where clauses, more interesting ways of selecting text, not just similarity between the question was asked and the snippet.
Joe Hellerstein (10:33.283) So maybe I can just jump in and give a little bit of background for some folks who haven't been hanging out in this space. So RAG stands for retrieval aware generation. If I got that right, Joey. Augmented generation. Yeah. Sorry. Excuse me. And, you know, basically it's like having a search engine connected to your, to your LLM and
Joey Gonzalez (10:42.987) Retrieval augmented, but close enough. Yeah, yeah.
Joe Hellerstein (10:52.731) When we talk about vector search and that sort of thing, we're mostly talking about the kind of similarity search that you do in a search engine. So in a search engine, you type in a bunch of words and it gives you documents that sort of are like the bunch of words you typed. And so similarly here, when the LLM wants to answer questions, it might wanna answer them in the context of stuff that you've indexed in a manner that's pretty similar to information retrieval to search engines.
Joey Gonzalez (11:16.686) So I can tell a fun story about that. So let's say I'm looking for interesting products, right? And so I want products. I'm shopping for my daughter for Christmas and she wants maybe a little drum set. So I'm interested in drums, right? If I just search, I'm interested in drums. Drums is a useful keyword. If I do vector similarity, that might find useful things. I might have a price range. So it'd be nice if there were like where clauses and find at least products that closely match. And then I'd like to feed in the products into an LLM, which would try to provide a summary of the kinds of products that I want.
Eldad (11:20.664) and social media.
Joey Gonzalez (11:47.114) Things that we've discovered that make a big difference. What if you asked the LM before it's read anything to make up a product description? So I'm looking for drums in that you ask LM, what would drums look like? What are kinds of things, you know, is it a metallic drums by looking for, you know, wooden drums, modern, contemporary. So having those descriptions and augmenting what I look up makes a big difference. And so we've since learned that this kind of idea you ask an LM once. To make up an answer, then you look up the made up answer to find relevant documents, then you ask the alum again, now here are correct documents, now answer the question again. So you start chaining these, so you start to have these workflows, it's not just one call, but many calls that go back and forth between different data systems. And that's kind of the other thing that's starting to merge is how to orchestrate alums in this more complex process.
Benjamin (12:34.118) So in this process, when I then go to my vector database to figure out relevant documents or something like that, like what actually generates the query for the vector database, right? Like, okay, with like Looker or Tableau, it's like the kind of program itself, generating SQL to query the underlying data store. How does it work here?
Joe Hellerstein (12:53.927) So think about this much more like a search engine right now, at least in the current state of the art, than you think about like a relational database. So it's literally just going to send essentially your question in to find relevant documents the way a search engine would take your free text and find relevant web pages. Now, as Joey alludes to, we may want to improve this over time with things like where clause predicates that we're familiar with in SQL that filter out stuff that's irrelevant, where you can have more logic in your queries.
Where we are today is pretty much what's called vector search or near neighbor search, which the simplest way to think about it is it's like web search. You're just finding relevant things in a ranked list. And when you look at stuff like PG vector, you know, where they've basically put vector search into a relational database, it's very much like the text search extensions to Postgres. In fact, it's built on the text search extensions to Postgres. So again, even though you're in the relational database, that part of the query, is just a text search lookup and it's finding stuff that is similar to what you asked about.
Joey Gonzalez (13:57.32) So, and this is just.
Eldad (13:58.95) So it's yet another data type, Benji.
Benjamin (14:02.574) and relational systems will swallow it all. Yeah. What's your take on that aesthetic controversial thing? What, where do you see this moving?
Joe Hellerstein (14:14.531) I don't find it controversial in the sense that I don't think relational systems are relational systems really anymore. They're just like the databases that do 90% of the work and 90% of what they do is relational. But as Postgres showed in the late 80s and has been showing until now, and it's pretty much been adopted broadly in the industry.
You can extend the relational model with all these data types and plugins and indexes and various things. So people had looked into plugging in search before into Postgres and other relational databases. And it works OK. And so you can use it for this purpose. Other folks will say it's really good to use a traditional search engine for this purpose. Some people will go use Elastic for a rag. It's kind of heavyweight. It's not clear that it's the state of the art.
And those systems are architected a little differently than a relational database. They're tuned for their workload. But the techniques under the covers are all textbook stuff we all know. So I don't foresee that, for instance, vector databases for LLMs is a viable market, if that's sort of the question, because it's just one feature of data management that you can integrate into a bigger platform.
But this is not relational maximalism. It's not like everything must be a table and you must use SQL. It's just much more about like sharing a backend across all your data.
Joey Gonzalez (15:31.211) Yeah, I'll go a step further, because right now we're really excited about text, because, well, that's what LMs read, but...
Often you have data that's not text, and it would be nice if the alum would look at that too, like pricing information, product features, do comparisons. It can be represented as text, but it's rows in a table or a document store. I think we're gonna get to a world where text is SQL or query engines from, I have a question, you generate a query that might dig up the right information and stuff that into the context as well, and start to use alums to do... query generation, maybe query refinements, so get some results back and then adjust it. And so programming with these like small calls to linguistic logic that extracts pieces of code, not just documents might actually be a big part of the workflow, and which means that we're gonna start to pull back in other parts of data engineering, tool chain into the LM answers, which should be really exciting.
Benjamin (16:29.99) Okay. So this world then would look more similar to actually what we also see today in the data warehousing space, where you would have then like an ecosystem company that just focuses on how do I generate like great SQL queries based on the text input I have to feed it into snowflake, redshift, or whatever, whatever type of system. Gotcha. Okay. So, no, sorry, go ahead.
Joey Gonzalez (16:46.086) Yep. This we have to... Yeah, keep going. I was just saying, one of the funny things about LLMs is you have to understand, they just predict the next word by guessing. And so the basic technology, all it's doing is saying, given all the words I have so far, what's the probably next word? Having read everything on the internet and as much code as possible. Which means that writing SQL queries, yeah, it's amazing it works, and it doesn't.
Eldad (17:08.914) It's amazing.
Eldad (17:13.249) People are very terrible at guessing. Get something better to guess. So I've been thinking about what you said on the stack. So obviously, there's this whole opportunity for new products, right? Isn't that a bit boring if we just try to squeeze everything into a database? Joe, you mentioned it. There's no need to squeeze everything into a database just because it's a database. So.
Joey Gonzalez (17:18.898) That's what we have is a powerful machine for guessing.
Eldad (17:39.748) we'll have this wave of new startups trying to change how data professionals you work and build solutions, right? So today you're setting the database to the enterprise. Thanks to Snowflake, we now have a market where data engineers can build stuff that a few years ago, only big engineering groups could do. So they kind of simplified the stack, right? They removed the stack. Just write SQL.
We'll do all the magic behind the scenes, including the hardware, software, whatever. Then there is the more serious projects done by longer term engineering, machine learning, companies that have real expertise. So I would guess those machine learning scientists would evolve and start building LLM stacks with the tools and vector database and everything. You've mentioned Hadoop, would it look like
Joey Gonzalez (18:35.893) I'm going to take a few minutes to get this done.
Eldad (18:36.052) the terrible hadoopiers, by the way, which we all thought were almost all of us thought was amazing. I hope that we will not go through the same retro, right? Like I hope that we will shortcut some of the stuff. I really hope that we are not going to have yet another cycle where we basically have data scientists and are doing the same thing with a new name for a very similar stack. And we're really going to see some changes here. Obviously on many of those fronts, maybe that new front can change that. But really, it would be sad to see Hadoop mindset taking over again. It would be nice to see startups or companies going to remove a lot of that complexity. If we can tell you one thing from experience is people haven't learned. Like you take the simplest thing, you need to join two tables and that's still a big problem for human beings. And
Joe Hellerstein (19:21.627) So.
Eldad (19:32.936) Yeah. So kind of my, my question to you is, is how do you foresee the market? Joe said a very big, I gave a very big statement, which I love, which is there is no future for vector database just for the purpose of, of doing vector search, which I completely buy into that future. What else, like what other things will exist if not that maybe that takes us to what you're building.
Joe Hellerstein (19:56.903) So yes, I was exactly going to try to build that bridge in the conversation. So at some level, I think a couple of years ago, where we started the company was in an environment where there was a ton of mayhem around the stack, lots of players building little things. But unlike the Hadoop years, these were, you know, early 2020, 2021, highly capitalized startups building little things.
Benjamin (20:01.71) Thank you.
Joe Hellerstein (20:24.003) So what happened in the Hadoop years was everybody had their own pet open source project and Cloudera and Hortonworks rolled them up. And we just had mayhem because every open source project was its own ship. And we were trying to build a fleet out of that. But two years ago in the machine learning space, it was the same thing, except these ships weren't open source. They were startups. Some were open source, some weren't. But each one of them had its own funding and was pursuing the market and telling themselves if we build this piece of the stack, eventually we'll own the whole stack.
And that was not going so well. So we wanted to make that very simple as a SAS environment for users so they wouldn't have to assemble this crazy stack. What's changed with LLMs is a large fraction of use cases can be handled with a single sort of brand of inference.
So the machine learning engineer's job is no longer, you know, what crazy ML packages are you trying to do machine learning with? It's more around how do I enhance the value of the LLMs, whether they're the open source LLMs or the commercial LLMs. And that gave us a fair bit of clarity.
One of the things I think that's it's kind of common agreement in Silicon Valley right now, but we lived it is that building the sort of picks and shovels company for the gold rush and machine learning is not going to be a strategy for at least the foreseeable future. Due to market changes, you know, it's hard to compete with the big players there. And also, just the stack probably is too fragmented to go do that.
So everybody wants to understand kind of what business value they're adding. We're not building just tools. We're building a thing that solves a business problem.
Joe Hellerstein (22:00.215) And so at RunLLM, what we decided was that the clearest space to make a difference here is in the developer community, because they're going to be early adopters. And unlike, say, medical applications or self-driving cars or stuff, the barrier to a viable product is much lower. If it helps the developer be more productive, great. If it's sometimes wrong, it's not the end of the world. There's a human in the loop, and they don't do something that's hopefully life-threatening on a daily basis.
So we've been focused on that developer environment and building augmented models as well as user experiences around that so that developers can be more productive, particularly with complex APIs. So a lot of what Relatin LLM is now focused on is based on work in Joey's team and a project called Gorilla that was focused on using LLMs to make complex APIs more approachable. And maybe with that, I should hand it off to Joey, because it's really his baby.
Joey Gonzalez (22:53.358) Yeah, I think maybe the easiest way to tell the story here is that developers have access to lots of things that they need to read to get their job done. And if we can bring these ideas of RAG and actually fine tuning, which we didn't talk about together with proper integration, not into just your documentation, but Slack discussions, GitHub, everything, to give you, to give our AIs visibility into what's going on. So you as a developer can ask questions and maybe even be notified when things happen that might change the way you're approaching stuff.
Eldad (23:18.426) And I think that's the best way to do it. And I hope that helps. Thank you.
Joey Gonzalez (23:23.394) In some sense, it's helping developers be more productive, which is something that, let's say, Copilot does. But our focus is more on the integration with all the data and the things that happen around you, your development context. Copilot helps you write code. But if you're trying to write a design document, trying to think through a bigger process, or maybe just looking up other ways you could do the thing you were trying to do already, we've been building a tool around that. And...
Our focus is kind of interesting. We started out as a picks and shovels company and realized that, you know, really focusing on a solution, a product would help us both bring LLMs to the group that are most likely to adopt them early and, and build them into their workflow in the future, and also give us more visibility into how this, this ecosystem will settle. It is true that there's currently a wild west of small, you know, LLM innovations from vector stores to like lane chain to rag add-ons.
and how they fit together. We once thought one could describe an LM stack, but that keeps changing so quickly that it was more clear to us to focus on just building a solution to a real problem with that technology to understand how it fits. And maybe someday we come back to the broader LM, run LM process itself. As someone who pioneered inference technology, it's kind of disappointing that we're not racing to build the world's best serving technology.
But if you look at the price war that's already ongoing, it's not a fun place to be. As Joe pointed out, these models have converged to a basic set of architectures, a basic set of APIs. You can win on price and speed. And that's currently a race to the bottom to bring the prices down. To the player right now, the prices that some people are proposing are below the power costs of the GPU running at like idle. So there's no way there's a sustainable prices.
Eldad (25:09.99) We love those, the Nvidia sponsoring growth bottom up. Sometimes that's the only way if you're in Nvidia. Super interesting.
Joey Gonzalez (25:10.013) Yeah, not a fun place to be.
Joey Gonzalez (25:16.867) Yeah.
Benjamin (25:22.129) So how exactly, like one thing I was curious about is you mentioned this can help me write a design doc, right? So I have a clear picture in my head of how to use copilot. Like I'm kind of hacking and in my IDE, it's kind of recommending code. Now I'm coming in and I say, Hey, I want to build a new, I don't know, hash aggregation into my database system.
Joey Gonzalez (25:37.489) Mm-hmm.
Benjamin (25:42.622) And I need to write a design doc because I want to figure out like, surface it with the team, kind of get thoughts. Like how does it, how does the interface here actually look like does run LLM propose the entire paragraph to me? Does it just show me relevant other things? Like what's actually happened?
Joey Gonzalez (25:44.473) Yeah.
Joey Gonzalez (25:47.778) Mm-hmm. That's it.
Joey Gonzalez (25:57.848) So it's a great question. And this has been the fun part of the product. One of the first realizations is it looks different depending on who you are. So some people really want to work in a copilot like environment. They want stuff on the side of copilot that maybe gives them the right documentation, suggested context, ways to approach what they're currently doing. Others wanted something more like a chat bot that I can ask questions. How would I do this with Terraform? Is there a faster way to get this Kubernetes deployment going?
And here's my current setup. Give me suggestions on how to improve it. And then others we've kind of talked to are looking at something more like a form. So imagine something like Stack Overflow, which was some of our inspiration. I should be able to ask a question and then get hundreds of LLMs to come together and try to answer the question, provide guidance, rate each other, give feedback on each other's answers. Like, wow, that's a much better way to approach this. And my question was totally wrong.
But how did that happen in real time, not in weeks? And be able to interact with that process with threads. So there's a lot of opportunity to build more than one interface. As a startup, we're trying to focus on the easiest things to integrate with first, the kinds of Slack, basic chat, very basic kind of message interface, message interface. But yeah.
Eldad (27:05.728) Quick question on that. You've mentioned Slack, kind of as a data source. I'm building a startup. I'm building a product. And now I'm coming to a user and ask them to open up all of their Slack history. So today, that feels similar to give Gmail approval to whatever, share your picture. Now I go now.
Joey Gonzalez (27:27.725) Yeah.
Eldad (27:31.22) I don't know, I'm registering to a podcast tool and it asks me, well, if I have access to your enterprise data, your Slack, your Salesforce or this and that, wouldn't we need as users kind of an abstraction there because every startup that does AI will ask for the same Slack feed. So kind of how would that work? Will we start upload that? How is the learning process even working? Because ML is very,
Joey Gonzalez (27:39.486) Okay. So, we're going to go ahead and get started. So, we're going to go ahead and
Eldad (28:00.424) vertical, right? We took a few scientists, they were sitting for a year, they own all the cleansing, all the hard stuff. How does it work? Like if I connect to a product, how does it work?
Joey Gonzalez (28:10.937) So I'll start narrow. Yep. Okay, great job. Yep.
Joe Hellerstein (28:11.195) So maybe I can take a crack at this one, Joey. So I think unpacking some of your question there, part of it was around authorization and security. Part of it was around how does this actually work. And so let me see if I can unpack that a little bit one thing at a time. So having spent time in the enterprise space at a couple of companies, I would say a lot of the discourse in the newspaper and on X about security is often just not really a problem inside a company. So let's take your startup, right? Your Slack is already visible to your employees. And if you're publishing it like dev help channel on your Slack, you probably aren't going to object if the dev help channel is being indexed and here.
Benjamin (28:55.752) Thank you.
Joe Hellerstein (28:58.059) indexed is indexed into the LLM, right? So there's gonna be places where we're gonna get very comfortable with your employer saying, these are public repositories. When you participate in this stuff, your comments are being indexed and shared.
Right. And then there'll be a bot in the Dev Help channel that will answer questions alongside your fellow engineers, right? And so I don't necessarily feel like the auth problem is all that complicated for a lot of these enterprise use cases that we're looking at. Now, if you talk about patient records or inside a hospital or something, life gets much more complicated. And again, this is why doing the Dev facing use cases first is just a lot easier for everyone to digest.
Eldad (29:36.752) Yeah.
Joe Hellerstein (29:37.871) So there's that piece. The second piece, which I think you're hinting at, which is part of the value proposition of a company like RunLLM, is you've got lots of really useful data in your enterprise that chat GPT knows nothing about. You've got all your documentation, you have your private GitHub repos, you have your Slack channels. These are all things that as developers, you guys share and you've authorized each other to use so you'd be happy to quote unquote index or to fine tune a model or do rag on. And so...
Then that gets to your operational question. How do I take this corpus of stuff inside my enterprise and do something that integrates the wonders of chat GPT with my private stuff?
And the answers to that technically these days are twofold. There's rag as we've described, so you index the stuff. And there's also fine tuning the model where you add another layer of inference on top of open AI or your open source model that is trained for your data set. And maybe with that, I'll hand it back to you, Joey, as to like, you know, what actually happens when I point run LLM at say my GitHub repo.
Joey Gonzalez (30:42.013) Yeah. So.
It's great questions. We mix these two technologies, two approaches, RAG, again, the index retrieval story that we discussed for a while, and then fine tuning, which maybe I'll give a very quick description. Maybe think of it as I'm studying for an exam, and how would I study for an exam? I could generate some practice questions and then give the model some example answers, and then tell the model, adjust your weight so that rather than predicting what you would have thought was the next word, predict the next word that more aligns with this answer.
Benjamin (31:00.311) Thank you.
Joey Gonzalez (31:10.478) So this is instruction fine-tuning, and we give actually, we generate practice question answers for your documentation, for your code, using external models like GPT, and then we can retarget our models. In fact, we can retarget hosted solutions like OpenAI's fine-tuned services against your documentation through these questions and answers. And we go even further to make it so that if we retrieve documentation, it gets good at reading the documentation and answering questions.
So we practice on the sort of an open book exam, if you will, and get better at even reading your documentation when being asked interesting questions. And so then we can specialize models for your domain. And we don't just yet, we could specialize models for your tone and how your company likes to speak about stuff to capture even the style of how you approach stuff, let's say, in a message board. So that's one aspect, so making the models more aware through, again, RAG and fine tuning. The other thing is just as an interface, where do I chat?
I chat on Slack most of the time today. So if I want to ask a question, it'd be nice if a bot would jump in and say, yeah, someone asked a similar question. Um, here's a more detailed answer. And even better tomorrow, I get a message from the bot, Hey, actually someone's figured out a much better solution to the problem that you originally asked, um, here's the better solution. So it chats back. Um, it's not just a matter of me going to a website and asking a question. Co-pilot doesn't come back to you and say, Hey, is there a better way to write your code now?
Um, and that's kind of where we want to be. So the interaction, the surface, uh, Slack provides a surface that people are already used to chatting with. Um, we're building other services. Well, uh, getting to the higher point, uh, companies in the space have to find ways to bring LMS back to where everyone is today, um, so that it can integrate in how people work with the data and the tools that they use. Um, and so a lot of it is, uh, integration, a little bit of it is a clever, you know, kind of combination of rag, fine tuning and other tricks that we're studying and research.
Benjamin (33:01.994) So, but like the, the fine tuning itself, this is then also done by OpenAI at the end of today, right? So kind of your, like, or you're the ones actually fine tuning the model.
Joey Gonzalez (33:02.305) Bye.
Joe Hellerstein (33:02.375) some.
Joey Gonzalez (33:09.822) Thank you. Yeah, it's complicated. So let me answer in two parts. So OpenAI and every other fix and shovel company now offers a fine tuning service. That fine tuning service takes strings with question answer pairs. That's it. And then it runs gradient descent, the boring algorithm that we've been doing for training on that data with a loss that everyone uses. The art in fine tuning now is not the algorithm itself, but the data that you construct.
That's what we do. So we construct very clever datasets that allow the vanilla fine-tuning techniques that everyone sells at under cost of the GPU power to be able to make the models that they host for us better. So we then own the model inside of OpenAI and then can serve predictions. But if we decide to use together or any scale, or I forget where the other recent startup announcements, any of their fine-tuning services would be able to take the same dataset and make a model better.
I'd love to run our own fine tuning. We've been studying the systems for both fine tuning and serving these models, but right now, because of the VC money being thrown at models and companies, it's cheaper to use other people's solutions to do it. So the innovation actually isn't in the technology stack. It's really in the data and how it gets put together.
Benjamin (34:28.621) Do you see that change? Like, do you see that changing? I mean, obviously if I'm open-air, I like, I'm not going to give you my model weight so you can start fine tuning, right? Like this is actually a big part of my secret sauce, right? You do?
Eldad (34:28.812) Well, we'll see. Go ahead.
Joey Gonzalez (34:39.802) You do. So you don't let me look at the weights, but you let me give you data, and then you fine tune your model for me. And you love it because I stay inside of your ecosystem. And you charge me a little bit more once I fine tuned my special model, my special version of your model. I never get to see the weights on my model. Yeah. Oh, good question. Yeah.
Benjamin (34:53.005) But I got confused by the eyes and use now, like I didn't know who was off my eye anymore and who was.
Eldad (34:59.168) So what you're saying is OpenAI, part of that is building an ecosystem that's actually made up of that optimizing, fine tuning the model for the vertical need. But will OpenAI own the indexing? Will Slack own the indexing? Will the startups ingest all of that Slack, like the GitHub or the Slack data, like all of that huge data set?
Joey Gonzalez (35:25.642) Great questions.
Eldad (35:27.117) Where is it? Who owns it? Who pays for it? How does it work?
Joey Gonzalez (35:30.654) So right now we do that. So we manage the data. If you like, you can have OpenAI do it for you at an extremely high price. They charge a very high premium to stick your documents in there, whatever they have doing their indexing. So yeah, so there is multiple solutions to this. We've put together a lot of commercial solutions that are currently low price, but based on open source projects. So we can easily swap in and out if we needed to. But right now the economics of all of these hosted solutions are pretty favorable.
So that's where we go today. It is true that there's this integration of different data, tools, models. That's kind of where people are trying to differentiate today, since the model itself is hard to differentiate on other than the big commercial providers. And the serving infrastructure itself, like what used to be the exciting stuff, is now pretty horizontal and hard to differentiate. So, yeah.
Benjamin (36:22.498) So to play a bit like kind of like devil's advocate and like push a bit, maybe it's like if I was working at Slack now, right. It's like, and if I'm seeing what you guys are doing with run LLM, I guess a lot of like the kind of cool stuff you're doing then goes around, hey, how do you kind of crawl Slack, Google docs, kind of build good ways to augment your models? If I was a Slack executive, I would really want a piece of that pie, right? Like I completely see the world moving into this direction that you're describing and outlining. But.
Eldad (36:44.028) And I'm going to be talking about the importance of the
Benjamin (36:52.53) I probably over time want to make it harder for you to kind of index Slack so that I can be the one indexing and kind of, uh, cause I own the data, right? Like I actually, like, I'm in the kind of like as Slack in the great position that I have all of this data. Um, I'd want to make it harder for a company like you to index it so that I can actually generate some of the value that, that you're kind of then trying to grab as run LLM. What's your, what are your thoughts on that?
Joe Hellerstein (37:17.103) Yeah, I have, I have a pretty clear sense that that's not going to happen for the same reasons that, um, these, uh, end user tools don't tend to own your search or other infrastructure, right? Uh, uh, inside the walled garden of Slack are many good things. And they also are constantly asking you to connect your Google docs, connect your GitHub, et cetera. Um, but
Typically, if you're even an individual, but certainly in an organization, you don't offload your data management to a chat tool. It's just, it's a type error that people usually don't engage in. So Slack, if they're smart, we'll have chat bots in there that understand the Slack data and they'll ask you to wire stuff up. But the idea that Slack is going to do a great job on helping developers do, uh, you know, API exploration and all the stuff that we're doing at, at run LLM just seems a little farfetched. that the competition will be with Amazon than with Slack, with the infrastructure providers, because it's a horizontal problem across all your data silos. And so you want a solution that goes across the data silos and presents you probably multiple user experiences, as Joey's been alluding to. So the Slack user experience is not the only one you'll want, and the data in Slack is not the only data you'll want. And so I don't see that as really the way the industry will evolve.
Benjamin (38:36.886) My follow-up question to that would be then maybe Slack is just like losing relevance over time, right? It's like, maybe there is a better tool where you do your chatting, writing design docs that's like wired into everything where kind of then you can like answer all of those questions. Right. Just like.
Eldad (38:50.342) But actually, we're, I think we're not discussing Slack, the IDE that serves users to chat. This has nothing to do with it. We're discussing Slack as a data source.
Benjamin (38:55.809) Right.
Eldad (39:02.652) How do we access it? In your case, do you need to pay them to ingest the data? If someone registers, are they offering you ways? Because there isn't an easy way today to treat Slack as a data source, but there are needs like as a data source versus humans that just search it.
Joe Hellerstein (39:22.663) So let me shift the conversation just slightly to provide perspective, not to dodge the question, but to provide a version of the question that maybe resonates better, because I know where you're going with this. Stack Overflow is maybe a better example, because Stack Overflow usage has been plummeting since ChatGPT and Copilot came out. So that was the place where there was a data set, a user experience, and a community around certain tasks, and a better, more intelligent place to do that has emerged in the public domain. So you can ask yourself, I guess, if I'm not gonna use Stack Overflow to answer my coding questions, I'm gonna use something else. Can I get that something else to also understand the internal stuff in my organization? And I think that evolution seems pretty natural to everyone. It's like, oh, I want a thing that's better than Stack Overflow because it uses LLMs, and I don't want it just on open source stuff. I also want it on our closed source stuff.
Right, so that's not really a user experience question, although we are certainly playing with the idea of what would Stack Overflow look like if it was dynamic. Stack Overflow is a materialized view of human knowledge. But if you replace it with an LLM, you've got an unmaterialized view of an infinite Stack Overflow, right? You can have seven different chatbots that are fine-tuned in different ways or use different foundation models answering questions and arguing with each other. And then that can be almost instantaneously dynamic. Instantaneous is strong, but within a minute or so. And so you can ask anything and get hundreds of comments. So that's kind of the shift I think that's interesting. The question of what's the API to extract data from Slack, and do they make it public, and do they charge for it? This is in there, but it's kind of a mechanical question. It's not this kind of like, paradigm shift sort of question, I guess. And the way I look at this space is, even today, you can build tools that will crawl your Salesforce and so on and build an enterprise search solution, for example. These kind of horizontal tools are going to be desirable. People who have interesting data will be incented to share that data because they probably can't own the universe. And there will be APIs to do that stuff. And perhaps,
Joe Hellerstein (41:43.003) people with interesting data get a slice of the pie. They get to charge you something for using the APIs. But I don't believe in a end user tool lock-in future at all. And I think there's room for these kind of more use case horizontals, let's call it. It's still a vertical because it's all about developer experience, let's say, in our scenario. In a health care org, it might be all about patient records. But the point is you pick a domain and you fine tune to it and you get really good at it. And you find the touch points in that domain where there's apps that people like to gather around. And that's kind of the ecosystem you want to integrate. But I think surely big players are not blind to this opportunity. But then they get analysis paralysis around which verticals should we go after. And maybe we should go after everything all at once. But then we're competing with OpenAI, so maybe we shouldn't do that. And you can see how taking a slice with a startup-sized focus, you can really make a dent.
Benjamin (42:41.19) That's a great answer, Joe. Nice. I think my closing question on the run LLM side would be, how do you see this evolving like over the next kind of one in two years, right? Like at the moment you're super focused on this internal developer side. When I'm thinking now is like, is there a path to open this, for example, to customers of your customers or something? Cause we're building a database. There's a lot of like just SQL in our Slack that's kind of relevant to the system we're building, right? This is something where you could in the next step then, I don't know. Write a chat bot for our customers. Like, yeah, like take us, take us through the evolution in which you guys have planned for, for the 2024.
Joey Gonzalez (43:19.759) I can start. So again, we may be overindexed on Slack. So a lot of our, in fact, right now we're not indexing Slack in our first cut of it. It's just a chat surface. And within the thread that we're in, we have that information, but we're not pulling all the information. In fact, we're focused more on a lot of your design docs, your internal GitHub documentation which means that companies that have lots of complicated APIs that might want their customers to have quicker help on those APIs and faster turnaround with lots of information that's kind of tailored to them might want our product not just to help themselves, but to help their customers. And one thing we're kind of playing with is the possibility of making this interface, this next generation of kind of the stacker of flow, something that's public that anyone could go to and then helping people find the right technologies to solve their higher, you know, higher level technical problems and having a, you know, a cogent discussion about why that's a good idea from a security perspective and from a performance perspective with different LMS giving that kind of insight. And I think that could be fun. And I think again, a lot of companies that have interesting APIs that they'd like to get in their customers hands might want to work with us to help make that process easier.
So I think there's an opportunity just in the developer world to broaden beyond just, you know, internal help to external help. Then of course, we've already started talking to you like, this is great. Can you help us with other parts that aren't development oriented? And, you know, in some cases, those other parts are a little further down the road, might require, I don't know, HEPA approval, like going through different processes, but getting started at something that's easy, again, in developer space and growing out is something that I think is a good path for us as a small company to grow and succeed.
Joe Hellerstein (45:01.819) So I don't know if you were looking for a role as a product person at OpenLM, but definitely the suggestion you're making of, wouldn't it be nice to have a well-trained chat bot for the API we're exposing to our customer, is something very much on our minds. I think if you're familiar with OpenAPI or Swagger, we're trying to auto-generate API docs as it is. It's pretty primitive, but it's something. But that's a good point kind of spirit, you put an LLM behind that and you perhaps populate it with more information than just what's in your open AI spec. And you get something that's a chat about on your web page that can answer your customers' questions, right? So yes, we are thinking about that. Joey alluded to the idea that maybe there's a Stack Overflow open source kind of thing that might be interesting. And of course, we talked about internal usage as well in companies. These are all things that are on our minds as we look into 2024.
And I think we'll pursue some subset of them with product announcements, you know, but we'll see exactly what those look like as they emerge.
Benjamin (46:08.062)
Awesome. We're, we're excited to see those. Um, thank you so much, Joe and Joey. Like really, I mean, this was great. Thanks for bearing with LDAT and me for all of our rookie questions about LLMs. Uh, it was great to kind of listen to two experts in the field telling us more about that, the ecosystem run LLM. So yeah, thanks a bunch for being on the show. Any closing words from your end?
Joey Gonzalez (46:14.82) Yeah.
Joey Gonzalez (46:30.554) Well, I guess I can end, you know, 2022 was a crazy year. 2024, I suspect will be twice as crazy. And we'll see a lot of startups hopefully succeed. A lot of startups are going to fail. And, and I actually, I would guess that by the end of 2024, we're using vastly different workflows, models, and kind of exciting new ways. I just, it's insane the rate at which research is happening now and, and how it's immediately being translated to products and a lot of cases. So.
Eldad (46:56.389) Thanks for watching!
Benjamin (46:58.846) That's awesome. So we'll have you back at the, we'll have, we'll have you back at end of 2024 and then see, see how that age Joey. Awesome. Yeah. Thanks. Thanks so much for being on.
Eldad (46:59.656) We'll take your prediction to the bank.
Joey Gonzalez (47:02.242) Yeah, we'll see. Yep.
Joey Gonzalez (47:08.026) Yeah, sounds good. Yeah, thank you. Yeah, bye. Thank you.
Eldad (47:12.692) Thank you. Thank you.
Joe Hellerstein (47:14.375) Thanks, guys.
# Pruning even more data with late materialization (/blog/pruning-even-more-data-with-late-materialization)
Have you ever explored data with a quick query like the following?
```sql
SELECT * FROM hits ORDER BY EventTime DESC LIMIT 10;
```
You want to see all columns of the first events in this table. Using SELECT \* can be convenient. But loading all columns can be slow because the system needs to read a lot of data from all these columns. Consider the performance difference:
```sql
SELECT * FROM hits ORDER BY EventTime DESC LIMIT 10
WITH late_materialization_max_rows=0;
-- Execution Time: 16 s, Scanned Data: 87 GB
SELECT Title, EventTime FROM hits ORDER BY EventTime DESC LIMIT 10
WITH late_materialization_max_rows=0;
-- Execution Time 1.3 s, Scanned Data: 10 GB
```
The query that returns all columns scans over 8x more data and has more than 10x execution time. Note that the query turns off late materialization with late\_materialization\_max\_rows=0.

Let's see what happens with late materialization:
```sql
SELECT * FROM hits ORDER BY EventTime DESC LIMIT 10;
-- Execution Time: 0.5 s, Scanned Data: 1.5 GB
```
This is over 30x faster than the original version and scans over 50x less data! It is even faster than selecting only the Title and EventTime columns. Since Firebolt version 4.28, this is the default behavior, so you automatically benefit from the performance improvement. You do not have to do anything to get these performance improvements.
Late materialization is not only useful for exploratory queries. The following, more complex analytical query also becomes 5x faster:
```sql
-- Without late materialization
SELECT watchid, userid, eventtime, url, title, referer,
(responseendtiming - responsestarttiming) AS server_response_ms
FROM hits
WHERE
eventdate BETWEEN DATE '2013-07-07' AND DATE '2013-07-14'
AND ismobile = 1 AND responsestarttiming > 0
ORDER BY server_response_ms DESC LIMIT 10
WITH late_materialization_max_rows=0;
-- Execution Time: 1 s, Scanned Data 5 GB
-- With late materialization
SELECT watchid, userid, eventtime, url, title, referer,
(responseendtiming - responsestarttiming) AS server_response_ms
FROM hits
WHERE
eventdate BETWEEN DATE '2013-07-07' AND DATE '2013-07-14'
AND ismobile = 1 AND responsestarttiming > 0
ORDER BY server_response_ms DESC LIMIT 10;
-- Execution Time: 0.2 s, Scanned Data 0.7 GB
```
## Most data can be pruned in top-K queries [#most-data-can-be-pruned-in-top-k-queries]
But how is it possible to reduce the amount of scanned data so drastically? The answer is simple: Most data is not actually needed to compute the result. For the queries above, only 10 rows are needed in the result. So for most columns, the system only needs to access these 10 rows, a tiny fraction of the data. Without late materialization, however, the system reads all rows just to discard most of the data.
But the system also needs to compute *which* 10 rows are part of the result. Fortunately, not all columns are needed to compute which rows qualify for the result. In the first example above, only the column EventTime (which has only 8 bytes per element) is needed to compute the qualifying rows. There are about 100M rows in this table. So the minimum amount of data that must be read to find the qualifying rows is about 800 MB.
Once the qualifying rows are known, the respective data from all remaining selected columns needs to be read.

## How data access works in Firebolt [#how-data-access-works-in-firebolt]
We previously published a blog post about data pruning in Firebolt: [Making Firebolt Fast By Doing Practically Nothing](https://www.firebolt.io/blog/making-firebolt-fast-with-pruning#a-real-life-use-case). Here we repeat the most relevant points for late materialization.
Firebolt divides tables into multiple tablets. The rows in each tablet are sorted by the primary index columns. The primary index is a sparse index, so it only has entries for granules instead of individual rows. A granule is a group of 8,192 consecutive rows by default ([documentation](https://docs.firebolt.io/reference-sql/commands/data-definition/create-fact-dimension-table#storage-parameters)).
**Tablet Pruning:** Each tablet stores min-max sketches for all columns. So if you have a query of the form WHERE col\_b = 10, Firebolt can check whether a tablet is guaranteed to be irrelevant to that query. For example, if the \[min, max] of col\_b in a tablet is \[100, 200], it cannot hold the value 10 and can be pruned.
**Primary Index Granule Pruning:** If col\_b is a column in the primary index, Firebolt can check for each granule whether it is guaranteed that no row qualifies. Note that this is slightly more difficult if the primary index has more than one column. For a primary index on (col\_a, col\_b), a granule with an index range of \[("a", 100), ("a", 150)] can be safely pruned. A sorted range of \[("a", 100), ("b", 150)], however, might contain the element ("b", 10) and cannot be ignored.
**Row Lookup Pruning:** A row is identified by the columns $tablet\_id and $tablet\_row\_number in Firebolt's managed storage. If you want to scan a specific set of rows from a table, Firebolt can also perform tablet pruning and granule pruning to only scan tablets and granules that contain rows you need. This is what Firebolt uses for late materialization.
## Late materialization [#late-materialization]
Late materialization is an optimization that tries to reduce the amount of scanned data as much as possible by delaying the scans of columns when possible. This optimization is made possible by storing data in a column-oriented format instead of a row-oriented format. Late materialization was researched at MIT over 18 years ago ([paper](https://dspace.mit.edu/bitstream/handle/1721.1/34929/MIT-CSAIL-TR-2006-078.pdf)) and has also recently been introduced to ClickHouse ([blog post](https://clickhouse.com/blog/clickhouse-gets-lazier-and-faster-introducing-lazy-materialization)).
### How late materialization is implemented in Firebolt [#how-late-materialization-is-implemented-in-firebolt]
In Firebolt, late materialization can be expressed using query plans that leverage existing pruning techniques. When Firebolt's optimizer recognizes that it can apply late materialization it chooses a plan that loads eligible columns later. Notably, there is no new implementation for late materialization in the runtime.
The general idea is to create two scans on the same table and join them on their row ids. The first scan only provides the columns needed to compute which rows should be in the result. The second scan provides all selected columns that are not part of the first scan already. The second scan can be heavily pruned using the IDs of rows that should be in the result. After joining the result of these two scans, there is another sort to restore the order that will be changed by the join.
Here are the plans for the example query:
```sql
-- Without late materialization
EXPLAIN
SELECT * FROM hits ORDER BY EventTime DESC LIMIT 10
WITH late_materialization_max_rows=0;
```
```text
[0] [Projection] hits.watchid ... \_[1] [Sort] OrderBy: [hits.eventtime Descending First] Limit:
[10] \_[2] [StoredTable] Name: "hits" [Types]: hits.watchid ...
```
The plan without late materialization is rather simple.
* \[2] scans all columns of the table hits,
* \[1] sorts and limits all rows, and
* \[0] returns the result.
Now see the same query with late materialization:
```sql
-- With late materialization
EXPLAIN
SELECT * FROM hits ORDER BY EventTime DESC LIMIT 10;
```
```javascript
[0] [Projection] hits.watchid ...
\_[1] [Sort] OrderBy: [hits.eventtime Descending First] Limit: [10]
\_[2] [Projection] hits.watchid ...
\_[3] [Join] Mode: Inner [(tuple(hits.$tablet_id, hits.$tablet_row_number) = tuple(hits.$tablet_id, hits.$tablet_row_number)), (hits.$tablet_id = hits.$tablet_id), (hits.$tablet_row_number = hits.$tablet_row_number)]
\_[4] [StoredTable] Name: "hits"
| [Types]: hits.watchid ...
\_[5] [Sort] OrderBy: [hits.eventtime Descending First] Limit: [10] Relaxed Limit
\_[6] [StoredTable] Name: "hits"
[Types]: hits.eventtime: timestamp not null, hits.$tablet_id: text not null, hits.$tablet_row_number: bigint not null
```
The plan with late materialization looks much more complicated. Here's the step-by-step breakdown:
* \[6] The first scan only includes the EventTime column and the row identifiers $tablet\_id and $tablet\_row\_number.
* \[5] The output of this scan is sorted and limited to 10 rows.
* \[4] The second scan provides all remaining columns, but is heavily pruned by the set of row ids from \[5]. This kind of pruning is also known as sideways information passing or semi join reduction.
* \[3]The scans are joined on row ids.
* \[2]Removes row id columns which are no longer needed.
* \[1]Sort because the join can change the sort order from \[5].
* \[0]The result is returned.
If you wonder whether your query already uses late materialization, you can check the EXPLAIN output of your query for this self join pattern. You can also manually control whether late materialization should be used if you want to override Firebolt's optimizer.
## How you can control late materialization in Firebolt [#how-you-can-control-late-materialization-in-firebolt]
Firebolt only applies late materialization for queries with a limit of 10 or smaller by default. If you choose a limit higher than that, you will observe that late materialization will not be triggered anymore:
```sql
SELECT * FROM hits ORDER BY EventTime DESC LIMIT 100;
-- Execution Time: 16 s, Scanned Data: 87 GB
```
If you want to override this decision and apply late materialization anyway, you can use the late\_materialization\_max\_rows [system setting](https://docs.firebolt.io/reference/system-settings#control-late-materialization):
```sql
SELECT * FROM hits ORDER BY EventTime DESC LIMIT 100
WITH late_materialization_max_rows=100;
-- Execution Time: 0.8 s, Scanned Data: 2 GB
```
You can either set this setting per query using the WITH clause as outlined in the example above or apply it to the whole session using a SET statement:
```sql
SET late_materialization_max_rows=100;
SELECT * FROM hits ORDER BY EventTime DESC LIMIT 100;
-- Execution Time: 0.8 s, Scanned Data: 2 GB
```
If you want to turn it off completely, you can set the value to 0:
```sql
SELECT * FROM hits ORDER BY EventTime DESC LIMIT 10
WITH late_materialization_max_rows=0;
-- Execution Time: 16 s, Scanned Data: 87 GB
```
## When should you use late materialization? [#when-should-you-use-late-materialization]
Why should you ever want to turn late materialization off? Because it is not equally effective in all cases. There are two main properties a query needs to have to benefit from late materialization:
* **Column Size:** The larger the byte size of the columns that are loaded late, the larger the improvement from late materialization. If there is a huge string column where each element is thousands of characters long on average, it will be much faster to only load parts of this data using effective pruning. In contrast, if all your strings are only one character long, there is not much to gain.
* **Row Count Difference:** Late materialization is only effective if it can prune data better than a regular scan. If your table has only 10 rows, you don't need late materialization. Note that Firebolt also pushes filters into the scan and prunes reads using these filters. If only 10 rows are returned from the scan, there is nothing to gain from applying late materialization. However, if the scan returns millions of rows, there is a lot to gain.
Running a more complex query plan with two scans and a join adds a small overhead in execution time. Whether late materialization is beneficial to a query depends on whether it prunes enough data to make up for this overhead. For some queries, this easily pays off and results in immense speedups. Other queries might not benefit at all and can even get a few milliseconds slower. This is why Firebolt applies late materialization only for limits of 10 and smaller by default, a very conservative threshold.
### An ideal use case [#an-ideal-use-case]
Here's another example where late materialization offers tremendous benefits. In this use case the session table contains a history of API requests for a cloud service. Say you are interested in investigating the requests with the worst response time on a specific day:
```sql
-- Without late materialization
SELECT ResponseTime, DebugTrace
FROM session
WHERE EventDate = '2025-09-01'
ORDER BY ResponseTime DESC LIMIT 10
WITH late_materialization_max_rows = 0;
-- Execution Time: 11 s, Scanned Data: 75 GB
-- With late materialization
SELECT ResponseTime, DebugTrace
FROM session
WHERE EventDate = '2025-09-01'
ORDER BY ResponseTime DESC LIMIT 10
WITH late_materialization_max_rows = 10;
-- Execution Time: 0.2 s, Scanned Data: 0.2 GB
```
A very useful column is DebugTrace, which stores execution log details. Elements in this column are usually very large. Also, there are many queries every day. Consequently, column size and row count difference are very large, and this query benefits strongly from late materialization.
Now, see examples where late materialization does not help:
### Too small columns [#too-small-columns]
```sql
-- Without late materialization
SELECT UserAgentMinor
FROM hits
ORDER BY EventTime LIMIT 10
WITH late_materialization_max_rows=0;
-- Execution Time: 0.4 s, Scanned Data: 1.9 GB
-- With late materialization
SELECT UserAgentMinor
FROM hits
ORDER BY EventTime LIMIT 10;
-- Execution Time: 0.4 s, Scanned Data: 1.5 GB
```
UserAgentMinor is a very short string column. There are usually only two characters per element. This is so little data that late materialization just does not pay off. Even though the version with late materialization scans a tiny bit less data, the execution time is equal for both versions.
### Too little difference in row counts [#too-little-difference-in-row-counts]
```sql
-- Without late materialization
SELECT *
FROM hits
WHERE EventTime <= '2013-07-01 20:00:00'
ORDER BY EventTime LIMIT 10
WITH late_materialization_max_rows=0;
-- Execution Time: 0.5 s, Scanned Data: 0.9 GB
-- With late materialization
SELECT *
FROM hits
WHERE EventTime <= '2013-07-01 20:00:00'
ORDER BY EventTime LIMIT 10;
-- Execution Time: 0.4 s, Scanned Data: 1.5 GB
```
Here you have an extremely selective filter. Only 39 elements have an event time that is as early as 2013-07-01 20:00:00. The row count difference is too small. Here you can also see that late materialization pruning is not perfect. The version with late materialization scans a little more data than the version with a single scan that prunes with the filter predicate.
## Conclusion [#conclusion]
Late materialization is an optimization that can offer tremendous speedups for top-K queries by reducing the amount of data that needs to be scanned. In Firebolt, you benefit from it automatically.
This blog post described the opportunity to prune data in top-K queries and how Firebolt uses it with late materialization. It explained how late materialization is implemented in Firebolt and how you can see it in the query plan. Finally, it showed in which cases late materialization can improve performance and when it doesn't help.
If you want to know more, check out the [documentation on late materialization](https://docs.firebolt.io/performance-and-observability/query-planning/late-materialization) or get [started with Firebolt](https://docs.firebolt.io/guides/getting-started) yourself.
# Querying Apache Iceberg with Sub-Second Performance (/blog/querying-apache-iceberg-with-sub-second-performance)
As the industry is moving toward open table formats, Apache Iceberg is rapidly gaining popularity as the preferred technology to manage massive datasets in data lakes. By bringing ACID transactions to data lakes, it enables multiple systems to concurrently and safely operate on the same data without vendor lock-in. For someone using Iceberg, this opens up new ways to optimize latency and cost. You can mix-and-match different query engines and choose the best one for each workload. Today, we are adding Firebolt to the list of tools to work with your Iceberg tables, releasing a **public preview** of native support for querying Iceberg tables with the extreme performance and efficiency you've come to expect from us.
Firebolt's native deployment on AWS enables seamless integration with AWS-based data lakes and Apache Iceberg tables stored in Amazon Simple Storage Service (Amazon S3), bringing enterprise-grade performance and reliability to open table format analytics.
Iceberg emerged as an open table format for large-scale batch processing. This means that it wasn't optimized for low latency. Yet being able to deliver low-latency analytics has become essential to serve many modern data applications. At the moment, this means that people have to ingest a copy of their Iceberg tables into low-latency query accelerators. This blog post gives a deep dive on the work we're doing at Firebolt to bring low latency and high performance directly on top of your Iceberg tables.
## The Elephant in the Room: Why is Low Latency on Iceberg So Hard? [#the-elephant-in-the-room-why-is-low-latency-on-iceberg-so-hard]
Querying Iceberg tables with the speed needed for interactive analytics or AI-driven exploration runs into a few common roadblocks:
**Metadata Overhead**: Iceberg's architecture involves a hierarchy of metadata: a catalog points to the current metadata file, which points to a manifest list, which in turn points to manifest files that finally track the actual data files. Each step in traversing this hierarchy, especially over a network to cloud storage, introduces latency. For large tables with many partitions and files, this can easily sum up to multiple seconds before even touching the data. But even for smaller tables, object storage latency and the sheer number of hops means this can easily add 0.5 seconds to every query.
**Data Access Latency**: Iceberg tables often reside in cloud object storage like Amazon S3, which has orders of magnitude higher latencies for accessing data compared to local storage. Reading numerous small files or even just the footers of larger Parquet files can quickly become a bottleneck.
**Query Engine Inefficiencies**: Traditional query engines might not be optimized for the fine-grained metadata understanding and aggressive caching required to mitigate these latencies effectively. Firebolt's AWS-native architecture addresses these challenges by leveraging Amazon Elastic Compute Cloud (Amazon EC2) instances with high-performance networking and intelligent caching strategies that minimize round-trips to Amazon S3, significantly reducing the latency bottlenecks inherent in cloud object storage access. Combine this with optimisations of how and when to deploy these resources to match the customer workload and you have a differentiated best in class solution.
## Firebolt's Approach: Engineered for Speed on Iceberg [#firebolts-approach-engineered-for-speed-on-iceberg]
We've engineered our Iceberg integration from the ground up to tackle these challenges head-on, enabling you to unlock the value in your Iceberg tables at blistering speeds. At the heart of this is our new `READ_ICEBERG` table-valued function (TVF).
```sql
SELECT *
FROM READ_ICEBERG(
LOCATION => my_iceberg_location,
MAX_STALENESS => INTERVAL '10 seconds'
)
LIMIT 100;
```
This function serves as your gateway to Iceberg data, allowing you to read Iceberg tables from file-based and REST catalogs (e.g., Snowflake Polaris) as well as the Databricks Unity Catalog. Firebolt recommends using a LOCATION object, which encapsulates Iceberg parameters and credentials, for secure, reusable access – more on that a bit later.
### Intelligent Metadata Management: Slashing Latency at the Source [#intelligent-metadata-management-slashing-latency-at-the-source]
Constantly re-fetching and re-parsing Iceberg metadata is a big source of slowdowns. Firebolt employs several strategies to get metadata access out of the critical query path:
* **In-Memory Caching of Iceberg Snapshots**: We aggressively cache Iceberg metadata(like metadata files, manifest lists, and manifest files). Not only do we cache the raw files on disk, but to avoid re-reading them over and over, we also cache the deserialized metadata in memory using our subresult reuse machinery, which transparently handles Iceberg in a fully transactional way. This means that after the initial read, subsequent queries can often find the metadata they need almost instantly, moving metadata resolution out of the query path and cutting latencies dramatically.
* **Configurable Data Freshness with MAX\_STALENESS**: For many interactive use cases, having data that is a few seconds or minutes out of date is perfectly acceptable if it means dramatically faster queries. The MAX\_STALENESS parameter allows you to define this tolerance (e.g., INTERVAL '30 seconds'). When specified, Firebolt can use a cached metadata file if it is within the allowed staleness, avoiding costly catalog checks. The default is 0 seconds, ensuring queries always see the latest snapshot if no staleness is specified.
* **Asynchronous Snapshot Refreshing (coming soon)**: To keep the metadata cache warm and up-to-date without impacting foreground query latency, Firebolt will soon asynchronously refresh snapshots before they expire.
### Accelerated Data Access: Bringing Data Closer, Faster [#accelerated-data-access-bringing-data-closer-faster]
Beyond metadata, accessing the actual Parquet data files efficiently is key:
* **Optimized Parquet Reading**: Firebolt is designed for efficient processing of Parquet files, the de-facto standard for columnar data in Iceberg tables. Our Parquet reader has fully decoupled (network & disk) I/O and file decoding, and only keeps as much data in memory as is required to achieve maximum throughput (provided that row group sizes are reasonable). It has an extensible architecture that allows us to bring exciting features previously only known on managed tables to Iceberg and Parquet workloads, with some launching today and more coming in future releases.
* **Multi-Tiered Caching**: Firebolt utilizes multi-tiered caching for data, spanning memory and local disk (SSD). This means frequently accessed data from your Iceberg tables can be served from much faster local tiers, significantly reducing the need to go to remote object storage for every request. Because the security of your data is paramount, cached data is tied to the credentials used for access. Data caching is also fully independent of MAX\_STALENESS and catalog caching, as data files for Iceberg are immutable and new snapshots typically reuse the vast majority of the previous snapshots' data files. Added some new rows? Firebolt will read those from object storage and serve the rest from cache.
### Smart Query Optimization: Leveraging Iceberg's Structure [#smart-query-optimization-leveraging-icebergs-structure]
Firebolt's query optimizer is now Iceberg-aware, using the rich metadata within Iceberg tables to make smarter decisions:
* **Subresult Caching:** Because Iceberg gives ACID guarantees, Iceberg queries can leverage Firebolt's [subresult and result caches (FireCache)](https://docs.firebolt.io/Overview/queries/understand-query-performance-subresult.html) just like managed tables can. If a portion of a query, e.g. a complex join build side, has been computed before and the underlying data (considering MAX\_STALENESS) hasn't changed, Firebolt can reuse this cached subresult. This is incredibly effective for dashboarding workloads or iterative query refinement, where queries often share common sub-structures. Of course, the same applies for Firebolt's query result cache. Combined, these are hugely powerful and far exceed what you could achieve with your own caching on top of Firebolt.
**Co-Located Joins and Aggregations:**
* When Iceberg tables share compatible partitioning schemes, Firebolt can leverage this information to perform co-located joins, eliminating data shuffling and dramatically speeding up join operations. For example, for an inner join between tables that are both bucketed on the join key, Firebolt can distribute entire partitions to nodes and run the join fully locally.
* Similar to joins, if an aggregation key matches the Iceberg table's partitioning key, Firebolt can optimize the aggregation process, potentially making local aggregations more efficient and even eliminating entire global aggregation stages.
* Use the enable\_iceberg\_partitioned\_scan setting to enable these optimizations when the number of partitions is high enough and their size is sufficiently balanced to make this worthwhile. In the future, we want to apply these optimizations automatically.
**Metadata Used In Query Planning:**
* Firebolt's optimizer applies its state-of-the-art join ordering algorithm on Iceberg tables. The optimizer gets Iceberg table row counts from the metadata in manifest files. The join ordering works for all scenarios, be it joins between Iceberg tables, or joins between Iceberg tables and managed tables.
* Iceberg table metadata like table row counts are also used in smart query rewrites. For example, a count(\*) query on an Iceberg table is answered purely using the metadata. No actual data file is scanned in this case.
* **File-Level Data Pruning**: Iceberg metadata often includes statistics about columns within each data file (e.g., min/max values). Firebolt leverages these statistics to perform aggressive data pruning. For instance, if a query filters the TPC-H orders table on o\_orderdate = '1998-01-01', Firebolt will only read data files whose order date ranges include this date, potentially skipping massive amounts of irrelevant data.
### Simplicity and Interoperability [#simplicity-and-interoperability]
Getting started is straightforward. Here's an example reading from a public Iceberg table in a file-based catalog on S3. **This is a real example you can try!**
```sql
SELECT *
FROM READ_ICEBERG(
URL => 's3://firebolt-publishing-public/help_center_assets/firebolt_sample_iceberg/tpch/iceberg/lineitem',
MAX_STALENESS => INTERVAL '30 seconds'
)
LIMIT 5;
```
Firebolt also supports reading from Iceberg REST catalogs, including the ability to query tables in Databricks Unity Catalog via its Iceberg REST API interface. You can create a LOCATION object to store the location and credential information, with either a generic REST catalog syntax, or with syntax specializations tailored for the type of catalog. For example, you can create a LOCATION for a table in Databricks Unity:
```sql
-- Example: Creating a Databricks Unity table LOCATION
CREATE LOCATION my_uc_location_name WITH
SOURCE = 'ICEBERG'
CATALOG = 'DATABRICKS_UNITY'
CATALOG_OPTIONS = (
WORKSPACE_INSTANCE = '.cloud.databricks.com'
CATALOG = 'my_uc_catalog_name'
SCHEMA = 'my_uc_schema_name'
TABLE = 'my_table_name'
)
CREDENTIALS = (
OAUTH_CLIENT_ID = ''
OAUTH_CLIENT_SECRET = ''
);
```
Which can be queried using the simple syntax:
```sql
SELECT *
FROM READ_ICEBERG(
LOCATION => 'my_uc_location_name',
MAX_STALENESS => INTERVAL '60 seconds'
)
```
For AWS customers, this seamless integration means you can leverage existing Amazon S3 buckets, and IAM permissions without complex data migration or credential management, enabling rapid deployment within your existing AWS environment.
## Current Support and Transparency [#current-support-and-transparency]
We are launching support for Iceberg tables with Parquet data files in S3 and support the most important features of versions 1 and 2 of the Apache Iceberg specification. It's important to be transparent about current capabilities. As of this release, the following are **not yet supported**:
* Row-level deletes (position or equality deletes).
* Schema evolution.
* Partition evolution.
* Reading past snapshots with time travel (but you can specify an older metadata file manually).
* Iceberg v3 features and data types.
It's also important to note that cross-region costs may be incurred based on the storage location of underlying Iceberg files.
## Conclusion: Unlock Your Iceberg Data with Firebolt [#conclusion-unlock-your-iceberg-data-with-firebolt]
Firebolt's new `READ_ICEBERG` capability does a lot of heavy lifting to provide low-latency access to your Iceberg tables. It leverages multiple different caching tiers, from MAX\_STALENESS to cache catalogs to completely transactionally caching the metadata lists and files on disk and their contents in memory, to cut metadata overhead to the absolute minimum. With Firebolt's subresult caching, commonly used join hash tables are transparently cached and reused while maintaining full transactional integrity. And our advanced query planner will make sure to optimize complex query plans on Iceberg.
We are confident that our approach delivers a significant leap in performance, and we're excited for you to experience it. We believe in building exceptional systems with an exceptional team, and telling this story. This public preview is just the beginning of our Iceberg journey, and we are committed to continuously enhancing our support and performance.
Built on AWS, Firebolt brings enterprise-grade performance to your Iceberg data lakes while maintaining the flexibility and cost-effectiveness that made you choose open table formats in the first place. Available through AWS Marketplace, Firebolt's innovation and engineering obsession for performance optimization provides a streamlined path to modernize your data analytics.
If you want to learn and hear more, tune in to a talk that Firebolt's VP of Engineering, Benjamin Wagner, gave at Data & AI Summit:
When you're done watching that, dive into the [documentation](https://docs.firebolt.io/reference-sql/functions-reference/table-valued/read_iceberg), try out the `READ_ICEBERG` function on your own tables, and let us know what you think! [Sign up now](https://go.firebolt.io/signup) for $200 in free credits and try it out – no card needed.
# Revolutionizing Data Governance with DataStrato’s Unified Open Source Approach (/blog/revolutionizing-data-governance-with-datastratos-unified-open-source-approach)
In this episode of The Data Engineering Show, the bros sit with Lisa Cao, Product Manager at DataStrato, to explore data catalogs and Apache Gravitino, a unified metadata lake used to manage access and perform data governance for all data sources. They discuss data catalogs and how they refine the data management process.
Listen on [Spotify](https://spoti.fi/3RLHWtD) or [Apple Podcasts](https://apple.co/425cAEs)
#### Episode Highlights [#episode-highlights]
##### What is Apache Gravitino? (01:24) [#what-is-apache-gravitino-0124]
Apache Gravitino is a meta-catalog that serves as a unified data governance and security layer used to manage different data systems. Lisa shares that Gravitino was the first to release an iceberg rest catalog and ended up open sourcing for the general community to use and as time passed, Polaris and Unity Catalog were also announced in open source. She highlights that although Gravitino, Polaris and Unity Catalog are very similar, Gravitino differs in that it is able to support multiple catalogs.
##### Unifying AI/ML and Big Data Stack (03:15) [#unifying-aiml-and-big-data-stack-0315]
One of the interesting things about Gravitino is that it offers more than just a catalog of data models and these model catalogs are the first step into looking at how to merge two worlds of AI and ML catalogs. Lisa shares the goal of effective management, that is, creating a system that can store and manage different types of data models, track changes to the models, and control access to the models.
##### Simplifying Data Governance (10:49) [#simplifying-data-governance-1049]
Think of Gravitino as a "traffic cop" that helps to manage and secure data from multiple sources. It is crucial to have a system that provides unified access control across all data sources, allowing teams to manage access and data governance so that ML teams don't have to worry about access. Lisa says that Apache Gravitino is the system that makes data accessible to different teams and users while making sure that it is secure and governed appropriately.
##### The Gravitino's Query Engine Solution (21:34) [#the-gravitinos-query-engine-solution-2134]
Every query engine has its own way of managing data, which makes it difficult to switch between engines - you have to reconfigure everything. Lisa highlights that Gravitino solves the problem by providing a single layer of data governance that works across multiple query engines.
##### Navigating the Fast-Paced World of Data Engineering (24:41) [#navigating-the-fast-paced-world-of-data-engineering-2441]
Lisa talks about how fast the data engineering space is moving and shares some insights to catching up;
* Don't try to learn everything at once.
* Don't get too deep into every tool
* Look for real-world adoption
She warns against the social media hype that can amplify the messaging around new tools, making it seem everyone is using it, when in reality, that can't be easily seen.
# Robust and efficient geospatial operations using snap rounding (Part III) (/blog/robust-and-efficient-geospatial-operations-using-snap-rounding-part-iii)
In parts I and II of this series of blog posts, we talked about the foundational building blocks of Firebolt's GEOGRAPHY data type. In part III, we will explore in more detail how Firebolt implements robust operations on geospatial data.
## What is snap rounding? [#what-is-snap-rounding]
One of the fundamental challenges we faced when implementing geospatial functions using the S2 Geometry library was the use of snap rounding. To understand snap rounding, let's first have a look at the problem it solves. For this example, we ignore the curvature of the Earth in order to keep it simple and understandable. Consider a line between the points at coordinates (0, 0.1) and (1, 0.3). We might have a point at coordinates (0.5, 0.2) and want to check whether it is on the line. Since the number 0.2 is not representable exactly as a [floating point number](https://standards.ieee.org/ieee/754/6210/) (and neither are 0.1 or 0.3), an exact calculation would always return that the point is not on the line. Such errors might also happen because of slight imprecisions in the data due to measurements or because points are constructed from other objects. For example, a point could have been constructed by intersecting two lines, which will introduce some inaccuracies due to floating point calculations. The S2 library goes even further: Because it cannot guarantee that an exact calculation always works, it will always return that the point is not on the line, except if it is exactly equal to one of the points used to define the line. Snap rounding is S2's solution to this problem: We "snap" two inputs together if they are less than a certain distance apart from each other (in Firebolt's case, we use 1 micron). During this process, we transform the line from (0, 0.1) to (1, 0.3) into a line that starts at (0, 0.1), passes through (0.5, 0.2) and ends at (1, 0.3). Now, (0.5, 0.2) is exactly the same as one of the points used to define the line, so we return that it lies on it.

## Mitigating the performance penalty of snap rounding [#mitigating-the-performance-penalty-of-snap-rounding]
While snap rounding is a great tool for building robust operations, it comes with a performance hit for actually performing snap rounding. S2 does not natively support a way to only perform snap rounding when it's required. So in order to get fast *and* robust operations, we need to take optimization into our own hands. We use several fast paths that are applied per tuple to circumvent performing the actual snap rounding.
Many of these checks use S2 cells and coverings. We already explained these in detail in part II. In short: S2 cells are a hierarchical partitioning of the Earth's surface, designed to represent spatial data at various levels of precision. A covering, in this context, refers to a set of S2 cells that together encompass a given region. One of the key advantages of using S2 cells is that they allow very efficient intersection and containment tests. Since each cell is represented by a single integer, determining whether one cell intersects with another can be done using simple integer comparisons, which are extremely fast.
One example is that if one of the inputs is constant in the query (for example when querying all points from a table that are inside a specific polygon), we compute a covering of S2 cells around the polygon. We can then perform a fast check whether a point from a table is within that covering. If it is not, we can immediately deduce that the point is not inside the polygon. The covering used here is chosen such that it includes everything that would be inside the polygon after snap rounding. Since the snap rounding distance of 1 microns is much smaller than even the most accurate S2 cells, this usually does not actually lead to a different covering than if we ignored snap rounding. It is still important to account for it because otherwise we might falsely determine that a point is outside of the polygon based on the covering even though it would be inside it after snap rounding.
In the example below, the point on the right is outside of the polygon's covering, so we can immediately return that it is outside of the polygon.

We use a similar technique for clearly positive cases: By computing an *interior* S2 cell covering, i.e. a set of S2 cells that are guaranteed to be completely inside a polygon, we can quickly determine that a point is inside the polygon.
The example below shows the same polygon and point set as the previous one. Here we see that the left-most point is inside the interior covering of the polygon so we can immediately return that it is inside the polygon.

In the examples above, we still have two points left that we need to check for containment in the polygon. The next optimization we apply here makes use of the rough covering consisting of only one S2 cell that is stored for each point (or any other GEOGRAPHY type, but we're focussing on points for these examples). This covering is configured such that it contains everything within the snap rounding distance around the point. For a containment query in a polygon, we can then check whether the polygon fully contains that cell. If the cell is contained, then snap rounding cannot change that fact so we can safely return that the polygon contains the point. If the cell is fully outside of the polygon, then the point cannot be snap rounded to be contained in the polygon, so we can safely return that as well. Only if the point is very close to the polygon's boundary, snap rounding will be applied and an exact containment test on the snap-rounded polygon and point will be performed. This is more expensive than the previous two checks which only used interior and exterior coverings of the polygon, but it is still cheaper than the exact containment test using snap rounding.
In the image below, we see the two remaining points from the example with their respective one-cell coverings. The top left point has a covering that is completely contained in the polygon, so we can return that it is contained in the polygon. The bottom right point's covering intersects the polygon, so we have to perform snap rounding only for this one remaining point.

## Handling degenerate shapes [#handling-degenerate-shapes]
Since snap rounding is also applied within a single shape, it can also lead to *degenerate* shapes: Consider a very narrow polygon in a diamond shape, where two of the opposite vertices are under 1 micrometer apart. During snap rounding, these two vertices are snapped together, so the result is a *degenerate* polygon that has an area of 0.

There are two approaches to handle such degeneracies: Dimension reduction where this polygon would be converted into a line string, or native support of degeneracies where this is treated as an infinitely small polygon (in terms of area). Since there are some functions that behave differently for polygons and linestrings (because they have different definitions of what their boundaries are), Firebolt uses S2's native support of degeneracies when these occur during snapping.
One example of a function that behaves differently depending on how degeneracies are handled, is ST*CONTAINS. In comparison to the very similar ST\_COVERS function, ST\_CONTAINS(A, B) only returns true if B is completely inside A, \_and* the interiors of A and B intersect. The interior of the degenerate polygon from the example above is empty, so it contains nothing according to the definition of ST*CONTAINS. However, if we reduce the polygon to a linestring, this changes: By definition, a line string's interior is the entire linestring except for the first and last vertex. So if we were to convert the polygon into a linestring, it would contain infinitely many points. Because of this subtlety, we recommend using ST\_COVERS instead of ST\_CONTAINS in most cases. It is easier to understand, less expensive to compute, and does the same as ST\_CONTAINS in \_almost* all cases.
This wraps up our three-part series of blog posts about geospatial support in Firebolt. We went from fundamentals in part I to data format and pruning in part II all the way down to the details of snap rounding and performance optimizations in part III. Hopefully you learned a thing or two about the fascinating engineering problems associated with geospatial support in a database. I would also like to thank Google and Eric Veach again for making S2 geometry available. It was an invaluable tool to build our geospatial support. Now that you know all about how it works, you can try out Firebolt's geospatial functionality yourself using our [demo project](https://github.com/firebolt-db/firebolt-demo/tree/main/geospatial) or have a first look using our [live-demo.](https://demo.docs.firebolt.io/geospatial/)
# Tech Stacks and Tradeoffs: Xudo's Founder on Picking the Right Tools for BI Success (/blog/tech-stacks-and-tradeoffs-xudos-founder-on-picking-the-right-tools-for-bi-success)
Wouter Trappers is the founder of Xudo and shares his slightly unconventional path from philosopher to data consultant and engineer with the Bros in this latest episode of The Data Engineering Show. Wouter's grounding in philosophy has proved to be a shaping influence on his approach to business intelligence, much more than just a software solution, for Wouter, BI is all about change management and aligning leadership with data projects.
Listen on [Spotify](https://spoti.fi/3OstS6P) or [Apple Podcasts](https://apple.co/3Va77Zj)
Intro/Outro - 00:00:04:
The Data Engineering Show is brought to you by Firebolt, the cloud data warehouse for low-latency analytics. Get $200 credits and start your free trial at firebolt.io.
Benjamin - 00:00:15: Hi everyone, and welcome back to The Data Engineering Show. Today we're super happy to have Wouter joining, who leads the data analytics consultancy called Xudo. Wouter, do you want to say a few intro words and introduce yourself to the audience?
Wouter - 00:00:31: Yeah, so my name is Wouter Trappers. I'm working from Ghent, Belgium, and I'm indeed the founder of Xudo. With Xudo, the intention is to guide people through their business intelligence journey. But myself, I'm a philosopher from background. So I really am an autodidact and I rolled into this profession from job to job and then learning every time a little bit more. So yeah, that's a bit my background.
Benjamin - 00:00:57: Backgrounds. So from philosopher to basically a BI and data analytics consultant, that's a super, super interesting journey. Maybe tell us a bit more about that in your background.
Wouter - 00:01:09: Yeah, so like you may expect, there's not really a job that really fits the philosophy backgrounds one-on-one. So after you study abroad, study like philosophy still needs to orient yourself a bit in the job markets. So first start, it's working as a teacher. But then I found out this was not really my thing. And I started working in the private sector with basic admin jobs. And then I discovered I had a knack for analysis and for Excel. And basically that started my data journey as maybe a lot of other people as well.
Eldad - 00:01:44: Benjamin, once upon a time, Excel was that thing that ignited all the passion on data and crunching and analyzing data. And it still is. So thanks for the intro.
Wouter - 00:01:55: I think there are still a lot of people who start with Excel, mainly in other jobs than IT. People in marketing or in finance or supply chain who discovered the power of data through Excel. So I still think it's still a gateway into data for a lot of people.
Benjamin - 00:02:13: Definitely. So then after you got hooked on data analytics through Excel, basically, tell us a bit about your journey that ultimately got you to then start your own company now.
Wouter - 00:02:23: Yes, I started indeed in Excel. And I started in a financial controlling department of a large retail store here in Belgium. And the European retail store is also active in the US. There I used Excel and then also Access to automate all the processes for controlling and reporting. And also using VBA. And then when the data sets became too large for Excel, getting into Access and learning about the databases, how you write SQL code, and then going from there. Once you know the front-end functions of Excel and the back-end SQL that's behind it, you can basically learn any BI tool on the market. So after that, I switched to a company that builds software for medical professions like pharmacies, general practitioners. And so on, everything except for hospitals. It's called Corillus. And there I was introduced to QlikView, now Qlik Sense, a product of the Qlik company, where I started applying the principles that I learned during my time working with Excel and Access, but then in a more professional business intelligence environment. There I automated all the BI flows for the different departments, like after-sales, customer care sales, finance, and HR. And then the people finally knew what was going on in their company. Yeah.
Eldad - 00:03:48: MS Access, VBA.
Wouter - 00:03:51: Yes.
Eldad - 00:03:52: Those are terms people mostly don't remember anymore. But let me tell you this. This is how data engineering started. Before VBA, we had to get out of our environment to write code. And VBA was that first programming language that was embedded within our business apps. So Excel, MS Access. Microsoft Word for some reason as well. So people took that skill set, Visual Basic, an amazing language, and started to apply it within their business environment. Therefore, the first generation of data engineering is born. And nobody called that data engineering back then. But I remember that period and I love that period. And yes, we've made great progress since then, but the foundation stayed the same.
Benjamin - 00:04:41: I always love this recurrent segment of the show. Where you explain to me who doesn't maybe know all of these things that people did back in the day.
Eldad - 00:04:50: Ah, and one last thing, QlikView, which turned Qlik Sense. I think it's a Swedish, originated in Sweden, one of the first BI companies that actually built a full stack. So you had the engine, in-memory engine, which was revolutionary from a slice and dice perspective. And you've had the UI in it. And you've had their own kind of VBA language, which was a bit different, but same concept. And it was a huge hitch back then. The reason I know it so well, because in my previous startup early on, we competed with QlikView. So QlikView was our nemesis and everything we did was compared to how the QlikView guys did. So thanks you for mentioning that. Great history.
Benjamin - 00:05:36: Also definitely a first on The Data Engineering. I never heard QlikView mentioned before. Jumping a couple of years into the present. Right? Tell us more about basically what types of challenges you're facing today and then helping companies solve.
Wouter - 00:05:49: Yeah. So I'm a sole consultant at Xudo. The idea first was to help people and companies to come up with a data strategy. In my definition, it is finding out how data can help you to improve your company by making more impact on revenue, by reducing costs and identifying redundancies or most important use case. In my view. For business intelligence is the peace of mind that you can trust the data that you're looking at to make decisions. So I thought I would build data strategy roadmaps and puts building blocks on the roadmap and then execute a roadmap. But unfortunately, I didn't find any customers for this offering. So I pivoted to concept enhance on working in the companies. So I usually have two types of companies I work for at the same time, one larger company that makes me money. And then a smaller company that I can really help build their data foundations from the ground up, and the larger companies, I take like role in the data teams of the company themselves. And in the smaller companies, I work with the tools that the company has at hand or I suggest my own tools that they can work with. And then I prefer the smaller customers. But of course, the bigger customers pay the bills.
Benjamin - 00:07:07: Right. This is actually super interesting and something we haven't explored that much on the show before. In many cases, we have like data engineers. Coming who work at companies that have like very established kind of data stacks, right? Ten different query engines, Iceberg, 5-BI tools and so on. Like when you're maybe smaller and also like, I don't know, like company from a more traditional sector kind of bootstrapping that data engineering or data warehousing stack. Tell us more about the types of challenges that are common there and what maybe the most common struggles are.
Wouter - 00:07:40: Yeah, I think if people are looking to start their data journey, they usually wind up in their research on the websites of the technology vendors and they are selling business intelligence projects. As a software that will solve their problems. But in fact, business intelligence It's always a change journey. So you have to approach it like this. If you do BI well, then you will use insights from the dashboards to improve your company, to improve the processes, and then also to maybe adapt the systems to those new processes. And then you go into a business intelligence cycle. So I think it's very important to make sure that the people who want to start using data are aware that if it's done well, it's always a transformational change for the organization. And that's something else than just implementing a software. So I put a lot of emphasis on change management as well. And also starting from the top leadership alignment on which are the KPIs of your business, which KPIs haven't you thought of yet, which are the obvious ones you want to tackle first. And then building kind of a roadmap starting there. But then also, of course, looking at the systems and the data that is in place. Where is the data? Can we Access the data? And then building the vision from the top and starting from the data at the bottom and meeting in the middle with nice dashboards to help the business move forward.
Benjamin - 00:09:08: At what point are you usually engaging with companies then, right? Like say they reach out to you if they're interested in your services. I guess at that point they already got into. Wanting to become data-driven and wanting some data solutions or is it in many cases also the chore actually the one pitching the company, hey, you're not using any BI tools or anything right now. Tell us more about that very early part of that process.
Wouter - 00:09:34: Most of those clients come inbound. So they find me and they already have an idea that they want to start using data. And usually they already have some kind of tool that they want to use as a source. And sometimes they also have some kind of like, one of these clients is called on a fine clever. It's like a nonprofit organization who helps people with disabilities, and they have a lot of data. They had done large CRM projects with a custom built application on top to help their coaches, help these people. And then they knew they could do more with their data and the tech stack they used to build the CRM and application was Zoho, it's a very modular program. It's an Indian company. And Zoho, also as an analytics module. So what they then did is they activate the analytics module. I studied it. I learned the new, the new module and I started building there. But of course, starting from the basics, building a data model, building an architecture, making sure that I could communicate what I was doing with it. Because the intention was to leave the company after a couple of months so that they could run it themselves and also extend. So a little bit hard, of course, because if you've never worked with tools like this can be daunting. So I basically taught them the basics of SQL. Then they use ChatGPT to write extensions to the code. And then once a year they call me in because they don't find the answer in ChatGPT. And then I solve their problems. For instance.
Benjamin - 00:11:08: That's actually super interesting how GenAI is already influencing workloads or especially for maybe less technical users who aren't as exposed to that technology.
Wouter - 00:11:18: Yeah. What I use GenAI myself, is to write simple standalone PowerShell scripts, for instance. Very powerful to do that. I haven't used it in larger deployments because it has to be integrated. It has to be persistent. I haven't used it myself that way. But for simple standalone scripts, it's very useful.
Benjamin - 00:11:38: Yeah. I feel the same way. I've done some benchmarking of Firebolt over the past couple of days and just bootstrapped some Python scripts. And it's really total, magic at this point. How easy it becomes, especially if it's closely integrated with your IDE. During my day to day job is mostly low level C++ development. There it's a bit less useful still in many cases, but especially for bootstrapping these standalone scripts, it's really crazy the amount of utility you can get from it.
Eldad - 00:12:05: It's also crazy how many tools, apps, and systems you can apply with GenAI.
Wouter - 00:12:12: Yeah.
Eldad - 00:12:12: Like you mentioned Zoho. Like I wouldn't have never imagined that. ChatGPT would be able to even remember that. And that's so widely used in very less technical environments. So now having a GenAI means we can start using a lot of what we already have. That's especially important for people and companies that have challenges catching up and making sure they're properly aligned on the right stack. So it's not that needed anymore. You can just use GenAI to cover the gaps and it's fascinating to see actually, to hear it happen in real life.
Wouter - 00:12:48: I think it's important in this case to understand that in Zoho you have different options to configure your setup. And I chose the option where you can write your own SQL so that you don't get locked into some type of a proprietary syntax of Zoho where there are a lot less experts who know how to work with this.
Benjamin - 00:13:08: One thing I'm super curious about is how do you think that actually changes these types of consulting projects, right? Because in the past you basically. At the end would have done a handover to their engineering team so they can solve their own problems. Now, in many cases, you're handing over to the AI who maintains things, does minor fixes themselves. Do you think that, for example, changes the way you would create like a knowledge base of the project before you give it into other hands? Like, give us your perspective on that. I'm super curious.
Wouter - 00:13:37: I think you still need a knowledge base to document things like architecture and high level approach. I think the GenAI in this case solves some minor type of. Issues they may have or the questions they may want to try to answer using their data and extending the queries a little bit. But I think in this case, the power of GenAI is to give the data environment in the hands of non-technical people who can think logically and who can sort of describe the prompts they need to get the codes they can use. So I think you mentioned handing it over to the engineering team, but this nonprofit doesn't have an engineering team. So I'm their engineering team. And I'm not interested, to be fair, to take up these less interesting support tasks so then they can work themselves. And then once a year I can help them do a review or put them on track a little bit.
Benjamin - 00:14:31: Right. Super interesting. So as you bootstrap then these data stacks, right? I think the company you talked about before already had like an existing solution where they had data in it. And then when you bootstrap the data warehousing stack you need to adjust to data. Is it, like, you have a blueprint for a data and BI stack that you can apply in many cases. Or is it really completely different, the actual solution you converge to from company to company?
Wouter - 00:15:01: The companies that engage me usually already have a tech stack. So I don't really have the opportunity to recommend the stack at that time. And I learned the environment of the customers.
Eldad - 00:15:12: Would you fire a customer or a prospect for having the wrong stack or having a stack that you're absolutely not willing to work with?
Wouter - 00:15:19: Yeah, I think if I'm not comfortable to work with, I will honestly admit it and probably the customer will not hire me. So that's the way it works.
Eldad - 00:15:27: So there's still the tooling fit between the consultant and the client and there needs to be a match. Not every tool we work with every consultant. And as you say, you will not start something with a stack that you don't trust, for example.
Wouter - 00:15:44: No, because I'm a philosopher in the end. I'm not that technical and I don't want to pretend that I'm something I'm not.
Benjamin - 00:15:51: I think after like a decade plus in the industry, at some point you are also a philosopher, but also a data engineer, a BI engineer. Super interesting. Cool. So when you then work with these larger clients, like how much of the problem basically translates? Like, do you feel like there's specific problems companies have that haven't done BI or data warehousing before? And then at the larger company, it's completely different? Or does it actually translate across company sizes?
Wouter - 00:16:25: I think the first point that I mentioned that BI projects have to be approached as change projects also translates to larger customers, where the issue usually is not that they don't have enough technology, but they have too much. So then they have to start pruning and putting emerging overlapping dashboards together or making sure that the message and the analysis that's possible in each dashboard is very clear. And then you go more down the governance track. But I think behind this thing I see at larger companies is still the same issue, namely that they are still approaching BI projects as a software project and not as a change project.
Benjamin - 00:17:09: One thing I'm curious about here that you mentioned is that data lineage and data quality is also something that really matters in these transformation processes, because obviously you want to make sure that the dashboard you look at, and make business decisions and change your strategy maybe as the right data. One thing that I find curious about is that if you think about tools for this, like Monte Carlo, which does data quality and these types of things, to me it always felt like something that actually happens quite far down your data journey. So you add your data stack and then at some point you have your first incidents, right? And you realize, oh damn, this dashboard I looked at actually represents something that doesn't reflect reality. It's interesting to me that this is something you already push for that early on in the process. It makes intuitively a lot of sense, but I would actually expect that for many companies, this is something that comes much later and not by the time they bootstrap their BI or data warehousing stack.
Wouter - 00:18:10: You mean governance?
Benjamin - 00:18:11: I mean like these kind of governance, lineage aspects, kind of data quality aspects and so on. I always felt like that might be more of a reaction to issues, you run into once you establish that the first time. And then for many companies, this is actually not something that they think about on day one.
Wouter - 00:18:29: Yeah, I think if you talk about data lineage, you can do it in two ways. Or you start building your data model where you draw your lines between your source, your intermediate layers, and then your visualization layers. And then you have some kind of plan upfront, and that's the data lineage that you then have to maintain when you start building it, and when you start expanding this environment. And in an ideal world, everyone documents everything and then you can find exactly what data is where and how it is used. Of course, in practice, it's usually not that clear. The people start building and then they have something that works and the business relies on it until there's a question, where is this data coming from? And nobody knows. And then you have to start documenting like the re-engineering where the data is coming from, what different steps are between the source and the visualization and can be extremely complex. So that's two ways I look at data lineage, one upfront and one after the fact. The second one is data lineage that I've done a couple of times for larger clients who have a lot of data, a lot of complex transformations, just to communicate the complexity and to also make them understand how difficult it is to change something and to build upon the existing data. And then, of course, you have tools to automate this kind of lineage, but in my experience, it's never really able to capture all the complexity that's going on. For instance, if you have a select star somewhere, I don't know any tool can go back in the code and find the exact fields that are used in the star, for instance.
Eldad - 00:20:14: It also depends on the source, right? Like, usually people think logs, data points that engineers generate while building apps. But most of the world still depends on a CRM, ERP, complex models, very business-driven vertical data sources. Federated sources. So a lot of stuff that many times we tend to forget. When we go with the kind of cloud native logs data. So I think modern lineage and quality is focusing on the latter, like on logs generated data, where as we build the apps and generate the logs, we make mistakes as engineers, and that needs to be fixed and that needs to be cataloged. But if you just go a few years back and look at existing data models, this is a whole different game. And that requires the human factor. And I connect to that a lot because people in companies spend fortunes of time and effort in building those things. The business runs on those things. And it's not easy to just go in and say, oh, let's just replace it. So it's like a surgery one little step at a time, trying to figure out how to do most impact with as less harm as possible. And again, it's also an educational and technological aspect, right? If you're born into data, into tech, from day zero, and you've went through the usual cycles and education, it's one thing, but it's a very different thing if you already have a huge business and you've operated for many years on previous generations of IT systems. So there's a lot of work ahead. And most companies need consultants, not just to consult, but to really guide them from one era to another era. There's a lot of philosophy in it. So a lot of human factor, a lot of convincing and reasoning, and it has nothing to do with technology.
Wouter - 00:22:13: Sometimes I try to make the connection between philosophy and data. And then the case that you would describe, Benjamin, is, for instance, a way of looking from philosophy through a lens of philosophy to this kind of data project. It's like you have to talk the same language. So in order to understand each other, often that's already an issue within large companies. You have different teams with different cultures. Sometimes the people who built the legacy systems are still there. They're completely talking a different language than the younger hires who are cloud native. And then already to recognize that this can be a problem, that there is a different culture and a different way of using words can help to unlock some difficulties that those companies have.
Benjamin - 00:23:00: So talking about cloud native hires, in how many cases are the companies you're working with actually on the cloud already? Like is most of the work you're doing on-prem? Or are these companies usually in the cloud already? And if they're not, is this like maybe also part of the journey?
Wouter - 00:23:16: Yeah, like the nonprofit I was talking about, they started in the cloud. Like SOA is a cloud platform. So they already started in the cloud. The larger companies usually started on-prem and then migrated some of their tools already to the cloud. Or maybe they are in the process of doing so. That are the two pieces I see at this time.
Benjamin - 00:23:37: Cool. Awesome, Wouter. Anything else at the moment, you see popping up a lot in the industry that you're passionate about, that you wanted to chat about today on the podcast.
Wouter - 00:23:48: I think that what I'm most excited about are the people who are on social media talking about going back to basic. A couple of years ago, everything wanted to do new things. But I think we have to understand that the old things were there for a reason. And I think people are rediscovering some of the basic concepts now. And I'm glad to see that.
Benjamin - 00:24:07: Nice. That's, by the way, a very recurring theme and I think kind of sentiment that a lot of people in the industry have.
Eldad - 00:24:14: Data is like fashion. It goes in cycles. We always get to the same starting point. In a good way, of course.
Benjamin - 00:24:20: Yeah. Very cool. Wouter, thank you so much for being on the show today. Seriously, this was a totally amazing perspective. We really never had a guest who works that much with companies in the early stages of their data journey. So it was amazing to have you on. All of the best with the clients you're working with. And thank you so much for joining.
Eldad - 00:24:40: Thank you.
Wouter - 00:24:41: My pleasure. Thanks for having me.
Intro/Outro - 00:24:44:
The Data Engineering Show is brought to you by Firebolt, the cloud data warehouse for low-latency analytics. Get $200 credits and start your free trial at firebolt.io.
# Technical Deep Dive: Automated Column Statistics (/blog/technical-deep-dive-automated-column-statistics)
## Abstract [#abstract]
We just introduced a new feature called automated column statistics in Firebolt. Automated column statistics provide the query optimizer with up-to-date information on statistical properties of the values in a column. These statistics can help to give better plans and ultimately achieve better performance without having to modify query texts. Once enabled, the statistics are collected transparently and kept up to date incrementally during table updates. In this blog post, we explain how to use the new feature by using a simple example query and then go into detail on how automated column statistics work under the hood.
## Background: Join Ordering [#background-join-ordering]
In the example query in this blog post, we mainly look at how to improve join performance using automated column statistics. To join two tables, Firebolt builds a hash table of one of the two inputs, with the hash key being the join key. It then streams through the other input, probes the hash table with each key, and optionally outputs the results. In order to keep the hash table fast and in-memory, the query optimizer tries to make sure that we build it from the smaller input. If it chooses the wrong ordering, the hash table might not even fit into the RAM of the engine and needs to be spilled to the SSD. Join ordering has a major influence on the performance of SQL queries. For deciding on the join ordering, the optimizer tries to estimate the cardinality of the two inputs. These inputs can potentially be filtered, aggregated or are the output of another join. Therefore, without knowing the distribution of the data, it is hard for the optimizer to give good estimates even when the input is a table with a simple filter on top. To illustrate this, consider the following example.
## Example [#example]
Let's say that we want to perform analytical queries for a video streaming platform. Our fictional platform has 100 million subscribers (Netflix has about 277 million) and has a catalog of 100'000 titles (for some countries, Netflix has 10k). So our database schema could be as follows.
```sql
CREATE TABLE movies (title TEXT, duration INT); -- 100'000 rows
CREATE TABLE views (movie_title TEXT, customer_name TEXT, device_type TEXT); -- 5 billion rows
```
The column `device_type` has 4 distinct values: `TV`, `Phone`, `PC`, and `Other`. The customer\_name column has 100 million distinct values due to the same number of customers. The `movie_title` column matches one of the titles in the `movies` table, so it has 100'000 distinct values. We are now interested in the following two queries:
```sql
Query 1: SELECT avg(duration) FROM views JOIN movies ON views.movie_title = movies.title WHERE views.device_type = 'TV';
Query 2: SELECT avg(duration) FROM views JOIN movies ON views.movie_title = movies.title WHERE views.customer_name = 'Alice';
```
In query 1, we want to know the average duration of movies watched on a TV screen. In query 2, we might want to build a customer-facing dashboard showing statistics about the movie watching behavior of individual customers. In this example, we look at the average duration of shows that Alice watches. Note that both queries follow the exact same syntactic structure, they just filter based on different columns. As a human, it is easy to guess that the selectivities of those two filters are vastly different: A decent percentage of all users likely watches their shows on a TV, so the number of rows left after the filter is likely still very large. In contrast, customer `Alice` likely only watched a couple of movies, and certainly not the entire movie library. Therefore, we have two syntactically equivalent queries where we still expect widely different cardinalities of intermediate results.
So what does Firebolt's query planner do? Without more information, it cannot do better than guessing. Firebolt does this by using a basic textbook estimation (inverse square root) for selectivity of the filter. We end up with the following plan structure for both of the queries. In Firebolt's explain output, the probe side of the join is printed first, and then the build side below. The optimizer puts the `views` table on the build side because it assumes that the filter is very selective and removes most rows.
```sql
QUERY PLAN TEXT
[0] [Projection] avg_0
\_[1] [Aggregate] GroupBy: [] Aggregates: [avg_0: avg(movies.duration)]
| [Types]: avg_0: double precision null
\_[2] [Projection] movies.duration
\_[3] [Join] Mode: Inner [(movies.title = views.movie_title)]
\_[4] [StoredTable] Name: "movies"
| [Types]: movies.title: text null, movies.duration: integer null
\_[5] [Projection] views.movie_title
\_[6] [Filter] (views.device_type = 'TV')
\_[7] [StoredTable] Name: "views"
[Types]: views.movie_title: text null, views.device_type: text null
```
For query 2, the filter for Alice is indeed very selective. After it, not many views are left, and putting the filtered views table on the build side is an optimal plan. For query 1, however, this is a bad plan. The filter by TV lets a significant number of rows pass through. We end up putting the large views table on the build side of the join. This hash table might end up being larger than the available RAM in the engine, forcing it to spill to the SSD. Because the optimizer does not know the semantics of those columns, we need to give it a better idea of the distribution of values in the columns.
## How To Guide The Optimizer [#how-to-guide-the-optimizer]
In Firebolt, there are several ways to guide the optimizer. If you need absolute control, you can enable the [user-guided optimizer mode](https://docs.firebolt.io/performance-and-observability/query-planning/user-guided-mode). In this mode, the optimizer does not perform join ordering, aggregate push-down, or redundant join removal. Only the user is responsible for writing efficient queries. This is useful for highly tuned heavy-duty queries.
The second option is to use the [no\_join\_ordering](https://docs.firebolt.io/performance-and-observability/query-planning/query-hints/no-join-ordering) hint. With this hint, the optimizer performs as its optimizations but skips the join ordering. The user is responsible for writing the joins in an efficient way.
However, those two options assume that the query texts can be changed and that users know how to tune them. Especially in the case of queries generated by an LLM, one cannot rely on them being well tuned. Additionally, for large queries, it might not even be feasible for a human to properly tune the join order. Firebolt now offers two additional ways to aid the optimizer: **automated column statistics** and **history based statistics**! In this blog post, we focus on automated column statistics, while a future blog post will go into detail about history based statistics.
## Automated Column Statistics [#automated-column-statistics]
Coming back to our example of the video streaming platform, how can we improve the join ordering? With automated column statistics, this is now extremely easy. Just run "alter table … add statistics" and you are ready to go. In our example above, this could look like the following.
```sql
ALTER TABLE views ADD STATISTICS (device_type) TYPE ndistinct;
ALTER TABLE views ADD STATISTICS (customer_name) TYPE ndistinct;
```
Before we dive into the details of what this does under the hood, let us look at the effect that this has on the join order. As mentioned above, for query 2 (filtering by customer\_name), the plan is already optimal. It does not change when enabling automated column statistics. However, for query 1 (filtering by device\_type), we now get the following changed query plan. The optimizer now knows that filtering for views made on a TV reduces the cardinality, but not as much as it initially assumed when it had no data on the distribution of values in the column. As you can see in the following explain output, the large views table is now on the probe side (top) of the join.
```sql
[0] [Projection] avg_0
\_[1] [Aggregate] GroupBy: [] Aggregates: [avg_0: avg(movies.duration)]
| [Types]: avg_0: double precision null
\_[2] [Projection] movies.duration
\_[3] [Join] Mode: Inner [(views.movie_title = movies.title)]
\_[4] [Projection] views.movie_title
| \_[5] [Filter] (views.device_type = 'TV')
| \_[6] [StoredTable] Name: "views"
| [Types]: views.movie_title: text null, views.device_type: text null
\_[7] [StoredTable] Name: "movies"
[Types]: movies.title: text null, movies.duration: integer null
```
This switched join order has a massive effect on the query performance. We performed a short experiment on a single-node engine, with result caching disabled. In query 2 (filtering by customer\_name), the plan already was optimal, and we do not see a change in the query time. This shows that the inference of automated column statistics does not slow down the query planning process. However, for query 1 (filtering by device\_type) we get massively improved performance due to the swapped join order. Note that this experiment was conducted on a developer workstation with an i9-14900K processor and not on our cloud offering. Nevertheless, Firebolt conveniently handles the 5 billion row views table in just a couple of seconds.
| Query | Time without ACS | Time with ACS | Improvement |
| --------------------------------- | ---------------- | ------------- | ----------- |
| 1: views.device\_type = 'TV' | 23.8 s | 7.8 s | 3x 🚀 |
| 2: views.customer\_name = 'Alice' | 2.4 s | 2.4 s | 1x ✅ |
With just these two simple alter table statements, we massively improved the performance of our example query. Additionally, all future queries can profit from the additional statistics – no changes to the query texts are required. Firebolt keeps the statistics up to date transparently and transactionally consistent.
## Statistics Collection [#statistics-collection]
In the following sections, we look under the hood of automated column statistics. We look at how Firebolt collects and reads the statistics that are so easy to add to a table. We also look at the query plans and cardinality estimates in more detail to understand the choices that our optimizer makes during query planning.
Automated column statistics are built on top of Firebolt's powerful [aggregating indexes](https://docs.firebolt.io/overview/indexes/aggregating-index). These indexes can pre-compute arbitrary aggregations on top of your tables and keep them up to date transparently during inserts and updates. Our aggregating indexes are transactionally consistent across all engines, so automated column statistics inherit this property. If you modify a table to collect automated column statistics, Firebolt automatically creates a system-managed aggregating index for maintaining the statistics. By default, it collects the distinct count of a column. However, you can specify other statistics you want to collect. You can spot these aggregating indexes in the show indexes command, indicated by the created\_by=system column:
```sql
show indexes;
index_name | table_name | type | created_by | expression
----------------+----------------+-------------+-------------+-------------------------------------------------------
views__STATS | views | aggregating | SYSTEM | [counting_hll_count_distinct("device_type"),count(*)]
views__STATS_1 | views | aggregating | SYSTEM | [counting_hll_count_distinct("customer"),count(*)]
```
As you can see, the index uses the [counting\_hll\_count\_distinct](https://docs.firebolt.io/reference-sql/functions-reference/aggregation/counting-hll-count-distinct) aggregate function. The function provides an approximate distinct count but in contrast to the [hll\_count\_distinct](https://docs.firebolt.io/reference-sql/functions-reference/aggregation/hll-count-distinct) function, it supports deletions. We built this aggregate function specifically for automated column statistics, but it is useful for all aggregating indexes when there are frequent deletions. We now look at the differences between them.
**HyperLogLog Sketches.** Our approximate distinct counts use a simple and elegant probabilistic data structure called [HyperLogLog (HLL) sketches](https://algo.inria.fr/flajolet/Publications/FlFuGaMe07.pdf). It builds on the fact that hash values of a good hash function can be considered random. They then have a probability of 2^-k to have the leading k bits set to 0. In other words, if we hash distinct keys, a share of 2^-k is expected to have k leading 0 bits. Conversely, if we hash all keys in a set and find a maximum of k leading 0-bits, we can approximate the number of distinct input keys by 2^k. In practice, it is more stable to hash each key to one of multiple buckets (using the first bits of the hash value) and counting the maximum leading zeroes (of the remaining hash) for each bucket individually. Finally, the maximum stored in all buckets can then be combined for example using the harmonic mean. Because we only need to store very small counters storing numbers smaller than 64 per bucket, this is very space efficient and fast to evaluate.
**Counting HyperLogLog Sketches.** A disadvantage of HLL sketches is that their aggregate state does not support deletions. Removing a value that has the same number of leading 0-bits as its bucket does not mean that no other values had the same hash. Therefore, if we used classical HLL sketches in an aggregating index, after each deletion we would need a full table scan to keep the index up to date. A solution for this problem are [*counting* HLL sketches](https://www.sciencedirect.com/science/article/pii/0022000085900418), introduced in 1985. Instead of just remembering the maximum leading zeroes of hash values, counting sketches keep counters of how often it saw each number of leading zeroes. To avoid this increasing the space consumption of aggregate states significantly, we use an enhancement [presented at the CIDR 2019](https://15799.courses.cs.cmu.edu/spring2025/papers/12-costmodel/p23-freitag-cidr19.pdf) conference. The idea is to count the first 128 values exactly. If the counter passes this value, it continues counting the logarithm of stored keys and updates the value probabilistically. This way, we only need one byte for each counter. While this still increases the space consumption compared to basic HLL sketches, it enables fast deletions without additional table scans. In order to make automated column statistics work well with deletions, we automatically create *counting* HLL sketches when adding statistics collection to a table.
## Statistics Inference [#statistics-inference]
To make the statistics available to the Firebolt optimizer, we extract all tables that a query references and detect whether they have automated column statistics enabled. If they have, we internally run a simple "select counting\_hll\_count\_distinct(column) from table" query. Because we know that an aggregating index with this function exists, we can rely on our automatic aggregating index matching to efficiently calculate the result without a full table scan. By confirming the size of the aggregating index from our metadata, we ensure to only run inference if it will be fast, given that this adds to the planning time. Additionally, we cache the results locally on each node. In our experiments, the statistics inference usually takes less than 20 ms when not cached. At the same time, they can potentially reduce the overall query time massively, for example due to better join ordering.
**Reading User-Created Indexes.** In addition to creating the aggregating indexes under the hood through "alter table … add statistics", we can also infer statistics from existing user-created aggregating indexes. To read those during query planning, enable [the corresponding query-level setting](https://docs.firebolt.io/performance-and-observability/query-planning/automated-column-statistics/user-created-indexes) by running "set infer\_statistics\_from\_indexes=true". On top of indexes directly using an aggregation that counts distinct values, we can infer distinct counts from additional indexes. For example, from an aggregating index defined as "on views (device\_type, count(\*))", we can derive the number of distinct device\_type values from the number of groups in the index. Reading user-created aggregating indexes can also help to assess the effect of enabling automated column statistics for the entire table without modifying the table definition.
## Cardinality Estimation in Firebolt [#cardinality-estimation-in-firebolt]
If you want to understand the reasoning behind decisions taken by the optimizer, you can use explain(statistics). Let us first look at a query plan when no automated column statistics are available. The optimizer knows the number of rows in each input table, but all other numbers are estimates. In our example queries (this is the same for query 1 and 2), the selectivity of the filter is most interesting. Here, the planner uses a textbook estimate for the filter selectivity: 1/sqrt(input rows). This means we choose the filtered views table as the build side (bottom input) of the join and the movies table as the probe side (top input). For query 2, this is optimal, but for query 1, it is not.
```sql
QUERY PLAN TEXT
[0] [Projection] avg_0
| [Logical Profile]: [est. #rows=1, column profiles={[avg_0: #distinct=1]}, source: estimated]
\_[1] [Aggregate] GroupBy: [] Aggregates: [avg_0: avg(movies.duration)]
| [Types]: avg_0: double precision null
| [Logical Profile]: [est. #rows=1, column profiles={[avg_0: #distinct=1]}, source: estimated]
\_[2] [Projection] movies.duration
| [Logical Profile]: [est. #rows=6.48341e+07, source: estimated]
\_[3] [Join] Mode: Inner [(movies.title = views.movie_title)]
| [Logical Profile]: [est. #rows=6.48341e+07, source: estimated]
\_[4] [StoredTable] Name: "movies"
| [Types]: movies.title: text null, movies.duration: integer null
| [Logical Profile]: [est. #rows=100000, source: metadata]
\_[5] [Projection] views.movie_title
| [Logical Profile]: [est. #rows=70711, source: estimated]
\_[6] [Filter] (views.device_type = 'TV')
| [Logical Profile]: [est. #rows=70711, column profiles={[views.device_type: #distinct=1]}, source: estimated]
\_[7] [StoredTable] Name: "views"
[Types]: views.movie_title: text null, views.device_type: text null
[Logical Profile]: [est. #rows=5e+09, source: metadata]
```
After altering the table to add ndistinct statistics to both columns, we now get the following plan for the query views.customer\_name = 'Alice'. The order that was guessed above is already optimal for this query. As you can see, our optimizer detects that and keeps the join order. In the explain(statistics) output, you can see that the optimizer now has the distinct count available. We assume uniform distribution across the distinct values, so our estimated number of rows after the filter is 50. This aligns well with our example: while in reality there will be some difference in viewing behavior between customers, we do not expect that individual customers watch multiple orders of magnitude more than other customers. In the future, we plan to support detecting non-uniform distributions, see later section on future steps.
```sql
QUERY PLAN TEXT
[0] [Projection] avg_0
| [Logical Profile]: [est. #rows=1, column profiles={[avg_0: #distinct=1]}, source: estimated]
\_[1] [Aggregate] GroupBy: [] Aggregates: [avg_0: avg(movies.duration)]
| [Types]: avg_0: double precision null
| [Logical Profile]: [est. #rows=1, column profiles={[avg_0: #distinct=1]}, source: estimated]
\_[2] [Projection] movies.duration
| [Logical Profile]: [est. #rows=67235, source: estimated]
\_[3] [Join] Mode: Inner [(movies.title = views.movie_title)]
| [Logical Profile]: [est. #rows=67235, source: estimated]
\_[4] [StoredTable] Name: "movies"
| [Types]: movies.title: text null, movies.duration: integer null
| [Logical Profile]: [est. #rows=100000, source: history]
\_[5] [Projection] views.movie_title
| [Logical Profile]: [est. #rows=50, source: estimated]
\_[6] [Filter] (views.customer = 'Alice')
| [Logical Profile]: [est. #rows=50, column profiles={[views.customer: #distinct=1]}, source: estimated]
\_[7] [StoredTable] Name: "views"
[Types]: views.movie_title: text null, views.customer: text null
[Logical Profile]: [est. #rows=5e+09, column profiles={[views.customer: #distinct=1.00193e+08]}, source: automated column statistics]
```
The situation is more interesting for the second query with the filter views.device\_type = 'TV'. With statistics, we now get a swapped join order there, making the plan significantly faster to execute. Still assuming uniform distribution of the values, we now estimate 1.25e9 rows after the filter, making it much larger than the number of movies. We again assume that all distinct values appear equally often. This is of course a crude approximation, but it is the best that any query planner can do when no further statistics are available. As mentioned above, we plan to add histogram sketches in the future to capture any bias in the value distribution.
```sql
QUERY PLAN TEXT
[0] [Projection] avg_0
| [Logical Profile]: [est. #rows=1, column profiles={[avg_0: #distinct=1]}, source: estimated]
\_[1] [Aggregate] GroupBy: [] Aggregates: [avg_0: avg(movies.duration)]
| [Types]: avg_0: double precision null
| [Logical Profile]: [est. #rows=1, column profiles={[avg_0: #distinct=1]}, source: estimated]
\_[2] [Projection] movies.duration
| [Logical Profile]: [est. #rows=8.00787e+11, source: estimated]
\_[3] [Join] Mode: Inner [(views.movie_title = movies.title)]
| [Logical Profile]: [est. #rows=8.00787e+11, source: estimated]
\_[4] [Projection] views.movie_title
| | [Logical Profile]: [est. #rows=1.25e+09, source: estimated]
| \_[5] [Filter] (views.device_type = 'TV')
| | [Logical Profile]: [est. #rows=1.25e+09, column profiles={[views.device_type: #distinct=1]}, source: estimated]
| \_[6] [StoredTable] Name: "views"
| [Types]: views.movie_title: text null, views.device_type: text null
| [Logical Profile]: [est. #rows=5e+09, column profiles={[views.device_type: #distinct=4]}, source: automated column statistics]
\_[7] [StoredTable] Name: "movies"
[Types]: movies.title: text null, movies.duration: integer null
[Logical Profile]: [est. #rows=100000, source: history]
```
As you can see in the examples, automated column statistics give the planner better insight into the data distribution. This can lead to much better plans and better performance. These lead to better performance without having to manually tweak your queries.
## Future Steps [#future-steps]
### Histogram Sketches [#histogram-sketches]
Fundamentally, query optimizers cannot know all properties of the input data without running the query. Therefore, they always have to rely on estimates. Currently, if the Firebolt optimizer knows the distinct count of a column, it assumes that the values are uniformly distributed. If we do not have more data, this is a sensible assumption. However, looking at our example queries above, the number of users watching movies might be skewed: We might have more users who watch their movies on their TV than on the phone. In the future, we plan to support histogram sketches of the data, which can reflect a possible bias in the distribution of column values.
### History-Based Statistics [#history-based-statistics]
A future blog post will cover [history-based statistics](https://docs.firebolt.io/overview/queries/understand-query-performance-hbs). By recording example queries, the optimizer can then learn cardinalities and base its estimates on them.
### Apache Iceberg Tables [#apache-iceberg-tables]
[Apache Iceberg](https://iceberg.apache.org/) is an open table format that gains traction because of its wide adoption in different systems. Due to the Iceberg format, one can freely choose a query engine without having to worry about data migration. The Iceberg format supports specifying statistics about the stored data, including distinct counts. However, whether these are available depends on the writer. Like for [many other writer improvements covered previously on our blog](https://www.firebolt.io/blog/unlocking-faster-iceberg-queries-the-writer-optimizations-you-are-missing), one cannot rely on the writer providing these.
Here at Firebolt, we embrace the Iceberg format and handle it as a first-class citizen in our database engine. We will soon release aggregating indexes on Iceberg tables. With this, Firebolt manages the indexes for your existing Iceberg catalog. This means that we also immediately get support for automated column statistics. In the near future, you will therefore be able to enable automatic statistics collection for Iceberg tables, and profit from better statistics in the optimizer, eventually leading to more informed decisions and better plans.
## Conclusion [#conclusion]
Firebolt can now transparently collect and access statistics about your data, built on top of aggregating indexes. These statistics help the optimizer to make more informed decisions, which can lead to better query plans and therefore significant performance improvements for your queries. Automated column statistics are available now, and the feature will soon be compatible with Apache Iceberg tables. [Get started](https://docs.firebolt.io/performance-and-observability/query-planning/automated-column-statistics) with automated column statistics now!
# Technical Deep Dive: Efficient and ACID Compliant Vector Search Indexes in Firebolt (/blog/technical-deep-dive-efficient-and-acid-compliant-vector-search-indexes-in-firebolt)
## Introduction [#introduction]
Modern AI applications heavily rely on semantic search through embeddings: High-dimensional numerical vectors that represent entities such as text, images, and audio. Efficiently searching these embeddings for the "top K closest points" in a multi-dimensional vector space, known as *vector search* or *similarity search*, is crucial. This deep dive explores how Firebolt implements native vector search indexing and search capabilities to enable these demanding AI workloads.
## The Need for Vector Search in Firebolt [#the-need-for-vector-search-in-firebolt]
For a long time, most of the vector search use cases powered by Firebolt involved moderately sized datasets, on the order of a few million embeddings. In these scenarios, computing exact distances between the query's target vector and all stored embeddings through a full table scan was both simple and fast enough. There was no pressing need for specialized indexing or approximate methods.
That changed when the [Similarweb](https://www.similarweb.com/) team brought us a very different challenge. Their workload involved running a semantic similarity search over a table containing hundreds of millions of embeddings. For them, query latencies of 20–30 seconds with brute-force nearest neighbor search were unacceptable. They needed sub-second search times to support interactive, production-grade applications, which is exactly what Firebolt is built for.
This is precisely the kind of scenario that traditional full-scan approaches cannot handle efficiently. At this scale, computing exact distances for every query quickly becomes prohibitively expensive. To meet these new requirements, Firebolt introduced native vector search indexing based on Approximate Nearest Neighbor (ANN) techniques, using the [Hierarchical Navigable Small World (HNSW)](https://arxiv.org/pdf/1603.09320) graph technique. ANN methods trade ingest performance and a small amount of accuracy for dramatically improved lookup performance, a trade-off that is often more than acceptable in real-world, analytical AI applications.
With Firebolt's vector search indexes, Similarweb was able to improve their production performance by 100x, reducing query execution times from tens of seconds to around 300 ms (99th percentile), enabling fast and scalable semantic search on massive datasets, all within the same compute cluster that powers their analytical workloads.
## How Vector Search Indexes Are Used in Firebolt [#how-vector-search-indexes-are-used-in-firebolt]
For a full rundown of how to use vector search indexes in Firebolt, see our [documentation](https://docs.firebolt.io/overview/indexes/vector-search-index#vector-search-index). Here's an example usage that we will use throughout this blog post.
Let's assume we have a table documents that stores embeddings as an array of real numbers as well as some id for these embeddings:
```sql
CREATE TABLE documents
(id int, embedding array(real not null) not null);
```
We can now [create a vector search index](https://docs.firebolt.io/reference-sql/commands/data-definition/create-vector-search-index#create-vector-search-index) using HNSW for these 256-dimension embeddings with the following query:
```sql
CREATE INDEX documents_vec_index
ON documents
USING HNSW ( embedding vector_cosine_ops )
WITH ( dimension = 256 );
```
Assuming the embeddings were generated using the AWS Bedrock amazon.titan-embed-text-v2:0 model, we can now search for the 10 documents that are semantically most similar to the search term "low latency analytics" using the [vector\_search table-valued function](https://docs.firebolt.io/reference-sql/functions-reference/vector/vector-search#vector_search) like this:
```sql
SELECT
*
FROM
vector_search (
INDEX documents_vec_index,
target_vector => (
SELECT
AI_EMBED_TEXT (
MODEL => 'amazon.titan-embed-text-v2:0',
INPUT_TEXT => 'low latency analytics',
LOCATION => 'bedrock_location_object',
DIMENSION => 256
)
),
top_k => 10
);
```
We are using the AI\_EMBED\_TEXT function to obtain a vector embedding for our search term in this example, but you could also get it using any other subquery on your database or simply pass it in as a constant value. The result of the vector\_search table-valued function is (at most) top\_k rows with the same columns as the original documents table. So the query above will return columns id and embedding.
## High-Level Architecture [#high-level-architecture]
Let's have a look at how vector search indexes are implemented in Firebolt at a high level. Each [tablet](https://docs.firebolt.io/overview/data-management#creating-tables) within Firebolt maintains its own dedicated vector index file for every vector search index defined on the table, this makes the indexes fully transactional and ACID compliant. These files each represent an HNSW index created using the [USearch library](https://github.com/unum-cloud/usearch). For every vector, Firebolt stores the row number of the corresponding row in the tablet inside the USearch index. This association between the vector's position in the index and its row number, together with the tablet identifier, provides a precise mapping from entries in the vector index files back to the actual table rows.
When a vector search query is executed, Firebolt initiates searches across each tablet's vector index file using USearch. This process identifies the relevant row numbers corresponding to the closest vectors. More precisely, when searching for the top K closest vectors to a target vector, this will return up to K results per tablet. So we need to condense this down to the overall top K closest vectors using an ORDER BY … LIMIT K operator in the execution pipeline. Once these row numbers are obtained, Firebolt performs an optimized base-table scan of these rows. Critically, this scan is not a full table scan; instead, it selectively reads only the rows identified by the vector index, significantly reducing the amount of data processed.
The following image illustrates the difference between vector search using a vector search index and brute-force vector search without an index. Because the index structure is optimized for this specific task and doesn't need to inspect each vector stored in the index, searching the vector search indexes is orders of magnitude faster than a full table scan. The table scan performed afterwards only needs to load a fraction of the data stored in the table because we know their precise location from the vector index.

## Caching and In-Memory Processing of Vector Search Indexes [#caching-and-in-memory-processing-of-vector-search-indexes]
With the architectural foundations in place, let's look at some of the performance characteristics of these indexes and how Firebolt manages them at runtime.
The performance of vector search depends not just on the algorithm, but also on how efficiently the system can access the index structure. Firebolt's vector search indexes are built using the USearch library, which is optimized for in-memory access. By default, Firebolt loads the full vector index files into memory when they are first queried and keeps them cached for subsequent searches. This strategy, controlled via the load\_strategy = in\_memory setting, delivers the best possible query latency once loaded: lookups happen entirely in RAM, and subsequent vector searches avoid any disk I/O overhead. Firebolt Compute Clusters that have enough memory to hold all relevant indexes can achieve highly predictable and consistently low latencies.
However, not all workloads or compute cluster configurations allow indexes to remain resident in memory. For those with tighter memory budgets, Firebolt supports a disk-backed mode via load\_strategy = disk. In this mode, index files are memory-mapped, which means the operating system decides when pages are loaded into memory and when they are evicted. This significantly reduces memory requirements, but at a cost: if a needed portion of the index is not available as a mapped in-memory page, it must be read from SSD again. As a result, queries may exhibit higher latency and less predictable performance compared to fully in-memory operations. This can happen for both recurring target vectors (the pages could have been evicted) and never-before-seen target vectors (part of the index structure has never been loaded)
In practice, users often choose the strategy based on workload characteristics. Latency-sensitive or high-throughput workloads typically run on compute clusters sized to keep all indexes in memory, ensuring optimal speed. Less critical or exploratory workloads can trade off some performance for cost efficiency by using disk-backed mode. The section on Best Practices for High-Performance Vector Indexing explains how to monitor cache behavior and configure Firebolt accordingly.
## Maintaining Vector Search Indexes: ACID Compliance and Speed [#maintaining-vector-search-indexes-acid-compliance-and-speed]
Some systems, including specialized vector databases, trade off transactional guarantees to simplify index maintenance, which can lead to inconsistencies between the data and the index. Firebolt takes a different approach: vector search indexes are fully ACID-compliant and transactionally consistent with the underlying table data. This ensures that searches always reflect the exact state of the database at the time of the query, even under concurrent writes or failures.
A key element of this design is Firebolt's tablet-based storage model. Inserts always create new tablets, and each tablet has its own corresponding vector index file. When data is inserted, Firebolt builds the vector index for the new tablet as part of the same transaction that writes the base table data. Once the transaction commits, the new data and its index appear together in the same snapshot. Deletes are handled through a separate per-tablet delete file. Vector search operations transparently consult this file to exclude deleted rows, maintaining correct results without requiring index rebuilds.
When ingesting data in bulk, Firebolt first writes the incoming data into intermediate tablets. These tablets are then merged into larger, finalized tablets before being uploaded to cloud storage and committed. Vector indexes are only built for the finalized tablets, which avoids the overhead of repeatedly building indexes on temporary intermediates while still preserving full transactional consistency.
In some cases, tablets may exist without an associated vector index file, for example, when an index is created after data has already been inserted. When this happens, users can bring the table fully up to date by running VACUUM (REINDEX=TRUE), which identifies any tablets without vector index files and builds the missing index files. This operation is fully transactional: it integrates cleanly with concurrent workloads and ensures that the index state always corresponds exactly to a well-defined snapshot of the table. For an explanation of how Firebolt handles tablets without index files during searches, see the section on handling unindexed tablets.
Firebolt's internal bookkeeping ensures that the presence or absence of index files is always tracked precisely. Queries run against a coherent view of which tablets are indexed and which are not, even as data is inserted, deleted, merged, or reindexed in parallel. As a result, the system consistently returns correct results and maintains strong transactional semantics, without requiring users to manage index consistency manually or forfeit data consistency.
## Best Practices for High-Performance Vector Indexing [#best-practices-for-high-performance-vector-indexing]
Getting the best out of Firebolt's vector search indexes comes down to two main aspects: making sure the index files fit in memory, and structuring the table in a way that minimizes the amount of data scanned after the index lookup.
The first step is to ensure that your engine has enough memory to keep the vector index files fully in memory. You can check the size of each index through the uncompressed\_bytes column in information\_schema.indexes:
```sql
SELECT index_name, uncompressed_bytes, number_of_tablets
FROM information_schema.indexes
WHERE table_name = 'documents' and index_type = 'vector_search';
```
Compare the reported index sizes with the defined cache size limit for vector search indexes.
```sql
SELECT sum(pool_size_limit_bytes) AS vector_index_cache_size_bytes
FROM information_schema.engine_caches
WHERE pool = 'vector_index' and cluster_ordinal = 1;
```
Firebolt lets you configure how much of that memory should be used for caching vector index files using the engine-level VECTOR\_INDEX\_CACHE\_MEMORY\_FRACTION configuration:
```sql
ALTER ENGINE my_engine
SET VECTOR_INDEX_CACHE_MEMORY_FRACTION = 0.6;
-- up to 60% of the each node's memory can be used for caching vector search indexes
```
If an index is too large to fit into memory after appropriate engine sizing, the next lever to consider is adjusting the HNSW index parameters to reduce its memory footprint. The M parameter controls the connectivity of the graph used for nearest neighbor search: lowering M reduces memory consumption and index size, at the cost of slightly lower recall. Using a different QUANTIZATION in the index can also substantially reduce its size by storing compressed vector representations, with a small decrease in accuracy. These adjustments give you flexibility to fit indexes for large tables into memory by trading off memory requirements for accuracy.
Once the engine and indexes are properly sized and configured, you should verify that they are actually being served from memory. Run a representative query with EXPLAIN (ANALYZE) and inspect the metrics for the read\_top\_k\_closest\_vectors node. After a warm-up run, the metric should show:
index files loaded from disk: 0/\
Any non-zero values indicate that some index files are being read from disk, which will cause higher and less predictable query latencies.
Warming up the engine is a critical step to getting predictable performance. Run any vector search on the index to load the index files into memory:
```sql
SELECT *
FROM vector_search(
INDEX documents_vec_index,
target_vector => ,
top_k => 1
);
```
To ensure that the base table data is cached on the engine's local SSD, perform a full table scan using a CHECKSUM query on the relevant columns:
```sql
SELECT CHECKSUM(*)
FROM documents;
```
This is intentionally a heavy operation — it ensures that all relevant data is read and cached locally, which pays off in lower latency for subsequent vector searches.
Another important factor is [index granularity](https://docs.firebolt.io/overview/indexes/primary-index#advanced-option%3A-index-granularity), which determines how many rows are stored per granule in the table. This directly affects how much data Firebolt reads when scanning the base table rows identified by the vector search index. Lowering the granularity means that less data is read for each matching row, which can significantly reduce I/O for selective top-K queries. The default granularity is 8192 rows, but for workloads that typically retrieve only a small number of rows, using a granularity of 128 can improve performance. You set this when creating the table:
```sql
CREATE TABLE documents (
id INT,
embedding ARRAY(REAL NOT NULL) NOT NULL
)
WITH ( index_granularity = 128 );
```
Finally, consider how many tablets your table has. Each tablet corresponds to one vector index file, and each file carries some amount of latency overhead during searches. Consolidating data into fewer, larger tablets – using [VACUUM](https://docs.firebolt.io/reference-sql/commands/data-management/vacuum) reduces the number of index files that need to be searched, which lowers per-query overhead and improves overall performance.
## A Detailed View on Vector Search Query Plans [#a-detailed-view-on-vector-search-query-plans]
When looking at a query plan for vector search for the first time, you might get a bit overwhelmed, so let's slowly build it up and understand what each step does. We are going to look at the plan for the following query using the table and index created in the examples above:
```sql
SELECT *
FROM vector_search(
INDEX documents_vec_index,
target_vector => ,
top_k => 10
);
```
We will go over the query plan with a simplified visualization as well as the output of EXPLAIN (PHYSICAL) on a single-node engine so you can reference this when running your own queries.
#### The Core: Efficiently Obtaining Results Using Vector Search Indexes [#the-core-efficiently-obtaining-results-using-vector-search-indexes]
Let's first look at a simplified version of the sub-plan of how we get tablet-id and the row numbers for the top 10 closest vectors to our target vector (we have removed some parts that will become important later). It is generally a good idea to read these query plans in the way data flows through them: From bottom to top.
```sql
[12] [Sort] OrderBy: [read_top_k_closest_vectors.distance Ascending Last] Limit: [10]
\_[13] [TableFuncScan] read_top_k_closest_vectors.tablet_id: $0.tablet_id, read_top_k_closest_vectors.tablet_row_number: $0.tablet_row_number, read_top_k_closest_vectors.distance: $0.distance
| $0 = read_top_k_closest_vectors(index => documents_vec_index, target_vector => target_vector, top_k => 10, ef_search => 64, load_strategy => in_memory, tablets => tablet)
| [Types]: read_top_k_closest_vectors.tablet_id: text not null, read_top_k_closest_vectors.tablet_row_number: bigint not null, read_top_k_closest_vectors.distance: double precision not null
\_[14] [Projection] tablet, target_vector: ARRAY
| [Types]: target_vector: array(double precision not null) null
\_[15] ...
\_[16] ...
\_[17] [TableFuncScan] tablet: $0.tablet
$0 = list_tablets(table_name => documents)
[Types]: tablet: tablet not null
```
We start in node \[17] by listing all the tablets that are part of the documents table. After doing some things that will become important later, we add the target vector using a projection in node \[14] (if the target vector is not a constant value, this might be replaced by a cross join with the subquery computing the target vector). Node \[13] does the actual work: It is a TableFuncScan using the read\_top\_k\_closest\_vectors table-valued function (TVF). The read\_top\_k\_closest\_vectors TVF is an internal function that takes tablets and a target vector as input, reads the physical index files for each tablet, and searches them by the target vector. As output, we get up to top\_k rows for each tablet with columns tablet\_id and tablet\_row\_number to identify the rows in the base table, as well as distance, which is the distance of the corresponding vector in the table to the target vector. This distance is then used in node \[12] to distill the top\_k rows per tablet down to the overall top\_k rows by ordering by the distance and keeping only the top\_k closest rows. Conceptually, our plan now looks like this:

Next, we need to read the corresponding rows from the actual table:
```shell
[5] [Join] Mode: Semi [(tuple_0 = tuple_1)]
\_[6] [Projection] documents.id, documents.embedding, documents.$tablet_id, documents.$tablet_row_number, tuple_0: tuple(documents.$tablet_id, documents.$tablet_row_number)
| | [Types]: tuple_0: tuple(text not null, bigint not null) not null
| \_[7] [TableFuncScan] documents.id: $0.id, documents.embedding: $0.embedding, documents.$tablet_id: $0.$tablet_id, documents.$tablet_row_number: $0.$tablet_row_number
| | $0 = read_tablets(table_name => documents, tablet, tuple(documents.$tablet_id, documents.$tablet_row_number) in (SUBQUERY{0}))
| | [Types]: documents.id: integer null, documents.embedding: array(real not null) not null, documents.$tablet_id: text not null, documents.$tablet_row_number: bigint not null
| \_[8] [TableFuncScan] tablet: $0.tablet
| $0 = list_tablets(table_name => documents)
| [Types]: tablet: tablet not null
\_[9] ...
\_[10] ...
\_[11] [Projection] tuple_1: tuple(read_top_k_closest_vectors.tablet_id, read_top_k_closest_vectors.tablet_row_number)
| [Types]: tuple_1: tuple(text not null, bigint not null) not null
\_[12] [Sort] OrderBy: [read_top_k_closest_vectors.distance Ascending Last, read_top_k_closest_vectors.tablet_id Ascending Last, read_top_k_closest_vectors.tablet_row_number Ascending Last] Limit: [10]
```
We continue with the same node \[12] from above, which returns the tablet\_id and tablet\_row\_number of the rows we want to read. In node \[5], this result is used as the right side of a left semi-join, where the left side is a table scan of the documents table (nodes \[7] and \[8] are how table scans are represented in Firebolt's physical query plans). The join condition requires that the $tablet\_id and $tablet\_row\_number from the documents table equal the values returned by the read\_top\_k\_closest\_vectors TVF. (The fact that this is modeled as a tuple in nodes \[11] and \[6] is an implementation detail that might change in the future.)
With this plan, we now have access to all columns of the base table for the exact rows returned by the vector search index:

Very observant readers might have noticed the in (SUBQUERY{0}) in node \[7]. This is what allows us to scan only the rows returned by the read\_top\_k\_closest\_vectors TVF: we pass the result of node \[9] (which contains the tablet\_id and tablet\_row\_number columns) into the scan and use it for pruning. In the literature, this technique is often called sideways information passing or semi-join reduction:
```sql
SUBQUERY{0}:
[23] ...
\_[24] [Projection] tuple_1
\_Recurring Node --> [9]
```
Our visualized plan now looks like this:

Because of this optimization, you would see something like granules: 10/1000000 for read\_tablets in node \[7] in the output of EXPLAIN(ANALYZE). This means that out of 1,000,000 granules in the table, only 10 were actually scanned—all thanks to read\_top\_k\_closest\_vectors doing the heavy lifting and semi-join reduction allowing us to prune the actual table scan.
This covers the core of how vector index queries are evaluated in Firebolt: we use the read\_top\_k\_closest\_vectors TVF and a \[Sort] OrderBy: \[distance] Limit: \[10] operator to obtain the tablet IDs and row numbers of the 10 closest vectors to the target vector. We then use a semi-join with semi-join reduction to selectively scan the actual rows from the base table.
#### What About the Rest? Getting Results From Unindexed Tablets [#what-about-the-rest-getting-results-from-unindexed-tablets]
But looking at the query plan, there is more going on. This is because Firebolt might produce tablets that don't have vector index files. This could either be because the index was created on an already populated table without triggering index recreation or because Firebolt decided that a tablet is too small to be worth indexing. To see how this is handled, let's look at nodes \[15] and \[16] that we left out before:
```shell
[13] [TableFuncScan] read_top_k_closest_vectors.tablet_id: $0.tablet_id, read_top_k_closest_vectors.tablet_row_number: $0.tablet_row_number, read_top_k_closest_vectors.distance: $0.distance
| $0 = read_top_k_closest_vectors(index => documents_vec_index, target_vector => target_vector, top_k => 10, ef_search => 64, load_strategy => in_memory, tablets => tablet)
| [Types]: read_top_k_closest_vectors.tablet_id: text not null, read_top_k_closest_vectors.tablet_row_number: bigint not null, read_top_k_closest_vectors.distance: double precision not null
\_[14] [Projection] tablet, target_vector: ARRAY
| [Types]: target_vector: array(double precision not null) null
\_[15] [Filter] contains(vector_search_index_files_in_tablet, 'documents_vec_index_547a5e0e_5bd0_4c98_8ec8_0874b248d511')
\_[16] [Projection] tablet, vector_search_index_files_in_tablet: json_pointer_extract_keys(vector_search_index_sizes, '')
| [Types]: vector_search_index_files_in_tablet: array(text not null) null
\_[17] [TableFuncScan] tablet: $0.tablet, vector_search_index_sizes: $0.vector_search_index_sizes
$0 = list_tablets(table_name => documents)
[Types]: tablet: tablet not null, vector_search_index_sizes: text null
```
Instead of just the tablets, we also retrieve the list of all vector index files (as well as their sizes) for each tablet from list\_tablets in node \[17]. Because this column is a JSON object, we use the json\_pointer\_extract\_keys function in node \[16] to extract an array containing the names of all vector search indexes that have a file for the given tablet and then in node \[15] filter out any tablet that doesn't have a file for the vector index documents\_vec\_index\_547a5e0e\_5bd0\_4c98\_8ec8\_0874b248d51 which is the internal name Firebolt used for the documents\_vec\_index vector index. Note that we only need to perform this json operation once per tablet of which there are orders of magnitude fewer than rows in the table.
Conceptually, this is simply a filter that only keeps tablets that have a vector index file for the vector search index used in our query:

This means that everything we looked at in the previous section only obtains results from tablets that actually have a vector index file. We now need to make sure we also include rows from tablets without vector index files in the result.
The following subplan scans all tablets not covered by a vector index. The pattern is very similar to the one above, but this time we filter out any rows that *do* have a vector index file in node \[20] and scan only the tablets that don't in node \[18].
```shell
[18] [TableFuncScan] documents.id: $0.id, documents.embedding: $0.embedding, documents.$tablet_id: $0.$tablet_id, documents.$tablet_row_number: $0.$tablet_row_number
| $0 = read_tablets(table_name => documents, tablet)
| [Types]: documents.id: integer null, documents.embedding: array(real not null) not null, documents.$tablet_id: text not null, documents.$tablet_row_number: bigint not null
\_[19] [Projection] tablet
\_[20] [Filter] (not contains(vector_search_index_files_in_tablet, 'documents_vec_index_547a5e0e_5bd0_4c98_8ec8_0874b248d511'))
\_[21] [Projection] tablet, vector_search_index_files_in_tablet: json_pointer_extract_keys(vector_search_index_sizes, '')
| [Types]: vector_search_index_files_in_tablet: array(text not null) null
\_[22] [TableFuncScan] tablet: $0.tablet, vector_search_index_sizes: $0.vector_search_index_sizes
$0 = list_tablets(table_name => documents)
[Types]: tablet: tablet not null, vector_search_index_sizes: text null
```

After this, we have all rows from tablets *without* vector index files after node \[18] as well as up to 10 (the top*k parameter we used in the vector\_search TVF) results from tablets \_with* vector index files after the join in node \[5]. We now need to retrieve the overall result from these two:
```shell
[0] [Projection] documents.id, documents.embedding
\_[1] [Sort] OrderBy: [distance_to_target Ascending Last, documents.$tablet_id Ascending Last, documents.$tablet_row_number Ascending Last] Offset: [0] Limit: [10]
\_[2] [Projection] documents.id, documents.embedding, documents.$tablet_id, documents.$tablet_row_number, distance_to_target: vector_cosine_distance(cast(documents.embedding as array(double precision)), ARRAY )
| [Types]: distance_to_target: double precision null
\_[3] [Union]
```
Node \[3] combines the results of nodes \[5] and \[18]. In node \[2], we calculate the distance of each scanned row to the target vector using the vector\_cosine\_distance function (because the documents\_vec\_index index was created with the vector\_cosine\_ops operation). In node \[1], we order by this distance and keep only the top 10 overall results. Finally, node \[0] projects only the columns selected in the query (i.e., all columns of the documents table).

Putting it all together, this is what the complete query plan looks like:
```shell
[0] [Projection] documents.id, documents.embedding
\_[1] [Sort] OrderBy: [distance_to_target Ascending Last, documents.$tablet_id Ascending Last, documents.$tablet_row_number Ascending Last] Offset: [0] Limit: [10]
\_[2] [Projection] documents.id, documents.embedding, documents.$tablet_id, documents.$tablet_row_number, distance_to_target: vector_cosine_distance(cast(documents.embedding as array(double precision)), ARRAY )
| [Types]: distance_to_target: double precision null
\_[3] [Union]
\_[4] [Projection] documents.id, documents.embedding, documents.$tablet_id, documents.$tablet_row_number
| \_[5] [Join] Mode: Semi [(tuple_0 = tuple_1)]
| \_[6] [Projection] documents.id, documents.embedding, documents.$tablet_id, documents.$tablet_row_number, tuple_0: tuple(documents.$tablet_id, documents.$tablet_row_number)
| | | [Types]: tuple_0: tuple(text not null, bigint not null) not null
| | \_[7] [TableFuncScan] documents.id: $0.id, documents.embedding: $0.embedding, documents.$tablet_id: $0.$tablet_id, documents.$tablet_row_number: $0.$tablet_row_number
| | | $0 = read_tablets(table_name => documents, tablet, tuple(documents.$tablet_id, documents.$tablet_row_number) in (SUBQUERY{0}))
| | | [Types]: documents.id: integer null, documents.embedding: array(real not null) not null, documents.$tablet_id: text not null, documents.$tablet_row_number: bigint not null
| | \_[8] [TableFuncScan] tablet: $0.tablet
| | $0 = list_tablets(table_name => documents)
| | [Types]: tablet: tablet not null
| \_[9] [Shuffle] Loopback with disjoint readers
| | [Affinity]: single node
| \_[10] [Aggregate partial] GroupBy: [tuple_1] Aggregates: []
| \_[11] [Projection] tuple_1: tuple(read_top_k_closest_vectors.tablet_id, read_top_k_closest_vectors.tablet_row_number)
| | [Types]: tuple_1: tuple(text not null, bigint not null) not null
| \_[12] [Sort] OrderBy: [read_top_k_closest_vectors.distance Ascending Last, read_top_k_closest_vectors.tablet_id Ascending Last, read_top_k_closest_vectors.tablet_row_number Ascending Last] Limit: [10]
| \_[13] [TableFuncScan] read_top_k_closest_vectors.tablet_id: $0.tablet_id, read_top_k_closest_vectors.tablet_row_number: $0.tablet_row_number, read_top_k_closest_vectors.distance: $0.distance
| | $0 = read_top_k_closest_vectors(index => documents_vec_index, target_vector => target_vector, top_k => 10, ef_search => 64, load_strategy => in_memory, tablets => tablet)
| | [Types]: read_top_k_closest_vectors.tablet_id: text not null, read_top_k_closest_vectors.tablet_row_number: bigint not null, read_top_k_closest_vectors.distance: double precision not null
| \_[14] [Projection] tablet, target_vector: ARRAY
| | [Types]: target_vector: array(double precision not null) null
| \_[15] [Filter] contains(vector_search_index_files_in_tablet, 'documents_vec_index_547a5e0e_5bd0_4c98_8ec8_0874b248d511')
| \_[16] [Projection] tablet, vector_search_index_files_in_tablet: json_pointer_extract_keys(vector_search_index_sizes, '')
| | [Types]: vector_search_index_files_in_tablet: array(text not null) null
| \_[17] [TableFuncScan] tablet: $0.tablet, vector_search_index_sizes: $0.vector_search_index_sizes
| $0 = list_tablets(table_name => documents)
| [Types]: tablet: tablet not null, vector_search_index_sizes: text null
\_[18] [TableFuncScan] documents.id: $0.id, documents.embedding: $0.embedding, documents.$tablet_id: $0.$tablet_id, documents.$tablet_row_number: $0.$tablet_row_number
| $0 = read_tablets(table_name => documents, tablet)
| [Types]: documents.id: integer null, documents.embedding: array(real not null) not null, documents.$tablet_id: text not null, documents.$tablet_row_number: bigint not null
\_[19] [Projection] tablet
\_[20] [Filter] (not contains(vector_search_index_files_in_tablet, 'documents_vec_index_547a5e0e_5bd0_4c98_8ec8_0874b248d511'))
\_[21] [Projection] tablet, vector_search_index_files_in_tablet: json_pointer_extract_keys(vector_search_index_sizes, '')
| [Types]: vector_search_index_files_in_tablet: array(text not null) null
\_[22] [TableFuncScan] tablet: $0.tablet, vector_search_index_sizes: $0.vector_search_index_sizes
$0 = list_tablets(table_name => documents)
[Types]: tablet: tablet not null, vector_search_index_sizes: text null
SUBQUERY{0}:
[23] [AggregateState] GroupBy: [] Aggregates: [make_pruning_set_0: make_pruning_set(tuple_1, 5000000)]
| [Types]: make_pruning_set_0: aggregatefunction(make_pruning_set(5000000), tuple(text not null, bigint not null) not null) not null
\_[24] [Projection] tuple_1
\_Recurring Node --> [9]
```
Or, using our simplified visualization:

## Benchmarks [#benchmarks]
To evaluate the performance of our vector index under a realistic load, we ran a sustained workload against the unchanged production dataset from one of our customers who uses Firebolt's vector index already in production on the same data.
### Setup [#setup]
The base table's size is 1.64TiB (compressed) and consists of 435,488,243 embeddings of dimension 1024.
```sql
CREATE TABLE embeddings (
id TEXT NOT NULL,
embedding ARRAY(REAL NOT NULL) NOT NULL
) PRIMARY INDEX id;
```
The vector search index on this table is 691GiB in size.
```sql
CREATE INDEX embeddings_vec_index
ON embeddings
USING hnsw (embedding vector_cosine_ops)
WITH
dimension=1024
quantization='i8'
connectivity=16 --default
ef_construction=128 --default
;
```
The engine we used was an 8-node, single-cluster, storage-optimized (SO) one in the AWS region us-east-1. This engine configuration was chosen regarding its capability of caching the whole index, which is required for optimal performance (the vector index cache size limit was increased).
```sql
CREATE ENGINE
WITH
nodes=8
clusters=1
family=SO
type=M
-- 75% of the engine's memory can be used for caching vector index
vector_index_cache_memory_fraction=0.75
-- 5% of the engine's memory can be used for caching query (sub-)results
query_cache_memory_fraction=0.05
;
```
As a client, we used an EC2 instance in the same region to minimize network overhead.
### Result [#result]
To evaluate our index, we look at two important metrics: latency at QPS and recall. We did not tune the index for any of these metrics and used the parameters' default values. Furthermore, the build time of the index and warming up the engine are not part of the presented numbers. Before we ran the simulation, we warmed up the engine to ensure (1) all base table data is cached on the engine's SSD and (2) the vector index is fully loaded into the engine's memory
For measuring latency at QPS, we selected 10,000 random embeddings from the base table as target vectors and simulated a steady, multi-user workload. We simulated 3000 users, where each one executes a vector search query and then waits \[1, 30] seconds before sending the next query.
We purposely disabled result caching to simulate a real-world use case where each target vector originates from an embedding model, which makes it likely unique and therefore unusable for result-reusage anyway.
```sql
SELECT keyword from vector_search (
INDEX embeddings_vec_index,
target_vector => ,
top_k => 10,
ef_search => 16 -- default value
) with (enable_subresult_cache = false);
```
At 3000 users with thinking time, we plateaued at a client-side-measured QPS of 190-200 with a p99 latency of 390ms (no query failures occurred). These numbers are measured on the client-side and therefore contain: HTTP request/response, (potential) query queueing, and query processing latency.
* p50: 170ms
* p60: 180ms
* P70: 200ms
* p80: 220ms
* p90: 260ms
* p95: 290ms
* p99: 390ms
Running this kind of concurrent workload without the index is not possible. On the same engine, a vector search without the index (single user, without any other queries running on the system) takes 80-90 seconds, while the search via the index is 1000-fold faster at 80ms.
The same (untuned) index achieved a recall\@10 of 87% (measured across 1500 different, randomly selected target vectors).
### Cost [#cost]
Firebolt's concept of compute-storage separation allows us to calculate the exact cost for this benchmark. Therefore, we simply have to add up the S3 storage cost with the compute cost for every minute the engine is up and running.
* Base table with full precision data vectors: 1.64TiB \~= $41.47 / month
* Index with quantized data vectors: 691GiB \~= $17.06 / month
* Engine: 8M SO $44.80 / hour
Firebolt also supports running large indexes on much smaller engines without the need to load them into memory fully. This option is available as a parameter directly in the TVF itself with `load_strategy='disk'`. We'll try to keep as much as possible in memory on a best-effort basis. Through parallel load on the engine, memory and cache eviction can happen, which will cause unpredictable latency for this strategy, though.
```sql
SELECT keyword from vector_search (
INDEX embeddings_vec_index,
target_vector => ,
top_k => 10,
ef_search => 16,
load_strategy => 'disk'
) with (enable_subresult_cache = false);
```
## Wrapping Up: Vector Search in Firebolt [#wrapping-up-vector-search-in-firebolt]
With vector search indexes, Firebolt makes it possible to run low-latency, high-throughput semantic search directly inside your data warehouse — without maintaining separate vector databases or compromising on ACID guarantees. By combining HNSW indexes with Firebolt's tablet-based storage and execution engine, we can search hundreds of millions of high-dimensional embeddings in milliseconds, while staying fully transactional and tightly integrated with SQL.
In this post, we explored the architectural building blocks behind Firebolt's vector search: how indexes are built and cached, how transactional consistency is maintained, and how query execution plans make use of these indexes to deliver fast top-K searches. We also covered best practices for engine sizing, index tuning, and table design to help you get the best performance for your workloads.
If you're building semantic search, retrieval-augmented generation (RAG), or other AI-powered applications, Firebolt's native vector search capabilities let you keep your data and your queries in one place, fast, reliable, and SQL-native.
Start using Firebolt's vector search indexes today:
* Explore the Firebolt documentation (coming soon) to learn more about CREATE INDEX, vector\_search, and tuning parameters.
Try out your own workloads with vector search indexes to see how far you can push performance with [$200 in free credits](https://go.firebolt.io/signup) to get you started.
# The $100M Problem: How Lyft's Data Platform Prevents ML Failures with Ritesh Varyani at Lyft (/blog/the-100m-problem-how-lyfts-data-platform-prevents-ml-failures-with-ritesh-varyani-at-lyft)
What if your data platform could serve AI-native workloads while scaling reliably across your entire organization? In this episode, Benjamin sits down with Ritesh, Staff Engineer at Lyft, to explore how to build a unified data stack with Spark, Trino, and ClickHouse, why AI is reshaping infrastructure decisions, and the strategies powering one of the industry's most sophisticated data platforms. Whether you're architecting data systems at scale or integrating AI into your analytics workflow, this conversation delivers actionable insights into reliability, modernization, and the future of data engineering. Tune in to discover how Lyft is balancing open-source investments with cutting-edge AI capabilities to unlock better insights from data.
Listen on [Spotify](https://tinyurl.com/7pzmx4tn) or [Apple Podcasts](https://tinyurl.com/ypdsb54s)
**Benjamin (00:01.337)**
All right. Hello everyone and welcome back to the Data Engineering Show. It's really good to have you. Elad couldn't make it today because he's out traveling. Yeah, but we're super happy to have Ritesh on the show today. Ritesh is a staff data engineer at Lyft. Yeah, we're actually staff software engineers, sorry. It's great to have you on the show. Kind of welcome. Do you want to quickly introduce yourself, kind of your journey into data and what you're working on today, Ritesh?
**Ritesh Varyani (00:30.624)**
Perfect. So I'm Ritesh. I'm a staff engineer here at Lyft. I've been at Lyft for about six years. Before that, my software journey actually began at Microsoft within the data space itself, doing some SaaS products at that point in the CRM space and then in the Azure space I did some platform work with Hadoop.
Then at Lyft, have primarily been within the data platform space in the six years and I've jumped across products and at this point kind of lead the products of Trino, Spark and Clickhouse at Lyft.
**Benjamin (01:13.701)**
Awesome. Okay. Very, very cool. So looking at the Lyft data stack, right? Like I assume you guys are processing kind of crazy amounts of data. Give us a bit more info on where your team actually works. Are you kind of like build driving, customer facing analytics? Are you kind of focused on internal analytics as well? Would love to understand the overall stack, where in the company your team works, et cetera.
**Ritesh Varyani (01:32.334)**
Ciao.
**Ritesh Varyani (01:37.91)**
Yeah, so the goal of our platform is to give our users access to the data as fast as possible so that they can drive the meaning from the data that they are kind of getting and take better data-driven decisions because of how very different data products are and they carry different kinds of goals around real-time accuracy, around latencies, around approximations, around predictions. So we serve kind of a variety of customers through these products like Spark being a massive batch processing engine. Through that, are primarily focused on customers who want to do.
ML training jobs who want to kind of process a lot of data and get a larger meaning out of it who want to run any kind of GDPR jobs that are going to do full scans across lots of your tables through years and being able to execute GDPR operations on it. So any kind of massive data parallel processing, you would have customers that would be using Spark as a platform. So we serve in that space ML engineers, software engineers, data engineers who are writing these pipelines or who writing these training jobs and trying to get meaning out of the data to make these decisions. If we go towards the Trino side of things, a typical customer anywhere from a data scientist to a data analyst to a software engineer who is either dashboarding who is doing I would say an average sized ETL like not too heavy like Trino ETL works really great for us
**Ritesh Varyani (03:27.352)**
There in Trino we have experimented with project R degrade, which is Trino's fault tolerant execution. We didn't eventually turn it on. It yielded good results. There was no problem in the technology per se or that particular feature. We just didn't feel the need and we were able to successfully do Trino ETL without kind of turning it on. The primary reason for that being we developed our own ETL kind of a logic within the orchestration layer for Trino. So you would be able to kind of write into a temp table, you would be able to do a dequeue, then you would be able to do an atomic insert override from the orchestration layer rather than the engine layer per se. So we had already done the work before those features came out and we were able to kind of do a lot.
I would say average sized ETLs through Trino. Another very important use case that Trino kind of serves is the BI dashboarding. Any person, any engineer or any software engineer, even an exec that's able to write SQL, they can kind of write Trino SQL and get the access to the data that they want and they can kind of product managers, they can get the data out and basically get the meaning out of it.
Clickhouse remains a little bit in a different space. It's engineered for a space where it's able to do excellently well. It's a sub-second query latency space. It's able to do excellently well if you know what you are querying so that you can kind of organize the data accordingly. The speed that we have personally seen getting delivered for us from Clickhouse is if we know how the data is going to be queried, we can certainly kind of help our customers, must store the data accordingly that gives them like sub-second latencies for their OLAP workloads so they are able to quickly slice and dice the data. Primary use cases for this would be something like in the marketplace where you want to immediately be able to see like how the last couple of hours in different regions is our supply demand and a bunch of other marketplace metrics doing. Couple of other use cases would be if anybody wants to forecast something, if anybody wants to feed it in a real time
**Ritesh Varyani (05:48.584)**
Pipeline and we have some of those use cases downstream that actually go off of Clickhouse and use that data to feed into the downstream RTML systems.
**Benjamin (05:59.653)**
Okay, awesome. And then in terms of the underlying infrastructure, do you folks run your own data centers? Is it all on public clouds? Is, yeah.
**Ritesh Varyani (06:08.91)**
Yeah, so Lyft is an AWS shop. So primarily, let's just consider 99 % of our infrastructure and platform is on AWS. For Clickhouse, there are a few technologies for which we have gone to our vendor. So Clickhouse, we use Clickhouse Cloud, but Spark and Trino, we continue to run in-house. We are a Kubernetes and an AWS shop. So it's everything on EKS at this point for us within the AWS environment.
**Benjamin (06:11.557)**
Okay.
**Benjamin (06:38.725)**
Okay. And then in terms of the like underlying data flow, are you folks dumping it all into a data lake that serves as like the source of truth for all of these systems? Are you using open table format? It's like, how do you actually move data around between the different systems basically?
**Ritesh Varyani (06:55.437)**
Sure. So almost all of our data not all the data But let's say almost all of the data that's used for offline processing today not online again is dumped into s3 And we have our parquet tables defined on top of it. We primarily are a high format shop We are going to be moving to other open table formats in the future But at this point we are a high table format you dump your data and you define a table on top of the set of parquet files and you're able to query that data through Spark and Reno. Through Clickhouse, you need to kind of make the data flow into their kind of ecosystem so that it's arranged in a way which gives you the performance of your queries. There are a few other places where data is outside of S3. also use very heavily Redshift, but it's more around the financial infrastructure space rather than the data platform space, but there is a lot of integration between the two spaces moving data across these two links.
**Benjamin (08:00.463)**
Okay, super cool, nice. So yeah, that's a super mature stack and it's kind of cool how many different pieces of data infrastructure you folks use to serve different workloads within the organization. What has your main focus been throughout this year?
**Ritesh Varyani (08:17.774)**
This year the primary goal that we particularly had in our mind was to understand our reliability problems, understand our scaling problems and future proofing our platform. We have for a good amount of time, we were doing a lot of open source work that kind of spun off into a lot of, I would say, products to manage, maintain, deploy, upgrade, ensure it's reliable, you're not falling behind the industry. We did a lot of that around Trino, around Spot even around Clickhouse with some Kubernetes operators when we were in-house. But our main goal at this point is primarily understanding how do we see the data platform running, let's say, five years from now, three years from now, and how we are able to future-proof it. And in this world of AI, we should not be falling behind in any way. And bringing AI in the right places within our platform.
We look at AI as something that integrates across the company. It doesn't remain as like a particular org or a particular domain or a particular kind of a product within the company. And across the domains of the company, we already have had some initiatives outside of data platform, couple of them within data platform. The goal is to kind of integrate wherever it makes sense so that you are able to kind of derive insights from it.
**Ritesh Varyani (09:53.6)**
second aspect is modernizing and unifying is what we are kind of trying to think at this point so that you are able to, again, the goal of giving a reliable platform to the users so that they are able to get the meaning out of the data consistently and reliably without kind of, you know, suffering through any of the hurdles in the picture.
**Benjamin (10:15.877)**
Okay, super interesting. So let me drill into the AI part, right? Because that's something I think that's on the mind of a lot of listeners. How do you think about that as a data platform team? Because fundamentally, right, like a lot of the work you folks are doing is kind of giving the other teams at Lyft the right tools for the job to get value out of their data. So how do you think about that? How do you think about it in the context of actually running multiple different systems? Is it like, hey, now we infuse?
Pre-know and give it like a, let's say, I don't know, like make it ready for rag workloads. Then you do it with Spark. Is it something you're actually driving more at the data lake layer, kind of how to open table formats connect, like a million questions here basically. And, and would love your take on all of those.
**Ritesh Varyani (11:02.094)**
Yeah, so a couple of things that, I mean, there are a lot more, but a couple of things that we immediately see kind of working out probably for us, at least in short to midterm, is specifically around how we kind of view semantic layer within the data space.
We have a semantic layer aspect or a product at left internally. We want to kind of make it ready for the AI native side of things. So think of it like semantic layer V2 with AI native support so that marketplace, a bunch of other orgs are able to drive the best meaning possible from the data that they see. It could be the data through the dashboards. It could be data that's… joined across through some ETL pipelines and some derived tables. It could be some real-time data that they are seeing through ClickHouse as an example. So that's one layer where we see direct interfacing and direct investment kind of happening from our side because we want to kind of be at the forefront of that and ensure that we are able to get the best predictions possible as fast for our use cases primarily. The second aspect that I see is big data systems are very complicated. They are distributed systems by nature. There is a lot going on under the hood. There is a curve to understanding all of those things. How do you simplify that for a new hire? How do you simplify that on a large scale for a platform with changing workloads? I'd lift probably at other companies as well.
**Ritesh Varyani (12:48.622)**
the kinds of queries that are being written, the kinds of data that is being queried, the kinds of data that is being brought into the system will always keep on changing. It's not about growth. Growth is a separate aspect, but it will always keep on changing from every six months to a year. So you planned for, you did a certain capacity planning as an example for certain kinds of workloads. And now those workloads are different. Now your capacity planning doesn't work. Now either you are reactive or you're, it's not, auto scaling? Do you have the right spark queues? Do you have the right greener queues? Do you have the right resourcing limit set as an example? So auto scaling at a platform level is a different problem. Yes, it can be a little bit reactor based on how much either you see the queue size, how much you see CPU usage, memory usage and so on and so forth. But where AI can help you is very clearly understand how the A, are the patterns changing? If the patterns are changing, what is like a good action to take on those patterns? Which is like an agentic workflow.
**Benjamin (13:48.387)**
Right. Which is super interesting. That basically means you almost have these two orthogonal topics in which you're leveraging AI. Like A, you have this topic around how can a user infuse their analytics journey with AI, right? How can I ask a question about the data and kind of the data platform just is able to give me that answer, which is where the semantic models come in. And then you also have the second dimension around how can you leverage AI as a team to run your systems in a more efficient way?
**Ritesh Varyani (14:17.411)**
How do you run your platform better?
**Benjamin (14:19.171)**
Exactly. think for us, right, that kind of like building firewall, we've actually seen a similar pattern that there's almost these two.
Kind of layers in which you want AI. So we have things like an agent and the product that you can use to ask questions about your data or like learn more about Firewall. So that's more of this like a user interacting with your product gets AI infused into the experience. And then there's of course the AI for the person building on top of that, right? Like how can we integrate with LLM providers in our SQL dialect? Kind of how can we add vectors or support and all of these things? So it's quite interesting that like for many, think like people working around data, as they think about AI, there's like different lenses with which AI comes into the overall picture.
**Ritesh Varyani (15:03.148)**
Yes, and to actually add to this, there is a, I would say there is an overlapping use case as well here. So think of the users who are using this platform to get their data. Now, how do you help them out if, as an example, a massive… 2000 line job or a PySpark job goes wrong. So some of the things that have been pitched internally is also having like let's say an AI agent for Spark. We have not kind of productionized it at this point, but let's say you have a job that ran for two and a half hours and it suddenly failed. Was it a platform issue? Was it a networking issue? Yes, you can go through a big chunk of log file and kind of try to determine which is what we do today, which is what Anon called us today in the team. But can you just take that entire log file? Can you just take the set of Spark configs that the users passed and just dump it in an LLM and try to get the best meaning out of it? You share your Spark version, you internally, you kind of share a bunch of metrics around that job that failed and get a set of response that, hey, probably was a config an error or was the nodes an error? So… you have like a two-pronged solution there is A, either you go fix-con-fix or B, is your probably the auto-scaling slow at that point? Did you not get enough kind of capacity to kind of run that kind of a workload? So those are the aspects where we are also thinking at this point to kind of invest and see if better meaning can be driven for users. Again, hitting the point of reliability and ensuring the users are able to use AI in a more integrated fashion and able to get the insights consistently for the data.
**Benjamin (16:52.547)**
Right. Okay. Super interesting. But then in terms of, in terms of your overall strategy on adopting, let's say like a new vendor is somewhere within your data stack, right? Like how does AI change? You're buying decisions as well. Like, would you say you have new requirements of platforms in terms of explainability, telemetry, kind of all of these things in terms of feature sets? I'd be super interested in how you basically think about that as you bring in new technology that doesn't even have to be commercial, right? Same thing with like an open source project you're bringing in if your lens on that is shifting at all.
**Ritesh Varyani (17:31.926)**
Yeah, so 100 % yes. It does not mean every decision for a vendor product revolves around AI. It becomes like an important factor for us to consider is, and to be honest, like all the vendor products that you will go out today to kind of research, everybody will give you an AI part in their system. Because people across the industry are really, really passionate and driven and from the business perspective as well, like they want to do something around that space. Now, does that mean that when you look out for a vendor product, you are kind of looking of, yeah, sorry. Yeah, so does that mean when you're looking for a vendor product, you're only interested in what AI work they are doing? No, but it becomes one of the, I would say, part of your RFP package to really kind of understand, hey, do you kind of,
**Benjamin (18:13.413)**
Sorry.
**Ritesh Varyani (18:30.122)**
Let's say, give any idea bug ability aspect. Do you give your own semantic layer? Or do you provide a few other features that allow for data discovery in a much better way as an example, rather than a simple search? Where that integration lies, how much it is fruitful, you need to also test it to see how much fruitful it is, to be honest.
**Benjamin (18:51.843)**
Right, okay, that's super interesting. then, kind of, yeah.
**Benjamin (19:01.838)**
In terms of your open source strategy, right? Like, and how that relates to AI, is that something that you feel is becoming?
More important because you also adopted some rights, managed offerings, and moved a bit away from being fully built around open source. But you're, of course, also contributing to Trino. I'd be really interested in how that is changing and also how your perspective on maybe open source has been changing over the time, whether AI actually plays a role in that, whether that's completely unrelated. Would love to learn more there.
**Ritesh Varyani (19:33.422)**
So a couple of things that come to my mind around this is Lyft has also evolved in all of this time. it's a question of kind of really prioritizing of what work is the most business critical work and how do you kind of go 100 % behind it and kind of deliver it for your customers. The second aspect to this then becomes is how do you accelerate that with AI?
So rather than thinking of this as an AI versus an open source thing, it's about a question of, hey, this is our goal. We want to modernize the platform. We want to build a platform that serves for the next three to five years or the next generation for Lyft. How do we get there given what we see, the current and the future or like short-term future developments that are happening in the industry. Open source investments are still very heavy from our side, like we do and want to keep on contributing to it. But at the same time, we want to kind of really understand where the business is going and also kind of ensure that our investments are in alignment with that. So rather than viewing this as a conversation.
As AI versus open source, would be more around, hey, according to the out-business strategy, where does it make sense to invest in the long term? And how do we kind of fulfill it with AI in the picture?
**Benjamin (21:05.293)**
That's super interesting. And then one thing I actually see. If you think through it, right, is like. Where does your engineering time go as a team? Like if you look at it at the portfolio level, it's like fundamentally more and more of that time is probably now going into like unlocking kind of value through AI, right? Because all of a sudden you woke up one morning, you have a new tool in your toolbox and you can all of a sudden build things that haven't been possible to you before that unlock different types of value eats into other capacity you might have, right? Like on, for example, evolving your cell phone, so data platform or contributing back to the core pre-no query engine. this like basically something you're then trying to balance in your overall strategy to also figure out, how can we get more cycles to actually focus engineering time on our AI strategy?
**Ritesh Varyani (22:00.46)**
So the short answer to this is not everybody is working on AI initiatives at this point. And to be honest, it won't make sense if everybody.
**Benjamin (22:09.731)**
Right, of course. Have a business to run, right? Like kind of customers, customers to serve, kind of, yeah.
**Ritesh Varyani (22:15.854)**
Exactly. So where does it make sense? We as a company have taken a public shift as well towards being like an AI native company. So if it makes sense according to our business strategy, if it falls in that particular purview, if it aligns with it, then obviously we go and invest. in the initial phase of this, was about, if you are the one who is going to take on the initiative, probably spend a few hours outside of what you're already working on. That is how you will kind of discover AI, the tooling for it. And then you can kind of probably get an engineering leadership buy into, if it makes sense to kind of support it. Like Lift overall as a company is a very, it's a bottom-up company and a top-down company, both like we need to align with business initiatives, but we are very much a bottom-up company that way therein wherein if hey I find some value in an experiment that I did or a few hours over a weekend or a couple of weekends that I spent that can kind of help towards the larger goal like okay what's the larger goal for data platform it's the reliability space it's making better sense of the data okay what AI initiatives you can bring in for it does it make sense do a POC let's see how much it succeeds what are the success metrics looking like will it fly in production. So you go through those cycles on your own and then all of a sudden it can very easily then get funded in your planning cycle because you're able to show that, it aligns well with what we want to do. It's the future. And this is how it kind of plays out in the space. That kind of works out very well. The second aspect to this is the leadership also has recognized it already and which is where comes the public shift of being an AI native company. So which is where.
**Ritesh Varyani (24:15.438)**
Currently we are in the phase wherein we have some dedicated use cases all across Lyft that kind of have AI tooling, have AI chat ports, AI agents that are already in production that are working. So if it aligns, again, coming back to the same point of, if it aligns with the strategy, if you're able to show, if you have the passion to work as an engineer, because it's a difficult space as well to get started with. The second aspect to that is it's changing a lot.
**Benjamin (24:40.367)**
Right.
**Ritesh Varyani (24:43.822)**
So you need to be abreast with everything that's happening in the space as much as possible so that you are not kind of left behind as well as an org, as a company or as a software engineer yourself as well. So if you are passionate enough, if you find the space, there's obviously encouragement to kind of go and see. If you see the results, we'll fund it.
**Benjamin (25:08.261)**
Okay, very cool. Interesting. Yeah. It's like, it's super fascinating, right? To like kind of see these very like large infrastructure stacks basically now adapting to the transformation. mean, it's incredibly difficult to run these types of data platforms, right? And there's very few companies that actually have workloads at that scale. And then all of a sudden you wake up one morning, there's new technology and you have to figure out kind of how to, yeah, kind of like have it bring value within your overall stack. And that's such an interesting transformation and like seems like a massively exciting engineering challenge to be involved with as well.
**Ritesh Varyani (25:44.672)**
It is, it is. And I mean, a few of the teams were very open to the AI space. Like we had, we have a deep integration with our agentic framework within our ML space as well. The ML org is doing pretty well on that particular side of deeply integrating and providing AI. And it's been a company wide initiative now as well at this point to ensure that it's deeply integrated with the developer tooling. It's deeply integrated within your GitHub PRs. It's deeply integrated as customer facing tooling agents that can help out customers. Not only that, it's deeply integrated in, let's say, a UI tool internally used at Lyft for people to kind of derive meaning as to what they are seeing from the data that's in front of them.
**Benjamin (26:31.439)**
Right.
**Ritesh Varyani (26:31.928)**
So it's these different spaces where people have, people initially kind of invested their own time, off in their own arcs in kind of trying to experiment, see if it works. And now slowly what we are trying to, what's kind of happening as a company is it's kind of consolidating into like a single direction of kind of providing different kinds of models so that you are easily able to integrate, you're easily able to kind of use an LLM prompt and get the result you want and showcase to the end user so that you are less worried about, how do I establish the connectivity to those models? How do I establish, how do I get the business subscription? How do I get through this entire approval process, procurement process and everything? All of that kind of gets now consolidated. And what you can kind of focus on is, hey, I just want to integrate this and I want to test and measure this if this kind of provides value to my customers.
**Benjamin (27:30.063)**
Right, okay, that's super exciting. Nice, very cool. kind of closing out, you think ahead, so 2026 is coming up, we're recording this on the 9th of December. What are you excited about next year? Kind of like at Lyft and the data space in general. Yeah, kind of what's on your mind?
**Ritesh Varyani (27:48.098)**
So two things, point number one is how are we able to work on unification of our data stack so that we are able to serve our customers better and probably not go towards an approach which becomes harder to manage, maintain, upgrade and kind of be reliable for our customers. The second aspect to this is use AI wherever it's pertaining within our data platform space to be able to serve our customers better, like add that particular layer of faster iteration development right from our engineers within the company to within the data space, you are able to kind of derive better meaning from the data that you see in front of you.
**Benjamin (28:37.189)**
Okay. Nice. Yeah. That's cool. That's exciting. Well, let's catch up like kind of let's do another episode a year from now, kind of see how your stack evolved. That would be amazing to learn about. It was really like super awesome having you on the podcast today is like kind of I learned a lot about the stack in Lyft. I'm sure our kind of listeners did it well. The work you folks are doing is super impressive. Any closing words from your side?
**Ritesh Varyani (29:03.534)**
We are hiring for the positions, so please take a look at our website. The positions are very much up to date on the website, so if you're interested in joining Lyft, please do apply. Thank you, Benjamin, for the time. It was amazing collaborating with you and talking about all these stacks between companies.
**Benjamin (29:25.763)**
Yeah, amazing. Love it. It seems like a great kind of team to work with. So if you're on the job market, kind of go, go apply to be a part of Retached's team or the data or get lift. And then, yep. See you around. Awesome.
**Ritesh Varyani (29:37.774)**
See you have a good one.
# The Resurgence of SQL: Insights from Ryanne Dolan from LinkedIn (/blog/the-resurgence-of-sql-insights-from-ryanne-dolan-from-linkedin)
In this episode of *The Data Engineering Show*, Ryanne Dolan from LinkedIn joins the Bros to discuss LinkedIn's [Hoptimator](https://github.com/linkedin/Hoptimator) project. Ryanne explains how they're simplifying complex data workflows by automating them through SQL queries, integrating Kubernetes, Kafka, and Flink. The conversation highlights the shift towards a consumer-driven data model and the future of data engineering.
Listen on [Spotify](https://spoti.fi/4eASwgw) or [Apple Podcasts](https://apple.co/4dh4E5d)
Transcript:
Into/Outro (00:00:04) The Data Engineering Show is brought to you by Firebolt, the cloud data warehouse for low latency analytics. Get $200 credits and start your free trial at firebolt.io.
Benjamin (00:00:15) Hi, everyone, and welcome back to the Data Engineering Show. It's been a while, but we're back today and with an amazing guest, Ryanne Dolan. Welcome to the show. So good to have you.
Ryanne (00:00:25) Thanks. Happy to be here.
Benjamin (00:00:26) Ryanne, is a Senior Staff Software Engineer at LinkedIn, and he's been there for closing in on three years at the end of this year, at least, and kind of also outspoken kind of in the community, giving talks, thinking a lot about Kubernetes, databases, and a bunch of things we want to chat about today. Ryanne, do you just want to quickly introduce yourself, talk about kind of your story?
Ryanne (00:00:47) First off, I started at LinkedIn almost exactly 10 years ago, but I left and came back. So I'm one of those big tech boomerangs that everyone wants to be. I think that lends me sort of an interesting perspective, right? Because I've been around for long enough that I've seen the company sort of grow and change. And that's kind of what I want to talk about today is, you know, sort of like industry trends that I've seen sort of evolve and change over time. The Kubernetes as a database talk. Is one such like something I couldn't really imagine talking about 10 years ago. Things changed so quickly. And I joined LinkedIn, like I said, 10 years ago from an acquisition. And I stayed for four years left to join Word and Works, Cloudera, Twitter, and came back to LinkedIn. So that's sort of been my growth story.
Benjamin (00:01:33) It's going to be a fun show because like I'm like a database internalist guy, right? So kind of when I prepped for the show, it was like, oh, database, just the Kubernetes. I was like, have some fun. I want to learn more about this. In the database community, it's the other way around. Like people are like, oh, databases should start replacing file systems, operating systems, your network stack, whatever. So looking forward to this.
Eldad (00:01:54) Kubernetes is just a compute of a database, right?
Ryanne (00:01:57) I would say everything could be a database. I mean, if you think about it, a database is like just built on what, like spinning rust. So at some point, right, it shouldn't be surprising you could build a database on practically anything, right? And sort of the trendy thing these days is, yeah. You know, pick your most reliable, easiest, cheapest way to store data and then build a database on top of it, right? If that's S3, go for it, right? If it's your favorite blob store, go for it. And I think that makes a lot of sense. And so with Kubernetes, it's kind of a, it's so ubiquitous, right? Like it's obvious that Kubernetes is there. You just take it, take for granted that it's there. It's the computer. It's, you might as well think of it as hardware or as a RAM or CPU, which is part of the stack, at least at some companies. So if you just sort of take for granted that Kubernetes is there. It starts to work. It starts to kind of make sense as sort of the operating system of the cloud or big distributed computer sort of thing. So, yeah, it shouldn't be that surprising that it's easy, sort of, or at least interesting to build a whole database right on top of that as, you know, sort of like a sort of obvious substrate.
Benjamin (00:02:59) So you already hinted at it. It's like in your introduction, right? But one thing that's on your mind is like this trends over the past 10 years, right? And having been at LinkedIn back then and kind of being there now, there's something we'd love to learn more about. Kind of. Give us your take on that, right? What data challenges do you have today that you still had years ago? Which ones are much better? Would love to hear more about it.
Ryanne (00:03:20) I think a big one, like I noticed this right away, almost day one when I came back and it absolutely shocked me. And I'm not sure it's industry-wide, but my guess is this is industry-wide. So I sort of like stepped out of the industry and sort of focused on startups and enterprise-y stuff for a while and then kind of got back into big tech. And when I came back, SQL was popular. Right? Like when I was in, you know, undergrad, grad school. You sort of learn about SQL and it's like, okay, it's this way that you talk to databases, but no one uses those things.
Benjamin (00:03:47) Old guy on the corner, like, what, who is this? Is it web scale?
Ryanne (00:03:51) Exactly. Relational was like, you know, something from mainframe, old and slow. But yeah, when I came back to LinkedIn, I was literally shocked to find that there were, you know, these projects which were 100% SQL. Like there would be a repos that were just hundreds of files of thousands of lines of SQL, which first of all, I didn't even know was SQL. It was a thing. My experience of SQL at the time was that here's a query, here's a response, here's a query, here's a response. I never thought of it as sort of a programming language and sort of a way to sort of like this like data engineering concept of moving data around and manipulating and making models and all that stuff. Completely foreign to me. So I think that that's been sort of the biggest change that I've seen, right, is suddenly SQL is popular. And sort of in that same vein, it used to be that relational databases were considered slow, unscalable and great for enterprise use. You know, I used to think, right, that like if you use a relational database, you're talking about like maybe a hospital, right? You're not talking about Twitter, right?
Eldad (00:04:51) The worst part of the hospital, of course. Exactly.
Ryanne (00:04:54) Exactly.
Benjamin (00:04:55) CrowdStrike and the database two things.
Eldad (00:04:58) SQL Server 6.5 running on. That's what we all imagine when we think databases, or at least that was sometimes.
Ryanne (00:05:07) Absolutely. And that was my prejudice. That was my sort of, you know, you should training. I mean, as you as you progress through grad school, undergrad, early career, you sort of trained to focus on scalability, right? Like my first resume was like at the very top. It's like I'm interested in scalability. Like that's what you had to say. So that's always front and center, especially sort of breaking into big tech is everything has to be scalable. So it was my sort of assumption that SQL was intrinsically not scalable. Relational databases were intrinsically not scalable. And fast forward years. I feel like that's really been challenged. There's a sort of like year window when all these big companies were building their own no SQL databases. Basically on the assumption that SQL was fundamentally broken and fundamentally not scalable. At Twitter, we had Manhattan. At LinkedIn, we have Espresso. Most big tech companies have like multiple such databases that they fill for different use cases. But we're sort of getting to the point where there's so much benefit from the relational model, so much benefit from SQL. That any sort of performance. Penalty is probably worth the penalty. Right? So I think a lot of that's being driven by AI revolution where I think fundamentally sort of the value of writing code is sort of diminished. So you want to do as much work with as little code as possible because if you have a bunch of like block of code that you want to generate, just have ChatGPT do it. Right? They don't have to get paid to do that anymore. So people are trying to like climb on top of this high hill, this pile of stuff. And if you climb high enough, you get to SQL. Right? SQL is like this abstraction layer that microservices, data pipelines, caches, all these things that we used to build by hand kind of fall out of like, let's just subtract this as SQL. So, yeah, that's been a big change. So suddenly SQL is popular and suddenly no SQL is not popular. Right? We don't need a specialized database just for key value store. You can build, you know, just a bunch of stuff. You can build distributed SQLs for relational databases with real constraints, real consistency. And you get value out of that, which, of course, we knew in the 1950s. But it's something that we forgot about, I think, for 15-20 years.
Eldad (00:07:23) But it's all Silicon Valley's fault. See? So as you describe kind of history and evolution, it all happened for a good reason. It all happened because databases reached an endpoint. They sucked. They just couldn't scale to cloud scale. And all the talent moved to big tech. And big tech could not wait for a third party to stop, to depend on. So they just started to build and build and build for the purpose of releasing features. And then kind of that whole open source came to be. Then it took 15 years for everything to converge back to normal, realizing that SQL is just marketing. And it's just a language. It's not even a real language. It's just an intermediate language, as you've mentioned, AI today. And research and startups and companies in databases caught up. So as you said, you get consistency at scale. There's so much innovation happening way below the SQL. But you're right. The biggest advantage of SQL is the optimizer that runs it. And a smart optimizer running on a great stack, a scalable stack, can do a lot. And that's kind of you coming back to LinkedIn. It's crazy. Because LinkedIn was one of those companies. Right? Hadoop. And all of the evolution. No SQL. Right? Hadoop, no SQL. Strata takes me back to those days. So now everything is mixed and just like good fashion, retro. There was always good retro. And SQL is retro, I think. So I'm very happy about it.
Ryanne (00:08:48) It's also sort of like this sort of like old culture of just giving raw materials to developers. You know, it used to be engineers just thought, okay, give me a bunch of spinning disks. Give me a bunch of SSDs. And I'll build my database. I'll build my product. Right? Give me enough RAM and I can do anything. And I think AI, to some extent, has sort of lessened the value of just writing code for code's sake. Just building things for building things' sake. And it's more like we want to build something as easily as possible. We want to lower the friction. We want to make something quickly. We want to build this new pipeline in seconds, not quarters. That shift has made SQL sort of like an obvious choice where it wasn't before. I think that's even a... Sort of a culture change on individuals, you know, individual engineers. I want to write one line of SQL and get something running. I don't want to spend a quarter building my own database from scratch with my own REST APIs from scratch and all this stuff, which is, again, sort of like mind blowing looking back. Because like when I was in college and looking for my first jobs, that's what the job was. The job was, here's some raw materials, go build a scalable thing. And like I thought, I thought like my whole purpose of being was to build the most scalable thing purpose built for me. My whatever project I was working on. And now it's all it's like in the big tech industry, we're talking about like doing more with less. Every company wants to do more with less. And I think that's sort of an easy way to phrase what's happening. But it's really about we want to accelerate the pace of accomplishing our next goal because AI especially is moving so fast. Right? We can't even fathom the next product that we want to work on in a couple of quarters. Right? None of the stuff that we're working on today could have been fathomed five years ago. So or even two years.
Eldad (00:10:34) The only thing that will survive in five years, which is like infinity for AI is a SQL. It's like everything looks different. But for some reason, AI sticks to SQL.
Ryanne (00:10:45) I'm telling you, AI loves SQL. There's a huge corpus of SQL code floating around. So it's and it's it's easily parsed and it's easily understood. And Benjamin to your point, you don't even have to be that good at writing SQL because there's the optimizer. There are decades of optimizer improvements. So if you can have AI plus SQL plus a great optimizer, you can actually get performance out of very little human input. Right? So I think that's like fundamentally why things like SQL becoming popular, things like databases on top of everything, the relational models becoming interesting again and not just building stuff out of raw materials like we used to do when we were, you know, when I started years ago.
Eldad (00:11:27) I will tell you something that maybe will a bit blow your mind. At Firebolt, even compute is part of SQL, meaning you're thinking Kubernetes, your units of compute, whether it's for a query or workload, whatever, it's part of the language. So we are so strong and not just us. Everyone from the database community is so behind SQL that it's not just expanding the type system.
Benjamin (00:11:51) Eldad is getting so worked up about SQL. Everything is like throwing his microphone around.
Eldad (00:11:56) Exactly. I'm so excited about SQL. Finally, for years I've been here. Thank you for sharing the love for SQL. And I think Benjamin agrees as well.
Benjamin (00:12:06) Totally agree. And I think it's funny, right? Because we sent out with this no SQL trend and kind of the industry scrambling around, figuring out, hey, here's this like, I don't know, I need to take things to crazy scale. What technology can actually cater to that? And in a sense, you're seeing a similar scrambling now around AI, right? And if there's, oh, everyone wants to build an AI-enabled applications, and then you have a million different ways popping up on how to do that, right? Kind of dedicated vector databases, kind of other types of systems, whatever. And as a relational database kind of guy, I of course think, okay, kind of it's all the same thing all over again and it will all be relational databases in the end. So love that you're seeing it the same way, Ryanne. If we look at the data stack then that you're using today at LinkedIn, basically, right? Kind of take us through it. Like what are the actual technologies you're leveraging every day?
Ryanne (00:12:55) Oh, I mean, I can't get in too many details, but typical stack. So we are moving to Kubernetes, we've talked a little bit about that. We are in the migration phase right now where we kind of like an old stack and a new stack. And the new stack is very much Kubernetes based. We sort of built things on top of it, which is sort of typical. We of course use Kafka all over the place. We of course use distributed SQL all over the place. But yeah, it's hard to talk about a stack when you're at a company like LinkedIn because it's been around for long enough that there is no stack in the sense of, like-
Eldad (00:13:30) Engineering is the stack.
Ryanne (00:13:31) Yes, exactly. It's, we're not talking about like LAMP, you know, we're talking about like a bunch of different teams making their own choices, right? There are some commonalities, but I mean, I think generally the commonalities are more out of convenience than constraint in the sense that Kafka is there. So people will pick it up and start using it. No one's saying we have to use Kafka. So yeah, the stack is sort of nebulous and hard to define, which is not a very gratifying answer, but in reality,
Eldad (00:13:59) But all the complexities is behind SQL.
Ryanne (00:14:02) So that's sort of like what I'm working on actually is trying to sort of roll up all this complexity that's been human made, you know, taking quarters of engineering effort and sort of rolling them up into abstractions that we can easily just express in SQL. So I spent many years working in data pipelines, building, you know, streaming data pipelines, batch data pipelines. And it turns out like if you zoom out, right, all these sort of pipelines look pretty much the same. And I think like Flink especially has sort of opened a lot of people's eyes to the reality that very complex work can be done in a few lines of SQL, or what used to be very complex work can be done in a few lines of SQL. I remember, well, like, 10 years ago, I built this sort of like multi-stage pipeline that was an individual Samza job, and all it did was consume records and repartition them. And then this other one like was, you know, there was like, you know, adding a little bit of data, maybe adding a column to a route, Each of these stages was hand-written in Java using structs essentially, right? Like not even any sort of like relational concept, but just like, here's your class, serialize it, deserialize it. And then each of these things had their own metrics, their own on-call responsibility, all this stuff.
Eldad (00:15:16) ORM. That's kind of when you know you've went.
Ryanne (00:15:18) Exactly. This was all very hand-built. It was sort of the theme I'm going after here, right? Like it's like hand-built stuff that scaled really well, or at least felt really low level even at the time. But yeah, if you zoom out and you think about these like multi-stage pipelines that took quarters of engineering effort, they're just essentially like insert into select from is like so much of what we built years ago was like one liners in SQL, especially with something like Flink. And then like another, so LinkedIn traditionally has had, I think this is pretty common, but there's been traditionally sort of like two sides to the company. There's been the offline grid side where we did things like Spark, Pig, all these bash things on Hadoop. And then there's like the streaming side of the company, which is Kafka, Samza. And there was no commonality almost at all between these two sides of the company, not in terms of from an organizational perspective or a tooling perspective. And I was a tech lead many years ago on a team that owned a bunch of pipelines. And we literally split the team down the middle and had engineers focused on building the back end, the batch version of the pipeline and the streaming version of the pipeline. And it was like different languages. There was, I mean, it was different tooling, different experiences, different skill sets entirely. And yeah, again, zooming out, all of that work could have just been select star. It's essentially what we were doing, right? So that's sort of what I've been working on recently is trying to like roll up those complex tasks that we used to build piecemeal and get them to the level where we can actually just write, insert into select star and actually get that machinery, running. Again, I'll point to AI as sort of like a motivating factor here. If you can express, I want to data pipeline to AI and it generates some SQL and that SQL generates a whole, like all the machinery, all the databases, all the web services, all that stuff. I don't think we're far away from that level of complexity under the hood where the actual developer experience is plain English. It takes seconds, not quarters.
Benjamin (00:17:17) That's crazy. So one of the then open source LinkedIn frameworks you're contributing a lot to is this, is this Kafkameter framework. Is this, kind of does this tie into the bigger story here?
Ryanne (00:17:30) Yeah, absolutely. I mean, Kafkameter is sort of an experiment and it's right in that, the vein I just described where we take literally just an ANSI select query. If you think the SQL, the simplest, most narrowest definition of SQL is like an ANSI query. That's it, right?
Eldad (00:17:49) Select one.
Ryanne (00:17:50) Yeah, exactly. We're not talking about DML. We're not talking about all this fancy stuff. We're just talking about selecting data and then like having a pluggable catalog of how do I get data from Kafka? How do I get data from Espresso? How do I get, whatever. And then just presenting that select as like a user experience. Like the user gives us just the data that they want. And from that literal query, we can stand up Flink jobs. We can set up materializers. One of the interesting things we do internally is wire up all the ACLs. If there's a new Flink job running from day zero, that Flink job doesn't have access to anything. Sort of like the traditional way of building a new service is to like send a bunch of emails around. Like, can I have access to your data? And so we've gotten to a point where it's all automated. So like when a new stage of a pipeline is deployed for the first time, it'll actually go and request ACLs and get permission and all that stuff totally automated. So that's what we've built with Kafkameter Project. It's all about what can we get into production from just a select query? And it ends up being quite a bit. Like you would, there's sort of like this easy mental leap, a SQL query and like a Flink job, but it's not just a Flink job. It's like all the surrounding things that, that Flink job has to touch. We'll go and we'll create Kafka topics. We'll go and, you know, request ACLs and set up schemas and all that stuff. The reason it's called Kafkameter, if you, you know, anyone who hasn't heard the project, which I imagine is almost everyone, right, that day behind the Kafkameter is it's automating these like multi-hop data pipelines. So we're not just talking about one-off jobs. We're talking about sequences of jobs, complex, you know, workflows, streaming workflows, and even sort of like batch and streaming at the same time, right? Like going back to my earlier example, we had a team that was just split down the middle. This is the batch side, this is the streaming side, completely different tooling. Your job is to do bootstrapping every day, backfilling every day, and that sort of thing. And your job is to stream the data as quickly as possible. Again, both of those sort of roles could be expressed in what is the data that the end user wants? So the end user just wants this data. So can they write a select query? And if they can, then we should be able to automate the process of standing up all the stuff that we built by hand, the streaming ingest, the backfill, you know, all this stuff.
Benjamin (00:20:01) This is super cool, but like it's also kind of meta, right? Because really what you're building is part of a database now because you're taking like an input query and you're actually figuring out, hey, what steps do I have to kind of perform to serve that SQL query at the end of the day in terms of splitting it into multiple pipelines? Like at the heart of this is understanding SQL, optimizing your SQL, just in the end, in your case, you have this multi-step pipeline instead of, okay, your traditional like operator graph that you then kind of plug into your runtime.
Eldad (00:20:32) But there is a problem, Benjamin. Unlike what you're imagining in your head, which is kind of the optimizer and the runtime are playing well together, here it's kind of reversed. It's first you build a complete system that has zero awareness for SQL. Zero. Like never knew no SQL. And then afterwards you're coming and saying, oh, we're going to wrap that in SQL. And it's going to be correct, and it's going to be consistent to a degree, hopefully. But then when you run your select, it spins up a Kubernetes job. It connects Kafka. It gets out of kind of catalogs the schema. It finds the bucket. Security is a nightmare. Just think about security. But if that actually works, if that works, Ryanne, then there is a future to engineering. Like we can reuse and recycle all of that without rebuilding. That's important and that's actually super fascinating.
Ryanne (00:21:29) That's exactly the idea. I mean, basically the essence of the project is, okay, we built all this low-level stuff that's not going anywhere overnight, right? Like these sort of, you can even consider them legacy systems to some extent. For so long, anywhere overnight. What we want to get to is this relational SQL-driven world. So can we just layer SQL on top of it? Right? And again, that surprising that that's possible, at least to some extent. Because again, at the end of the day, like a relational database is just spinning rust, right? It's like you're layering something on top of unreliable, inconsistent metal.
Eldad (00:22:08) But Ryanne, just one warning. Don't connect Looker to that system. They like, every dashboard refresh, every BI dashboard refresh will... Spin up those things. You'll need an admission controller and resource manager. But that's the beauty. Like, it could actually work.
Ryanne (00:22:26) Yeah, and what we've been doing so far, I call this out in my original design doc. It's like, you know, we could do this. The risk is, okay, people are just running random queries and we're like standing up, you know, like big tech level jobs that are like thousands of machines in service of their requests just because the optimizer is like, this is how I get the result in this amount of time, right? So that is a real risk.
Eldad (00:22:49) Just knowing what people need, having that interface in one place, like you have a single version of the truth, not of the data, of the queries. So you get the query history of all the outside system people you don't know about. And you get this one place where it tells you like, this is how people interact with our legacy, with our huge multi-complex system. It's beautiful.
Ryanne (00:23:12) Yeah, I actually, I'd love to drill more on that concept because what you just alluded to is something, I call producer-driven versus consumer-driven. I've read about this a couple of times, but it's sort of like, I don't know if it's trendy, but it's sort of obvious that you want to enable data producers to produce data. And so like everything we've built in the history of big tech, and I think the industry at large has been to enable producers to produce more data. But you have data products and things like that, which is all around, okay, you're the owner of the data. You're producing the data. You have the responsibility, and you own the schemas, and you store the data. You build the APIs. That whole model works to a point. What? What you're alluding to is what you really need is to know what all the consumers need. It doesn't actually help if everyone just throws in data all over the place.
Eldad (00:24:04) It won't work. We're not producers by nature. We're consumers by nature.
Ryanne (00:24:08) Exactly. But in development mode, when it's your job to own data, your whole mental model is to be a producer. I want to build this new API. I want to store this data somewhere. Your whole mindset is to produce data, and so we've built things around that structure.
Eldad (00:24:22) But look what happened to you. You turned a production system. So someone says, oh, I'm a producer creating a data pipeline. But this is a select query. Yes, the select query might generate a billion records as a report, but it's still a select. So producing data is actually running a select versus. Generating a data pipe with a lot of complexity that ends up, as you said, as an isolated mart.
Ryanne (00:24:44) Yes. Exactly. What I want to do is sort of like turn that model on its head. What you want to do is enable your consumers to express what data they need. And then the whole producer aspect largely goes away. And, you know, if you sort of like zoom out in the perfect world, you have these sort of like source of truth databases that someone has to produce at some level because otherwise you don't get any data. You have to have producers somewhere. You want to have a smallish number of high quality source of truth databases. And then you don't want to use that same technology, the same level of effort to build all of your data products, for lack of a better term, where all of your data engineers are building all these components that all have their own schemas and their own databases and their own SLAs and all this stuff. What you really want is you want like your source of truth databases with this high quality. And then you have your consumers, which can be front ends, can be mobile devices, can be other pipelines. And if you can wire up the consumer's requirements. To what the producers have and spin up pipelines automatically, then you sort of eliminate a whole lot of engineering. You don't get into this like, especially like in a streaming sense, you have you end up with like Kafka topics on top of Kafka topics on top of Kafka topics. They just keep going forever and ever and ever. No one ever says, let's delete this Kafka topic and just refactor this into one. You know, let's take these and collapse it. No one ever does that. It just keeps building and building and building and building. But if you zoom out and say, like, who are the end consumers? Well, you know, I have one application that needs status messages. Right. And then everything in between can be theoretically auto-generated from these sources of truths to the consumers. And questions like, well, should we use Kafka? Should we use a database? Should we use key value stores? Should we use relational database? Those are all optimization problems that a relational database could solve.
Eldad (00:26:33) They could be abstracted from the user and handled at the back end because it's about optimizing data access. So there are many ways to solve it and many products. But the abstraction with SQL, moving from being a producer mindset to a consumer mindset is super, super interesting and feels like that's exactly where the market is heading. But it's actually really nice to see it's also heading on those huge, large, impossible to change tech companies like LinkedIn, which is amazing.
Ryanne (00:27:01) I mean, it'll take time, but we're getting there. Yeah, we have some use cases in production at this point. So I'll say that Hoppimator is largely experimental. It sort of has always been experimental. But we've taken some interesting learnings from it and actually starting to deliver some real value. So we'll see.
Benjamin (00:27:17) Super exciting. And then to close out this conversation, I guess this is then where it also all connects to Kubernetes style control plane things, right? Because in the end, you just write SQL and then a tool like Hoppimator kind of figures out the rest. And it becomes a bit like your Kubernetes operator in a sense, kind of which knows how to spin up compute, kind of do like your request. So it's a bit of a reconciliation and so on, I think.
Ryanne (00:27:43) It's funny, like if you come at a control plane problem with the experience that sort of like skipped over Kubernetes, you know, like you're talking about people who may be a decade out of phase with me, right? You hit the industry when there were relational databases and things, and maybe there was Kafka before Kubernetes. If that sort of generation of engineers think about this problem, it's like you're just talking about data pipelines again. You're just talking about here's a bunch of metadata, and you're trying to get to the next state, right? Which is what you're describing as a Kubernetes operator is just a transformation of like this is the input state, this is the output state. And if you sort of like think about Kubernetes as just a metadata store, and you think of it as operators and controllers, as just these things that take an input state and produce an output state, which is largely what they do, right? Not but largely what they do, then you can sort of solve that problem in many ways. It's just a complex state machine. It's a workflow. It's a function from inputs to outputs. Incidentally, the Hopp-to-Mator project, I kind of backed into using the operator and controller model. That's sort of native to Kubernetes to solve this problem. It wasn't like the only way to do it. But we, again, we sort of like we're using Kubernetes. Flink is running on Kubernetes at LinkedIn. So sort of a natural fit. So the sort of like weird thing we arrived at was we've got SQL at the very top. And when you, you know, when you write the SQL query, a bunch of Kubernetes YAML gets generated. So we kind of have this like SQL at the top and YAML is sort of this intermediate layer. And then a bunch of controllers go and, you know, spin up. And it ends up being really easy to throw like database tech, you know, something like Apache Calcite is what we use, right? To turn SQL into YAML. It's actually not hard at all. And so that's sort of like the core of what Kafkameter is. And Kubernetes as a database talk is sort of, you know, spun out of that. Like if we can take SQL and turn it into YAML, why not just treat all of Kubernetes resources as SQL tables? Like why can't I select star from pods and get all my pods? And it turns out you totally can. And like, is that valuable? I'm not sure. But it's definitely cool and interesting.
Eldad (00:29:51) It's valuable if you force everyone to use only that to access Kubernetes.
Ryanne (00:29:56) If the only language you've got is SQL, then it's super valuable. And, you know, it's probably beneficial to like stay away from YAML as much as possible. So like if you give developers SQL as an authoring experience or a querying experience instead of having them write YAML and, you know, like checking in YAML and deploying YAML scripts and all that stuff. I think it's really cool. I think it actually might be a better long-term strategy. But yeah, it's sort of like a happy accident, right? That was we were playing with Kubernetes in YAML and SQL and it's like, we can translate between these different media. Why not treat Kubernetes as a database and a meta store? And that's basically what we do with Kafkameter . Again, we take SQL, we generate YAML, and then the Kafkameter project doesn't know anything about Flink. It doesn't know about Flink syntax, doesn't know anything about Flink. It just knows how to generate YAML that the Flink operator can then pick up. You can just like extrapolate from there. If any system just has an interface exposed as a YAML spec, then it's easy to wire up. Okay, this sort of table, this sort of pipeline is implemented by this sort of YAML and then just leave a controller to actually make that happen. Sort of manifest those, you know, those actual pipelines or those actual tables. And so that's sort of the direction we've headed with the Kafkameter project. And the Kubernetes as a database is just funny. Well, let's take this to the absurd level of, what if we just store metadata? What if we just query pods and query deployments and everything as a table? It's kind of a fun thing you can do, which I'm not convinced it's useful, but it's definitely interesting.
Eldad (00:31:26) Everything can get into information schema. You'll be surprised. Any metadata can be squeezed in there. Information schema is endless. That's how it can expand forever. If you look at the typical database, yeah, it's that place where everyone throws in all those weird tables that usually you don't want to scan, but you need to scan to understand the system from time to time. So having Kubernetes as a first-class information schema citizen would be interesting to see how that rolls out and what people will do with it. Super interesting, really. Super exciting.
Benjamin (00:31:58) Agree. Thank you so much, Ryanne, really, for taking us through this journey from the evolution of data pipelines to data at LinkedIn, Kafkameter, and then Kubernetes. Really awesome work. Love that you're such a SQL guy. Any closing words from your side?
Ryanne (00:32:14) I don't have any top-of-mind thoughts to close with, other than I'll correct you that I'm not, I haven't traditionally been a SQL guy. I've sort of like come around to it. So I can, I cannot say that I've always been right on this point. Like, I cannot say I told you so.
Benjamin (00:32:29) Real growth is realizing your past mistakes and then learning from them. Awesome. And then thank you so much for being on the show. Was an absolute pleasure having you and excited to see where Kafkameter goes.
Ryanne (00:32:43) Pleasure. Thanks so much, guys.
Eldad (00:32:45) Thank you.
Intro/Outro (00:32:48) The Data Engineering Show is brought to you by Firebolt, the cloud data warehouse for low latency analytics. Get $200 credits and start your free trial at firebolt.io.
# Transitioning from software engineering to data engineering (/blog/transitioning-from-software-engineering-to-data-engineering)
This time on The Data Engineering Show, Xiaoxu Gao is an inspiring Python and data engineering expert with 10.6K followers on Medium. She's a data engineer at Adyen with a software engineering background, and she met the bros to talk about why both software and data engineering skills are so important. Without software engineering skills you'll be limited to the rigid capabilities of your stack. But without data engineering skills you'll find it hard to be cost effective and see the bigger picture.
Listen on [Spotify](https://open.spotify.com/episode/2fyoKva0ISkzVrquo5dHsC) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/transitioning-from-software-engineering-to-data-engineering/id1561927688?i=1000635692108)
Benjamin (00:03.101) Hi, everyone. Welcome back to the Data Engineering Show. We have Xiaoxu joining us today. She's a very well-known data engineering blogger, would you say that's correct, Xiaoxu, and thought leader, influencer. You can tell us all about that in a second. And she's a data engineer at Adyen which is a financial technology platform, basically, which I'm sure we'll also learn a lot about. Yeah, do you want to say a few words as intro to Xiao Shu?
Xiaoxu (01:06.814) Yeah, so I actually started my career in 2017. I first started my career as a software engineer, and then I joined International Bank for a few years. And then I moved to a startup as an official data engineer. And then right now, I'm at Ardian, which is a fintech company.
What we do is we do payment worldwide. So as you can imagine, data is everywhere in the company. We need data for reporting, fraud detection, machine learning model, et cetera. And on the side, I'm also blogging, as you mentioned. And right now I'm busy with my first data engineering online course, which hopefully can be finished soon. But yeah, I love writing content and then do knowledge sharing with the community.
Yeah, that's in a nutshell about myself.
Benjamin (02:04.958) So you're hosting the data engineering course basically like you're or you're taking it
Xiaoxu (02:11.026) It's an online, let's say self-paced learning materials. It will be published on a platform called Educative, where you can basically learning things through reading content, not really video, but through reading. So yeah, I'm busy with the final reviews, et cetera, so hopefully it can be finished soon.
Benjamin (02:34.617) Awesome sounds awesome. We look forward to looking at that once it's out. Nice. Sounds great.
Eldad (02:40.327) Obviously if there is any material we can link and share to our listeners.
Xiaoxu (02:44.53) Yeah, thank you, really appreciate it, but not yet.
Eldad (02:47.989) Okay.
Benjamin (02:49.509) We'll just have you back. We'll have you back in a couple of months once everything's ready. So what got you into blogging basically, right? So you have more than 10,000 followers on Medium. Like tell us about that journey kind of going from software engineer to data engineer to then being like a thought leader in this space.
Xiaoxu (02:51.762) Yeah. Thank you so much.
Xiaoxu (03:09.054) Yeah, sure. So I can start with my writing journey. I started writing in May 2020, basically right at the beginning of the pandemic. I always call my medium blog as my pandemic baby. So the thing is before pandemic our team had this weekly knowledge sharing session where you know like everybody shared their insights, findings, products in this meeting.
And during pandemic, this type of meeting became extremely important because we can't really see each other anymore. And that meeting was basically the only moment that we can learn from each other. So whenever it was my turn to do a presentation, I always put in a lot of effort to make the presentation by giving a bit more context. and then give more examples and then prepare a Jupyter notebook so people can reproduce it at home whenever they want. So I remember it all started with one presentation. I was doing the thing about data class and the name tuple in Python. I did a session because I couldn't really find the things I want online, so I did my own experimentations and then I did a session. It went pretty well.
And after the session, I thought, okay, maybe I can share with a larger audience online because nothing was really confidential. And also at that time, things we were staying at home, I was also trying to find a way to be connected with the rest of the world. So I quickly summarized my note and then published my article in just two days. And yeah, sorry.
Benjamin (04:54.873) So this is back when you were at ING, right? And back then you were still a software engineer. So was your blog about software engineering initially, and then you transitioned to data engineering at some point?
Xiaoxu (05:06.398) Yes, exactly. So the first few blogs were written during my ING time. So it was all about Python and those little details about Python and comparing different packages and the optimization, etc. So indeed, it was all started with the software engineering content. And yeah, after I published my first article, I was really surprised because like...
On the next day, it got so many views. Like, I don't know if it should work like that or not. I was super excited and also motivated. And then I collect my other notes and then I published two other articles in the first week. So that was the most productive week in terms of my writing. And then the more I write, the more I enjoy it. So I kind of like make it as a habit and I start producing more and more content since then.
Benjamin (06:02.357) So you went viral right away basically.
Xiaoxu (06:04.91) I don't know, it's still a puzzle for me. But as I write more articles, I realized that not every article should be like that. So it was definitely a luck. Maybe other engineers also had that problem at that time, I guess. Yeah, I was really surprised.
Benjamin (06:31.801) Eldad, your mic's not working.
Eldad (06:35.199) Thank you for that Benjamin. Tell us, so when did you start feeling like, no actually no, I haven't said anything. When did you start feel like a data engineer and less of a software engineer and how did it feel like?
Benjamin (06:39.429) He was talking all the time. He was talking.
Xiaoxu (06:49.502) Yeah, that's a very good question. I actually intended to be a data engineer. So my career change was from, that was the moment from ING to DOT, which is the first time I became a data engineer. So the thing is when I was at ING, I was working with a data integration platform team as a software engineer. And then my main responsibility was to help the team
the data integration platform from scratch. So although we were busy with the platform work, like setting up infrastructure, writing software, I was also writing software to integrate data into different systems. And as you can tell, it was already a sort of like a data engineering work, although I wasn't really realizing that. But then a few years later, I started to read a bit more about data engineering. I'm trying to recall why, but I think it was because all the news from the cloud providers, or maybe just more people on my LinkedIn start to have data engineer title. And then I started to read a bit more and then I was really interested in those cloud services. And then I followed some tutorials here and there. I was really surprised because essentially what I was...
doing in the data integration platform can be requested at a service in the cloud providers. And it was kind of mind-blown for me. I don't mean like it can be technically like physically replaced, but like on a conceptual level, because we were doing like a Kafka cluster. We were building our own Kafka cluster streaming engine. And then we were building our own scheduling system.
And then we were doing our own monitoring, alert, et cetera. All these can actually be requested as a service from the cloud providers. And then I really wanted to get more into that area. So I started to look for like jobs, which uses the cloud services, and then like also provide those official data engineer title.
Xiaoxu (09:10.334) And that's how I, yeah, transit from a software engineer to a data engineer.
Benjamin (09:17.921) How, how\... Go ahead, Eldad.
Eldad (09:18.003) Would you say that, go ahead Benji, would you say that being a software engineer gave you some advantage in becoming a data engineer? Is that related to the project you've been involved with as a software engineer or it's just your journey? What have you seen out there? What can you share?
Xiaoxu (09:36.55) Yeah, that's a good question. So I think software engineer and data engineer are very similar. So if people are thinking about changing their career from one to the other, it's definitely possible as many people have done this in my friend circle and I'm also an example. But at the same time, they are like also very different. They have their own strengths and skills.
and they can also learn from each other. And for me, I feel like there are two things which really helps me in my data engineering journey. One is that as a software engineer, we have this mentality of building things from scratch. We love coding and we love writing unit tests to make sure our software works perfect. But.
The testing part is not really a standard in data engineering, especially the unit testing part. It is what I felt. People would write more data validation test as part of the production pipeline rather than unit test as part of the CSCD. But when you start to create really complicated transformation logics, you need unit test to make sure that you know what you have written and you are confident on your code. And when I was at DOT, we were mostly using SQL to write transformation logic. And then writing unit test for SQL was such a pain. But I did see the value of doing that. So, and also one thing is we were using dbt, so everything was in SQL. So I came up with this unit testing framework in dbt that allow us to do unit testing.
in dbt which works really well in the end. This is one let's say advantage I can see that we can introduce more software engineering basic practices into data engineering field which makes it more yeah more robust and then more correct and also I'm happy to see that dbt will soon support unit testing natively as well.
It is also a trend and means that people want to introduce more software engineering best practices to the data engineering.
Benjamin (12:09.253) So when you're talking about unit testing, dbt, then like, this is something like, okay, here's a part of the pipeline. Then here's some input data. Here's some output data and kind of making sure that the transformation works as expected.
Xiaoxu (12:16.988) Exactly.
Yeah, it doesn't really run in the production pipeline. It's the task next to the regular data transformation, but it runs in the CSCE. And it can block the release if there is something going on there. So indeed, yeah.
Eldad (12:37.263) Benji, it's like user spinning up engines.
Xiaoxu (12:40.746) Yep.
Benjamin (12:41.197) And it runs on the same data warehouse that you're running on anyways, just with like a then significantly reduced data set basically.
Xiaoxu (12:51.398) So for the unit test, I usually prepare my own input and output. So I don't usually use the production data for unit test because one, it can change, and two, it's usually quite big, the data set, so it can take a longer time. So I prepare my own data set and I also know what is testing in the data set. So I have more control over it.
Benjamin (13:17.213) but it runs on the same system. So if you're running on BigQuery, for example, your unit tests would also then be a DBT job running kind of on BigQuery orchestrated by whatever CI system you use.
Xiaoxu (13:19.102) Yeah, run another system. Yeah.
Xiaoxu (13:25.351) Exactly.
Xiaoxu (13:29.726) Yeah, so for me, it will create a data set dedicated for unit testing BigQuery, for example.
Benjamin (13:37.073) That's good. So we talked a bit about then, I love this angle by the way, kind of your like software background, kind of giving you like allowing you to think about testing and all of those things kind of maybe in a rigorous way. Like you finished university in 2017, right? Like how would you say your like traditional computer science degree in a sense prepared you for like a data engineering job now? Like, do you think universities are doing a good job here? Do you think...
It didn't help at all. What are your thoughts here?
Xiaoxu (14:08.99) Yeah, I think university did help me because in my bachelor's I didn't really study computer science. I was studying electrical engineering. So it was more like on the hardware side. Master program was the first time that I started to know about computer science and learn coding, et cetera. So it definitely helped me for the long term. But of course,
what we learned in a university was very, let's say low level, it's, and very, let's say basic. And a lot of things I still learn through the work self. So learning by doing it.
Eldad (14:52.996) Life, Benjamin.
Benjamin (14:54.641) life.
Xiaoxu (14:55.44) Yep. Learning by doing. Yeah.
Eldad (14:56.243) You don't learn much at school. But it's great, it's good for the CV. Where did you study?
Xiaoxu (15:06.282) Where? My bachelor was in China in Shanghai and then my master was in the Netherlands in Delft, Delft University.
Benjamin (15:19.105) and you fell in love with the Netherlands and kind of now we're in Amsterdam. That's awesome.
Xiaoxu (15:23.834) Yeah, except for the weather. The weather is crazy these days. But yeah, for the rest, so far so good.
Benjamin (15:26.449) Thank you.
Eldad (15:29.223) You know, Amsterdam has a very long tradition of innovation on databases. CWI in Amsterdam, MonetDB, like a lot of stuff, Duck DB.
Benjamin (15:33.729) Peace.
Xiaoxu (15:41.234) And also duck DB duck DB and Python as well. Yeah. Dutch people like a program.
Eldad (15:46.655) Yes, yes.
Eldad (15:52.719) Wait till you visit Munich, by the way. But it is, it is like, yeah, it is a lot of innovation coming from there. We have like, we have a big office in Munich and we love those places. And so, yeah, it brings me warm memories. What is it like you've mentioned bringing migrating software engineering practices to data engineering? Is that a good thing?
Xiaoxu (15:54.908) Okay.
Eldad (16:21.827) Is there a way to actually rethink software engineering practices, given the fact that data engineers now interact a lot with new stakeholders, with users, with the business? I've always thought about data engineering as an evolution in many ways of software engineering, because you're much closer to the final outcome, the value of the stuff you're building.
Is it changing? Are we still kind of doing the same data engineering stuff all over? Or what do you see there? Are there new trends coming on how data engineering should evolve? Or is it basically, yes, we trust software engineering. Because if you do a subset of the data, then you do a unit test on it. And will it reflect actually that should you run it on the production? Reason I am mentioning it, we get those questions all the time. And frankly.
We don't know what to answer. So usually we don't try to intervene in how people perceive the data lifecycle. We just try to build a product or kind of to support their journey. There's so many ways to get things done right, uh, with data. But I was wondering, are we trying to rebuild our software engineering stack? Or are we also going to improve it or simplify it so we can have a bigger data engineering community because it's a lot about the community as well.
Xiaoxu (17:46.014)
Yeah, so what I see is they are not really conflicting. I feel like the software engineering skills only gives me benefits rather than, let's say I had a stereotype on something and I couldn't accept what is being done in the data engineering side. But well, to be fair, when I just...
became a data engineer in the first two months, I was indeed really surprised by the actual work the data engineer was doing. Because before I was doing this low level programming, hardcore everyday, but when I just switched, I actually spent most of the time like reading and understanding the systems rather than coding. And...
At the end of the day, what you need to do is maybe just changing a few configurations. And for me, this change at the beginning was a bit strange. Like I used to do this coding, I don't know how many lines per day, this type of thing. But in the end, it was just like, okay, changing it. Yeah. I like changing a YAML file or a click a button in the cloud. But then the more I learn, I feel like there's another dimension in it.
Eldad (18:52.939) PRs.
Xiaoxu (19:08.046) the theory and how the system works under the hood, which is also very interesting for me. So that's why in the end, I also very, very enjoyed this work. But coming back to your question, how this data engineering skills influence software engineering, another point I can see is that
the software engineering skills brings a lot of potential to a data engineering team, because of course we have those modern data stack that we can leverage without reinventing the wheel or without having a lot of engineering effort. But sometimes we are also constrained by these two links. And when we are constrained, then we need people with coding skill to expand this capability.
And I can give you like one example. When I first joined, I was trying to, like the team was trying to make a connection between dbt cloud and the airflow. At that time, the dbt cloud operator was not here yet. So I implemented our own customer operator and stuff like that. In the end, it worked well. And of course, a few months later, the operator came out and then we replaced that one with the official one.
But we can make this possible a few months earlier because we have someone in the team who can do this for us. So I feel like a data team should have at least one data engineer with software engineering background because they can really bring a lot of potentials to the data team. Yeah, this is how I feel about it.
Eldad (20:55.531) Absolutely, absolutely. Data apps, data platforms, they run 24-7. They're very unpredictable, even though we think of them as totally predictable with the modeling, and we've tested everything. But they're not. They serve other users, and those users do take those data apps to their extremes. And you're right. Keeping business continuity by having this depth in the team.
even though it's not needed on a daily basis, like a software engineering team, that can make the whole difference. And I've seen that many times spot on.
Xiaoxu (21:32.906) Yeah, and another good example is the on-call culture. A while ago, I also wrote an article about it. In the software engineering world, we do a lot of on-call duties because those softwares are mostly like APIs. They are exposed to the outside world. And if the API is done, then we need to do duty, et cetera. But in the data engineering world, the on-call culture is very important.
culture isn't really a thing there. But at the same time, I feel like we do need to have this culture. It doesn't mean that we need to be on duty 24-7. If a dashboard breaks, probably it doesn't really matter. But it's more about how data engineers should handle on-demand requests, should handle unplanned work, should handle all those kind of requests from the other teams, like incidents.
Um, on that point we can learn a lot from how the software, software engineering teams do the, do the on-call, for example.
Benjamin (22:43.177) I loved your comment on the dashboard being broken not being important. Our last guest, Wim Wischwichte, said he never saw a dashboard that had positive ROI, so it seems like everyone's...
Xiaoxu (22:46.102) I'm sorry.
Eldad (22:49.395) Hahaha
Eldad (22:53.28) He never saw a dashboard that's not broken Even if it returns results, it's deep inside broken
Xiaoxu (22:56.034) Ha ha!
Xiaoxu (23:00.808) Well, that's a true story everywhere, I guess.
Benjamin (23:07.229) The consistent data engineering show theme. No one cares about dashboards. Nice. I love that. I love that angle. So one thing I'm curious about is you have quite a journey behind you now in terms of data engineering. How do you notice today that you've actually gotten better? So for a software engineer, you get better at designing clean interfaces, getting into new code bases, all of those things.
How does growth actually look like as a data engineer to you? Like, how do you know that today you're better or more senior than you were a couple of years back?
Xiaoxu (23:43.872) You mean like how do I know if the day that becomes better or I become a day better?
Benjamin (23:49.361) with you as a data engineer, so like kind of personal growth as a data engineer.
Xiaoxu (23:53.966) Ah, okay, personal growth, not the data itself, okay. Yeah, that's a very good question. I evaluated myself from different dimensions. One is the, let's say, the landscape I know on the data technologies, although I know there are so many stuff out there. It's basically impossible to catch up with everything. But I see a few\...
core technologies which are really important for me. And I try to learn them and go deep into them as far as I can. So for example, the cloud stuff, like you don't really need to learn every single cloud provider, but getting production experience with one cloud provider can really help you learn a lot as a data engineer. And also another thing is Spark.
I found Spark really interesting and a powerful, that's a data processing engine. This is also what we do a lot at Ardian, like we have a really Spark cluster and that is my first time to learn Spark as well. So there are like so many new stuff out there as well. And another stack is Airflow, like the OG.
for the data orchestration, you must master it. There's no question. And some other data transformation, let's say, toolings, I love dbt. I feel really lucky. When I joined Dot, which was my first data engineer company, I was exposed to most of the tools that I wanted. Even now, I learned cloud, dbt, and airflow, etc.
So one is on the tooling side and another dimension is that I think in the end we try to make data as a product. So there's also a lot of connection with the business, like how I would make sure that my data fit its purpose, how it helps the business, how it impacts the business. This is something that I'm trying to.
Xiaoxu (26:09.93) yeah, be better on it as well, either by developing toolings. Like I was dreaming about this data status page. I did a POC when I was at DOT. So basically, every data pipeline or data product has its own SLA and SLO. And then we have a status page which can show the status. And then we can communicate with the stakeholders. So either by developing toolings or building my own
software skills and then to talk to stakeholders and then be better and storytelling, etc. So yeah, two parts. One is on the technology side and one is on the, let's say, data acceptance. Yeah, and also improve the data quality to be better used by the users.
Benjamin (26:59.365) So those like business skills of understanding how your data is used, kind of understanding what value it provides and so on. Would you say getting better at that is transferable between companies or like it gets hard reset every time you switch companies?
Xiaoxu (27:04.364) Yeah.
Xiaoxu (27:13.022) Yeah, it definitely gets better. So when I first joined DOT, it was like a small data team. We had six people, a data engineer, and then a few analytics scientists. So the number of stakeholders were not that much. So for me, it was more about learning new technologies rather than, let's say, scaling my business skills. But now I moved to a company which has 250 data people and a lot of stakeholders because for me I work in a product team. So in a product team I need to talk to my stakeholders on a daily basis. So I really need to understand their feelings and to make sure that what I do fits their purposes and also talk to the other data people within the company so the scale is much bigger. So I feel like in my current role I'm building my business skills much more than the technical skills.
Eldad (28:18.223) By the way, I've heard, I've seen, and I know people that actually also moved to product. So they started in software engineering, they went to data engineering, then they actually, they ended up in product because data engineering puts you in front of users much more frequently than software engineering. As you said, it puts you in front of stakeholders, you find yourself talking product and being product lead growth. And the product is the business. So
more and more companies operate like that. I think this is the big revolution that goes into how we run businesses. It's kind of, we see that from consensus over dashboards to really driving the business with a data product. And I think that's where data engineering will be pivotal and, and you're one of them, like you're the, like, really, like, uh, to me, at least kind of, uh, the perfect definition of the future of data engineering. Um, really, I'm sorry. I have to say it.
Xiaoxu (29:11.531) Oh, oh, oh my gosh.
Eldad (29:16.14) I agree with everything you're saying and with kind of your philosophy and mindset.
What can I say, I hope many more will follow
Xiaoxu (29:22.834) Yeah, but I also. Yeah, thank you. I also agree with what you say that some data people will become product people, because in the end, like you said, right now I'm in like a product team. So although I'm a data engineer, but I need to know a lot of stuff about the product. So my team is doing report for the financial controller. So I need to know all the financial products that we do at the company and all those business logics are maintained by us. So maybe one day when I get enough business knowledge, then I can easily switch to a product team and then do product. Yeah, it's definitely possible.
Benjamin (30:06.441) Awesome. Sounds good. Xiaoxu, any closing words on your end?
Eldad (30:06.527) Sounds good.
Xiaoxu (30:14.243) Yeah, so maybe like we were talking about switching from like for me, I switched from software engineer to a data engineer. But maybe like some of the audience want to do the other way around, like switching to a software engineer from a data engineer. And they might wonder like what their past experience can help them in the future.
So maybe two closing remarks on that. One is, I think as a data engineer, we have very strong end-to-end ownership because we don't just look at one single component, we look at the entire chain from end to end. And I feel like the end-to-end ownership sometimes is lacking in a software engineer. But having that really helps because we been discussing about data contract. It's mostly a contract between backend team and data team. But it can also be a backend team and a backend team. So to have those close connection between your upstream and downstreams can help you and help the company prevent a lot of issues. This is one, let's say, advantage. And another is that maybe not that common is if you work with cloud, you probably let's say know a lot about the cost of optimization side. And you can also bring this mentality to a software engineering team by optimizing the resources and having this kind of mindset can also help the team grow as well. So it works both way. Because it's never the case that one is the subset of the other. Like people can always switch around and learn from each other.
Eldad (31:59.619) Impossible. It's data engineers have cost consciousness embedded in their DNA. Software engineers, they work in budgets and come and have a lot of software excuses. Just kidding. I love software engineers. Thank you. Thank you for the insightful last kind of ending comments.
Xiaoxu (32:05.474) Hehehehe
Xiaoxu (32:09.331) Yeah.
Benjamin (32:19.237) Definitely. Thanks. Thanks so much for being on the show. We had a great time. Enjoy your evening. And yeah, we look forward to reading your learning course soon.
Xiaoxu (32:30.334) Yeah, thank you so much. Thanks for the podcast. Thank you.
# Transitioning Scopely’s 5.5 PB Data Platform to the Modern Data Stack (/blog/transitioning-scopelys-5-5-pb-data-platform-to-the-modern-data-stack)
Should data engineering AND BI be handled by the same people? According to Jonathan Palmer, VP Data Platform at Scopely – YES. By Analytics Engineers.
His team of Analytics Engineers is in the final stages of transitioning 5.5 PBs of data which include 15B events per day to the modern data stack. Tune in to learn how they did it.
Listen on [Apple Podcasts](https://podcasts.apple.com/us/podcast/transitioning-scopelys-5-5-pb-data-platform-to-the/id1561927688?i=1000557232989) or [Spotify](https://open.spotify.com/episode/5VqeaVAihFQTTOvAMs0dwm?si=BJR4RAjGS9aZtx6_AIZ4LQ)
Boaz: Hello, everybody. Welcome to another episode of the Data Engineering Show. I am Boaz. I am here alone today without Eldad. But with me is Jonathan Palmer. Hi, Jonathan, how are you?
Jonathan: Hey, I am good. How are you doing?
Boaz: Very good. Jonathan is the VP of the data platform at Scopely. Scopely, if you do not know, is a gaming company, a mobile-first gaming company. They did games like the Walking Dead and Scrabble Go, Looney Tunes: World of Mayhem, and many others. Prior to that, Jonathan was head of BI at GoCardless. Before that, he also spent more years in gaming at King. And in his background, combines both business intelligence and data engineering.
Jonathan, what did I miss about your background that is important to mention, or did I get it right?
Jonathan: I think you got it pretty much spot on.
Boaz: Awesome. So, tell us, what do you do Scopely?
Jonathan: I work for a part of Scopely called Playgami. Playgami is the platform that powers our games, including their analytics capabilities, but enables them to do everything from building the game provisioning infrastructure to experimentation to CRM, push messaging campaigns, and everything in-between, most of which are powered by data in one form or another. So, it is my job to kind of manage the strategy for how that data platform scales and grows and powers all the games, the tech vendors we work with, the people we hire, the products, ideas that we have from one way or the other, that is my job to make sure that all kind of works together.
Boaz: We say platform in the context of what you do, essentially the data platform or the entire platform on top of which the games run?
Jonathan: Yes, my bit is the data platform but sits within the entire platform on which the games run.
Boaz: You mentioned, we see behind you in the background the Playgami. So, tell us a little bit about that. Are you guys positioning this as something to become a publicly for-out, about kind of name for the data platform? What is the story of Playgami under Scopely?
Jonathan: Playgami is an internal brand at this point, but as Scopely works with, we have our organic titles, once that was built by Scopely studios, but we also publish games, third parties
and we also have kind of acquired a few businesses over recent years, like, FoxNext with Marvel StrikeForce and GSN games just recently. So, Playgami in some ways is part of the offering and the value proposition of Scopely. It is a super cool platform that enables you to manage the full life cycle of your game and data is kind of the lifeblood of that.
Boaz: I would love to spend more time on that. Before that, walk us a little bit through your background, you did an interesting journey. You started from more on the business intelligence side, then touched many things, ended up now being sort of VP of the platform. Walk us through your journey.
Jonathan: I had a strange start, so I studied ancient history. So, I had a kind of zero tech background.
Boaz: So you studied Hadoop.
Jonathan: No Microsoft Access. Yeah, I got my break in tech from a company called Spark Data in Bristol, UK, which deliberately employed people from non-tech backgrounds and taught them how to be software engineers, and that is working on data-driven systems. That is where I learned SQL and over a variety of years, I got closer and closer to the data side. And then, the kind of inflection point I guess was when I joined King, as you mentioned, then I was a clickfree developer. But, I became more and more interested in the life cycle of data, how it works in enormous complex, fast-paced businesses, like in gaming, and that took me into some kind of roles like a principal engineer and product manager and then Director of data analytics side. More and more thinking about the strategy and how all of this tech fits together with and ultimately what the business is trying to do.
Boaz: That is amazing. You do not meet many data/historians.
Jonathan: No, there were not many jobs in ancient Rome to go to. So, I went down on a different path.
Boaz: Yeah. I know it is the kind of thing you hear about companies or people trying to encourage each other. Let us give a shot to people who do not come directly from tech or from engineering education to take on these positions. But, the truth of the matter is that it remains very rare still. So, very interesting to hear, and it has worked amazingly well for you. I wish we would see that more often.
Jonathan: Yeah, it is an interesting model. It definitely gives a high level of imposter syndrome, I can tell you that.
Boaz: I am sure it does. At Scopely let us dive into the platform which is super interesting. Before that, how many employees does Scopely have worldwide and how big are the data-related teams?
Jonathan: Scopely is growing super fast and by the time I say a number it is probably already out of date, but last time I looked it was like about 1700 people in 17 markets all over the UK, everywhere from LA to Dublin, London, and Barcelona where I am. And, the data part of that I would say is roughly between 60 and 70 people in data-related roles that cover a data infrastructure team, data applications, core BI, data science, embed, mini data teams or product analysts and analytics engineers into games and verticals as well. So, across all of that organization, it is about 60 to 70 people.
Boaz: And at Playgami as the platform you described, is that the new initiative, was that there from the outset.
Jonathan: Every game requires a platform to run on it. So, the platform itself is I guess is as old as Scopely, but over the last couple of years we have really refined how we think about that platform and increasingly sort of delivering more and more competitive differentiating features and so on, and it is becoming a stronger and stronger brand within Scopely itself.
Boaz: So, when did the Playgami brand sort of launch
Jonathan: I think in the last 12 to 18 months, I would say.
Boaz: So, what does the data stack look like?
Jonathan: We are in the final stages of moving from a legacy stack to our current stack. I mean future proof stack, so to speak, is principally BigQuery, DBT, ELT with Airflow orchestrating DBT jobs, and then the main kind of user touchpoint is Looker. But, we ingest data originally as Parquet files into S3 on AWS and then ship everything across on a kind of micro-batch basis over to BigQuery.
Boaz: What does the legacy stack look like?
Jonathan: Before we made this change, the data ingested, the raw events are ingested into S3 and then, we used EMR and we kind of managed Spark and then parlor ourselves on AWS infrastructure, for ETL and also used as querying the data directly and then the principal means of access was Tableau.
Boaz: Okay. So, Tableau or the old stack to the current new one.
Jonathan: Yes. Yeah, exactly.
Boaz: What were the tipping points? Can you describe that? The moment in time when the company decided we have to start looking or build out a new stack because this one does not cut it anymore.
Jonathan: So, I joined Scopely a couple of years ago. And, when I joined, one of the first things I did was to kind of look around and see, okay, how is everything scaling both in terms of the data warehouse, the computational workloads, also I used to experience and the team had done a super good job both running the Spark and Impala set up, but we were seeing a fairly growing number of outages and one of my Northstars, especially as we already made the decision, the Scopely made the decision before I joined to move towards Looker, average career duration in Looker is one of my kind of Northstar metrics. And, when we started off, it was around five minutes, using our spark infrastructure, which is clearly not what the user wants, and moving to BigQuery, we have now got that to 30 seconds and in terms of the actual sort of distribution of that, like most of our queries are in single figures.
Boaz: You guys have been in the AWS shop, I am sure making the decision to go into the GCP direction was not an easy one. Tell us a little bit about that decision process there, and then decide to go for it.
Jonathan: Yeah. I do not think it was too controversial, I think Scopely had ambitions for a while to kind of consider a multi-cloud approach, and this was one of the more obvious opportunities to do that. I had worked with the Google stack in King and GoCardless. So, I already rated it pretty highly. But, fundamentally, we went through a sort of a super rigorous market search. You know, we looked at all of the big players, Databricks, Snowflake, and Redshift. We put them all through their paces based on our particular use case. At the end of that evaluation process, there was a clear winner across a variety of dimensions, not just speed, but scalability, cost efficiency, a big one for me is a barrier to entry. So, Spark is super powerful, but it is quite hard to hire people who straight away get Spark and can be effective in that world whereas the barrier to entry, I think BigQuery is a lot lower.
Boaz: The intention though is to stay multi-cloud and keep using both.
Jonathan: Yeah. I mean, Scopely leverages AWS for all sorts of parts of the gaming side of the infrastructure and does it extremely well. So, at the moment, I think we have got a good balance between using the strengths of the various platforms in the right place.
Boaz: Another interesting thing with the migration, is they move from Tableau to Looker. Oftentimes, we see people, not that they have anything against moving out of Tableau, but the sheer amount of reports that already exist makes it tough for people to agree on moving to a new tool. How did you go about that? Is there a process of recreating or are you saying new ones in the new platform, old one, staying Tableau?
Jonathan: It is funny. I think for me it was non-controversial because the sheer number of Tableau reports was one of the main reasons to move. I think we deprecated somewhere between 80% and 90% of all the Tableau dashboards created in Scopely's history as part of the move to Looker. There was a huge amount of debt there. That was not being used. Basically creates confusion, that sort of single source of truth becomes impossible. So, it was a very compelling reason to tear it up and start again.
Boaz: It is like moving apartments, you start, going drawer by drawer.
Jonathan: Exactly. It is the housekeeping.
Boaz: The trash bin there and then what goes into the trash and what stays in the drawers.
Jonathan: Exactly, everybody talks about spring cleaning, but rarely gets round to doing it. So, it is the best opportunity to do that.
Boaz: It takes courage too. At the end of the day, deprecating 80%, 90%, many companies, it takes courage to some extent. Even though it might be the completely rational thing to do, and the correct thing to do, I feel many companies or maybe we get emotionally attached to so many work hours behind those dashboards. Throwing them out oftentimes makes people feel what do I do it for?
Jonathan: That is true. I would say it is the courage you could see but in our case, we had a relatively new team and I think in a pretty clear sense that the way we were operating was not going to scale for a Scopely that is 10X, what it was then. So, we had to do something and the value of Looker in that was really clear to the team, I think.
Boaz: From a skill set perspective, how did you go about that? I mean, did you guys need to hire a lot of new people, or were you able to use the same people who were in charge of the existing stack, take over and also move to the new one happily.
Jonathan: We have been hiring rapidly anyway, so it is a kind of mixture of both. But, all the people who are working on that old stack, transition pretty seamlessly and impressively to the new stack. I put a lot of time and effort personally in that beginning stage of the sort of evangelization, education piece so that when people started working it, they could operate pretty autonomously. Because just having me as the only person who knows how to look in a large organization, is not very scalable. So, I distributed that knowledge super fast, but also getting that sort of buy-in and now feeling like, okay, this is actually going to make our lives better and we want to learn and we want to adapt this. So, generally, that went pretty well and people have grabbed it and really embraced the opportunity. And when it comes to hiring, if you only target people who are working with your stack, that technology is already, like your target addressable market is super small even today, but generally, we are looking for different attitudes and abilities, we are not restricting it to just the people who worked on that.
Boaz: What data volumes are you guys dealing with? For example, in BigQuery, how much data are you looking at?
Jonathan: I believe there are about five and a half petabytes. I mean we ingest somewhere between like 12 and 15 billion events a day. And, then in terms of kind of consumption, we have got about one and a half thousand Looker users and they are running about 40,000 queries a day on top of that data.
Boaz: Wow. That is impressive. And, so what kinds of use cases are running on top of the platform?
Jonathan: A variety of things the most digestible of which is the portfolio reporting. So the Northstar KPIs that the whole business uses to understand retention engagement, acquisition, monetization, are the default standard of this is how these are the gold standard measure of those things. Then, there is this deep-dive analysis in the games themselves looking at Live-Ops performance, level progression, characters, battles, depending on the genre. We have user acquisition and ads use cases. So whether it is managing the sort of budget, allocation with UA campaigns or actually we are also powering the forwarding of other events onto UA networks and things like that.
Boaz: I stop you right here for a second. Across these three use cases, for example, how are the people split? Are we talking about analysts embedded in the respective departments? Are we talking about and then joint horizontal data engineering that takes care of, particularly under the hood, or walk us through the human structure.
Jonathan: Got yes. Our sort of centralized structure is around the data infrastructure, fundamentally running the platforms that power the ingestion, the power, how we think about, the relationship between AWS and GCP. Then there is the core BI team, their job is to build those kinds of baseline abstractions, those things that power those portfolios, KPIs, for example, but also the kind of building blocks that other people can extend, Looker explores and things like that, the responsibility of that team.
Boaz: So, for example, if one of the Northstar KPIs, you want to introduce a new one, which people are involved? Who is the analyst that defines the KPI and where he or she sits.
Jonathan: Fundamentally, the engineering will be done by the core BI team and the guidance of that team comes from the product manager of that team. And, they partner closely with this kind of strategic analytics organization, which is not a part of my organization, is part of the commercial side of the business and their job, part of what they do is thinking about what we think about what success looks like in Scopely and how do we measure that? And so typically it will be then just the driving that KPI, for example, at the moment we are working on bringing kind of a better lens to reactivation of players. So they are the ones thinking about the definition of that. And, then my core BI team will actually be the ones who are going to build that.
Boaz: Amazing. So, in the core BI team, there are both data engineering skill sets and BI skill sets?
Jonathan: Yeah. And, like all the technical roles generally called analytics engineers, and in line with that kind of move towards people who can move pretty seamlessly between building data models or writing LookML and constructing Looker products. So, daily people have different specializations, but our mission is to have people who are pretty able to move across the three DBT.
Boaz: Okay. So, it is not like, person A stops at BigQuery, and person B starts at Looker. You have the same people who could do both.
Jonathan: Yeah. And like, we have come from that model where it was more of a kind of relay handover and more and more I am trying to bring it together so there is greater diversity across.
Boaz: Interesting. Do you consider the transition over? Is it done? Is it still in progress? How much is left?
Jonathan: The sort of backbone of the transition is done. From the moment we signed a deal with Google to kind of first business value was three months. So, we worked backward through the migration. So we moved all whether that was table dashboards that still existed at that point, Looker, or direct SQL query access. All of that happened first so that we could take the pressure off teams, kind of keeping the lights on the old stack and delivering value as early as possible relative to the meter that started running. And, then we worked back upstream through ETL. So, all of our ETL, whether it is kind of game-specific or the core side of things, was all migrated in less than 12 months. We have deprecated pretty much all of the AWS side of things. And the only thing, I would say it is not so much related to migration, but the thing we are working on now is Now streaming ingestion directly into BigQuery so that the data is as fresh as possible and as fast as possible.
Boaz: What other new initiatives are lined out. Where do you want to see Scopely as the platform a year from now?
Jonathan: We have got a couple of big focuses, but I think one of the first ones at this point like I am super happy with the scalability of our platform. It does what it needs to do. We will scale naturally as Scopely grows and we add more and more use cases, more games, more studios, more data. Focus now is kind of on quality and observability. So, we can do great things for the data, but we rely heavily on that data being clean, and fundamentally that comes a lot down to the sort of tracking side of things. And at the moment, it is pretty easy for a game to implement their tracking completely incorrectly, and then we will throw analytics engineers at the problems to tidy it all up, and get into the shape we need it. So, we are going kind of upstream now and looking at the way we do tracking whether it is the semantics or the kind of the tooling that we give to game teams to help with that so they can get it easily, right like the first time, and the kind of time to insight, which is kind of another, my Northstar metrics is as quick as possible.
Boaz: But, how are you literally doing that? Are you using any tools, libraries in these new approaches? How are you going about that change?
Jonathan: Some of it is about education in the process, some of it is about building tooling. For example, with that focusing on a sort of streaming ingestion into BigQuery, the faster data available in BigQuery, the more we can leverage tools like Looker to enable people to QA and inspect their events, using the same business logic. Before it was a bit like, if you want to see the raw event, in real-time here, you can, but if you want basic business logic as well. Well, that is a batch, that is over here. We are bringing those two worlds together. But, we are also looking at, I mean, we have had conversations with a variety of vendors in the space over the years. No decisions have been made, but I think as interesting companies like Montecarlo and between them are attacking this thing from a different angle, either looking at commoditizing how companies build tracking plans and making that super easy, rather than having to build that yourself or the Montecarlo side, that observability piece that plugs into your existing stack. So, all of these things are kind of we are looking at them and working out, where the kind of build, the tradeoff is and what it is that is the kind of Scopely singularity about the problem.
Boaz: Yeah. I mean, definitely looking into quality observability is taking the data engineering world by storm so to say with vendors like Montecarlo, Databand, many others. It is interesting to follow how that plays out and who sort of will stay on top and, but most importantly, that seeing everybody adopting these processes and mindsets. Because at the end of the day, it is something that can make a difference. So, much time is wasted by figuring things out downstream when they are wrong. It is a natural next step for our market, I guess.
Jonathan: Absolutely. And it is great to see and it is interesting working with software engineers and noticed that when they are talking about observability of platforms, three to five years ago, and then you start to see the same thing crop up in data engineering certainly afterward. I think there is this kind of virtuous cycle of these patterns that come over from software engineering and get applied to data engineering to our use cases as well.
Boaz: What about data science? Is there a dedicated science department? Tell us a little bit about that?
Jonathan: Yeah. We have a data science department. They work principally on our kind of predictive offering of the platform. So, they are using machine learning, many working with Dataproc on GCP to do things like LTV predictions, churn predictions but we are moving more and more to a world where we do not just kind of predict what is going to happen but prescribe what would be the best course of action to avoid the negative outcome as well?
Boaz: How big is that?
Jonathan: Data science is relatively small. It is like less than 10 people. But working with software engineers and certain product designers around them to build these products.
Boaz: Well, you have mentioned earlier there are 60 people dealing with data, I guess. What sort of processes are in place to stay horizontally aligned. Is there sort of a weekly Guild meeting or anything like that, how do you make sure everybody knows what is happening around with data?
Jonathan: Yeah. So, with these kinds of embedded teams, we have a matrix management structure. So, they both report into the games, for example, so that their direction is coming mainly from the roadmap of the games and then the horizontal management is kind of looking at that and making sure we move in a consistent, coherent kind of way. There are always challenges when you have got kind of two dimensions competing, but by and large, I think we have done a pretty good job, especially with things like the degree migration, for example, coordinating that across multiple embedded teams and doing it in a consistent way, kind of definitely stress-tested in that organization and the outcome has been pretty good, I would say.
Boaz: Awesome. Specifically, what does your direct team look like on the platform?
Jonathan: Typically, I have a few directors who report to me, both those horizontal leads that we are talking about or directors of analytics engineering working in core BI, for example, and the leadership of the data infrastructure team. Plus I have, a group of product managers who report to me who cover a variety of things from, core data products to data applications, going back to that point about the Playgami platform, the most visible product of the Playgami platform is Playgami console, which is the fundamentally the UI that you go to work with everything that a game team needs. So, how we embed data applications, capabilities into that, that also got product focus as well.
Boaz: Awesome. So within that journey from Scopely in the last two years, maybe specifically with the GCP move or in general, what in retrospect, if you rewind, would you have avoided? What would you do a little bit differently to make it even smoother?
Jonathan: Great question. Well, I think one thing was when we were migrating, whether we were doing the moving things like tableau dashboards and Looker products or game-specific data pipelines, or core data pipelines that created bottlenecks and pressure on different profiles and parts of the team. If someone is migrating, your ETL is not building the things you need to leverage in Looker, which affects roadmaps. By and large, I think we got the balance okay. But, one of the reasons, one of the big drivers for why we were moving to this analytics engineering mindset and having people who can kind of reverse the whole stack more easily is to avoid that, and do not intend to migrate daily warehouses every couple of years. But, it brought home that point around bottlenecks and kind of linear dependencies and making sure. So, if I could do it again, it would have been, I am not sure the circumstances of that to do it, but it would have been great to kind of get that analytics engineering thing first and then the stack second but that is still an argument. That is a bit the sort of tail wagging the dog.
Boaz: Yeah. You mentioned you had been using Google at King as well and in between, you know, a few years past, and then you took on Google again at Scopely. How did you feel the platform's improvement over those years? What is better at the GCP stack today versus when you were still at King?
Jonathan: One of the big things I guess, is when I was still at King, Luca was still an independent company, and I was not super surprised when the acquisition happened. But you can really see now how those things are converging and complementing each other. I would say there has also been progress made in a bunch of different areas. Just this, the simple things like the BigQuery UI and how easy that is to work within, but also some of the things under the hood around, security and privacy permissions, all of these things have come on some since I first started playing around with BigQuery what must have been good four or five years ago now.
Boaz: Okay. This has been super, super interesting, Jonathan, I cannot thank you enough, maybe a good note to finish with would be what advice would you have for people like yourself who are not educated with the technical background and are considering a move into tech?
Jonathan: Great question. One thing I think is super cool now, which did not happen so much at the start of my career, but you have so many Bootcamp companies now who can help accelerate that and take different views on the problem, both organizations like Codop for example, who works with women and underrepresented minorities, there are so many doors that are opening now. So I think that is one avenue. The other thing I would say is working or aiming at companies that use them on the data stack. They use tools like DBT or BigQuery or snowflake, for example, and they tend to be startups, and those companies tend to move super fast, super agile like in my experience, someone at the start of their career can learn more in a year working in that sort of environment than they can five years working in a sort of big monolithic legacy organization. So I would aim for those roles, and they also tend to be much more open-minded about skill, I think, in their hiring as well.
Boaz: Amazing. Great piece of advice. I wish I could go back in time, study history, and take your advice.
Jonathan: Also by accident rather than by design but it has worked out okay.
Boaz: Jonathan, thank you so much.
Jonathan: Thank you.
Boaz: See you around.
Thank you for listening, everybody.
# Unlock Conversational Data Interaction: Firebolt MCP Server for Advanced LLM Integration (/blog/unlock-conversational-data-interaction-firebolt-mcp-server-for-advanced-llm-integration)
**TL;DR**
The Firebolt MCP Server enables seamless integration between Firebolt cloud data warehouses and AI tools like Claude or Copilot via the Model Context Protocol (MCP). It allows LLMs to securely query data, explore schemas, and access documentation, streamlining tasks like SQL generation and code automation. This unlocks faster, smarter workflows and paves the way for autonomous AI agents to perform high-speed, data-driven research directly within Firebolt. Click [here](https://github.com/firebolt-db/mcp-server) for access to the GitHub repo.
As data engineers, we constantly seek ways to streamline workflows, optimize query performance, and accelerate the path from raw data to actionable insights within our cloud data warehouse environments. Today, we're excited to introduce a powerful new way to do just that: the **Firebolt MCP Server** — a bridge between your Firebolt cloud data warehouse and the AI tools you already use, like Claude, Copilot, Cursor, and others.
This new offering implements the Model Context Protocol (MCP), an open standard designed to securely and effectively connect Large Language Models (LLMs) to diverse data sources and tools. Think of MCP as a standardized API layer for AI, enabling seamless, contextual communication between language models and external systems — including your Firebolt environment.
With the Firebolt MCP Server, LLMs can go beyond general-purpose tasks to perform specific, high-value interactions directly with your Firebolt databases. Whether you're querying data, exploring schemas, or building intelligent assistants, this is a major step forward in making AI a first-class citizen in your data engineering workflows.
## Why MCP? Standardizing LLM Tool Interaction [#why-mcp-standardizing-llm-tool-interaction]
Before MCP, integrating LLMs with specific tools often required bespoke, complex solutions. The MCP ([Model Context Protocol](https://modelcontextprotocol.io/)) introduces a standardized interface, simplifying how AI models discover and utilize available capabilities. For data engineers using Firebolt, this translates to a secure, reliable way to grant LLMs controlled access to specific functionalities.
The Firebolt MCP Server exposes a curated set of [tools](https://modelcontextprotocol.io/docs/concepts/tools) tailored for data engineering workflows:
1. *firebolt\_docs*: Provides the LLM with direct, programmatic access to Firebolt's official documentation, including SQL reference, function definitions, data type specifications, and architectural guides. This ensures the LLM uses accurate, up-to-date information when assisting with syntax or concepts.
2. *firebolt\_connect*: Allows the LLM to discover the user's accessible Firebolt environment, listing available accounts, databases, and compute engines. This is crucial for contextual awareness before attempting operations.
3. *firebolt\_query*: Enables the LLM to execute SQL queries directly against a specified Firebolt database using an appropriate engine. The server manages secure connections and returns results to the LLM.
Authentication is handled securely via [Firebolt service accounts](https://docs.firebolt.io/Guides/managing-your-organization/service-accounts.html) (client ID and secret), ensuring credentials are never stored persistently by the server and access adheres to configured permissions. The server itself can be run easily via Docker or as a standalone binary.
## Practical Applications for Data Engineers: AI-Accelerated Workflows [#practical-applications-for-data-engineers-ai-accelerated-workflows]
Let's move beyond the theoretical and explore concrete ways the Firebolt MCP Server, powering an LLM like Claude, can become an indispensable part of your toolkit.
### 1. Context-Aware Documentation Lookup [#1-context-aware-documentation-lookup]
Navigating extensive documentation during development or troubleshooting can interrupt flow. The MCP server allows the LLM to fetch precise technical information on demand.
**Technical Flow:** A user asks, "What's the difference between a login and a user in Firebolt?" Instead of relying on memory, the LLM uses `firebolt_docs` to search and read the relevant documentation in real time. It extracts and summarizes the key distinction: logins handle authentication, while users define access within specific accounts. The result is accurate, contextual guidance drawn straight from the source.
### 2. Natural Language to High-Performance SQL [#2-natural-language-to-high-performance-sql]
While proficient SQL is essential, translating complex business questions into optimized queries takes time. MCP enables LLMs to act as intelligent translators, bridging natural language requirements with Firebolt's SQL dialect.
**Technical Flow:** A user asks, "Which age category of players experiences the most error codes during gameplay in the UltraFast database?" The LLM retrieves documentation via `firebolt_docs`, connects to Firebolt, and crafts a SQL query joining gameplay logs with player demographics. Using `firebolt_query`, it computes total and average errors per player by age group, ranks the results, and returns a clear summary: users aged 46–55 see the most errors. It then drills down into top error types by age, surfacing actionable insights — all from a single prompt.
### 3. Accelerated Client Code Generation [#3-accelerated-client-code-generation]
Integrating Firebolt queries into applications or scripts often involves repetitive boilerplate code for establishing connections, executing queries, and handling results.
**Technical Flow:** An engineer asks the LLM to connect to the UltraFast database, find all tables with "playstats" in the name, and generate a Go program that fetches and scans data into typed structs. The LLM retrieves Go SDK documentation via `firebolt_docs`, connects to Firebolt, queries `information_schema` for relevant tables and columns, and builds `models.go` and `main.go`. The result: fully functional Go code that connects to Firebolt, executes SQL, and processes results — all tailored to the live schema.
## Getting Started with Firebolt MCP Server [#getting-started-with-firebolt-mcp-server]
Integrating this powerful AI capability into your workflow is straightforward:
**1. Prerequisites:** Ensure you have a Firebolt service account with its client ID and secret.
**2. Installation:** Deploy the Firebolt MCP Server using the provided Docker image or download the appropriate binary for your environment. Pass your credentials securely as environment variables or command-line arguments.
```bash
Example using Docker
docker run --rm -i \
-e FIREBOLT_MCP_CLIENT_ID="your-client-id" \
-e FIREBOLT_MCP_CLIENT_SECRET="your-client-secret" \
ghcr.io/firebolt-db/mcp-server:latest
```
**3. LLM Client Configuration:** Configure your preferred MCP-compatible client (e.g., Claude Desktop, GitHub Copilot Chat in VSCode, Cursor editor) to connect to the running MCP server instance. Refer to the [MCP Server README](https://github.com/firebolt-db/mcp-server) and client-specific documentation for detailed steps.
**4. Engage:** Start interacting with your Firebolt data warehouse through your AI assistant!
## Expanding Horizons: Enabling Autonomous AI Agents with High-Speed Analytics [#expanding-horizons-enabling-autonomous-ai-agents-with-high-speed-analytics]
While the above use cases significantly enhance data engineering productivity, the combination of MCP and Firebolt unlocks a more profound potential: enabling autonomous AI agents capable of conducting complex, data-driven research at machine speed.
Imagine an AI research agent tasked with uncovering subtle correlations within petabytes of IoT sensor data or web traffic data stored in Firebolt. MCP provides the crucial interface for this agent to interact with the data warehouse autonomously:
* **Self-Directed Exploration:** The agent can understand the available datasets and schemas within Firebolt.
* **Hypothesis Generation & Testing:** Based on its objectives and initial data exploration, the agent formulates hypotheses. It then designs and executes sequences of complex SQL queries via Firebolt MCP Server to test these hypotheses against massive datasets.
* **Iterative Refinement:** The agent analyzes the query results returned through MCP. Critically, Firebolt's ultra-fast query performance is essential here. Sub-second response times on terabyte-scale datasets allow the agent to iterate rapidly – adjusting hypotheses, formulating new queries, and diving deeper into the data far faster than any human analyst could.
* **Knowledge Synthesis:** The agent leverages Firebolt documentation to ensure its generated queries are syntactically correct and utilize Firebolt features optimally (e.g., geospatial functions, array processing, index utilization). It integrates findings from multiple queries to build a comprehensive understanding.
This autonomous loop – explore, hypothesize, query, analyze, refine – powered by the MCP interface and accelerated by Firebolt's high-performance analytics engine, transforms the scale and speed at which data-driven research can occur. The agent isn't just retrieving data; it's performing in silico experiments, discovering patterns, and generating novel insights directly from the underlying data warehouse, pushing the boundaries of automated scientific discovery and business intelligence.
Our release of the Firebolt MCP Server is just the beginning. We're closely following the latest advancements in AI and working on new features specifically designed to power the next generation of data and AI applications. Whether you're building analytics pipelines, creating intelligent data products, or enabling AI agents to autonomously explore your data warehouse, Firebolt is the high-performance foundation you need.
We're excited about what's next. Stay tuned — more powerful capabilities are on the way
# Unlocking Simplicity and Security: Firebolt’s New LOCATION Object (/blog/unlocking-simplicity-and-security-firebolts-new-location-object)
**TL;DR**
In the past, working with external data in Firebolt — such as [External Tables](https://docs.firebolt.io/sql_reference/commands/data-definition/create-external-table.html), [COPY FROM](https://docs.firebolt.io/sql_reference/commands/data-management/copy-from.html), [COPY TO](https://docs.firebolt.io/sql_reference/commands/data-management/copy-to.html), and [TVF](https://docs.firebolt.io/sql_reference/functions-reference/table-valued/)s — meant manually embedding cloud credentials (like AWS keys) into every query or table definition. This was a tedious and risky practice. The new LOCATION object replaces this with a secure, reusable way to store and manage credentials across databases, engines, and operations. It simplifies workflows, strengthens security, and lays the foundation for scalable, future-proof integrations.
## The Credential Problem We Had to Solve [#the-credential-problem-we-had-to-solve]
Until now, working with external data in Firebolt required manually embedding credentials into every external table, COPY statement, or TVF query.
This approach made life harder for engineers and security teams alike:
* Credentials were duplicated across queries and projects.
* Rotation of secrets was complex and error-prone.
* It was impossible to separate **who can see** the credentials from **who can use** them — creating unnecessary exposure risks.
It was clear that we needed a better model — one that combined **security**, **simplicity**, and **operational efficiency** at scale.
## The Breakthrough: Introducing the LOCATION Object [#the-breakthrough-introducing-the-location-object]
Firebolt's LOCATION object provides a clean, powerful solution.
A LOCATION is a secure, reusable object that stores:
* Credentials for external access.
* Source-specific configuration (such as S3 URLs).
* Optional descriptive metadata.
Instead of repeating credentials across SQL scripts, users define a LOCATION once — and reference it wherever needed:
```sql
COPY INTO my_table
FROM my_data_location
WITH
OBJECT_PATTERN = '*.parquet'
TYPE = PARQUET;
```
Currently, LOCATION supports **AWS S3** as the first storage integration.
But the design is **extensible by nature** — built to seamlessly support a wide variety of authenticated sources in the future, including:
* Apache Iceberg REST catalogs
* Apache Kafka clusters
* Google Cloud Storage
* Third-party APIs requiring credentials
As we extend Firebolt's language features, we'll also keep extending LOCATION to integrate with more and more technologies. So LOCATION is the foundation for how Firebolt will securely and efficiently connect to external systems going forward.
## A Closer Look at LOCATION [#a-closer-look-at-location]
### Centralized, Account-Level Scope [#centralized-account-level-scope]
LOCATION objects are created at the account level, not the database level.
This design ensures:
* Easy sharing across multiple databases and engines.
* Simplified credential management across teams and environments.
* One place to rotate credentials safely, affecting all dependents automatically.
#### Example: Using a LOCATION across databases [#example-using-a-location-across-databases]
```sql
-- While Using Database A: Create a LOCATION object. This isn't tied to the database, but the account.
USE DATABASE a;
CREATE LOCATION shared_data WITH
SOURCE = 'AMAZON_S3',
CREDENTIALS = (AWS_ROLE_ARN = 'arn:aws:iam::123456789012:role/S3Access'),
URL = 's3://cross-db-example/';
-- While Using Database B: Referencing the same LOCATION just works.
USE DATABASE b;
CREATE EXTERNAL TABLE ext_table_b (
id INT,
name TEXT
)
LOCATION = shared_data
OBJECT_PATTERN = '*.parquet'
TYPE = PARQUET;
```
🔐 Even though databases a and b are isolated, they can reference the same LOCATION — enabling secure, centralized access management.
### RBAC-Driven Access Control — With Credential Separation [#rbac-driven-access-control--with-credential-separation]
Firebolt extends its RBAC system to LOCATION objects, introducing fine-grained privileges:
* CREATE LOCATION
* MODIFY LOCATION
* USAGE LOCATION
This is not just about permissions — it fundamentally improves credential security:
* Only creators and modifiers of a LOCATION can view or change the credentials.
* Users with USAGE privilege can access the external data without ever seeing the credentials themselves.
#### Example: Restricting LOCATION modification, enabling safe usage [#example-restricting-location-modification-enabling-safe-usage]
```sql
-- Admin creates a secure LOCATION object
CREATE LOCATION secure_loc WITH
SOURCE = 'AMAZON_S3',
CREDENTIALS = (AWS_ACCESS_KEY_ID = 'key' AWS_SECRET_ACCESS_KEY = 'secret'),
URL = 's3://secure-data/';
-- Admin grants only usage rights to analysts
GRANT USAGE ON LOCATION secure_loc TO role_data_analyst;
-- Analysts can query external data:
-- (they don't see credentials and cannot modify the location)
COPY INTO analytics_table
FROM secure_loc
WITH OBJECT_PATTERN = '*.parquet' TYPE = PARQUET;
```
🧑💼 If an analyst tries to alter the LOCATION:
```sql
ALTER LOCATION secure_loc SET URL = 's3://oops/';
-- ERROR: location 'secure_loc' does not exist or not authorized.
```
🔐 This ensures a clean separation:
* **Admins** manage credentials.
* **Users** access data — without compromising security.
There is now a clear, enforced boundary between the people who manage secrets and the people who use the data.
This separation dramatically strengthens credential hygiene in large organizations, ensuring that only authorized users handle sensitive authentication details.
Administrators can easily grant or restrict access at both object-level and account-level granularity, aligning with security best practices.
### Encryption Everywhere [#encryption-everywhere]
Every LOCATION object:
* Encrypts credentials at rest using Firebolt's KMS integration.
* Caches decrypted credentials securely (for a limited window) to balance performance and security.
* Masks sensitive information even in system metadata views.
Credentials are never stored or surfaced in plaintext after creation.
### Safe Dependency Management [#safe-dependency-management]
Firebolt makes sure that managing LOCATIONs is safe:
* Locations can't be dropped if they are still in use by external tables or other objects.
* Concurrency between create/alter/drop operations is safely controlled.
This ensures strong consistency and safety even in complex environments.
## LOCATION in Action: Simplifying Daily Operations [#location-in-action-simplifying-daily-operations]
Here's what changes for users:
| Before | After |
| ------------------------------------------ | ------------------------------------- |
| Embed credentials inside every query | Reference a secure LOCATION object |
| Rotate secrets manually across many places | Rotate once at the LOCATION level |
| Risk of leaking credentials in SQL | Credentials hidden and encrypted |
| Difficult permission management | Fine-grained RBAC separation of roles |
External tables, COPY TO/FROM operations, and TVFs are all LOCATION-aware making external access easier, safer, and more scalable across the platform.
## Future-Proof Extensibility: Building for Tomorrow's Needs [#future-proof-extensibility-building-for-tomorrows-needs]
```sql
SELECT *
FROM read_csv(
location => my_s3_location,
object_pattern => 'daily/*.csv',
header => TRUE
);
```
Although today LOCATION supports only AWS S3 storage, it is designed to grow into Firebolt's universal external authentication layer.
Over time, we plan to extend LOCATION support with:
* Apache Iceberg REST-based catalog authentication (for next-generation lakehouses)
* Apache Kafka integrations with SASL/SSL credentials
* Azure Blob Storage and Google Cloud Storage secure access
* Any third-party APIs that require token, key, or role-based authentication
RBAC, encryption, dependency management, and system integration are ready to scale with new source types.
By investing early in LOCATION's flexibility, Firebolt ensures that connecting to new data sources — no matter how diverse — will always remain secure, simple, and efficient.
## Conclusion: A Foundation for Simpler, Safer, Scalable Access [#conclusion-a-foundation-for-simpler-safer-scalable-access]
Managing credentials in Firebolt has just become a lot simpler.
The LOCATION object is a foundational improvement to Firebolt's data access model. It removes the friction of embedding secrets in every query, enables clean separation of duties through RBAC, and allows credentials to be managed centrally and reused safely across your entire account — from external tables to COPY statements and TVFs.
It's a small change to how you write SQL — but a big leap in how Firebolt handles external access.
And as support grows for new source types like Iceberg, Kafka, and cloud object stores, LOCATION will make those connections seamless, secure, and easy to use.
We're excited about where this takes us — and even more excited to see what you build with it.
# Vector Databases Won’t Replace SQL - Andy Pavlo (/blog/vector-databases-wont-replace-sql---andy-pavlo)
SQL's slow. SQL's stupid. We hear these claims every time a new shiny tool enters the market, only to realize five years later when the hype dies down that SQL is actually a good idea. In this super techie episode of the Data Engineering Show, Andy Pavlo, Associate Professor at Carnegie Mellon University, joins the bros to delve into database internals and optimization. Andy discusses leveraging ML for autonomous database optimization, using Postgres for practical applications, tuning production databases safely, and why SQL is here to stay.
Listen on [Spotify](https://open.spotify.com/episode/4atvh29aha4IP63r69TTik) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/vector-databases-wont-replace-sql-andy-pavlo/id1561927688?i=1000657766164)
Transcript:
Benjamin (03:22.99): Hi, everyone, and welcome back to another episode of the Data Engineering Show. We're super fortunate to have Andy Pavlo on today. Andy's professor at Carnegie Mellon University CMU in Pittsburgh in the US. Many of you might know him from either Auditune or his kind of super well-known lecture series on database internals on YouTube.
Welcome to the show, Andy. Great to have you.
Andy Pavlo (03:54.235) Hey Ben, thanks for having me.
Benjamin (03:56.782) Yeah, it's a pleasure. Last time we chatted, I was there kind of giving a guest lecture in your advanced database systems, or no, the like vaccination seminar series. So it's super nice to have you on today on the podcast chatting about all things databases. Yeah. What have you been up to?
Andy Pavlo (04:05.979) Seminar series. Yeah.
Andy Pavlo (04:17.178) Uh, well, it's the middle of the semester. So there's obviously I'm teaching. Um, I mean, there's like research Andy and then there's like auto tune Andy. Uh, and we can talk about both of those. Um, so research Andy, uh, we have, that's the thing you always get, you get both. Uh, and I will say, I try very hard to keep, uh, research Andy, like university Andy and auto tune Andy separated for legal reasons. Uh, which we, which we can talk about it. You want as well. Um,
Benjamin (04:29.614) who's in the room with us right now.
Yeah.
Andy Pavlo (04:47.29) And so research, Andy, we've been focused still on applying machine learning to optimize database systems. But in the prior efforts in previous years, we actually were building a system from scratch called NoisePage, where we were designing from the internals to be entirely run by machine learning components and autonomous operation. For a variety of reasons, we've scuttled the project and just focused now primarily on Postgres. It's good and bad.
present concept to this approach in terms of research. But we've been mostly focusing on training, sort of constructing an environment we call sort of like a gym. Like basically instead of building a simulator for the database, as often occurs in like robotics and other sort of autonomous operation, you know, research areas, we just use the database as a simulator itself. And the idea is like, you just instrument the system and collect all the telemetry that you then train machine learning models to.
and determine the best way to add indexes, what knobs to tune, you know, to have your table partitioning, everything you could possibly human could tune, we try to automate inside of Postgres. And so one of the cool things we do at the gym is we instrument Postgres to accelerate queries faster than they actually would really run, by cutting things off earlier sampling, because you don't actually need the real results if you're just training stuff in a machine learning environment, or machine learning models.
And then I have another student that's looking into doing more, more advanced models for automatically tuning databases. So instead of having like these bespoke models, like here's my index tuner model and here's my query tuning model. We think we can build a single model that can accomplish everything. And what's cool about it, it can exploit the similarities between certain actions. Like if you have an index on A and B, it's very similar to an index on ABC.
So rather than treating each one as a discrete action, you can learn the sort of latent relationship between all of them. So there's still a huge effort we're doing on applying machine learning to automatically attune databases. On the system side, we have two ongoing projects. One is looking at applying query compilation techniques or compilation methods and optimization methods to user-defined functions. So there was a whole line of research at a Microsoft, actually at ships and SQL server today, where they can take a, like a,
Andy Pavlo (07:11.096) a UDF written in T-SQL or PL SQL. And you convert that inline into either SQL or relational algebra and inject that into the call and query. So now the optimizer sees a bunch of queries or sort of a subset of a query or a subquery rather than an opaque box UDF. And so we've been looking at going further than Microsoft and trying to handle all possible UDFs and do a combination of identifying what part needs to be inline and what part needs to be compiled out.
And so that one actually we're doing a DuckDB, because it has a, as I said before, a German style query optimizer that's based on Hyper out of Munich, which Ben, I know you're very familiar with. So they have probably the best optimizers that's out of the source we could use for this project now. And then I have one last student that's looking at BPF optimizations for databases. So this would be the Extended Berkeley Packet Filter in Linux. You can write these... you can write these basic verifiable programs that you can then load into the kernel and run in a safe manner to avoid the overhead of copying data back and forth between the user space. So we did some initial work of using it for making a Postgres proxy run faster, like making PG Bounce run faster. But now the basic is that, okay, if you put the entire database in the kernel with BPF, what can you do? All right, so that's a lot there. Any questions about any of that? Again, that's all just the research side of things.
Benjamin (08:38.35) Nice, that's awesome. Let's wrap up kind of the full circle first, kind of tell us about industry, Andy, as well, and then we'll start digging into some questions.
Andy Pavlo (08:47.64) Absolutely, yeah. Yes. So, industry Andy is still with AutoTune. This is a project that started that we spun out of the University of Research lab where we were using machine learning to optimize databases. That one we were taking an opaque box approach where you assume you can't touch the internals of the database system, but you can't modify anything on the inside of the source code. How much can you optimize?
Eldad (08:48.112) opening up everything.
Andy Pavlo (09:16.856) just through external methods, external APIs. And.
Benjamin (09:19.662) what would be examples of things you would kind of change then in the kind of system configuration.
Andy Pavlo (09:25.24) So this would be in terms of like the internals of the source code or what we can tune.
Benjamin (09:30.734) what you can tune. So you can tune things like indexes and all of that stuff, right? It's just you're not touching Postgres source code, basically. Gotcha.
Andy Pavlo (09:37.912) Correct, yes, or we also support MySQL. But Auditorium originally focused just on knobs. So these are these configuration parameters that pretty much every data system exposes, like caching policies, buffer pool sizes, work memory, stuff like that. And so we can train machine learning models that can predict how the system's behavior will change as you start changing these knobs. And then you can sort of do a search to look for... a set of configurations for the given workload running on a given database that'll maximize, you know, minimize latency, maximize performance and so forth. And so basically what happened was while this was a research project, a bunch of people emailed us and said, we had the exact problem. We'll give you money to fly a student out and set it up for us. And this happens so many times you figured, okay, let's spin it off as a startup.
Benjamin (10:27.95) That's awesome. And especially on Postgres, it's like such a, like the differences can be so huge. Like I remember when we wrote that scheduling paper for SIGMOD and we benchmarked Postgres and like you take the out of box configuration, it's just horrible. And like, I mean, as a researcher, you should do a good job kind of measuring the other systems in the fair environment. So it was actually a lot of manual tuning there as well, just to kind of get Postgres to perform. And then it's an amazing system, of course.
Andy Pavlo (10:40.44) Yeah.
Andy Pavlo (10:52.92) I mean, notice they're immune to this. They all have this problem. Now, when you run on a hosted environment, like for example, Amazon has RDS, they have a host of version of Postgres, Aurora is sort of a similar thing. They've done some tuning for you. Like they know that you're running on this instance size that has this many cores and this amount of memory. And so they've set the parameters to some general purpose of lowest common denominator setting for that environment. But if you tune it for exactly how the... the applications using the database, like what queries are running, what the database, what the data looks like, then you get much better performance. And I would say Postgres is particularly tricky, not just for scheduling, but the auto-vacuum is always the thing that trips people up, and it's just an artifact of Postgres's architecture. So tuning that makes a big difference.
Benjamin (11:42.414) Cool. So kind of then to now close the loop and go back to research and you write like one thing I'm curious about is like this transition from building your own system, right? Like kind of Peloton and then noise page now to using some of the existing kind of open source databases and mainly doing research on those. And I feel like that's also becoming more common. Like you see so many papers nowadays implementing things into doc DB into post-press and so on. Like. Tell us about that transition, right?
Andy Pavlo (12:13.207) Why do we go down this path? Why do we abandon writing system? Yeah, I would say the reason why we abandoned the writing system from scratch was threefold. One is it was right when we forked off the startup and I was just sort of spreading myself too thin. Two is the pandemic. And the artifact of that one that caused the problem was not so much that we were remote and work from home. It was that... I took on way more students to work on the project than I should have. Basically that first year in 2020, any student that contacted me at CMU and said, I lost my internship. Can I just, can I do something, right? Can I work on something so that like, I don't have a gap in my, my resume. So we took a bunch of students who had maybe not taken our database class before, maybe weren't the strongest C++ programmer, but it was like my way of just trying to help.
And so we ballooned to like 35 people and it was not sustainable. The code really suffered. Correct, yes, of varying quality, yes. So, you know, so it was that and the quality of the code sort of tanked. And as part of this also too, like we didn't have actually a good query optimizer. So like, even though we could do all this great machine learning stuff, in the end of the day, if you're picking crap, crappy query plans.
Eldad (13:19.428) an army of interns.
Andy Pavlo (13:42.263) It's all for not, right? And then the last one is my wife and I had a baby who was colic and it was just screaming 24 seven. And I was like, I can't do this. So we looked around and said PostGus was readable at the target. And like I said, there's pros and cons to all of these. Like the auto vacuum is the probably the, I mean, for the sum of stuff we're doing, it doesn't cause too much problems just because we know it's something we need to tune and automate.
It just needs to be careful when you run experiments, make sure it doesn't like, you know, start making crazy changes and slow things down. In terms of DuckDB, we chose that one, as I said before, for the user defined function researcher because they could do, it's one of the few systems that implements the Thomas Norman paper on arbitrary unnesting of subqueries. Now they didn't support lateral joins and my students last year actually submitted a pull request to.
DuckDB to get that in there. So now we can unlast lateral joins. And that's the last piece you need to use it up functions. Like, so because hyper is not open source, umber is not open source, we didn't have anything we could use for this. So we chose DuckDB. And I say, I miss it. I miss building a new system. It was, how does, we were, we were, it's not that I was competing with the, the, the data researchers at Munich, but, you know, I was definitely heavily inspired by a lot of stuff they were doing and like, I know that they're working their deepest all the time and sort of kept me kept me excited and want to keep working on our system. But I think I guess at this point in my career, it wasn't sort of too much. Now, then where we want to go next is, of course, I want to go back and build a new system. But when you look around the landscape of what the Davis market looks like right now, I don't see anything immediately obvious. I'm like, oh, this is a very this is a niche we could we could we could go down. That would be set us apart from, the duck DBs, the click houses, the fire bolts, the snowflakes and so forth. So where I'm at right now is actually rather than building a new data system from scratch, we're actually gonna build a query optimizer first. And we have an ongoing project on that being sort of like something like to replace calcite. And then that way, yes, yes, yes. It's called Opti, it's very, very, very early. So I don't publicly really talk about it because like, you know, we're,
Benjamin (15:56.782) Nice. And that's built in, that's built in Rust, right? Kind of from, from what I know. Nice. That's awesome.
Andy Pavlo (16:08.215) we're doing basic things right now, like predicate pushdown and so forth. But the goal is, and eventually have a sort of clean code base for this, and then we can just sort of keep rolling with it and hopefully, you know, get far enough along where we can expose it to other people. And then, again, the same way that people have sort of adopted Calcite, we hope that this will help people out as well. So that's sort of, that's where my current, I was gonna say also too, query optimization is the thing I know the least about in databases. And so to me, I'm naturally, gravitation towards that because it's hard. And so we'll see how that goes.
Benjamin (16:39.374) But it's like a kind of top-down cascade scale optimizer, right? From what I saw. Okay, it's funny because like we here at Firebolt, we also do like bottom-up kind of DT and so on. So cascades, like it never clicked with me. Like up to this day, I've tried multiple times in my life to really understand it. And I've always, how the heck do you build this without just running in circles or whatever? Like it's always...
Andy Pavlo (16:44.567) Yes.
Andy Pavlo (17:00.727) It's. I'm the exact opposite. To me, when I read the hypergraph paper, again, from Thomas Neumann, again, this is gonna sound like we're just talking about Thomas nonstop, but he's prolific, writes a lot of papers. When I read those papers, I'm like, I guess, right? But to me, the Cascades one, because it's all this unified model and he threw everything in, that makes more sense. Now, I will say the inventor of Cascades made a passing comment at Sigmod in... 2017, 2018, when he won sort of the test time awards. And he sort of mentioned at the end, he's like, yeah, if I had to do it all over again, I would use Cascades to do all the initial planning, but then run the bottoms up DP to do joint ordering. So we might eventually get to that step, like have the extra thing at the end to do the DP, but right now it's just purely a Cascades.
Benjamin (17:55.534) Nice. That's awesome. Anyways, I'm super excited to see how that project will go. Like for us as well, when we rebuilt our query optimizer kind of from scratch, it was like, right, like kind of looking, hey, kind of what's there? What can we draw inspiration from? And on the C++ side, for example, like there's nothing like Calcite, right? And we wanted to stay all C++. So kind of going with Calcite wasn't an option. So when we rebuilt everything kind of, we actually like took many of the kind of design choices that Calcite did in terms of modeling the algebra and so on. But we had to build it all from scratch kind of in C++. So yeah, that's going to be a kind of awesome project and very nice for the community as well.
Andy Pavlo (18:26.967) Mm-hmm.
Andy Pavlo (18:32.279) There's two other sort of query optimizers that are floating out there now. There's from the data fusion guys, they have something. But when we looked at it, I think it's very much rule-based and heuristic based. I don't think it's cost-based. My understanding is a student in China sent them a pull request to add Cascades and I think they've rejected it. And then the Velox team at Emetta, they have an experimental branch from Velox for a query optimizer called Verax.
I don't know again, how robust it is. We haven't looked at it yet, but those are two other projects I think that might float around that might go somewhere.
Benjamin (19:10.254) Thanks for the hints. That's awesome. Cool. So transitioning a bit to kind of in -
Eldad (19:15.748) Since we're all throwing around optimizers and projects around, so let me take you back a few years back to the Huffler-Platner Institute in Germany. So Firebolt actually started by taking that weird, amazing little project from there. Obviously the team rewrote everything, but yeah, your discussion kind of reminded me on picking the right component at the right time. But yeah, lots has evolved.
Andy Pavlo (19:18.583) Yes.
Andy Pavlo (19:43.063) I mean, you're referring to high rise, right? Yeah.
Benjamin (19:45.486) Yes, exactly. So the optimizer originally in Firebolt was part of that. And we connected it to ClickHouse Runtime, but yeah, not anymore.
Eldad (19:46.66) Yes.
Andy Pavlo (19:55.287) We looked at high rise originally, I think a few years ago, and I think what spooked us was the license, the source code license. We're not lawyers, but like, because it wasn't Apache, it was like this very, it was Hasno Plot, it was something very bizarre. We were like, we're not touching this.
Benjamin (20:12.686) But they also have this high rise too, like they actually have two versions, right? Like at some point they rewrote it completely. And I think kind of the newer one is under kind of standard open source license. But anyways, I think we got very deep now into like, uh, query.
Eldad (20:16.164) Yes. Yes.
Eldad (20:26.692) They did it after most of the team went to work for Snowflake and they kind of had the generation shift and one of the decisions was yes, let's open the license.
Benjamin (20:33.87) Yes. And now all of the high-rise two people are also at Snowflake and they're all great. So, nice. In terms of industry, Andy, right? Like it seems like actually many of the things you're doing in academia now, I guess, are inspired by the like real industry problems you've then seen at AutoTune and so on. Like how's the kind of connection between those two? Right? In the beginning, you said, okay, you're trying to keep them separate.
Eldad (20:40.196) Hehehehehe
Benjamin (21:02.286) from your story now, it seems hard to keep them separate because they are working on similar things.
Andy Pavlo (21:08.311) Yeah.
Eldad (21:09.732) That's anti-lawyer, he was not invited to the podcast.
Andy Pavlo (21:14.391) I had to sign this contract, this license contract with the university that basically says like the autotune research project at CMU is dead and everything is over on the, everything's now in the startup. So I would say that like the problems we're solving at the startup are certainly, I mean, they're related, but there's things you have to deal with real world production databases that honestly in the university and most of academia don't touch.
So, and I also give an example of something we saw in the startup that we said, Hey, this is an interesting problem. We don't have the resources to solve it. And then that went, you know, I brought along back to the university. Um, so with autotune, the original idea was that you, the customer or the user would clone their database, make a snapshot of it, collect a workload trace, and then have a spare machine or backup. They would then run experiments on tune that figure out the best configuration, then apply it to the production database. And when we did initial pilots with the university project, one was like at a big bank and we talked to other people at like the US Patent Office. Like these had full time DBAs that had the know how and the time to actually just do this setup. And when you read all the papers on doing ML for databases, they're all pretty much making the same assumption as well. Like you're going to run on the spare machine because you don't want to interfere with the production database. When we then spun it out as a startup, you know, sure, some of the initial people we were first talking to could do this. Uh, but the majority of people cannot, cannot like, you know, Amazon does make it easy to make, to take a snapshot and clone the database. But the workload capture and replay tools for the open source database is like Postgres and MySQL pale a comparison to the commercial ones like, like Oracle's tools, for example. So even if you could capture the workload trace, you're not really going to re simulate the production environment on the clone. And then furthermore, some people are running on like really expensive machines on the cloud. Um, you know, like, you know, $40 ,000 a month, uh, you know, instance with backups and replicas and provision, I have all that. Like no one's going to spend about $40 ,000 a month, uh, instance just to run an experiments on. So a lot of what we did in auto-tune since spinning it out has been to make the machine learning models more safe and a bit more conservative because we know people are pointing at us at production databases. Like in the very beginning, when we first came out of stealth,
Andy Pavlo (23:41.623) everybody we talked to, we'd get on a call and like, okay, like just so you know, you don't want to run this in production, right? Cause it's machine learning has to learn. You want to do this on a replica and everybody that would listen to us like, oh yeah, okay. I do want this, but I can't, I can't do that. I'll just jack up my instance size, pay Amazon more. Cause that that'll solve my current performance problems. For some reason, everyone heated our warnings except for Brazilians. Uh, like in the very beginning, Brazilians for some reason, like we don't care. We'll point it at production database and we're like a brand new startup. Like I wouldn't put, I mean, the very beginning, I would not put it on alternate out of production database. And now I think it's safe, but in the very beginning it was like, but for some reason, Brazilians is like driving without a seatbelt. They were okay with it. Um, so since then we've, we've set the service up to, to get, like I said, be more conservative, be mindful that we're running production databases, avoid, you know, wild swings and performance do things incrementally. Uh, we need to see 24, our observation window is 24 hours, so we sort of see the day night pattern. We automatically skip weekends and holidays, right? So the book, like all those are the things we need to do to make this people comfortable with using a machine learning based tool for the databases. There's other stupid thing, yes. Yes.
Eldad (24:50.404) So Andy, quick question. Let's fast forward to data warehouses having full workload isolation, metadata decoupled, every compute spin up, every cluster spin up, basically looking and being able to redo the same exact same workload up to the query history that remembers everything you've run. So kind of retrying the same idea behind the project on that architecture, which has almost none of the limitations that you've mentioned from the past where you're like, sure nothing. Yes, you run one system in production. Now you say, okay, so how do I reproduce? It's kind of like you do auto-scale, like you do a version change while running a workload on a snowflake, right? That project might actually be much more acceptable and safe because you cannot hurt the system. The warehouse, right? The cluster, the warehouse that runs, it just runs. And then the warehouse that learns just emulates, but it's the exact same thing. So you spin up another warehouse on the side, it's isolated. It kind of solves a lot of that.
Andy Pavlo (25:51.383) Yes.
Andy Pavlo (26:03.991) But who's paying for that? Who's paying for that compute? That's the challenge, right? If you have, if you.
Eldad (26:08.068) So you as a user, you do, yes.
Andy Pavlo (26:12.279) Right, so some people don't wanna do that potentially, right? So to your point, yes, it does make things easier. Microsoft basically does something very similar for their auto indexer in SQL Azure. They call them B instances, they'll spin up a replica, do a snapshot, split the traffic so it gets a mirror of the traffic as well. But again, Microsoft is doing that underneath the covers, the user's not aware. And again, Microsoft is just eating that cost. Going back to University of Andy research, one of the things we may want to try to be able to do is to transfer learning, be able to take like a T3 micro, the smallest you can get, train some models to learn from how the workload is going to respond on a scaled down version of it, and then apply those same changes to the larger production instance. Because then like T3 micros cost nothing. But right now, none of the work I think is, at least in our own work, is getting into this space. A lot of the research out there is like, has to be exact same hardware every single time. But to your point, yes, there are aspects of shared architecture that would make this a lot easier. But in the Postgres MySQL world, that doesn't exist.
Eldad (27:24.9) Exactly. That's a whole different problem. In fact, this is the problem. So kind of building a solution that can tune, change knobs while flying is a post-growth problem. It's kind of a, you know.
Andy Pavlo (27:27.223) Yes.
Andy Pavlo (27:38.615) Yes. And that's, that's what we were trying to do with the noise page project, like to build the system. So that if you change the comfortable size, you don't have to restart. Like you're doing, you know, in me, that was the whole goal. Uh, we just, at the end, like I said, we.
Eldad (27:49.82) But then there's a whole range of problems that you've solved that go beyond solving the architecture challenge. So once, yes, we can afford running for 30 minutes, assuming it takes 30 minutes to get to an optimized result recommendation. So I spin up a snowflake warehouse, it runs for 30 minutes. It's exact same warehouse definition. I have no clue about the hardware. I just kind of... I assume that the warehouse at Snowflake of type S will be the same if I spin it two, three, 10 times. I run it. Yes, I completely ignore all the infra architecture kind of post-gress challenges, but then it goes straight to learning and really improving pure database knobs on that new architecture, which could be amazing. Even as a project for students that can experiment their research on modern products, like, as you say, like Microsoft has that since last year, complete work with isolation, Snowflake has that, BigQuery obviously, yeah.
Andy Pavlo (28:51.055) Yes.
Benjamin (28:51.182) I mean, you even have products around stuff like like Kibo now kind of do these kind of cost optimization things for these, like the couple storage and compute systems.
Andy Pavlo (29:02.575) Yeah, so Kibo is a startup founded by a Danish friend, Barzama Safari, a professor at University of Michigan and his current student who's now a professor at University of Illinois. So they're doing two things. One, Snowflake doesn't expose that many knobs. So I think it's like three knobs they can actually tune. So one is the search space is much less than what we have to deal with at Postgres and MySQL for autotune.
The other thing they do is also sampling. So they'll rewrite your query to hit a sample table rather than the full table. And they can provide some statistical guarantees about the error rate. So like they're doing a bit more than autotune in terms of like doing query verifying because they're sitting in front of the snowflake. And we currently can't do that autotune because it requires additional infrastructure. But I was to say, there's other stupid things we had to do that again, it's not research, it's not deep science to make this work. So one example was, a bunch of customers started complaining saying that the values of the recommendations looked like things that a human wouldn't generate. Like there was too many decimals. So we rounded them up, right?
Benjamin (30:09.39) Nice.
Eldad (30:11.351) That's it exists. We have it. Benjamin, you see this abuse of telemetry, really the decimal numbers. Thank you Andy for everyone out there. Please don't abuse decimal numbers.
Andy Pavlo (30:26.157) But like, so we added that, right? Like I said, that, and those are the things, because we know we're running production databases, we just gotta be more cautious for. Yeah, so I forgot what the original question was, but like the, it's, you know, actually the challenges of putting stuff in production and then what we've learned about, you know, from maybe Autotune or the research and how we sort of cross-pollinate them. So let me give an example of something that we saw a lot in Autotune that we then brought back to the research side. We found a lot of people that were using proxies in front of Postgres. And it's mostly PG Bouncer. If you're running on Amazon RDS, then they have their own proprietary version called RDS proxy. And from a research side of things, they're very, very, I couldn't find any papers about the modern incarnations of proxies in front of database systems. And so we... when we were looking at sort of this BPF research, we said, oh, this is clearly something we could accelerate with BPF, the proxy, because all it's really doing is packet shows up, read the header, figure out where it needs to go, and then shove it right back out in the NIC. So instead of copying up from the NIC through the kernel up into user space, you just intercept the packet as it comes into the kernel, figure out where it needs to go, and then shove it right back out. And you can scale this thing much higher than PG Bowser can do. Because the mem copy is what slows you down. So that's a good example of like, this is something like, hey, it'd be nice if we could solve this problem, but we just can't do it at the startup. So we brought it back to the university and it worked.
Benjamin (32:05.838) Nice. That's awesome. Super cool. Sweet. So one other thing in kind of transitioning a bit away from this, right? So I follow you on Twitter, of course, a lot of people follow you on Twitter. One thing that keeps coming up is your, your take on SQL as a whole, right? So like zooming out a bit specifically from these tuning aspects, like SQL databases. And I think you also have a paper kind of coming out with Mike Stonebreaker. soon what goes around comes around and around where you kind of say, hey, like, this has been around for a really long time, kind of nothing's replacing SQL anytime soon. But AI is going to eat it. Like, there's no more SQL in five years, really. I think you're totally wrong. No.
Andy Pavlo (32:50.539) Sure, yes.
Eldad (32:56.02) It will end up as a Netflix Doctor.
Andy Pavlo (32:56.139) Don't you work at a SQL company?
Benjamin (32:58.094) Yeah.
Andy Pavlo (33:00.779) You work at a SQL database company, what are you talking about?
Benjamin (33:03.438) No, of course, it wasn't, it wasn't serious. Um, but so there are changes, right? And like, okay, you have vector-based databases becoming super popular now, all of those things. My personal take is that like all of this will just be in the longer term kind of soaked up by SQL system. Like there's no good reason why this wouldn't be okay. A vector database type kind of in your relational SQL database, you have some aggregate functions, you have some scalar functions, whatever.
Andy Pavlo (33:05.355) Yes.
Benjamin (33:33.326) So. Yeah, like give us a, your take on kind of SQL as a language kind of surviving for so long and still being strong. And then maybe especially now in the context of all of the generative AI stuff going on.
Andy Pavlo (33:46.827) Yeah, so let me give some background about this particular paper you were referencing. So the first edition of this paper was called What Goes Around Comes Around, and that was written by Mike Sturmbricker with also Joe Hellerstein out of Berkeley in 2005. And it's basically Mike's recounting of the history of data models since the very beginning of the 1960s and up until 2005. And he sort of goes through like, you know, in the beginning, it was the network model, the hierarchical model, the codosil, and then the relational model comes around, and then there's all the sentience and variations of them. I think he gets up to XML databases, which became in vogue in late 1990s, early 2000s. And the main takeaway from that paper is basically how the relational model is the superior model, and it can account for, actually, very specific, the object relational model, which... Stonemaker coined as part of the Postgres project, that's what he considers to be the superior model. But object-relational, he just really means extensible relational model, whether it's XML or JSON or arrays and so forth. And so again, this paper lays out, here's all the things that people tried and none of them has overcome SQL. And they're all basically making the same arguments every time something new comes out.
Oh, SQL slow, SQL stupid, relational model is slow and stupid. Here's this new thing that's shiny that's so much better. Oh, turns out, you know, five years later after all the hype dies down, oh, SQL and relational model is actually a good idea. So we've, we've. It's, it.
Eldad (35:25.869) Nobody read the paper. That's why we ended up with mongos and with no sequels and new sequels and all sorts of shit for 15 years, by the way, 15 years. Like, like nobody wanted to listen to people who know SQL. That was the thing, right? Like big data, go away, SQL guys. And it took 15 years to take it back. Really.
Andy Pavlo (35:32.587) Yeah. Yes.
Andy Pavlo (35:40.907) Yes.
Andy Pavlo (35:47.787) So it's what goes around comes around. So that was the impetus of writing this paper, the follow-up, because actually there was a hacker news, there was a hacker news comment that I saw where someone writes, I don't know why people want to use relational database. Everything should just be a graph database. And what they wrote, yeah, but it literally was almost like a verbatim quote of what was in the original 2005 paper. So I emailed Mike and I was like, we gotta write the follow-up, because it's, you know,
Benjamin (36:05.294) Is it web scale? Is it web scale?
Andy Pavlo (36:17.643) 15, 20 years have passed and it's what goes around comes down all over again. As you said, the new single guys basically are making the same arguments that the XML database people made, that the object database people made. And so the paper is coming out I think later this year, depending on when this podcast is put out, the paper should be available. But basically we go through the last 20 year history of other alternative data models, like the document model for the JSON stuff, key value stores, although they've been around since the 1990s, late 80s, but now there's key value systems. Graph databases, obviously, array databases, vector databases, of course. And the MapReduce really isn't a data model, but as I said, the big data movement was also around that architecture. And we basically just say, look, here's the same thing as people tried before. They make the same arguments that SQL slow, SQL stupid. You don't wanna use that. And then it turns out, oh, there's actually value to using the relational model and a declarative language like SQL. And although, and then the proponents of these other alternative data models, their systems end up morphing into supporting SQL. And actually, I would say that part is unique because the object relational database guys and the semantic database people in the eighties, they never adopted SQL. They just died and withered on the vine. The new SQL systems, Lee saw the writing on the wall and said, oh yeah, SQL's inflation multiply actually good idea. And they've all pretty much morphed into being something that looks like a SQL database. And so now, again, as they said, the hot thing is vector databases. You see some wild claims basically saying the same thing, like how vector databases are gonna destroy, you know, relational databases, SQL stupid, SQL slow, yada, yada, yada, right? And what will happen is of course, these systems will either, you know, will morph over time the vector databases.
And add something that looks like relational model and add something that looks like SQL with obviously these vector indexes to accelerate things. And then likewise, SQL will evolve over time and add support for vector primitives or vector built-ins and vector functions. And you're already seeing that now pretty much every database that's out there, I don't know if Firebolt has anything yet, but like every database has some kind of vector index.
Benjamin (38:34.862) So in 15 years, when we do kind of season 27 of the data engineering show, will we then have the same conversation of the past 15 years? We're just about vector databases. Like, do you think it's another 15 year cycle or do you think it's becoming shorter?
Andy Pavlo (38:51.21) I think, that's a good question. I have the cycles probably five, 10 years for like when something new comes out, a lot of excitement. And then, uh, and then like, again, SQL expands and evolves. Like I would say what's different this time is, um, like with no SQL and the Jason databases, so in particular couch, GB, Mongo, and so forth. Uh, it took a while before the relational databases, the, the, the incumbent systems add a support for Jason. Oh, which is surprising because they add, they add a support for XML and Jason isn't that, but that much farther from it. They could, they could add it pretty, pretty easily. And then the SQL standard added support for JSON, I think 2016. I think the most recent version they had for their data types in 2023. But it took a while for that to happen. What surprised me with the vector index stuff is you look at chat TBT, chat TBT blew up November 2022, right? Like December 2022. I mean, it's obviously been around before then, but that's when everyone was talking about it. It was on the news. Look out, you know, like. It was in the zeitgeist of humanity, right? And not just like a tech bro, tech only thing. So, but within a year or less than a year of chat CPT becoming super hot, again, all these common databases add a support for vector indexes. And that's, I think it's a combination of two things. That's one, it's either because the hype was so much, like how could you ignore this? And they were just sort of trying to ride the wave.
And also too, there was enough open source tools or vector indexes out there that could do approximate nearest neighbor search that were good enough that people could just adopt them and download them and integrate them. Like disk ANN from Microsoft, FICE or FAST from Meta Facebook. Like most of these people didn't roll their own vector index. Now you have the change, like the query operator is like, there's some changes up above, but it's, from my perspective, it's just another index. It's not a, it's not.
Andy Pavlo (40:51.562) It's not, does not necessitate writing a completely brand new database system architecture, right? It's basically all of these, these vector databases, they're going to do vector search faster than the relational guys, but is that, is it fast enough to, and better enough than to give them a moat to protect them from other people encroaching.
Eldad (41:12.958) There's so much value in kind of just talking to users the way they picture it is, look, we are using Chatch UDP as a part of our data, new data pipeline. So user goes, sends a query or UI sends a query. We get results. We need to send it to Chatch UDP. We need to get the result back. Can you do UDF? Can you do a lateral join? And for every value, can you cache the result of Chatch UDP? Like those are the kind of questions that people ask and for good reasons, it's just like another data pipeline project for them. And it should end up as a type index and an extension to a lot of good stuff that already exists within database systems. Maybe not necessarily just SQL, but I think it makes sense for users to view that as yet another type or index. Absolutely. To what you're saying. So.
Andy Pavlo (42:10.058) But I would say the vector database is what they do better than most of what I've seen out there for other relational databases is they have better integration with the AI ML tooling, like Lang chain or TachyBT, Lama index, like all that you can call directly within the data system itself rather than having to bring it back to the application of Python code. I mean, you could run the Python UDF, but as far as I know, I haven't seen any, any vendor providing like, hey, here's our language integration, UDF package, or things like that. Then you also then go back to the problem I was saying before. Now, if you're calling UDF, depending on where it is in your query plan, that's gonna cause problems with the query optimizer, because it's gonna see us as an opaque box. So there's no free lunch. It poses new challenges, but at the end of the day at Ohio, I don't think it's significantly different.
Eldad (43:03.773) interesting to see how the planner evolves.
Benjamin (43:04.814) head nodding all around.
Andy Pavlo (43:08.554) Yeah, sorry. Yeah, the planners at intranet one too also as well, because the, you know, you have the pre-filter, post-filter problem where, you know, your where clause, assuming again, the SQL, you're doing some lookup, you know, to find the tens, the 10 highest rankings for some embedding. But then you also want to like filter like, for the additional metadata, like attributes and, you know, like, find me all the people that promote someone to this person above this age or something like that, right? So when do you actually... When is it better to get the rankings then filter? Should you filter first? Then how do you incorporate that into the search index? The Pinecone has some claims how they can sort of integrate all that once. Same with Weav8 has something as well. I haven't looked deeply enough to say what they're actually doing, whether they're maintaining, if it's using HNSW, whether that metadata is embedded in that graph as well. That part I don't know entirely. But as you said, like. Now you have this index that you can't maybe get back selectivity because it's not like an exact lookup. It's like, hey, give me 10 things that are close enough. And so how do you cost that when you do that versus another predicate is a very interesting problem. I think it's still open.
Eldad (44:23.511) Exactly. And those kind of problems are, this is why at least I believe that it cannot just stay within a vector database or kind of a vector database needs to grow itself or we'll see some interesting things happening in databases in the future.
Andy Pavlo (44:38.41) But there's two paths that could go down, right? And I did the WeV8 podcast and I basically told the CTO, you guys are gonna have to support SQL in like five years. And so that's what they go down the path of evolving, sort of similar to how MongoDB has evolved to their hosted version, Atlas, support SQL. But going from like, oh, JSON only, JSON API, no SQL. Now they've had SQL and Xpand because people want to start hooking up MongoDB to more stuff. So they could either go down the path of becoming more relational like the other NoSQL systems so they can fit into the rest of the data ecosystem that's out there. Or they go down the path of still maintaining their proprietary query language and their API and then just going down the path of something like Elastic where you don't get SQL, right? Because it's not meant to be the primary database of record, it's the separate thing that you can do your analytics on or searches on. So I think that's the two paths that have to go down. And they go down the first one, they'll have to support SQL.
Eldad (45:49.142) And the market will define it for them, right? Like you think Elastic, they had great timing. They had like those few years where nobody competed with them. Nobody was interested into what they're doing. And obviously anyone that was building a database was somewhere else. So that opened up new users. And then you're saying SQL is back, at least from what we're seeing, it's like, yes, it's because the people who own the budget now want to write SQL. And a few years back, that was a... a cloud ops engineer using node .js and that was the great thing to do to become a human optimizer and just write down your planner. That was it, that was the job. Because their SQL database couldn't handle it. So that opened up right like this dynamic in market, like write the stars. Having that opportunity open for elastic that crazy run, which is now obviously.
Andy Pavlo (46:25.409) Yes.
Eldad (46:46.196) much more complicated as everyone is trying to turn any vertical into a data type. So whether it's a search, whether it's just another type. So yeah, it's going to be very interesting and competitive in the next few years to redefine that, even the SQL, right? Like the barrier of entry going forward to build sophisticated solution on SQL are dropping with AI.
Andy Pavlo (46:56.385) Yeah.
Andy Pavlo (47:13.249) Yes.
Eldad (47:14.386) There's no need to reinvent the universe. Just plain dead simple. Generate boilerplate to kind of parse JSON, right? So it's an ugly task that someone would spend a few days on and there was forums and people spending and now it's done, over. It's just, so yeah, nice.
Andy Pavlo (47:35.169) So the Weaviate guys make an argument, which I, if I can get for their perspective, it makes sense, is like, they don't see themselves potentially having to support SQL because you can just put a transformer in front of that that rewrites SQL into whatever API that they have. So they don't really need a parser or anything like that. We'll see how that plays out. But that was their perspective, like why, like, SQL to them is just another interface that could be natural language, could be SQL, could be... whatever you want. But we'll see.
Benjamin (48:07.502) LLMs all the way down. I think those are some great closing thoughts, Andy. Thank you so much for being on the show. It was a pleasure having you. Good luck. All the best both to industry Andy and academia Andy. So we're rooting for you. Yeah. Thanks for being on.
Andy Pavlo (48:09.345) Yes.
Andy Pavlo (48:22.656) Yes.
Andy Pavlo (48:26.239) Hey, thanks for having us, this was fun.
Benjamin (48:28.078) Awesome.
# Vin Vashishta explains why we should stop using dashboards (/blog/vin-vashishta-explains-why-we-should-stop-using-dashboards)
Vin Vashishta, the guy we all love to follow, has never seen a dashboard with positive ROI. This time on The Data Engineering Show, he met the bros to talk about the difference between BI dashboards and analytics that actually introduce knowledge. It's no longer just about the data volume, it's about quality and relevance.
Listen on [Spotify](https://open.spotify.com/episode/1DVL3pTZDgEJ3H1TeSxjP7) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/vin-vashishta-explains-why-we-should-stop-using-dashboards/id1561927688?i=1000630165552)
**Benjamin:** Hi everyone and welcome back for another episode of the Data Engineering Show. So today we have Vin Vasishta joining us, which is awesome. Good to have you on the podcast, Vin.
**Eldad:** Hey.
**Vin Vashishta:** Thanks for having me. I appreciate it.
**Benjamin:** Awesome. So for everyone who hasn't heard of Vin and Vin is really well known in the kind of data thought leader space, he's an AI advisor, kind of co-founded V Squared, which is a data and AI consultancy and wrote a really well known book called From Data to Profit. So great to have you on today. We're super excited to chat and joining me of course is the data bro, Eldad. Good to have you back after being away for the last two episodes. So welcome back on the show.
**Eldad:** Thank you for giving me another shot on the show.
**Benjamin:** Yeah, we discussed really like a long time before what we said. Okay, we'll, we'll have you back. Awesome. Uh, cool. Vin, do you just want to quickly introduce yourself, uh, kind of, uh, yeah. Uh, and tell us what you're up to these days.
**Eldad:** Yes.
**Vin Vashishta:** Yeah, definitely. So you covered the high points. Vin Vashishta, been in the technology space for almost 30 years in data science and machine learning for the last 11 and a half years. Went to school to do data science and graduated in the 90s when no one wanted it. So I had to go into software development, software engineering roles, built and led teams, got really close to implementation and execution. We had to deliver more than most data science teams do.
because the technology was more mature. So there's a closer connection in my background to implementation, to execution. I brought that with me in starting vSquared. Really quickly realized that being successful with digital and with software and with cloud and all of the other technology trends, very different than being successful with data and AI. So I built my consulting practice around not only implementing solutions, but all of the other components that are necessary.
Really the book is the culmination of 11 years journey through how do you make this, this technology actually result in some cashflow for people. So that's, uh, that's my background. What am I up to today? A little bit too much. I'm not going to go into everything. I'm doing too many things. I'm going to put it that way. I need to take a vacation.
**Benjamin:** Nice. So you said that you.
**Eldad:** Data is hard, data is not easy. It's not like the 90s.
**Vin Vashishta:** Not everyone can do it. No, I think data was harder in the nineties because we didn't have any of the infrastructure and none of the libraries. I mean, we were trying to build models and see it was, it was painful. Yes.
**Eldad:** It was all about being creative in the nineties, like the movies in the nineties, like, you didn't have a budget, you had to be creative and then, you know, make assumptions about the data and every now it's all about facts and it's boring. And, but yeah, I mean, that's good.
**Vin Vashishta:** Ha ha ha ha ha. Yeah, no more shooting from the hip anymore. You have to actually be able to back up what you can say you can deliver.
**Benjamin:** So one thing I'm curious about is you said you like over the time at V squared and also before that realized that running successful data projects is kind of very different from running successful software projects from a more traditional software background. Like tell us more about that, right? How is it different? Like what are, what are your lessons there?
**Vin Vashishta:** Well, I think the biggest difference is the way that they're monetized. So the companies that I worked with in the early 2010s, they were solving the technology problem. And the first couple of years that I'd started vsquared, I got swept up into the same thing. I was helping them solve technology problems. We would deliver really successful initiatives, but trying to maintain momentum, trying to get that, you know, the hype, the, the tens of millions, hundreds of millions, billions.
It was more complex because we needed bigger budgets and we had to then justify. We had to support it better. And what you realize very quickly is you have last mile problems, which are the technology problems, but you also have first mile problems, which connect to the business strategy data is unique. It has unique monetization properties. We have to gather it in different ways. We've been gathering it, really thinking about business intelligence use cases.
And to use it for analytics, for machine learning use cases, completely different data gathering, data generation, data engineering. So all of those are really enterprise wide problems. When I broke it down into pieces, realized you have to fix problems that start at the strategy level, move into culture and then go into technology. And data teams inherit. I don't know.
10 to 20 years of technical debt from every other organization and every other part of the business. And we're expected to solve all of these problems. Some of them are technical and we can solve them, but other ones are cultural and strategic. So we need new roles. That's what I began filling myself and now I'm teaching people how to fill those roles and I teach companies how to build teams in a way that really connects all of the dots, not just the technical dots.
So very different monetization, very different development. If you try doing it the digital way, it's a failure. It just won't work.
**Benjamin:** And so just for me to get this, when you're talking about monetizing data products in this case, is this always something customer facing, right? Say a new ML recommendation system for products in your online shop or something, or is it also for internal use cases?
**Vin Vashishta:** I think we need to treat all data and AI products like they're customer facing, even if they are internal facing, but it's both.
**Benjamin:** So I get it for the shopping card recommender. You can say, hey, we made this much money and kind of feel like this user group, I do an A-B test, I see, hey, 10% profit, easy, kind of, right? You can show the numbers, say, hey, we did amazing here. How is that for internal data projects? Because there it seems much harder, right? I'm kind of working on some dashboards for someone else to use. Like how can you approach monetization? They are just kind of putting some...
Business value onto that.
**Vin Vashishta:** When you look at dashboards, it's a good case for self-service data tools, because I've never seen a dashboard that had positive ROI. Your data team, the cost, yep, whoops, whoops.
**Eldad:** What? Wait, we need to freeze, we need to pause for a moment, yeah.
**Benjamin:** What? Eldar, Sisense was useless.
**Eldad:** We never cut anything out of our data pro blog post, but you're so right. You're so right. Dashboards are just destined to die eventually. And then, but they stick, they stick around. I've seen dashboards that were running for years, years and exactly for the reasons you've mentioned, culture, alignment, you know, having three Salesforce accounts with having three teams updating those three different accounts. And then you need a... to make sense of a data lake. And it's hard, it's hard. And you're right, it's about culture and we all wanna be data-driven, but it's not enough anymore to be just data-driven. Yeah.
**Vin Vashishta:** Yeah, I think the important part about the dashboard, especially that component of it, we should give users the power to do some of these things themselves. And we underestimate users, consistently underestimate them. I have trained users who have never done any sort of engineering before. I've trained people in marketing teams, especially, to build out dashboards using Python, Jupyter
connecting to SQL, it is so simple now. And we underestimate what users can do on their own. Low code, no code tools are good, but we can actually, when we talk about data literacy and AI literacy, we can train frontline users to do so much more. And we can then support them in a way that's feasible. Behind dashboards, that's where we can deliver some ROI. It is...
Sometimes those models powering the dashboard that you never see, never think about, never hear, those are where the true insights come from. When you look at what it is that data scientists, data engineers, data analysts should be doing, it's the more complex. That's what we take on. If you want someone who simply does a select statement and displays the output, there's no way to make that a positive ROI initiative.
And so data, you know, frontline users should have the power to be able to do that for themselves when it comes to. Oh yes. Yes. Let them have it.
**Eldad:** But they're so addicted to dashboards. They're so addicted. Because we taught them, everyone taught them consensus on data is more important than anything else. So everyone is waiting for a dashboard because of that consensus that needs to happen, right? The single version of the truth, like all of those philosophies. And it's pretty depressing if you're on the go-to-market team and you need to support your customers.
**Vin Vashishta:** Mm-hmm.
**Eldad:** And you need the data to do it and you get a dashboard and you ask, like, can I drill into the numbers? Like I see that average, right? Like nobody wins on averages and dashboards are all about showing average numbers. So people want to drill in, people are looking for a way to express themselves with data. And that's a big question because there are so many religions out there. You've mentioned Python, like the languages, right? How do we approach data? Should I like, and that's changing. And technology gurus tell us like, oh, we just had Scala thrown out the window, like just, right? Like it feels like 20 years, it happens a few. So it's SQL, the way to do it is, is there, is AI going to help us kind of figure that out, how we move to natural language? Cause language is barrier, right? Like interfaces is barrier for users and cost obviously, but cost is being tackled by technology, so we should assume cost will be lower. But.
**Vin Vashishta:** No. Ha ha.
**Benjamin:** Thank you.
**Eldad:** interfaces don't really change. And that's the dashboard. We were stuck for 20 years with the same pie chart. And then I love pie charts.
**Vin Vashishta:** Well, when you think about dashboards, that's the BI mentality. That's the digital data mentality. You gather data to display it. And so it's gathered close to the people who will need it the most. And it never escapes that silo. So dashboards aren't data science. Dashboards aren't analytics. That's BI. And we have to stop saying dashboards are data science. Dashboards are analytics. They're not. They really aren't. I mean, an average.
**Eldad:** What makes BI from data from your experience, if you look at the data in the metadata only? You don't even know who the users are. What has changed? It's the same builders, the same builders that build the BI stacks before. Now they're moving forward to do something else, something bigger. Taking over engineering, owning the business is owning the data.
**Vin Vashishta:** Mm-hmm.
**Eldad:** So kind of from your experience talking to so many companies over the last few years, what's changing there in terms of data politics and like...
**Vin Vashishta:** Data politics is a thing now, but when you look at, so the challenge here is, and I don't wanna make anyone else mad, I've already probably made about 90% of your audience mad, and I think I'm gonna go get the other 10% right now. The people who created this problem are now trying to sell you the solution to it. The people who created all of these digital data silos with BI tools are now realizing that analytics and data science, yeah, hoo.
**Eldad:** Who's that? Who did that?
**Vin Vashishta:** No, not micros. I mean, no, but looking at, you know, all of these companies that created a digital data infrastructure. Where if I'm sales, I gather sales data and I keep sales data in the sales database connected to a sales app or an ecosystem of sales apps, because why? Why the heck would HR need any of that? Why would anyone else need it? It's sales data. It's us. And there was very limited cross pollination. That's BI. That's digital.
When you come into data for analytics, for machine learning, now we're looking at longer chains. Now we are looking at, instead of gathering data, we are gathering domain knowledge and building out a domain graph or building out a knowledge graph is the best way to manage for our use cases because that's the value. Why go from BI and dashboards to...
analytics and machine learning models. Well, what's the ROI? And that's really the, what's the difference? Why should I spend more for an analyst than for somebody who's doing BI? Why should I buy this new application? And I think it's important for companies to ask those questions because those are the only ways you start the discussions that lead you to this is different. This is going to find patterns that aren't obvious. This is going to introduce domain knowledge into workflows that we haven't had before.
Sometimes that means we're going to be able to make better decisions. Sometimes that means we're going to be able to see forward further. Sometimes that means we're going to be able to take in all of these complex symptoms and diagnose a root cause and understand the implications of potential fixes. That's not something BI handles. And in order to make the move, we have to stop thinking about it in silo monolith BI dashboard.
and begin to think about holistics, systems, knowledge graphs, domain knowledge that's new, not that's existing, that's new, being extracted from the data using patterns that aren't obvious. They're there, but it's so much complexity, so much trash, so much, you know, all of this other stuff that it's not obvious to people. So we use the math to tease that domain knowledge out. And introduce it back to the business or introduce it into automation so it can handle different parts of workflows for business users, for customers. So that's different and we have to justify it in order to start the conversation. And if business leaders start asking that question, why? Why can't I just use BI? Why can't I just code this up? Why do I have to go to this next level? Then we have those conversations. You're forcing us to start with value.
**Benjamin:** So in that world, like how, how does a data team actually operate? Right. Then where do these types of projects then come from? Right. Now I have a business leader coming saying, Hey, I need a dashboard for, for XYZ in this kind of a different world you're describing how, what types of questions kind of who asked these questions.
**Vin Vashishta:** Yep. The biggest problem is if your boss asks for a dashboard, just say no. And aggressive staring. No.
**Eldad:** Or more important, how do I answer my boss that they don't need a dashboard?
**Benjamin:** Hehehe
**Eldad:** Let's quit.
**Vin Vashishta:** So the best way to answer that question is to explain none of this sounds like data science, does it? None of this sounds like data engineering, does it? Doesn't sound like analytics either. So let's stop forcing data scientists to do all this stuff. This is one of the primary problems. We need data literate business users. We need AI literate business users. We also need business literate data scientists and data engineers and analysts. So that's important. But we can't, and we're doing this in both sides of the equation, we can't overstep this and start asking all of our business users to become data scientists and analysts. We can't ask all of our data scientists to become product managers and strategists.
We need new roles. We need new roles in frontline organizations that are non-technical who are these hybrids where they are domain experts and they have the ability to build these dashboards. They have coding capabilities. They understand SQL. They understand hopefully one of the no SQL databases too. Let's be a little modern here. And you know, that would, that's.
**Eldad:** So you're a company and you spent the last year searching for data scientists because like, you know, that's what they said last year and there's no one to find, right? Like your whole business is based on cold calls. And right. And then that's how you feel, at least as a unit manager and somewhere, even at the most techie company. And how do you transition teams? How is the future going to look like? How do we going to like, what should we ask for when we recruit?
**Vin Vashishta:** We have to pick up these new roles. You can't expect to adopt an entirely different paradigm of product without product managers. You can't implement technical strategy, an entirely new type of strategy, without technical strategists. We can't expect CXOs to magically become these things that they've never been before, especially when the role is as big as these are. So we need new roles. Like I said, in the frontline teams, we need people who are hybrid.
**Benjamin:** Thanks for watching!
**Vin Vashishta:** They are domain experts with technical capabilities. We need in data teams and data organizations, we need people who are product managers and strategists. We need a top layer, a C level layer of leadership for the data team, who is not just a people leader, but also a leader of strategy, where they can take the AI strategy and implement it. We have to accept that this is not just a technology problem, but it's also not enough
To say technology team, all you do is technology. They also need to own product. They also need to own strategy, whether that lives in the data team or in a product organization or in a strategy organization. That's, you know, organizational structure is really what works best for the business. Big businesses, small businesses structured differently, different industries, there are different structures that you're going to put together. But.
**Benjamin:** Thank you.
**Vin Vashishta:** You have to acknowledge there are new roles that are necessary to succeed. Isn't enough to just throw technologists at this problem. It won't solve the value side of the equation. It just gets you a whole bunch of technology.
**Eldad:** Because you're not throwing enough technology at the problem. Just throwing a... Just enough technology will not solve it.
**Benjamin:** So.
**Vin Vashishta:** Let's get some AI into the product manager space and into the strategist. Let's hire GPT as our product strategist and see how that works. You and me, we got a patent. We got to get together after this. We can get some funding for that.
**Eldad:** It will be a big thing, it will be a big thing, I'm telling you.
**Benjamin:** Nice. So like, say I'm a listener, I'm listening to this podcast and saying, Hey, this sounds cool. And I'm giving you a call, right? I'm saying, Hey, like Vin kind of V squared, like come to us, implement these things. Like, how do you actually approach this? Right? Cause like what you're describing at a high level seems to make sense, but actually getting this into an organization, like there's going to be a lot of friction. Right? You need kind of buy in from the highest levels. You're kind of talking about big changes. How is that actually, I guess that actually working?
**Vin Vashishta:** Mm-hmm.
**Eldad:** Nah.
**Eldad:** So Vin is coming in after the changes and then everyone is ready to listen. I think like, right? No, I can need to know when to... Yeah, go ahead.
**Vin Vashishta:** Yeah, exactly. Yeah. If you are in a data organization and you're thinking about bringing me in, you're at the wrong place. It won't succeed. You really need a strategist, need a product manager to begin the process with your C-level leaders. You need buy-in from the top level. And that can come from two different directions. One, you can.
establish a track record of success, just deliver some, you know, they, they call them quick wins, but they're not so quick. And the win is much bigger than most people expect it to be. Deliver a couple of those that actually have top and bottom line impacts. You get attention from CXOs. They, they will show up because growth isn't easy anymore and they're on the hook to save costs. So if you're doing both, they will come to you and that's when you can make the pitch and say, we need more.
If you really want the top level value, you've seen nothing so far. We can do a lot more, but we need new roles and we need people to come in who can help establish this connection across the enterprise. The other way you can do it is by scaring the daylights out of your CXOs, because they already are scared. If your C level or founder has made some, some wild claims about how they're using AI and going to
drive growth with it. Um, we all know many of those claims are not materializing as fast as they've promised and they're, they're running out of time. Investors are punishing companies, especially if you're a startup. Oh yeah. If you're a startup or if you have, uh, you know, if you're a multi-billion dollar company, if you've said the AI, yeah, if you've said the AI story, but you have not delivered the AI results yet, and that means, you know, money, cash.
**Eldad:** Listen to that Benjamin.
**Vin Vashishta:** It has to show up. If you haven't done it yet, your CEO right now is sweating. The board's calling. Investors are asking for more than a story. Yep.
**Eldad:** Everyone is all in on AI. Everyone is all in on AI. It's like has to succeed for everyone
**Vin Vashishta:** Yeah, they're all in, but now they want to see some cash. They want to see the chips show up. It's, you know, you can push all your chips into the middle of the table. If they just disappear, they're not going to keep coming. No one will continue to fund you if you keep throwing chips at nothing. And that's the, there's the duality. On the one hand, you can demonstrate how powerful the opportunity is. And on the other side, you can demonstrate almost like a lifeline or a life preserver.
**Eldad:** Thanks. Yes.
**Vin Vashishta:** where you reach out to the business and say, look, we've never had this relationship before, but we need to build it now. You need to bring me your problems, I'll solve them. Don't bring me a technology, don't tell me generative AI. I figured that out, remember, I'm the technology person. We will get to AI, but we're not going to get to it the way you thought we were. But we're still going to give you cash, and that's what your investors want. And that is the life vest I am going to give you. Would you like to take it? That's the second way in is a little bit of fear.
**Eldad:** Classic selling, you know, always fear the customer, always scare the customer.
**Vin Vashishta:** Mm-hmm. Well, I think if the opportunity doesn't work, you have to. I start with the carrot. Yeah, I start with the carrot and say this is the size of the opportunity. And if that doesn't work, then well, let me show you this scary monster in the closet.
**Eldad:** Yes. One way or another. Yes.
**Eldad:** Makes perfect sense.
**Benjamin:** So when, when you said have some quick wins in the beginning, get your CXO on board and so on, like that actually already assumes that you can show that your quick wins are driving profit in some way, right? Like that seems a bit cyclic. Uh, so how, like, how do I get to that in the first place? Cause then you're saying, okay, then you're
**Vin Vashishta:** Yes. Why would you start an initiative if you didn't know what it was going to return? I mean if I walked up to you on the street and said, give me five bucks, what would be your first question be? I mean, really, right? Yeah, that's, but I mean if you're on the street, I walk up to you, I'm just somebody that you've seen before a couple of times, and I say, hey, I need five bucks. You're not just going to shell it out. If I come over to you and say, hey, I need your help, come on, let's go do this thing. You're gonna say, what thing? Why am I doing this? Wait, hold on, slow down.
**Benjamin:** Hell no.
**Eldad:** Why just five?
**Vin Vashishta:** That's what we should do. We shouldn't just start working. We shouldn't take all of this, this stream of consciousness that's coming to us from the business and accepted at face value as being something that'll generate returns. We have sort of this magic aura around us right now because the business needs us. And it's a true need. So being able to push back is one of our superpowers. We can say, look, I want to return value to you. There's one of me, there's 80 of you.
let's figure out what we should be working on to get the highest returns. And if you all want something done tomorrow, let's hire some people. By the way, that's going to cost a lot. So you'd better have some ROI behind this. Let's start talking about this just in basic business terms. And if you don't have the ability to have that conversation, it's okay. You're a data engineer. You're an analyst, you're a data scientist. You don't have an MBA. That's okay. You weren't hired to have an MBA.
Talk about this in pragmatic terms and say, look, I don't do ROI, I don't do product strategy, I don't do AI strategy. Maybe we should get someone, maybe we should hire someone. And it doesn't have to be a consultant. I mean, I know this almost sounds like I'm pitching my services, but really I'm not. Train somebody into the role. There's tons of people in the business, in the data organization who want to go into these new roles, who want to own the product more.
who want to have more control over the direction of strategy, of that high level, where are we going to go with this technology, of developing a vision. There are people who want to do this, just upskill them. Don't pay me a really large number of dollars per hour to be in your business for over a year. Don't do that, hire some people. You will be happier in the end with that approach. And so that's the...
say this more often than I really pitch my own services, give your people career paths, they'll stay longer. Explain that we need new talent and you'll be more successful. Don't try to take on a role that you're not qualified for and half the time don't want to do in the first place. Don't feel pressured into that. Just start the conversation with value. And when the questions come up, don't be the engineer. We want to be the one with all the answers. Just say, look, I don't do this. It's like asking you to write something in you know, pick a programming language that no one ever uses anymore. If someone were to ask you to do that, yeah. Yeah. If somebody said, I need you to write an enterprise app and I don't know, Fortran and you'd probably, um, maybe we need someone else for this. I mean, I'll take a look, but
**Eldad:** VBA.
**Benjamin:** What's I never even heard of that? No, just kidding.
**Eldad:** on purpose, that's why I used it.
**Vin Vashishta:** Look at it the same way. If you're not a strategist, don't feel pushed into it.
**Eldad:** I'll take it to the bank.
**Vin Vashishta:** I'll actually get the five bucks. Can I have that five?
**Eldad:** I'll take that.
**Benjamin:** this. We, to maybe hit you with a more controversial question, we had Joe Rice on the last episode and he actually said, I'm tired of talking about profit all the time and ROI, right? And it's showing that as an, exactly that kind of as an industry, we're talking too much about this and actually shows that we're not delivering enough value and enough ROI because we need to keep cycling on the same thing over and over. Like what you, yeah.
**Vin Vashishta:** Mm-hmm. Yep. I watched that one. Yep. Yes. Yep. Mm-hmm.
**Vin Vashishta:** Right.
**Benjamin:** What would be your take on that?
**Vin Vashishta:** Yeah, he's right. I mean, it should be like chewing gum. Chewing gum doesn't need to advertise. It just tells you what flavor it is. Pepsi doesn't advertise. It's just a Pepsi. It is sitting right there on the shelf. It doesn't say here's my value proposition. It just advertises somebody drinking a Pepsi and looking happy. Like that's where I would want to get, but, you know, unlike Joe, I don't live in that world, Joe's kind of a rock star. And his book, let's just say we both have books, but his book appears to be doing a just...
just a little better than mine. So he's, yes, yes. But the people that I talk to don't understand it. And the reason why it's really two-sided, one side of the problem is they don't know how much it's going to cost them and how much work it is. It's enterprise wide. When you're looking at data engineering, just as a niche.
**Eldad:** Wait, wait! Patience, patience!
**Vin Vashishta:** You think it's just an engineering thing. If we move our data from all these other places to one place, everything's cool, but it isn't. And that's the, you know, if you read Joe's book, he's actually in that book. He gives you that bad news a few times. Yeah. Guess what? Most of the data you have worthless. And when you centralize it, it's still worthless. So you've just spent a whole bunch of money centralizing, you know, landfill. It's.
**Eldad:** to have one version of this centralized worthless truth. So at least it's one version versus.
**Benjamin:** It's the best landfill.
**Vin Vashishta:** But yes, I mean, it has to be truth, though. If there's trash in the data, it's not truth. You know, it's and centralized. Oh, yes. It's dangerous. Well, and centralization is wonderful, but we have to talk about centralization differently. We're centralizing knowledge, not data. The data isn't a display element.
**Eldad:** No, but that's true. The perception that data is oil is nice. It's a nice thing. Yeah.
**Vin Vashishta:** the data contains more complex domain knowledge. And we're centralizing it so that it is easier to take that domain knowledge out and begin to use it in multiple ways for customers and internally so that we can create and deliver about value more efficiently. That's the goal. And if we spend most of our time moving bad data around, it costs the same to move bad data around.
as it does to gather new good data. So why don't we just gather good new data? Well, it's because the business doesn't understand the value of it. It's expensive. And they think, but we already have data at home. And unfortunately, the data that they have at home is kind of like the dollar store knockoff version of whatever candy bar you wanted to buy in the store, where you wanted to buy the good bubble gum.
**Eldad:** You know, remember, remember swatch the watches. And remember when we used to collect them as if those will be like Rolex. You still die. Just some still do.
**Vin Vashishta:** Oh yes, yes. Used to. Wait, used to? I am the last member of Members Only. I just want to let everybody know.
**Eldad:** So yeah, so welcome to the club. And it turned out to be worthless, but it was a lot of fun. And the thing is people spend tons of energy on, go ahead. People spent tons of energy on building data pipelines. And now you're coming and you're telling them like, why did you do that in the first place? Like, why did you build all of those data pipelines that are designed to scale, designed to build lakes, designed to offload and unload and transform.
And that's strategy. So that was 10 years, all about that was the strategy, like getting the data in. Now people start to realize, okay, so that's worthless, or at least it's not as equal. As oil, so it's not more data, more value. It's just more resolution. And as you said, without the domain, it's pixels. Nobody can understand them. So.
How do you negotiate, given the fact that data is only becoming like managing data and utilizing data really becomes super complicated and users are going to work every day. They need to get the job done. So.
**Vin Vashishta:** Yep. Strategy has to be lightweight. Strategy should be a framework for decision-making that informs and improves decision-making across the enterprise about data and AI. That's what strategy should do. If you think that gathering data is a strategy, we're in trouble. And this is the problem. You know, when Joe says, I don't want to talk about value anymore. We want to get to the point where we understand the value of data gathering.
And that means we have to do that education step where we explain, you have to gather data differently. Gathering it for BI has BI value, but you didn't hire me to do BI. You hired me to do data science. Data science requires a different type of data. So if you force your data engineers into a digital paradigm and give them digitally gathered data, well, guess what you're going to get in that data repository. So if we start with tactics, if we start with technology.
All we get is more tactics and technology. If we start tactical, you're going to hire a ton of individual contributors. And a McKinsey study is, I think two months ago, where they looked at what companies who have several successful deployments in production versus companies who haven't been able to succeed with a deployment in production. The lower maturity, less production deployments.
valued AI researchers, data scientists, individual contributors the most when it came to roles. The people who had a ton, the companies who had several multiple production deployments, leaders, translators, strategists, it's a different kind of problem. So there's, we want to get to that point. And I think that was Joe's point is we want to get there.
to where you can just say, I'm going to build a model and everyone goes, yeah, I got it. Where it is so ubiquitous and we've delivered so much value. And he's also right about, we've been talking about delivering value for so long, we're in danger of hitting, you know, sort of crypto web three territory, where there's a whole lot of talk, there's a whole lot of hype, there's not a whole lot of delivery. There's not a whole lot of value creation.
And even if you're not a scam, you can sure get labeled as one.
**Eldad:** So you're saying everything that runs on the GPU is a scam? Everything you run? Any data-related project that ends on the GPU? We should be worried. But I'm not correlating anything. We love you, NVIDIA.
**Vin Vashishta:** I'm sorry, Nvidia. Just wait. Yeah, I'm sorry, Nvidia. I didn't say, please give me access to GPUs. Please don't take those away from me. Yeah. It's interesting that you look at some of what was being developed for Web3, legitimate. Had use cases, had applications, but there was so much over promise under deliver, just across the board, even where there could have been value.
**Eldad:** I'm sorry, I apologize.
**Vin Vashishta:** that technology that had potential was swept aside and all of it's a con. And that's the danger we're in right now.
**Eldad:** So in 100 years from now, they will look at us and say, oh, it took them so fast to get AI running. Nobody will remember the reports and the dashboards. And everything we've done, it will all be replaced by just a system that works. So that's the next step for everyone. And we're just here to support it, each with its own value in the big data value chain. So thank you for that optimistic outlook for everyone.
**Vin Vashishta:** We're putting in the pieces. Yeah. We're putting in the pieces and laying the foundation. And if we listen to the right voices, you know, the Joe Reis, if we listen to those people, you're going to have a more successful approach than if we listen to some of the Thread Boys and the people who get way too much attention, media-wise, social media-wise.
It's great to have an announcement. I want to see a product that works. If it, if announcement does not produce product that works, we have to start discounting those people and letting those voices just kind of fade into nowhere. And focus on being the people that build things, being the people that deliver value. There's, there's this tendency of, Oh, we want to make awareness. You know, we've got to show that these people are no, just believe me, they will fade into nowhere.
They just go away. That's the great thing about it is if you don't deliver two or three times, people just ignore you. So all you have to do is focus on delivering and we'll be fine.
**Benjamin:** Thank you.
**Eldad:** I'll boom to that. This is like so important. This is so true for everything we do in life and for data sometimes as well.
**Benjamin:** Awesome. I think that was a great closing statement, Vin. Thank you so much for being on the show today. It was super fun having you. Yeah, and all the best going forward with all of the other companies where you can have impact.
**Vin Vashishta:** Thank you so much for having me. I appreciate it. And letting me say a few wild things here and there.
**Eldad:** Thank you, thank you.
# Zach Wilson on what makes a great data engineer (/blog/zach-wilson-on-what-makes-a-great-data-engineer)
How good you are at Spark or Flink ≠ how good you are at data engineering. After years of data engineering experience at Airbnb, Netflix, and Facebook, Zach Wilson is now focused on spreading the knowledge in EcZachly and all over social media. He met Benjamin Wagner to explain why data modeling and storytelling are more important than the actual tech, why data engineering is going to see more job growth than data science, and what brought him to start creating content, reaching over 250K followers on LinkedIn.
Listen on [Spotify](https://open.spotify.com/episode/18duS0TwlWuZGZWuRK7TKL?si=W3Er9xs6RXqfqTwne5E8_w) or [Apple Podcasts](https://podcasts.apple.com/us/podcast/zach-wilson-on-what-makes-a-great-data-engineer/id1561927688?i=1000610838344)
Transcript:
Zach Wilson: I've been doing big data for a long time, like, uh, since like about 2014, 2015, like I actually started my career in like data science, doing like analytics stuff, doing like Tableau and SQL and stuff like that. But I realized that I like to build more than I like to like build models and query data and stuff like that. And that's where I found this nice intersection with data engineering. And I got this job at Teradata back in 2015. And like, I've mostly been on the data engineering train since then. I've like, kind of fallen off every once in a while where I've been like, I'm done with pipelines, I don't wanna do this anymore. But then I like, I end up coming back around, I end up coming back around. Like it happened, it especially happened near the end of my time at Facebook where I was like, I'm done with data engineering, I'm out. But then I came back around. And so like, it's been great. And then more recently, last month or so, I decided it was time to do content full time and to also try some other things out. So I'm actually doing, Three separate things I'm trying to do right now. One is content. I make content on five platforms, LinkedIn, Twitter, Instagram, TikTok, and YouTube. Those are the five that I'm on right now. And
Benjamin: Awesome.
Zach Wilson: it's been fun. Really trying to get those rhythms going and get that pattern going of just video. Because I have a theory that actually all social media is gonna merge into short form one to two minute videos that they're all kind of going that way. It's interesting because I've even seen the short form videos perform really well on LinkedIn as well. Where it's like,
Benjamin: Okay.
Zach Wilson: I'm like, wow, like, ooh, this is cool. Like, I'm like, and it makes it easier for me because then I'm like, oh, I can make one video, post it five times. Whereas before it was like, I had to make a LinkedIn post and then convert it to a video. And it was just like, there was a lot more like processing and all these steps to it. Right. So I think that that's like one of the things for me is like, I really do think that video is going to be a big part of the future. But like, so that's one. The second thing is I'm doing a data engineering boot camp. So I have taken on about 50 students, and I'm doing a six week boot camp for them. I call it like, it's kind of like the good to great boot camp. So this is not like a breaking into data engineering boot camp. This is you have a job as a data engineer, and you want to learn the skills that are going to take you to the next level. And they're going to make you a better data engineer. And so that is definitely what I've been working on. The very first session of that is tomorrow. So like we're starting. day one, week one, tomorrow. And it's been a lot of planning and stuff like that. And that's been really fun. And the third thing I have is I have this platform called Tech Creator, which is going to be a platform that makes a lot of managing your job as a tech creator easier. And it's going to have these chat GPT plugins to make content creation easier, especially content creation in your voice. That's the part that's gonna be really, really exciting. That stuff should all be done in the next couple months as well. So be on the lookout for that stuff. So techcreator.io, it's pretty cool. Yeah,
Benjamin: Nice.
Zach Wilson: so yeah, that's kind of what I've been doing in the last month or so.
Benjamin: Awesome. So I'll definitely be on your lookout. And like, that's a, that's a lot of stuff to talk about. So like, let's start right in the beginning, right? Kind of like then a month back, kind of you made the decision, okay, I'll leave kind of my like comfy job at Airbnb kind of become full time, focus on kind of content creation and all of these other things you talked about, uh, like take, take us through that journey. Right? So like at this point, I think you have like more than 250,000 followers on LinkedIn. So that's obviously a kind of huge audience. What, what actually got you started kind of like? creating content,
Zach Wilson: Yeah,
Benjamin: sharing your knowledge, all of that stuff.
Zach Wilson: that's good. I think there's kind of like two journeys here that are important, and one of them is like a content creation journey, and one of them is like a confidence journey, but I think they both matter. So content creation, I've actually been doing content for a very long time. So I started posting almost daily in 2010, 2011, when I was like 16, 17 years old, and a lot of that content is really terrible. It's all on Facebook. I only have like 500 people there and I intentionally keep it small because there's so much bad content like I think I feel like I
Benjamin: Thanks for watching.
Zach Wilson: need to go back through and delete some of it because I feel like I'm going to get cancelled but but anyways.
Benjamin: But so that was also technical already, or kind of like
Zach Wilson: my thoughts and my perspectives. It just kind of shifted over time. For a while, it was just my thoughts about school and math. And then it was more like, especially after that, it was more into politics. Because three or four years, I got way into politics for a while. I was way about that. 2013 to 2016, I was obsessed. And that's all I wrote about was politics. And then I would say after that, especially after I got the job at Facebook, it started to get more technical. And then I realized, one of the things I realized about it was, it was actually in 2020, this is something that happened when I quit my job at Netflix. I remember telling people, I was like, when I quit, I was like, you're gonna know my name. I'm gonna come back around and you're gonna see me, I promise. Even though at that point I had no big following or anything like that, but I had the feeling that I knew the skills that I needed to get engagement and all that stuff. And then I really started making content consistently a couple months before I got in Airbnb. It was like December, 2020. And then that was when, yeah, things have been going pretty well since then. Like I like, it's been consistency is the main thing, right? Where... So now since then, since December 2020, I show up and post almost every day for the last two and a half years. I've missed, I think, 12 times. I've not posted 12 days in the last two and a half years. So consistency is important. And I try to shoot for like 95-ish percent of the days. So I get like one day a month, right, where I'm like, yeah, I don't need to make content today. So, but like that's, that consistency is I think one of the big important factors of like why I've experienced so much growth because that's the part that's hard. That's the part that is challenging. That's the part that like separates me from a lot of people is like, uh, is that. And I know I'd say that that's like been the main thing. And now, now it's been like, now that I have more time, it's been like very interesting because it's like, uh, there's like so many different strategies on how to like grow, like cross platform and all that stuff, but Yeah, I'd say that's my journey for sure.
Benjamin: Okay, gotcha, super cool. So, but I mean, that's interesting, right? So you actually kind of started out kind of like just posting a random stuff or kind of whatever
Zach Wilson: Mm-hmm.
Benjamin: And then kind of, I guess kind of like over time, you found more of your technical audience in a sense. That's super interesting. Like at least on my end, right? I always feel like there's a certain barrier to writing something, right? Do I actually have something to say, which is smart enough?
Zach Wilson: Yeah
Benjamin: will people care?
Zach Wilson: Definitely. And that's what I was saying about these two journeys, right? One of them is a content journey and one of them is a confidence journey. And I feel like for me, that confidence journey was like, especially after I got the staff tech lead role at Airbnb, then I was like, okay, I have the credential now, right? People are gonna care about what I have to say now. right? And that made a big boost for me and just in like my, uh, my own perspective of my own thoughts, I guess, in that like, wow, okay, I can actually do these things. I think that is actually that is tricky.
Benjamin: woke up one day, looked at your CV and you're like, hell yeah.
Zach Wilson: Yep.
Benjamin: Kind of like I have the credibility now. I can start writing about this.
Zach Wilson: Yeah, and like, I realize now though that like, that that was actually kind of stupid of me, that like actually the correct way to go is to just put your thoughts out there and let, let the algorithms and let the audience decide if it's valuable, not whether you think you are credible already. Because I mean, I see these, for example, there's a couple people on LinkedIn I follow who are in high school, right, they're 18 ish years old, and they have 25,000 followers on LinkedIn, and they're in high school, right? And it's like, okay, there's no way they have the credibility or the established they're in high school, they're still learning, right. And that's like, but that's the thing that like I learned in my journey here is that like, really, it's more about just putting your thoughts out there, putting your opinions out there and letting the feedback come. And then you can know if you know what you're talking about or not, because, you'll get the feedback from the internet about it. And like, and I think that's actually one of the beautiful things, that for me, I think was actually an unnecessary barrier. And for me, I'm like, I wish I would have started earlier. I wish I would, because I had the, I feel like I had the content writing skills for many years, for many, many years. I think I learned those in like three or four years, kind of the copywriting skills of how to write content that's engaging, as opposed to content that's maybe educational is a separate skill, but engaging content is pretty universal. And so for me, I think that that was actually a mistake and a thinking error. where I'm like, oh, I have to build up all this credibility and resume before I should speak. And that's actually like, definitely one of the things that I want anyone watching this podcast, make content, just like put your thoughts out there. Doesn't matter, doesn't matter where you're at. Like remember, there's people on LinkedIn who have 25,000 followers who are in high school and they're making money from LinkedIn. And like, there's no way they have more credibility than you, like, and so, like, and so definitely just remember that, right?
Benjamin: Awesome. Super cool. So another thing I find interesting, now you're super focused on data engineering, both in terms of technical content, right? I watched some of your TikToks today about window functions, those types of things. And really digging into how to be a successful data engineer. And then also things around how to build a career in data engineering. How do you make the decision? What's worthwhile content?
Zach Wilson: Yeah, that's great. I think some of it is there's kind of two sides to it is that like some of it is around I get feedback from people. This is where TikTok is really amazing, where LinkedIn needs to do better on this. But like TikTok, like the people in the comment section, they give me so many ideas. They give me so many ideas where I'm like, yeah, that's totally right. That's totally like the thing that I need to be talking about right now. And like, because for me, that's actually interesting and kind of perplexing about content is that the content that I really like to write about is the content that is like, I would say more like intermediate to advanced data engineering. But if you post that stuff, like a lot of times it kind of bombs because it's just not relevant to most people. And so and that makes it so the content doesn't do well. But like on the flip side, I don't want to do just like 101 content because I just feel that that market is already tabbed. that there's enough creators teaching select and group by and where, and I don't need to be teaching that. Right? So for me, I'm trying to find a way to make these more intermediate advanced concepts more approachable. And that's the tricky part. Right? Because usually my V1, I write it and I'm like, ah, no, this is more for practitioners, not for people learning. And so it's an interesting balancing act. But then there's the other side of it of like, okay, more than just the technical content, like of how to do data engineering, like the actual nuts and bolts of it, but how to grow your career, how to be, because I feel it's how to do data engineering and how to be a data engineer. Those are separate things, right? And so I definitely think that that's... One of the things for me is like, especially on LinkedIn, this is something I learned. There's this lady named Leah Turner on LinkedIn. Definitely follow her, she's amazing. And she has this framework where there's three, there's three types of content, right? You have show content, which is just explaining something, a concept, right? Then you have grow content, which is gonna be more around like inspiring people to take action or like really speaking to their emotions. Right? And then there's a third piece of content which is like get to know, right? And that's how they can learn more about me. They can learn more about my story and my journey. And I found that like if you kind of those three buckets, like for me personally, I think the correct breakdown of those three buckets is show content should be like half, grow content should be like 30% or 40%. and get to know content should be like 10%. It's like the spice that goes in there a little bit much. Because you want to be careful to get to know stuff because if you post it too often, people just unfollow, they boo you. But if you post it, but infrequently though, sometimes those posts can be your most viral pieces though. And it's like, but you just don't want to talk about yourself all the time, otherwise people are like, this guy's full of himself, he only wants to talk about himself. And so you want to be careful with that type of content because it's powerful, but also dangerous. So
Benjamin: say it's kind of one of those three pillars, like, right. So on my end, like I'm just kind of like C++ database nerd. It kind of, I care about how to build concurrent index structures. I don't know how to build a fast multi-threaded join and so on. And whenever I look at data engineering, right. And I interface with a lot of data engineers and Firebolt kind of, it's always seems really daunting just because like the kind of breath of the field in terms of the, like just amount of different technologies you have, just seems so crazy, right? So how do you decide on what to focus on in terms of actually teaching people skills that are valuable for their career?
Zach Wilson: Yeah, well, that's great. That's actually a great, great question. So for my boot camp, for example, there's six weeks. Only two of those weeks are actually tech specific. The other four weeks are not tech specific. They're tech agnostic. So where, because I think there's a couple things and a couple philosophies that are really important in data engineering that are actually like. They apply regardless of if you're using Spark or Snowflake or Databricks or Presto or Flink or like whatever, you know, tech you want to use for it. And there's a couple of them, like one is like around like data modeling, how to do, how to do proper data modeling for dimensions and facts and how to really get those things like compacted down. And there's a lot of trade-offs in that space that is very art, very, it's not a science and it's a lot more art and you have to understand like your consumers, and they have to have that empathy. And that part is powerful. For example, for me, when I was working at Airbnb, I would say 80% to 90% of the impact I had was going to be in two buckets. It was in the leadership bucket of inspiring other people and helping them grow. And the other one is data modeling and making robust data models that can then be used by a large number of people downstream. What wasn't as important was like how good I was at Spark. Like, I mean, I go look at Spark as more of like a means of accomplishing something or like as like a, it's just one kind of path forward and that like you just, it solves the problem and maybe it can be a little bit faster. Maybe it can solve those things, but really this is the fundamental thing that I think a lot of data engineers need to remember is your product that you sell is data. It's not a pipeline. A pipeline, it can help in terms of maintenance and pain and suffering. If your pipeline sucks, then the data is going to be annoying. But generally speaking, the value you're providing is in the data sets that you provide. And if those data sets are not modeled properly, then that's where you can have a lot of unnecessary costs. And in big tech companies, these mistakes actually cost them millions and millions and millions of dollars a year. Because if you don't model things the right way, then downstream the compression doesn't work the same way. And then it can blow the data up again. And there's a lot of interesting tricky things that I've noticed with how data modeling works. So that's one. I'd say another kind of tech agnostic thing is around data quality and understanding Again, there's technologies here like Amazon DQ and great expectations, and there's going to be 10 trillion more coming. And but like, it's more again around like how to test data like, of like, is this quality? Is it not quality? How to validate it? Right? That's very like agnostic of like the tech that you're using. And you should definitely be able to do that. And then the last bucket of things that are as tech agnostic is storytelling, right? Can you tell a compelling story? Can you like make some cool charts? Can you persuade people to give you the time to make this data and other things like that? Because there's like the story, there's like the before data story, and then there's also the after you have data story. And both of those stories matter. And being able to construct those narratives in a compelling way, very important persuasion. And then the tech, and then the tech is the last one. And like, in some ways, I think the last one, but also not as important. But it's tricky because, and this is the thing I hate about industry in some regards, is If you go into an interview, right, and you go into the interview, like, 80% of the questions are going to be on like Spark or Flink or like, and be like, Oh, do you know this very specific minor detail about Spark? And it's like, dude, like, this doesn't matter that much, actually, in the end. But like, but that's how it's things are tested, right? And I hope that industry changes in that way.
Benjamin: You need to be able to tune the number of shuffle partitions and those types of things.
Zach Wilson: Yeah, exactly.
Benjamin: Gotcha. Looking at this bootcamp, because you're framing it in that context, you said in the beginning, the goal is from good to great. So usually those will be people who already know Spark, know how to maybe write Scala code for their Spark stuff, those types of things. On the other end, and if there's hard to say from good to great, is you have someone, maybe that... Okay, not the influencer with 25,000 followers talking about data engineering, but just someone wrapping up high school who wants to get into data engineering. Kind of, you have to pick up some technology, right? And like there, it just seems kind of like daunting to in a sense, like make that choice, right? Kind of what horses do you bet on? Kind of with what do you get started to actually start with that career?
Zach Wilson: I mean, I totally agree. I think it's similar to, so I'm a pretty athletic, sporty guy, right? And one of the things that I remember as a kid growing up, my parents were always like, you gotta do sports, right? And then I was like, okay. And then I tried a bunch of them. I tried soccer, I tried basketball, I tried baseball. I tried all these different sports that were all interesting and different and like. But one of the things that I learned about it was, especially going through that process, was like, yeah, you just got to pick one and be like, I'm going to get good at this one. And for me, that was basketball. I mean, I got lucky. I'm tall. I'm 6'2". So basketball was the easy, obvious choice. And that's one of the things that's tricky about tech sometimes is that the choices aren't so obvious, right? They're not like, oh, yeah, this one is 6'7", and this one's 4'2". So we should go with the taller one or whatever, right? It's not that obvious a lot of the time. And so I think there's kind of a couple pieces there, like on how to pick technologies is, one is gonna be like, okay, what do you see on social media? I'm not gonna say all social media because I still don't trust TikTok here. But on LinkedIn, if you have enough of a network on LinkedIn, do a poll. Polls are broken on LinkedIn. Like if you do a poll on LinkedIn, like even if you have no followers, It's going to be seen by like 10,000 people because polls are broken and they're very good and the reach they get is too good. And so you can learn, you can ask, right? And I found that there's like two or three really high fidelity sources of like where to get like good information. You have LinkedIn. The thing about LinkedIn though is it doesn't give you very, it's not as good about negative feedback. So if you're like asking someone to be like, hey, say why my video sucks. Most people aren't going to do that in the comments section on LinkedIn because they're like, I don't want to look like an asshole. And so that's one. Reddit, Reddit's better for that. Reddit's almost too good for that. If you want people to tear you down, go to Reddit. Reddit or blinds even better if you really want to get... But those places, you can also get the more guidance on like, okay, should I learn these texts or these texts? And I think really the big things are going to be just like... learning the languages first, SQL Python, right? If you really want to break into the SQL Python, learn the languages first, and then you can learn the tech after that actually. Cause you can do, like, if you just do like Postgres to learn SQL, just like a very basic database, then you can build into the more complicated ones like Snowflake or Firebolt or Spark or whatever you want to use, right? And then like, you can kind of go into the cloud that way and like... I found that kind of building more locally first and kind of learning languages that way. I've had more success with some students kind of teaching them that way, as opposed to being like, okay, here's how you set up an EC2 instance on AWS and now you have a computer in the cloud and it's gonna crunch the data for you. And a lot of that feels like a lot more complexity that they don't really need yet until they have more confidence in their own skills.
Benjamin: Right. Yeah, that makes perfect sense. I mean, like looking at SQL databases, I guess kind of one of the good things is looking at the space right now is like so many systems are actually kind of converging around the postgres dialect. There's not like you need to kind of learn like seven different kind of flavors of like SQL or like window function syntaxes, whatever. Actually, like a lot of the system tend to behave at least more similar to data than maybe they used to some time ago. So... Um, that's, that's super cool. So thanks, thanks for the insights there. I mean, I'm sure a lot of listeners will appreciate that. So zooming out a bit, right from the kind of specific technology, like, do you see any kind of big kind of trends at the moment? Like when I look at LinkedIn, for example, like one thing that keeps coming up is like data observability, kind of data monitoring, data quality, like those seem to be what some of the things kind of generating a lot of, a lot of buzz. What else is out there?
Zach Wilson: Yeah, like data monitoring, ML ops, data versioning, there's all sorts of interesting things. And then there's a couple of them that come back around sometimes, like data mesh. I hear data mesh once every three months or something like that. And I'm like, hey, it's there. It's a thing. Right? But I think a couple of things that I really am seeing is definitely data observability of like, yo, how is this data changing over time? And it's very closely linked with data quality. And honestly, they should be closely linked because if you aren't aware of how your data, like the shape of your data, what it looks like over time, then you really don't have good data quality checks because you haven't done your due diligence on looking at what is normal and what is abnormal. I mean, there's data quality checks out there that are very easy to know, or what is normal and not normal. Like, is there any data? No data is abnormal, right? Or like this column's null when it should never be null, that's abnormal, very easy check. But then things, it could get like what you define as normal versus abnormal, it gets more and more complicated as like you look at more and more different data points in together. And that's where, you know, if you look like week over week row counts, that's going to be one that what is abnormal versus normal, it depends. Because a lot of times those week over week row counts, like on Christmas day, they fail because there's not as much data or there's too much data. And like, and it's just, it's actually not like you're looking at the wrong period instead of like week over week. You really should be looking year over year and looking at it kind of on like zooming out to find the actual pattern that matters the most. And I mean, that's why people do week over week, right, instead of day over day, because of like the Sunday, Monday phenomena. And Sunday, Monday is super annoying as well. That one's very common to like trip people up. That's why week over week is better, but it still misses the holiday patterns, right? So I think that those kinds of observability things are super important because it's linked to quality and that is linked to trust. And because it's without quality, you don't have trust, right? And definitely, I think that that I would say is the big thing that I definitely been seeing. I've also been seeing a little bit more of a push towards like streaming and trying to get more people involved with like, I've been hearing about ClickHouse like so much recently, like everyone's like, you got to try ClickHouse, you got to try ClickHouse. I have not tried ClickHouse yet, but I need to just because it's been I've seen it in like every single comment section of all my posts. So yeah, for sure.
Benjamin: Nice. Okay. Super, super cool. Super interesting. So on, on that kind of up observability space slash kind of data quality, like that as well, kind of looking in from the outside, like that seems extremely kind of fragmented, like there's a lot of companies around this kind of open source frameworks, etc. Do you see that converging in any way, or is it actually getting worse and kind of there's a new thing popping up every day.
Zach Wilson: Yeah, I think it's going to be similar to a distributed compute environment. Especially when a competitor gets tested, then you see this proliferation happen. And this is also happening with orchestration, because you have Airflow is the giant incumbent. And then there's all these other orchestration layers that are competing. You can see it proliferating a little bit. And then, generally, what happens... In distributed compute, this happened with... You started with Java MapReduce. MapReduce was the king. And then you had all of these things. You had Hive and Pig and Drill. There were seven other ones. And then Spark came along. And then everyone just realized that Spark was... the spark was just so much better than all of those other ones that like these companies made these big pushes of like, okay, we need to everyone get on Spark. And that's great when you have that consistency, right? Where it's like, hey, everyone's doing the same thing. And we're all talking about the same stuff. And I think that like, one of the companies that I find that is probably going to be leading there is definitely dbt. I think dbt is one of the ones that really has a good like, like foothold in that kind data quality environment. It's so good, it's crazy. So I live by this place called Death by Taco, and it's dbt, right? And I'm always, every time I walk by, it's like, I'm just like, it's so funny. And I get that, I get a laugh about that almost every day now. So, and like, but yeah, I think that that's what's gonna happen though, is that there's gonna be this kind of, same thing, proliferation, people learn from each other, and then there's gonna probably be more of a consolidation sort of thing that happens. But maybe not because the thing is, is like the consolidation stuff really only ever happens when you have a technology that is like in order of magnitude better, right? If it's not in order of magnitude better, then there's not enough of a motivation to get people to switch. And that's, and I think that's one of the things that's tricky with Airflow is that Airflow is like, is kind of annoying to work with, but the replacements aren't good enough. to make the switch, to pay the price to make the switch. And I think that's one of the things that these other orchestration companies are realizing is that like that is the hard, that's the hard part for sure.
Benjamin: Right, gotcha. So looking back, right, when you started in 2014, you were also working on data pipelines to Tetra. It's not like data quality didn't matter back then, right? It was equally important in a sense. So how did you actually tackle these problems back then? Why do you only have these platforms, in a sense, emerging now?
Zach Wilson: Yeah, I mean, also back then, like pipeline development was a lot slower. Like that was, I guess that was one of the things, you couldn't like, uh, yet at when I worked at Teradata, right. Uh, the pipelines we worked on, they didn't even have Hive yet adopted yet at Teradata, even though Hive existed, they didn't have it yet. So we just had to use Java MapReduce for everything. And so, uh, that's like where you essentially have to write your own. query engine parser thing. It feels like almost one layer down, kind of the work that you do. But you do it for every single job, right? And that's how it used to be. Right. And so, um, like, at least from the perspective of the pipeline development, that was a part of it, but the testing part, you're totally right in the fact that like after the big data part is crunched, that part has been pretty similar, right? Where it's like, then you have this pattern, right? You have a thing called, and I've seen this pattern like for my whole career, essentially. It's crazy, I've also seen it at companies that don't use this. Like Facebook actually was an exception to this pattern. So this pattern called write audit publish, where you have your big data pipeline right to a staging table, then you run your audits to test it. And then if they pass, you move the data from staging to production. And that's the contract, right? And the audits, are your guarantees. And that's where you can run some really lightweight kind of SQL queries or sample. Back then, we didn't even do it on the full data set. We just did sampling because the SQL, we didn't have Presto. We didn't have the nice distributed SQL queries. And we weren't going to write another MapReduce job to test the data that we already wrote with the other MapReduce job, which was so painful. So like what we did instead is like we just did sampling and we're like, okay, that's good enough. That's one of the things I like about this new world though is that like with these new technologies, you do get guarantees. Like you actually do get full guarantees of like this data is high quality. And so that's like where you do have this interesting, that's why I like this new world, even though there are so many tools and like I hate the proliferation of things because it's like, why can't we all just agree on something? ut I also think that I really like the up level and quality for sure.
Benjamin: Right, so in a sense, the maturation of the space and some things becoming easier or less work lifts you up to not at the point where you can actually focus more on those quality aspects. Kind of.
Zach Wilson: Mm-hmm.
Benjamin: OK, interesting. Cool. All right. Nice. So yeah. Look, any kind of closing words on your end, right? Anything you wanted to talk about
Zach Wilson: Yeah.
Benjamin: that we didn't get to that.
Zach Wilson: Yeah. Yeah, for sure. So, I mean, I think that there's a lot in this data engineering world that is important to remember. And, this is a couple of closing things that I think here are like, one is that getting into this field is not like data science. It's easier. And if you are someone who's considering getting into it, definitely, try it out. It's something that is maybe a six month to a one year road. Like you don't even really need a computer science degree. There's a lot of people who I know who are getting into data engineering with no degree or like an unrelated degree. And so you can really get pretty far into this field. If you could just get the right learnings and get the right teachings from the right people. And that is something that can really change your life because data engineering pays pretty well and it's going to be a big part of the future. There's a lot of job growth in this area. My opinion, it's actually the job that is going to see way more job growth than data science because people are realizing that a lot of the data science roles that they hired for were actually data engineering roles. And that's a big kind of shift that's happening is making data engineering grow really quickly. So yeah, that's my main closing thing is like, if you have any thoughts about breaking in and you want to try it, definitely try it. Yeah, for sure.
Benjamin: Awesome, perfect. Then thank you so much for joining in, Zach. It was a pleasure having you.
Zach Wilson: Awesome. Yeah, it was great being here.
# Bigabid Slashes Latency and Boosts Query Performance 400x Using AWS and Firebolt (/customers/bigabid-slashes-latency-and-boosts-query-performance-400x-using-aws-and-firebolt)
## Executive Summary [#executive-summary]
Israel-based Bigabid is a digital advertising technology company that uses big data and machine learning (ML) to help application developers increase the take-up and usage of their apps while optimizing advertising spend. Its infrastructure was built using massive SQL databases running on multiple Amazon Web Services (AWS), but they weren't able to query data fast enough to produce the required results. In 2022, Bigabid started working with AWS Partner Firebolt to improve performance and consolidate its internal business intelligence and analytics platforms. Bigabid has now improved its search performance 400 times. It now only takes a few seconds to deliver analytics results that would previously have taken days or weeks to calculate.
## Building on an AWS Foundation [#building-on-an-aws-foundation]
Digital advertising technology company [Bigabid](https://www.bigabid.com/) uses big data and machine learning (ML) to drive app growth for developers. Founded in Israel in 2016, Bigabid's platform processes vast amounts of data and connects with multiple ad suppliers and ad exchanges in near real-time to provide clients with insights into how their apps are performing. This allows Bigabid's clients to target highly specific audiences, increasing app usage, and optimizing their advertising spend.
Bigabid's business intelligence (BI) platform analyzes how a client's app is being used, measuring everything, including impressions, ad clicks, app installs, and in-app purchases. It also has a separate internal data analysis platform that is used by its developers and campaign managers to continuously optimize performance. Bigabid's AdTech platform was built using AWS. It uses [Amazon Simple Storage Solution](https://aws.amazon.com/s3/) (Amazon S3) object storage, built to retrieve any amount of data from anywhere, for its data lakes, and [Amazon Elastic Compute Cloud](https://aws.amazon.com/ec2/) (Amazon EC2), which provides secure and resizable compute capacity for virtually any workload. Yaron Cohen-Leo, BI team lead at Bigabid, states that they chose AWS because of its reputation for reliability and its diverse range of managed services, explaining "We use many AWS tools for various cases. Having such a robust feature set helps simplify our data efforts across the company."
Firebolt is a massive improvement to our BI efforts. Using the same test dataset of 100 million
records, other databases took minutes, Firebolt analyzed in seconds.
## AWS Partner Firebolt Blows Away the Competition in Evaluations [#aws-partner-firebolt-blows-away-the-competition-in-evaluations]
Bigabid's analytical databases were originally based on MySQL and were falling short of the performance the company required. It was taking days to generate data insights and the company struggled to view data older than three months due to process-heavy data aggregations, that made it impossible to compare results seasonally, or year to year. The company wanted to do more. It wanted to be able to analyze data for a million ad auctions every second and manipulate data lakes containing hundreds of terabytes of data in near real-time while accessing tables with billions of rows to create hundreds of live dashboards.
To do this, Bigabid undertook the ambitious project of building a high-performance big data infrastructure. This would require it to find a high-performance database and merge its internal BI and data analysis platforms into a central data platform. Bigabid evaluated several high-performance database options and was impressed with [AWS Partner](https://partners.amazonaws.com/partners/0010h00001cBotKAAS/Firebolt%20Analytics%20Ltd) [Firebolt](https://www.firebolt.io/home-alt4). "Firebolt is a massive improvement to our BI efforts," says Cohen-Leo. "Using the same test dataset of 100 million records, other databases took minutes, Firebolt analyzed in seconds."
In August 2022, Bigabid chose to adopt Firebolt's analytics using its existing Amazon S3 data lake and merged its BI and analytics systems. The results were so impressive that by the start of 2023 the company had completed its migration project and was optimizing its systems and building new dashboards. "We now rely on Firebolt and AWS," says Cohen-Leo. "We don't have to manage anything—everything is in one place. We can query a database containing 30 billion records and receive results in a second."
We can query a database containing 30 billion records and receive results in a second.
## Query Performance Improved 400x with Firebolt [#query-performance-improved-400x-with-firebolt]
A year on from starting the project, Bigabid has significantly optimized its data warehouse resources using AWS and Firebolt. Using Firebolt, the query response times for searching a 31 TB table have improved 400 times over the previous system. Firebolt uses its own compression system, which has reduced the amount of storage needed for the same table from 31 TB to 7 TB. "The Firebolt compression means we need less storage, which also reduces our costs," says Cohen-Leo.
The migration has also solved Bigabid's challenge of analyzing older data. Now, the data is delivered to dashboards in near real-time, and there is no longer a limit on how far back the data can be queried. "Now we can easily analyze data for seasonal and year-on-year changes, which produces more valuable business insights," says Cohen-Leo. "We have a single BI dashboard that reveals the real-time state of our business at a glance—and it can be customized in less than a second."
Now, a single table serves both the analytics and BI roles and is updated in near-real time. From this master dashboard, the company has created many dashboards to provide more granular information to support different clients and provide insights into other parts of the business. Looking to the future, Bigabid expects to be able to squeeze even more power out of the Firebolt and AWS infrastructure. "As we grow rapidly, our requirements increase, so we need to continue optimizing the platform and gain a better understanding of the increasing volume of data," says Cohen-Leo.
# How Similarweb uses Firebolt to deliver sub-second analytics over more than 1 trillion rows (/customers/how-similarweb-delivers-sub-second-analytics-over-1-trillion-rows-with-firebolt-over-aws)
## Abstract [#abstract]
Similarweb, hosted on AWS, provides detailed analytics on how end customer audiences interact with websites. This requires ingestion and processing of large volumes of clickstream data in an AWS data lake on a daily basis. The challenges of performing segmentation analysis on big data combined with the need for sub-second end user response times for customer dashboards led Similarweb to evaluate modern analytics platforms. In this session, Similarweb and Firebolt, an AWS Technology Partner, will share how they delivered sub-second, high concurrency analytics cost effectively.
Firebolt's cloud-native architecture runs on AWS, enabling organizations like Similarweb to leverage Amazon Web Services' global infrastructure, elastic compute capabilities, and enterprise-grade security to power demanding analytics workloads at scale.
## About Similarweb [#about-similarweb]
Imagine analyzing how the entire internet is used, like website analytics for all websites everywhere. That's Similarweb.
Similarweb is a big data powerhouse that collects enormous amounts of web-related data to help marketers, brands, salespeople, and other professionals analyze how audiences interact with websites. With SimilarWeb, you can easily track keyword searches that drive traffic to your site, identify which pages visitors land on, discover when they visit competitor websites instead, analyze their geographic location and mobile OS preferences, and determine whether they clicked organic or paid links.
#### The User Experience [#the-user-experience]
The first thing visitors see on SimilarWeb is a search box for researching any website. For Similarweb, delivering an exceptional user experience where users can analyze, slice and dice, and extract insights through diverse analytical views is fundamental to their mission.

For example, users can see a comprehensive overview of [godaddy.com](http://godaddy.com)'s web behavior,

then dive deeper with additional analytical views:

One particularly powerful feature allows head-to-head comparisons between multiple sites, such as [godaddy.com](https://www.godaddy.com) versus [wix.com](http://wix.com).

#### The Technical Challenge [#the-technical-challenge]
Delivering these experiences requires a purpose-built data stack that can simultaneously ingest, store, process, and analyze massive amounts of data while maintaining consistent performance.
## Similarweb's data architecture [#similarwebs-data-architecture]
Similarweb operates an AWS-based data lake that unifies raw datasets from multiple sources, including public data, partner feeds, and proprietary SimilarWeb collection methods. This centralized architecture enables efficient data processing, cleaning, and privacy protection in preparation for downstream analytics.
#### Processing Pipeline [#processing-pipeline]
The data pipeline centers on Spark and Airflow, ingesting 5TB of data daily. Machine learning models then analyze the cleansed data to create predictive insights, enabling SimilarWeb to draw conclusions about internet-wide behavior patterns from partial data points.
This foundational architecture supports diverse customer use cases across SimilarWeb's platform.
## The challenge of the 'Segment Analysis' use case [#the-challenge-of-the-segment-analysis-use-case]
Imagine you're a marketer at FootLocker who wants to compare FootLocker.com's performance against Amazon.com. It's not an apples-to-apples comparison since Amazon sells much more than shoes, but both companies compete in the footwear market, so valuable insights exist.
Similarweb wanted to let users analyze specific segments within larger websites. This would allow comparing FootLocker.com traffic directly with shoe-related searches on Amazon.com. The company recognized this as one of their most analytically complex features ever.
#### The Technical Challenges [#the-technical-challenges]
Similarweb faced several major hurdles:
**Data Scale**: The volumes are massive, making ETL processes costly and time-intensive to develop and maintain.
**Dynamic User Input**: Users can create exponentially complex combinations for comparison. Pre-processing every possible combination would be impossible.
**Query Performance**: Amazon alone generates 150GB of daily data in Similarweb's system. When users want to analyze two years of historical data using dynamic URL patterns, the system must scan enormous datasets. This becomes painfully slow, especially since multiple URLs are grouped into arrays for each user session.
## Solutions considered [#solutions-considered]
Similarweb considered a number of technology solutions before they selected Firebolt
**Presto**: Similarweb initially considered Presto since they already used it for internal analytics. However, they quickly ruled it out because it couldn't deliver the sub-second latency required for a responsive end-user experience.
**NoSQL Key-Value Databases**: While these databases offer fast document storage, Similarweb rejected this approach because of poor SQL compatibility and lack of support for dynamic grouping operations.
**Auto-Scaled Serverless Compute**: Similarweb tested a more complex approach that would trigger one serverless function for each day in the user's query range. This solution had multiple fatal flaws:
* **Storage Requirements**: The approach required converting their existing ORC format data to JSON, essentially creating a full duplicate of their 1PB dataset to support all possible user segment requests.
* **Performance Issues**: Despite parallelization, some serverless functions consistently ran slower than others, creating unacceptable wait times for the overall query.
* **Limited Functionality**: The solution lacked SQL support for the additional grouping and aggregations that users needed.
## Selection of Firebolt [#selection-of-firebolt]
Firebolt is an analytical database that combines the cost and efficiency benefits of cloud-native architecture with sub-second performance at terabyte scale. It gives engineers the performance, flexibility, and control needed to power production-grade data and AI workloads.
#### Cloud-Native Architecture [#cloud-native-architecture]
Built on AWS, Firebolt uses decoupled storage and compute principles to eliminate traditional challenges around provisioning, scaling, and resource utilization. The platform leverages key AWS services including Amazon S3 for durable object storage and Amazon EKS with EC2 for elastic compute engines. This AWS-native approach ensures analytics solutions align with existing cloud strategy and security standards.
#### The Final Evaluation [#the-final-evaluation]
Similarweb narrowed their choice to BigQuery and Firebolt, then conducted performance benchmarks. Firebolt emerged as the clear winner for several reasons:

**Superior Performance**: Firebolt delivered the best results thanks to performance optimization engineering at every database layer. Importantly, it required no additional pre-processing—raw data could be loaded and immediately queried with sub-second performance.
**Workload Isolation**: The decoupled storage and compute architecture enabled easy workload isolation. Similarweb could isolate their new feature workloads to deliver consistently fast, predictable queries while continuing development on separate compute clusters (called "engines") without affecting production.
**Cost Efficiency**: Firebolt offered the best price-performance ratio and lowest total cost of ownership.
**Rapid Implementation**
Within weeks, Firebolt was fully integrated into production using Airflow and Firebolt's REST APIs. The platform now powers features like analyzing PlayStation 5 traffic patterns on Amazon.com—scanning multiple terabytes of data with dynamic "PS5" filtering while delivering \~1-second UI load times.
The seamless integration with Similarweb's existing AWS infrastructure enabled rapid deployment and immediate value realization.

For organizations looking to modernize their analytics infrastructure, Firebolt runs on AWS and is available through AWS Marketplace, providing a streamlined path to deploy high-performance analytics capabilities within existing AWS environments.
# Why IQVIA relies on Firebolt for Life Science analytics (/customers/iqvia)
IQVIA uses Firebolt to accelerate Life Science analytics, ensuring consistent sub-second query performance regardless of concurrent user load. Jeremy Stroud, Director of IT Architecture at IQVIA, highlights that maintaining fast response times across 100 to 250 simultaneous BI tool users is critical to their operations, making Firebolt a key strategic partner for their analytics infrastructure.
Whether we have 100, 200, or 250 users accessing a BI tool, we need consistent sub-second query
performance. Firebolt is a key partner for us.
# How Primer Accelerates Query Performance with SQL Only (/customers/primer)
Fintech company Primer moved to Firebolt to achieve faster query performance while using only SQL. Aaron Rank, Head of Data & Analytics, explains that as a customer-facing product, Primer cannot tolerate 3–4 second delays when switching between different time period views. With Firebolt, they achieve instant performance that meets their product requirements.
As a customer-facing product, we can't have a 3-4 second delay when trying to move from a 6-month
to a 12-month view. I need that instantly.
# How Sweet Security Delivers Sub-Second Threat Detection Analytics with Firebolt (/customers/sweet)
Sweet Security delivers runtime threat detection analytics to hundreds of enterprise customers, ingesting 1TB of security events daily. Firebolt powers sub-second query latency across 350+ continuously updating tables, with workload isolation that keeps complex multi-join policy queries fast even under continuous ingestion.
# Athena vs Redshift (/comparison/athena-vs-redshift)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Athena | Redshift |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| Separation of storage and compute | Yes, serverless with optional provisioned capacity. Workloads can be isolated through Workgroups and Capacity Reservations | RA3 instances enable separation of compute and storage, but limited workload isolation compared to other platforms |
| Supported cloud infrastructure | AWS only | AWS only |
| Isolated tenancy – option for dedicated resources | • Multi-tenant pooled resources by default • Dedicated compute resources available via Provisioned Capacity • VPC endpoint connections supported | • Isolated tenant & resources • Runs in your VPC |
| Control vs abstraction of compute | • Serverless by default with no infrastructure control • Optional Provisioned Capacity allows dedicated DPU allocation (minimum 24 DPUs) • Two pricing models: on-demand ($5/TB scanned) or provisioned ($0.30/DPU-hour) | • Configurable cluster size • Configurable compute types |
| Self-hosted and hybrid deployment options | No self-hosted options – serverless only | Limited hybrid options with Redshift Serverless |
| ACID Compliance and Transactions | No ACID compliance – eventual consistency model | ACID compliant at table level with some limitations on concurrent operations |
**Athena** is serverless and built on a decoupled storage and compute architecture that queries data directly in S3, without the need to ingest/copy the data. It runs in multi-tenancy with shared resources. Users do not have control over the compute resources Athena chooses to allocate per query from the shared resource pool. For folks requiring additional or dedicated resources, they can reserve dedicated processing capacity in the form of Data Processing Units (DPU), with each DPU providing 4 vCPU and 16 GB RAM. RPU allocation ranges from 24 - 1000 per region.
**Redshift** has the oldest architecture, being the first Cloud DW in the group. Its architecture wasn't designed to separate storage & compute. While it now has RA3 nodes which allow you to scale compute and only cache the data you need locally, all compute still operates together. You cannot separate and isolate different workloads over the same data, which puts it behind other decoupled storage/compute architectures. Redshift runs as an isolated tenant per customer, and unlike other cloud data warehouses, it is deployed in your VPC. Redshift offers a serverless option which is based on an abstracted unit called Redshift Processing Unit (RPU) ranging from 8 to 512 in increments of 8. Each RPU provides 2 vCPU and 16GB RAM. Thus, 8 RPU is equivalent to 16 vCPU / 128GB RAM. The minimum RPU is 8.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Athena | Redshift |
| --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| Elasticity – Scaling for larger data volumes and faster queries | • Fully abstracted on-demand scaling • Provisioned Capacity allows manual scaling of DPUs for predictable performance • Capacity reservations can be adjusted with minimum 1-hour billing periods | Available via Elastic Resize – slow and limited, downtime required |
| Elasticity – Scaling for higher concurrency | • Default limit of 25 concurrent DML queries and 20 DDL queries (adjustable via service quotas) • Provisioned Capacity enables higher concurrency with dedicated DPUs • Query queuing available when capacity is exceeded | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries |
**Athena** is a shared multi-tenant resource, with no guarantees on the amount or availability of the resources allocated for your queries. From a data volume perspective, it can scale to large volumes, but large data volumes can suffer from very long run times and frequent time outs. Query concurrency is maxed at 20. If scalability is a top priority, Athena is probably not the best choice.
**Redshift** is limited in scale because even with RA3, it cannot distribute different workloads across clusters. While it can scale to up to 10 clusters automatically to support query concurrency, it can only handle a maximum of 50 queued queries across all clusters by default.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Athena | Redshift |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | No traditional indexes – relies on partition pruning and data organization in S3. Uses columnar formats and compression for optimization | None |
| Compute tuning | • No compute tuning in on-demand mode • Provisioned Capacity allows DPU allocation control (4 vCPU and 16GB RAM per DPU) • Minimum 24 DPUs with scaling in 4-DPU increments | Choice over number of nodes and their type |
| Storage format | Supports multiple formats: Parquet, ORC, Avro, JSON, CSV, TSV on S3. Native support for open table formats including Apache Iceberg, Apache Hudi, and Delta Lake | Columnar & compressed storage (RA3 nodes) |
| Table-level partition & pruning techniques | • User-defined table-level partitions with Hive-style partitioning • Pruning at partition level • Partition projection for advanced performance optimization • Supports open table formats with built-in partitioning | • No table partitions • User-defined distribution & sort keys are used to optimize for speed |
| Result cache | Query result caching for up to 30 days with configurable retention. Results reuse supported across workgroups | Yes |
| Warm cache (SSD) | No local caching – queries data directly from S3. Relies on S3's performance characteristics and intelligent tiering | Only with RA3 nodes at partition-level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, comprehensive JSON support including Lambda expressions, array functions, and native nested data handling | Limited |
| Vector Search and AI Capabilities | No native AI or vector search capabilities | Limited AI capabilities – primarily through integrations |
| Query Optimizations | • Cost-based optimizer (CBO) in Athena engine v3 • Query result caching (up to 30 days) • Partition projection for advanced optimization • CTAS for precomputed queries • Join reordering and aggregation pushdown • Automatic parallel query execution • Support for columnar formats (Parquet, ORC) • Integration with AWS Glue Data Catalog | • Basic query optimizer • Materialized views • Result caching • ANALYZE for table statistics • Workload management (WLM) • Automated materialized views (AutoMV) • AI-driven scaling (Serverless preview) |
**Athena**, (and Presto) are designed to query data where it is, sacrificing storage-compute optimizations. This makes it very convenient for easy and immediate querying but at the expense of performance. This typically puts Athena behind cloud data warehouses in terms of performance. But Athena still does relatively well in performance benchmarks, especially when external storage is managed by experts. While it supports partitions, there is no support for indexing, and together with the fact that resources are pooled from a shared multi-tenant service, low-latency and consistent performance are not Athena's sweet spot. A cloud data warehouse be more performant better than Athena in most cases.
**Redshift** does provide a result cache for accelerating repetitive query workloads and also has more tuning options than some others. But it does not deliver much faster compute performance than other cloud data warehouses in benchmarks. Sort keys can be used to optimize performance, but their contribution is limited. There is no support for indexes, and low-latency analytics at large data volumes is hard to achieve. Because Redshift decoupling of storage & compute is limited compared to other cloud data warehouses, it doesn't support isolating workloads, which means performance can degrade under pressure and competition for resources
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Athena | Redshift |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • Seconds to minutes response times for interactive dashboards • Performance varies based on data partitioning, file formats, and query optimization • Provisioned Capacity can improve consistency for dashboard workloads • Best suited for analytical dashboards rather than sub-second operational dashboards | • Seconds to tens of seconds load times at 100s of GB scale • Can achieve faster performance with Concurrency Scaling and proper tuning |
| Enterprise BI | • Good integration with AWS ecosystem BI tools (QuickSight, etc.) • Standard SQL compatibility enables most BI tool connections • Cost-effective for variable workloads and ad-hoc analytics • JDBC/ODBC drivers support enterprise BI tools • Limited advanced BI features compared to dedicated data warehouses | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Strong AWS ecosystem integration |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Default concurrency limits (25 DML/20 DDL queries) may require service quota increases • Provisioned Capacity enables higher concurrency with dedicated resources • Seconds-level response times typical • Cost-effective for customer-facing analytics with proper optimization • Best suited for analytical rather than operational workloads • No native AI capabilities | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries • Seconds-level response times typical • Automatic scaling for burst workloads • Limited AI application support |
| Ad hoc | • Purpose-built for ad-hoc analytics on data lakes • Serverless with zero infrastructure management • Direct querying of S3 data without ETL • Cost-effective pay-per-query model ideal for exploratory analysis • Strong support for multiple data formats and federated queries • Apache Spark integration for advanced analytics | • Performance dependent on predefined distribution & sort keys • Elastic Resize enables adding compute resources • Typically subset of data loaded for ad-hoc analysis |
**Athena** is a great choice for Ad-Hoc analytics. You can keep the data where it is, and start querying without worrying about hardware or pretty much anything else, given that Athena is serverless and takes care of everything behind the scenes. However, it is not a great fit when you need consistent and fast query performance, and/or high concurrency. This is why it is typically not the best choice for operational and customer-facing applications. It can be also easily and flexibly used for batch processing, which is often leveraged for ML use cases.
**Redshift** was originally designed to support traditional internal BI reporting and dashboard use cases for analysts. As such, it is typically used as a general-purpose Enterprise data warehouse. With deep integrations into the AWS ecosystem, it can also leverage AWS ML service, making it also useful for ML projects. However, given the coupling of storage & compute, and the difficulty in delivering low-latency analytics at scale, it is less suited for operational use cases and customer-facing use cases like Data Apps. The coupling of storage and compute, together with the need to predefine sort & dist keys for optimal performance, make it challenging to use for Ad-Hoc analytics.
# BigQuery vs Athena (/comparison/bigquery-vs-athena)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | BigQuery | Athena |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Fully serverless with complete separation of compute (Dremel) and storage (Colossus), powered by Jupiter network and Borg orchestration | Yes, serverless with optional provisioned capacity. Workloads can be isolated through Workgroups and Capacity Reservations |
| Supported cloud infrastructure | Google Cloud only | AWS only |
| Isolated tenancy – option for dedicated resources | • Multi-tenant pooled resources • VPC Service Controls provide enhanced security and connectivity isolation to customer VPCs • Cross-region disaster recovery for enterprise workloads | • Multi-tenant pooled resources by default • Dedicated compute resources available via Provisioned Capacity • VPC endpoint connections supported |
| Control vs abstraction of compute | Fully serverless with no control over compute resources – BigQuery automatically allocates computing resources as needed with intelligent workload management and dynamic slot allocation | • Serverless by default with no infrastructure control • Optional Provisioned Capacity allows dedicated DPU allocation (minimum 24 DPUs) • Two pricing models: on-demand ($5/TB scanned) or provisioned ($0.30/DPU-hour) |
| Self-hosted and hybrid deployment options | No self-hosted options – fully managed service only | No self-hosted options – serverless only |
| ACID Compliance and Transactions | Limited ACID support – eventual consistency model with some transactional capabilities | No ACID compliance – eventual consistency model |
**BigQuery** was one of the first decoupled storage and compute architectures. It is a unique piece of engineering and not a typical data warehouse in part because it started as an on-demand serverless query engine. It runs in multi-tenancy with shared resources, allocated as "slots" which represent a virtual CPU that executes SQL. BigQuery determines how many slots a query requires, without the ability of the user to control it. BigQuery can be priced on a $/TB scanned basis or through slot reservations. A slot in BigQuery is logically equivalent to 0.5 vCPU and 0.5GB of RAM. There are multiple models to allocate slots in BigQuery.
**Athena** is serverless and built on a decoupled storage and compute architecture that queries data directly in S3, without the need to ingest/copy the data. It runs in multi-tenancy with shared resources. Users do not have control over the compute resources Athena chooses to allocate per query from the shared resource pool. For folks requiring additional or dedicated resources, they can reserve dedicated processing capacity in the form of Data Processing Units (DPU), with each DPU providing 4 vCPU and 16 GB RAM. RPU allocation ranges from 24 - 1000 per region.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | BigQuery | Athena |
| --------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Fully automated serverless scaling – BigQuery automatically determines resource allocation and scales to petabytes without user intervention. Can dynamically burst beyond baseline slot allocations for performance optimization | • Fully abstracted on-demand scaling • Provisioned Capacity allows manual scaling of DPUs for predictable performance • Capacity reservations can be adjusted with minimum 1-hour billing periods |
| Elasticity – Scaling for higher concurrency | Dynamic concurrency management with query queueing supporting up to 1,000 interactive queries and 20,000 batch queries per project per region. Automatic fair scheduling and slot distribution across workloads | • Default limit of 25 concurrent DML queries and 20 DDL queries (adjustable via service quotas) • Provisioned Capacity enables higher concurrency with dedicated DPUs • Query queuing available when capacity is exceeded |
**BigQuery** scales very well to large data volumes, and automatically assigns more compute resources when needed behind the scenes, in the form of "slots". BigQuery works either in an "on-demand pricing model", where slot assignment is completely in the hands of BigQuery and the state of the shared resource pool, or in "flat-rate pricing model" where slots are reserved in advanced. With reserved slots there is more control over compute resources, thus making scaling more predictable. Concurrency is limited to 100 users by default.
**Athena** is a shared multi-tenant resource, with no guarantees on the amount or availability of the resources allocated for your queries. From a data volume perspective, it can scale to large volumes, but large data volumes can suffer from very long run times and frequent time outs. Query concurrency is maxed at 20. If scalability is a top priority, Athena is probably not the best choice.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | BigQuery | Athena |
| ------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | Search indexes (GA) for efficient text search optimization on STRING, JSON, and array columns. Support for LOG\_ANALYZER, NO\_OP\_ANALYZER, and PATTERN\_ANALYZER with column-level granularity for improved query performance and cost efficiency | No traditional indexes – relies on partition pruning and data organization in S3. Uses columnar formats and compression for optimization |
| Compute tuning | Serverless architecture with automatic resource optimization – no manual tuning required. Intelligent workload management with AI-powered resource allocation and dynamic slot distribution | • No compute tuning in on-demand mode • Provisioned Capacity allows DPU allocation control (4 vCPU and 16GB RAM per DPU) • Minimum 24 DPUs with scaling in 4-DPU increments |
| Storage format | Columnar & compressed storage (Capacitor format) with support for open table formats including Apache Iceberg, Delta Lake, and Hudi. Intelligent tiering with automatic long-term storage cost reduction after 90 days | Supports multiple formats: Parquet, ORC, Avro, JSON, CSV, TSV on S3. Native support for open table formats including Apache Iceberg, Apache Hudi, and Delta Lake |
| Table-level partition & pruning techniques | • Automatic table organization with intelligent micro-partitioning • Clustering keys for data organization • Automatic partition pruning optimization • Supports time-based and custom partitioning strategies | • User-defined table-level partitions with Hive-style partitioning • Pruning at partition level • Partition projection for advanced performance optimization • Supports open table formats with built-in partitioning |
| Result cache | Yes, with cross-user result caching and intelligent cache management for up to 24 hours | Query result caching for up to 30 days with configurable retention. Results reuse supported across workgroups |
| Warm cache (SSD) | BI Engine provides in-memory caching and acceleration for frequently accessed data and dashboards | No local caching – queries data directly from S3. Relies on S3's performance characteristics and intelligent tiering |
| Support for semi-structured data & JSON functions within SQL | Yes, including advanced JSON functions, Lambda expressions, and native support for nested and repeated fields | Yes, comprehensive JSON support including Lambda expressions, array functions, and native nested data handling |
| Vector Search and AI Capabilities | • BigQuery ML for in-database machine learning • Vertex AI integration • Natural language querying with Gemini AI • Limited vector search capabilities | No native AI or vector search capabilities |
| Query Optimizations | • Advanced query optimizer with Dremel engine • Search indexes with column-level granularity • Materialized views with smart refresh and automatic query rewriting • BI Engine in-memory acceleration • Gemini AI-powered query optimization and natural language querying • Cross-user result caching • Automatic partitioning and clustering optimization • Cost-based optimization with intelligent workload management | • Cost-based optimizer (CBO) in Athena engine v3 • Query result caching (up to 30 days) • Partition projection for advanced optimization • CTAS for precomputed queries • Join reordering and aggregation pushdown • Automatic parallel query execution • Support for columnar formats (Parquet, ORC) • Integration with AWS Glue Data Catalog |
**BigQuery** lines up in benchmarks in the same ballpark as other cloud data warehouses but does come in consistently last in most queries. Beyond implementing according to best practices, there is little you can do to accelerate BigQuery performance, as it determines the amount of resources (slots) the query needs for you. BigQuery can be used together with "BigQuery BI Engine" for lower latency analytics. However, BI Engine is limited in terms of scale because it runs in memory. Its maximum capacity is 100GB.
**Athena** (and Presto) are designed to query data where it is, sacrificing storage-compute optimizations. This makes it very convenient for easy and immediate querying but at the expense of performance. This typically puts Athena behind cloud data warehouses in terms of performance. But Athena still does relatively well in performance benchmarks, especially when external storage is managed by experts. While it supports partitions, there is no support for indexing, and together with the fact that resources are pooled from a shared multi-tenant service, low-latency and consistent performance are not Athena's sweet spot. A cloud data warehouse be more performant better than Athena in most cases.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | BigQuery | Athena |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • Sub-second to seconds response times at TB+ scale with BI Engine acceleration • Search indexes and materialized views provide significant performance improvements for dashboard queries • Intelligent caching reduces query costs for repeated dashboard access • Dynamic concurrency management supports high user loads | • Seconds to minutes response times for interactive dashboards • Performance varies based on data partitioning, file formats, and query optimization • Provisioned Capacity can improve consistency for dashboard workloads • Best suited for analytical dashboards rather than sub-second operational dashboards |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set with native Google Cloud ecosystem integration • Strong integration with Looker, Looker Studio, and major BI tools • Gemini AI integration for natural language insights • Cross-cloud analytics capabilities • Advanced governance with Dataplex Universal Catalog | • Good integration with AWS ecosystem BI tools (QuickSight, etc.) • Standard SQL compatibility enables most BI tool connections • Cost-effective for variable workloads and ad-hoc analytics • JDBC/ODBC drivers support enterprise BI tools • Limited advanced BI features compared to dedicated data warehouses |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Dynamic concurrency supporting 1,000+ interactive queries with intelligent queuing and fair scheduling • Sub-second response times with BI Engine acceleration and search indexes • Serverless architecture eliminates infrastructure management overhead • Advanced caching and materialized views optimize repeated queries • AI-powered optimization for customer-facing applications • BigQuery ML for AI workloads | • Default concurrency limits (25 DML/20 DDL queries) may require service quota increases • Provisioned Capacity enables higher concurrency with dedicated resources • Seconds-level response times typical • Cost-effective for customer-facing analytics with proper optimization • Best suited for analytical rather than operational workloads • No native AI capabilities |
| Ad hoc | • Excellent for ad-hoc analytics with serverless architecture requiring zero infrastructure management • Intelligent query optimization handles unpredictable workloads automatically • Gemini AI integration enables natural language querying for business users • Advanced JSON support and schema inference enable flexible data exploration • Cross-cloud analytics capabilities for federated queries | • Purpose-built for ad-hoc analytics on data lakes • Serverless with zero infrastructure management • Direct querying of S3 data without ETL • Cost-effective pay-per-query model ideal for exploratory analysis • Strong support for multiple data formats and federated queries • Apache Spark integration for advanced analytics |
**BigQuery** is a mature general-purpose data warehouse, which lends itself well to internal BI & reporting. The fact that it's serverless in nature and tightly integrated with GCP, makes it very convenient for Ad-Hoc analytics and ML use cases on GCP. On the other hand, because BigQuery makes resource allocation decisions for you, it is not always the best fit for operational use cases and Data Apps where performance needs to be consistent and predictable.
**Athena** is a great choice for Ad-Hoc analytics. You can keep the data where it is, and start querying without worrying about hardware or pretty much anything else, given that Athena is serverless and takes care of everything behind the scenes. However, it is not a great fit when you need consistent and fast query performance, and/or high concurrency. This is why it is typically not the best choice for operational and customer-facing applications. It can be also easily and flexibly used for batch processing, which is often leveraged for ML use cases.
# BigQuery vs ClickHouse (2025) (/comparison/bigquery-vs-clickhouse)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | BigQuery | ClickHouse |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Fully serverless with complete separation of compute (Dremel) and storage (Colossus), powered by Jupiter network and Borg orchestration | Yes – SharedMergeTree engine in ClickHouse Cloud enables full separation of storage and compute, with compute-compute separation through Warehouses feature (introduced 2025) allowing multiple isolated compute services sharing the same data |
| Supported cloud infrastructure | Google Cloud only | AWS, GCP, Azure, cloud service and on-premises |
| Isolated tenancy – option for dedicated resources | • Multi-tenant pooled resources • VPC Service Controls provide enhanced security and connectivity isolation to customer VPCs • Cross-region disaster recovery for enterprise workloads | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client in cloud |
| Control vs abstraction of compute | Fully serverless with no control over compute resources – BigQuery automatically allocates computing resources as needed with intelligent workload management and dynamic slot allocation | Configurable cluster size and compute types in ClickHouse Cloud with granular control over nodes (1-128 nodes) and node characteristics. Warehouses feature enables multiple isolated read-only compute environments. |
| Self-hosted and hybrid deployment options | No self-hosted options – fully managed service only | Self-managed deployments available with full control over infrastructure |
| ACID Compliance and Transactions | Limited ACID support – eventual consistency model with some transactional capabilities | Limited ACID compliance with MergeTree engine family. |
**BigQuery** was one of the first decoupled storage and compute architectures. It is a unique piece of engineering and not a typical data warehouse in part because it started as an on-demand serverless query engine. It runs in multi-tenancy with shared resources, allocated as "slots" which represent a virtual CPU that executes SQL. BigQuery determines how many slots a query requires, without the ability of the user to control it. BigQuery can be priced on a $/TB scanned basis or through slot reservations. A slot in BigQuery is logically equivalent to 0.5 vCPU and 0.5GB of RAM. There are multiple models to allocate slots in BigQuery.
**ClickHouse** was originally developed at Yandex, the Russian search engine, as an OLAP engine for low latency analytics. It was built as an on-premise solution with coupled storage & compute, and a large variety of tuning options in the form of indexes and and merge trees. ClickHouse's architecture is famous for its focus on performance and low-latency queries. The tradeoff is that it is considered very difficult to work with. SQL support is very limited, and tuning/running it requires significant engineering resources.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | BigQuery | ClickHouse |
| --------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Fully automated serverless scaling – BigQuery automatically determines resource allocation and scales to petabytes without user intervention. Can dynamically burst beyond baseline slot allocations for performance optimization | Automatic horizontal and vertical scaling in ClickHouse Cloud with SharedMergeTree architecture. Manual scaling for self-managed deployments with cluster rebalancing capabilities |
| Elasticity – Scaling for higher concurrency | Dynamic concurrency management with query queueing supporting up to 1,000 interactive queries and 20,000 batch queries per project per region. Automatic fair scheduling and slot distribution across workloads | Supports high concurrency with proper resource allocation and configuration. Vertical auto-scaling and horizontal manual scaling. Additional warehouses can idle to zero billing. Primary service always on in multi-warehouse configurations. |
**BigQuery** scales very well to large data volumes, and automatically assigns more compute resources when needed behind the scenes, in the form of "slots". BigQuery works either in an "on-demand pricing model", where slot assignment is completely in the hands of BigQuery and the state of the shared resource pool, or in "flat-rate pricing model" where slots are reserved in advanced. With reserved slots there is more control over compute resources, thus making scaling more predictable. Concurrency is limited to 100 users by default.
**ClickHouse** doesn't offer any dedicated scaling features or mechanisms. While it can deliver linearly scalable performance for some types of queries, scaling itself has to be done manually. Hardware is self-managed in ClickHouse. This means that to scale you would have to provision a cluster and migrate.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today. While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | BigQuery | ClickHouse |
| ------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | Search indexes (GA) for efficient text search optimization on STRING, JSON, and array columns. Support for LOG\_ANALYZER, NO\_OP\_ANALYZER, and PATTERN\_ANALYZER with column-level granularity for improved query performance and cost efficiency | • Primary indexes • Skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • MergeTree indexes • Incremental Materialized views |
| Compute tuning | Serverless architecture with automatic resource optimization – no manual tuning required. Intelligent workload management with AI-powered resource allocation and dynamic slot distribution | Configurable compute resources in cloud offering |
| Storage format | Columnar & compressed storage (Capacitor format) with support for open table formats including Apache Iceberg, Delta Lake, and Hudi. Intelligent tiering with automatic long-term storage cost reduction after 90 days | Columnar, supports sorted, compressed, encoded & sparsely indexed files with native Apache Iceberg support. |
| Table-level partition & pruning techniques | • Automatic table organization with intelligent micro-partitioning • Clustering keys for data organization • Automatic partition pruning optimization • Supports time-based and custom partitioning strategies | Partitioning by date/time and custom partitions with MergeTree indexes. |
| Result cache | Yes, with cross-user result caching and intelligent cache management for up to 24 hours | Yes, results cache with TTL and query condition cache. |
| Warm cache (SSD) | BI Engine provides in-memory caching and acceleration for frequently accessed data and dashboards | Yes, at indexed data-range level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, including advanced JSON functions, Lambda expressions, and native support for nested and repeated fields | Yes, including Lambda expressions and native JSON data type (GA in v25.3) |
| Vector Search and AI Capabilities | • BigQuery ML for in-database machine learning • Vertex AI integration • Natural language querying with Gemini AI • Limited vector search capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference |
| Query Optimizations | • Advanced query optimizer with Dremel engine • Search indexes with column-level granularity • Materialized views with smart refresh and automatic query rewriting • BI Engine in-memory acceleration • Gemini AI-powered query optimization and natural language querying • Cross-user result caching • Automatic partitioning and clustering optimization • Cost-based optimization with intelligent workload management | • Primary indexes (ORDER BY) • Data skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • Materialized views • Projections • PREWHERE optimization • Query analysis tools • Automatic global join reordering (v25.9) • Enhanced JSON query optimization • Streaming secondary indices |
**BigQuery** lines up in benchmarks in the same ballpark as other cloud data warehouses but does come in consistently last in most queries. Beyond implementing according to best practices, there is little you can do to accelerate BigQuery performance, as it determines the amount of resources (slots) the query needs for you. BigQuery can be used together with "BigQuery BI Engine" for lower latency analytics. However, BI Engine is limited in terms of scale because it runs in memory. Its maximum capacity is 100GB.
**ClickHouse** is famous for being one of the fastest local runtimes ever built for OLAP workloads. Its columnar storage, compression and indexing capabilities make it a consistent leader in benchmarks. Its lack of support for standard SQL and lack of query optimizer means that it's less suitable for traditional BI workloads, and more suitable for engineering managed workloads. While fast, it requires a lot of tuning and optimization.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | BigQuery | ClickHouse |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second to seconds response times at TB+ scale with BI Engine acceleration • Search indexes and materialized views provide significant performance improvements for dashboard queries • Intelligent caching reduces query costs for repeated dashboard access • Dynamic concurrency management supports high user loads | • Sub-second load times at TB+ scale with proper indexing • ClickHouse Cloud reduces engineering overhead with managed service • Proven low-latency performance (120ms at 2500 QPS in benchmarks) • Purpose-built for low-latency OLAP and real-time analytics |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set with native Google Cloud ecosystem integration • Strong integration with Looker, Looker Studio, and major BI tools • Gemini AI integration for natural language insights • Cross-cloud analytics capabilities • Advanced governance with Dataplex Universal Catalog | • Growing ecosystem with 50+ integrations including major BI tools • Native MySQL protocol support enables broad BI tool compatibility • Strong SQL compliance with PostgreSQL compatibility • Best suited for modern analytical workloads and engineering-managed use cases |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Dynamic concurrency supporting 1,000+ interactive queries with intelligent queuing and fair scheduling • Sub-second response times with BI Engine acceleration and search indexes • Serverless architecture eliminates infrastructure management overhead • Advanced caching and materialized views optimize repeated queries • AI-powered optimization for customer-facing applications • BigQuery ML for AI workloads | • Sub-second response times at TB+ scale • Supports 1000 concurrent users per replica • Strong price-performance on customer-facing applications • Native vector search and embeddings |
| Ad hoc | • Excellent for ad-hoc analytics with serverless architecture requiring zero infrastructure management • Intelligent query optimization handles unpredictable workloads automatically • Gemini AI integration enables natural language querying for business users • Advanced JSON support and schema inference enable flexible data exploration • Cross-cloud analytics capabilities for federated queries | • Good for ad-hoc queries with ClickHouse Cloud’s separated storage/compute architecture • Join optimizations enable more query complexity • Strong sampling capabilities (TABLESAMPLE) for exploratory analysis • Resource management through user quotas prevents query interference • Materialized views offer performance improvements for common aggregation patterns, ad-hoc users specify directly in SQL |
**BigQuery** is a mature general-purpose data warehouse, which lends itself well to internal BI & reporting. The fact that it's serverless in nature and tightly integrated with GCP, makes it very convenient for Ad-Hoc analytics and ML use cases on GCP. On the other hand, because BigQuery makes resource allocation decisions for you, it is not always the best fit for operational use cases and Data Apps where performance needs to be consistent and predictable.
**ClickHouse** was not designed to be a data warehouse, but rather a low-latency query execution runtime. Managing it typically requires significant engineering overhead. Hence, it's a good fit for engineering managed operational use cases and customer-facing data apps, where low latency matters. It is not a good fit for a general purpose data warehouse, nor for Ad-Hoc analytics or ELT.
# BigQuery vs Databricks (2025) (/comparison/bigquery-vs-databricks)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | BigQuery | Databricks |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Fully serverless with complete separation of compute (Dremel) and storage (Colossus), powered by Jupiter network and Borg orchestration | Yes |
| Supported cloud infrastructure | Google Cloud only | AWS, Azure, GCP. Marketplaces and BYOC |
| Isolated tenancy – option for dedicated resources | • Multi-tenant pooled resources • VPC Service Controls provide enhanced security and connectivity isolation to customer VPCs • Cross-region disaster recovery for enterprise workloads | • Control plane in Databricks account • Data plane in customer VPC (optional) • Storage in customer VPC • Serverless SQL runs in Databricks account with private connectivity |
| Control vs abstraction of compute | Fully serverless with no control over compute resources – BigQuery automatically allocates computing resources as needed with intelligent workload management and dynamic slot allocation | • Configurable clusters and instance types • Serverless SQL warehouses (GA 2025) run in Databricks account with private connectivity, no public IPs • Pro/Classic warehouses run in customer VPC |
| Self-hosted and hybrid deployment options | No self-hosted options – fully managed service only | • Databricks on customer cloud accounts • Unity Catalog for hybrid governance |
| ACID Compliance and Transactions | Limited ACID support – eventual consistency model with some transactional capabilities | • ACID transactions with Delta Lake • Time travel and versioning • Concurrent read/write operations |
**BigQuery** was one of the first decoupled storage and compute architectures. It is a unique piece of engineering and not a typical data warehouse in part because it started as an on-demand serverless query engine. It runs in multi-tenancy with shared resources, allocated as “slots” which represent a virtual CPU that executes SQL. BigQuery determines how many slots a query requires, without the ability of the user to control it. BigQuery can be priced on a $/TB scanned basis or through slot reservations. A slot in BigQuery is logically equivalent to 0.5 vCPU and 0.5GB of RAM. There are multiple models to allocate slots in BigQuery.
**Databricks** was built by the founders of Spark as an analytics platform to support machine learning use cases. It leverages the Spark framework to process data residing in a data lake and is supported on AWS, GCP and Azure. Databricks coined the marketing term “Lakehouse '' architecture to illustrate the unification of data lake and data warehouse use cases. Customers still manage Spark clusters that process data residing in a Delta lake. Conversion of data to Delta Lake format is required to leverage the functionality of Delta Lake. Databricks Sql is a relatively new addition to simplify access to data stored in a data lake.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | BigQuery | Databricks |
| --------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Fully automated serverless scaling – BigQuery automatically determines resource allocation and scales to petabytes without user intervention. Can dynamically burst beyond baseline slot allocations for performance optimization | Autoscaling clusters based on workload demand. Serverless SQL warehouses provide near-instant scaling (2-6 seconds startup) |
| Elasticity – Scaling for higher concurrency | Dynamic concurrency management with query queueing supporting up to 1,000 interactive queries and 20,000 batch queries per project per region. Automatic fair scheduling and slot distribution across workloads | • 10 concurrent queries per cluster limit • Scales up to 40 clusters per warehouse (400 total concurrent queries) • Serverless SQL warehouses provide near-instant autoscaling • Pro/Classic warehouses take several minutes to provision new clusters • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on complexity |
**BigQuery** scales very well to large data volumes, and automatically assigns more compute resources when needed behind the scenes, in the form of “slots”. BigQuery works either in an “on-demand pricing model”, where slot assignment is completely in the hands of BigQuery and the state of the shared resource pool, or in “flat-rate pricing model” where slots are reserved in advance. With reserved slots there is more control over compute resources, thus making scaling more predictable. Concurrency is limited to 100 users by default.
**Databricks** allow for autoscaling of clusters based on utilization. Additionally, increasing concurrency associated with a sql endpoint can be accomplished through the addition of clusters. Query concurrency per cluster is maxed at 10. However, scaling with additional clusters for concurrency is possible. Databricks provides a choice of instance types.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | BigQuery | Databricks |
| ------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | Search indexes (GA) for efficient text search optimization on STRING, JSON, and array columns. Support for LOG\_ANALYZER, NO\_OP\_ANALYZER, and PATTERN\_ANALYZER with column-level granularity for improved query performance and cost efficiency | None |
| Compute tuning | Serverless architecture with automatic resource optimization – no manual tuning required. Intelligent workload management with AI-powered resource allocation and dynamic slot distribution | Choice of cluster type, node types including SSD-optimized instances. Serverless provides automatic resource allocation with Intelligent Workload Management (IWM) |
| Storage format | Columnar & compressed storage (Capacitor format) with support for open table formats including Apache Iceberg, Delta Lake, and Hudi. Intelligent tiering with automatic long-term storage cost reduction after 90 days | • Delta Lake format with Liquid Clustering (February 2025 – replaces Z-ordering and traditional partitioning) • Cannot use Liquid Clustering alongside Z-ordering on same table • Allows for sorted data in Delta Lake • Requires Optimize to maintain ordering |
| Table-level partition & pruning techniques | • Automatic table organization with intelligent micro-partitioning • Clustering keys for data organization • Automatic partition pruning optimization • Supports time-based and custom partitioning strategies | • Table level partitioning • Liquid Clustering for improved query performance and reduced data skew (February 2025) • Z-ordering (legacy, replaced by Liquid Clustering) • Periodic optimization of storage required |
| Result cache | Yes, with cross-user result caching and intelligent cache management for up to 24 hours | Multi-layered caching: local in-memory cache per cluster plus remote result cache (serverless only) that persists across all warehouses in workspace |
| Warm cache (SSD) | BI Engine provides in-memory caching and acceleration for frequently accessed data and dashboards | Yes. Delta cache for data read by queries at file level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, including advanced JSON functions, Lambda expressions, and native support for nested and repeated fields | Yes |
| Vector Search and AI Capabilities | • BigQuery ML for in-database machine learning • Vertex AI integration • Natural language querying with Gemini AI • Limited vector search capabilities | • MLflow integration and Databricks ML platform • Native vector search in Delta Lake (Vector Search) • AI and ML workloads optimized |
| Query Optimizations | • Advanced query optimizer with Dremel engine • Search indexes with column-level granularity • Materialized views with smart refresh and automatic query rewriting • BI Engine in-memory acceleration • Gemini AI-powered query optimization and natural language querying • Cross-user result caching • Automatic partitioning and clustering optimization • Cost-based optimization with intelligent workload management | • Photon engine (C++ vectorized engine providing 3-8x average speedups, maximum speedups over 10x) • Automated stats collection (January 2025) enables cost-based optimization • Predictive I/O for faster point lookups and data updates • Liquid Clustering (February 2025) • Intelligent Workload Management (IWM) with AI-powered resource allocation • Delta cache • Materialized views support |
**BigQuery** lines up in benchmarks in the same ballpark as other cloud data warehouses but does come in consistently last in most queries. Beyond implementing according to best practices, there is little you can do to accelerate BigQuery performance, as it determines the amount of resources (slots) the query needs for you. BigQuery can be used together with the “BigQuery BI Engine” for lower latency analytics. However, BI Engine is limited in terms of scale because it runs in memory. Its maximum capacity is 100GB.
**Databricks** is designed to leverage the Spark framework for processing large volumes of data. It leverages compressed Parquet files in a Delta Lake. To reduce the amount of data processed, it uses data pruning on partitions and Parquet file metadata. Databricks does not provide any indexes.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | BigQuery | Databricks |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second to seconds response times at TB+ scale with BI Engine acceleration • Search indexes and materialized views provide significant performance improvements for dashboard queries • Intelligent caching reduces query costs for repeated dashboard access • Dynamic concurrency management supports high user loads | • Sub-second to seconds load times at TB+ scale • Enhanced by Photon engine (3-8x average speedups) and Delta cache • Serverless SQL warehouses provide rapid startup (2-6 seconds) • Performance depends on cluster configuration |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set with native Google Cloud ecosystem integration • Strong integration with Looker, Looker Studio, and major BI tools • Gemini AI integration for natural language insights • Cross-cloud analytics capabilities • Advanced governance with Dataplex Universal Catalog | • Strong for data science and ML workloads • Unified analytics platform approach • Growing traditional BI integrations • Serverless SQL warehouses improve accessibility • Delta sharing capabilities |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Dynamic concurrency supporting 1,000+ interactive queries with intelligent queuing and fair scheduling • Sub-second response times with BI Engine acceleration and search indexes • Serverless architecture eliminates infrastructure management overhead • Advanced caching and materialized views optimize repeated queries • AI-powered optimization for customer-facing applications • BigQuery ML for AI workloads | • 10 concurrent queries per cluster, scaling to 400 total concurrent queries per warehouse • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on workload complexity • Serverless provides near-instant autoscaling • Photon engine delivers 3-8x performance improvements • Strong ML and AI platform integration |
| Ad hoc | • Excellent for ad-hoc analytics with serverless architecture requiring zero infrastructure management • Intelligent query optimization handles unpredictable workloads automatically • Gemini AI integration enables natural language querying for business users • Advanced JSON support and schema inference enable flexible data exploration • Cross-cloud analytics capabilities for federated queries | • Excellent for ad-hoc with decoupled storage/compute • Serverless SQL warehouses provide instant provisioning • Intelligent Workload Management handles unpredictable workloads automatically • Strong for exploratory data analysis and ML workloads • Automated stats collection improves query planning |
**BigQuery** is a mature general-purpose data warehouse, which lends itself well to internal BI & reporting. The fact that it’s serverless in nature and tightly integrated with GCP, makes it very convenient for Ad-Hoc analytics and ML use cases on GCP. On the other hand, because BigQuery makes resource allocation decisions for you, it is not always the best fit for operational use cases and Data Apps where performance needs to be consistent and predictable.
**Databricks** is a mature Spark based platform proven for processing streaming data. It is widely used for Machine Learning use cases by data scientists through the use of integrated notebooks. From a low latency query perspective, while it offers features like Delta Cache, it does not provide specialized indexes that can deliver low latency queries.
# BigQuery vs Druid (2025) (/comparison/bigquery-vs-druid)
The H2 headings, prose, and tables converted to clean MDX body:
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | BigQuery | Druid |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Fully serverless with complete separation of compute (Dremel) and storage (Colossus), powered by Jupiter network and Borg orchestration | No |
| Supported cloud infrastructure | Google Cloud only | Can be installed anywhere |
| Isolated tenancy – option for dedicated resources | • Multi-tenant pooled resources • VPC Service Controls provide enhanced security and connectivity isolation to customer VPCs • Cross-region disaster recovery for enterprise workloads | Single tenant |
| Control vs abstraction of compute | Fully serverless with no control over compute resources – BigQuery automatically allocates computing resources as needed with intelligent workload management and dynamic slot allocation | • Complex configuration of compute tier with multiple role-specific nodes • Configurable node count • Configurable compute types (virtual machines or kubernetes) |
| Self-hosted and hybrid deployment options | No self-hosted options – fully managed service only | Self-managed deployment required |
| ACID Compliance and Transactions | Limited ACID support – eventual consistency model with some transactional capabilities | Limited ACID support with eventual consistency |
**BigQuery** was one of the first decoupled storage and compute architectures. It is a unique piece of engineering and not a typical data warehouse in part because it started as an on-demand serverless query engine. It runs in multi-tenancy with shared resources, allocated as “slots” which represent a virtual CPU that executes SQL. BigQuery determines how many slots a query requires, without the ability of the user to control it. BigQuery can be priced on a $/TB scanned basis or through slot reservations. A slot in BigQuery is logically equivalent to 0.5 vCPU and 0.5GB of RAM. There are multiple models to allocate slots in BigQuery.
**Druid** is an OLAP engine designed to provide fast real time analytics. Druid adopts a clustered architecture with servers that host various role specific processes. These processes address real time and batch ingestion, indexing, querying of historical and real time data. Apache Druid can be deployed as a virtual machine or a Kubernetes based cluster. Druid does not support a decoupled compute & storage architecture. Deep storage in the form of object storage is used to replicate data to.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | BigQuery | Druid |
| --------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Fully automated serverless scaling – BigQuery automatically determines resource allocation and scales to petabytes without user intervention. Can dynamically burst beyond baseline slot allocations for performance optimization | Scale-up of nodes requires careful planning and downtime. Addition of new nodes for scale-out is possible |
| Elasticity – Scaling for higher concurrency | Dynamic concurrency management with query queueing supporting up to 1,000 interactive queries and 20,000 batch queries per project per region. Automatic fair scheduling and slot distribution across workloads | Supports 100s to 100,000s queries per second (1000+ QPS) with proper configuration and scaling |
**BigQuery** scales very well to large data volumes, and automatically assigns more compute resources when needed behind the scenes, in the form of “slots”. BigQuery works either in an “on-demand pricing model”, where slot assignment is completely in the hands of BigQuery and the state of the shared resource pool, or in “flat-rate pricing model” where slots are reserved in advance. With reserved slots there is more control over compute resources, thus making scaling more predictable. Concurrency is limited to 100 users by default.
**Druid** provides the ability to handle fast ingest and high concurrency. Custom sizing and cluster tuning are required to balance the compute, memory, storage needs of each process within Druid and to provide high concurrency. Druid clusters can be grown by adding nodes with automatic rebalancing of storage segments assigned to nodes. Self hosted Druid on Kubernetes is an option that users leverage to simplify scaling. Additionally, Cloud based managed Druid offerings are being rolled out. However, these managed offerings are limited in scale and scaling is not granular.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | BigQuery | Druid |
| ------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| Indexes | Search indexes (GA) for efficient text search optimization on STRING, JSON, and array columns. Support for LOG\_ANALYZER, NO\_OP\_ANALYZER, and PATTERN\_ANALYZER with column-level granularity for improved query performance and cost efficiency | Compressed bitmap indexes for data access and roll-ups to manage aggregations |
| Compute tuning | Serverless architecture with automatic resource optimization – no manual tuning required. Intelligent workload management with AI-powered resource allocation and dynamic slot distribution | On-premises, self-managed hardware. Druid requires infrastructure management and leverages commonly available instance types |
| Storage format | Columnar & compressed storage (Capacitor format) with support for open table formats including Apache Iceberg, Delta Lake, and Hudi. Intelligent tiering with automatic long-term storage cost reduction after 90 days | Columnar storage format with time-based sorting |
| Table-level partition & pruning techniques | • Automatic table organization with intelligent micro-partitioning • Clustering keys for data organization • Automatic partition pruning optimization • Supports time-based and custom partitioning strategies | Restrictive time-based partitioning. Can partition based on other secondary columns |
| Result cache | Yes, with cross-user result caching and intelligent cache management for up to 24 hours | Ability to support caching on broker (set to off by default) |
| Warm cache (SSD) | BI Engine provides in-memory caching and acceleration for frequently accessed data and dashboards | Yes, at much larger segment level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, including advanced JSON functions, Lambda expressions, and native support for nested and repeated fields | Recommend flattening JSON or translate to array prior to loading. No support for JSON parsing at query runtime |
| Vector Search and AI Capabilities | • BigQuery ML for in-database machine learning • Vertex AI integration • Natural language querying with Gemini AI • Limited vector search capabilities | No native AI or vector search capabilities |
| Query Optimizations | • Advanced query optimizer with Dremel engine • Search indexes with column-level granularity • Materialized views with smart refresh and automatic query rewriting • BI Engine in-memory acceleration • Gemini AI-powered query optimization and natural language querying • Cross-user result caching • Automatic partitioning and clustering optimization • Cost-based optimization with intelligent workload management | • Compressed bitmap indexes • Roll-up aggregations • Time-based optimization • Query optimization requires manual tuning |
**BigQuery** lines up in benchmarks in the same ballpark as other cloud data warehouses but does come in consistently last in most queries. Beyond implementing according to best practices, there is little you can do to accelerate BigQuery performance, as it determines the amount of resources (slots) the query needs for you. BigQuery can be used together with the “BigQuery BI Engine” for lower latency analytics. However, BI Engine is limited in terms of scale because it runs in memory. Its maximum capacity is 100GB.
**Druid** provides high performance through columnar storage format, parallel processing, bitmap indexes and roll-ups. Druid, however, recommends a denormalized data model for performance needs. Join operations in Druid are a relatively new feature with various limitations, especially if there is a need to join large datasets.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | BigQuery | Druid |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • Sub-second to seconds response times at TB+ scale with BI Engine acceleration • Search indexes and materialized views provide significant performance improvements for dashboard queries • Intelligent caching reduces query costs for repeated dashboard access • Dynamic concurrency management supports high user loads | • Sub-second load times optimized for time-series and real-time analytics • Built for high-concurrency interactive dashboards • Requires denormalized data model |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set with native Google Cloud ecosystem integration • Strong integration with Looker, Looker Studio, and major BI tools • Gemini AI integration for natural language insights • Cross-cloud analytics capabilities • Advanced governance with Dataplex Universal Catalog | • Limited integrations with traditional Enterprise BI tools • Strong for real-time operational dashboards • Requires specialized visualization tools |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Dynamic concurrency supporting 1,000+ interactive queries with intelligent queuing and fair scheduling • Sub-second response times with BI Engine acceleration and search indexes • Serverless architecture eliminates infrastructure management overhead • Advanced caching and materialized views optimize repeated queries • AI-powered optimization for customer-facing applications • BigQuery ML for AI workloads | • Built for high concurrency (1000+ QPS) with distributed architecture • Sub-second response times for time-series data • Optimized for real-time operational applications • No AI capabilities |
| Ad hoc | • Excellent for ad-hoc analytics with serverless architecture requiring zero infrastructure management • Intelligent query optimization handles unpredictable workloads automatically • Gemini AI integration enables natural language querying for business users • Advanced JSON support and schema inference enable flexible data exploration • Cross-cloud analytics capabilities for federated queries | • Not optimized for ad-hoc queries • Requires predefined roll-ups and data modeling • Limited flexibility for exploratory analysis |
**BigQuery** is a mature general-purpose data warehouse, which lends itself well to internal BI & reporting. The fact that it’s serverless in nature and tightly integrated with GCP, makes it very convenient for Ad-Hoc analytics and ML use cases on GCP. On the other hand, because BigQuery makes resource allocation decisions for you, it is not always the best fit for operational use cases and Data Apps where performance needs to be consistent and predictable.
**Druid** is designed as an OLAP engine to provide fast access to aggregations that are run against large volumes of data. Druid is typically used for customer facing analytics and streaming data processing. Druid is used as an add-on with other data warehousing products that are efficient at scaling, joining, and filtering large volumes of data. It is not a suitable option for data warehouse replacement.
# BigQuery vs Redshift (/comparison/bigquery-vs-redshift)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | BigQuery | Redshift |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------ |
| Separation of storage and compute | Fully serverless with complete separation of compute (Dremel) and storage (Colossus), powered by Jupiter network and Borg orchestration | RA3 instances enable separation of compute and storage, but limited workload isolation compared to other platforms |
| Supported cloud infrastructure | Google Cloud only | AWS only |
| Isolated tenancy – option for dedicated resources | • Multi-tenant pooled resources • VPC Service Controls provide enhanced security and connectivity isolation to customer VPCs • Cross-region disaster recovery for enterprise workloads | • Isolated tenant & resources • Runs in your VPC |
| Control vs abstraction of compute | Fully serverless with no control over compute resources – BigQuery automatically allocates computing resources as needed with intelligent workload management and dynamic slot allocation | • Configurable cluster size • Configurable compute types |
| Self-hosted and hybrid deployment options | No self-hosted options – fully managed service only | Limited hybrid options with Redshift Serverless |
| ACID Compliance and Transactions | Limited ACID support – eventual consistency model with some transactional capabilities | ACID compliant at table level with some limitations on concurrent operations |
**BigQuery** was one of the first decoupled storage and compute architectures. It is a unique piece of engineering and not a typical data warehouse in part because it started as an on-demand serverless query engine. It runs in multi-tenancy with shared resources, allocated as "slots" which represent a virtual CPU that executes SQL. BigQuery determines how many slots a query requires, without the ability of the user to control it. BigQuery can be priced on a $/TB scanned basis or through slot reservations. A slot in BigQuery is logically equivalent to 0.5 vCPU and 0.5GB of RAM. There are multiple models to allocate slots in BigQuery.
**Redshift** has the oldest architecture, being the first Cloud DW in the group. Its architecture wasn't designed to separate storage & compute. While it now has RA3 nodes which allow you to scale compute and only cache the data you need locally, all compute still operates together. You cannot separate and isolate different workloads over the same data, which puts it behind other decoupled storage/compute architectures. Redshift runs as an isolated tenant per customer, and unlike other cloud data warehouses, it is deployed in your VPC. Redshift offers a serverless option which is based on an abstracted unit called Redshift Processing Unit (RPU) ranging from 8 to 512 in increments of 8. Each RPU provides 2 vCPU and 16GB RAM. Thus, 8 RPU is equivalent to 16 vCPU / 128GB RAM. The minimum RPU is 8.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | BigQuery | Redshift |
| --------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| Elasticity – Scaling for larger data volumes and faster queries | Fully automated serverless scaling – BigQuery automatically determines resource allocation and scales to petabytes without user intervention. Can dynamically burst beyond baseline slot allocations for performance optimization | Available via Elastic Resize – slow and limited, downtime required |
| Elasticity – Scaling for higher concurrency | Dynamic concurrency management with query queueing supporting up to 1,000 interactive queries and 20,000 batch queries per project per region. Automatic fair scheduling and slot distribution across workloads | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries |
**BigQuery** scales very well to large data volumes, and automatically assigns more compute resources when needed behind the scenes, in the form of "slots". BigQuery works either in an "on-demand pricing model", where slot assignment is completely in the hands of BigQuery and the state of the shared resource pool, or in "flat-rate pricing model" where slots are reserved in advanced. With reserved slots there is more control over compute resources, thus making scaling more predictable. Concurrency is limited to 100 users by default.
**Redshift** is limited in scale because even with RA3, it cannot distribute different workloads across clusters. While it can scale to up to 10 clusters automatically to support query concurrency, it can only handle a maximum of 50 queued queries across all clusters by default.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today. While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | BigQuery | Redshift |
| ------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | Search indexes (GA) for efficient text search optimization on STRING, JSON, and array columns. Support for LOG\_ANALYZER, NO\_OP\_ANALYZER, and PATTERN\_ANALYZER with column-level granularity for improved query performance and cost efficiency | None |
| Compute tuning | Serverless architecture with automatic resource optimization – no manual tuning required. Intelligent workload management with AI-powered resource allocation and dynamic slot distribution | Choice over number of nodes and their type |
| Storage format | Columnar & compressed storage (Capacitor format) with support for open table formats including Apache Iceberg, Delta Lake, and Hudi. Intelligent tiering with automatic long-term storage cost reduction after 90 days | Columnar & compressed storage (RA3 nodes) |
| Table-level partition & pruning techniques | • Automatic table organization with intelligent micro-partitioning • Clustering keys for data organization • Automatic partition pruning optimization • Supports time-based and custom partitioning strategies | • No table partitions • User-defined distribution & sort keys are used to optimize for speed |
| Result cache | Yes, with cross-user result caching and intelligent cache management for up to 24 hours | Yes |
| Warm cache (SSD) | BI Engine provides in-memory caching and acceleration for frequently accessed data and dashboards | Only with RA3 nodes at partition-level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, including advanced JSON functions, Lambda expressions, and native support for nested and repeated fields | Limited |
| Vector Search and AI Capabilities | • BigQuery ML for in-database machine learning • Vertex AI integration • Natural language querying with Gemini AI • Limited vector search capabilities | Limited AI capabilities – primarily through integrations |
| Query Optimizations | • Advanced query optimizer with Dremel engine • Search indexes with column-level granularity • Materialized views with smart refresh and automatic query rewriting • BI Engine in-memory acceleration • Gemini AI-powered query optimization and natural language querying • Cross-user result caching • Automatic partitioning and clustering optimization • Cost-based optimization with intelligent workload management | • Basic query optimizer • Materialized views • Result caching • ANALYZE for table statistics • Workload management (WLM) • Automated materialized views (AutoMV) • AI-driven scaling (Serverless preview) |
**BigQuery** lines up in benchmarks in the same ballpark as other cloud data warehouses but does come in consistently last in most queries. Beyond implementing according to best practices, there is little you can do to accelerate BigQuery performance, as it determines the amount of resources (slots) the query needs for you. BigQuery can be used together with "BigQuery BI Engine" for lower latency analytics. However, BI Engine is limited in terms of scale because it runs in memory. Its maximum capacity is 100GB.
**Redshift** does provide a result cache for accelerating repetitive query workloads and also has more tuning options than some others. But it does not deliver much faster compute performance than other cloud data warehouses in benchmarks. Sort keys can be used to optimize performance, but their contribution is limited. There is no support for indexes, and low-latency analytics at large data volumes is hard to achieve. Because Redshift decoupling of storage & compute is limited compared to other cloud data warehouses, it doesn't support isolating workloads, which means performance can degrade under pressure and competition for resources.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | BigQuery | Redshift |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • Sub-second to seconds response times at TB+ scale with BI Engine acceleration • Search indexes and materialized views provide significant performance improvements for dashboard queries • Intelligent caching reduces query costs for repeated dashboard access • Dynamic concurrency management supports high user loads | • Seconds to tens of seconds load times at 100s of GB scale • Can achieve faster performance with Concurrency Scaling and proper tuning |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set with native Google Cloud ecosystem integration • Strong integration with Looker, Looker Studio, and major BI tools • Gemini AI integration for natural language insights • Cross-cloud analytics capabilities • Advanced governance with Dataplex Universal Catalog | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Strong AWS ecosystem integration |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Dynamic concurrency supporting 1,000+ interactive queries with intelligent queuing and fair scheduling • Sub-second response times with BI Engine acceleration and search indexes • Serverless architecture eliminates infrastructure management overhead • Advanced caching and materialized views optimize repeated queries • AI-powered optimization for customer-facing applications • BigQuery ML for AI workloads | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries • Seconds-level response times typical • Automatic scaling for burst workloads • Limited AI application support |
| Ad hoc | • Excellent for ad-hoc analytics with serverless architecture requiring zero infrastructure management • Intelligent query optimization handles unpredictable workloads automatically • Gemini AI integration enables natural language querying for business users • Advanced JSON support and schema inference enable flexible data exploration • Cross-cloud analytics capabilities for federated queries | • Performance dependent on predefined distribution & sort keys • Elastic Resize enables adding compute resources • Typically subset of data loaded for ad-hoc analysis |
**BigQuery** is a mature general-purpose data warehouse, which lends itself well to internal BI & reporting. The fact that it's serverless in nature and tightly integrated with GCP, makes it very convenient for Ad-Hoc analytics and ML use cases on GCP. On the other hand, because BigQuery makes resource allocation decisions for you, it is not always the best fit for operational use cases and Data Apps where performance needs to be consistent and predictable.
**Redshift** was originally designed to support traditional internal BI reporting and dashboard use cases for analysts. As such, it is typically used as a general-purpose Enterprise data warehouse. With deep integrations into the AWS ecosystem, it can also leverage AWS ML service, making it also useful for ML projects. However, given the coupling of storage & compute, and the difficulty in delivering low-latency analytics at scale, it is less suited for operational use cases and customer-facing use cases like Data Apps. The coupling of storage and compute, together with the need to predefine sort & dist keys for optimal performance, make it challenging to use for Ad-Hoc analytics.
# ClickHouse vs Athena (2025) (/comparison/clickhouse-vs-athena)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | ClickHouse | Athena |
| ------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Yes – SharedMergeTree engine in ClickHouse Cloud enables full separation of storage and compute, with compute-compute separation through Warehouses feature (introduced 2025) allowing multiple isolated compute services sharing the same data | Yes, serverless with optional provisioned capacity. Workloads can be isolated through Workgroups and Capacity Reservations |
| Supported cloud infrastructure | AWS, GCP, Azure, cloud service and on-premises | AWS only |
| Isolated tenancy – option for dedicated resources | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client in cloud | • Multi-tenant pooled resources by default • Dedicated compute resources available via Provisioned Capacity • VPC endpoint connections supported |
| Control vs abstraction of compute | Configurable cluster size and compute types in ClickHouse Cloud with granular control over nodes (1-128 nodes) and node characteristics. Warehouses feature enables multiple isolated read-only compute environments. | • Serverless by default with no infrastructure control • Optional Provisioned Capacity allows dedicated DPU allocation (minimum 24 DPUs) • Two pricing models: on-demand ($5/TB scanned) or provisioned ($0.30/DPU-hour) |
| Self-hosted and hybrid deployment options | Self-managed deployments available with full control over infrastructure | No self-hosted options – serverless only |
| ACID Compliance and Transactions | Limited ACID compliance with MergeTree engine family. | No ACID compliance – eventual consistency model |
**ClickHouse** was originally developed at Yandex, the Russian search engine, as an OLAP engine for low latency analytics. It was built as an on-premise solution with coupled storage & compute, and a large variety of tuning options in the form of indexes and and merge trees. ClickHouse's architecture is famous for its focus on performance and low-latency queries. The tradeoff is that it is considered very difficult to work with. SQL support is very limited, and tuning/running it requires significant engineering resources.
**Athena** is serverless and built on a decoupled storage and compute architecture that queries data directly in S3, without the need to ingest/copy the data. It runs in multi-tenancy with shared resources. Users do not have control over the compute resources Athena chooses to allocate per query from the shared resource pool. For folks requiring additional or dedicated resources, they can reserve dedicated processing capacity in the form of Data Processing Units (DPU), with each DPU providing 4 vCPU and 16 GB RAM. RPU allocation ranges from 24 - 1000 per region.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | ClickHouse | Athena |
| --------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Automatic horizontal and vertical scaling in ClickHouse Cloud with SharedMergeTree architecture. Manual scaling for self-managed deployments with cluster rebalancing capabilities | • Fully abstracted on-demand scaling • Provisioned Capacity allows manual scaling of DPUs for predictable performance • Capacity reservations can be adjusted with minimum 1-hour billing periods |
| Elasticity – Scaling for higher concurrency | Supports high concurrency with proper resource allocation and configuration. Vertical auto-scaling and horizontal manual scaling. Additional warehouses can idle to zero billing. Primary service always on in multi-warehouse configurations. | • Default limit of 25 concurrent DML queries and 20 DDL queries (adjustable via service quotas) • Provisioned Capacity enables higher concurrency with dedicated DPUs • Query queuing available when capacity is exceeded |
**ClickHouse** doesn't offer any dedicated scaling features or mechanisms. While it can deliver linearly scalable performance for some types of queries, scaling itself has to be done manually. Hardware is self-managed in ClickHouse. This means that to scale you would have to provision a cluster and migrate.
**Athena** is a shared multi-tenant resource, with no guarantees on the amount or availability of the resources allocated for your queries. From a data volume perspective, it can scale to large volumes, but large data volumes can suffer from very long run times and frequent time outs. Query concurrency is maxed at 20. If scalability is a top priority, Athena is probably not the best choice.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | ClickHouse | Athena |
| ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Primary indexes • Skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • MergeTree indexes • Incremental Materialized views | No traditional indexes – relies on partition pruning and data organization in S3. Uses columnar formats and compression for optimization |
| Compute tuning | Configurable compute resources in cloud offering | • No compute tuning in on-demand mode • Provisioned Capacity allows DPU allocation control (4 vCPU and 16GB RAM per DPU) • Minimum 24 DPUs with scaling in 4-DPU increments |
| Storage format | Columnar, supports sorted, compressed, encoded & sparsely indexed files with native Apache Iceberg support. | Supports multiple formats: Parquet, ORC, Avro, JSON, CSV, TSV on S3. Native support for open table formats including Apache Iceberg, Apache Hudi, and Delta Lake |
| Table-level partition & pruning techniques | Partitioning by date/time and custom partitions with MergeTree indexes. | • User-defined table-level partitions with Hive-style partitioning • Pruning at partition level • Partition projection for advanced performance optimization • Supports open table formats with built-in partitioning |
| Result cache | Yes, results cache with TTL and query condition cache. | Query result caching for up to 30 days with configurable retention. Results reuse supported across workgroups |
| Warm cache (SSD) | Yes, at indexed data-range level granularity | No local caching – queries data directly from S3. Relies on S3's performance characteristics and intelligent tiering |
| Support for semi-structured data & JSON functions within SQL | Yes, including Lambda expressions and native JSON data type (GA in v25.3) | Yes, comprehensive JSON support including Lambda expressions, array functions, and native nested data handling |
| Vector Search and AI Capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference | No native AI or vector search capabilities |
| Query Optimizations | • Primary indexes (ORDER BY) • Data skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • Materialized views • Projections • PREWHERE optimization • Query analysis tools • Automatic global join reordering (v25.9) • Enhanced JSON query optimization • Streaming secondary indices | • Cost-based optimizer (CBO) in Athena engine v3 • Query result caching (up to 30 days) • Partition projection for advanced optimization • CTAS for precomputed queries • Join reordering and aggregation pushdown • Automatic parallel query execution • Support for columnar formats (Parquet, ORC) • Integration with AWS Glue Data Catalog |
**ClickHouse** is famous for being one of the fastest local runtimes ever built for OLAP workloads. Its columnar storage, compression and indexing capabilities make it a consistent leader in benchmarks. Its lack of support for standard SQL and lack of query optimizer means that it's less suitable for traditional BI workloads, and more suitable for engineering managed workloads. While fast, it requires a lot of tuning and optimization.
**Athena** (and Presto) are designed to query data where it is, sacrificing storage-compute optimizations. This makes it very convenient for easy and immediate querying but at the expense of performance. This typically puts Athena behind cloud data warehouses in terms of performance. But Athena still does relatively well in performance benchmarks, especially when external storage is managed by experts. While it supports partitions, there is no support for indexing, and together with the fact that resources are pooled from a shared multi-tenant service, low-latency and consistent performance are not Athena's sweet spot. A cloud data warehouse be more performant better than Athena in most cases.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | ClickHouse | Athena |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • Sub-second load times at TB+ scale with proper indexing • ClickHouse Cloud reduces engineering overhead with managed service • Proven low-latency performance (120ms at 2500 QPS in benchmarks) • Purpose-built for low-latency OLAP and real-time analytics | • Seconds to minutes response times for interactive dashboards • Performance varies based on data partitioning, file formats, and query optimization • Provisioned Capacity can improve consistency for dashboard workloads • Best suited for analytical dashboards rather than sub-second operational dashboards |
| Enterprise BI | • Growing ecosystem with 50+ integrations including major BI tools • Native MySQL protocol support enables broad BI tool compatibility • Strong SQL compliance with PostgreSQL compatibility • Best suited for modern analytical workloads and engineering-managed use cases | • Good integration with AWS ecosystem BI tools (QuickSight, etc.) • Standard SQL compatibility enables most BI tool connections • Cost-effective for variable workloads and ad-hoc analytics • JDBC/ODBC drivers support enterprise BI tools • Limited advanced BI features compared to dedicated data warehouses |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Sub-second response times at TB+ scale • Supports 1000 concurrent users per replica • Strong price-performance on customer-facing applications • Native vector search and embeddings | • Default concurrency limits (25 DML/20 DDL queries) may require service quota increases • Provisioned Capacity enables higher concurrency with dedicated resources • Seconds-level response times typical • Cost-effective for customer-facing analytics with proper optimization • Best suited for analytical rather than operational workloads • No native AI capabilities |
| Ad hoc | • Good for ad-hoc queries with ClickHouse Cloud's separated storage/compute architecture • Join optimizations enable more query complexity • Strong sampling capabilities (TABLESAMPLE) for exploratory analysis • Resource management through user quotas prevents query interference • Materialized views offer performance improvements for common aggregation patterns, ad-hoc users specify directly in SQL | • Purpose-built for ad-hoc analytics on data lakes • Serverless with zero infrastructure management • Direct querying of S3 data without ETL • Cost-effective pay-per-query model ideal for exploratory analysis • Strong support for multiple data formats and federated queries • Apache Spark integration for advanced analytics |
**ClickHouse** was not designed to be a data warehouse, but rather a low-latency query execution runtime. Managing it typically requires significant engineering overhead. Hence, it's a good fit for engineering managed operational use cases and customer-facing data apps, where low latency matters. It is not a good fit for a general purpose data warehouse, nor for Ad-Hoc analytics or ELT.
**Athena** is a great choice for Ad-Hoc analytics. You can keep the data where it is, and start querying without worrying about hardware or pretty much anything else, given that Athena is serverless and takes care of everything behind the scenes. However, it is not a great fit when you need consistent and fast query performance, and/or high concurrency. This is why it is typically not the best choice for operational and customer-facing applications. It can be also easily and flexibly used for batch processing, which is often leveraged for ML use cases.
# ClickHouse vs Databricks (2025) (/comparison/clickhouse-vs-databricks)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | ClickHouse | Databricks |
| ------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Yes – SharedMergeTree engine in ClickHouse Cloud enables full separation of storage and compute, with compute-compute separation through Warehouses feature (introduced 2025) allowing multiple isolated compute services sharing the same data | Yes |
| Supported cloud infrastructure | AWS, GCP, Azure, cloud service and on-premises | AWS, Azure, GCP. Marketplaces and BYOC |
| Isolated tenancy – option for dedicated resources | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client in cloud | • Control plane in Databricks account • Data plane in customer VPC (optional) • Storage in customer VPC • Serverless SQL runs in Databricks account with private connectivity |
| Control vs abstraction of compute | Configurable cluster size and compute types in ClickHouse Cloud with granular control over nodes (1-128 nodes) and node characteristics. Warehouses feature enables multiple isolated read-only compute environments. | • Configurable clusters and instance types • Serverless SQL warehouses (GA 2025) run in Databricks account with private connectivity, no public IPs • Pro/Classic warehouses run in customer VPC |
| Self-hosted and hybrid deployment options | Self-managed deployments available with full control over infrastructure | • Databricks on customer cloud accounts • Unity Catalog for hybrid governance |
| ACID Compliance and Transactions | Limited ACID compliance with MergeTree engine family. | • ACID transactions with Delta Lake • Time travel and versioning • Concurrent read/write operations |
**ClickHouse** was originally developed at Yandex, the Russian search engine, as an OLAP engine for low latency analytics. It was built as an on-premise solution with coupled storage & compute, and a large variety of tuning options in the form of indexes and merge trees. ClickHouse's architecture is famous for its focus on performance and low-latency queries. The tradeoff is that it is considered very difficult to work with. SQL support is very limited, and tuning/running it requires significant engineering resources.
**Databricks** was built by the founders of Spark as an analytics platform to support machine learning use cases. It leverages the Spark framework to process data residing in a data lake and is supported on AWS, GCP and Azure. Databricks coined the marketing term "Lakehouse '' architecture to illustrate the unification of data lake and data warehouse use cases. Customers still manage Spark clusters that process data residing in a Delta lake. Conversion of data to Delta Lake format is required to leverage the functionality of Delta Lake. Databricks Sql is a relatively new addition to simplify access to data stored in a data lake.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | ClickHouse | Databricks |
| --------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Automatic horizontal and vertical scaling in ClickHouse Cloud with SharedMergeTree architecture. Manual scaling for self-managed deployments with cluster rebalancing capabilities | Autoscaling clusters based on workload demand. Serverless SQL warehouses provide near-instant scaling (2-6 seconds startup) |
| Elasticity – Scaling for higher concurrency | Supports high concurrency with proper resource allocation and configuration. Vertical auto-scaling and horizontal manual scaling. Additional warehouses can idle to zero billing. Primary service always on in multi-warehouse configurations. | • 10 concurrent queries per cluster limit • Scales up to 40 clusters per warehouse (400 total concurrent queries) • Serverless SQL warehouses provide near-instant autoscaling • Pro/Classic warehouses take several minutes to provision new clusters • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on complexity |
**ClickHouse** doesn't offer any dedicated scaling features or mechanisms. While it can deliver linearly scalable performance for some types of queries, scaling itself has to be done manually. Hardware is self-managed in ClickHouse. This means that to scale you would have to provision a cluster and migrate.
**Databricks** allow for autoscaling of clusters based on utilization. Additionally, increasing concurrency associated with a sql endpoint can be accomplished through the addition of clusters. Query concurrency per cluster is maxed at 10. However, scaling with additional clusters for concurrency is possible. Databricks provides a choice of instance types.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today. While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | ClickHouse | Databricks |
| ------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Primary indexes • Skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • MergeTree indexes • Incremental Materialized views | None |
| Compute tuning | Configurable compute resources in cloud offering | Choice of cluster type, node types including SSD-optimized instances. Serverless provides automatic resource allocation with Intelligent Workload Management (IWM) |
| Storage format | Columnar, supports sorted, compressed, encoded & sparsely indexed files with native Apache Iceberg support. | • Delta Lake format with Liquid Clustering (February 2025 – replaces Z-ordering and traditional partitioning) • Cannot use Liquid Clustering alongside Z-ordering on same table • Allows for sorted data in Delta Lake • Requires Optimize to maintain ordering |
| Table-level partition & pruning techniques | Partitioning by date/time and custom partitions with MergeTree indexes. | • Table level partitioning • Liquid Clustering for improved query performance and reduced data skew (February 2025) • Z-ordering (legacy, replaced by Liquid Clustering) • Periodic optimization of storage required |
| Result cache | Yes, results cache with TTL and query condition cache. | Multi-layered caching: local in-memory cache per cluster plus remote result cache (serverless only) that persists across all warehouses in workspace |
| Warm cache (SSD) | Yes, at indexed data-range level granularity | Yes. Delta cache for data read by queries at file level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, including Lambda expressions and native JSON data type (GA in v25.3) | Yes |
| Vector Search and AI Capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference | • MLflow integration and Databricks ML platform • Native vector search in Delta Lake (Vector Search) • AI and ML workloads optimized |
| Query Optimizations | • Primary indexes (ORDER BY) • Data skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • Materialized views • Projections • PREWHERE optimization • Query analysis tools • Automatic global join reordering (v25.9) • Enhanced JSON query optimization • Streaming secondary indices | • Photon engine (C++ vectorized engine providing 3-8x average speedups, maximum speedups over 10x) • Automated stats collection (January 2025) enables cost-based optimization • Predictive I/O for faster point lookups and data updates • Liquid Clustering (February 2025) • Intelligent Workload Management (IWM) with AI-powered resource allocation • Delta cache • Materialized views support |
**ClickHouse** is famous for being one of the fastest local runtimes ever built for OLAP workloads. Its columnar storage, compression and indexing capabilities make it a consistent leader in benchmarks. Its lack of support for standard SQL and lack of query optimizer means that it's less suitable for traditional BI workloads, and more suitable for engineering managed workloads. While fast, it requires a lot of tuning and optimization.
**Databricks** is designed to leverage the Spark framework for processing large volumes of data. It leverages compressed Parquet files in a Delta Lake. To reduce the amount of data processed, it uses data pruning on partitions and Parquet file metadata. Databricks does not provide any indexes.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | ClickHouse | Databricks |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second load times at TB+ scale with proper indexing • ClickHouse Cloud reduces engineering overhead with managed service • Proven low-latency performance (120ms at 2500 QPS in benchmarks) • Purpose-built for low-latency OLAP and real-time analytics | • Sub-second to seconds load times at TB+ scale • Enhanced by Photon engine (3-8x average speedups) and Delta cache • Serverless SQL warehouses provide rapid startup (2-6 seconds) • Performance depends on cluster configuration |
| Enterprise BI | • Growing ecosystem with 50+ integrations including major BI tools • Native MySQL protocol support enables broad BI tool compatibility • Strong SQL compliance with PostgreSQL compatibility • Best suited for modern analytical workloads and engineering-managed use cases | • Strong for data science and ML workloads • Unified analytics platform approach • Growing traditional BI integrations • Serverless SQL warehouses improve accessibility • Delta sharing capabilities |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Sub-second response times at TB+ scale • Supports 1000 concurrent users per replica • Strong price-performance on customer-facing applications • Native vector search and embeddings | • 10 concurrent queries per cluster, scaling to 400 total concurrent queries per warehouse • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on workload complexity • Serverless provides near-instant autoscaling • Photon engine delivers 3-8x performance improvements • Strong ML and AI platform integration |
| Ad hoc | • Good for ad-hoc queries with ClickHouse Cloud's separated storage/compute architecture • Join optimizations enable more query complexity • Strong sampling capabilities (TABLESAMPLE) for exploratory analysis • Resource management through user quotas prevents query interference • Materialized views offer performance improvements for common aggregation patterns, ad-hoc users specify directly in SQL | • Excellent for ad-hoc with decoupled storage/compute • Serverless SQL warehouses provide instant provisioning • Intelligent Workload Management handles unpredictable workloads automatically • Strong for exploratory data analysis and ML workloads • Automated stats collection improves query planning |
**ClickHouse** was not designed to be a data warehouse, but rather a low-latency query execution runtime. Managing it typically requires significant engineering overhead. Hence, it's a good fit for engineering managed operational use cases and customer-facing data apps, where low latency matters. It is not a good fit for a general purpose data warehouse, nor for Ad-Hoc analytics or ELT.
**Databricks** is a mature Spark based platform proven for processing streaming data. It is widely used for Machine Learning use cases by data scientists through the use of integrated notebooks. From a low latency query perspective, while it offers features like Delta Cache, it does not provide specialized indexes that can deliver low latency queries.
# Databricks vs Athena (/comparison/databricks-vs-athena)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Databricks | Athena |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Yes | Yes, serverless with optional provisioned capacity. Workloads can be isolated through Workgroups and Capacity Reservations |
| Supported cloud infrastructure | AWS, Azure, GCP. Marketplaces and BYOC | AWS only |
| Isolated tenancy – option for dedicated resources | • Control plane in Databricks account • Data plane in customer VPC (optional) • Storage in customer VPC • Serverless SQL runs in Databricks account with private connectivity | • Multi-tenant pooled resources by default • Dedicated compute resources available via Provisioned Capacity • VPC endpoint connections supported |
| Control vs abstraction of compute | • Configurable clusters and instance types • Serverless SQL warehouses (GA 2025) run in Databricks account with private connectivity, no public IPs • Pro/Classic warehouses run in customer VPC | • Serverless by default with no infrastructure control • Optional Provisioned Capacity allows dedicated DPU allocation (minimum 24 DPUs) • Two pricing models: on-demand ($5/TB scanned) or provisioned ($0.30/DPU-hour) |
| Self-hosted and hybrid deployment options | • Databricks on customer cloud accounts • Unity Catalog for hybrid governance | No self-hosted options – serverless only |
| ACID Compliance and Transactions | • ACID transactions with Delta Lake • Time travel and versioning • Concurrent read/write operations | No ACID compliance – eventual consistency model |
**Databricks** was built by the founders of Spark as an analytics platform to support machine learning use cases. It leverages the Spark framework to process data residing in a data lake and is supported on AWS, GCP and Azure. Databricks coined the marketing term "Lakehouse '' architecture to illustrate the unification of data lake and data warehouse use cases. Customers still manage Spark clusters that process data residing in a Delta lake. Conversion of data to Delta Lake format is required to leverage the functionality of Delta Lake. Databricks Sql is a relatively new addition to simplify access to data stored in a data lake.
**Athena** is serverless and built on a decoupled storage and compute architecture that queries data directly in S3, without the need to ingest/copy the data. It runs in multi-tenancy with shared resources. Users do not have control over the compute resources Athena chooses to allocate per query from the shared resource pool. For folks requiring additional or dedicated resources, they can reserve dedicated processing capacity in the form of Data Processing Units (DPU), with each DPU providing 4 vCPU and 16 GB RAM. RPU allocation ranges from 24 - 1000 per region.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Databricks | Athena |
| --------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Autoscaling clusters based on workload demand. Serverless SQL warehouses provide near-instant scaling (2-6 seconds startup) | • Fully abstracted on-demand scaling • Provisioned Capacity allows manual scaling of DPUs for predictable performance • Capacity reservations can be adjusted with minimum 1-hour billing periods |
| Elasticity – Scaling for higher concurrency | • 10 concurrent queries per cluster limit • Scales up to 40 clusters per warehouse (400 total concurrent queries) • Serverless SQL warehouses provide near-instant autoscaling • Pro/Classic warehouses take several minutes to provision new clusters • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on complexity | • Default limit of 25 concurrent DML queries and 20 DDL queries (adjustable via service quotas) • Provisioned Capacity enables higher concurrency with dedicated DPUs • Query queuing available when capacity is exceeded |
\*\*Databricks \*\*allow for autoscaling of clusters based on utilization. Additionally, increasing concurrency associated with a sql endpoint can be accomplished through the addition of clusters. Query concurrency per cluster is maxed at 10. However, scaling with additional clusters for concurrency is possible. Databricks provides a choice of instance types.
**Athena** is a shared multi-tenant resource, with no guarantees on the amount or availability of the resources allocated for your queries. From a data volume perspective, it can scale to large volumes, but large data volumes can suffer from very long run times and frequent timeouts. Query concurrency is maxed at 20. If scalability is a top priority, Athena is probably not the best choice.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Databricks | Athena |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | None | No traditional indexes – relies on partition pruning and data organization in S3. Uses columnar formats and compression for optimization |
| Compute tuning | Choice of cluster type, node types including SSD-optimized instances. Serverless provides automatic resource allocation with Intelligent Workload Management (IWM) | • No compute tuning in on-demand mode • Provisioned Capacity allows DPU allocation control (4 vCPU and 16GB RAM per DPU) • Minimum 24 DPUs with scaling in 4-DPU increments |
| Storage format | • Delta Lake format with Liquid Clustering (February 2025 – replaces Z-ordering and traditional partitioning) • Cannot use Liquid Clustering alongside Z-ordering on same table • Allows for sorted data in Delta Lake • Requires Optimize to maintain ordering | Supports multiple formats: Parquet, ORC, Avro, JSON, CSV, TSV on S3. Native support for open table formats including Apache Iceberg, Apache Hudi, and Delta Lake |
| Table-level partition & pruning techniques | • Table level partitioning • Liquid Clustering for improved query performance and reduced data skew (February 2025) • Z-ordering (legacy, replaced by Liquid Clustering) • Periodic optimization of storage required | • User-defined table-level partitions with Hive-style partitioning • Pruning at partition level • Partition projection for advanced performance optimization • Supports open table formats with built-in partitioning |
| Result cache | Multi-layered caching: local in-memory cache per cluster plus remote result cache (serverless only) that persists across all warehouses in workspace | Query result caching for up to 30 days with configurable retention. Results reuse supported across workgroups |
| Warm cache (SSD) | Yes. Delta cache for data read by queries at file level granularity | No local caching – queries data directly from S3. Relies on S3’s performance characteristics and intelligent tiering |
| Support for semi-structured data & JSON functions within SQL | Yes | Yes, comprehensive JSON support including Lambda expressions, array functions, and native nested data handling |
| Vector Search and AI Capabilities | • MLflow integration and Databricks ML platform • Native vector search in Delta Lake (Vector Search) • AI and ML workloads optimized | No native AI or vector search capabilities |
| Query Optimizations | • Photon engine (C++ vectorized engine providing 3-8x average speedups, maximum speedups over 10x) • Automated stats collection (January 2025) enables cost-based optimization • Predictive I/O for faster point lookups and data updates • Liquid Clustering (February 2025) • Intelligent Workload Management (IWM) with AI-powered resource allocation • Delta cache • Materialized views support | • Cost-based optimizer (CBO) in Athena engine v3 • Query result caching (up to 30 days) • Partition projection for advanced optimization • CTAS for precomputed queries • Join reordering and aggregation pushdown • Automatic parallel query execution • Support for columnar formats (Parquet, ORC) • Integration with AWS Glue Data Catalog |
**Databricks** is designed to leverage the Spark framework for processing large volumes of data. It leverages compressed Parquet files in a Delta Lake. To reduce the amount of data processed, it uses data pruning on partitions and Parquet file metadata. Databricks does not provide any indexes.
**Athena** (and Presto) are designed to query data where it is, sacrificing storage-compute optimizations. This makes it very convenient for easy and immediate querying but at the expense of performance. This typically puts Athena behind cloud data warehouses in terms of performance. But Athena still does relatively well in performance benchmarks, especially when external storage is managed by experts. While it supports partitions, there is no support for indexing, and together with the fact that resources are pooled from a shared multi-tenant service, low-latency and consistent performance are not Athena’s sweet spot. A cloud data warehouse is more performant than Athena in most cases.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Databricks | Athena |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • Sub-second to seconds load times at TB+ scale • Enhanced by Photon engine (3-8x average speedups) and Delta cache • Serverless SQL warehouses provide rapid startup (2-6 seconds) • Performance depends on cluster configuration | • Seconds to minutes response times for interactive dashboards • Performance varies based on data partitioning, file formats, and query optimization • Provisioned Capacity can improve consistency for dashboard workloads • Best suited for analytical dashboards rather than sub-second operational dashboards |
| Enterprise BI | • Strong for data science and ML workloads • Unified analytics platform approach • Growing traditional BI integrations • Serverless SQL warehouses improve accessibility • Delta sharing capabilities | • Good integration with AWS ecosystem BI tools (QuickSight, etc.) • Standard SQL compatibility enables most BI tool connections • Cost-effective for variable workloads and ad-hoc analytics • JDBC/ODBC drivers support enterprise BI tools • Limited advanced BI features compared to dedicated data warehouses |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 10 concurrent queries per cluster, scaling to 400 total concurrent queries per warehouse • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on workload complexity • Serverless provides near-instant autoscaling • Photon engine delivers 3-8x performance improvements • Strong ML and AI platform integration | • Default concurrency limits (25 DML/20 DDL queries) may require service quota increases • Provisioned Capacity enables higher concurrency with dedicated resources • Seconds-level response times typical • Cost-effective for customer-facing analytics with proper optimization • Best suited for analytical rather than operational workloads • No native AI capabilities |
| Ad hoc | • Excellent for ad-hoc with decoupled storage/compute • Serverless SQL warehouses provide instant provisioning • Intelligent Workload Management handles unpredictable workloads automatically • Strong for exploratory data analysis and ML workloads • Automated stats collection improves query planning | • Purpose-built for ad-hoc analytics on data lakes • Serverless with zero infrastructure management • Direct querying of S3 data without ETL • Cost-effective pay-per-query model ideal for exploratory analysis • Strong support for multiple data formats and federated queries • Apache Spark integration for advanced analytics |
\*\*Databricks \*\*is a mature Spark based platform proven for processing streaming data. It is widely used for Machine Learning use cases by data scientists through the use of integrated notebooks. From a low latency query perspective, while it offers features like Delta Cache, it does not provide specialized indexes that can deliver low latency queries.
**Athena** is a great choice for Ad-Hoc analytics. You can keep the data where it is, and start querying without worrying about hardware or pretty much anything else, given that Athena is serverless and takes care of everything behind the scenes. However, it is not a great fit when you need consistent and fast query performance, and/or high concurrency. This is why it is typically not the best choice for operational and customer-facing applications. It can be also easily and flexibly used for batch processing, which is often leveraged for ML use cases.
# Databricks vs Snowflake (/comparison/databricks-vs-snowflake)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Databricks | Snowflake | Firebolt |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Separation of storage and compute | Yes | Yes | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. |
| Supported cloud infrastructure | AWS, Azure, GCP. Marketplaces and BYOC | AWS, Azure, GCP with full feature parity across all three major clouds | AWS (GCP coming soon) & anywhere (Firebolt Core) |
| Isolated tenancy – option for dedicated resources | • Control plane in Databricks account • Data plane in customer VPC (optional) • Storage in customer VPC • Serverless SQL runs in Databricks account with private connectivity | • Multi-tenant pooled resources • Isolated tenancy available via VPS tier | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client |
| Control vs abstraction of compute | • Configurable clusters and instance types • Serverless SQL warehouses (GA 2025) run in Databricks account with private connectivity, no public IPs • Pro/Classic warehouses run in customer VPC | • Configurable warehouse sizes (XS to 6XL) • Multi-cluster warehouses with auto-scaling • Choice between Generation 1 and Generation 2 standard warehouses • MAX\_CONCURRENCY\_LEVEL parameter for resource allocation | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. |
| Self-hosted and hybrid deployment options | • Databricks on customer cloud accounts • Unity Catalog for hybrid governance | Snowflake for Government Cloud and private cloud options available | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS |
| ACID Compliance and Transactions | • ACID transactions with Delta Lake • Time travel and versioning • Concurrent read/write operations | Full ACID compliance with Time Travel and zero-copy cloning capabilities | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications |
**Databricks** was built by the founders of Spark as an analytics platform to support machine learning use cases. It leverages the Spark framework to process data residing in a data lake and is supported on AWS, GCP and Azure. Databricks coined the marketing term “Lakehouse '' architecture to illustrate the unification of data lake and data warehouse use cases. Customers still manage Spark clusters that process data residing in a Delta lake. Conversion of data to Delta Lake format is required to leverage the functionality of Delta Lake. Databricks Sql is a relatively new addition to simplify access to data stored in a data lake.
**Snowflake** was one of the first decoupled storage and compute architectures, making it the first to have nearly unlimited compute scale and workload isolation, and horizontal user scalability. It runs on AWS, Azure and GCP. It is multi-tenant over shared resources in nature and requires you to move data out of your VPC and into the Snowflake cloud. “Virtual Private Snowflake” (VPS) is its highest-priced tier, and can run a dedicated isolated version of Snowflake. Its virtual warehouses can be T-shirt sized along an XS/S/M…/4XL axis, where each discrete T-shirt size is bundled with fixed HW properties that are abstracted from the users. Snowflake has recently added support for Snowflake managed Iceberg tables.
**Firebolt** is built on a natively decoupled storage & compute architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A “Firebolt Engine” can be granularly configured across # of nodes and different CPU/RAM/SSD combinations.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Databricks | Snowflake | Firebolt |
| --------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Autoscaling clusters based on workload demand. Serverless SQL warehouses provide near-instant scaling (2-6 seconds startup) | • Instant warehouse resize (XS to 6XL) with no downtime • Multi-cluster auto-scaling • Generation 2 warehouses provide \~2x performance improvement over Generation 1 | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. |
| Elasticity – Scaling for higher concurrency | • 10 concurrent queries per cluster limit • Scales up to 40 clusters per warehouse (400 total concurrent queries) • Serverless SQL warehouses provide near-instant autoscaling • Pro/Classic warehouses take several minutes to provision new clusters • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on complexity | • Single warehouse supports many concurrent queries (MAX\_CONCURRENCY\_LEVEL=8 controls resource allocation per query, not query limit) • Multi-cluster warehouses enable thousands of concurrent queries with auto-scaling • Unlimited virtual warehouses can be created | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. |
**Databricks** allow for autoscaling of clusters based on utilization. Additionally, increasing concurrency associated with a sql endpoint can be accomplished through the addition of clusters. Query concurrency per cluster is maxed at 10. However, scaling with additional clusters for concurrency is possible. Databricks provides a choice of instance types.
**Snowflake** scales very well both for data volumes and query concurrency. The decoupled storage/compute architecture supports resizing clusters without downtime, and in addition, supports auto-scaling horizontally for higher query concurrency during peak hours.
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn’t instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Databricks | Snowflake | Firebolt |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | None | • Search Optimization Service for point lookups and selective queries (additional cost) • Clustering keys for data organization and automatic clustering • Materialized views • Snowflake Optima automatic indexing on Generation 2 warehouses (no additional cost) • No traditional database indexes | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage |
| Compute tuning | Choice of cluster type, node types including SSD-optimized instances. Serverless provides automatic resource allocation with Intelligent Workload Management (IWM) | • Warehouse T-shirt sizing (XS to 6XL) • Multi-cluster configuration and scaling policies • Generation 1 vs Generation 2 warehouse selection • MAX\_CONCURRENCY\_LEVEL parameter tuning • Query Acceleration Service for long-running queries | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. |
| Storage format | • Delta Lake format with Liquid Clustering (February 2025 – replaces Z-ordering and traditional partitioning) • Cannot use Liquid Clustering alongside Z-ordering on same table • Allows for sorted data in Delta Lake • Requires Optimize to maintain ordering | Columnar micro-partitioned & compressed storage | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support |
| Table-level partition & pruning techniques | • Table level partitioning • Liquid Clustering for improved query performance and reduced data skew (February 2025) • Z-ordering (legacy, replaced by Liquid Clustering) • Periodic optimization of storage required | • Data automatically divided into micro-partitions • Automatic pruning at micro-partition level • Clustering keys for data organization with automatic clustering • Snowflake Optima provides additional automatic pruning optimization on Gen2 warehouses | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. |
| Result cache | Multi-layered caching: local in-memory cache per cluster plus remote result cache (serverless only) that persists across all warehouses in workspace | Yes | Yes, results and sub-results cache with transactional spoiling. |
| Warm cache (SSD) | Yes. Delta cache for data read by queries at file level granularity | Yes, at micro-partition level granularity | Yes, at indexed data-range level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes | Yes | Yes, including Lambda expressions and native nested array structures |
| Vector Search and AI Capabilities | • MLflow integration and Databricks ML platform • Native vector search in Delta Lake (Vector Search) • AI and ML workloads optimized | AI integration through Cortex AI and Snowpark ML | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference |
| Query Optimizations | • Photon engine (C++ vectorized engine providing 3-8x average speedups, maximum speedups over 10x) • Automated stats collection (January 2025) enables cost-based optimization • Predictive I/O for faster point lookups and data updates • Liquid Clustering (February 2025) • Intelligent Workload Management (IWM) with AI-powered resource allocation • Delta cache • Materialized views support | • Search Optimization Service for point lookups (additional cost) • Query Acceleration Service (QAS) for long-running and unpredictable workloads • Snowflake Optima automatic optimization on Generation 2 warehouses (no additional cost) • Automatic clustering with background maintenance • Materialized views with automatic refresh • Result cache (24hrs) • Cost-based optimization with dynamic query rewriting | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry |
**Databricks** is designed to leverage the Spark framework for processing large volumes of data. It leverages compressed Parquet files in a Delta Lake. To reduce the amount of data processed, it uses data pruning on partitions and Parquet file metadata. Databricks does not provide any indexes.
**Snowflake** typically comes on top for most queries when it comes to performance in public TPC-based benchmarks when compared to BigQuery and Redshift, but only marginally. Its micro partition storage approach effectively scans less data compared to larger partitions. The ability to isolate workloads over the decoupled storage & compute architecture lets you avoid competition for resources compared to multi-tenant shared resource solutions, and the ability to increase warehouse sizes can often enhance performance (for a higher price), but not always linearly. Snowflake’s recently released “Search optimization service” delivers index-like behavior for point queries, but comes at an additional cost.
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges that are significantly smaller. In addition, Firebolt lets users accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Databricks | Snowflake | Firebolt |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second to seconds load times at TB+ scale • Enhanced by Photon engine (3-8x average speedups) and Delta cache • Serverless SQL warehouses provide rapid startup (2-6 seconds) • Performance depends on cluster configuration | • Sub-second to seconds response times at TB+ scale with proper clustering and optimization • Enhanced by Query Acceleration Service and Search Optimization Service • Generation 2 warehouses provide \~2x performance improvement over Generation 1 • Snowflake Optima provides automatic optimization | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications |
| Enterprise BI | • Strong for data science and ML workloads • Unified analytics platform approach • Growing traditional BI integrations • Serverless SQL warehouses improve accessibility • Delta sharing capabilities | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Multi-cloud deployment options with consistent experience • Strong SQL compliance and wide ecosystem support • Zero-copy data sharing capabilities | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 10 concurrent queries per cluster, scaling to 400 total concurrent queries per warehouse • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on workload complexity • Serverless provides near-instant autoscaling • Photon engine delivers 3-8x performance improvements • Strong ML and AI platform integration | • Multi-cluster warehouses support thousands of concurrent users with auto-scaling • Individual warehouses support many concurrent queries (not limited to 8 concurrent queries) • Sub-second to seconds response times with proper optimization • Generation 2 warehouses provide significant performance improvements for high-concurrency workloads • AI integration through Cortex AI | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings |
| Ad hoc | • Excellent for ad-hoc with decoupled storage/compute • Serverless SQL warehouses provide instant provisioning • Intelligent Workload Management handles unpredictable workloads automatically • Strong for exploratory data analysis and ML workloads • Automated stats collection improves query planning | • Excellent for ad-hoc with decoupled storage/compute • Auto-scaling and instant compute provisioning • Minimal predefined optimization required • Query Acceleration Service handles unpredictable workloads automatically • Snowflake Optima provides automatic optimization for recurring patterns | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer |
**Databricks** is a mature Spark based platform proven for processing streaming data. It is widely used for Machine Learning use cases by data scientists through the use of integrated notebooks. From a low latency query perspective, while it offers features like Delta Cache, it does not provide specialized indexes that can deliver low latency queries.
**Snowflake** is a well rounded general purpose cloud data warehouse, that can also span beyond traditional BI & Analytics use cases into Ad-Hoc and ML use cases. Thanks to the flexible decoupeld storage & compute architecture that allows you to isolate and control the amount of compute per workload, it’s possible to tackle a broad spectrum of workloads. However, like its close siblings Redshift & BigQuery, it struggles to deliver low-latency query performance at scale, making it a lesser fit for operational use cases and customer-facing data apps.
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It’s great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
# Druid vs Athena (2025) (/comparison/druid-vs-athena)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Druid | Athena |
| ------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | No | Yes, serverless with optional provisioned capacity. Workloads can be isolated through Workgroups and Capacity Reservations |
| Supported cloud infrastructure | Can be installed anywhere | AWS only |
| Isolated tenancy – option for dedicated resources | Single tenant | • Multi-tenant pooled resources by default • Dedicated compute resources available via Provisioned Capacity • VPC endpoint connections supported |
| Control vs abstraction of compute | • Complex configuration of compute tier with multiple role-specific nodes • Configurable node count • Configurable compute types (virtual machines or kubernetes) | • Serverless by default with no infrastructure control • Optional Provisioned Capacity allows dedicated DPU allocation (minimum 24 DPUs) • Two pricing models: on-demand ($5/TB scanned) or provisioned ($0.30/DPU-hour) |
| Self-hosted and hybrid deployment options | Self-managed deployment required | No self-hosted options – serverless only |
| ACID Compliance and Transactions | Limited ACID support with eventual consistency | No ACID compliance – eventual consistency model |
**Druid** is an OLAP engine designed to provide fast real time analytics. Druid adopts a clustered architecture with servers that host various role specific processes. These processes address real time and batch ingestion, indexing, querying of historical and real time data. Apache Druid can be deployed as a virtual machine or a Kubernetes based cluster. Druid does not support a decoupled compute & storage architecture. Deep storage in the form of object storage is used to replicate data to.
**Athena** is serverless and built on a decoupled storage and compute architecture that queries data directly in S3, without the need to ingest/copy the data. It runs in multi-tenancy with shared resources. Users do not have control over the compute resources Athena chooses to allocate per query from the shared resource pool. For folks requiring additional or dedicated resources, they can reserve dedicated processing capacity in the form of Data Processing Units (DPU), with each DPU providing 4 vCPU and 16 GB RAM. RPU allocation ranges from 24 - 1000 per region.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Druid | Athena |
| --------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Scale-up of nodes requires careful planning and downtime. Addition of new nodes for scale-out is possible | • Fully abstracted on-demand scaling • Provisioned Capacity allows manual scaling of DPUs for predictable performance • Capacity reservations can be adjusted with minimum 1-hour billing periods |
| Elasticity – Scaling for higher concurrency | Supports 100s to 100,000s queries per second (1000+ QPS) with proper configuration and scaling | • Default limit of 25 concurrent DML queries and 20 DDL queries (adjustable via service quotas) • Provisioned Capacity enables higher concurrency with dedicated DPUs • Query queuing available when capacity is exceeded |
**Druid** provides the ability to handle fast ingest and high concurrency. Custom sizing and cluster tuning are required to balance the compute, memory, storage needs of each process within Druid and to provide high concurrency. Druid clusters can be grown by adding nodes with automatic rebalancing of storage segments assigned to nodes. Self hosted Druid on Kubernetes is an option that users leverage to simplify scaling. Additionally, Cloud based managed Druid offerings are being rolled out. However, these managed offerings are limited in scale and scaling is not granular.
**Athena** is a shared multi-tenant resource, with no guarantees on the amount or availability of the resources allocated for your queries. From a data volume perspective, it can scale to large volumes, but large data volumes can suffer from very long run times and frequent timeouts. Query concurrency is maxed at 20. If scalability is a top priority, Athena is probably not the best choice.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Druid | Athena |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | Compressed bitmap indexes for data access and roll-ups to manage aggregations | No traditional indexes – relies on partition pruning and data organization in S3. Uses columnar formats and compression for optimization |
| Compute tuning | On-premises, self-managed hardware. Druid requires infrastructure management and leverages commonly available instance types | • No compute tuning in on-demand mode • Provisioned Capacity allows DPU allocation control (4 vCPU and 16GB RAM per DPU) • Minimum 24 DPUs with scaling in 4-DPU increments |
| Storage format | Columnar storage format with time-based sorting | Supports multiple formats: Parquet, ORC, Avro, JSON, CSV, TSV on S3. Native support for open table formats including Apache Iceberg, Apache Hudi, and Delta Lake |
| Table-level partition & pruning techniques | Restrictive time-based partitioning. Can partition based on other secondary columns | • User-defined table-level partitions with Hive-style partitioning • Pruning at partition level • Partition projection for advanced performance optimization • Supports open table formats with built-in partitioning |
| Result cache | Ability to support caching on broker (set to off by default) | Query result caching for up to 30 days with configurable retention. Results reuse supported across workgroups |
| Warm cache (SSD) | Yes, at much larger segment level granularity | No local caching – queries data directly from S3. Relies on S3’s performance characteristics and intelligent tiering |
| Support for semi-structured data & JSON functions within SQL | Recommend flattening JSON or translate to array prior to loading. No support for JSON parsing at query runtime | Yes, comprehensive JSON support including Lambda expressions, array functions, and native nested data handling |
| Vector Search and AI Capabilities | No native AI or vector search capabilities | No native AI or vector search capabilities |
| Query Optimizations | • Compressed bitmap indexes • Roll-up aggregations • Time-based optimization • Query optimization requires manual tuning | • Cost-based optimizer (CBO) in Athena engine v3 • Query result caching (up to 30 days) • Partition projection for advanced optimization • CTAS for precomputed queries • Join reordering and aggregation pushdown • Automatic parallel query execution • Support for columnar formats (Parquet, ORC) • Integration with AWS Glue Data Catalog |
**Druid** provides high performance through columnar storage format, parallel processing, bitmap indexes and roll-ups. Druid, however, recommends a denormalized data model for performance needs. Join operations in Druid are a relatively new feature with various limitations, especially if there is a need to join large datasets.
**Athena** (and Presto) are designed to query data where it is, sacrificing storage-compute optimizations. This makes it very convenient for easy and immediate querying but at the expense of performance. This typically puts Athena behind cloud data warehouses in terms of performance. But Athena still does relatively well in performance benchmarks, especially when external storage is managed by experts. While it supports partitions, there is no support for indexing, and together with the fact that resources are pooled from a shared multi-tenant service, low-latency and consistent performance are not Athena’s sweet spot. A cloud data warehouse is more performant than Athena in most cases.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Druid | Athena |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • Sub-second load times optimized for time-series and real-time analytics • Built for high-concurrency interactive dashboards • Requires denormalized data model | • Seconds to minutes response times for interactive dashboards • Performance varies based on data partitioning, file formats, and query optimization • Provisioned Capacity can improve consistency for dashboard workloads • Best suited for analytical dashboards rather than sub-second operational dashboards |
| Enterprise BI | • Limited integrations with traditional Enterprise BI tools • Strong for real-time operational dashboards • Requires specialized visualization tools | • Good integration with AWS ecosystem BI tools (QuickSight, etc.) • Standard SQL compatibility enables most BI tool connections • Cost-effective for variable workloads and ad-hoc analytics • JDBC/ODBC drivers support enterprise BI tools • Limited advanced BI features compared to dedicated data warehouses |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Built for high concurrency (1000+ QPS) with distributed architecture • Sub-second response times for time-series data • Optimized for real-time operational applications • No AI capabilities | • Default concurrency limits (25 DML/20 DDL queries) may require service quota increases • Provisioned Capacity enables higher concurrency with dedicated resources • Seconds-level response times typical • Cost-effective for customer-facing analytics with proper optimization • Best suited for analytical rather than operational workloads • No native AI capabilities |
| Ad hoc | • Not optimized for ad-hoc queries • Requires predefined roll-ups and data modeling • Limited flexibility for exploratory analysis | • Purpose-built for ad-hoc analytics on data lakes • Serverless with zero infrastructure management • Direct querying of S3 data without ETL • Cost-effective pay-per-query model ideal for exploratory analysis • Strong support for multiple data formats and federated queries • Apache Spark integration for advanced analytics |
**Druid** is designed as an OLAP engine to provide fast access to aggregations that are run against large volumes of data. Druid is typically used for customer facing analytics and streaming data processing. Druid is used as an add-on with other data warehousing products that are efficient at scaling, joining, and filtering large volumes of data. It is not a suitable option for data warehouse replacement.
**Athena** is a great choice for Ad-Hoc analytics. You can keep the data where it is, and start querying without worrying about hardware or pretty much anything else, given that Athena is serverless and takes care of everything behind the scenes. However, it is not a great fit when you need consistent and fast query performance, and/or high concurrency. This is why it is typically not the best choice for operational and customer-facing applications. It can be also easily and flexibly used for batch processing, which is often leveraged for ML use cases.
# Druid vs ClickHouse (2025) (/comparison/druid-vs-clickhouse)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Druid | ClickHouse |
| ------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | No | Yes – SharedMergeTree engine in ClickHouse Cloud enables full separation of storage and compute, with compute-compute separation through Warehouses feature (introduced 2025) allowing multiple isolated compute services sharing the same data |
| Supported cloud infrastructure | Can be installed anywhere | AWS, GCP, Azure, cloud service and on-premises |
| Isolated tenancy – option for dedicated resources | Single tenant | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client in cloud |
| Control vs abstraction of compute | • Complex configuration of compute tier with multiple role-specific nodes • Configurable node count • Configurable compute types (virtual machines or kubernetes) | Configurable cluster size and compute types in ClickHouse Cloud with granular control over nodes (1-128 nodes) and node characteristics. Warehouses feature enables multiple isolated read-only compute environments. |
| Self-hosted and hybrid deployment options | Self-managed deployment required | Self-managed deployments available with full control over infrastructure |
| ACID Compliance and Transactions | Limited ACID support with eventual consistency | Limited ACID compliance with MergeTree engine family. |
**Druid** is an OLAP engine designed to provide fast real time analytics. Druid adopts a clustered architecture with servers that host various role specific processes. These processes address real time and batch ingestion, indexing, querying of historical and real time data. Apache Druid can be deployed as a virtual machine or a Kubernetes based cluster. Druid does not support a decoupled compute & storage architecture. Deep storage in the form of object storage is used to replicate data to.
**ClickHouse** was originally developed at Yandex, the Russian search engine, as an OLAP engine for low latency analytics. It was built as an on-premise solution with coupled compute & storage, and a large variety of tuning options in the form of indexes and merge trees. ClickHouse's architecture is famous for its focus on performance and low-latency queries. The tradeoff is that it is considered very difficult to work with. SQL support is very limited, and tuning/running it requires significant engineering resources.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Druid | ClickHouse |
| --------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Scale-up of nodes requires careful planning and downtime. Addition of new nodes for scale-out is possible | Automatic horizontal and vertical scaling in ClickHouse Cloud with SharedMergeTree architecture. Manual scaling for self-managed deployments with cluster rebalancing capabilities |
| Elasticity – Scaling for higher concurrency | Supports 100s to 100,000s queries per second (1000+ QPS) with proper configuration and scaling | Supports high concurrency with proper resource allocation and configuration. Vertical auto-scaling and horizontal manual scaling. Additional warehouses can idle to zero billing. Primary service always on in multi-warehouse configurations. |
**Druid** provides the ability to handle fast ingest and high concurrency. Custom sizing and cluster tuning are required to balance the compute, memory, storage needs of each process within Druid and to provide high concurrency. Druid clusters can be grown by adding nodes with automatic rebalancing of storage segments assigned to nodes. Self hosted Druid on Kubernetes is an option that users leverage to simplify scaling. Additionally, Cloud based managed Druid offerings are being rolled out. However, these managed offerings are limited in scale and scaling is not granular.
**ClickHouse** doesn't offer any dedicated scaling features or mechanisms. While it can deliver linearly scalable performance for some types of queries, scaling itself has to be done manually. Hardware is self-managed in ClickHouse. This means that to scale you would have to provision a cluster and migrate.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Druid | ClickHouse |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | Compressed bitmap indexes for data access and roll-ups to manage aggregations | • Primary indexes • Skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • MergeTree indexes • Incremental Materialized views |
| Compute tuning | On-premises, self-managed hardware. Druid requires infrastructure management and leverages commonly available instance types | Configurable compute resources in cloud offering |
| Storage format | Columnar storage format with time-based sorting | Columnar, supports sorted, compressed, encoded & sparsely indexed files with native Apache Iceberg support. |
| Table-level partition & pruning techniques | Restrictive time-based partitioning. Can partition based on other secondary columns | Partitioning by date/time and custom partitions with MergeTree indexes. |
| Result cache | Ability to support caching on broker (set to off by default) | Yes, results cache with TTL and query condition cache. |
| Warm cache (SSD) | Yes, at much larger segment level granularity | Yes, at indexed data-range level granularity |
| Support for semi-structured data & JSON functions within SQL | Recommend flattening JSON or translate to array prior to loading. No support for JSON parsing at query runtime | Yes, including Lambda expressions and native JSON data type (GA in v25.3) |
| Vector Search and AI Capabilities | No native AI or vector search capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference |
| Query Optimizations | • Compressed bitmap indexes • Roll-up aggregations • Time-based optimization • Query optimization requires manual tuning | • Primary indexes (ORDER BY) • Data skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • Materialized views • Projections • PREWHERE optimization • Query analysis tools • Automatic global join reordering (v25.9) • Enhanced JSON query optimization • Streaming secondary indices |
**Druid** provides high performance through columnar storage format, parallel processing, bitmap indexes and roll-ups. Druid, however, recommends a denormalized data model for performance needs. Join operations in Druid are a relatively new feature with various limitations, especially if there is a need to join large datasets.
**ClickHouse** is famous for being one of the fastest local runtimes ever built for OLAP workloads. Its columnar storage, compression and indexing capabilities make it a consistent leader in benchmarks. Its lack of support for standard SQL and lack of query optimizer means that it's less suitable for traditional BI workloads, and more suitable for engineering managed workloads. While fast, it requires a lot of tuning and optimization.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Druid | ClickHouse |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second load times optimized for time-series and real-time analytics • Built for high-concurrency interactive dashboards • Requires denormalized data model | • Sub-second load times at TB+ scale with proper indexing • ClickHouse Cloud reduces engineering overhead with managed service • Proven low-latency performance (120ms at 2500 QPS in benchmarks) • Purpose-built for low-latency OLAP and real-time analytics |
| Enterprise BI | • Limited integrations with traditional Enterprise BI tools • Strong for real-time operational dashboards • Requires specialized visualization tools | • Growing ecosystem with 50+ integrations including major BI tools • Native MySQL protocol support enables broad BI tool compatibility • Strong SQL compliance with PostgreSQL compatibility • Best suited for modern analytical workloads and engineering-managed use cases |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Built for high concurrency (1000+ QPS) with distributed architecture • Sub-second response times for time-series data • Optimized for real-time operational applications • No AI capabilities | • Sub-second response times at TB+ scale • Supports 1000 concurrent users per replica • Strong price-performance on customer-facing applications • Native vector search and embeddings |
| Ad hoc | • Not optimized for ad-hoc queries • Requires predefined roll-ups and data modeling • Limited flexibility for exploratory analysis | • Good for ad-hoc queries with ClickHouse Cloud's separated storage/compute architecture • Join optimizations enable more query complexity • Strong sampling capabilities (TABLESAMPLE) for exploratory analysis • Resource management through user quotas prevents query interference • Materialized views offer performance improvements for common aggregation patterns, ad-hoc users specify directly in SQL |
**Druid** is designed as an OLAP engine to provide fast access to aggregations that are run against large volumes of data. Druid is typically used for customer facing analytics and streaming data processing. Druid is used as an add-on with other data warehousing products that are efficient at scaling, joining, and filtering large volumes of data. It is not a suitable option for data warehouse replacement.
**ClickHouse** was not designed to be a data warehouse, but rather a low-latency query execution runtime. Managing it typically requires significant engineering overhead. Hence, it's a good fit for engineering managed operational use cases and customer-facing data apps, where low latency matters. It is not a good fit for a general purpose data warehouse, nor for Ad-Hoc analytics or ELT.
# Druid vs Databricks (2025) (/comparison/druid-vs-databricks)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Druid | Databricks |
| ------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | No | Yes |
| Supported cloud infrastructure | Can be installed anywhere | AWS, Azure, GCP. Marketplaces and BYOC |
| Isolated tenancy – option for dedicated resources | Single tenant | • Control plane in Databricks account • Data plane in customer VPC (optional) • Storage in customer VPC • Serverless SQL runs in Databricks account with private connectivity |
| Control vs abstraction of compute | • Complex configuration of compute tier with multiple role-specific nodes • Configurable node count • Configurable compute types (virtual machines or kubernetes) | • Configurable clusters and instance types • Serverless SQL warehouses (GA 2025) run in Databricks account with private connectivity, no public IPs • Pro/Classic warehouses run in customer VPC |
| Self-hosted and hybrid deployment options | Self-managed deployment required | • Databricks on customer cloud accounts • Unity Catalog for hybrid governance |
| ACID Compliance and Transactions | Limited ACID support with eventual consistency | • ACID transactions with Delta Lake • Time travel and versioning • Concurrent read/write operations |
**Druid** is an OLAP engine designed to provide fast real time analytics. Druid adopts a clustered architecture with servers that host various role specific processes. These processes address real time and batch ingestion, indexing, querying of historical and real time data. Apache Druid can be deployed as a virtual machine or a Kubernetes based cluster. Druid does not support a decoupled compute & storage architecture. Deep storage in the form of object storage is used to replicate data to.
**Databricks** was built by the founders of Spark as an analytics platform to support machine learning use cases. It leverages the Spark framework to process data residing in a data lake and is supported on AWS, GCP and Azure. Databricks coined the marketing term "Lakehouse '' architecture to illustrate the unification of data lake and data warehouse use cases. Customers still manage Spark clusters that process data residing in a Delta lake. Conversion of data to Delta Lake format is required to leverage the functionality of Delta Lake. Databricks Sql is a relatively new addition to simplify access to data stored in a data lake.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Druid | Databricks |
| --------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Scale-up of nodes requires careful planning and downtime. Addition of new nodes for scale-out is possible | Autoscaling clusters based on workload demand. Serverless SQL warehouses provide near-instant scaling (2-6 seconds startup) |
| Elasticity – Scaling for higher concurrency | Supports 100s to 100,000s queries per second (1000+ QPS) with proper configuration and scaling | • 10 concurrent queries per cluster limit • Scales up to 40 clusters per warehouse (400 total concurrent queries) • Serverless SQL warehouses provide near-instant autoscaling • Pro/Classic warehouses take several minutes to provision new clusters • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on complexity |
**Druid** provides the ability to handle fast ingest and high concurrency. Custom sizing and cluster tuning are required to balance the compute, memory, storage needs of each process within Druid and to provide high concurrency. Druid clusters can be grown by adding nodes with automatic rebalancing of storage segments assigned to nodes. Self hosted Druid on Kubernetes is an option that users leverage to simplify scaling. Additionally, Cloud based managed Druid offerings are being rolled out. However, these managed offerings are limited in scale and scaling is not granular.
**Databricks** allow for autoscaling of clusters based on utilization. Additionally, increasing concurrency associated with a sql endpoint can be accomplished through the addition of clusters. Query concurrency per cluster is maxed at 10. However, scaling with additional clusters for concurrency is possible. Databricks provides a choice of instance types.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Druid | Databricks |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | Compressed bitmap indexes for data access and roll-ups to manage aggregations | None |
| Compute tuning | On-premises, self-managed hardware. Druid requires infrastructure management and leverages commonly available instance types | Choice of cluster type, node types including SSD-optimized instances. Serverless provides automatic resource allocation with Intelligent Workload Management (IWM) |
| Storage format | Columnar storage format with time-based sorting | • Delta Lake format with Liquid Clustering (February 2025 – replaces Z-ordering and traditional partitioning) • Cannot use Liquid Clustering alongside Z-ordering on same table • Allows for sorted data in Delta Lake • Requires Optimize to maintain ordering |
| Table-level partition & pruning techniques | Restrictive time-based partitioning. Can partition based on other secondary columns | • Table level partitioning • Liquid Clustering for improved query performance and reduced data skew (February 2025) • Z-ordering (legacy, replaced by Liquid Clustering) • Periodic optimization of storage required |
| Result cache | Ability to support caching on broker (set to off by default) | Multi-layered caching: local in-memory cache per cluster plus remote result cache (serverless only) that persists across all warehouses in workspace |
| Warm cache (SSD) | Yes, at much larger segment level granularity | Yes. Delta cache for data read by queries at file level granularity |
| Support for semi-structured data & JSON functions within SQL | Recommend flattening JSON or translate to array prior to loading. No support for JSON parsing at query runtime | Yes |
| Vector Search and AI Capabilities | No native AI or vector search capabilities | • MLflow integration and Databricks ML platform • Native vector search in Delta Lake (Vector Search) • AI and ML workloads optimized |
| Query Optimizations | • Compressed bitmap indexes • Roll-up aggregations • Time-based optimization • Query optimization requires manual tuning | • Photon engine (C++ vectorized engine providing 3-8x average speedups, maximum speedups over 10x) • Automated stats collection (January 2025) enables cost-based optimization • Predictive I/O for faster point lookups and data updates • Liquid Clustering (February 2025) • Intelligent Workload Management (IWM) with AI-powered resource allocation • Delta cache • Materialized views support |
**Druid** provides high performance through columnar storage format, parallel processing, bitmap indexes and roll-ups. Druid, however, recommends a denormalized data model for performance needs. Join operations in Druid are a relatively new feature with various limitations, especially if there is a need to join large datasets.
**Databricks** is designed to leverage the Spark framework for processing large volumes of data. It leverages compressed Parquet files in a Delta Lake. To reduce the amount of data processed, it uses data pruning on partitions and Parquet file metadata. Databricks does not provide any indexes.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Druid | Databricks |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second load times optimized for time-series and real-time analytics • Built for high-concurrency interactive dashboards • Requires denormalized data model | • Sub-second to seconds load times at TB+ scale • Enhanced by Photon engine (3-8x average speedups) and Delta cache • Serverless SQL warehouses provide rapid startup (2-6 seconds) • Performance depends on cluster configuration |
| Enterprise BI | • Limited integrations with traditional Enterprise BI tools • Strong for real-time operational dashboards • Requires specialized visualization tools | • Strong for data science and ML workloads • Unified analytics platform approach • Growing traditional BI integrations • Serverless SQL warehouses improve accessibility • Delta sharing capabilities |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Built for high concurrency (1000+ QPS) with distributed architecture • Sub-second response times for time-series data • Optimized for real-time operational applications • No AI capabilities | • 10 concurrent queries per cluster, scaling to 400 total concurrent queries per warehouse • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on workload complexity • Serverless provides near-instant autoscaling • Photon engine delivers 3-8x performance improvements • Strong ML and AI platform integration |
| Ad hoc | • Not optimized for ad-hoc queries • Requires predefined roll-ups and data modeling • Limited flexibility for exploratory analysis | • Excellent for ad-hoc with decoupled storage/compute • Serverless SQL warehouses provide instant provisioning • Intelligent Workload Management handles unpredictable workloads automatically • Strong for exploratory data analysis and ML workloads • Automated stats collection improves query planning |
**Druid** is designed as an OLAP engine to provide fast access to aggregations that are run against large volumes of data. Druid is typically used for customer facing analytics and streaming data processing. Druid is used as an add-on with other data warehousing products that are efficient at scaling, joining, and filtering large volumes of data. It is not a suitable option for data warehouse replacement.
**Databricks** is a mature Spark based platform proven for processing streaming data. It is widely used for Machine Learning use cases by data scientists through the use of integrated notebooks. From a low latency query perspective, while it offers features like Delta Cache, it does not provide specialized indexes that can deliver low latency queries.
# Druid vs Snowflake (2025) (/comparison/druid-vs-snowflake)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Druid | Snowflake |
| ------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | No | Yes |
| Supported cloud infrastructure | Can be installed anywhere | AWS, Azure, GCP with full feature parity across all three major clouds |
| Isolated tenancy – option for dedicated resources | Single tenant | • Multi-tenant pooled resources • Isolated tenancy available via VPS tier |
| Control vs abstraction of compute | • Complex configuration of compute tier with multiple role-specific nodes • Configurable node count • Configurable compute types (virtual machines or kubernetes) | • Configurable warehouse sizes (XS to 6XL) • Multi-cluster warehouses with auto-scaling • Choice between Generation 1 and Generation 2 standard warehouses • MAX\_CONCURRENCY\_LEVEL parameter for resource allocation |
| Self-hosted and hybrid deployment options | Self-managed deployment required | Snowflake for Government Cloud and private cloud options available |
| ACID Compliance and Transactions | Limited ACID support with eventual consistency | Full ACID compliance with Time Travel and zero-copy cloning capabilities |
**Druid** provides the ability to handle fast ingest and high concurrency. Custom sizing and cluster tuning are required to balance the compute, memory, storage needs of each process within Druid and to provide high concurrency. Druid clusters can be grown by adding nodes with automatic rebalancing of storage segments assigned to nodes. Self hosted Druid on Kubernetes is an option that users leverage to simplify scaling. Additionally, Cloud based managed Druid offerings are being rolled out. However, these managed offerings are limited in scale and scaling is not granular.
**Snowflake** was one of the first decoupled storage and compute architectures, making it the first to have nearly unlimited compute scale and workload isolation, and horizontal user scalability. It runs on AWS, Azure and GCP. It is multi-tenant over shared resources in nature and requires you to move data out of your VPC and into the Snowflake cloud. “Virtual Private Snowflake” (VPS) is its highest-priced tier, and can run a dedicated isolated version of Snowflake. Its virtual warehouses can be T-shirt sized along an XS/S/M…/4XL axis, where each discrete T-shirt size is bundled with fixed HW properties that are abstracted from the users. Snowflake has recently added support for Snowflake managed Iceberg tables.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Druid | Snowflake |
| --------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Scale-up of nodes requires careful planning and downtime. Addition of new nodes for scale-out is possible | • Instant warehouse resize (XS to 6XL) with no downtime • Multi-cluster auto-scaling • Generation 2 warehouses provide \~2x performance improvement over Generation 1 |
| Elasticity – Scaling for higher concurrency | Supports 100s to 100,000s queries per second (1000+ QPS) with proper configuration and scaling | • Single warehouse supports many concurrent queries (MAX\_CONCURRENCY\_LEVEL=8 controls resource allocation per query, not query limit) • Multi-cluster warehouses enable thousands of concurrent queries with auto-scaling • Unlimited virtual warehouses can be created |
**Druid** provides the ability to handle fast ingest and high concurrency. Custom sizing and cluster tuning are required to balance the compute, memory, storage needs of each process within Druid and to provide high concurrency. Druid clusters can be grown by adding nodes with automatic rebalancing of storage segments assigned to nodes.
**Snowflake** scales very well both for data volumes and query concurrency. The decoupled storage/compute architecture supports resizing clusters without downtime, and in addition, supports auto-scaling horizontally for higher query concurrency during peak hours.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today. While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Druid | Snowflake |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | Compressed bitmap indexes for data access and roll-ups to manage aggregations | • Search Optimization Service for point lookups and selective queries (additional cost) • Clustering keys for data organization and automatic clustering • Materialized views • Snowflake Optima automatic indexing on Generation 2 warehouses (no additional cost) • No traditional database indexes |
| Compute tuning | On-premises, self-managed hardware. Druid requires infrastructure management and leverages commonly available instance types | • Warehouse T-shirt sizing (XS to 6XL) • Multi-cluster configuration and scaling policies • Generation 1 vs Generation 2 warehouse selection • MAX\_CONCURRENCY\_LEVEL parameter tuning • Query Acceleration Service for long-running queries |
| Storage format | Columnar storage format with time-based sorting | Columnar micro-partitioned & compressed storage |
| Table-level partition & pruning techniques | Restrictive time-based partitioning. Can partition based on other secondary columns | • Data automatically divided into micro-partitions • Automatic pruning at micro-partition level • Clustering keys for data organization with automatic clustering • Snowflake Optima provides additional automatic pruning optimization on Gen2 warehouses |
| Result cache | Ability to support caching on broker (set to off by default) | Yes |
| Warm cache (SSD) | Yes, at much larger segment level granularity | Yes, at micro-partition level granularity |
| Support for semi-structured data & JSON functions within SQL | Recommend flattening JSON or translate to array prior to loading. No support for JSON parsing at query runtime | Yes |
| Vector Search and AI Capabilities | No native AI or vector search capabilities | AI integration through Cortex AI and Snowpark ML |
| Query Optimizations | • Compressed bitmap indexes • Roll-up aggregations • Time-based optimization • Query optimization requires manual tuning | • Search Optimization Service for point lookups (additional cost) • Query Acceleration Service (QAS) for long-running and unpredictable workloads • Snowflake Optima automatic optimization on Generation 2 warehouses (no additional cost) • Automatic clustering with background maintenance • Materialized views with automatic refresh • Result cache (24hrs) • Cost-based optimization with dynamic query rewriting |
**Druid** provides high performance through columnar storage format, parallel processing, bitmap indexes and roll-ups. Druid, however, recommends a denormalized data model for performance needs. Join operations in Druid are a relatively new feature with various limitations, especially if there is a need to join large datasets.
**Snowflake** typically comes on top for most queries when it comes to performance in public TPC-based benchmarks when compared to BigQuery and Redshift, but only marginally. Its micro partition storage approach effectively scans less data compared to larger partitions. The ability to isolate workloads over the decoupled storage & compute architecture lets you avoid competition for resources compared to multi-tenant shared resource solutions, and the ability to increase warehouse sizes can often enhance performance (for a higher price), but not always linearly. Snowflake’s recently released “Search optimization service” delivers index-like behavior for point queries, but comes at an additional cost.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Druid | Snowflake |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • Sub-second load times optimized for time-series and real-time analytics • Built for high-concurrency interactive dashboards • Requires denormalized data model | • Sub-second to seconds response times at TB+ scale with proper clustering and optimization • Enhanced by Query Acceleration Service and Search Optimization Service • Generation 2 warehouses provide \~2x performance improvement over Generation 1 • Snowflake Optima provides automatic optimization |
| Enterprise BI | • Limited integrations with traditional Enterprise BI tools • Strong for real-time operational dashboards • Requires specialized visualization tools | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Multi-cloud deployment options with consistent experience • Strong SQL compliance and wide ecosystem support • Zero-copy data sharing capabilities |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Built for high concurrency (1000+ QPS) with distributed architecture • Sub-second response times for time-series data • Optimized for real-time operational applications • No AI capabilities | • Multi-cluster warehouses support thousands of concurrent users with auto-scaling • Individual warehouses support many concurrent queries (not limited to 8 concurrent queries) • Sub-second to seconds response times with proper optimization • Generation 2 warehouses provide significant performance improvements for high-concurrency workloads • AI integration through Cortex AI |
| Ad hoc | • Not optimized for ad-hoc queries • Requires predefined roll-ups and data modeling • Limited flexibility for exploratory analysis | • Excellent for ad-hoc with decoupled storage/compute • Auto-scaling and instant compute provisioning • Minimal predefined optimization required • Query Acceleration Service handles unpredictable workloads automatically • Snowflake Optima provides automatic optimization for recurring patterns |
**Druid** is designed as an OLAP engine to provide fast access to aggregations that are run against large volumes of data. Druid is typically used for customer facing analytics and streaming data processing. Druid is used as an add-on with other data warehousing products that are efficient at scaling, joining, and filtering large volumes of data. It is not a suitable option for data warehouse replacement.
**Snowflake** is a well rounded general purpose cloud data warehouse, that can also span beyond traditional BI & Analytics use cases into Ad-Hoc and ML use cases. Thanks to the flexible decoupeld storage & compute architecture that allows you to isolate and control the amount of compute per workload, it’s possible to tackle a broad spectrum of workloads. However, like its close siblings Redshift & BigQuery, it struggles to deliver low-latency query performance at scale, making it a lesser fit for operational use cases and customer-facing data apps.
# Firebolt vs Athena (2025) (/comparison/firebolt-vs-athena)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Firebolt | Athena |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. | Yes, serverless with optional provisioned capacity. Workloads can be isolated through Workgroups and Capacity Reservations |
| Supported cloud infrastructure | AWS (GCP coming soon) & anywhere (Firebolt Core) | AWS only |
| Isolated tenancy – option for dedicated resources | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client | • Multi-tenant pooled resources by default • Dedicated compute resources available via Provisioned Capacity • VPC endpoint connections supported |
| Control vs abstraction of compute | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. | • Serverless by default with no infrastructure control • Optional Provisioned Capacity allows dedicated DPU allocation (minimum 24 DPUs) • Two pricing models: on-demand ($5/TB scanned) or provisioned ($0.30/DPU-hour) |
| Self-hosted and hybrid deployment options | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS | No self-hosted options – serverless only |
| ACID Compliance and Transactions | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications | No ACID compliance – eventual consistency model |
**Firebolt** is built on a natively decoupled storage & compute architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A "Firebolt Engine" can be granularly configured across # of nodes and different CPU/RAM/SSD combinations.
**Athena** is serverless and built on a decoupled storage and compute architecture that queries data directly in S3, without the need to ingest/copy the data. It runs in multi-tenancy with shared resources. Users do not have control over the compute resources Athena chooses to allocate per query from the shared resource pool. For folks requiring additional or dedicated resources, they can reserve dedicated processing capacity in the form of Data Processing Units (DPU), with each DPU providing 4 vCPU and 16 GB RAM. RPU allocation ranges from 24 - 1000 per region.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Firebolt | Athena |
| --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. | • Fully abstracted on-demand scaling • Provisioned Capacity allows manual scaling of DPUs for predictable performance • Capacity reservations can be adjusted with minimum 1-hour billing periods |
| Elasticity – Scaling for higher concurrency | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. | • Default limit of 25 concurrent DML queries and 20 DDL queries (adjustable via service quotas) • Provisioned Capacity enables higher concurrency with dedicated DPUs • Query queuing available when capacity is exceeded |
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn't instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
**Athena** is a shared multi-tenant resource, with no guarantees on the amount or availability of the resources allocated for your queries. From a data volume perspective, it can scale to large volumes, but large data volumes can suffer from very long run times and frequent time outs. Query concurrency is maxed at 20. If scalability is a top priority, Athena is probably not the best choice.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Firebolt | Athena |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage | No traditional indexes – relies on partition pruning and data organization in S3. Uses columnar formats and compression for optimization |
| Compute tuning | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. | • No compute tuning in on-demand mode • Provisioned Capacity allows DPU allocation control (4 vCPU and 16GB RAM per DPU) • Minimum 24 DPUs with scaling in 4-DPU increments |
| Storage format | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support | Supports multiple formats: Parquet, ORC, Avro, JSON, CSV, TSV on S3. Native support for open table formats including Apache Iceberg, Apache Hudi, and Delta Lake |
| Table-level partition & pruning techniques | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. | • User-defined table-level partitions with Hive-style partitioning • Pruning at partition level • Partition projection for advanced performance optimization • Supports open table formats with built-in partitioning |
| Result cache | Yes, results and sub-results cache with transactional spoiling. | Query result caching for up to 30 days with configurable retention. Results reuse supported across workgroups |
| Warm cache (SSD) | Yes, at indexed data-range level granularity | No local caching – queries data directly from S3. Relies on S3's performance characteristics and intelligent tiering |
| Support for semi-structured data & JSON functions within SQL | Yes, including Lambda expressions and native nested array structures | Yes, comprehensive JSON support including Lambda expressions, array functions, and native nested data handling |
| Vector Search and AI Capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference | No native AI or vector search capabilities |
| Query Optimizations | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry | • Cost-based optimizer (CBO) in Athena engine v3 • Query result caching (up to 30 days) • Partition projection for advanced optimization • CTAS for precomputed queries • Join reordering and aggregation pushdown • Automatic parallel query execution • Support for columnar formats (Parquet, ORC) • Integration with AWS Glue Data Catalog |
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges, that are significantly smaller. In addition, Firebolt lets user accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
**Athena** (and Presto) are designed to query data where it is, sacrificing storage-compute optimizations. This makes it very convenient for easy and immediate querying but at the expense of performance. This typically puts Athena behind cloud data warehouses in terms of performance. But Athena still does relatively well in performance benchmarks, especially when external storage is managed by experts. While it supports partitions, there is no support for indexing, and together with the fact that resources are pooled from a shared multi-tenant service, low-latency and consistent performance are not Athena's sweet spot. A cloud data warehouse be more performant better than Athena in most cases.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Firebolt | Athena |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications | • Seconds to minutes response times for interactive dashboards • Performance varies based on data partitioning, file formats, and query optimization • Provisioned Capacity can improve consistency for dashboard workloads • Best suited for analytical dashboards rather than sub-second operational dashboards |
| Enterprise BI | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem | • Good integration with AWS ecosystem BI tools (QuickSight, etc.) • Standard SQL compatibility enables most BI tool connections • Cost-effective for variable workloads and ad-hoc analytics • JDBC/ODBC drivers support enterprise BI tools • Limited advanced BI features compared to dedicated data warehouses |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings | • Default concurrency limits (25 DML/20 DDL queries) may require service quota increases • Provisioned Capacity enables higher concurrency with dedicated resources • Seconds-level response times typical • Cost-effective for customer-facing analytics with proper optimization • Best suited for analytical rather than operational workloads • No native AI capabilities |
| Ad hoc | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer | • Purpose-built for ad-hoc analytics on data lakes • Serverless with zero infrastructure management • Direct querying of S3 data without ETL • Cost-effective pay-per-query model ideal for exploratory analysis • Strong support for multiple data formats and federated queries • Apache Spark integration for advanced analytics |
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It's great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
**Athena** is a great choice for Ad-Hoc analytics. You can keep the data where it is, and start querying without worrying about hardware or pretty much anything else, given that Athena is serverless and takes care of everything behind the scenes. However, it is not a great fit when you need consistent and fast query performance, and/or high concurrency. This is why it is typically not the best choice for operational and customer-facing applications. It can be also easily and flexibly used for batch processing, which is often leveraged for ML use cases.
# Firebolt vs BigQuery (2025) (/comparison/firebolt-vs-bigquery)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Firebolt | BigQuery |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Separation of storage and compute | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. | Fully serverless with complete separation of compute (Dremel) and storage (Colossus), powered by Jupiter network and Borg orchestration |
| Supported cloud infrastructure | AWS (GCP coming soon) & anywhere (Firebolt Core) | Google Cloud only |
| Isolated tenancy – option for dedicated resources | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client | • Multi-tenant pooled resources • VPC Service Controls provide enhanced security and connectivity isolation to customer VPCs • Cross-region disaster recovery for enterprise workloads |
| Control vs abstraction of compute | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. | Fully serverless with no control over compute resources – BigQuery automatically allocates computing resources as needed with intelligent workload management and dynamic slot allocation |
| Self-hosted and hybrid deployment options | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS | No self-hosted options – fully managed service only |
| ACID Compliance and Transactions | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications | Limited ACID support – eventual consistency model with some transactional capabilities |
**Firebolt** is built on a natively decoupled storage & compute architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A "Firebolt Engine" can be granularly configured across # of nodes and different CPU/RAM/SSD combinations
**BigQuery** was one of the first decoupled storage and compute architectures. It is a unique piece of engineering and not a typical data warehouse in part because it started as an on-demand serverless query engine. It runs in multi-tenancy with shared resources, allocated as "slots" which represent a virtual CPU that executes SQL. BigQuery determines how many slots a query requires, without the ability of the user to control it. BigQuery can be priced on a $/TB scanned basis or through slot reservations. A slot in BigQuery is logically equivalent to 0.5 vCPU and 0.5GB of RAM. There are multiple models to allocate slots in BigQuery.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Firebolt | BigQuery |
| --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. | Fully automated serverless scaling – BigQuery automatically determines resource allocation and scales to petabytes without user intervention. Can dynamically burst beyond baseline slot allocations for performance optimization |
| Elasticity – Scaling for higher concurrency | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. | Dynamic concurrency management with query queueing supporting up to 1,000 interactive queries and 20,000 batch queries per project per region. Automatic fair scheduling and slot distribution across workloads |
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn't instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
**BigQuery** scales very well to large data volumes, and automatically assigns more compute resources when needed behind the scenes, in the form of "slots". BigQuery works either in an "on-demand pricing model", where slot assignment is completely in the hands of BigQuery and the state of the shared resource pool, or in "flat-rate pricing model" where slots are reserved in advance. With reserved slots there is more control over compute resources, thus making scaling more predictable. Concurrency is limited to 100 users by default.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Firebolt | BigQuery |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage | Search indexes (GA) for efficient text search optimization on STRING, JSON, and array columns. Support for LOG\_ANALYZER, NO\_OP\_ANALYZER, and PATTERN\_ANALYZER with column-level granularity for improved query performance and cost efficiency |
| Compute tuning | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. | Serverless architecture with automatic resource optimization – no manual tuning required. Intelligent workload management with AI-powered resource allocation and dynamic slot distribution |
| Storage format | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support | Columnar & compressed storage (Capacitor format) with support for open table formats including Apache Iceberg, Delta Lake, and Hudi. Intelligent tiering with automatic long-term storage cost reduction after 90 days |
| Table-level partition & pruning techniques | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. | • Automatic table organization with intelligent micro-partitioning • Clustering keys for data organization • Automatic partition pruning optimization • Supports time-based and custom partitioning strategies |
| Result cache | Yes, results and sub-results cache with transactional spoiling. | Yes, with cross-user result caching and intelligent cache management for up to 24 hours |
| Warm cache (SSD) | Yes, at indexed data-range level granularity | BI Engine provides in-memory caching and acceleration for frequently accessed data and dashboards |
| Support for semi-structured data & JSON functions within SQL | Yes, including Lambda expressions and native nested array structures | Yes, including advanced JSON functions, Lambda expressions, and native support for nested and repeated fields |
| Vector Search and AI Capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference | • BigQuery ML for in-database machine learning • Vertex AI integration • Natural language querying with Gemini AI • Limited vector search capabilities |
| Query Optimizations | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry | • Advanced query optimizer with Dremel engine • Search indexes with column-level granularity • Materialized views with smart refresh and automatic query rewriting • BI Engine in-memory acceleration • Gemini AI-powered query optimization and natural language querying • Cross-user result caching • Automatic partitioning and clustering optimization • Cost-based optimization with intelligent workload management |
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges, that are significantly smaller. In addition, Firebolt lets user accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
**BigQuery** lines up in benchmarks in the same ballpark as other cloud data warehouses but does come in consistently last in most queries. Beyond implementing according to best practices, there is little you can do to accelerate BigQuery performance, as it determines the amount of resources (slots) the query needs for you. BigQuery can be used together with "BigQuery BI Engine" for lower latency analytics. However, BI Engine is limited in terms of scale because it runs in memory. Its maximum capacity is 100GB.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Firebolt | BigQuery |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications | • Sub-second to seconds response times at TB+ scale with BI Engine acceleration • Search indexes and materialized views provide significant performance improvements for dashboard queries • Intelligent caching reduces query costs for repeated dashboard access • Dynamic concurrency management supports high user loads |
| Enterprise BI | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem | • Mature and comprehensive Enterprise DW feature set with native Google Cloud ecosystem integration • Strong integration with Looker, Looker Studio, and major BI tools • Gemini AI integration for natural language insights • Cross-cloud analytics capabilities • Advanced governance with Dataplex Universal Catalog |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings | • Dynamic concurrency supporting 1,000+ interactive queries with intelligent queuing and fair scheduling • Sub-second response times with BI Engine acceleration and search indexes • Serverless architecture eliminates infrastructure management overhead • Advanced caching and materialized views optimize repeated queries • AI-powered optimization for customer-facing applications • BigQuery ML for AI workloads |
| Ad hoc | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer | • Excellent for ad-hoc analytics with serverless architecture requiring zero infrastructure management • Intelligent query optimization handles unpredictable workloads automatically • Gemini AI integration enables natural language querying for business users • Advanced JSON support and schema inference enable flexible data exploration • Cross-cloud analytics capabilities for federated queries |
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It's great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
**BigQuery** is a mature general-purpose data warehouse, which lends itself well to internal BI & reporting. The fact that it's serverless in nature and tightly integrated with GCP, makes it very convenient for Ad-Hoc analytics and ML use cases on GCP. On the other hand, because BigQuery makes resource allocation decisions for you, it is not always the best fit for operational use cases and Data Apps where performance needs to be consistent and predictable.
# Firebolt vs ClickHouse (2025) (/comparison/firebolt-vs-clickhouse)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Firebolt | ClickHouse |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. | Yes – SharedMergeTree engine in ClickHouse Cloud enables full separation of storage and compute, with compute-compute separation through Warehouses feature (introduced 2025) allowing multiple isolated compute services sharing the same data |
| Supported cloud infrastructure | AWS (GCP coming soon) & anywhere (Firebolt Core) | AWS, GCP, Azure, cloud service and on-premises |
| Isolated tenancy – option for dedicated resources | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client in cloud |
| Control vs abstraction of compute | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. | Configurable cluster size and compute types in ClickHouse Cloud with granular control over nodes (1-128 nodes) and node characteristics. Warehouses feature enables multiple isolated read-only compute environments. |
| Self-hosted and hybrid deployment options | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS | Self-managed deployments available with full control over infrastructure |
| ACID Compliance and Transactions | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications | Limited ACID compliance with MergeTree engine family. |
**Firebolt** is built around three architectural bets that compound: disaggregated storage and compute with sub-second cold start, a vectorized query engine with SIMD-optimized kernels, and sparse indexes that make full scans almost obsolete at scale. When you combine partition pruning with sparse indexes, you're often touching \<1% of the data for typical analytical queries. Disaggregated storage isn't just a cost story, it fundamentally changes what's possible for multi-cluster concurrency without the data movement overhead that kills P99 latency in shuffle-heavy systems.
**ClickHouse** is a column-oriented OLAP database built around a shared-nothing architecture where each node owns its storage and compute, scaling horizontally through explicit sharding and replication. The MergeTree engine family is the core primitive - sorted, sparse-indexed storage with variants (ReplacingMergeTree, AggregatingMergeTree) handling different update and aggregation patterns. Query execution is vectorized with LLVM JIT compilation for hot paths. The tight coupling of storage and compute that makes single-node performance exceptional becomes operationally non-trivial at very large scale.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Firebolt | ClickHouse |
| --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. | Automatic horizontal and vertical scaling in ClickHouse Cloud with SharedMergeTree architecture. Manual scaling for self-managed deployments with cluster rebalancing capabilities |
| Elasticity – Scaling for higher concurrency | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. | Supports high concurrency with proper resource allocation and configuration. Vertical auto-scaling and horizontal manual scaling. Additional warehouses can idle to zero billing. Primary service always on in multi-warehouse configurations. |
**Firebolt** separates storage and compute entirely. Data lives in object storage, engines are stateless and spin up in seconds. Multiple engine clusters can read and write from the same storage simultaneously, isolating workloads without contention or data movement. Scaling out is just provisioning; no rebalancing, no shuffle overhead, no distributed consensus. Uses simple SQL commands to scale vertically, horizontally and in cluster count with zero down-time.
**ClickHouse** Cloud runs on SharedMergeTree, a closed-source engine where data lives in object storage and compute nodes are fully stateless. No sharding needed; you scale by adding compute nodes against shared storage, with vertical autoscaling and scale-to-zero built in. Compute-compute separation lets multiple isolated node groups share the same data without extra copies, useful for isolating reads from writes. The main caveats: metadata coordination through ClickHouse Keeper introduces concurrency limits.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Firebolt | ClickHouse |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage | • Primary indexes • Skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • MergeTree indexes • Incremental Materialized views |
| Compute tuning | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. | Configurable compute resources in cloud offering |
| Storage format | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support | Columnar, supports sorted, compressed, encoded & sparsely indexed files with native Apache Iceberg support. |
| Table-level partition & pruning techniques | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. | Partitioning by date/time and custom partitions with MergeTree indexes. |
| Result cache | Yes, results and sub-results cache with transactional spoiling. | Yes, results cache with TTL and query condition cache. |
| Warm cache (SSD) | Yes, at indexed data-range level granularity | Yes, at indexed data-range level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, including Lambda expressions and native nested array structures | Yes, including Lambda expressions and native JSON data type (GA in v25.3) |
| Vector Search and AI Capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference |
| Query Optimizations | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry | • Primary indexes (ORDER BY) • Data skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • Materialized views • Projections • PREWHERE optimization • Query analysis tools • Automatic global join reordering (v25.9) • Enhanced JSON query optimization • Streaming secondary indices |
**Firebolt**'s performance advantage starts with how it reads data. While most data warehouses fetch entire partitions over the network, Firebolt works with indexed data ranges that are dramatically smaller, resulting in aggressive pruning that scans a fraction of what other systems touch. Multiple index types (aggregating indexes, join indexes) let users push this further for specific query patterns. Decoupled storage and compute means workloads can be isolated to guarantee consistent latency. A heavy analytical job doesn't degrade concurrent dashboard queries. A single engine handles hundreds of concurrent queries without needing to scale out, which makes it particularly well-suited for operational and customer-facing analytics where sub-second response times are non-negotiable.
**ClickHouse's** performance story is built on its columnar storage, compression, and indexing capabilities, which make it a consistent benchmark leader for raw query execution speed. The MergeTree engine family and sparse primary indexes are highly effective at minimizing I/O for queries that align well with the sort key. Where it gets complicated is that this performance is not automatic, it requires significant engineering investment to tune table engines, indexes, and merge strategies for each workload. The lack of a cost-based query optimizer means query performance is sensitive to how SQL is written, and standard BI tooling that generates arbitrary SQL will often underperform. For engineering-managed workloads where the query patterns are known and controlled, ClickHouse is extremely fast.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Firebolt | ClickHouse |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications | • Sub-second load times at TB+ scale with proper indexing • ClickHouse Cloud reduces engineering overhead with managed service • Proven low-latency performance (120ms at 2500 QPS in benchmarks) • Purpose-built for low-latency OLAP and real-time analytics |
| Enterprise BI | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem | • Growing ecosystem with 50+ integrations including major BI tools • Native MySQL protocol support enables broad BI tool compatibility • Strong SQL compliance with PostgreSQL compatibility • Best suited for modern analytical workloads and engineering-managed use cases |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings | • Sub-second response times at TB+ scale • Supports 1000 concurrent users per replica • Strong price-performance on customer-facing applications • Native vector search and embeddings |
| Ad hoc | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer | • Good for ad-hoc queries with ClickHouse Cloud’s separated storage/compute architecture • Join optimizations enable more query complexity • Strong sampling capabilities (TABLESAMPLE) for exploratory analysis • Resource management through user quotas prevents query interference • Materialized views offer performance improvements for common aggregation patterns, ad-hoc users specify directly in SQL |
**Firebolt** is purpose-built for operational and customer-facing analytics where consistent, sub-second query latency is a hard requirement at scale. The combination of aggressive data pruning, multiple index types, and workload isolation makes it the strongest fit for data apps and AI applications serving end users directly, scenarios where P99 latency matters as much as P50, and where concurrent workloads need to be isolated to guarantee SLA consistency. It's less suited for ad-hoc analytics or general-purpose enterprise BI, where ecosystem maturity matter more than raw performance.
**ClickHouse** is best suited for engineering-managed, high-throughput analytical workloads where query patterns are known in advance and teams have the expertise to tune schema design accordingly. It has a strong track record in observability, telemetry, event analytics, and time-series use cases - scenarios where data volumes are massive, ingestion rates are high, and the queries are well-understood. The tradeoff is operational overhead: getting the most out of ClickHouse requires deliberate investment in sort keys, projections, and materialized views. For teams willing to make that investment, it delivers exceptional raw performance. For teams expecting a more managed, self-optimizing experience, that overhead becomes a liability.
# Firebolt vs Databricks (2025) (/comparison/firebolt-vs-databricks)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Firebolt | Databricks |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. | Yes |
| Supported cloud infrastructure | AWS (GCP coming soon) & anywhere (Firebolt Core) | AWS, Azure, GCP. Marketplaces and BYOC |
| Isolated tenancy – option for dedicated resources | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client | • Control plane in Databricks account • Data plane in customer VPC (optional) • Storage in customer VPC • Serverless SQL runs in Databricks account with private connectivity |
| Control vs abstraction of compute | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. | • Configurable clusters and instance types • Serverless SQL warehouses (GA 2025) run in Databricks account with private connectivity, no public IPs • Pro/Classic warehouses run in customer VPC |
| Self-hosted and hybrid deployment options | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS | • Databricks on customer cloud accounts • Unity Catalog for hybrid governance |
| ACID Compliance and Transactions | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications | • ACID transactions with Delta Lake • Time travel and versioning • Concurrent read/write operations |
**Firebolt** is built on a natively decoupled storage & compute architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A "Firebolt Engine" can be granularly configured across # of nodes and different CPU/RAM/SSD combinations.
**Databricks** was built by the founders of Spark as an analytics platform to support machine learning use cases. It leverages the Spark framework to process data residing in a data lake and is supported on AWS, GCP and Azure. Databricks coined the marketing term "Lakehouse '' architecture to illustrate the unification of data lake and data warehouse use cases. Customers still manage Spark clusters that process data residing in a Delta lake. Conversion of data to Delta Lake format is required to leverage the functionality of Delta Lake. Databricks Sql is a relatively new addition to simplify access to data stored in a data lake.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Firebolt | Databricks |
| --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. | Autoscaling clusters based on workload demand. Serverless SQL warehouses provide near-instant scaling (2-6 seconds startup) |
| Elasticity – Scaling for higher concurrency | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. | • 10 concurrent queries per cluster limit • Scales up to 40 clusters per warehouse (400 total concurrent queries) • Serverless SQL warehouses provide near-instant autoscaling • Pro/Classic warehouses take several minutes to provision new clusters • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on complexity |
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn't instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
**Databricks** allow for autoscaling of clusters based on utilization. Additionally, increasing concurrency associated with a sql endpoint can be accomplished through the addition of clusters. Query concurrency per cluster is maxed at 10. However, scaling with additional clusters for concurrency is possible. Databricks provides a choice of instance types.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Firebolt | Databricks |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage | None |
| Compute tuning | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. | Choice of cluster type, node types including SSD-optimized instances. Serverless provides automatic resource allocation with Intelligent Workload Management (IWM) |
| Storage format | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support | • Delta Lake format with Liquid Clustering (February 2025 – replaces Z-ordering and traditional partitioning) • Cannot use Liquid Clustering alongside Z-ordering on same table • Allows for sorted data in Delta Lake • Requires Optimize to maintain ordering |
| Table-level partition & pruning techniques | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. | • Table level partitioning • Liquid Clustering for improved query performance and reduced data skew (February 2025) • Z-ordering (legacy, replaced by Liquid Clustering) • Periodic optimization of storage required |
| Result cache | Yes, results and sub-results cache with transactional spoiling. | Multi-layered caching: local in-memory cache per cluster plus remote result cache (serverless only) that persists across all warehouses in workspace |
| Warm cache (SSD) | Yes, at indexed data-range level granularity | Yes. Delta cache for data read by queries at file level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, including Lambda expressions and native nested array structures | Yes |
| Vector Search and AI Capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference | • MLflow integration and Databricks ML platform • Native vector search in Delta Lake (Vector Search) • AI and ML workloads optimized |
| Query Optimizations | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry | • Photon engine (C++ vectorized engine providing 3-8x average speedups, maximum speedups over 10x) • Automated stats collection (January 2025) enables cost-based optimization • Predictive I/O for faster point lookups and data updates • Liquid Clustering (February 2025) • Intelligent Workload Management (IWM) with AI-powered resource allocation • Delta cache • Materialized views support |
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges, that are significantly smaller. In addition, Firebolt lets users accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
**Databricks** is designed to leverage the Spark framework for processing large volumes of data. It leverages compressed Parquet files in a Delta Lake. To reduce the amount of data processed, it uses data pruning on partitions and Parquet file metadata. Databricks does not provide any indexes.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Firebolt | Databricks |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications | • Sub-second to seconds load times at TB+ scale • Enhanced by Photon engine (3-8x average speedups) and Delta cache • Serverless SQL warehouses provide rapid startup (2-6 seconds) • Performance depends on cluster configuration |
| Enterprise BI | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem | • Strong for data science and ML workloads • Unified analytics platform approach • Growing traditional BI integrations • Serverless SQL warehouses improve accessibility • Delta sharing capabilities |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings | • 10 concurrent queries per cluster, scaling to 400 total concurrent queries per warehouse • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on workload complexity • Serverless provides near-instant autoscaling • Photon engine delivers 3-8x performance improvements • Strong ML and AI platform integration |
| Ad hoc | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer | • Excellent for ad-hoc with decoupled storage/compute • Serverless SQL warehouses provide instant provisioning • Intelligent Workload Management handles unpredictable workloads automatically • Strong for exploratory data analysis and ML workloads • Automated stats collection improves query planning |
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It's great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
**Databricks** is a mature Spark based platform proven for processing streaming data. It is widely used for Machine Learning use cases by data scientists through the use of integrated notebooks. From a low latency query perspective, while it offers features like Delta Cache, it does not provide specialized indexes that can deliver low latency queries.
# Firebolt vs Druid (2025) (/comparison/firebolt-vs-druid)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Firebolt | Druid |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. | No |
| Supported cloud infrastructure | AWS (GCP coming soon) & anywhere (Firebolt Core) | Can be installed anywhere |
| Isolated tenancy – option for dedicated resources | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client | Single tenant |
| Control vs abstraction of compute | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. | • Complex configuration of compute tier with multiple role-specific nodes • Configurable node count • Configurable compute types (virtual machines or kubernetes) |
| Self-hosted and hybrid deployment options | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS | Self-managed deployment required |
| ACID Compliance and Transactions | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications | Limited ACID support with eventual consistency |
**Firebolt** is built on a natively decoupled compute & storage architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A "Firebolt Engine" can be granularly configured across # of nodes and different CPU/RAM/SSD combinations.
**Druid** is an OLAP engine designed to provide fast real time analytics. Druid adopts a clustered architecture with servers that host various role specific processes. These processes address real time and batch ingestion, indexing, querying of historical and real time data. Apache Druid can be deployed as a virtual machine or a Kubernetes based cluster. Druid does not support a decoupled compute & storage architecture. Deep storage in the form of object storage is used to replicate data to.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Firebolt | Druid |
| --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. | Scale-up of nodes requires careful planning and downtime. Addition of new nodes for scale-out is possible |
| Elasticity – Scaling for higher concurrency | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. | Supports 100s to 100,000s queries per second (1000+ QPS) with proper configuration and scaling |
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn't instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
**Druid** provides the ability to handle fast ingest and high concurrency. Custom sizing and cluster tuning are required to balance the compute, memory, storage needs of each process within Druid and to provide high concurrency. Druid clusters can be grown by adding nodes with automatic rebalancing of storage segments assigned to nodes. Self hosted Druid on Kubernetes is an option that users leverage to simplify scaling. Additionally, Cloud based managed Druid offerings are being rolled out. However, these managed offerings are limited in scale and scaling is not granular.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today. While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Firebolt | Druid |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage | Compressed bitmap indexes for data access and roll-ups to manage aggregations |
| Compute tuning | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. | On-premises, self-managed hardware. Druid requires infrastructure management and leverages commonly available instance types |
| Storage format | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support | Columnar storage format with time-based sorting |
| Table-level partition & pruning techniques | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. | Restrictive time-based partitioning. Can partition based on other secondary columns |
| Result cache | Yes, results and sub-results cache with transactional spoiling. | Ability to support caching on broker (set to off by default) |
| Warm cache (SSD) | Yes, at indexed data-range level granularity | Yes, at much larger segment level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, including Lambda expressions and native nested array structures | Recommend flattening JSON or translate to array prior to loading. No support for JSON parsing at query runtime |
| Vector Search and AI Capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference | No native AI or vector search capabilities |
| Query Optimizations | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry | • Compressed bitmap indexes • Roll-up aggregations • Time-based optimization • Query optimization requires manual tuning |
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges, that are significantly smaller. In addition, Firebolt lets users accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
**Druid** provides high performance through columnar storage format, parallel processing, bitmap indexes and roll-ups. Druid, however, recommends a denormalized data model for performance needs. Join operations in Druid are a relatively new feature with various limitations, especially if there is a need to join large datasets.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Firebolt | Druid |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications | • Sub-second load times optimized for time-series and real-time analytics • Built for high-concurrency interactive dashboards • Requires denormalized data model |
| Enterprise BI | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem | • Limited integrations with traditional Enterprise BI tools • Strong for real-time operational dashboards • Requires specialized visualization tools |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings | • Built for high concurrency (1000+ QPS) with distributed architecture • Sub-second response times for time-series data • Optimized for real-time operational applications • No AI capabilities |
| Ad hoc | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer | • Not optimized for ad-hoc queries • Requires predefined roll-ups and data modeling • Limited flexibility for exploratory analysis |
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It's great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
**Druid** is designed as an OLAP engine to provide fast access to aggregations that are run against large volumes of data. Druid is typically used for customer facing analytics and streaming data processing. Druid is used as an add-on with other data warehousing products that are efficient at scaling, joining, and filtering large volumes of data. It is not a suitable option for data warehouse replacement.
# Firebolt vs Redshift (2025) (/comparison/firebolt-vs-redshift)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Firebolt | Redshift |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------ |
| Separation of storage and compute | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. | RA3 instances enable separation of compute and storage, but limited workload isolation compared to other platforms |
| Supported cloud infrastructure | AWS (GCP coming soon) & anywhere (Firebolt Core) | AWS only |
| Isolated tenancy – option for dedicated resources | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client | • Isolated tenant & resources • Runs in your VPC |
| Control vs abstraction of compute | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. | • Configurable cluster size • Configurable compute types |
| Self-hosted and hybrid deployment options | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS | Limited hybrid options with Redshift Serverless |
| ACID Compliance and Transactions | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications | ACID compliant at table level with some limitations on concurrent operations |
**Firebolt** is built on a natively decoupled storage & compute architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A "Firebolt Engine" can be granularly configured across # of nodes and different CPU/RAM/SSD combinations
**Redshift** has the oldest architecture, being the first Cloud DW in the group. Its architecture wasn't designed to separate storage & compute. While it now has RA3 nodes which allow you to scale compute and only cache the data you need locally, all compute still operates together. You cannot separate and isolate different workloads over the same data, which puts it behind other decoupled storage/compute architectures. Redshift runs as an isolated tenant per customer, and unlike other cloud data warehouses, it is deployed in your VPC. Redshift offers a serverless option which is based on an abstracted unit called Redshift Processing Unit (RPU) ranging from 8 to 512 in increments of 8. Each RPU provides 2 vCPU and 16GB RAM. Thus, 8 RPU is equivalent to 16 vCPU / 128GB RAM. The minimum RPU is 8.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Firebolt | Redshift |
| --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| Elasticity – Scaling for larger data volumes and faster queries | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. | Available via Elastic Resize – slow and limited, downtime required |
| Elasticity – Scaling for higher concurrency | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries |
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn't instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
**Redshift** is limited in scale because even with RA3, it cannot distribute different workloads across clusters. While it can scale to up to 10 clusters automatically to support query concurrency, it can only handle a maximum of 50 queued queries across all clusters by default.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Firebolt | Redshift |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage | None |
| Compute tuning | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. | Choice over number of nodes and their type |
| Storage format | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support | Columnar & compressed storage (RA3 nodes) |
| Table-level partition & pruning techniques | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. | • No table partitions • User-defined distribution & sort keys are used to optimize for speed |
| Result cache | Yes, results and sub-results cache with transactional spoiling. | Yes |
| Warm cache (SSD) | Yes, at indexed data-range level granularity | Only with RA3 nodes at partition-level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, including Lambda expressions and native nested array structures | Limited |
| Vector Search and AI Capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference | Limited AI capabilities – primarily through integrations |
| Query Optimizations | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry | • Basic query optimizer • Materialized views • Result caching • ANALYZE for table statistics • Workload management (WLM) • Automated materialized views (AutoMV) • AI-driven scaling (Serverless preview) |
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges, that are significantly smaller. In addition, Firebolt lets user accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
**Redshift** does provide a result cache for accelerating repetitive query workloads and also has more tuning options than some others. But it does not deliver much faster compute performance than other cloud data warehouses in benchmarks. Sort keys can be used to optimize performance, but their contribution is limited. There is no support for indexes, and low-latency analytics at large data volumes is hard to achieve. Because Redshift decoupling of storage & compute is limited compared to other cloud data warehouses, it doesn't support isolating workloads, which means performance can degrade under pressure and competition for resources.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Firebolt | Redshift |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications | • Seconds to tens of seconds load times at 100s of GB scale • Can achieve faster performance with Concurrency Scaling and proper tuning |
| Enterprise BI | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Strong AWS ecosystem integration |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries • Seconds-level response times typical • Automatic scaling for burst workloads • Limited AI application support |
| Ad hoc | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer | • Performance dependent on predefined distribution & sort keys • Elastic Resize enables adding compute resources • Typically subset of data loaded for ad-hoc analysis |
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It's great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
**Redshift** was originally designed to support traditional internal BI reporting and dashboard use cases for analysts. As such, it is typically used as a general-purpose Enterprise data warehouse. With deep integrations into the AWS ecosystem, it can also leverage AWS ML service, making it also useful for ML projects. However, given the coupling of storage & compute, and the difficulty in delivering low-latency analytics at scale, it is less suited for operational use cases and customer-facing use cases like Data Apps. The coupling of storage and compute, together with the need to predefine sort & dist keys for optimal performance, make it challenging to use for Ad-Hoc analytics.
# Firebolt vs Snowflake (2025) (/comparison/firebolt-vs-snowflake)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Firebolt | Snowflake |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. | Yes |
| Supported cloud infrastructure | AWS (GCP coming soon) & anywhere (Firebolt Core) | AWS, Azure, GCP with full feature parity across all three major clouds |
| Isolated tenancy – option for dedicated resources | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client | • Multi-tenant pooled resources • Isolated tenancy available via VPS tier |
| Control vs abstraction of compute | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. | • Configurable warehouse sizes (XS to 6XL) • Multi-cluster warehouses with auto-scaling • Choice between Generation 1 and Generation 2 standard warehouses • MAX\_CONCURRENCY\_LEVEL parameter for resource allocation |
| Self-hosted and hybrid deployment options | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS | Snowflake for Government Cloud and private cloud options available |
| ACID Compliance and Transactions | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications | Full ACID compliance with Time Travel and zero-copy cloning capabilities |
**Firebolt** is built on a natively decoupled storage & compute architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A "Firebolt Engine" can be granularly configured across # of nodes and different CPU/RAM/SSD combinations.
**Snowflake** was one of the first decoupled storage and compute architectures, making it the first to have nearly unlimited compute scale and workload isolation, and horizontal user scalability. It runs on AWS, Azure and GCP. It is multi-tenant over shared resources in nature and requires you to move data out of your VPC and into the Snowflake cloud. "Virtual Private Snowflake" (VPS) is its highest-priced tier, and can run a dedicated isolated version of Snowflake. Its virtual warehouses can be T-shirt sized along an XS/S/M…/4XL axis, where each discrete T-shirt size is bundled with fixed HW properties that are abstracted from the users. Snowflake has recently added support for Snowflake managed Iceberg tables.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Firebolt | Snowflake |
| --------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. | • Instant warehouse resize (XS to 6XL) with no downtime • Multi-cluster auto-scaling • Generation 2 warehouses provide \~2x performance improvement over Generation 1 |
| Elasticity – Scaling for higher concurrency | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. | • Single warehouse supports many concurrent queries (MAX\_CONCURRENCY\_LEVEL=8 controls resource allocation per query, not query limit) • Multi-cluster warehouses enable thousands of concurrent queries with auto-scaling • Unlimited virtual warehouses can be created |
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn't instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
**Snowflake** scales very well both for data volumes and query concurrency. The decoupled storage/compute architecture supports resizing clusters without downtime, and in addition, supports auto-scaling horizontally for higher query concurrency during peak hours.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Firebolt | Snowflake |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage | • Search Optimization Service for point lookups and selective queries (additional cost) • Clustering keys for data organization and automatic clustering • Materialized views • Snowflake Optima automatic indexing on Generation 2 warehouses (no additional cost) • No traditional database indexes |
| Compute tuning | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. | • Warehouse T-shirt sizing (XS to 6XL) • Multi-cluster configuration and scaling policies • Generation 1 vs Generation 2 warehouse selection • MAX\_CONCURRENCY\_LEVEL parameter tuning • Query Acceleration Service for long-running queries |
| Storage format | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support | Columnar micro-partitioned & compressed storage |
| Table-level partition & pruning techniques | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. | • Data automatically divided into micro-partitions • Automatic pruning at micro-partition level • Clustering keys for data organization with automatic clustering • Snowflake Optima provides additional automatic pruning optimization on Gen2 warehouses |
| Result cache | Yes, results and sub-results cache with transactional spoiling. | Yes |
| Warm cache (SSD) | Yes, at indexed data-range level granularity | Yes, at micro-partition level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes, including Lambda expressions and native nested array structures | Yes |
| Vector Search and AI Capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference | AI integration through Cortex AI and Snowpark ML |
| Query Optimizations | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry | • Search Optimization Service for point lookups (additional cost) • Query Acceleration Service (QAS) for long-running and unpredictable workloads • Snowflake Optima automatic optimization on Generation 2 warehouses (no additional cost) • Automatic clustering with background maintenance • Materialized views with automatic refresh • Result cache (24hrs) • Cost-based optimization with dynamic query rewriting |
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges, that are significantly smaller. In addition, Firebolt lets user accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
**Snowflake** typically comes on top for most queries when it comes to performance in public TPC-based benchmarks when compared to BigQuery and Redshift, but only marginally. Its micro partition storage approach effectively scans less data compared to larger partitions. The ability to isolate workloads over the decoupled storage & compute architecture lets you avoid competition for resources compared to multi-tenant shared resource solutions, and the ability to increase warehouse sizes can often enhance performance (for a higher price), but not always linearly. Snowflake's recently released "Search optimization service" delivers index-like behavior for point queries, but comes at an additional cost.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Firebolt | Snowflake |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications | • Sub-second to seconds response times at TB+ scale with proper clustering and optimization • Enhanced by Query Acceleration Service and Search Optimization Service • Generation 2 warehouses provide \~2x performance improvement over Generation 1 • Snowflake Optima provides automatic optimization |
| Enterprise BI | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Multi-cloud deployment options with consistent experience • Strong SQL compliance and wide ecosystem support • Zero-copy data sharing capabilities |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings | • Multi-cluster warehouses support thousands of concurrent users with auto-scaling • Individual warehouses support many concurrent queries (not limited to 8 concurrent queries) • Sub-second to seconds response times with proper optimization • Generation 2 warehouses provide significant performance improvements for high-concurrency workloads • AI integration through Cortex AI |
| Ad hoc | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer | • Excellent for ad-hoc with decoupled storage/compute • Auto-scaling and instant compute provisioning • Minimal predefined optimization required • Query Acceleration Service handles unpredictable workloads automatically • Snowflake Optima provides automatic optimization for recurring patterns |
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It's great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
**Snowflake** is a well rounded general purpose cloud data warehouse, that can also span beyond traditional BI & Analytics use cases into Ad-Hoc and ML use cases. Thanks to the flexible decoupeld storage & compute architecture that allows you to isolate and control the amount of compute per workload, it's possible to tackle a broad spectrum of workloads. However, like its close siblings Redshift & BigQuery, it struggles to deliver low-latency query performance at scale, making it a lesser fit for operational use cases and customer-facing data apps.
# Redshift vs ClickHouse (/comparison/redshift-vs-clickhouse)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Redshift | ClickHouse |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | RA3 instances enable separation of compute and storage, but limited workload isolation compared to other platforms | Yes – SharedMergeTree engine in ClickHouse Cloud enables full separation of storage and compute, with compute-compute separation through Warehouses feature (introduced 2025) allowing multiple isolated compute services sharing the same data |
| Supported cloud infrastructure | AWS only | AWS, GCP, Azure, cloud service and on-premises |
| Isolated tenancy – option for dedicated resources | • Isolated tenant & resources • Runs in your VPC | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client in cloud |
| Control vs abstraction of compute | • Configurable cluster size • Configurable compute types | Configurable cluster size and compute types in ClickHouse Cloud with granular control over nodes (1-128 nodes) and node characteristics. Warehouses feature enables multiple isolated read-only compute environments. |
| Self-hosted and hybrid deployment options | Limited hybrid options with Redshift Serverless | Self-managed deployments available with full control over infrastructure |
| ACID Compliance and Transactions | ACID compliant at table level with some limitations on concurrent operations | Limited ACID compliance with MergeTree engine family. |
**Redshift** has the oldest architecture, being the first Cloud DW in the group. Its architecture wasn't designed to separate storage & compute. While it now has RA3 nodes which allow you to scale compute and only cache the data you need locally, all compute still operates together. You cannot separate and isolate different workloads over the same data, which puts it behind other decoupled storage/compute architectures. Redshift runs as an isolated tenant per customer, and unlike other cloud data warehouses, it is deployed in your VPC. Redshift offers a serverless option which is based on an abstracted unit called Redshift Processing Unit (RPU) ranging from 8 to 512 in increments of 8. Each RPU provides 2 vCPU and 16GB RAM. Thus, 8 RPU is equivalent to 16 vCPU / 128GB RAM. The minimum RPU is 8.
**ClickHouse** was originally developed at Yandex, the Russian search engine, as an OLAP engine for low latency analytics. It was built as an on-premise solution with coupled storage & compute, and a large variety of tuning options in the form of indexes and and merge trees. ClickHouse's architecture is famous for its focus on performance and low-latency queries. The tradeoff is that it is considered very difficult to work with. SQL support is very limited, and tuning/running it requires significant engineering resources.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
\| Feature | |
\| --- | |
\| Elasticity – Scaling for larger data volumes and faster queries | |
\| Elasticity – Scaling for higher concurrency | |
**Redshift** is limited in scale because even with RA3, it cannot distribute different workloads across clusters. While it can scale to up to 10 clusters automatically to support query concurrency, it can only handle a maximum of 50 queued queries across all clusters by default.
**ClickHouse** doesn't offer any dedicated scaling features or mechanisms. While it can deliver linearly scalable performance for some types of queries, scaling itself has to be done manually. Hardware is self-managed in ClickHouse. This means that to scale you would have to provision a cluster and migrate.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Redshift | ClickHouse |
| ------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | None | • Primary indexes • Skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • MergeTree indexes • Incremental Materialized views |
| Compute tuning | Choice over number of nodes and their type | Configurable compute resources in cloud offering |
| Storage format | Columnar & compressed storage (RA3 nodes) | Columnar, supports sorted, compressed, encoded & sparsely indexed files with native Apache Iceberg support. |
| Table-level partition & pruning techniques | • No table partitions • User-defined distribution & sort keys are used to optimize for speed | Partitioning by date/time and custom partitions with MergeTree indexes. |
| Result cache | Yes | Yes, results cache with TTL and query condition cache. |
| Warm cache (SSD) | Only with RA3 nodes at partition-level granularity | Yes, at indexed data-range level granularity |
| Support for semi-structured data & JSON functions within SQL | Limited | Yes, including Lambda expressions and native JSON data type (GA in v25.3) |
| Vector Search and AI Capabilities | Limited AI capabilities – primarily through integrations | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference |
| Query Optimizations | • Basic query optimizer • Materialized views • Result caching • ANALYZE for table statistics • Workload management (WLM) • Automated materialized views (AutoMV) • AI-driven scaling (Serverless preview) | • Primary indexes (ORDER BY) • Data skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • Materialized views • Projections • PREWHERE optimization • Query analysis tools • Automatic global join reordering (v25.9) • Enhanced JSON query optimization • Streaming secondary indices |
**Redshift** does provide a result cache for accelerating repetitive query workloads and also has more tuning options than some others. But it does not deliver much faster compute performance than other cloud data warehouses in benchmarks. Sort keys can be used to optimize performance, but their contribution is limited. There is no support for indexes, and low-latency analytics at large data volumes is hard to achieve. Because Redshift decoupling of storage & compute is limited compared to other cloud data warehouses, it doesn't support isolating workloads, which means performance can degrade under pressure and competition for resources.
**ClickHouse** is famous for being one of the fastest local runtimes ever built for OLAP workloads. Its columnar storage, compression and indexing capabilities make it a consistent leader in benchmarks. Its lack of support for standard SQL and lack of query optimizer means that it's less suitable for traditional BI workloads, and more suitable for engineering managed workloads. While fast, it requires a lot of tuning and optimization.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Redshift | ClickHouse |
| ---------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Seconds to tens of seconds load times at 100s of GB scale • Can achieve faster performance with Concurrency Scaling and proper tuning | • Sub-second load times at TB+ scale with proper indexing • ClickHouse Cloud reduces engineering overhead with managed service • Proven low-latency performance (120ms at 2500 QPS in benchmarks) • Purpose-built for low-latency OLAP and real-time analytics |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Strong AWS ecosystem integration | • Growing ecosystem with 50+ integrations including major BI tools • Native MySQL protocol support enables broad BI tool compatibility • Strong SQL compliance with PostgreSQL compatibility • Best suited for modern analytical workloads and engineering-managed use cases |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries • Seconds-level response times typical • Automatic scaling for burst workloads • Limited AI application support | • Sub-second response times at TB+ scale • Supports 1000 concurrent users per replica • Strong price-performance on customer-facing applications • Native vector search and embeddings |
| Ad hoc | • Performance dependent on predefined distribution & sort keys • Elastic Resize enables adding compute resources • Typically subset of data loaded for ad-hoc analysis | • Good for ad-hoc queries with ClickHouse Cloud's separated storage/compute architecture • Join optimizations enable more query complexity • Strong sampling capabilities (TABLESAMPLE) for exploratory analysis • Resource management through user quotas prevents query interference • Materialized views offer performance improvements for common aggregation patterns, ad-hoc users specify directly in SQL |
**Redshift** was originally designed to support traditional internal BI reporting and dashboard use cases for analysts. As such, it is typically used as a general-purpose Enterprise data warehouse. With deep integrations into the AWS ecosystem, it can also leverage AWS ML service, making it also useful for ML projects. However, given the coupling of storage & compute, and the difficulty in delivering low-latency analytics at scale, it is less suited for operational use cases and customer-facing use cases like Data Apps. The coupling of storage and compute, together with the need to predefine sort & dist keys for optimal performance, make it challenging to use for Ad-Hoc analytics.
**ClickHouse** was not designed to be a data warehouse, but rather a low-latency query execution runtime. Managing it typically requires significant engineering overhead. Hence, it's a good fit for engineering managed operational use cases and customer-facing data apps, where low latency matters. It is not a good fit for a general purpose data warehouse, nor for Ad-Hoc analytics or ELT.
# Redshift vs Databricks (/comparison/redshift-vs-databricks)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Databricks | Redshift | Firebolt |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Separation of storage and compute | Yes | RA3 instances enable separation of compute and storage, but limited workload isolation compared to other platforms | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. |
| Supported cloud infrastructure | AWS, Azure, GCP. Marketplaces and BYOC | AWS only | AWS (GCP coming soon) & anywhere (Firebolt Core) |
| Isolated tenancy – option for dedicated resources | • Control plane in Databricks account • Data plane in customer VPC (optional) • Storage in customer VPC • Serverless SQL runs in Databricks account with private connectivity | • Isolated tenant & resources • Runs in your VPC | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client |
| Control vs abstraction of compute | • Configurable clusters and instance types • Serverless SQL warehouses (GA 2025) run in Databricks account with private connectivity, no public IPs • Pro/Classic warehouses run in customer VPC | • Configurable cluster size • Configurable compute types | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. |
| Self-hosted and hybrid deployment options | • Databricks on customer cloud accounts • Unity Catalog for hybrid governance | Limited hybrid options with Redshift Serverless | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS |
| ACID Compliance and Transactions | • ACID transactions with Delta Lake • Time travel and versioning • Concurrent read/write operations | ACID compliant at table level with some limitations on concurrent operations | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications |
**Redshift** has the oldest architecture, being the first Cloud DW in the group. Its architecture wasn't designed to separate storage & compute. While it now has RA3 nodes which allow you to scale compute and only cache the data you need locally, all compute still operates together. You cannot separate and isolate different workloads over the same data, which puts it behind other decoupled storage/compute architectures. Redshift runs as an isolated tenant per customer, and unlike other cloud data warehouses, it is deployed in your VPC. Redshift offers a serverless option which is based on an abstracted unit called Redshift Processing Unit (RPU) ranging from 8 to 512 in increments of 8. Each RPU provides 2 vCPU and 16GB RAM. Thus, 8 RPU is equivalent to 16 vCPU / 128GB RAM. The minimum RPU is 8.
**Databricks** was built by the founders of Spark as an analytics platform to support machine learning use cases. It leverages the Spark framework to process data residing in a data lake and is supported on AWS, GCP and Azure. Databricks coined the marketing term "Lakehouse '' architecture to illustrate the unification of data lake and data warehouse use cases. Customers still manage Spark clusters that process data residing in a Delta lake. Conversion of data to Delta Lake format is required to leverage the functionality of Delta Lake. Databricks Sql is a relatively new addition to simplify access to data stored in a data lake.
**Firebolt** is built on a natively decoupled storage & compute architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A "Firebolt Engine" can be granularly configured across # of nodes and different CPU/RAM/SSD combinations.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Databricks | Redshift | Firebolt |
| --------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Autoscaling clusters based on workload demand. Serverless SQL warehouses provide near-instant scaling (2-6 seconds startup) | Available via Elastic Resize – slow and limited, downtime required | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. |
| Elasticity – Scaling for higher concurrency | • 10 concurrent queries per cluster limit • Scales up to 40 clusters per warehouse (400 total concurrent queries) • Serverless SQL warehouses provide near-instant autoscaling • Pro/Classic warehouses take several minutes to provision new clusters • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on complexity | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. |
**Redshift** is limited in scale because even with RA3, it cannot distribute different workloads across clusters. While it can scale to up to 10 clusters automatically to support query concurrency, it can only handle a maximum of 50 queued queries across all clusters by default.
**Databricks** allow for autoscaling of clusters based on utilization. Additionally, increasing concurrency associated with a sql endpoint can be accomplished through the addition of clusters. Query concurrency per cluster is maxed at 10. However, scaling with additional clusters for concurrency is possible. Databricks provides a choice of instance types.
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn't instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Databricks | Redshift | Firebolt |
| ------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | None | None | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage |
| Compute tuning | Choice of cluster type, node types including SSD-optimized instances. Serverless provides automatic resource allocation with Intelligent Workload Management (IWM) | Choice over number of nodes and their type | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. |
| Storage format | • Delta Lake format with Liquid Clustering (February 2025 – replaces Z-ordering and traditional partitioning) • Cannot use Liquid Clustering alongside Z-ordering on same table • Allows for sorted data in Delta Lake • Requires Optimize to maintain ordering | Columnar & compressed storage (RA3 nodes) | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support |
| Table-level partition & pruning techniques | • Table level partitioning • Liquid Clustering for improved query performance and reduced data skew (February 2025) • Z-ordering (legacy, replaced by Liquid Clustering) • Periodic optimization of storage required | • No table partitions • User-defined distribution & sort keys are used to optimize for speed | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. |
| Result cache | Multi-layered caching: local in-memory cache per cluster plus remote result cache (serverless only) that persists across all warehouses in workspace | Yes | Yes, results and sub-results cache with transactional spoiling. |
| Warm cache (SSD) | Yes. Delta cache for data read by queries at file level granularity | Only with RA3 nodes at partition-level granularity | Yes, at indexed data-range level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes | Limited | Yes, including Lambda expressions and native nested array structures |
| Vector Search and AI Capabilities | • MLflow integration and Databricks ML platform • Native vector search in Delta Lake (Vector Search) • AI and ML workloads optimized | Limited AI capabilities – primarily through integrations | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference |
| Query Optimizations | • Photon engine (C++ vectorized engine providing 3-8x average speedups, maximum speedups over 10x) • Automated stats collection (January 2025) enables cost-based optimization • Predictive I/O for faster point lookups and data updates • Liquid Clustering (February 2025) • Intelligent Workload Management (IWM) with AI-powered resource allocation • Delta cache • Materialized views support | • Basic query optimizer • Materialized views • Result caching • ANALYZE for table statistics • Workload management (WLM) • Automated materialized views (AutoMV) • AI-driven scaling (Serverless preview) | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry |
**Redshift** does provide a result cache for accelerating repetitive query workloads and also has more tuning options than some others. But it does not deliver much faster compute performance than other cloud data warehouses in benchmarks. Sort keys can be used to optimize performance, but their contribution is limited. There is no support for indexes, and low-latency analytics at large data volumes is hard to achieve. Because Redshift decoupling of storage & compute is limited compared to other cloud data warehouses, it doesn't support isolating workloads, which means performance can degrade under pressure and competition for resources.
**Databricks** is designed to leverage the Spark framework for processing large volumes of data. It leverages compressed Parquet files in a Delta Lake. To reduce the amount of data processed, it uses data pruning on partitions and Parquet file metadata. Databricks does not provide any indexes.
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges that are significantly smaller. In addition, Firebolt lets users accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Databricks | Redshift | Firebolt |
| ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second to seconds load times at TB+ scale • Enhanced by Photon engine (3-8x average speedups) and Delta cache • Serverless SQL warehouses provide rapid startup (2-6 seconds) • Performance depends on cluster configuration | • Seconds to tens of seconds load times at 100s of GB scale • Can achieve faster performance with Concurrency Scaling and proper tuning | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications |
| Enterprise BI | • Strong for data science and ML workloads • Unified analytics platform approach • Growing traditional BI integrations • Serverless SQL warehouses improve accessibility • Delta sharing capabilities | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Strong AWS ecosystem integration | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 10 concurrent queries per cluster, scaling to 400 total concurrent queries per warehouse • Real-world performance degradation typically occurs at 50-150 concurrent queries depending on workload complexity • Serverless provides near-instant autoscaling • Photon engine delivers 3-8x performance improvements • Strong ML and AI platform integration | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries • Seconds-level response times typical • Automatic scaling for burst workloads • Limited AI application support | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings |
| Ad hoc | • Excellent for ad-hoc with decoupled storage/compute • Serverless SQL warehouses provide instant provisioning • Intelligent Workload Management handles unpredictable workloads automatically • Strong for exploratory data analysis and ML workloads • Automated stats collection improves query planning | • Performance dependent on predefined distribution & sort keys • Elastic Resize enables adding compute resources • Typically subset of data loaded for ad-hoc analysis | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer |
**Redshift** was originally designed to support traditional internal BI reporting and dashboard use cases for analysts. As such, it is typically used as a general-purpose Enterprise data warehouse. With deep integrations into the AWS ecosystem, it can also leverage AWS ML service, making it also useful for ML projects. However, given the coupling of storage & compute, and the difficulty in delivering low-latency analytics at scale, it is less suited for operational use cases and customer-facing use cases like Data Apps. The coupling of storage and compute, together with the need to predefine sort & dist keys for optimal performance, make it challenging to use for Ad-Hoc analytics.
**Databricks** is a mature Spark based platform proven for processing streaming data. It is widely used for Machine Learning use cases by data scientists through the use of integrated notebooks. From a low latency query perspective, while it offers features like Delta Cache, it does not provide specialized indexes that can deliver low latency queries.
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It's great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
# Redshift vs Druid (2025) (/comparison/redshift-vs-druid)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Redshift | Druid |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | RA3 instances enable separation of compute and storage, but limited workload isolation compared to other platforms | No |
| Supported cloud infrastructure | AWS only | Can be installed anywhere |
| Isolated tenancy – option for dedicated resources | • Isolated tenant & resources • Runs in your VPC | Single tenant |
| Control vs abstraction of compute | • Configurable cluster size • Configurable compute types | • Complex configuration of compute tier with multiple role-specific nodes • Configurable node count • Configurable compute types (virtual machines or kubernetes) |
| Self-hosted and hybrid deployment options | Limited hybrid options with Redshift Serverless | Self-managed deployment required |
| ACID Compliance and Transactions | ACID compliant at table level with some limitations on concurrent operations | Limited ACID support with eventual consistency |
**Redshift** has the oldest architecture, being the first Cloud DW in the group. Its architecture wasn't designed to separate storage & compute. While it now has RA3 nodes which allow you to scale compute and only cache the data you need locally, all compute still operates together. You cannot separate and isolate different workloads over the same data, which puts it behind other decoupled storage/compute architectures. Redshift runs as an isolated tenant per customer, and unlike other cloud data warehouses, it is deployed in your VPC. Redshift offers a serverless option which is based on an abstracted unit called Redshift Processing Unit (RPU) ranging from 8 to 512 in increments of 8. Each RPU provides 2 vCPU and 16GB RAM. Thus, 8 RPU is equivalent to 16 vCPU / 128GB RAM. The minimum RPU is 8.
**Druid** is an OLAP engine designed to provide fast real time analytics. Druid adopts a clustered architecture with servers that host various role specific processes. These processes address real time and batch ingestion, indexing, querying of historical and real time data. Apache Druid can be deployed as a virtual machine or a Kubernetes based cluster. Druid does not support a decoupled compute & storage architecture. Deep storage in the form of object storage is used to replicate data to.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Redshift | Druid |
| --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | Available via Elastic Resize – slow and limited, downtime required | Scale-up of nodes requires careful planning and downtime. Addition of new nodes for scale-out is possible |
| Elasticity – Scaling for higher concurrency | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries | Supports 100s to 100,000s queries per second (1000+ QPS) with proper configuration and scaling |
**Redshift** is limited in scale because even with RA3, it cannot distribute different workloads across clusters. While it can scale to up to 10 clusters automatically to support query concurrency, it can only handle a maximum of 50 queued queries across all clusters by default.
**Druid** provides the ability to handle fast ingest and high concurrency. Custom sizing and cluster tuning are required to balance the compute, memory, storage needs of each process within Druid and to provide high concurrency. Druid clusters can be grown by adding nodes with automatic rebalancing of storage segments assigned to nodes. Self hosted Druid on Kubernetes is an option that users leverage to simplify scaling. Additionally, Cloud based managed Druid offerings are being rolled out. However, these managed offerings are limited in scale and scaling is not granular.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Redshift | Druid |
| ------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| Indexes | None | Compressed bitmap indexes for data access and roll-ups to manage aggregations |
| Compute tuning | Choice over number of nodes and their type | On-premises, self-managed hardware. Druid requires infrastructure management and leverages commonly available instance types |
| Storage format | Columnar & compressed storage (RA3 nodes) | Columnar storage format with time-based sorting |
| Table-level partition & pruning techniques | • No table partitions • User-defined distribution & sort keys are used to optimize for speed | Restrictive time-based partitioning. Can partition based on other secondary columns |
| Result cache | Yes | Ability to support caching on broker (set to off by default) |
| Warm cache (SSD) | Only with RA3 nodes at partition-level granularity | Yes, at much larger segment level granularity |
| Support for semi-structured data & JSON functions within SQL | Limited | Recommend flattening JSON or translate to array prior to loading. No support for JSON parsing at query runtime |
| Vector Search and AI Capabilities | Limited AI capabilities – primarily through integrations | No native AI or vector search capabilities |
| Query Optimizations | • Basic query optimizer • Materialized views • Result caching • ANALYZE for table statistics • Workload management (WLM) • Automated materialized views (AutoMV) • AI-driven scaling (Serverless preview) | • Compressed bitmap indexes • Roll-up aggregations • Time-based optimization • Query optimization requires manual tuning |
**Redshift** does provide a result cache for accelerating repetitive query workloads and also has more tuning options than some others. But it does not deliver much faster compute performance than other cloud data warehouses in benchmarks. Sort keys can be used to optimize performance, but their contribution is limited. There is no support for indexes, and low-latency analytics at large data volumes is hard to achieve. Because Redshift decoupling of storage & compute is limited compared to other cloud data warehouses, it doesn't support isolating workloads, which means performance can degrade under pressure and competition for resources.
**Druid** provides high performance through columnar storage format, parallel processing, bitmap indexes and roll-ups. Druid, however, recommends a denormalized data model for performance needs. Join operations in Druid are a relatively new feature with various limitations, especially if there is a need to join large datasets.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Redshift | Druid |
| ---------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Low-latency dashboards | • Seconds to tens of seconds load times at 100s of GB scale • Can achieve faster performance with Concurrency Scaling and proper tuning | • Sub-second load times optimized for time-series and real-time analytics • Built for high-concurrency interactive dashboards • Requires denormalized data model |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Strong AWS ecosystem integration | • Limited integrations with traditional Enterprise BI tools • Strong for real-time operational dashboards • Requires specialized visualization tools |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries • Seconds-level response times typical • Automatic scaling for burst workloads • Limited AI application support | • Built for high concurrency (1000+ QPS) with distributed architecture • Sub-second response times for time-series data • Optimized for real-time operational applications • No AI capabilities |
| Ad hoc | • Performance dependent on predefined distribution & sort keys • Elastic Resize enables adding compute resources • Typically subset of data loaded for ad-hoc analysis | • Not optimized for ad-hoc queries • Requires predefined roll-ups and data modeling • Limited flexibility for exploratory analysis |
**Redshift** was originally designed to support traditional internal BI reporting and dashboard use cases for analysts. As such, it is typically used as a general-purpose Enterprise data warehouse. With deep integrations into the AWS ecosystem, it can also leverage AWS ML service, making it also useful for ML projects. However, given the coupling of storage & compute, and the difficulty in delivering low-latency analytics at scale, it is less suited for operational use cases and customer-facing use cases like Data Apps. The coupling of storage and compute, together with the need to predefine sort & dist keys for optimal performance, make it challenging to use for Ad-Hoc analytics.
**Druid** is designed as an OLAP engine to provide fast access to aggregations that are run against large volumes of data. Druid is typically used for customer facing analytics and streaming data processing. Druid is used as an add-on with other data warehousing products that are efficient at scaling, joining, and filtering large volumes of data. It is not a suitable option for data warehouse replacement.
# Snowflake vs Athena vs Firebolt (2025) (/comparison/snowflake-vs-athena)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Snowflake | Athena | Firebolt |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Separation of storage and compute | Yes | Yes, serverless with optional provisioned capacity. Workloads can be isolated through Workgroups and Capacity Reservations | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. |
| Supported cloud infrastructure | AWS, Azure, GCP with full feature parity across all three major clouds | AWS only | AWS (GCP coming soon) & anywhere (Firebolt Core) |
| Isolated tenancy – option for dedicated resources | • Multi-tenant pooled resources • Isolated tenancy available via VPS tier | • Multi-tenant pooled resources by default • Dedicated compute resources available via Provisioned Capacity • VPC endpoint connections supported | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client |
| Control vs abstraction of compute | • Configurable warehouse sizes (XS to 6XL) • Multi-cluster warehouses with auto-scaling • Choice between Generation 1 and Generation 2 standard warehouses • MAX\_CONCURRENCY\_LEVEL parameter for resource allocation | • Serverless by default with no infrastructure control • Optional Provisioned Capacity allows dedicated DPU allocation (minimum 24 DPUs) • Two pricing models: on-demand ($5/TB scanned) or provisioned ($0.30/DPU-hour) | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. |
| Self-hosted and hybrid deployment options | Snowflake for Government Cloud and private cloud options available | No self-hosted options – serverless only | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS |
| ACID Compliance and Transactions | Full ACID compliance with Time Travel and zero-copy cloning capabilities | No ACID compliance – eventual consistency model | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications |
**Snowflake** was one of the first decoupled storage and compute architectures, making it the first to have nearly unlimited compute scale and workload isolation, and horizontal user scalability. It runs on AWS, Azure and GCP. It is multi-tenant over shared resources in nature and requires you to move data out of your VPC and into the Snowflake cloud. "Virtual Private Snowflake" (VPS) is its highest-priced tier, and can run a dedicated isolated version of Snowflake. Its virtual warehouses can be T-shirt sized along an XS/S/M…/4XL axis, where each discrete T-shirt size is bundled with fixed HW properties that are abstracted from the users. Snowflake has recently added support for Snowflake managed Iceberg tables.
**Athena** is serverless and built on a decoupled storage and compute architecture that queries data directly in S3, without the need to ingest/copy the data. It runs in multi-tenancy with shared resources. Users do not have control over the compute resources Athena chooses to allocate per query from the shared resource pool. For folks requiring additional or dedicated resources, they can reserve dedicated processing capacity in the form of Data Processing Units (DPU), with each DPU providing 4 vCPU and 16 GB RAM. RPU allocation ranges from 24 - 1000 per region.
**Firebolt** is built on a natively decoupled storage & compute architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A "Firebolt Engine" can be granularly configured across # of nodes and different CPU/RAM/SSD combinations.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Snowflake | Athena | Firebolt |
| --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | • Instant warehouse resize (XS to 6XL) with no downtime • Multi-cluster auto-scaling • Generation 2 warehouses provide \~2x performance improvement over Generation 1 | • Fully abstracted on-demand scaling • Provisioned Capacity allows manual scaling of DPUs for predictable performance • Capacity reservations can be adjusted with minimum 1-hour billing periods | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. |
| Elasticity – Scaling for higher concurrency | • Single warehouse supports many concurrent queries (MAX\_CONCURRENCY\_LEVEL=8 controls resource allocation per query, not query limit) • Multi-cluster warehouses enable thousands of concurrent queries with auto-scaling • Unlimited virtual warehouses can be created | • Default limit of 25 concurrent DML queries and 20 DDL queries (adjustable via service quotas) • Provisioned Capacity enables higher concurrency with dedicated DPUs • Query queuing available when capacity is exceeded | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. |
**Snowflake** scales very well both for data volumes and query concurrency. The decoupled storage/compute architecture supports resizing clusters without downtime, and in addition, supports auto-scaling horizontally for higher query concurrency during peak hours.
**Athena** is a shared multi-tenant resource, with no guarantees on the amount or availability of the resources allocated for your queries. From a data volume perspective, it can scale to large volumes, but large data volumes can suffer from very long run times and frequent time outs. Query concurrency is maxed at 20. If scalability is a top priority, Athena is probably not the best choice.
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn't instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Snowflake | Athena | Firebolt |
| ------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Search Optimization Service for point lookups and selective queries (additional cost) • Clustering keys for data organization and automatic clustering • Materialized views • Snowflake Optima automatic indexing on Generation 2 warehouses (no additional cost) • No traditional database indexes | No traditional indexes – relies on partition pruning and data organization in S3. Uses columnar formats and compression for optimization | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage |
| Compute tuning | • Warehouse T-shirt sizing (XS to 6XL) • Multi-cluster configuration and scaling policies • Generation 1 vs Generation 2 warehouse selection • MAX\_CONCURRENCY\_LEVEL parameter tuning • Query Acceleration Service for long-running queries | • No compute tuning in on-demand mode • Provisioned Capacity allows DPU allocation control (4 vCPU and 16GB RAM per DPU) • Minimum 24 DPUs with scaling in 4-DPU increments | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. |
| Storage format | Columnar micro-partitioned & compressed storage | Supports multiple formats: Parquet, ORC, Avro, JSON, CSV, TSV on S3. Native support for open table formats including Apache Iceberg, Apache Hudi, and Delta Lake | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support |
| Table-level partition & pruning techniques | • Data automatically divided into micro-partitions • Automatic pruning at micro-partition level • Clustering keys for data organization with automatic clustering • Snowflake Optima provides additional automatic pruning optimization on Gen2 warehouses | • User-defined table-level partitions with Hive-style partitioning • Pruning at partition level • Partition projection for advanced performance optimization • Supports open table formats with built-in partitioning | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. |
| Result cache | Yes | Query result caching for up to 30 days with configurable retention. Results reuse supported across workgroups | Yes, results and sub-results cache with transactional spoiling. |
| Warm cache (SSD) | Yes, at micro-partition level granularity | No local caching – queries data directly from S3. Relies on S3's performance characteristics and intelligent tiering | Yes, at indexed data-range level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes | Yes, comprehensive JSON support including Lambda expressions, array functions, and native nested data handling | Yes, including Lambda expressions and native nested array structures |
| Vector Search and AI Capabilities | AI integration through Cortex AI and Snowpark ML | No native AI or vector search capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference |
| Query Optimizations | • Search Optimization Service for point lookups (additional cost) • Query Acceleration Service (QAS) for long-running and unpredictable workloads • Snowflake Optima automatic optimization on Generation 2 warehouses (no additional cost) • Automatic clustering with background maintenance • Materialized views with automatic refresh • Result cache (24hrs) • Cost-based optimization with dynamic query rewriting | • Cost-based optimizer (CBO) in Athena engine v3 • Query result caching (up to 30 days) • Partition projection for advanced optimization • CTAS for precomputed queries • Join reordering and aggregation pushdown • Automatic parallel query execution • Support for columnar formats (Parquet, ORC) • Integration with AWS Glue Data Catalog | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry |
**Snowflake** typically comes on top for most queries when it comes to performance in public TPC-based benchmarks when compared to BigQuery and Redshift, but only marginally. Its micro partition storage approach effectively scans less data compared to larger partitions. The ability to isolate workloads over the decoupled storage & compute architecture lets you avoid competition for resources compared to multi-tenant shared resource solutions, and the ability to increase warehouse sizes can often enhance performance (for a higher price), but not always linearly. Snowflake's recently released "Search optimization service" delivers index-like behavior for point queries, but comes at an additional cost.
**Athena**, (and Presto) are designed to query data where it is, sacrificing storage-compute optimizations. This makes it very convenient for easy and immediate querying but at the expense of performance. This typically puts Athena behind cloud data warehouses in terms of performance. But Athena still does relatively well in performance benchmarks, especially when external storage is managed by experts. While it supports partitions, there is no support for indexing, and together with the fact that resources are pooled from a shared multi-tenant service, low-latency and consistent performance are not Athena's sweet spot. A cloud data warehouse be more performant better than Athena in most cases.
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges that are significantly smaller. In addition, Firebolt lets users accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Snowflake | Athena | Firebolt |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second to seconds response times at TB+ scale with proper clustering and optimization • Enhanced by Query Acceleration Service and Search Optimization Service • Generation 2 warehouses provide \~2x performance improvement over Generation 1 • Snowflake Optima provides automatic optimization | • Seconds to minutes response times for interactive dashboards • Performance varies based on data partitioning, file formats, and query optimization • Provisioned Capacity can improve consistency for dashboard workloads • Best suited for analytical dashboards rather than sub-second operational dashboards | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Multi-cloud deployment options with consistent experience • Strong SQL compliance and wide ecosystem support • Zero-copy data sharing capabilities | • Good integration with AWS ecosystem BI tools (QuickSight, etc.) • Standard SQL compatibility enables most BI tool connections • Cost-effective for variable workloads and ad-hoc analytics • JDBC/ODBC drivers support enterprise BI tools • Limited advanced BI features compared to dedicated data warehouses | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Multi-cluster warehouses support thousands of concurrent users with auto-scaling • Individual warehouses support many concurrent queries (not limited to 8 concurrent queries) • Sub-second to seconds response times with proper optimization • Generation 2 warehouses provide significant performance improvements for high-concurrency workloads • AI integration through Cortex AI | • Default concurrency limits (25 DML/20 DDL queries) may require service quota increases • Provisioned Capacity enables higher concurrency with dedicated resources • Seconds-level response times typical • Cost-effective for customer-facing analytics with proper optimization • Best suited for analytical rather than operational workloads • No native AI capabilities | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings |
| Ad hoc | • Excellent for ad-hoc with decoupled storage/compute • Auto-scaling and instant compute provisioning • Minimal predefined optimization required • Query Acceleration Service handles unpredictable workloads automatically • Snowflake Optima provides automatic optimization for recurring patterns | • Purpose-built for ad-hoc analytics on data lakes • Serverless with zero infrastructure management • Direct querying of S3 data without ETL • Cost-effective pay-per-query model ideal for exploratory analysis • Strong support for multiple data formats and federated queries • Apache Spark integration for advanced analytics | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer |
**Snowflake** is a well rounded general purpose cloud data warehouse, that can also span beyond traditional BI & Analytics use cases into Ad-Hoc and ML use cases. Thanks to the flexible decoupeld storage & compute architecture that allows you to isolate and control the amount of compute per workload, it's possible to tackle a broad spectrum of workloads. However, like its close siblings Redshift & BigQuery, it struggles to deliver low-latency query performance at scale, making it a lesser fit for operational use cases and customer-facing data apps.
**Athena** is a great choice for Ad-Hoc analytics. You can keep the data where it is, and start querying without worrying about hardware or pretty much anything else, given that Athena is serverless and takes care of everything behind the scenes. However, it is not a great fit when you need consistent and fast query performance, and/or high concurrency. This is why it is typically not the best choice for operational and customer-facing applications. It can be also easily and flexibly used for batch processing, which is often leveraged for ML use cases.
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It's great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
# Snowflake vs BigQuery (/comparison/snowflake-vs-bigquery)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Snowflake | BigQuery | Firebolt |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Separation of storage and compute | Yes | Fully serverless with complete separation of compute (Dremel) and storage (Colossus), powered by Jupiter network and Borg orchestration | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. |
| Supported cloud infrastructure | AWS, Azure, GCP with full feature parity across all three major clouds | Google Cloud only | AWS (GCP coming soon) & anywhere (Firebolt Core) |
| Isolated tenancy – option for dedicated resources | • Multi-tenant pooled resources • Isolated tenancy available via VPS tier | • Multi-tenant pooled resources • VPC Service Controls provide enhanced security and connectivity isolation to customer VPCs • Cross-region disaster recovery for enterprise workloads | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client |
| Control vs abstraction of compute | • Configurable warehouse sizes (XS to 6XL) • Multi-cluster warehouses with auto-scaling • Choice between Generation 1 and Generation 2 standard warehouses • MAX\_CONCURRENCY\_LEVEL parameter for resource allocation | Fully serverless with no control over compute resources – BigQuery automatically allocates computing resources as needed with intelligent workload management and dynamic slot allocation | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. |
| Self-hosted and hybrid deployment options | Snowflake for Government Cloud and private cloud options available | No self-hosted options – fully managed service only | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS |
| ACID Compliance and Transactions | Full ACID compliance with Time Travel and zero-copy cloning capabilities | Limited ACID support – eventual consistency model with some transactional capabilities | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications |
**Snowflake** was one of the first decoupled storage and compute architectures, making it the first to have nearly unlimited compute scale and workload isolation, and horizontal user scalability. It runs on AWS, Azure and GCP. It is multi-tenant over shared resources in nature and requires you to move data out of your VPC and into the Snowflake cloud. "Virtual Private Snowflake" (VPS) is its highest-priced tier, and can run a dedicated isolated version of Snowflake. Its virtual warehouses can be T-shirt sized along an XS/S/M…/4XL axis, where each discrete T-shirt size is bundled with fixed HW properties that are abstracted from the users. Snowflake has recently added support for Snowflake managed Iceberg tables.
**BigQuery** was one of the first decoupled storage and compute architectures. It is a unique piece of engineering and not a typical data warehouse in part because it started as an on-demand serverless query engine. It runs in multi-tenancy with shared resources, allocated as "slots" which represent a virtual CPU that executes SQL. BigQuery determines how many slots a query requires, without the ability of the user to control it. BigQuery can be priced on a $/TB scanned basis or through slot reservations. A slot in BigQuery is logically equivalent to 0.5 vCPU and 0.5GB of RAM. There are multiple models to allocate slots in BigQuery.
**Firebolt** is built on a natively decoupled storage & compute architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A "Firebolt Engine" can be granularly configured across # of nodes and different CPU/RAM/SSD combinations.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Snowflake | BigQuery | Firebolt |
| --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | • Instant warehouse resize (XS to 6XL) with no downtime • Multi-cluster auto-scaling • Generation 2 warehouses provide \~2x performance improvement over Generation 1 | Fully automated serverless scaling – BigQuery automatically determines resource allocation and scales to petabytes without user intervention. Can dynamically burst beyond baseline slot allocations for performance optimization | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. |
| Elasticity – Scaling for higher concurrency | • Single warehouse supports many concurrent queries (MAX\_CONCURRENCY\_LEVEL=8 controls resource allocation per query, not query limit) • Multi-cluster warehouses enable thousands of concurrent queries with auto-scaling • Unlimited virtual warehouses can be created | Dynamic concurrency management with query queueing supporting up to 1,000 interactive queries and 20,000 batch queries per project per region. Automatic fair scheduling and slot distribution across workloads | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. |
**Snowflake** scales very well both for data volumes and query concurrency. The decoupled storage/compute architecture supports resizing clusters without downtime, and in addition, supports auto-scaling horizontally for higher query concurrency during peak hours.
**BigQuery** scales very well to large data volumes, and automatically assigns more compute resources when needed behind the scenes, in the form of "slots". BigQuery works either in an "on-demand pricing model", where slot assignment is completely in the hands of BigQuery and the state of the shared resource pool, or in "flat-rate pricing model" where slots are reserved in advanced. With reserved slots there is more control over compute resources, thus making scaling more predictable. Concurrency is limited to 100 users by default.
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn't instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Snowflake | BigQuery | Firebolt |
| ------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Search Optimization Service for point lookups and selective queries (additional cost) • Clustering keys for data organization and automatic clustering • Materialized views • Snowflake Optima automatic indexing on Generation 2 warehouses (no additional cost) • No traditional database indexes | Search indexes (GA) for efficient text search optimization on STRING, JSON, and array columns. Support for LOG\_ANALYZER, NO\_OP\_ANALYZER, and PATTERN\_ANALYZER with column-level granularity for improved query performance and cost efficiency | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage |
| Compute tuning | • Warehouse T-shirt sizing (XS to 6XL) • Multi-cluster configuration and scaling policies • Generation 1 vs Generation 2 warehouse selection • MAX\_CONCURRENCY\_LEVEL parameter tuning • Query Acceleration Service for long-running queries | Serverless architecture with automatic resource optimization – no manual tuning required. Intelligent workload management with AI-powered resource allocation and dynamic slot distribution | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. |
| Storage format | Columnar micro-partitioned & compressed storage | Columnar & compressed storage (Capacitor format) with support for open table formats including Apache Iceberg, Delta Lake, and Hudi. Intelligent tiering with automatic long-term storage cost reduction after 90 days | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support |
| Table-level partition & pruning techniques | • Data automatically divided into micro-partitions • Automatic pruning at micro-partition level • Clustering keys for data organization with automatic clustering • Snowflake Optima provides additional automatic pruning optimization on Gen2 warehouses | • Automatic table organization with intelligent micro-partitioning • Clustering keys for data organization • Automatic partition pruning optimization • Supports time-based and custom partitioning strategies | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. |
| Result cache | Yes | Yes, with cross-user result caching and intelligent cache management for up to 24 hours | Yes, results and sub-results cache with transactional spoiling. |
| Warm cache (SSD) | Yes, at micro-partition level granularity | BI Engine provides in-memory caching and acceleration for frequently accessed data and dashboards | Yes, at indexed data-range level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes | Yes, including advanced JSON functions, Lambda expressions, and native support for nested and repeated fields | Yes, including Lambda expressions and native nested array structures |
| Vector Search and AI Capabilities | AI integration through Cortex AI and Snowpark ML | • BigQuery ML for in-database machine learning • Vertex AI integration • Natural language querying with Gemini AI • Limited vector search capabilities | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference |
| Query Optimizations | • Search Optimization Service for point lookups (additional cost) • Query Acceleration Service (QAS) for long-running and unpredictable workloads • Snowflake Optima automatic optimization on Generation 2 warehouses (no additional cost) • Automatic clustering with background maintenance • Materialized views with automatic refresh • Result cache (24hrs) • Cost-based optimization with dynamic query rewriting | • Advanced query optimizer with Dremel engine • Search indexes with column-level granularity • Materialized views with smart refresh and automatic query rewriting • BI Engine in-memory acceleration • Gemini AI-powered query optimization and natural language querying • Cross-user result caching • Automatic partitioning and clustering optimization • Cost-based optimization with intelligent workload management | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry |
**Snowflake** typically comes on top for most queries when it comes to performance in public TPC-based benchmarks when compared to BigQuery and Redshift, but only marginally. Its micro partition storage approach effectively scans less data compared to larger partitions. The ability to isolate workloads over the decoupled storage & compute architecture lets you avoid competition for resources compared to multi-tenant shared resource solutions, and the ability to increase warehouse sizes can often enhance performance (for a higher price), but not always linearly. Snowflake's recently released "Search optimization service" delivers index-like behavior for point queries, but comes at an additional cost.
**BigQuery** lines up in benchmarks in the same ballpark as other cloud data warehouses but does come in consistently last in most queries. Beyond implementing according to best practices, there is little you can do to accelerate BigQuery performance, as it determines the amount of resources (slots) the query needs for you. BigQuery can be used together with "BigQuery BI Engine" for lower latency analytics. However, BI Engine is limited in terms of scale because it runs in memory. Its maximum capacity is 100GB.
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges that are significantly smaller. In addition, Firebolt lets users accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Snowflake | BigQuery | Firebolt |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second to seconds response times at TB+ scale with proper clustering and optimization • Enhanced by Query Acceleration Service and Search Optimization Service • Generation 2 warehouses provide \~2x performance improvement over Generation 1 • Snowflake Optima provides automatic optimization | • Sub-second to seconds response times at TB+ scale with BI Engine acceleration • Search indexes and materialized views provide significant performance improvements for dashboard queries • Intelligent caching reduces query costs for repeated dashboard access • Dynamic concurrency management supports high user loads | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Multi-cloud deployment options with consistent experience • Strong SQL compliance and wide ecosystem support • Zero-copy data sharing capabilities | • Mature and comprehensive Enterprise DW feature set with native Google Cloud ecosystem integration • Strong integration with Looker, Looker Studio, and major BI tools • Gemini AI integration for natural language insights • Cross-cloud analytics capabilities • Advanced governance with Dataplex Universal Catalog | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Multi-cluster warehouses support thousands of concurrent users with auto-scaling • Individual warehouses support many concurrent queries (not limited to 8 concurrent queries) • Sub-second to seconds response times with proper optimization • Generation 2 warehouses provide significant performance improvements for high-concurrency workloads • AI integration through Cortex AI | • Dynamic concurrency supporting 1,000+ interactive queries with intelligent queuing and fair scheduling • Sub-second response times with BI Engine acceleration and search indexes • Serverless architecture eliminates infrastructure management overhead • Advanced caching and materialized views optimize repeated queries • AI-powered optimization for customer-facing applications • BigQuery ML for AI workloads | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings |
| Ad hoc | • Excellent for ad-hoc with decoupled storage/compute • Auto-scaling and instant compute provisioning • Minimal predefined optimization required • Query Acceleration Service handles unpredictable workloads automatically • Snowflake Optima provides automatic optimization for recurring patterns | • Excellent for ad-hoc analytics with serverless architecture requiring zero infrastructure management • Intelligent query optimization handles unpredictable workloads automatically • Gemini AI integration enables natural language querying for business users • Advanced JSON support and schema inference enable flexible data exploration • Cross-cloud analytics capabilities for federated queries | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer |
**Snowflake** has broader support for use cases beyond traditional reporting and dashboards. Its decoupled storage and compute architecture enables you to isolate different workloads to meet SLAs, and it also supports high user concurrency. But Snowflake also does not provide interactive or ad hoc query performance because of inefficient data access along with a lack of extensive indexing and query optimization. Snowflake also cannot support streaming or low latency ingestion below one minute ingestion intervals. All of these limitations exclude Snowflake from many operational use cases and most customer-facing applications that require second-level performance.
**BigQuery**, like Snowflake, has broader support for use cases beyond reporting and dashboards. You can isolate workloads by assigning each workload to different reserved slots. Unlike Snowflake, Redshift, or Athena, BigQuery also supports low latency streaming. But like these other three technologies. BigQuery also lacks the performance to support interactive or ad hoc queries at scale. This eliminates BigQuery from being a great option for many operational and customer-facing use cases where the users demand a few seconds of wait at worst, which translates to sub-second query times for the data warehouse.
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It's great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
# Snowflake vs ClickHouse (/comparison/snowflake-vs-clickhouse)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Snowflake | ClickHouse |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Separation of storage and compute | Yes | Yes – SharedMergeTree engine in ClickHouse Cloud enables full separation of storage and compute, with compute-compute separation through Warehouses feature (introduced 2025) allowing multiple isolated compute services sharing the same data |
| Supported cloud infrastructure | AWS, Azure, GCP with full feature parity across all three major clouds | AWS, GCP, Azure, cloud service and on-premises |
| Isolated tenancy – option for dedicated resources | • Multi-tenant pooled resources • Isolated tenancy available via VPS tier | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client in cloud |
| Control vs abstraction of compute | • Configurable warehouse sizes (XS to 6XL) • Multi-cluster warehouses with auto-scaling • Choice between Generation 1 and Generation 2 standard warehouses • MAX\_CONCURRENCY\_LEVEL parameter for resource allocation | Configurable cluster size and compute types in ClickHouse Cloud with granular control over nodes (1-128 nodes) and node characteristics. Warehouses feature enables multiple isolated read-only compute environments. |
| Self-hosted and hybrid deployment options | Snowflake for Government Cloud and private cloud options available | Self-managed deployments available with full control over infrastructure |
| ACID Compliance and Transactions | Full ACID compliance with Time Travel and zero-copy cloning capabilities | Limited ACID compliance with MergeTree engine family. |
**Snowflake** was one of the first decoupled storage and compute architectures, making it the first to have nearly unlimited compute scale and workload isolation, and horizontal user scalability. It runs on AWS, Azure and GCP. It is multi-tenant over shared resources in nature and requires you to move data out of your VPC and into the Snowflake cloud. "Virtual Private Snowflake" (VPS) is its highest-priced tier, and can run a dedicated isolated version of Snowflake. Its virtual warehouses can be T-shirt sized along an XS/S/M…/4XL axis, where each discrete T-shirt size is bundled with fixed HW properties that are abstracted from the users. Snowflake has recently added support for Snowflake managed Iceberg tables.
**ClickHouse** was originally developed at Yandex, the Russian search engine, as an OLAP engine for low latency analytics. It was built as an on-premise solution with coupled storage & compute, and a large variety of tuning options in the form of indexes and and merge trees. ClickHouse's architecture is famous for its focus on performance and low-latency queries. The tradeoff is that it is considered very difficult to work with. SQL support is very limited, and tuning/running it requires significant engineering resources.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Snowflake | ClickHouse |
| --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | • Instant warehouse resize (XS to 6XL) with no downtime • Multi-cluster auto-scaling • Generation 2 warehouses provide \~2x performance improvement over Generation 1 | Automatic horizontal and vertical scaling in ClickHouse Cloud with SharedMergeTree architecture. Manual scaling for self-managed deployments with cluster rebalancing capabilities |
| Elasticity – Scaling for higher concurrency | • Single warehouse supports many concurrent queries (MAX\_CONCURRENCY\_LEVEL=8 controls resource allocation per query, not query limit) • Multi-cluster warehouses enable thousands of concurrent queries with auto-scaling • Unlimited virtual warehouses can be created | Supports high concurrency with proper resource allocation and configuration. Vertical auto-scaling and horizontal manual scaling. Additional warehouses can idle to zero billing. Primary service always on in multi-warehouse configurations. |
**Snowflake** scales very well both for data volumes and query concurrency. The decoupled storage/compute architecture supports resizing clusters without downtime, and in addition, supports auto-scaling horizontally for higher query concurrency during peak hours.
**ClickHouse** doesn't offer any dedicated scaling features or mechanisms. While it can deliver linearly scalable performance for some types of queries, scaling itself has to be done manually. Hardware is self-managed in ClickHouse. This means that to scale you would have to provision a cluster and migrate.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today. While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Snowflake | ClickHouse |
| ------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Search Optimization Service for point lookups and selective queries (additional cost) • Clustering keys for data organization and automatic clustering • Materialized views • Snowflake Optima automatic indexing on Generation 2 warehouses (no additional cost) • No traditional database indexes | • Primary indexes • Skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • MergeTree indexes • Incremental Materialized views |
| Compute tuning | • Warehouse T-shirt sizing (XS to 6XL) • Multi-cluster configuration and scaling policies • Generation 1 vs Generation 2 warehouse selection • MAX\_CONCURRENCY\_LEVEL parameter tuning • Query Acceleration Service for long-running queries | Configurable compute resources in cloud offering |
| Storage format | Columnar micro-partitioned & compressed storage | Columnar, supports sorted, compressed, encoded & sparsely indexed files with native Apache Iceberg support. |
| Table-level partition & pruning techniques | • Data automatically divided into micro-partitions • Automatic pruning at micro-partition level • Clustering keys for data organization with automatic clustering • Snowflake Optima provides additional automatic pruning optimization on Gen2 warehouses | Partitioning by date/time and custom partitions with MergeTree indexes. |
| Result cache | Yes | Yes, results cache with TTL and query condition cache. |
| Warm cache (SSD) | Yes, at micro-partition level granularity | Yes, at indexed data-range level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes | Yes, including Lambda expressions and native JSON data type (GA in v25.3) |
| Vector Search and AI Capabilities | AI integration through Cortex AI and Snowpark ML | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference |
| Query Optimizations | • Search Optimization Service for point lookups (additional cost) • Query Acceleration Service (QAS) for long-running and unpredictable workloads • Snowflake Optima automatic optimization on Generation 2 warehouses (no additional cost) • Automatic clustering with background maintenance • Materialized views with automatic refresh • Result cache (24hrs) • Cost-based optimization with dynamic query rewriting | • Primary indexes (ORDER BY) • Data skipping indexes (minmax, set, bloom filters, ngrambf\_v1, tokenbf\_v1) • Materialized views • Projections • PREWHERE optimization • Query analysis tools • Automatic global join reordering (v25.9) • Enhanced JSON query optimization • Streaming secondary indices |
**Snowflake** typically comes on top for most queries when it comes to performance in public TPC-based benchmarks when compared to BigQuery and Redshift, but only marginally. Its micro partition storage approach effectively scans less data compared to larger partitions. The ability to isolate workloads over the decoupled storage & compute architecture lets you avoid competition for resources compared to multi-tenant shared resource solutions, and the ability to increase warehouse sizes can often enhance performance (for a higher price), but not always linearly. Snowflake's recently released "Search optimization service" delivers index-like behavior for point queries, but comes at an additional cost.
**ClickHouse** is famous for being one of the fastest local runtimes ever built for OLAP workloads. Its columnar storage, compression and indexing capabilities make it a consistent leader in benchmarks. Its lack of support for standard SQL and lack of query optimizer means that it's less suitable for traditional BI workloads, and more suitable for engineering managed workloads. While fast, it requires a lot of tuning and optimization.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Snowflake | ClickHouse |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second to seconds response times at TB+ scale with proper clustering and optimization • Enhanced by Query Acceleration Service and Search Optimization Service • Generation 2 warehouses provide \~2x performance improvement over Generation 1 • Snowflake Optima provides automatic optimization | • Sub-second load times at TB+ scale with proper indexing • ClickHouse Cloud reduces engineering overhead with managed service • Proven low-latency performance (120ms at 2500 QPS in benchmarks) • Purpose-built for low-latency OLAP and real-time analytics |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Multi-cloud deployment options with consistent experience • Strong SQL compliance and wide ecosystem support • Zero-copy data sharing capabilities | • Growing ecosystem with 50+ integrations including major BI tools • Native MySQL protocol support enables broad BI tool compatibility • Strong SQL compliance with PostgreSQL compatibility • Best suited for modern analytical workloads and engineering-managed use cases |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Multi-cluster warehouses support thousands of concurrent users with auto-scaling • Individual warehouses support many concurrent queries (not limited to 8 concurrent queries) • Sub-second to seconds response times with proper optimization • Generation 2 warehouses provide significant performance improvements for high-concurrency workloads • AI integration through Cortex AI | • Sub-second response times at TB+ scale • Supports 1000 concurrent users per replica • Strong price-performance on customer-facing applications • Native vector search and embeddings |
| Ad hoc | • Excellent for ad-hoc with decoupled storage/compute • Auto-scaling and instant compute provisioning • Minimal predefined optimization required • Query Acceleration Service handles unpredictable workloads automatically • Snowflake Optima provides automatic optimization for recurring patterns | • Good for ad-hoc queries with ClickHouse Cloud's separated storage/compute architecture • Join optimizations enable more query complexity • Strong sampling capabilities (TABLESAMPLE) for exploratory analysis • Resource management through user quotas prevents query interference • Materialized views offer performance improvements for common aggregation patterns, ad-hoc users specify directly in SQL |
**Snowflake** is a well rounded general purpose cloud data warehouse, that can also span beyond traditional BI & Analytics use cases into Ad-Hoc and ML use cases. Thanks to the flexible decoupeld storage & compute architecture that allows you to isolate and control the amount of compute per workload, it's possible to tackle a broad spectrum of workloads. However, like its close siblings Redshift & BigQuery, it struggles to deliver low-latency query performance at scale, making it a lesser fit for operational use cases and customer-facing data apps.
**ClickHouse** was not designed to be a data warehouse, but rather a low-latency query execution runtime. Managing it typically requires significant engineering overhead. Hence, it's a good fit for engineering managed operational use cases and customer-facing data apps, where low latency matters. It is not a good fit for a general purpose data warehouse, nor for Ad-Hoc analytics or ELT.
# Snowflake vs Redshift vs Firebolt (2025) (/comparison/snowflake-vs-redshift)
## Architecture [#architecture]
The biggest difference among cloud data warehouses are whether they separate storage and compute, how much they isolate data and compute, and what clouds they can run on.
| Feature | Snowflake | Redshift | Firebolt |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Separation of storage and compute | Yes | RA3 instances enable separation of compute and storage, but limited workload isolation compared to other platforms | Yes, separation of storage and metadata as well as compute from compute with full workload isolation. |
| Supported cloud infrastructure | AWS, Azure, GCP with full feature parity across all three major clouds | AWS only | AWS (GCP coming soon) & anywhere (Firebolt Core) |
| Isolated tenancy – option for dedicated resources | • Multi-tenant pooled resources • Isolated tenancy available via VPS tier | • Isolated tenant & resources • Runs in your VPC | • Multi-tenant metadata layer • Isolated tenancy for compute & storage per client |
| Control vs abstraction of compute | • Configurable warehouse sizes (XS to 6XL) • Multi-cluster warehouses with auto-scaling • Choice between Generation 1 and Generation 2 standard warehouses • MAX\_CONCURRENCY\_LEVEL parameter for resource allocation | • Configurable cluster size • Configurable compute types | Uses engine abstraction: • Each engine has configurable cluster size (1-128 nodes) for horizontal scaling. • Configurable compute family (compute vs storage optimized) and type (XS, S, M, L, XL) for vertical scaling • Number of clusters for concurrency (auto)scaling. Provides full workload isolation across engines. |
| Self-hosted and hybrid deployment options | Snowflake for Government Cloud and private cloud options available | Limited hybrid options with Redshift Serverless | • Firebolt Core: Forever free, self-hosted edition with full query engine capabilities • Same performance and features as managed service • Deploy anywhere: local laptop, cloud, datacenter, Kubernetes • Production-grade distributed architecture • No usage restrictions except building competing SaaS |
| ACID Compliance and Transactions | Full ACID compliance with Time Travel and zero-copy cloning capabilities | ACID compliant at table level with some limitations on concurrent operations | • Full ACID compliance with snapshot isolation • Multi-statement transactions supported • Strong consistency across all operations • Supports concurrent reads and writes • Transactional integrity for data applications |
**Snowflake** was one of the first decoupled storage and compute architectures, making it the first to have nearly unlimited compute scale and workload isolation, and horizontal user scalability. It runs on AWS, Azure and GCP. It is multi-tenant over shared resources in nature and requires you to move data out of your VPC and into the Snowflake cloud. “Virtual Private Snowflake” (VPS) is its highest-priced tier, and can run a dedicated isolated version of Snowflake. Its virtual warehouses can be T-shirt sized along an XS/S/M…/4XL axis, where each discrete T-shirt size is bundled with fixed HW properties that are abstracted from the users. Snowflake has recently added support for Snowflake managed Iceberg tables.
**Redshift** has the oldest architecture, being the first Cloud DW in the group. Its architecture wasn’t designed to separate storage & compute. While it now has RA3 nodes which allow you to scale compute and only cache the data you need locally, all compute still operates together. You cannot separate and isolate different workloads over the same data, which puts it behind other decoupled storage/compute architectures. Redshift runs as an isolated tenant per customer, and unlike other cloud data warehouses, it is deployed in your VPC. Redshift offers a serverless option which is based on an abstracted unit called Redshift Processing Unit (RPU) ranging from 8 to 512 in increments of 8. Each RPU provides 2 vCPU and 16GB RAM. Thus, 8 RPU is equivalent to 16 vCPU / 128GB RAM. The minimum RPU is 8.
**Firebolt** is built on a natively decoupled storage & compute architecture, on AWS only. Data has to be copied outside of your VPC into the Firebolt, where both your compute and data run in a dedicated and isolated tenant. A “Firebolt Engine” can be granularly configured across # of nodes and different CPU/RAM/SSD combinations.
## Scalability [#scalability]
There are three big differences among data warehouses and query engines that limit scalability: decoupled storage and compute, dedicated resources, and continuous ingestion.
| Feature | Snowflake | Redshift | Firebolt |
| --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Elasticity – Scaling for larger data volumes and faster queries | • Instant warehouse resize (XS to 6XL) with no downtime • Multi-cluster auto-scaling • Generation 2 warehouses provide \~2x performance improvement over Generation 1 | Available via Elastic Resize – slow and limited, downtime required | Granular cluster resize with node types, number of nodes and number of clusters. Zero downtime. |
| Elasticity – Scaling for higher concurrency | • Single warehouse supports many concurrent queries (MAX\_CONCURRENCY\_LEVEL=8 controls resource allocation per query, not query limit) • Multi-cluster warehouses enable thousands of concurrent queries with auto-scaling • Unlimited virtual warehouses can be created | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries | A single engine can handle hundreds of concurrent queries. Engines auto-scale the number of clusters up and down base on resource usage thresholds. Idle engines scale down to zero billing. |
**Snowflake** scales very well both for data volumes and query concurrency. The decoupled storage/compute architecture supports resizing clusters without downtime, and in addition, supports auto-scaling horizontally for higher query concurrency during peak hours.
**Redshift** is limited in scale because even with RA3, it cannot distribute different workloads across clusters. While it can scale to up to 10 clusters automatically to support query concurrency, it can only handle a maximum of 50 queued queries across all clusters by default.
**Firebolt** can handle the largest data volumes and concurrency on a single comparable cluster size, thanks to its superior hardware efficiency. Thanks to its decoupled storage & compute architecture it scales very well to large data volumes. However, resizing an engine size isn’t instant and requires orchestration if avoiding downtime is necessary. A single Firebolt engine can support hundreds of concurrent queries, avoiding the need to scale out for most use cases. Scaling horizontally for even higher concurrency is manual.
## Performance [#performance]
Performance is the biggest challenge with most data warehouses today.
While decoupled storage and compute architectures improved scalability and simplified administration, for most data warehouses it introduced two bottlenecks; storage, and compute. Most modern cloud data warehouses fetch entire partitions over the network instead of just fetching the specific data needed for each query. While many invest in caching, most do not invest heavily in query optimization. Most vendors also have not improved continuous ingestion or semi-structured data analytics performance, both of which are needed for operational and customer-facing use cases.
| Feature | Snowflake | Redshift | Firebolt |
| ------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Indexes | • Search Optimization Service for point lookups and selective queries (additional cost) • Clustering keys for data organization and automatic clustering • Materialized views • Snowflake Optima automatic indexing on Generation 2 warehouses (no additional cost) • No traditional database indexes | None | • Sparse primary indexes • Aggregating indexes • Join indexes • Optimizer driven index usage |
| Compute tuning | • Warehouse T-shirt sizing (XS to 6XL) • Multi-cluster configuration and scaling policies • Generation 1 vs Generation 2 warehouse selection • MAX\_CONCURRENCY\_LEVEL parameter tuning • Query Acceleration Service for long-running queries | Choice over number of nodes and their type | SQL defined engines. Control number of nodes, node family and type per cluster, with one or more clusters per engine. Multiple engines isolate workloads. |
| Storage format | Columnar micro-partitioned & compressed storage | Columnar & compressed storage (RA3 nodes) | Columnar, sorted & compressed & sparsely indexed storage (F3 – Firebolt File Format) with native Apache Iceberg support |
| Table-level partition & pruning techniques | • Data automatically divided into micro-partitions • Automatic pruning at micro-partition level • Clustering keys for data organization with automatic clustering • Snowflake Optima provides additional automatic pruning optimization on Gen2 warehouses | • No table partitions • User-defined distribution & sort keys are used to optimize for speed | • User-defined table-level partitions are optional. • Data is automatically sorted, compressed and indexed into F3 format. • Pruning at indexed data-range level. |
| Result cache | Yes | Yes | Yes, results and sub-results cache with transactional spoiling. |
| Warm cache (SSD) | Yes, at micro-partition level granularity | Only with RA3 nodes at partition-level granularity | Yes, at indexed data-range level granularity |
| Support for semi-structured data & JSON functions within SQL | Yes | Limited | Yes, including Lambda expressions and native nested array structures |
| Vector Search and AI Capabilities | AI integration through Cortex AI and Snowpark ML | Limited AI capabilities – primarily through integrations | • Native vector search capabilities and embeddings • MCP Server for AI driven analytics • Natural Language to SQL • SQL based Inference |
| Query Optimizations | • Search Optimization Service for point lookups (additional cost) • Query Acceleration Service (QAS) for long-running and unpredictable workloads • Snowflake Optima automatic optimization on Generation 2 warehouses (no additional cost) • Automatic clustering with background maintenance • Materialized views with automatic refresh • Result cache (24hrs) • Cost-based optimization with dynamic query rewriting | • Basic query optimizer • Materialized views • Result caching • ANALYZE for table statistics • Workload management (WLM) • Automated materialized views (AutoMV) • AI-driven scaling (Serverless preview) | • Primary indexes, aggregating indexes, join indexes, sparse indexes • Sub-plan result caching • F3 storage format optimization • Automatic query optimizer with aggressive pruning • Late column materialization • Query analysis tools based on execution telemetry |
**Snowflake** typically comes on top for most queries when it comes to performance in public TPC-based benchmarks when compared to BigQuery and Redshift, but only marginally. Its micro partition storage approach effectively scans less data compared to larger partitions. The ability to isolate workloads over the decoupled storage & compute architecture lets you avoid competition for resources compared to multi-tenant shared resource solutions, and the ability to increase warehouse sizes can often enhance performance (for a higher price), but not always linearly. Snowflake’s recently released “Search optimization service” delivers index-like behavior for point queries, but comes at an additional cost.
**Redshift** does provide a result cache for accelerating repetitive query workloads and also has more tuning options than some others. But it does not deliver much faster compute performance than other cloud data warehouses in benchmarks. Sort keys can be used to optimize performance, but their contribution is limited. There is no support for indexes, and low-latency analytics at large data volumes is hard to achieve. Because Redshift decoupling of storage & compute is limited compared to other cloud data warehouses, it doesn’t support isolating workloads, which means performance can degrade under pressure and competition for resources.
**Firebolt** is the fastest when it comes to query performance when compared to cloud data warehouses and services like Athena. Its unique approach to storage and indexing results in highly aggressive data pruning that scans dramatically less data compared to other technologies. While other technologies scan partitions or micro-partitions, Firebolt works with indexed data ranges that are significantly smaller. In addition, Firebolt lets users accelerate queries further with multiple index types (Aggregating index, Join index), and using its decoupled storage & compute architecture workloads can be easily isolated to guarantee consistent performance.
## Use cases [#use-cases]
There are a host of different analytics use cases that can be supported by a data warehouse. Look at your legacy technologies and their workloads, as well as the new possible use cases, and figure out which ones you will need to support in the next few years.
| Feature | Snowflake | Redshift | Firebolt |
| ---------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Low-latency dashboards | • Sub-second to seconds response times at TB+ scale with proper clustering and optimization • Enhanced by Query Acceleration Service and Search Optimization Service • Generation 2 warehouses provide \~2x performance improvement over Generation 1 • Snowflake Optima provides automatic optimization | • Seconds to tens of seconds load times at 100s of GB scale • Can achieve faster performance with Concurrency Scaling and proper tuning | • 120ms query latency at 4000 QPS (FireScale benchmark 2025) • Sub-second performance at TB+ scale with proper indexing • Built for AI-driven analytics, dashboards, and real-time analytic applications |
| Enterprise BI | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Multi-cloud deployment options with consistent experience • Strong SQL compliance and wide ecosystem support • Zero-copy data sharing capabilities | • Mature and comprehensive Enterprise DW feature set • Extensive integrations with Enterprise BI ecosystem • Strong AWS ecosystem integration | • Growing ecosystem with focus on modern BI tools • Strong SQL compliance with PostgreSQL • Wire level compatibility drives expansion to PostgreSQL BI and ETL ecosystem |
| Data Apps and AI Applications (Customer-facing low-latency high concurrency) | • Multi-cluster warehouses support thousands of concurrent users with auto-scaling • Individual warehouses support many concurrent queries (not limited to 8 concurrent queries) • Sub-second to seconds response times with proper optimization • Generation 2 warehouses provide significant performance improvements for high-concurrency workloads • AI integration through Cortex AI | • 5 concurrent queries per WLM queue by default (up to 8 queues) • Concurrency Scaling enables thousands of concurrent queries • Seconds-level response times typical • Automatic scaling for burst workloads • Limited AI application support | • 120ms latency at 4000+ QPS proven performance at TB+ scale • Supports hundreds to thousands of concurrent queries on single engine • Price-performance leader (8x better than Snowflake, 18x vs Redshift) • Purpose-built for AI agents and data-intensive applications • Native vector search and embeddings |
| Ad hoc | • Excellent for ad-hoc with decoupled storage/compute • Auto-scaling and instant compute provisioning • Minimal predefined optimization required • Query Acceleration Service handles unpredictable workloads automatically • Snowflake Optima provides automatic optimization for recurring patterns | • Performance dependent on predefined distribution & sort keys • Elastic Resize enables adding compute resources • Typically subset of data loaded for ad-hoc analysis | • Excellent performance out-of-the-box with engine optimized for star and snowflake joins and aggregations • Self learning query plan optimizer • Full workload isolation prevents ad-hoc complexity from affecting real-time workloads • Aggregating indexes are automatically used by optimizer |
**Snowflake** is a well rounded general purpose cloud data warehouse, that can also span beyond traditional BI & Analytics use cases into Ad-Hoc and ML use cases. Thanks to the flexible decoupeld storage & compute architecture that allows you to isolate and control the amount of compute per workload, it’s possible to tackle a broad spectrum of workloads. However, like its close siblings Redshift & BigQuery, it struggles to deliver low-latency query performance at scale, making it a lesser fit for operational use cases and customer-facing data apps.
**Redshift** was originally designed to support traditional internal BI reporting and dashboard use cases for analysts. As such, it is typically used as a general-purpose Enterprise data warehouse. With deep integrations into the AWS ecosystem, it can also leverage AWS ML service, making it also useful for ML projects. However, given the coupling of storage & compute, and the difficulty in delivering low-latency analytics at scale, it is less suited for operational use cases and customer-facing use cases like Data Apps. The coupling of storage and compute, together with the need to predefine sort & dist keys for optimal performance, make it challenging to use for Ad-Hoc analytics.
**Firebolt** stands out by being the fastest cloud data warehouse when compared to Snowflake, Redshift, BigQuery and Athena. It’s great for delivering sub-second analytics at scale, while remaining hardware efficient and high concurrency friendly. This makes it a great choice for operational use cases and customer-facing data apps. Given that it is not as feature-rich and integration rich as the more mature data warehouses makes it a lesser fit for a general-purpose Enterprise data warehouse. It is also not the best fit for ad-hoc use cases, because of the need to predefine indexing at the table level.
# E-Commerce Analytics Primer (/free-sample-datasets/e-commerce)
This primer introduces data engineers to sample use cases for data-driven decision-making in
e-commerce. It walks through ingesting data and exploring performance optimizations using the
Firebolt Cloud Data Warehouse, with data from the Open Customer Data Platform (CDP) Project.
## Understanding the e-commerce data model [#understanding-the-e-commerce-data-model]
A well-thought-out schema is critical for efficiently managing the diverse data in an e-commerce
warehouse — user interactions, product details, and transaction records. This dataset uses a
single-table design:
| Property | Data Type | Description |
| --------------- | -------------- | --------------------------------------- |
| `event_time` | TIMESTAMPTZ | Time in UTC |
| `event_type` | TEXT | Customer event (view / cart / purchase) |
| `product_id` | BIGINT | ID of a product |
| `category_id` | TEXT | Product's category ID |
| `category_code` | TEXT | Product's category code |
| `brand` | TEXT | Product brand |
| `price` | NUMERIC(38, 9) | Price of a product |
| `user_id` | TEXT | Permanent user ID |
| `user_session` | TEXT | Temporary user session ID |
The schema supports user identification, event tracking, product trends, and session details for
personalization and segmentation.
## Working with the e-commerce dataset [#working-with-the-e-commerce-dataset]
The data comes from a Kaggle dataset loaded into a public Amazon S3 bucket
(`firebolt-sample-datasets-public-us-east-1`). It spans seven months of activity (October 2019 –
April 2020), roughly 32GB uncompressed as CSV, converted to Parquet for ingestion. Once ingested it
is approximately **412 million records** — 52GB uncompressed, 21GB compressed.
Start by creating a database and engine:
```sql
CREATE DATABASE ecommercedb WITH DESCRIPTION = 'ECommerce Analytics Primer';
CREATE ENGINE ecommerceEngine;
START ENGINE ecommerceEngine;
USE ENGINE ecommerceEngine;
USE ecommercedb;
```
Create the fact table:
```sql
CREATE FACT TABLE IF NOT EXISTS "ecommerce" (
"event_time" TIMESTAMPTZ NOT NULL,
"event_type" TEXT NOT NULL,
"product_id" BIGINT NOT NULL,
"category_id" TEXT NULL,
"category_code" TEXT NULL,
"brand" TEXT NULL,
"price" NUMERIC(38, 9) NULL,
"user_id" TEXT NULL,
"user_session" TEXT NULL
);
```
Ingest the data:
```sql
COPY INTO ecommerce FROM 's3://firebolt-sample-datasets-public-us-east-1/ecommerce_primer/parquet/'
WITH PATTERN='*.gz.parquet' TYPE = PARQUET;
SHOW TABLES;
```
## Sample analytics on the e-commerce dataset [#sample-analytics-on-the-e-commerce-dataset]
### Customer lifetime value (LTV) [#customer-lifetime-value-ltv]
Assess the total revenue generated by customers over their engagement with your brand.
```sql
SELECT
user_id,
SUM(price) AS total_revenue,
COUNT(DISTINCT user_session) AS total_sessions,
(SUM(price) / COUNT(DISTINCT user_session)) AS average_revenue_per_session
FROM ecommerce
WHERE event_type = 'purchase'
AND user_session IS NOT NULL
GROUP BY user_id
ORDER BY total_revenue DESC;
```
### Funnel analysis for conversion optimization [#funnel-analysis-for-conversion-optimization]
Track the customer journey from views to purchases and identify drop-off points.
```sql
SELECT
event_type,
COUNT(DISTINCT user_id) AS unique_users
FROM ecommerce
WHERE event_type IN ('view', 'cart', 'purchase')
GROUP BY event_type
ORDER BY 2 DESC;
```
### Product recommendations and cross-selling [#product-recommendations-and-cross-selling]
Analyze co-purchase data to reveal products frequently bought together.
```sql
WITH CoPurchaseCounts AS (
SELECT a.product_id AS product_A, b.product_id AS product_B, COUNT(*) AS purchase_count
FROM ecommerce a
JOIN ecommerce b ON a.user_session = b.user_session AND a.product_id < b.product_id
WHERE a.event_type = 'purchase' AND b.event_type = 'purchase'
GROUP BY a.product_id, b.product_id
)
SELECT *
FROM CoPurchaseCounts cp
WHERE purchase_count > 2
ORDER BY purchase_count DESC
LIMIT 10;
```
### Seasonal sales trends and insights [#seasonal-sales-trends-and-insights]
Analyze sales by month to understand seasonal variation.
```sql
SELECT
EXTRACT(MONTH FROM event_time) AS sales_month,
SUM(price) AS total_revenue
FROM ecommerce
WHERE event_type = 'purchase'
GROUP BY sales_month
ORDER BY sales_month;
```
### Customer segmentation for targeted marketing [#customer-segmentation-for-targeted-marketing]
Segment customers by activity level and average revenue.
```sql
WITH CustomerSegments AS (
SELECT
user_id,
COUNT(DISTINCT user_session) AS session_count,
SUM(price) AS total_revenue
FROM ecommerce
WHERE event_type = 'purchase'
GROUP BY user_id
)
SELECT
CASE
WHEN session_count <= 5 THEN 'Low Activity'
WHEN session_count <= 20 THEN 'Medium Activity'
ELSE 'High Activity'
END AS segment,
COUNT(user_id) AS user_count,
AVG(total_revenue) AS avg_revenue
FROM CustomerSegments
GROUP BY segment;
```
### Cart abandonment rate [#cart-abandonment-rate]
Identify where potential customers drop off in the purchase process.
```sql
WITH CartAbandonment AS (
SELECT event_time::DATE as event_date,
user_id,
COUNT(CASE WHEN event_type = 'cart' THEN 1 ELSE NULL END) AS cart_count,
COUNT(CASE WHEN event_type = 'purchase' THEN 1 ELSE NULL END) AS purchase_count
FROM ecommerce
WHERE event_type IN ('cart', 'purchase')
GROUP BY event_date, user_id
)
SELECT
event_date, COUNT(*) AS total_users,
SUM(CASE WHEN cart_count > 0 AND purchase_count = 0 THEN 1 ELSE 0 END) AS abandoned_carts,
(SUM(CASE WHEN cart_count > 0 AND purchase_count = 0 THEN 1 ELSE 0 END) * 100.0 / COUNT(*)) AS abandonment_rate
FROM CartAbandonment
GROUP BY event_date
ORDER BY event_date;
```
### User purchase frequency distribution [#user-purchase-frequency-distribution]
Analyze the distribution of user purchase frequencies for marketing planning.
```sql
WITH UserPurchaseFrequency AS (
SELECT
user_id,
COUNT(DISTINCT event_time::DATE) AS purchase_frequency
FROM ecommerce
WHERE event_type = 'purchase'
GROUP BY user_id
)
SELECT
purchase_frequency,
COUNT(user_id) AS user_count
FROM UserPurchaseFrequency
GROUP BY purchase_frequency
ORDER BY purchase_frequency;
```
## Optimizing performance [#optimizing-performance]
Use the `recommend_ddl` command to find an optimal primary index for the workload:
```sql
CALL recommend_ddl (
ecommerce,
(
SELECT
query_text
FROM
information_schema.engine_query_history
where query_text ilike 'select%'
and end_time > NOW() - INTERVAL '30 minutes'
)
);
```
For this workload, Firebolt recommends a primary index on `event_type` and `user_session`, which
achieves roughly 94% average data pruning. Recreate the table with that index and reload:
```sql
CREATE FACT TABLE IF NOT EXISTS "ecommerce_pi" (
"event_time" TIMESTAMPTZ NOT NULL,
"event_type" TEXT NOT NULL,
"product_id" BIGINT NOT NULL,
"category_id" TEXT NULL,
"category_code" TEXT NULL,
"brand" TEXT NULL,
"price" double precision NULL,
"user_id" TEXT NULL,
"user_session" TEXT NULL
) PRIMARY INDEX event_type, user_session;
INSERT INTO ecommerce_pi SELECT * FROM ecommerce;
```
Data analytics is a critical part of e-commerce operations today, and performance and efficiency are
essential in cloud-based analytics to avoid slower performance and cost overruns.
## Appendix: querying the data lake with external tables [#appendix-querying-the-data-lake-with-external-tables]
You can query data directly from S3 without loading it into the warehouse — useful for ad-hoc
analysis.
```sql
CREATE EXTERNAL TABLE IF NOT EXISTS ex_ecommerce (
event_time TIMESTAMPTZ NOT NULL,
event_type TEXT NOT NULL,
product_id BIGINT NOT NULL,
category_id TEXT NULL,
category_code TEXT NULL,
brand TEXT NOT NULL,
price NUMERIC(38, 9) NULL,
user_id TEXT NULL,
user_session TEXT NULL
) URL = 's3://firebolt-sample-datasets-public-us-east-1/ecommerce_primer/parquet/'
OBJECT_PATTERN = '*.gz.parquet' TYPE= (PARQUET);
```
```sql
SELECT event_type, count(*)
FROM ex_ecommerce
GROUP BY ALL;
```
Querying directly against the data lake does not leverage Firebolt's performance optimizations in
the form of columnar storage and indexes.
Spin up an engine and load this dataset yourself — Firebolt is free to try.
# Analyzing the NYC Parking Violations Dataset (/free-sample-datasets/nyc-parking-violations)
New York City's extensive parking-violation records provide rich analytical opportunities. This
guide explains the dataset schema, demonstrates loading the data into Firebolt, and provides sample
SQL queries for analysis.
## Understanding the schema [#understanding-the-schema]
The dataset contains 44 columns capturing violation details:
| Column Name | Data Type | Description |
| ----------------------------------- | --------- | ------------------------------------- |
| `summons` | Integer | Unique identifier for each violation |
| `plateid` | Text | License plate identifier |
| `registration` | Text | Vehicle registration identifier |
| `plate` | Text | License plate information |
| `issue_date` | Date | Date the violation was issued |
| `violation_code` | Integer | Code indicating the violation type |
| `vehicle_body` | Text | Vehicle body type |
| `vehicle_make` | Text | Vehicle manufacturer |
| `issuing_agency` | Text | Agency responsible for issuing |
| `street_code1` | Integer | Primary street code |
| `street_code2` | Integer | Secondary street code |
| `street_code3` | Integer | Another secondary street code |
| `vehicle_expiration` | Text | Vehicle registration expiration |
| `violation_location` | Integer | Violation location code |
| `violation_precinct` | Integer | Precinct where the violation occurred |
| `issuer_precinct` | Integer | Issuing officer precinct |
| `issuer_code` | Integer | Issuing officer code |
| `issuer_command` | Text | Issuing officer command |
| `issuer_squad` | Text | Issuing officer squad |
| `violation_time` | Text | Time of the violation |
| `time_first_observed` | Text | When the violation was first observed |
| `violation_county` | Text | County of the violation |
| `violation_infront_opposite` | Text | Position relative to the address |
| `house_integer` | Text | House number identifier |
| `street_name` | Text | Street name |
| `intersecting_street` | Text | Intersecting street name |
| `date_first_observed` | Text | Date first observed |
| `law_section` | Integer | Relevant law section |
| `sub_division` | Text | Subdivision information |
| `violation_legal_code` | Text | Legal code |
| `days_parking_in_effect` | Text | Days when parking restrictions apply |
| `from_hours_in_effect` | Text | Starting time of restrictions |
| `to_hours_in_effect` | Text | Ending time of restrictions |
| `vehicle_color` | Text | Vehicle color |
| `unregistered_vehicle` | Text | Unregistered vehicle indicator |
| `vehicle_year` | Text | Vehicle manufacture year |
| `meter_integer` | Text | Meter number |
| `feet_from_curb` | Integer | Distance from curb in feet |
| `violation_post_code` | Text | Postal code |
| `violation_description` | Text | Violation description |
| `no_standing_or_stopping_violation` | Text | Standing/stopping violation indicator |
| `hydrant_violation` | Text | Hydrant violation indicator |
| `double_parking_violation` | Text | Double parking indicator |
## Loading data into Firebolt [#loading-data-into-firebolt]
There are two ways to load the data.
### Option 1: COPY FROM [#option-1-copy-from]
Let Firebolt infer the schema and create the table automatically:
```sql
COPY INTO nyc_parkingviolations FROM 's3://firebolt-sample-datasets-public-us-east-1/nyc_sample_datasets/nycparking/parquet/'
WITH PATTERN="*.parquet" AUTO_CREATE=TRUE
TYPE=PARQUET;
```
### Option 2: External table [#option-2-external-table]
First create an external table over the Parquet files in S3:
```sql
CREATE EXTERNAL TABLE ex_nyc_parkingviolations (
summons Integer,
plateid Text,
registration Text,
plate Text,
issue_date TEXT,
violation_code Integer,
vehicle_body Text,
vehicle_make Text,
issuing_agency Text,
street_code1 Integer,
street_code2 Integer,
street_code3 Integer,
vehicle_expiration Text,
violation_location Integer,
violation_precinct Integer,
issuer_precinct Integer,
issuer_code Integer,
issuer_command Text,
issuer_squad Text,
violation_time Text,
time_first_observed Text,
violation_county Text,
violation_infront_opposit Text,
house_integer Text,
street_name Text,
intersecting_street Text,
date_first_observed Integer,
law_section Integer,
sub_division Text,
violation_legal_code Text,
days_parking_in_effect Text,
from_hours_in_effect Text,
to_hours_in_effect Text,
vehicle_color Text,
unregistered_vehicle Integer,
vehicle_year Text,
meter_integer Text,
feet_from_curb Integer,
violation_post_code Text,
violation_description Text,
no_standing_or_stopping_violation Text,
hydrant_violation Text,
double_parking_violation Text
)
URL = 's3://firebolt-sample-datasets-public-us-east-1/nyc_sample_datasets/nycparking/parquet/'
OBJECT_PATTERN = '*.parquet'
TYPE = (PARQUET);
```
Then create the internal table:
```sql
CREATE TABLE nyc_parkingviolations (
summons Integer,
plateid Text,
registration Text,
plate Text,
issue_date Date,
violation_code Integer,
vehicle_body Text,
vehicle_make Text,
issuing_agency Text,
street_code1 Integer,
street_code2 Integer,
street_code3 Integer,
vehicle_expiration Integer,
violation_location Integer,
violation_precinct Integer,
issuer_precinct Integer,
issuer_code Integer,
issuer_command Text,
issuer_squad Text,
violation_time Text,
time_first_observed Text,
violation_county Text,
violation_infront_opposit Text,
house_integer Text,
street_name Text,
intersecting_street Text,
date_first_observed Integer,
law_section Integer,
sub_division Text,
violation_legal_code Text,
days_parking_in_effect Text,
from_hours_in_effect Text,
to_hours_in_effect Text,
vehicle_color Text,
unregistered_vehicle Integer,
vehicle_year Text,
meter_integer Text,
feet_from_curb Integer,
violation_post_code Text,
violation_description Text,
no_standing_or_stopping_violation Text,
hydrant_violation Text,
double_parking_violation Text
);
```
Finally, insert from the external table, converting `issue_date` to a timestamp:
```sql
INSERT INTO nyc_parkingviolations
SELECT
summons ,
plateid ,
registration ,
plate ,
TO_TIMESTAMP(issue_date,'YYYY-MM-DD') ,
violation_code ,
vehicle_body ,
vehicle_make ,
issuing_agency ,
street_code1 ,
street_code2 ,
street_code3 ,
vehicle_expiration ,
violation_location ,
violation_precinct ,
issuer_precinct ,
issuer_code ,
issuer_command ,
issuer_squad ,
violation_time ,
time_first_observed ,
violation_county ,
violation_infront_opposit ,
house_integer ,
street_name ,
intersecting_street ,
date_first_observed ,
law_section ,
sub_division ,
violation_legal_code ,
days_parking_in_effect ,
from_hours_in_effect ,
to_hours_in_effect ,
vehicle_color ,
unregistered_vehicle ,
vehicle_year ,
meter_integer ,
feet_from_curb ,
violation_post_code ,
violation_description ,
no_standing_or_stopping_violation ,
hydrant_violation ,
double_parking_violation
FROM ex_nyc_parkingviolations;
```
## Sample SQL queries for analysis [#sample-sql-queries-for-analysis]
### Count of violations by violation code [#count-of-violations-by-violation-code]
Identify the most frequently occurring parking violations.
```sql
SELECT violation_code, violation_description, COUNT(*) AS violation_count
FROM nyc_parkingviolations
GROUP BY violation_code, violation_description
ORDER BY violation_count DESC;
```
### Violations per day [#violations-per-day]
Reveal temporal patterns in violation issuance.
```sql
SELECT issue_date, COUNT(*) AS daily_violations
FROM nyc_parkingviolations
GROUP BY issue_date
ORDER BY issue_date;
```
### Top vehicle makes with the most violations [#top-vehicle-makes-with-the-most-violations]
See which vehicle manufacturers are most frequently cited.
```sql
SELECT vehicle_make, COUNT(*) AS violation_count
FROM nyc_parkingviolations
GROUP BY vehicle_make
ORDER BY violation_count DESC
LIMIT 10;
```
### Violations by issuing agency [#violations-by-issuing-agency]
Categorize violations by the responsible enforcement agency.
```sql
SELECT issuing_agency, COUNT(*) AS violation_count
FROM nyc_parkingviolations
GROUP BY issuing_agency
ORDER BY violation_count DESC;
```
The dataset invites further exploration to uncover insights about enforcement patterns and vehicle
characteristics.
Spin up an engine and load this dataset yourself — Firebolt is free to try.
# Processing Semi-Structured Data with the NYC Philharmonic Dataset (/free-sample-datasets/nyc-philharmonic)
Not all data fits a neat tabular model. Firebolt can handle semi-structured JSON, and this tutorial
demonstrates three approaches: storing JSON as raw `TEXT`, partially flattening it into columns, or
fully extracting and flattening it into a structured table.
## Key JSON functions [#key-json-functions]
Firebolt provides these functions for working with semi-structured data:
* `JSON_POINTER_EXTRACT` — extracts the value for a key using a JSON pointer expression and an
expected data type.
* `JSON_POINTER_EXTRACT_ARRAY` — extracts an array of strings from the path provided.
* `JSON_VALUE` — converts to a scalar value.
* `UNNEST` — converts an array into a set of rows.
* Lambda functions for array processing.
## The NYC Philharmonic dataset [#the-nyc-philharmonic-dataset]
This dataset covers 175 years of NYC Philharmonic performances, with nested structures for seasons,
concerts, composers, soloists, and works. It lives in S3 at:
```
s3://firebolt-sample-datasets-public-us-east-1/nyc_sample_datasets/nycphilharmonic
```
Start by creating an external table that reads each JSON file as a single raw `TEXT` column:
```sql
CREATE EXTERNAL TABLE ex_nyc_phil (
raw_data TEXT
)
URL ='s3://firebolt-sample-datasets-public-us-east-1/nyc_sample_datasets/nycphilharmonic'
PATTERN = '*.json'
TYPE = (JSON PARSE_AS_TEXT = TRUE);
```
## Example queries [#example-queries]
### Program count [#program-count]
Count the programs in the raw JSON array.
```sql
SELECT LENGTH(JSON_POINTER_EXTRACT_ARRAY(raw_data,'/programs')) AS programs_arrays FROM ex_nyc_phil;
```
### Most popular composers [#most-popular-composers]
Unnest programs and works to rank composers by how often their works were performed.
```sql
WITH programs AS (
SELECT JSON_POINTER_EXTRACT_ARRAY(raw_data, '/programs') AS programs_arrays
FROM ex_nyc_phil),
works AS (
SELECT JSON_POINTER_EXTRACT_ARRAY(program, '/works') as works_array
FROM programs, UNNEST(programs_arrays) AS r(program))
SELECT
JSON_VALUE( JSON_POINTER_EXTRACT(work, '/composerName'))
as composer_name, count(*)
FROM works, UNNEST(works_array) AS f(work)
WHERE JSON_VALUE(JSON_POINTER_EXTRACT(work, '/composerName')) IS NOT NULL
GROUP BY ALL
ORDER BY count(*) DESC;
```
### Concert start times [#concert-start-times]
Unnest programs and concerts to see the most common concert start times.
```sql
WITH programs AS (
SELECT JSON_POINTER_EXTRACT_ARRAY(raw_data, '/programs') AS programs_arrays
FROM ex_nyc_phil
),
concerts AS (
SELECT JSON_POINTER_EXTRACT_ARRAY(program, '/concerts') as concerts_arrays
FROM programs, UNNEST(programs_arrays) AS r(program)
)
SELECT JSON_POINTER_EXTRACT(concert, '/Time') as concert_time, count(*)
FROM concerts ,UNNEST(concerts_arrays) AS f(concert)
GROUP BY ALL
ORDER BY count(*) DESC;
```
## Flattening the JSON into a table [#flattening-the-json-into-a-table]
For repeated, column-based querying, flatten the nested JSON into a structured table with a single
`CREATE TABLE AS SELECT` (CTAS). This unnests programs, concerts, works, and soloists into one wide
table:
```sql
CREATE TABLE nyc_phil AS
WITH programs AS (
SELECT JSON_POINTER_EXTRACT_ARRAY(raw_data, '/programs') AS programs_arrays
FROM ex_nyc_phil
), concerts_works AS (
SELECT
JSON_POINTER_EXTRACT(program, '/season') AS season,
JSON_POINTER_EXTRACT(program, '/orchestra') AS orchestra,
JSON_POINTER_EXTRACT_ARRAY(program, '/concerts') as concerts_array,
JSON_POINTER_EXTRACT(program, '/programID') as program_id,
JSON_POINTER_EXTRACT_ARRAY(program, '/works') as works_array
FROM programs,
UNNEST(programs_arrays) AS r(program)
), concerts_works_soloists AS (
SELECT
season,
orchestra,
JSON_VALUE(JSON_POINTER_EXTRACT(concert, '/Date'))::timestamptz as concert_date,
JSON_POINTER_EXTRACT(concert, '/eventType') as concert_event_type,
JSON_POINTER_EXTRACT(concert, '/Venue') as concert_venue,
JSON_POINTER_EXTRACT(concert, '/Location') as concert_location,
JSON_POINTER_EXTRACT(concert, '/Time') as concert_time,
program_id,
JSON_POINTER_EXTRACT(work, '/workTitle') as work_title,
JSON_POINTER_EXTRACT(work, '/ID') as work_id,
JSON_POINTER_EXTRACT(work, '/conductorName') as conduct_name,
JSON_POINTER_EXTRACT(work, '/composerName') as composer_name,
CASE WHEN JSON_POINTER_EXTRACT_ARRAY(work, '/soloists') = [] THEN ['No soloists'] ELSE JSON_POINTER_EXTRACT_ARRAY(work, '/soloists') END as soloists_array
FROM concerts_works,
UNNEST (concerts_array) as f(concert),
UNNEST (works_array) as p(work)
)
SELECT
JSON_VALUE(season) as season,
JSON_VALUE(orchestra)as orchestra,
concert_date,
JSON_VALUE(concert_event_type) as concert_event_type,
JSON_VALUE(concert_venue) as concert_venue,
JSON_VALUE(concert_location)as concert_location,
JSON_VALUE(concert_time) as concert_time,
JSON_VALUE(program_id) as program_id,
JSON_VALUE(work_title) as work_title,
JSON_VALUE(work_id) as work_id,
JSON_VALUE( conduct_name) as conduct_name,
JSON_VALUE(composer_name) as composer_name,
JSON_POINTER_EXTRACT(soloist, '/soloistName') as soloist_name,
JSON_POINTER_EXTRACT(soloist, '/soloistRoles') as soloist_roles,
JSON_POINTER_EXTRACT(soloist, '/soloistInstrument') as soloist_instrument
FROM concerts_works_soloists, UNNEST (soloists_array) AS t(soloist);
```
Once flattened, you can query concert dates, venues, composers, conductors, and soloists directly as
columns.
Spin up an engine and load this dataset yourself — Firebolt is free to try.
# Analyzing the NYC Restaurant Inspections Dataset (/free-sample-datasets/nyc-restaurants)
New York City's Department of Health publishes the results of every restaurant inspection. This
tutorial covers the schema, loading the data into Firebolt, and sample SQL queries for extracting
insights from the records.
## Dataset schema [#dataset-schema]
The dataset contains 27 columns:
| Column Name | Data Type | Description |
| ----------------------- | --------- | -------------------------------------- |
| `camis` | INTEGER | Unique restaurant ID |
| `dba` | TEXT | Business name ("doing business as") |
| `boro` | TEXT | Borough |
| `building` | TEXT | Building number |
| `street` | TEXT | Street |
| `zipcode` | TEXT | ZIP code |
| `phone` | TEXT | Phone number |
| `cuisine_description` | TEXT | Cuisine type |
| `inspection_date` | TEXT | Date of inspection |
| `action` | TEXT | Action taken as a result of inspection |
| `violation_code` | TEXT | Violation code |
| `violation_description` | TEXT | Violation description |
| `critical_flag` | TEXT | Whether the violation was critical |
| `score` | NUMERIC | Inspection score |
| `grade` | TEXT | Letter grade |
| `grade_date` | DATE | Date the grade was issued |
| `record_date` | DATE | Date the record was created |
| `inspection_type` | TEXT | Type of inspection |
| `latitude` | NUMERIC | Latitude |
| `longitude` | NUMERIC | Longitude |
| `community_board` | TEXT | Community board |
| `council_district` | TEXT | Council district |
| `census_tract` | TEXT | Census tract |
| `bin` | TEXT | Building identification number |
| `bbl` | TEXT | Borough-block-lot identifier |
| `nta` | TEXT | Neighborhood tabulation area |
| `location_point1` | TEXT | Location point |
## Loading data into Firebolt [#loading-data-into-firebolt]
There are two ways to load the data.
### Option 1: COPY FROM [#option-1-copy-from]
Let Firebolt infer the schema and create the table automatically:
```sql
COPY INTO nyc_restaurant_inspections FROM 's3://firebolt-sample-datasets-public-us-east-1/nyc_sample_datasets/nyc_restaurant_inspections/parquet/'
WITH PATTERN="*.parquet" AUTO_CREATE=TRUE TYPE=PARQUET;
```
### Option 2: External table [#option-2-external-table]
Create an external table over the Parquet files in S3:
```sql
CREATE EXTERNAL TABLE ex_nyc_restaurant_inspections (
camis INTEGER,
dba TEXT,
boro TEXT,
building TEXT,
street TEXT,
zipcode TEXT,
phone TEXT,
cuisine_description TEXT,
inspection_date TEXT,
action TEXT,
violation_code TEXT,
violation_description TEXT,
critical_flag TEXT,
score NUMERIC,
grade TEXT,
grade_date DATE,
record_date DATE,
inspection_type TEXT,
latitude NUMERIC,
longitude NUMERIC,
community_board TEXT,
council_district TEXT,
census_tract TEXT,
bin TEXT,
bbl TEXT,
nta TEXT,
location_point1 TEXT
) URL = 's3://firebolt-sample-datasets-public-us-east-1/nyc_sample_datasets/nyc_restaurant_inspections/parquet/'
OBJECT_PATTERN = '*.parquet'
TYPE = (PARQUET);
```
Create the internal table:
```sql
CREATE TABLE nyc_restaurant_inspections(
camis INTEGER,
dba TEXT,
boro TEXT,
building TEXT,
street TEXT,
zipcode TEXT,
phone TEXT,
cuisine_description TEXT,
inspection_date TEXT,
action TEXT,
violation_code TEXT,
violation_description TEXT,
critical_flag TEXT,
score NUMERIC,
grade TEXT,
grade_date DATE,
record_date DATE,
inspection_type TEXT,
latitude NUMERIC,
longitude NUMERIC,
community_board TEXT,
council_district TEXT,
census_tract TEXT,
bin TEXT,
bbl TEXT,
nta TEXT,
location_point1 TEXT
);
```
Load the data from the external table:
```sql
INSERT INTO
nyc_restaurant_inspections
SELECT
camis,
dba,
boro,
building,
street,
zipcode,
phone,
cuisine_description,
inspection_date,
action,
violation_code,
violation_description,
critical_flag,
score,
grade,
grade_date,
record_date,
inspection_type,
latitude,
longitude,
community_board,
council_district,
census_tract,
bin,
bbl,
nta,
location_point1
FROM
ex_nyc_restaurant_inspections;
```
## Sample SQL queries for analysis [#sample-sql-queries-for-analysis]
### Total inspections by borough [#total-inspections-by-borough]
```sql
SELECT boro, COUNT(*) as total_inspections
FROM nyc_restaurant_inspections
GROUP BY boro
ORDER BY total_inspections DESC;
```
### Top 10 most common cuisine types [#top-10-most-common-cuisine-types]
```sql
SELECT cuisine_description, COUNT(*) as count
FROM nyc_restaurant_inspections
GROUP BY cuisine_description
ORDER BY count DESC
LIMIT 10;
```
### Top 10 highest-scoring restaurants [#top-10-highest-scoring-restaurants]
```sql
SELECT dba, boro, max(score) as score
FROM nyc_restaurant_inspections
GROUP BY ALL
ORDER BY score DESC
LIMIT 10;
```
### Violations by critical flag [#violations-by-critical-flag]
```sql
SELECT critical_flag, COUNT(*) as count
FROM nyc_restaurant_inspections
GROUP BY critical_flag;
```
### Average inspection score by cuisine type (top 10) [#average-inspection-score-by-cuisine-type-top-10]
```sql
SELECT cuisine_description, AVG(score) as avg_score
FROM nyc_restaurant_inspections
GROUP BY cuisine_description
ORDER BY avg_score DESC
LIMIT 10;
```
Customize these queries based on your specific analytical objectives.
Spin up an engine and load this dataset yourself — Firebolt is free to try.
# NYC Traffic Dataset Analysis: A SQL Warm-Up (/free-sample-datasets/nyc-traffic)
New York City maintains comprehensive traffic records through its Department of Transportation. This
tutorial walks through analyzing NYC traffic patterns using SQL queries in Firebolt.
## Dataset schema [#dataset-schema]
| Column Name | Data Type | Description |
| ------------ | --------- | ------------------------------------------ |
| `requestid` | Bigint | Unique identifier for each record |
| `boro` | Text | Borough the traffic data is from |
| `yr` | Int | Year of the data |
| `month` | Int | Month of the data |
| `dd` | Int | Day of the data |
| `hh` | Int | Hour of the data |
| `mm` | Int | Minute of the data |
| `vol` | Int | Traffic volume |
| `segmentid` | Bigint | Unique identifier for each segment |
| `wktgeom` | Text | Well-Known Text representation of geometry |
| `street` | Text | Street name |
| `fromstreet` | Text | Starting street |
| `tostreet` | Text | Ending street |
| `direction` | Text | Traffic direction |
## Loading NYC traffic data into Firebolt [#loading-nyc-traffic-data-into-firebolt]
### Option 1: COPY FROM [#option-1-copy-from]
Let Firebolt infer the schema and create the table automatically:
```sql
COPY INTO nyc_traffic FROM 's3://firebolt-sample-datasets-public-us-east-1/nyc_sample_datasets/nyctraffic/parquet/'
WITH PATTERN="*.parquet" AUTO_CREATE=TRUE TYPE=PARQUET;
```
This generates the following table structure:
```sql
CREATE TABLE "nyc_traffic" ("requestid" bigint NULL, "boro" text NULL, "yr" integer NULL, "month" integer NULL, "dd" integer NULL, "hh" integer NULL, "mm" integer NULL, "vol" integer NULL, "segmentid" bigint NULL, "wktgeom" text NULL, "street" text NULL, "fromstreet" text NULL, "tostreet" text NULL, "direction" text NULL)
```
### Option 2: External table [#option-2-external-table]
Create an external table over the Parquet files in S3:
```sql
CREATE EXTERNAL TABLE ex_nyc_traffic (
requestid bigint,
boro Text,
yr int,
month int,
dd int,
hh int,
mm int,
vol int,
segmentid bigint,
wktgeom Text,
street Text,
fromstreet Text,
tostreet Text,
direction Text
) URL = 's3://firebolt-sample-datasets-public-us-east-1/nyc_sample_datasets/nyctraffic/parquet/' OBJECT_PATTERN = '*.parquet'
TYPE = (PARQUET);
```
Create the internal table with a primary index tuned for the queries below:
```sql
CREATE TABLE nyc_traffic (
requestid bigint,
boro Text,
yr int,
month int,
dd int,
hh int,
mm int,
vol int,
segmentid bigint,
wktgeom Text,
street Text,
fromstreet Text,
tostreet Text,
direction Text) PRIMARY INDEX yr,month,dd, hh, mm, boro, street;
```
Load the data:
```sql
INSERT INTO nyc_traffic SELECT * FROM ex_nyc_traffic;
```
## Sample SQL queries [#sample-sql-queries]
### Total traffic volume by segment and borough [#total-traffic-volume-by-segment-and-borough]
Identify the top 10 segments with the highest traffic volume.
```sql
SELECT boro, street, segmentid, SUM(vol) AS TotalTrafficVolume
FROM nyc_traffic
GROUP BY boro, street, segmentid
ORDER BY TotalTrafficVolume DESC
LIMIT 10;
```
### Traffic volume by direction [#traffic-volume-by-direction]
Break down traffic volume by direction.
```sql
SELECT direction, SUM(vol) AS TotalTrafficVolume
FROM nyc_traffic
GROUP BY direction;
```
### Count of records by month and year [#count-of-records-by-month-and-year]
Identify trends across month-year combinations.
```sql
SELECT yr, month, COUNT(*) AS RecordCount
FROM nyc_traffic
GROUP BY yr, month
ORDER BY yr, month;
```
### Busiest streets by total traffic volume [#busiest-streets-by-total-traffic-volume]
Find the top 10 busiest streets for infrastructure planning.
```sql
SELECT street, SUM(vol) AS TotalTrafficVolume
FROM nyc_traffic
GROUP BY street
ORDER BY TotalTrafficVolume DESC
LIMIT 10;
```
### Peak traffic hour [#peak-traffic-hour]
Find the hour with the highest traffic volume.
```sql
SELECT hh, SUM(vol) AS TotalTrafficVolume
FROM nyc_traffic
GROUP BY hh
ORDER BY TotalTrafficVolume DESC
LIMIT 1;
```
### Comprehensive traffic analysis [#comprehensive-traffic-analysis]
Combine year, borough, streets, and direction into one view.
```sql
SELECT yr, boro, street, fromstreet, tostreet, direction, SUM(vol) AS TotalTrafficVolume
FROM nyc_traffic
GROUP BY ALL
ORDER BY yr, TotalTrafficVolume DESC;
```
Spin up an engine and load this dataset yourself — Firebolt is free to try.
# Ultra Fast Gaming: Firebolt Sample Dataset (/free-sample-datasets/ultra-fast-gaming)
UltraFast Gaming Inc. publishes online games across PlayStation, Xbox, PC, iOS, and Nintendo. The
company collects extensive data about games, levels, players, tournaments, play sessions, and
rankings to support data-driven decisions about development, tournaments, community initiatives, and
interactive leaderboards.
## Getting started [#getting-started]
### Create and select an engine [#create-and-select-an-engine]
Create a database engine before running queries — through the Firebolt UI or with SQL:
```sql
CREATE ENGINE "ultra_fast_engine" WITH
TYPE = "S"
NODES = 1
AUTO_STOP = 10
INITIALLY_STOPPED = false
AUTO_START = true
CLUSTERS = 1;
```
After creating the engine, select it using the dropdown in the query editor.
### Load the dataset [#load-the-dataset]
For new Firebolt accounts, the UltraFast Gaming dataset loads automatically into the `ultra_fast`
database. Access it with:
```sql
USE ultra_fast;
```
For existing accounts without the dataset, create the database and ingest the data:
```sql
CREATE DATABASE ultra_fast;
COPY INTO games FROM 's3://firebolt-sample-datasets-public-us-east-1/gaming/parquet/games/'
WITH PATTERN='*.snappy.parquet' TYPE = PARQUET;
COPY INTO levels FROM 's3://firebolt-sample-datasets-public-us-east-1/gaming/parquet/levels/'
WITH PATTERN='*.snappy.parquet' TYPE = PARQUET;
COPY INTO players FROM 's3://firebolt-sample-datasets-public-us-east-1/gaming/parquet/players/'
WITH PATTERN='*.snappy.parquet' TYPE = PARQUET;
COPY INTO playstats FROM 's3://firebolt-sample-datasets-public-us-east-1/gaming/parquet/playstats/'
WITH PATTERN='*.snappy.parquet' TYPE = PARQUET;
COPY INTO rankings FROM 's3://firebolt-sample-datasets-public-us-east-1/gaming/parquet/rankings/'
WITH PATTERN='*.snappy.parquet' TYPE = PARQUET;
COPY INTO tournaments FROM 's3://firebolt-sample-datasets-public-us-east-1/gaming/parquet/tournaments/'
WITH PATTERN='*.snappy.parquet' TYPE = PARQUET;
SHOW TABLES;
```
## Tables included [#tables-included]
| Table | Description |
| ------------- | ----------------------------------------------- |
| `Games` | Game titles, supported platforms, launch dates |
| `Levels` | Game level information |
| `Players` | Player profiles, platforms, subscription status |
| `PlayStats` | Session statistics, scores, playtime metrics |
| `Rankings` | Tournament rankings and player positions |
| `Tournaments` | Tournament details and metadata |
## Example queries [#example-queries]
### Platform support analysis [#platform-support-analysis]
Determine which games support a specific platform:
```sql
SELECT
Title AS GameTitle,
ARRAY_CONTAINS(SupportedPlatforms, 'PlayStation') AS SupportsPlayStation
FROM
Games;
```
Find the total number of platforms a game supports:
```sql
SELECT
Title AS GameTitle,
ARRAY_LENGTH(SupportedPlatforms) AS NumberOfPlatforms
FROM
Games
WHERE
Title = 'Johnny B. Quick';
```
Functions used: `ARRAY_CONTAINS`, `ARRAY_LENGTH`.
### Game launch date analysis [#game-launch-date-analysis]
Analyze launch dates, time elapsed, and recency:
```sql
SELECT
Title AS GameTitle,
DATE_DIFF('day', LaunchDate, CURRENT_DATE) AS DaysSinceLaunch,
DATE_TRUNC('month', LaunchDate) AS LaunchMonth,
CASE
WHEN LaunchDate >= CURRENT_DATE - INTERVAL '1 year' THEN 'Yes'
ELSE 'No'
END AS LaunchedWithinLastYear
FROM
Games;
```
Functions used: `DATE_DIFF`, `DATE_TRUNC`, and date arithmetic operators.
### Tournament player performance [#tournament-player-performance]
Use window functions to rank player performance within specific tournaments:
```sql
WITH PlayerScores AS (
SELECT
ps.GameID,
ps.PlayerID,
AVG(ps.CurrentScore) AS AvgScore
FROM
PlayStats ps
WHERE
ps.TournamentID IN (56, 16, 98)
GROUP BY ALL
),
RankedScores AS (
SELECT
ps.GameID,
ps.PlayerID,
ps.AvgScore,
RANK() OVER (PARTITION BY ps.GameID ORDER BY ps.AvgScore DESC) AS ScoreRank
FROM
PlayerScores ps
)
SELECT
g.Title AS GameTitle,
p.Nickname AS PlayerNickname,
rs.AvgScore,
rs.ScoreRank
FROM
RankedScores rs
JOIN
Games g ON rs.GameID = g.GameID
JOIN
Players p ON rs.PlayerID = p.PlayerID
WHERE
rs.ScoreRank <= 50;
```
### Subscription and platform impact on playtime [#subscription-and-platform-impact-on-playtime]
Analyze how subscription status and platform influence player engagement:
```sql
WITH PlayerPlayTime AS (
SELECT
ps.PlayerID,
ps.GameID,
p.IsSubscribedToNewsletter,
UNNEST(p.Platforms) AS Platform,
AVG(ps.CurrentPlayTime) AS AvgPlayTime,
RANK() OVER (PARTITION BY ps.GameID ORDER BY AVG(ps.CurrentPlayTime) DESC) AS PlayTimeRank
FROM
PlayStats ps
JOIN
Players p ON ps.PlayerID = p.PlayerID
WHERE
ps.StatTime BETWEEN '2020-12-01' AND '2021-02-01'
AND ps.TournamentID > 100
GROUP BY ALL
),
FilteredPlayers AS (
SELECT
ppt.PlayerID,
ppt.GameID,
ppt.IsSubscribedToNewsletter,
ppt.Platform,
ppt.AvgPlayTime,
ppt.PlayTimeRank,
CASE
WHEN ppt.IsSubscribedToNewsletter = TRUE THEN 'Subscribed'
ELSE 'Not Subscribed'
END AS SubscriptionStatus
FROM
PlayerPlayTime ppt
WHERE
ppt.AvgPlayTime > 0
)
SELECT
fp.GameID,
fp.Platform,
fp.SubscriptionStatus,
AVG(fp.AvgPlayTime) AS AvgPlayTime,
AVG(fp.PlayTimeRank) AS AvgPlayTimeRank
FROM
FilteredPlayers fp
GROUP BY ALL
HAVING
AVG(fp.AvgPlayTime) > 0;
```
SQL features used: `CASE WHEN`, `UNNEST`, `HAVING`, and window functions.
## Exploring the schema [#exploring-the-schema]
Discover all tables and their columns:
```sql
SELECT * FROM information_schema.tables;
SHOW COLUMNS IN
;
```
The dataset supports advanced analysis using subresult reuse for query optimization, and it can be
extended with custom data ingested from your own S3 sources.
Spin up an engine and explore this dataset yourself — Firebolt is free to try.
# What are aggregations? (/glossary-items/aggregations)
In general, aggregations are grouping functions. Instead of returning single-row or single-document data, the response is grouped by a certain criteria. The most common case is the faceted search on an ecommerce shop like amazon.com. After you have searched for something like a new TV, part of your results page contains information on how many products were found for each brand or for each display size — allowing you to reduce the result further by using brand or dimensions as a filter for your product selection. There are many technology solutions to handle aggregations, and each technology handles aggregations differently. Different implementations are covered in the following sections.

The above image features three different information retrieval parts. First, the list of TVs on the right side. Second, the brand selection and count, and third, the display size selection and count. The last two are aggregations.
## Examples of Aggregations [#examples-of-aggregations]
Use of relational databases to store and retrieve data using Structured Query Language (SQL) is a very prevalent and proven approach. In the SQL world, an aggregate function collects a set of values and returns a single number. The most common ones are probably COUNT, MAX, MIN, SUM and AVG in combination with a grouping criteria in a GROUP BY statement or using WINDOW functions.
Translating SQL into aggregations of other data stores or vice versa shows the complexity due to subtle differences. Starting with some SQL for the above ecommerce example, each query fetches a part for the web page to display. First, the products on the right, then the two aggregations based on brand and display size on the left:
Elasticsearch provides another approach to searching through piles of data. Taking a look at Elasticsearch (or OpenSearch aggregations for that matter), the above e-commerce use case can be implemented with a search request using the JSON query DSL (domain specific language). This request fetches data that contains aggregations for each criteria on the left as well as results.
NoSQL databases, such as MongoDB, are another approach to information storage and retrieval. In MongoDB, aggregations have a different syntax than the nested JSON approach from Elasticsearch, using function names preceded by a $ sign. There is a dedicated $facet aggregation for this very common use-case to simplify the way the query is written. This shows one query for the aggregated data and another for retrieving the search results.
## Challenges with aggregations [#challenges-with-aggregations]
Aggregations may need to read through a lot of data and can take significant time to compute. Pre-aggregating data, Materialized Views in SQL for example, is a common trick to speed up retrieval of aggregate data. The trade-off is that the aggregates may not reflect the current state of the system. In the use case above, if the inventory was pre-aggregated 12-hours ago, the aggregate search might reflect stale data and show that a product is available when it is really out-of-stock. When calculating aggregations, the trade-offs between fresh data, performance and technology choices should be considered. Not all technologies perform the same way when dealing with pre-aggregations.
Additionally, when dealing with large amounts of data, most technologies require significant compute and memory to calculate aggregations. Some implementations overflow to disk once a certain amount of data is exceeded in memory, other data stores abort the request to prevent running out of memory and remain responsive for other requests or allow for background execution instead of waiting for the request to be completed and blocking the connection. Check your data store implementation to be sure it can handle the amount of data.
Another approach to speeding up aggregations is the use of probabilistic data structures, also called data sketches. Usage of those results in counts, distinct counts or quantiles not being a hundred percent exact. The reason for this is the ability to run such calculations on data on different nodes and then merge the results together without having all the data centralized on a single system, which would be necessary using the naive approach. This also makes it harder to compare the output of two different data stores to ensure that you have written the same query for two different data stores.
Transactional consistency is another aspect that can have an impact on the aggregate data returned. Data modifications may have happened between two requests. In this example this might lead to different counts of your displayed results, which in this case is not a big issue.
# What is Airflow? (/glossary-items/airflow)
Airflow is a powerful open-source tool for orchestrating and managing workflows. As a software engineer, you can use Airflow to automate, schedule, and monitor complex data pipelines and machine learning workflows.
Before diving into the specifics of how to properly utilize Airflow, it's important to understand the concepts and components of the tool.
## Airflow concepts and components [#airflow-concepts-and-components]
Airflow is built around the concept of "DAGs," or directed acyclic graphs. A DAG represents a collection of tasks and the dependencies between them. Each node in the DAG represents a task, and the edges between the nodes represent the dependencies between the tasks.
Airflow uses a scheduler to determine the order in which tasks should be executed. The scheduler takes into account the dependencies between tasks, as well as other factors such as task concurrency and task retries.
In addition to the scheduler, Airflow also has a web interface that allows you to monitor and manage your workflows. The web interface provides a visual representation of your DAGs, as well as the status of each task and the logs for each task execution.
## How to use Airflow [#how-to-use-airflow]
Now that we have an understanding of the basic concepts and components of Airflow, let's dive into how to properly utilize the tool.
### Define your workflow as a DAG [#define-your-workflow-as-a-dag]
The first step in utilizing Airflow is to define your workflow as a DAG. A DAG is defined using Python code, and it should include all of the tasks that make up your workflow, as well as the dependencies between those tasks.
Here is an example of a simple DAG that defines a workflow with two tasks:

In this example, the DAG is defined with a schedule\_interval of one day, which means that it will run automatically once per day. The DAG has two tasks: task1 and task2. The `>>` operator is used to define the dependencies between the tasks, in this case task2 depends on task1.
### Use the Airflow web interface to monitor and manage your workflows [#use-the-airflow-web-interface-to-monitor-and-manage-your-workflows]
Once you have defined your DAG, you can use the Airflow web interface to monitor and manage your workflows. The web interface provides a visual representation of your DAGs, as well as the status of each task and the logs for each task execution.
You can use the web interface to trigger a manual execution of a task, or to view the logs for a task execution. You can also use the web interface to view the status of all of the tasks in your DAG, including information such as the start and end time for each task execution, and whether or not the task succeeded or failed.
### Utilize Airflow hooks and operators to interact with external systems [#utilize-airflow-hooks-and-operators-to-interact-with-external-systems]
One of the powerful features of Airflow is the ability to interact with external systems using hooks and operators. Hooks are used to connect to external systems, such as databases or cloud storage, while operators are used to perform specific actions, such as running a SQL query or uploading a file to cloud storage.
Here is an example of how to use the Airflow PostgresHook to run a SQL query on a Postgres database:

In this example, a PythonOperator is used to run a function called run\_query, which connects to a Postgres database using the PostgresHook and runs a SQL query.
### Utilize Airflow's built-in features for task retries, task concurrency, and task backfilling [#utilize-airflows-built-in-features-for-task-retries-task-concurrency-and-task-backfilling]
Airflow provides several built-in features that can be used to improve the robustness and efficiency of your workflows. For example, you can configure a task to automatically retry if it fails, or to limit the number of concurrently running tasks.
Here is an example of how to configure a task to automatically retry if it fails:

In this example, the task is configured to retry 3 times with a delay of 5 minutes between retries.
You can also use Airflow's backfilling feature to retroactively run a DAG for a specified time range. This can be useful for re-processing data or for catching up on missed task executions.
In conclusion, Airflow is a powerful tool for orchestrating and managing workflows. By properly utilizing the concepts of DAGs, the web interface, hooks and operators, and built-in features like task retries and backfilling, you can improve the robustness and efficiency of your data pipelines and machine learning workflows.
# The Cloud Data Warehousing Guide (/glossary-items/cloud-data-warehouse)
An Introduction to Cloud Data Warehousing describes how businesses store and manage data in the cloud. Cloud data warehouses represent a significant evolution in data storage, enabling flexibility, scalability, and affordability in managing increasingly large and complex data. In this whitepaper, we will use the terms "cloud data warehouse" and "data warehouse" interchangeably.
## An Introduction to Cloud Data Warehousing [#an-introduction-to-cloud-data-warehousing]
Every aspect of data management is conducted under the virtual roof of a cloud data warehouse. Most importantly, a cloud data warehouse transforms data into assets companies can use to improve their capabilities, fuel innovation, and enhance profits.
Traditional data warehouses are physical structures typically on-site. While these have served businesses well for many years, a series of challenges, including high costs and complexities with legacy hardware, have rapidly antiquated them. Cloud data warehouses are a robust and ultramodern alternative to traditional data warehouses, one that can lead to profound success for modern businesses. However, enterprises that have to adhere to special compliance or connectivity requirements still leverage on-premises solutions.
By 2026, the [market value of the cloud data warehousing industry](https://www.marketsandmarkets.com/Market-Reports/data-warehouse-as-a-service-market-191544663.html) is forecast to hit $12.9 billion, a compound annual growth rate of 22.3%. While North America and Europe hold the highest market share, the fastest-growing segment in cloud data warehousing is the Asia-Pacific region, powered by the booming megamarkets of China and India.
The key factor behind these numbers is that data now drives the world, including business. Cloud data warehouses are highly scalable and provide a safe, secure environment backed up by the expertise of leading high-tech companies.
Industries that benefit the most from cloud data warehousing include manufacturing, energy and utilities, healthcare, IT, government, retail, and BFSI (banking, financial services, and insurance). Since cloud data warehousing is such a flexible solution, the use cases are diverse. But one thing is certain: cloud data warehousing is now the norm and the foundation upon which the future will be constructed.
## A Brief Overview of Data Warehouse Architecture [#a-brief-overview-of-data-warehouse-architecture]
Data warehouse architecture comprises three tiers. The top tier represents the front-end client that offers results via analysis, reporting, data mining tools, and other management. The second or middle tier comprises ELT, which organizations use to access and analyze data. The third or bottom tier of data warehouse architecture is essentially the database server where enterprises load and store data.
Data can be stored in two ways. Organizations can either leverage high-speed storage to enable quick and frequent access or implement cheap object storage for infrequently accessed data. The data warehouse will move frequently accessed data into "fast" storage to optimize query speeds. Depending on access requirements, organizations can use different logic as well.


## Online Analytical Processing (OLAP) [#online-analytical-processing-olap]
Online analytical processing (OLAP) is a type of data processing that occurs in a data warehouse and serves different workloads and requirements.
Data consistency is optional for OLAP systems since they typically use data snapshots. OLAP systems handle large data volumes and use denormalized database designs leveraging star schema or snowflake schema. This approach increases data redundancy, improves query performance, and accelerates data-driven decision-making.
## How Does Data Warehousing Work? [#how-does-data-warehousing-work]
Data warehouses continuously collect and organize data into a dedicated comprehensive centralized repository. Data collected from various sources are systematically sorted into tables based on the data type and layout.
Insights harvested from a data warehouse help businesses better understand their target audience or customers and be alert to emerging trends. For example, enterprises can gain a competitive advantage by forecasting market changes, formulating a robust pricing strategy, or developing better products.
There are three layers of data warehousing:
1. An **enterprise data warehouse (EDW)** is a centralized repository providing decision-making support to various departments across the company. EDWs provide a comprehensive and consolidated approach to how companies organize and represent data. As such, data teams can classify the data based on the subject and grant access accordingly.
2. An **operational data store (ODS)** is often a go-to choice for organizations with data warehouse systems failing to satisfy their reporting requirements. As ODS can be refreshed in real time, it is a popular option for storing routine activities. In a large healthcare system, real-time or near-real-time data access is crucial for patient care. Due to their batch-processing nature, traditional data warehouses can't always meet this need. This is where an ODS comes into play. The ODS can integrate data from these various sources like electronic medical records (EMR), pharmacies, laboratory systems, and radiology to present a unified and current view of the patient's data.
For instance, a doctor can access the ODS to get the most recent patient data like lab results or medication history; and a pharmacist can verify prescriptions to prevent harmful drug interactions. Thus, the ODS provides timely, integrated data to healthcare providers, improving patient care.
1. A **data mart** is designed for specific business or industry verticals. For example, they are prevalent in finance, sales, and inventory. Moreover, data marts can quickly collect data from a source.
An EDW stores static data, whereas an ODS integrates dynamic operational data. The data mart creates specialized data views over the EDW.
Enterprises can configure data warehouses into one or multiple of the following system configurations:
* **Offline operational database:** Data will be copied periodically to a server from an ODS to load, process, and report. This approach is practical when data synchronization isn't a must.
* **Offline data warehouse:** Data is stored and regularly updated from the operational database and other sources to derive critical business insights.
* **Real-time data warehouse:** A real-time data warehouse is used for up-to-date insights and analysis based on the latest transactional data. All transactions in an operational database are updated in the data warehouse.
* **Integrated data warehouse:** The integrated data warehouse consolidates data into a unified view for analysis. All transactions occurring in the operational database are simultaneously updated in the data warehouse. Once updated, the data warehouse will generate transactions and forward them to the operational database.
Software tools and hardware used for storing, transforming, and analyzing data are called data warehouse appliances.
With the concepts introduced above, we can highlight three key ways in which a data warehouse can work:
1. **Basic data warehouse:** Organizations can eliminate data redundancy, which reduces the amount of data in storage. This, in turn, makes data clearer and more user-friendly. The key benefit here is that different departments from multiple sources can quickly access data directly from the warehouse.
2. **Data warehouse with staging area:** Organizations can clean data in "staging areas" before moving it to storage. This is one of the leading methods of ensuring that only relevant and valuable data is stored in the data warehouse.
3. **Data warehouse with data marts:** Organizations can enhance their data warehouse's customization level after data is processed, allowing them to streamline information to staff, teams, or departments that need it the most. This approach helps boost productivity and accelerate the decision-making pace.
## Next part: Top 3 use cases [#next-part-top-3-use-cases]
# Columnar Storage (/glossary-items/columnar-storage)
Columnar storage is a data organization technique where data is stored and retrieved by column rather than by row. While traditional row-based databases write complete records contiguously on disk, columnar databases group and store each column's values together — which has huge implications for performance, compression, and parallelism.
If you're new to databases or trying to squeeze better performance out of your analytics, you've probably heard about columnar storage. But what is it, really?
In other words, columnar storage means storing your data by column instead of by row. That might sound like a minor difference, but it has a huge impact on performance — especially when you're running analytical queries on large datasets.
## Why storage layout matters for analytical workloads [#why-storage-layout-matters-for-analytical-workloads]
To understand the impact of storage layout, think about how analytical queries work. These queries often:
* Scan billions of rows
* Touch only a small subset of columns
* Perform aggregations (e.g., `SUM`, `AVG`, `COUNT`)
In a **row-based system**, every row must be read in full, even if the query only needs a single column. That results in unnecessary I/O and slower performance.
In contrast, **columnar storage** reads only the relevant columns — skipping everything else. This dramatically reduces disk I/O, speeds up scans, and enables vectorized execution.
Columnar formats also benefit from:
* **High compression ratios** (since column values are similar in type and distribution)
* **SIMD optimizations** (processing many values at once in memory)
* **Late materialization** (delaying row reassembly until necessary)
These architectural features make columnar storage ideal for **OLAP-style** workloads — think dashboards, metrics reporting, data exploration, and ML feature prep — not transactional workloads.
**Row storage:**

**Column storage:**

## How columnar storage works [#how-columnar-storage-works]
At a high level, columnar storage organizes data so that each column is stored independently, often in its own physical block or file segment. This separation unlocks a series of optimizations that dramatically improve performance for read-heavy workloads — especially analytical queries. Let's break it down.
### Column-based layout: storage and retrieval [#column-based-layout-storage-and-retrieval]
In a columnar database, instead of storing complete rows one after another, each column's values are stored contiguously. This means:
* You can **read only the columns you need**, skipping the rest entirely (this is called column pruning).
* It's easier to **compress data**, since values in a column tend to be similar (more on that next).
* Execution engines can apply SIMD and vectorized processing, scanning through columnar data blocks with higher CPU efficiency.
For example, consider this table:

Instead of storing it row-by-row, a columnar engine stores it like this:

If you run a query like `SELECT AVG(age)`, only the Age column is scanned — saving I/O and compute.
### Compression techniques in columnar storage [#compression-techniques-in-columnar-storage]
Because columns contain similar data types and value ranges, columnar formats unlock high compression ratios. Some common techniques include:
* Run-Length Encoding (RLE): Stores repeated values as a single value + count.
* Example: `[US, US, US, UK]` becomes `[(US, 3), (UK, 1)]`
* Dictionary Encoding: Replaces values with a reference to a dictionary.
* Example: `[US, UK, US]` → Dictionary: `{0=US, 1=UK}` → Encoded: `[0,1,0]`
* Bit-Packing & Delta Encoding: Efficient for numeric columns with small value ranges or sorted data.
The result? Smaller data blocks, faster reads, and lower memory usage — all critical for analytical queries.
### Execution optimizations: why it's so fast [#execution-optimizations-why-its-so-fast]
Columnar engines typically support:
* Column Pruning: Only load columns referenced in a query.
* Predicate Pushdown: Apply WHERE clauses early to avoid loading irrelevant data.
* Vectorized Execution: Process data in batches (vectors) rather than row-by-row, boosting CPU throughput.
* Late Materialization: Delay reassembling rows until absolutely necessary, keeping the engine operating on compressed, columnar formats as long as possible.
All of this contributes to sub-second query performance — especially for aggregation-heavy workloads.
## Columnar vs row-based storage: what's the difference? [#columnar-vs-row-based-storage-whats-the-difference]
When evaluating database architectures, one of the most important choices is how data is physically laid out: row-based or columnar. Each has its strengths — and they serve very different purposes.
### Row-based storage: optimized for OLTP [#row-based-storage-optimized-for-oltp]
In row-based databases (think Postgres, MySQL, or SQL Server), each row is stored as a complete record. This is ideal for:
* Transactional workloads (OLTP) where the application frequently inserts, updates, or reads full records.
* Use cases like order processing, user authentication, or point-of-sale systems.
Example:
```sql
SELECT * FROM users WHERE user_id = 123;
```
A row store retrieves the entire record in one read — efficient and predictable. Also, transactional databases often normalize data into many tables and use joins, which row stores handle well on a per-row basis. The strength of row stores in OLTP comes from their ability to retrieve or modify entire records quickly, and to do so concurrently for many users. They ensure fast commit times and low-latency point queries. For example, a Postgres or MySQL database can easily handle thousands of short transactions per second for an e-commerce app, each reading or writing just a few rows.
### Columnar storage: optimized for OLAP [#columnar-storage-optimized-for-olap]
In columnar databases (like Firebolt, Redshift, or BigQuery), data is stored column-by-column, which:
* Reduces I/O by scanning only the columns needed
* Improves compression and query efficiency
* Enables vectorized execution, increasing throughput
This makes columnar storage the go-to choice for OLAP-style workloads — dashboards, aggregations, BI queries, and ML pipelines. These queries are read-heavy, often performing calculations (SUM, AVG, MAX, etc.) over large columns and involving filtering on certain dimensions (e.g. sales in Q4 for retail category). Columnar storage is optimized for this pattern: it can read through billions of values in a single column extremely fast, especially with compression and vectorized processing.
Example:
```sql
SELECT AVG(sales) FROM transactions WHERE region = 'West';
```
Only two columns (`sales`, `region`) are read and processed — no full-row overhead.
## Benchmark highlights: query speed & storage footprint [#benchmark-highlights-query-speed--storage-footprint]
Let's compare row vs columnar performance in typical analytical workloads:

[Explore Firebolt's benchmark results](https://www.firebolt.io/blog/introducing-firescale)
## Benefits of columnar storage [#benefits-of-columnar-storage]
When queries only touch a few columns out of potentially hundreds, columnar storage avoids reading unnecessary data. This leads to:
* Lower disk I/O
* Smaller data scans
* Faster execution times
Whether you're filtering millions of records or aggregating over time, columnar engines can deliver results in sub-seconds, even at scale.
Storing similar data types together allows columnar systems to use highly effective compression techniques like Run-Length Encoding (RLE) and dictionary encoding.
This results in:
* Smaller data footprints
* Lower storage costs
* Less memory consumption during query execution
Compressed data also means faster reads since there's less data to move from disk to memory.
Columnar storage lends itself naturally to vectorized execution, where operations are performed on batches of values instead of row-by-row.
It enables:
* SIMD (Single Instruction, Multiple Data) CPU acceleration
* Parallel scans across columns
* Better cache utilization
The result? Higher throughput and better performance on modern hardware — especially when running complex aggregations or filtering large datasets.
Columnar databases shine in OLAP scenarios, including:
* Real-time dashboards (fast metrics retrieval)
* Time-series analysis (e.g., tracking trends over time)
* Aggregations and filtering (e.g., SUM, AVG, GROUP BY)
If your workload involves slicing and dicing large volumes of data — think BI tools, product analytics, ML feature engineering — columnar is almost always the right choice.
## Columnar file formats: what you should know [#columnar-file-formats-what-you-should-know]
Columnar storage isn't just about how databases organize data internally — it also applies to how data is serialized and exchanged across distributed systems. That's where columnar file formats come in.
If you're working with tools like Spark, Presto, or data lake architectures, chances are you've come across formats like Parquet, ORC, or Arrow. Each was designed to optimize read performance, compression, and interoperability — but they have trade-offs worth knowing.
Let's break down the three most commonly used formats in modern data pipelines:

Each format is optimized for specific use cases:
| Format | Best For | Compression | Read Speed | In-Memory? | Tooling Support |
| ------- | ------------------------------ | ----------- | ---------- | ---------- | ----------------- |
| Parquet | General-purpose, lakehouses | ✅ High | ✅ Fast | ❌ No | ✅ Broad |
| ORC | Hive, numeric analytics | ✅ Very High | ✅ Fast | ❌ No | ⚠️ Ecosystem bias |
| Arrow | Real-time, in-memory pipelines | ⚠️ Lower | ⚡ Fastest | ✅ Yes | ✅ Growing |
Choose the format based on:
* Whether your workload is batch or real-time
* How important compression is for your storage layer
* Whether the system requires in-memory data sharing
## Columnar storage in Firebolt: what makes it different? [#columnar-storage-in-firebolt-what-makes-it-different]
Firebolt takes columnar storage to the next level by combining a proprietary file format, aggressive indexing, and a high-performance execution engine to deliver ultra-fast analytics at scale.
### Proprietary file format and performance layer [#proprietary-file-format-and-performance-layer]
Firebolt stores data in its own optimized columnar format (F3), designed for cloud-native performance. Data is automatically sorted, compressed, and indexed as it's ingested. Firebolt also adds a performance caching layer: recently accessed data is pulled from cloud storage into fast local SSD and RAM, minimizing latency and maximizing throughput.
### Indexes that accelerate performance [#indexes-that-accelerate-performance]
Unlike traditional columnar databases that rely mostly on brute-force scans, Firebolt uses indexes as a first-class performance booster:
* Sparse indexes (primary indexes) allow Firebolt to skip large chunks of irrelevant data by organizing tables in sorted order.
* Aggregating indexes precompute common aggregation queries like SUM, COUNT, and AVG, eliminating the need to scan entire tables during analytics.
This indexing approach drastically reduces the amount of data scanned and processed, leading to consistently sub-second queries even at massive scale.
### Firebolt architecture overview [#firebolt-architecture-overview]
Firebolt separates storage, compute, and metadata into distinct layers for maximum scalability. Data lives in low-cost object storage, while compute engines are spun up or down independently based on workload needs. During queries, Firebolt prunes data aggressively using indexes, retrieves only relevant compressed column chunks, and processes data using vectorized, massively parallel execution across compute nodes — delivering cloud-native analytics performance far beyond traditional columnar systems.
## Best practices for implementing columnar storage [#best-practices-for-implementing-columnar-storage]
Columnar storage unlocks major speed and efficiency gains for analytics, but realizing its full potential requires thoughtful implementation. From schema design to indexing strategy, making smart choices upfront ensures you maximize performance and minimize costs. Here are the key best practices every data engineer should follow:
### Schema design tips [#schema-design-tips]
Columnar databases work best when schema design aligns with how data is queried and scanned.
* **Favor wide tables**: Store related attributes together to reduce costly joins. Columnar storage handles wide tables efficiently because queries only scan needed columns.
* **Use efficient data types**: Smaller, more appropriate types (e.g., INT over BIGINT) save storage and enhance compression, speeding up queries.
* **Flatten where practical**: A moderate level of denormalization can drastically improve analytic query performance by reducing join complexity and avoiding excessive data movement.
### Partitioning and indexing [#partitioning-and-indexing]
Proper partitioning and indexing are critical for query pruning and fast scan performance in columnar databases.
* Partition by common filter columns: Choose partition keys that match typical query filters (like event\_date, region, or customer\_id) to skip irrelevant data quickly.
* Define primary/sparse indexes wisely: Set sort keys on columns most often used in filters or ranges. This enables the engine to prune data aggressively instead of scanning full columns.
* Leverage aggregating indexes: Precompute aggregates for frequent queries (e.g., daily sales totals) to deliver sub-second performance even on massive datasets.
### Handling large datasets and cold data [#handling-large-datasets-and-cold-data]
Scaling columnar storage effectively means balancing performance with storage optimization over time.
* Store cold data efficiently: Compress and archive historical or infrequently accessed data to reduce storage costs without sacrificing access when needed.
* Tier storage intelligently: Use cloud-native systems that automatically cache hot data in fast storage (like SSD or RAM) while keeping cold data in cost-effective object storage.
* Optimize refresh and maintenance: Tailor refresh rates, indexing, and compaction strategies for older data to minimize resource usage while ensuring analytics stay accurate and fast.
[Get expert help migrating to Firebolt's columnar storage](https://www.firebolt.io/book-a-demo)
## Common misconceptions about columnar storage [#common-misconceptions-about-columnar-storage]
While columnar storage is a proven powerhouse for analytics, there are still a few myths that cause confusion — especially among teams new to modern data architectures. Let's break down the most common misconceptions and set the record straight:
### "It's only for big data" [#its-only-for-big-data]
| Reality | Example |
| ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Columnar storage shines with large datasets, but its benefits aren't exclusive to "big data" scale. Even with moderate data volumes, columnar systems can dramatically speed up analytics by scanning only the columns you need, compressing data efficiently, and minimizing I/O. | A startup analyzing customer engagement metrics (with just a few million records) can see 10x faster query times on a columnar warehouse like Firebolt compared to a traditional row-store database. |
### "It doesn't support real-time" [#it-doesnt-support-real-time]
| Reality | Example |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Early columnar systems were optimized mainly for batch analytics, but modern engines like Firebolt have shattered this limitation. Today's columnar databases can handle real-time or near-real-time analytics by combining fast ingestion with aggressive indexing and caching strategies. | Firebolt's sparse indexes and performance layer allow users to query freshly ingested data within seconds — enabling real-time dashboards, personalization engines, and operational reporting without compromising speed. |
### "It's not developer-friendly" [#its-not-developer-friendly]
| Reality | Example |
| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Modern columnar databases are designed to be highly developer-friendly, often speaking standard SQL and supporting familiar data modeling patterns. Many even offer integrations with major BI tools, orchestration platforms, and SDKs to make development seamless. | Firebolt supports ANSI-SQL, offers native SDKs, REST APIs, and integrates smoothly with dbt — allowing developers to build complex analytical apps and pipelines without learning proprietary languages or complex abstractions. |
## Conclusion [#conclusion]
Columnar storage has transformed the way businesses approach analytics, unlocking faster query performance, greater scalability, and more efficient use of resources. By storing data by column instead of by row, organizations can dramatically reduce I/O, leverage compression, and accelerate analytical workloads — whether working with millions or billions of records.
As data volumes and user expectations continue to grow, adopting a columnar-first approach isn't just a performance boost — it's essential for building scalable, modern analytics platforms. Choosing the right columnar engine makes all the difference.
Firebolt sets itself apart with a proprietary file format, aggressive indexing strategies, and a cloud-native architecture built for extreme speed. With Firebolt, teams can power real-time analytics, deliver sub-second queries at massive scale, and optimize both cost and performance — all with a developer-friendly experience.
Ready to supercharge your analytics? [See Firebolt's columnar engine in action](https://www.firebolt.io/book-a-demo)
## FAQs [#faqs]
### What is columnar storage in a database? [#what-is-columnar-storage-in-a-database]
Columnar storage is a method of organizing data so that each column is stored separately, rather than storing entire rows together. This design is particularly effective for analytical queries, where users often need to scan specific columns across large datasets.
Instead of reading every field in every row, a columnar database can read just the relevant columns, significantly reducing I/O and improving performance.
### How is columnar storage different from row storage? [#how-is-columnar-storage-different-from-row-storage]
In row-based storage, each row is stored as a single unit — ideal for transactional systems where entire records are frequently inserted, updated, or retrieved.
In contrast, columnar storage groups all values of a column together. This makes it more efficient for analytics because:
* Only needed columns are scanned (column pruning)
* Data compresses better (homogeneous values)
* Execution engines can process data in batches (vectorization)
Row storage is best for OLTP workloads (e.g., banking systems), while columnar storage excels in OLAP workloads (e.g., dashboards, reports, machine learning pipelines).
### Why is columnar storage faster? [#why-is-columnar-storage-faster]
Columnar storage is faster for analytical queries because it:
* Reads less data by scanning only relevant columns
* Uses better compression, shrinking the amount of data moved from disk to memory
* Supports vectorized execution, allowing CPUs to process multiple values in parallel
* Delays row reassembly (late materialization), optimizing memory usage and execution flow
Together, these optimizations reduce I/O, improve CPU efficiency, and deliver sub-second query performance on large datasets.
### What companies use columnar storage? [#what-companies-use-columnar-storage]
Columnar storage is widely adopted by companies that rely on analytics, business intelligence, or large-scale data processing. You'll find it powering:
* Cloud data warehouses like Firebolt, Amazon Redshift, Snowflake, and BigQuery
* BI tools like Apache Druid and ClickHouse
* Data lake engines like Apache Parquet and Apache ORC (used in Spark, Hive, etc.)
From startups to Fortune 500 companies, columnar storage is a foundational technology behind real-time dashboards, product analytics, and AI pipelines.
### How does Firebolt utilize columnar storage? [#how-does-firebolt-utilize-columnar-storage]
Firebolt is built from the ground up as a columnar-first cloud data warehouse, designed to deliver sub-second performance at scale. Its columnar engine powers key features like:
* Column pruning and predicate pushdown to minimize I/O
* Vectorized query execution for efficient in-memory processing
* High compression ratios to reduce storage and speed up scans
* Native indexing for fast data access
* Integration with Apache Iceberg, bringing columnar performance to lakehouse data
# Customer Facing Analytics (/glossary-items/customer-facing-analytics)
With more data being collected on a day to day basis and data being treated as an asset, analytics is moving beyond internal use. This has resulted in a growing trend towards analytics being delivered to customers — both internal and external — in the form of customer facing analytics. While traditional BI enables ad hoc analysis and self service, customer facing analytics provides a packaged experience to the end user. Customer facing analytics provides an opportunity for the data producer (Enterprise that owns the data) and data consumer (end user or customer) to interact through data. In customer facing analytics, data models and visualizations are defined and delivered to the end user and enable interactive analysis within the set boundaries of a packaged experience. Customer facing analytics can take various shapes and forms and can be delivered as interactive dashboards, web or mobile apps. With this shift in analytics workloads, towards customer facing analytics, there are various aspects of analytics that become significant.
## Key Considerations for delivering Customer facing analytics [#key-considerations-for-delivering-customer-facing-analytics]
1. Define customer requirements in terms of specific insights and the data required to deliver these insights.
2. Create flexible data models that are prescriptive, yet extensible to address current and future customer needs.
3. Deliver analytics in various shapes and forms as needed by the customer.
4. Define clear expectations on freshness of data, response times. Performance is a critical element of customer facing analytics.
5. Ensure scalability, availability and provide secure access to insights.
Customer facing analytics might need service levels to be defined to ensure that expectations are set properly. Service levels can be used to agree on freshness of data, query response times, security of access etc. Examples of customer facing analytics vary depending on the end customer. A personalized dashboard in a gaming app or activity reporting in a fitness tracker or an order and inventory tracking application available to suppliers are all examples of customer facing analytics.
# Data Flattening and Data Unflattening (/glossary-items/data-flattening-and-data-unflattening)
Data flattening usually refers to the act of flattening semi-structured data, such as name-value pairs in JSON, into separate columns where the name becomes the column name that holds the values in the rows. Data unflattening is the opposite; adding nested structure to relational data.
If flattening does not sound great, you have good intuition. Flattening is mostly done to put data in a format that a relational data warehouse can use natively, and make sure queries perform well. But you do lose information; it's like having a shadow instead of the real object. When you flatten you lose information. This is one reason why a data lake should store the full, raw structure. You cannot unflatten data unless you have this extra information somewhere else.
Flattening data also limits your analytics. Often you want to be able to walk the actual JSON structure. When you flatten JSON, those analytics become nested queries, which is a horribly slow query to run on any relational database.
If you have mostly relational data, then flattening some data may make good sense. But if you have a lot, or plan to have a lot, you need native support for both relational and semi-structured data in the same data warehouse. For a great example of semi-structured data support you can read this [whitepaper on high-performance semi-structured analytics](https://www.firebolt.io/resources/semi-structured-analytics), and what it takes.
# What is a data-intensive application? (/glossary-items/data-intensive-application)
A data-intensive application is an application that makes an intense usage of data in all its heterogeneous forms. This earnestness of data handling can be measured in several ways. Nowadays, the vast majority of modern applications could be considered data-intensive. Generally speaking, we can call an application data-intensive if data is its primary challenge and from where almost all the business value comes. Furthermore, every application could become a data-intensive one, and probably, all or nearly applications that are not should strive to adopt a data-intensive approach.
A common trap when thinking about these kinds of applications is to focus on the size of the data sets handled. This, however, is not really what makes an application data-intensive. After all, if we had an application that used one petabyte of data, but all that data was static and never changed, we could probably get away with storing it on a single machine. The challenge with data-intensive applications is not necessarily the amount of data they use, but rather the fact that the data is constantly changing and often needs to be processed in real time.
Data-intensive applications are typically built around one or more core pieces of functionality that require access to large amounts of data. For example, a social networking site like Facebook needs to be able to quickly retrieve and process information about the relationships between different users. A search engine like Google needs to be able to index the billions of web pages on the Internet so that users can find the information they are looking for. And a fraud detection system like those used by credit card companies need to be able to analyze large numbers of transactions in real time to look for patterns that might indicate fraudulent activity.
In each of these cases, the functionality of the application is directly related to its ability to process large amounts of data quickly and effectively. These applications focus on packaging consumer-grade analytics experiences in a robust and responsive way. Evolving from traditional analytics and delivered by software engineering teams, the data and the experience around it is the product in this applications. Ultimately, they add value to existing products through purposeful analytics experiences intended to improve operations and efficiencies.
## Cloud offerings for data-intensive applications [#cloud-offerings-for-data-intensive-applications]
Applications benefit from cloud offerings including storage and delivering analytics platforms that support data apps. Cloud offerings also provide a number of other advantages for data-intensive applications, such as the ability to easily integrate with other cloud-based services and the ability to scale up or down quickly and easily in response to changes in demand.
For example, Facebook makes use of Amazon's Simple Storage Service (S3) to store images and videos uploaded by users, as well as data generated by the application itself. Facebook also uses Amazon's Elastic Compute Cloud (EC2) to run its web servers and database servers. Google uses a similar mix of S3 and EC2 for many of its applications, including its search engine, Gmail, and YouTube.
While the use of cloud services is not required for data-intensive applications, it can provide a number of significant advantages in terms of cost, speed, and scalability.
## Kinds of data [#kinds-of-data]
Every product, application, or service actively used by clients has access to various kinds of data:
* **Users' proprietary data**: consciously inserted and owned by the user, such as business data, profile data, or configuration data
* **Third-Party Data**: data retrieved from third-party systems (usually Data Management Platform) is used to enrich proprietary user data with insightful information.
* **Audit Data**: data generated during the usage of the application itself. Records what the user has done and how
It's easy to imagine how all those kinds of data could be aggregated to generate new business insight. Some examples include predictions about what the user needs, user experience customizations, or targeted decision-making strategies on the user's behalf.
The developer and the software architect of a data-intensive application combine several tools working with constantly evolving data: data from disparate systems, data of various types (structured, unstructured, binary, etc.), and varying speed, sizes, and shapes. Application developers are now becoming more and more data engineers; they should be accustomed to working with abstractions and virtualization of data systems to support the diversity of tools and structures, extending the capability of computations to multiple brands of products. This includes integrating with data platforms using APIs, SDKs, and SQL.
## Scalability, reliability, and performance [#scalability-reliability-and-performance]
Scalability, reliability, and performance are the three main concerns for any data system. Unfortunately, the more the data intensity grows, the more those fundamental characteristics of an excellent data-intensive application will be challenging to implement.
Scalability is what brought life to the idea of distributed data systems. Vertical and centralized data server scalability is an option, but not when we start requiring more than the standard commodity. After passing the standard-hardware commodity limit, specialized hardware costs draw an exceptionally sloping curve. We should consider that scalability for data-intensive applications could happen in various ways. It will depend, of course, on the specific need we are considering - more storage space needed or faster cluster data replication, rather than new geographically distributed nodes.
The level of reliability a standard data-intensive application needs depends on predefined SLAs (Service Level Agreements). Usually, we want the application to be able to handle any errors as fast as possible. Maintainability, instead, becomes increasingly complex as more and more tools and heterogeneous systems are added and aggregated into our application. Therefore, the aim should be to strive for simplicity and an effortless procedure for maintenance.
Finally, let's talk about performance. Usually, in most applications, we want real-time access and instant changes. There are several techniques we can adopt to optimize data access:
* **Indexes**: standard database indexes could help enhance data access. Developers should, however, pay close attention to how many indexes they want to create, or writing could become costly.
* [**Materialized Views**](https://www.firebolt.io/glossary-items/materialized-views): Also known as pre-computed queries. In this case, as in the previous one, developers should consider the use cases accordingly before accessing this kind of optimization. The flip side is a considerable increase in storage space used.
* **Caches**: database caching is similar to Materialized Views. Query results are pre-computed in both cases. The main difference is that caching is a static process: data cannot be cached if input filters change dynamically or if you need to have lots of cached data.
* **Geographically distributed database replicas**: a CDN-like structure in data systems is challenging to maintain but will grant substantial performance improvements.
# What is a Data Lake? Architecture, Best Practices & Implementation (/glossary-items/data-lake)
A data lake is a scalable repository, designed to collect and maintain structured, [semi-structured](https://www.firebolt.io/blog/a-primer-on-analyzing-semi-structured-data), and unstructured data from multiple sources, including enterprise databases, IoT sensors, mobile applications, and cloud platforms. Something traditional storage models are incapable of.
A data lake ingests data without immediate transformation, allowing flexibility for batch processing, low-latency queries, and streaming workflows. Governance frameworks, metadata, and schema layers can be added as needed to improve organization and access.
Organizations dealing with rapid data growth, fragmented systems, and outdated infrastructure stand to gain these advantages from data lakes:
* **Raw Data Preservation:** Legacy systems modify or restrict inputs, while data lakes store data in its raw form for flexible analysis.
* **Security and Compliance:** Access controls, encryption, and data masking sensitive information and meet regulatory requirements.
* **Data Integrity:** Automated pipelines, versioning, and error handling help maintain [data quality](https://www.firebolt.io/blog/matthew-weingarten-from-disney-streaming-about-data-quality-best-practices) while minimizing downstream issues.
* **Easy Access:** Standard APIs, metadata catalogs, and policy-based controls make data easy to discover and use.
## Core Components of Data Lake Architecture [#core-components-of-data-lake-architecture]
A data lake includes layers for storing raw data, executing queries, enforcing security, and controlling access, allowing users to retrieve and analyze data as needed. Here is how:
### Storage Layer [#storage-layer]
This layer serves as the foundation of a data lake, storing raw data from multiple sources in open file formats without requiring predefined schema structures. Cloud-native object storage solutions like Amazon S3 and Azure Blob provide virtually unlimited storage at a low cost, making them common choices.
A high-performance system efficiently ingests raw datasets from IoT devices, clickstreams, databases, and applications in its original formats such as JSON, Parquet, AVRO, ORC, and XML, and keeps it readily accessible for processing.
To maintain organization, staging zones segment data into three categories: raw, cleansed, and curated datasets, making processing faster and analytics easier.
Firebolt [integrates natively](https://www.firebolt.io/integrations) with data lake storage, providing direct access to common open file formats and making data more accessible.
### Processing Engine [#processing-engine]
The processing engine transforms, queries, and analyzes stored data. Many use a massively parallel processing (MPP) architecture to distribute computation across clusters. This speeds up batch jobs and queries using vectorized execution, [adaptive request optimization](https://www.firebolt.io/blog/5-steps-to-debug-your-complex-sql-queries-in-firebolt), and code generation.
A distributed SQL engine lets data engineers build ETL pipelines using standard ANSI SQL, eliminating the need for specialized programming languages. Optimized execution enables rapid analysis, even over petabyte-scale semi-structured datasets, simplifying exploratory analysis for data scientists.
Firebolt's [high-performance engine](https://www.firebolt.io/blog/a-comparison-of-data-warehouse-and-query-engines-on-amazon-web-services-aws) delivers rapid query execution, enabling instant insights for data-heavy applications.
### Security & Governance [#security--governance]
Governance controls safeguard data privacy, restrict access, and enforce security policies. Role-based [access](https://www.firebolt.io/data-security), encryption for data at rest and in transit, and audit logging provide enterprise-level security. [Columnar storage](https://www.firebolt.io/glossary-items/columnar-storage) and fine-grained permissions allow data masking and limit the exposure of sensitive information.
Building security, monitoring, and logging into the system from the start strengthens compliance, reduces costs, and ensures visibility into data activity.
Firebolt offers enterprise-grade [security features](https://www.firebolt.io/data-security), including end-to-end encryption, role-based access control, and data masking, ensuring compliance with industry standards and regulations.
## Data Lake vs Alternative Solutions [#data-lake-vs-alternative-solutions]
Data warehouses are built for structured data and optimized query performance, while lakehouses merge the flexibility of data lakes with the governance of warehouses. Understanding these differences helps organizations choose the right architecture for their workloads. Here's how data lakes compare to other models:
Here is a comparison of the primary distinctions between data lakes and data warehouses:
| Aspect | Data Lake | Data Warehouse |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Schema Flexibility | Uses a schema-on-read approach, storing raw data without predefined schemas. This allows handling of diverse data types and adapting schemas as needed. | Uses a schema-on-write approach, requiring predefined schemas before ingestion. This ensures consistency but limits flexibility. |
| Data Types | Stores structured, semi-structured, and unstructured data in native formats like JSON, XML, audio, video, and text files, making it versatile. | Designed primarily for structured data, with predefined columns and data types. Some modern warehouses support semi-structured formats like JSON. |
| Query Performance | Can be slower due to the lack of indexing and optimization, especially with large datasets. Performance improves with additional processing, indexing, and query engines. | Optimized for fast queries using indexing, partitioning, and other storage techniques. Supports complex analytics with efficient query execution. |
| Cost | Storage is generally cheaper, especially for large datasets, as data lakes use cost-effective object storage. However, processing and managing unstructured data may add costs. | Higher compute costs due to resource-intensive queries and data transformation. The structured format and fast query performance justify the investment. |
| Use Cases | Best for big data storage, machine learning, exploratory analysis, and raw data archiving. Common in IoT, logs, streaming, and unstructured data analysis. | Best for business intelligence (BI), structured reporting, and operational analytics where data consistency and fast queries matter. |
While data lakes excel in flexibility and scalability, they often struggle with performance and data management. In contrast, data warehouses are designed for structured data and faster querying but lack the adaptability to handle diverse data types.
Firebolt bridges this gap by combining the strengths of both architectures. Its advanced indexing and [decoupled storage and compute architecture](https://www.firebolt.io/blog/introducing-firebolts-next-gen-cloud-data-warehouse) allow for sub-second execution and high concurrency, even with large datasets. This design supports scalable resource allocation, allowing you to manage and analyze data without sacrificing speed or flexibility.
### Data Lake vs. Data Lakehouse [#data-lake-vs-data-lakehouse]
Data lakehouses introduce more structure to raw data without enforcing a rigid schema-on-write. By automatically capturing granular metadata and statistics, they enable querying through familiar database interfaces. This allows for complex transformations, joins, and analysis across decoupled storage and compute environments while maintaining the governance principles of traditional data warehouses.
[Lakehouses](https://www.firebolt.io/glossary-items/data-lakehouse) also incorporate ACID transaction support, overcoming a limitation of early data lakes. Separating storage from processing keeps data accurate, even under concurrent workloads. Improved data layouts allow for faster analysis of semi-structured data across distributed compute clusters.
Firebolt enhances lakehouse architectures with a high-performance distributed SQL engine for cloud object storage. It runs fast queries on raw and transformed data without moving it. Granular role-based access controls manage permissions at the column and table level, making workflows easier to control as lakehouse architectures scale.
## Common Data Lake Challenges [#common-data-lake-challenges]
Data lakes store large, diverse datasets, but keeping performance high, costs low, and data accurate at scale is challenging. Many architectures fail to support high-performance production workloads while maintaining governance across multiple teams, data sources, and compliance requirements.
Here are some common challenges in data lakes, and how Firebolt helps overcome them:
### Performance at Scale [#performance-at-scale]
As data lakes grow, query [latency](https://www.firebolt.io/low-latency) increases. Mixed data formats slow scans, rising user activity strains resources, and uneven data distribution across nodes further skew speed. Firebolt resolves these issues with an optimized engine that parallelizes execution and scales linearly. Indexes, caching, and code-free partitioning improve responsiveness for both ad-hoc and high-concurrency workloads.
### Cost Management [#cost-management]
Public cloud pricing is consumption-based, but costs can escalate without monitoring such as repository expenses growing as data accumulates, and over-provisioned resources inflating computing costs.
Firebolt improves price-performance with its serverless elastic engine, storage offloading, and built-in Cost Control tools. Granular metering optimizes configurations and [lowers costs by up to 40%](https://hi.firebolt.io/elasticity/cost-savings) compared to other cloud data warehouses.
### Data Quality [#data-quality]
Maintaining data quality in data lakes is difficult. Ingesting data from multiple sources causes inconsistencies, and weak governance amplifies errors, making accurate analytics harder to achieve.
Firebolt enforces schema flexibility, [metadata management](https://www.firebolt.io/blog/firebolt-features-effortless-metadata-management-for-faster-workflows), [validation](https://www.firebolt.io/blog/data-quality-with-dbt-and-firebolt), and lineage tracking to improve reliability. It also simplifies data integration with high-throughput ingestion and transformation capabilities.
## Best Practices for Data Lake Implementation [#best-practices-for-data-lake-implementation]
Applying best practices in architecture, data organization, and performance tuning improves the reliability and efficiency of a data lake. Here's how to approach each stage:
### Architecture Design [#architecture-design]
A well-planned architecture balances storage, compute, and security to ensure efficient data processing and scalability in a data lake. Consider the following:
* **Storage Layer Considerations:**
* Use managed cloud object stores like S3, GCS, or Azure Blob for durable and available raw data storage—these services scale capacity and throughput to handle growth.
* Decouple storage from compute to enable independent scaling, paying only for resources during active query processing.
* **Compute Resource Planning:**
* Use serverless query engines like Firebolt to scale on demand and eliminate overprovisioned resources.
* Choose cost-effective storage tiers (e.g., S3 Infrequent Access) while ensuring accessibility through Firebolt's querying capabilities.
* **Security Architecture Needs:**
* Encrypt data both at rest and in transit using AES-256 or SSL standards. Control access with identity federation through SSO and role-based policies.
* Comply with data governance regulations like GDPR while enabling auditing. Firebolt provides fine-grained logging for user activity tracking.
* **Integration Guidelines:**
* Connect to object storage using JDBC, ODBC, and native drivers, eliminating unnecessary ETL processes.
* Use Firebolt's SQL engine for ANSI compatibility, enabling the use of existing BI tools like Tableau for faster insights.
### Data Organization [#data-organization]
Effective data organization improves query performance, reduces storage costs, and simplifies data management in a data lake. Here are the key considerations:
* **Folder Structure Approaches:**
* Align with natural data hierarchies using a nested folder structure (e.g., /year=2022/month=January/day=1) for intuitive navigation.
* Partition by date columns or high-cardinality fields to improve filtering and query pruning.
* **File Format Selection:**
* Use columnar formats like Parquet instead of row-based formats for better compression and efficient column-based queries.
* [Firebolt's proprietary F3 format](https://www.firebolt.io/performance-at-scale) automatically indexes and sorts data during ingestion, optimizing query performance.
* **Partitioning Strategies:**
* Split data into multiple files by date, product, location, etc., to only access [partitions](https://www.firebolt.io/glossary-items/partitioning-and-sharding) matching query filters, reducing I/O.
* Keep partitions under 250MB to avoid overhead from too many small files or out-of-memory errors with large ones.
### Performance Optimization [#performance-optimization]
Optimizing query performance in a data lake requires efficient indexing, caching, and query execution techniques. Here are the key strategies:
* **Indexing Strategies:** [Firebolt's multi-dimensional indexing](https://www.firebolt.io/blog/firebolt-indexes-in-action) maps data patterns, relationships, and statistics for faster query execution. Selecting high-cardinality columns enhances performance.
* **Caching Approaches:** Grid cache keeps hot data in memory across all nodes, reducing disk I/O. Users get dedicated micro-caches as well.
* **Firebolt's Query Optimization Techniques:**
* Advanced cost-based optimizers translate SQL queries into ideal execution plans tailored to the engine based on data size, indexes, join types, etc.
* The solution's performance features include vectorized execution, which applies query instructions to batches of column data for speed. Modern CPUs handle this efficiently using SIMD parallelization.
* Code generation emits optimized C++ code for the query plan rather than interpreting it, speeding up computations.
## Firebolt's Data Lake Solution [#firebolts-data-lake-solution]
Carefully evaluating your analytics ecosystem and available technologies ensures you select the best solution.
While data lakes offer flexibility, they also present challenges. Applying best practices in data organization and performance optimization, along with advanced analytics engines, helps businesses get the most out of their data.
Firebolt natively integrates with cloud object storage platforms like Amazon S3, Google Cloud Storage, and Azure Blob Storage, providing direct, high-performance access to data. Its engine eliminates traditional ETL complexity and latency, allowing analysts to query data lakes natively through JDBC, ODBC, and REST API.
Here's what sets Firebolt apart:
* **Sub-Second Latency:** Firebolt combines dynamic indexing, code generation, vectorized execution, and aggregate caching to achieve [industry-leading speed](https://www.firebolt.io/resources/guide-to-sub-second-analytics) at any scale. Queries run 50-200x faster than existing data lake engines.
* **Support for 2,000+ Concurrent Users:** Firebolt scales to handle unlimited users without speed drops, thanks to its multi-tenant architecture and near-linear scalability.
* **3-Way Decoupled Architecture:** Compute, storage, and control layers operate independently for flexibility, high availability, and optimal resource use.
* **Postgres-Compatible SQL Dialect:** You can run [standard SQL](https://www.firebolt.io/blog/making-a-query-engine-postgres-compliant-part-i-functions) with ANSI syntax. Support for semi-structured data enables complex analytics.
By eliminating full dataset refreshes, reducing ETL overhead, and significantly improving query speed, Firebolt accelerates insights at any scale. Its cloud-native engine lowers costs, supports high concurrency, and delivers unmatched speed for data teams.
[Book a demo](https://www.firebolt.io/book-a-demo) today to see how Firebolt can transform your data lake performance.
# Data Lakehouse - Definition (/glossary-items/data-lakehouse)
A data lakehouse is a unified data platform that combines the low-cost, flexible storage of a data lake with the structure, governance, and performance of a [data warehouse](https://www.firebolt.io/blog/cloud-data-warehouse-solutions-for-big-data-analytics?). It collapses storage and analytics into a single system for holding raw and curated data, running analytics, and training models — so business intelligence and machine learning teams operate on the same source of truth.
Today's analytics and AI pipelines choke on fragmentation as data sprawls across lakes, warehouses, and silos. This leads to wasted hours moving data between systems just to get a basic model running or a dashboard updated.
A [data lakehouse](https://www.firebolt.io/blog/the-real-meaning-of-a-data-lake) is a means to overcome these challenges as it works by collapsing storage and analytics into one place, offering one system for storing raw and curated data, running analytics, and training models.
With unified storage, both business intelligence and machine learning teams operate on the same data — a single source of truth that supports high-throughput queries, fine-grained access control, and scalable compute from the same platform.
## What is a Data Lakehouse [#what-is-a-data-lakehouse]
It combines the low-cost, flexible storage of a data lake with the structure, governance, and performance of a data warehouse. It offers both structured and unstructured data while bringing reliable transactions, tight access control, and fast queries to flexible storage. This setup makes it easier for engineering, analytics, and AI teams to work from a shared foundation without copying data over multiple systems or managing conflicting pipelines. Its traits include:
* **Cloud-native storage**: Uses low-cost object storage to hold all data types such as raw, semi-structured, and structured.
* **Open table formats**: Works with Delta Lake, Apache Iceberg, or Hudi for managing large datasets with schema control and versioning.
* **Multi-language support**: Let's teams query and process data with SQL, Python, R, or Scala.
* **Hybrid data processing**: Handles both batch jobs and [streaming data](https://www.firebolt.io/glossary-items/stream-data-processing) in the same architecture.
* **Simplified stack**: Unifies the roles of data lakes and warehouses, cutting down on infrastructure sprawl and operational overhead.
## How a Data Lakehouse Works [#how-a-data-lakehouse-works]
A data [lakehouse](https://www.firebolt.io/blog/cloud-data-warehouse-vs-data-lake) builds on cloud object storage, but adds structure through table formats, metadata layers, and compute engines. The result is a unified system that offers analytics, machine learning, and streaming data without moving data between tools. Together, these layers create a platform where raw data can be ingested, transformed, queried, and used for machine learning, all in one place. Here's how they fit together:
* **Data ingestion**: Accepts batch, micro-batch, and streaming inputs from pipelines or external sources.
* **File storage**: Stores data in open formats like Parquet, ORC, or Avro for compatibility and compression.
* **Table formats**: Uses Delta Lake, Iceberg, or Hudi to add transactional guarantees, schema enforcement, and time travel.
* **Compute engines**: Query engines like Apache Spark, Trino, and Dremio process data directly from object storage, no staging required.
* **Catalogs**: Systems like Hive Metastore or Unity Catalog manage table definitions, access control, and schema versions.
* **APIs**: Expose data through standard interfaces like SQL, REST, JDBC, and notebooks to support BI tools and code-based workflows.
## Core Components [#core-components]
A lakehouse architecture is modular by design. It integrates open-source and cloud-native tools across storage, processing, and access layers.
* **Object Storage**: Stores all data using platforms like Amazon S3, Google Cloud Storage, Azure Blob, MinIO, or HDFS.
* **File Formats**: Supports Parquet, ORC, Avro, JSON, and CSV. These formats enable compression and compatibility across systems.
* **Table Formats**: Uses Delta Lake, Apache Iceberg, or Apache Hudi to manage schema evolution, enable transactions, and track data versions.
* **Query Engines**: Processes queries with engines such as Apache Spark, Trino, Dremio, PrestoDB, Flink, DuckDB, Starburst, or Athena.
* **Metadata and Catalogs**: Tools like Hive, Unity Catalog, AWS Glue, Nessie, and Amundsen handle table registration, schemas, and permissions.
* **Processing Layers**: Supports batch jobs (Spark, Dask) and streaming pipelines (Flink, Kafka, Spark Structured Streaming).
* **Machine Learning Support**: Connects directly to frameworks like TensorFlow, PyTorch, and scikit-learn.
* **Notebook and BI Access**: Compatible with Jupyter, Zeppelin, Tableau, Looker, Superset, and Power BI.
## Architecture Overview [#architecture-overview]
A data lakehouse stacks each layer to deliver a seamless experience across ingestion, storage, processing, and access. Each layer is decoupled but interoperable, so teams can swap tools or scale specific parts as needed.
* **Raw Data Layer**: Collects input from APIs, logs, events, and IoT devices.
* **Storage Layer**: Holds files in columnar formats on object storage.
* **Table Format Layer**: Applies versioning, partitioning, and schema control.
* **Metadata Layer**: Manages catalogs, access control, and table definitions.
* **Compute Layer**: Runs SQL, machine learning, ETL, and streaming queries.
* **Access Layer**: Serves notebooks, dashboards, and applications.
## File Formats in a Lakehouse [#file-formats-in-a-lakehouse]
Lakehouses rely on open file and table formats for flexibility and cross-platform compatibility.
**Columnar File Formats:**
* **Parquet**: Default in most lakehouses. Columnar, compressed, supports schema evolution.
* **ORC**: Similar to Parquet but optimized for Hive and Hadoop-based systems.
* **Avro**: Row-based format, useful for data interchange and schema evolution.
* **JSON**: Flexible for semi-structured data, but inefficient for analytics.
* **CSV**: Human-readable, rarely used for production-scale storage.
**Table Formats:**
* **Delta Lake**: Adds ACID transactions, schema enforcement, and time travel to Parquet.
* **Apache Iceberg**: Provides hidden partitioning, branching, rollback, and scalable metadata.
* **Apache Hudi**: Tailored for streaming ingestion, upserts, and incremental processing.
## Performance Optimization Strategies [#performance-optimization-strategies]
Lakehouses improve [performance](https://www.firebolt.io/blog/how-to-accelerate-looker-performance-on-redshift-snowflake-and-bigquery) by tightly managing how data is stored, accessed, and processed. These strategies help reduce scan time, improve compute efficiency, and speed up query results.
### Partition Pruning [#partition-pruning]
Partitions divide a table into logical segments based on a column like date, region, or `customer_id`. Partition pruning avoids full table scans by skipping irrelevant partitions.
* **Example**: A sales table partitioned by year will only scan data from 2024 if that's the only year requested.
* **Impact**: Reduces disk I/O and memory use, especially when querying narrow time windows or filtered dimensions.
### Column Pruning [#column-pruning]
Columnar formats like Parquet and ORC store data by column, not row. This allows engines to skip all non-selected columns during query execution.
* **Example**: If a dashboard query only asks for `total_revenue`, the engine ignores columns like `customer_name`, `sku`, or `promo_code`.
* **Impact**: Speeds up reads, reduces memory footprint, and shrinks intermediate data in pipelines.
### File Compaction [#file-compaction]
Write-heavy lakehouses often accumulate many small files. Compaction jobs consolidate these into larger, optimally sized files.
* **Example**: Delta Lake's `OPTIMIZE` command merges files to target 128 MB chunks, the sweet spot for parallel reads in cloud storage.
* **Impact**: Improves scan throughput and reduces cloud object store latency caused by metadata overhead.
### Z-Ordering [#z-ordering]
Z-ordering sorts data across multiple columns by interleaving values into a single ordering. It improves scan locality for multi-column filters.
* **Example**: Z-ordering a clickstream table by `session_id` and `event_time` groups related events close together on disk.
* **Impact**: Query engines can skip large file chunks when users filter on those fields, reducing scan time.
### Caching [#caching]
Caching frequently accessed tables or query results in memory or SSD avoids repeated reads from cold storage.
* **Example**: Spark's `persist(StorageLevel.MEMORY_AND_DISK)` stores popular datasets across nodes. Trino uses distributed caching for hot partitions.
* **Impact**: Improves response times for dashboards and interactive analysis.
### Predicate Pushdown [#predicate-pushdown]
Engines push filters down to the scan phase instead of applying them after data is loaded.
* **Example**: A query with `WHERE country = 'CA'` only reads blocks where the column statistics indicate `country = 'CA'` might exist.
* **Impact**: Reduces data volume early in the query plan, improving overall efficiency.
### Vectorized Execution [#vectorized-execution]
Columnar processing engines like Arrow, Spark, and DuckDB use vectorized execution to process data in batches.
* **Example**: Instead of calling a function 1 million times for each row, the engine processes 1024 values at once using SIMD-friendly operations.
* **Impact**: Increases CPU throughput and reduces function call overhead.
### Metadata Indexing [#metadata-indexing]
Table formats like Iceberg and Hudi maintain stats and min/max indexes for files and columns. Query engines use this metadata to avoid scanning irrelevant files.
* **Example**: A query for `amount > 10_000` will skip files where metadata shows `max(amount) < 5_000`.
* **Impact**: Reduces unnecessary scans and accelerates query planning.
### Auto-Tuning [#auto-tuning]
Some engines adjust execution parameters based on workload patterns, system memory, and concurrency.
* **Example**: Databricks Photon adapts join strategies and task parallelism based on observed query shapes. Trino adjusts spill thresholds and memory settings on the fly.
* **Impact**: Optimizes resource use without manual tuning.
### Materialized Views [#materialized-views]
Precomputes the result of expensive aggregations or joins and stores them as static tables for reuse.
* **Example**: A view that pre-aggregates sales by product and day lets BI tools load reports instantly without reprocessing raw data.
* **Impact**: Reduces repeated compute and shortens dashboard load times.
### Async Writes [#async-writes]
Async write paths ingest data in the background while still allowing concurrent reads on previously committed data.
* **Example**: Apache Hudi's "write optimized" mode writes data asynchronously to storage while exposing the latest snapshot to readers.
* **Impact**: Supports high-ingest workloads without blocking analytics or breaking consistency.
## Common Use Cases [#common-use-cases]
Lakehouses are versatile enough to power both analytics and AI workloads. They support a broad range of use cases without the need to move data between systems.
**Business Intelligence**
Teams build dashboards, KPIs, and custom reports directly from transaction, event, and operational data stored in tables like Delta Lake or Iceberg. SQL engines query fresh data without needing pre-aggregation into separate warehouses.
**Machine Learning**
Data scientists train, tune, and deploy models by connecting ML frameworks like TensorFlow or PyTorch directly to the underlying datasets. Feature engineering, training, and inference workflows operate on live production data without copying datasets to new environments.
**Event Analytics**
Lakehouses ingest high-throughput event streams such as app interactions, website clickstreams, and IoT sensor logs. These events are processed in real time and stored alongside historical data, enabling queries that span seconds to years without switching systems.
**Marketing Analytics**
Campaign data from ad platforms, CRM systems, and internal product usage feeds into the lakehouse for full-funnel attribution tracking. Marketers measure lead conversion rates, lifetime value by acquisition channel, and optimize spend based on unified performance data.
**Customer Segmentation**
User behavior, purchase history, geography, and account attributes are combined into dynamic cohorts. Segments are refreshed continuously based on updated data and used to drive targeted messaging, promotions, or personalized product experiences.
**Fraud Detection**
Lakehouses allow both real-time flagging of suspicious transactions and historical analysis of patterns. Machine learning models and rule-based systems operate on centralized log, transaction, and event data to detect anomalies and trigger automated responses.
**Product Analytics**
Teams track how users interact with application features, identify drop-off points in customer journeys, and monitor adoption rates. Clickstream events and user activity logs feed directly into analysis models without needing external product analytics tools.
**Centralized Governance**
Governance layers manage schema evolution, enforce access controls, track data lineage, and audit query activity across raw, intermediate, and curated datasets. Tools like Unity Catalog or AWS Glue catalog are used to secure and document all assets.
**Data Science Access**
Data scientists query operational, transactional, and event data directly in notebooks for exploration, hypothesis testing, and model development. Structured access controls and scalable compute resources enable experimentation without bottlenecks.
**Self-Serve Exploration**
Lakehouses expose SQL endpoints or BI connectors that allow business users to explore data sets, build reports, and create visualizations independently. Query performance and data permissions are managed centrally to maintain reliability and security.
## Data Lake vs Data Warehouse vs Lakehouse [#data-lake-vs-data-warehouse-vs-lakehouse]
These architectures differ in storage format, schema handling, cost structure, and supported workloads.
| Feature | Data Lake | Data Warehouse | Lakehouse |
| ------------------ | ------------------------ | ---------------------------- | ------------------------------------- |
| Storage | Object storage | Columnar storage engines | Object storage with format layer |
| Formats | Parquet, ORC, Avro | Proprietary (e.g., Redshift) | Open (Parquet, Delta, Iceberg) |
| Schema enforcement | Optional or late-binding | Strict | Configurable and flexible |
| Workloads | ML, batch processing | BI, dashboards | BI, ML, streaming, batch |
| Transactions | No | Yes | Yes (via Delta/Iceberg/Hudi) |
| Cost | Low | High | Low for storage, variable for compute |
| Governance | Limited, manual | Built-in | Varies, improving rapidly |
| Streaming support | Requires integration | Limited | Native in some engines |
| Tooling ecosystem | Fragmented | Mature | Growing and modular |
## Benefits of a Data Lakehouse [#benefits-of-a-data-lakehouse]
Data lakehouses combine the flexibility of data lakes with the performance and governance of data warehouses, offering several key advantages:
* **Cost Efficiency**: Utilize low-cost cloud object storage, reducing expenses compared to traditional data warehouses.
* **Unified Workflows**: Support both SQL analytics and machine learning workloads without the need to move data between systems.
* **Simplified Architecture**: Eliminate the need for multiple data platforms by consolidating storage and processing capabilities. [RightData](https://www.getrightdata.com/resources/private-chapter-5-data-lakehouse-challenges-and-benefits)
* **Open Standards**: Leverage open file and table formats, promoting interoperability and avoiding vendor lock-in.
* **Versatile Data Processing**: Handle both batch and streaming data, accommodating diverse data ingestion needs. [Medium](https://medium.com/data-engineering-with-dremio/the-data-lakehouse-the-benefits-implementation-challenges-and-implementation-solutions-5202925b4125)
* **Data Versioning**: Enable features like time travel and rollback, facilitating data auditing and recovery.
* **Cloud Integration**: Easily integrate with cloud-native services, enhancing scalability and flexibility.
* **Schema Flexibility**: Adapt to changing data schemas and pipelines without significant rework.
* **Unified Governance**: Implement consistent data governance and cataloging across all data assets.
## Challenges of a Data Lakehouse [#challenges-of-a-data-lakehouse]
While data lakehouses offer numerous benefits, they also present certain challenges:
* **Write Amplification**: Some table formats may experience write amplification, impacting performance.
* **Performance Tuning**: Achieving optimal performance often requires careful tuning and optimization.
* **Tooling Maturity**: The ecosystem for access control and data lineage is still evolving, potentially leading to gaps in functionality.
* **Learning Curve**: Managing hybrid workloads can involve a steep learning curve for teams accustomed to traditional systems.
* **ACID Compliance Variability**: Not all query engines fully support ACID transactions, which may affect data consistency.
* **Version Compatibility**: Ensuring compatibility across different tools and versions can be challenging.
* **Access Control Limitations**: Some stacks may lack advanced role-based access control (RBAC) or fine-grained access features.
* **Orchestration Requirements**: Separate orchestration and pipeline management tools are often necessary, adding complexity.
## Cost Considerations [#cost-considerations]
Data lakehouses cut down on storage costs by relying on cloud object storage. But total cost depends on how you manage compute and surrounding tools.
* **Storage**: Billed based on usage in services like S3, GCS, or Azure Blob, which are typically cheaper than warehouse storage.
* **Compute**: Costs vary by engine and usage. Trino charges per query execution, while Spark jobs are metered by runtime.
* **Caching**: Some platforms bill separately for memory usage or SSD-based caching layers.
* **Orchestration**: Requires third-party tools like [Airflow](https://www.firebolt.io/blog/analyzing-the-github-events-dataset-using-firebolt-incremental-updates-with-apache-airflow), Prefect, or dbt for managing pipelines and transformations.
* **Catalog and Governance**: Tools like Unity Catalog or AWS Glue may involve additional charges or licensing.
* **Maintenance**: Compared to managed warehouses, lakehouses often need more manual configuration and upkeep.
## Governance and Compliance [#governance-and-compliance]
Lakehouses offer strong security and compliance standards when set up across storage, compute, and metadata layers.
* **Access Control**: Tools like Unity Catalog or Hive Metastore allow catalog- and table-level permissions.
* **Column-Level Permissions**: Some ecosystems (Ranger, Delta Sharing) are introducing finer access granularity.
* **Encryption**: Encrypts data at rest and in transit using native cloud settings.
* **Auditing**: Tracks queries and access events through logs from compute engines like Spark or Trino.
* **Lineage**: Often added with external tools like Marquez or OpenLineage, or via custom metadata tracking.
* **Compliance Readiness**: Supports HIPAA, SOC 2, GDPR if properly configured, but not always out-of-the-box.
* **Policy Enforcement**: Ties into IAM systems, engine policies, and cloud storage permissions to enforce rules.
## When to Use a Lakehouse [#when-to-use-a-lakehouse]
A lakehouse fits when your workloads span both analytics and machine learning, and you want one platform.
Use a lakehouse when:
* You're working with structured and unstructured data together.
* You want SQL, notebooks, and model training access from a single dataset.
* You want to stop copying data between lakes and warehouses.
* Your pipelines include batch ETL, real-time streaming, and model training.
* You prefer open table formats over vendor-specific storage.
* You're scaling to terabytes or petabytes and want to keep storage costs low.
* You need flexible schemas to support evolving pipelines and schemas.
## Popular Lakehouse Platforms [#popular-lakehouse-platforms]
A range of vendors and open projects support the lakehouse model using modular compute and open table formats:
* **Databricks + Delta Lake**: Combines Spark-based compute with ACID transactions and ML integrations.
* **Apache Iceberg**: Backed by Dremio, AWS Athena, Starburst, and Snowflake for scalable table format support.
* **Apache Hudi**: Optimized for incremental loads and upserts, which are useful for streaming-heavy pipelines.
* **AWS Lake Formation**: Native AWS service that ties into Glue, S3, and Athena for managed governance and access.
* **Google BigLake**: Bridges BigQuery with GCS, providing a single query interface over structured and unstructured data.
* **Snowflake + Unistore**: Expanding toward lakehouse features with support for unstructured and transactional data.
* **Dremio**: Lakehouse platform built on Apache Arrow and Iceberg with strong SQL support.
* **Trino**: Query engine that works with Hive Metastore or Nessie catalogs for lakehouse-style analytics.
* **Starburst Galaxy**: A fully managed Trino platform designed for federated, lakehouse-compatible workloads.
* **OpenLakehouse**: Community project focused on interoperability across engines, formats, and catalogs.
## FAQs [#faqs]
### What makes a data lakehouse different from a traditional warehouse? [#what-makes-a-data-lakehouse-different-from-a-traditional-warehouse]
A lakehouse uses cloud object storage and open file formats, but adds table management and compute layers to support SQL and ML workloads.
### Can a lakehouse replace both a lake and a warehouse? [#can-a-lakehouse-replace-both-a-lake-and-a-warehouse]
Yes. It can handle structured and unstructured data, support multiple compute engines, and serve use cases across analytics and data science.
### Can it handle real-time data? [#can-it-handle-real-time-data]
Yes. Some formats like Hudi and Delta Lake support streaming ingest. Engines like Spark and Flink process updates continuously.
### What kind of data works best in a lakehouse? [#what-kind-of-data-works-best-in-a-lakehouse]
Lakehouses work well with semi-structured data like JSON, log files, clickstreams, and large-scale structured data in formats like Parquet.
### Do lakehouses support transactional consistency? [#do-lakehouses-support-transactional-consistency]
Yes, depending on the table format. Delta Lake, Apache Iceberg, and Apache Hudi all provide varying levels of ACID guarantees.
### How do lakehouses handle schema changes? [#how-do-lakehouses-handle-schema-changes]
Most support schema evolution. You can add or modify columns without rewriting historical data, though compatibility varies by format.
### Which tools can query a data lakehouse? [#which-tools-can-query-a-data-lakehouse]
Common engines include Spark, Trino, Dremio, Presto, and Snowflake. BI tools like Tableau and Power BI can connect via JDBC or ODBC.
### Is a lakehouse cheaper than a warehouse? [#is-a-lakehouse-cheaper-than-a-warehouse]
Storage is usually cheaper since it uses object storage. Compute costs depend on engine choice, workload size, and tuning.
### What governance tools are available? [#what-governance-tools-are-available]
You can use AWS Glue, Unity Catalog, Hive Metastore, Apache Ranger, or third-party tools for access control, auditing, and lineage.
### Can I use my existing data lake as a lakehouse? [#can-i-use-my-existing-data-lake-as-a-lakehouse]
If your data is in Parquet or ORC and stored in object storage, you can add a table format layer like Iceberg or Delta to start querying it like a lakehouse.
# Data Mesh - Definition (/glossary-items/data-mesh)
In 2019, Zhamak Dehghani had a stroke of genius, which brought to life an organizational, architectural, technology-agnostic framework for big data molded by the principles of mesh networking. This framework became known as data meshing—a model based on the concept of decentralization.
Literally, a data mesh is a dense network of nodes containing data. Unlike the centralized and monolithic architectures of data warehouses and [data lakes](https://www.firebolt.io/glossary-items/data-lake), a data mesh is a highly decentralized architecture.
The use of centralized data silos, aimed at creating a single source of truth (SSOT) usually populated with extract-transform-load ([ETL](https://www.firebolt.io/glossary-items/etl-and-elt)) processes, has historically led to big challenges, like bootstrapping the monolith and scaling in terms of both the number of consumers of data and computational resources. In addition, users of centralized monoliths typically have difficulty in finding quality data and interpreting them due to a lack of **data ownership**.
The first and most important principle of data meshes is the idea of *domain-oriented decentralized data ownership and architecture*. It aims at resolving a specific problem: moving the analysis of data to the same domain where the data itself was born. In this approach, the experts of the domain from which the data has come are the same responsible for its quality and interpretation. And due to their experience in the domain, the problems regarding interpretation and quality of the data automatically decrease.
Data mesh applies DDD (domain-driven design) to data-based architectures. The first implementation of DDD within software architecture was microservices. Just as microservices are software components that expose elementary application functionality, **data products** are software components that expose elementary data and analytical functionality.
As with microservices architectures, the data mesh disaggregated model requires a series of rules and tools to follow. In particular, data products must comply with the rules identified in the **DATSIS** acronym, which states that a product must be: **discoverable, addressable, trustworthy, self-describing, integrated, and secure.**
Moreover, in applying a data mesh architecture, the organization risks seeing an increase in the heterogeneity of the technologies, making the sourcing strategy very complex. It is, therefore, important for a company to equip itself with a **data infrastructure platform**—like PaaS cloud platform—which allows standardization of data product development and the introduction of a common language between the various domains.
Furthermore, it is very important that the company equips itself with **ecosystem governance tools** to provide comprehensibility and visibility of data products. On these tools, each data product can be registered and consequently searchable and reusable according to specific authentication and authorization policies.
Data meshing is not a strict technical implementation. It's a transformation of data software architecture that promotes moving from a monolithic data lake to a distributed data-driven architecture. Its features include moving from centralized data ownership to decentralized ownership, shifting from a pipeline-as-a-first-concern understanding to a domain data-as-a-first-concern understanding, moving from creating siloed data engineering teams to cross-functional domain data teams, and focusing on data as a product instead of a centralized data lake/[data warehouse](https://www.firebolt.io/glossary-items/data-warehouse).
At times, some people may confuse a data mesh and a data fabric. The difference between the two is that a data fabric uses a decidedly architectural approach to data access, whereas a data mesh architecture is more about connecting users with data processes.
## Advantages of Data Mesh Implementation [#advantages-of-data-mesh-implementation]
Let's now enucleate the principal advantages of implementing a data mesh architecture.
* **Agility**: As a data mesh is a distributed architecture, it enables decentralized data operations and increased team performance. This, in turn, improves time to market, scalability, and agility due to reduced complexities and IT backlogs.
* **Resilience to Technological Progress:** A data mesh paradigm represents a very strong guarantee against the risk of technological obsolescence. In the future, when new technologies emerge, any **data product** will be able to adopt them without a problem.
* **Speed**: Thanks to its well-governed and decentralized data storage, data mesh provides a simple, well-governed, and centralized infrastructure based on self-service for faster and secure access to data.
* **Decoupling**: As with microservices, data mesh brings the great advantages of a decoupled infrastructure, granting, as a result, easier scalability, technological independence, and organizational independence.
* **Flexibility**: A data mesh provides flexibility to enterprise organizations to become more vendor agnostic. Additionally, as a result of the inherently decentralized nature of data meshes, the individual domains become responsible for the quality, security, and transfers. This connectivity enables all kinds of users to have access regardless of their physical location, overcoming the problems of traditional data sharing that falls under the scrutiny of international guidelines.
* **Compliance**: The distributed architecture reconciles data ingestion with its sources and formats to allow businesses to control security and helps create a compliant data platform.
* **Access**: Data meshes prevent the creation of data silos, improving data access for cross-functional teams and transparency. Its unified infrastructure enables business domains to share the data products efficiently, enforcing standards.
* **Data Governance**: The distributed architecture of data meshes allows businesses and companies to control their security from the source, simplifying compliance with global data governance policies.
## Challenges of Data Mesh Implementation [#challenges-of-data-mesh-implementation]
Adopting the data mesh framework comes with several challenges caused primarily by the inhomogeneity of software systems that compose the nodes of the network.
* **Specialization**: Data mesh implementation requires specialists to create domain-specific ETL, data lake implementations with complex data systems, and so on.
* **Data Redundancy**: Data mesh makes data governance more difficult to manage due to its multi-cloud and hybrid infrastructure. Redundancy occurs when the data of one domain is repurposed to serve the business needs of another domain, which can impact resource usage and data management costs.
* **Adoption Costs**: Decentralizing data management to adopt data mesh implementation requires major changes when switching from a highly centralized data architecture. Ecosystem governance tools and data infrastructure platform tools needed to maintain a good quality data mesh solution come with the cost of bootstrapping and maintenance.
* **Complexity**: Enterprise-wide data models must be defined to merge various data products and make them available to authorized users in a central location.
# Data Nesting and Data Unnesting - Definition (/glossary-items/data-nesting-and-data-unnesting)
Data nesting and data unnesting are the paired processes of transforming data between a flat, relational structure and a hierarchical, nested one. Nesting composes fields into structured, multi-level records, while unnesting expands those nested structures back into flat rows. Together they are the inverse of data flattening and unflattening.
## What is data nesting and data unnesting [#what-is-data-nesting-and-data-unnesting]
Data nesting and data unnesting is the opposite of data flattening and data unflattening. In order to nest data from a "flattened" form such as a relational model, you need extra information about how to nest. When you unnest, you lose information.
# What is Data Observability? (/glossary-items/data-observability)
Data observability is concerned with the health of your data, which is obtained from a variety of sources. It entails much more than just keeping an eye on things. As companies are gathering data from a variety of sources, they have become much more reliant on their data for everyday operations and decision-making. As a result, it is vital to guarantee that data is delivered in a timely and high-quality manner. When a large amount of data is transported around an organization, typically for analytical purposes, data pipelines serve as the central repository for the data. Data observability contributes to the assurance of a dependable and effective flow of information.
There are five pillars of data observability that must be met. Let us take a quick look at them.
Freshness is vital for the business since we need to know how up-to-date your data is to make informed judgments about our products and processes. When it comes to decision-making, freshness is extremely crucial. As we all know, stale and obsolete data is a waste of time, money, and human resources.
Specifically, volume describes the completeness of your data tables and gives critical information about the health of your data sources, i.e. when garbage or poor data is sent out, it indicates that the data sources are out of current and that they must be updated. A vast number of organizations rely on data-processing platforms such as Google Big Query and Snowflake to process their data and generate meaningful insights for their business decisions. Data scientists can analyze massive amounts of data in real time using BigQuery, which is totally self-contained.
The accuracy of data is crucial for the development of high-quality and trustworthy data systems for an organization. When we talk about distribution, we are talking about the measure of diversity in the system. If the data in the system wildly differs from one another, it is possible that there is a problem with the data accuracy. The quality of data produced and consumed by the data system is the main focus of the distribution. With distribution as part of your data observability stack, you can keep an eye out for anomalies in your data values and prevent erroneous data values from being injected into your system.
Schema updates are unavoidable because every organization is always growing and adding new features, which in turn affects the application database. Changes to the schema, on the other hand, that are not thoroughly tested and controlled can result in downtime for your application. The schema in the data observability ensures that database schemas, such as data tables, fields, columns, and names, are accurate, up to date, and subject to regular auditing and validation.
To effectively manage and monitor your data system's health, you must have a complete picture of your data ecosystem. The ease with which data can be traced via our data system is referred to as lineage. A unified picture of your data, or a blueprint of your data, is created with the help of lineage.
## Advantages [#advantages]
* **Cost-effective**: Because of the elastic nature of the cloud, which stores and processes your data, if you need to load data faster or run a high volume of queries, you can scale up your [data warehouse](https://www.firebolt.io/glossary-items/data-warehouse) to take advantage of more processing capabilities, which will lower your costs. After that, you can minimize the size of the warehouse and only pay for the time that was spent in it. As a result, it becomes more cost effective as well. Once you have analyzed the data, you can eliminate the false positives to save even more resources.
* **Insights**: Data observability identifies circumstances that you aren't aware of or that would otherwise go undetected, allowing you to avert problems before they have a significant impact on the business. Observability of data can be used to track the relationships between individual issues and to offer context and relevant information for root cause analysis and repair.
* **Alerting**: Data observability ensures that teams are notified when there is a data issue, allowing them to resolve it swiftly and save everyone from the consequences of a data outage. These alerts should ideally be configured to send a notification as soon as an anomaly is detected, which will save a significant amount of time and resources in the long run.
* **Cross-Operability**: Because of the system's architecture, users are able to share information with one another. It is very simple for businesses to exchange information with every data consumer who uses them. The fact that these data availability tools are mostly used in the cloud, that they are dispersed across availability zones of the platform on which they operate—either Amazon Web Services or Microsoft Azure—and that they are designed to function continuously and to suffer component and network failures with minimal impact, the impact on customers is kept to an absolute bare minimum.
## Challenges [#challenges]
* **Data Vigilance**: There are a variety of tools available that may provide you with unlimited storage or allow you to pay only for what you use. Another concern that may develop in this situation is the use of bad data. In the event that you continue to provide data into the system, particularly without confirming it, this could prove to be rather costly for you. Therefore, rules must be established to review and authenticate data before it is uploaded to the cloud for storage.
* **Process Determination**: Before starting, it is also vital to establish the method we should use to process the data, such as extract, transform, load ([ETL](https://www.firebolt.io/glossary-items/etl-and-elt)) or extract, load, transform ([ELT](https://www.firebolt.io/glossary-items/etl-and-elt)). The transformation will be performed before the data is loaded, whereas with ELT, the transformation will be performed after the data has been loaded.
* **Performance**: When dealing with a large amount of data, performance is essential because the faster you can load the data, the faster you can complete operations. Even when dealing with large amounts of data, performance cannot always be guaranteed due to the nature of the material being handled.
* **Complexity**: The complexity of data may develop as an organization begins to collect information from a number of diverse sources, each of which may have a unique set of characteristics. A large number of organizations do not comprehend the complexity of the situation, and those organizations, as well as their instruments, are not prepared to manage the situation.
# What is Data Orchestration? (/glossary-items/data-orchestration)
Data orchestration is the automated process of combining data from multiple sources, transforming it to a common standard, and organizing it in a centralized location so that teams across an organization can access and use it. Data orchestration frameworks enforce the relationships and dependencies between tasks, allowing multiple tasks to run simultaneously while pipelines are created, monitored, scheduled, and alerted on programmatically.
An organization's data needs can be met from a variety of sources. The sources of data might include historical data, batch data, and streaming data. This data can be stored in different databases and in different forms. This creates silos, which act as barriers for teams to collaborate, leading to poor data practices. So, it is imperative to remove silos and make the data more accessible and available across the organization.
Before the advent of big data, data collection and transformation jobs were done manually. Developers used to run scripts and monitor the results of each job manually. But as the amount of data increased by multiple folds, the manual option became out of the question. So, developers started to use cron jobs for scheduling tasks. But after a point, even cron jobs weren't sufficient to handle the scale. Monitoring, handling dependencies, and rerunning failed jobs out of 1000s of jobs was tedious. A lot of time and effort was required to log details such as task runtime, the status of the job, dependencies required, and scheduling details. Data engineers were in need of a system that offered a holistic view of all these tasks in a centralized location.
These two wants led to the rise of data orchestration frameworks, which enabled developers to programmatically create, monitor, schedule, and alert pipelines with ease. With a data orchestration framework, we can combine data from multiple sources, transform them to a fixed standard and store them in a centralized location so that multiple teams can use this data to power their applications.
Data orchestration tools enforce the relationship and dependency between the tasks. Multiple tasks can be run simultaneously in this approach. This is done through the creation of a Directed Acyclic Graph or DAG, which is a series of tasks that can be triggered manually or through scheduling. A typical DAG involves waiting for data, collecting the data, sending the data to a different source for transformation, monitoring the state of the task, getting the resultant data after preprocessing and storing it in a centralized source.
There are many popular data orchestration tools available in the market. A few of these are:
* **Apache Airflow**: Developed by AirBnB, it is one of the most popular data orchestration tools available. It uses a web interface and command-line utilities to author workflows as directed acyclic graphs.
* **AWS Data Pipeline**: Developed by Amazon Web Services, it helps in copying and transforming data that is stored on AWS.
* **Google Cloud Dataflow**: Developed by Google, it helps in building and managing data pipelines. It can process both batch and stream data.
* **Microsoft Azure Data Factory**: Developed by Microsoft, it is a cloud-based data integration service that helps in creating data-driven workflows for orchestrating and automating data movement and transformation.
* **Luigi**: Developed by Spotify, it uses Python to build complex pipelines and handle dependencies programmatically.
These orchestration tools offer different features, and it is essential to choose the right tool based on the use case. But one of the most important factors while choosing an orchestration tool is its integration with other tools and technologies in the organization's data ecosystem. Another factor is the ability to monitor and track the progress of each task in the DAG. And finally, ease of use and debugging capabilities are some other factors that should be considered before choosing an orchestration tool.
## Advantages of Data Orchestration [#advantages-of-data-orchestration]
* **Cost and time-efficiency**: Managing thousands of pipelines is a very intensive operation, both in terms of time and manpower. Developers have to spend valuable time going through the logs, which could be better spent elsewhere. Since orchestration takes care of repetitive tasks, developers can focus on intuitive problem-solving. As there is very little to no human intervention, the possibility of human error decreases as well. This leads to improvement in overall productivity.
* **Breaking silos**: In addition to getting the right data we must ensure that we get the data at the right time. Having data silos within an organization delays the process of getting data at the right time. Due to the centralized nature of data orchestration, the concept of silos is broken. This in turn enables the organization to make informed decisions in a timely manner before it's too late.
* **Improved data quality**: An organization might get its data from [data lakes](https://www.firebolt.io/glossary-items/data-lake), [data warehouses](https://www.firebolt.io/glossary-items/data-warehouse), blobs, message queues, relational databases, NoSQL databases, streaming services, etc. Every database will have its own format. Some of the data may be structured, some might be semi-structured and the rest will have no structure at all. Even for structured data, each business follows its own convention.
For example, let's take the state of New York. Some of the different notations are NY, Newyork, NYC, New York City, and NYS. This leads to extreme disparities in the data. So we have to convert the data based on a gold standard. This makes data accessible across the organization. As data orchestration involves defining streamlined workflows, which take care of the transformations, it helps in maintaining the quality of data.
* **Monitoring**: Monitoring thousands of pipelines is extremely tedious. There are a lot of parameters and dependencies that need to be monitored. Since the orchestration tool provides centralized access to monitor all the workflows we can easily identify the tasks that require handling.
* **Easier migration**: Due to the rise of different cloud service providers, a lot of companies are looking to migrate their on-premises data to [cloud data warehouses](https://www.firebolt.io/blog/cloud-data-warehouse). Migration should be done efficiently without the loss of data. This is a complex process in itself, but if we are required to perform incremental migration then the complexity increases exponentially. As data orchestration tools have efficient checks and connect to multiple services, they make migration easier.
### Challenges of Data Orchestration [#challenges-of-data-orchestration]
* **Integration capabilities**: The data orchestration tool must be able to connect with a lot of services. It must have built-in plugins to connect to all the standard databases and also support custom plugins to connect to newer ones. This is extremely important because the orchestration tool must connect with a plethora of sources to get and send data. The number of sources an orchestration tool can support determines its reach.
* **Ease of setting up**: The ease of setting up and running an orchestration tool is another important challenge. We have to set up the storage and compute infrastructure for the orchestration tool to run. Additionally, we may have to configure the network as well. Setting up data orchestration tools can be complex. The orchestration tool must also have less maintenance overhead.
* **Regulations**: GDPR and HIPAA have rigorous requirements for the security, use of private data, encryption, and storage location. In a data orchestration system, the data keeps moving from one source to another so it will be a challenge to comply with these regulations.
# What is a Data Warehouse? A Deep Dive for Data Engineers (/glossary-items/data-warehouse)
Data warehouses are the analytical backbone of modern enterprises, enabling data engineers to store, transform, and analyze massive datasets at scale. Unlike traditional databases, a data warehouse is designed for high-speed analytics, data transformation, and decision intelligence.
## Why Data Engineers Must Master Data Warehousing [#why-data-engineers-must-master-data-warehousing]
* **Efficient Query Processing:** Optimizing query execution plans, indexing, and storage formats can reduce query times from minutes to milliseconds.
* **Scalability & Cost Optimization:** Choosing between columnar storage, partitioning, and indexing strategies can cut cloud costs by 5-10x.
* **High-Speed ETL/ELT Pipelines:** Mastering Extract, Transform, Load (ETL) vs. Extract, Load, Transform (ELT) ensures seamless data movement at scale.
* **Concurrency Handling:** Implementing strategies to support thousands of simultaneous queries efficiently.
* **Real-Time Analytics:** Leveraging low-latency query execution to support real-time dashboards and decision-making.
## Introduction: Why Data Warehousing is Critical for Data Engineers [#introduction-why-data-warehousing-is-critical-for-data-engineers]
In the modern data ecosystem, **data warehouses** serve as the analytical backbone of enterprises, enabling efficient data storage, transformation, and retrieval at scale. Unlike traditional relational databases optimized for transactional workloads, data warehouses are purpose-built for analytical queries, high-speed aggregations, and complex transformations over massive datasets.
For **data engineers**, mastering data warehousing is no longer optional—it's essential for building scalable, high-performance data pipelines that power real-time analytics, AI-driven insights, and business intelligence (BI) applications.
## Why Data Engineers Must Master Data Warehousing [#why-data-engineers-must-master-data-warehousing-1]
* **Efficient Query Processing:** Optimizing query execution plans, storage formats, and indexing techniques can reduce query times from minutes to milliseconds.
* **Scalability & Cost Optimization:** Strategic use of columnar storage, partitioning, and indexing can cut cloud compute costs by 5-10x.
* **High-Speed ETL/ELT Pipelines:** Choosing the right Extract, Transform, Load (ETL) vs. Extract, Load, Transform (ELT) approach ensures seamless data movement at scale.
* **Concurrency Handling:** Efficiently supporting thousands of simultaneous queries without performance degradation is crucial for large-scale analytics.
* **Low-Latency Analytics:** Leveraging distributed query execution, indexing, and caching allows for sub-second response times, enabling real-time dashboards and decision-making.
## Core Data Warehouse Architecture: Deep-Dive Analysis [#core-data-warehouse-architecture-deep-dive-analysis]
A well-optimized data warehouse follows a structured architecture to manage large-scale analytics workloads efficiently.
### The Three-Tier Data Warehouse Architecture [#the-three-tier-data-warehouse-architecture]
Data warehouses typically follow a three-tier architecture to optimize performance, scalability, and cost.
| Tier | Purpose | Examples |
| ----------- | -------------------------------------------------------------------- | ----------------------------------------------------- |
| Bottom Tier | Physical storage layer (data lakes, columnar storage, MPP databases) | Data lakes (S3, ADLS), columnar stores (Parquet, ORC) |
| Middle Tier | Query processing layer (OLAP engines, SQL optimizations, caching) | Firebolt, Snowflake, Redshift, BigQuery |
| Top Tier | BI & analytics tools (Looker, Tableau, Power BI, direct SQL access) | Looker, Tableau, Power BI |
**Optimization Strategies for High-Performance Data Warehousing**
* **Columnar Storage:** Store data in compressed columnar formats (Parquet, ORC) to minimize I/O and scan costs.
* **Multi-Layer Indexing:** Implement sparse indexing, partition pruning, and zone maps for ultra-fast lookups and data skipping.
* **Distributed Query Execution:** Use MPP engines like Firebolt, Snowflake, and BigQuery to parallelize query execution across nodes, reducing compute overhead.
* **Metadata Caching:** Optimize query metadata caching to avoid repeated computations, reducing query latency.
* **Data Lifecycle Management:** Automate tiered storage policies to balance cost and performance by offloading cold data to cost-efficient storage.
* **Real-World Example:** A global e-commerce company switched from PostgreSQL to a Firebolt-based cloud data warehouse, reducing query execution time by 85% by adopting a columnar storage model and sparse indexing.
## Query Performance Optimization: Making Queries Faster [#query-performance-optimization-making-queries-faster]
### 1. Partitioning & Clustering: Reducing Scan Costs [#1-partitioning--clustering-reducing-scan-costs]
Partitioning and clustering help reduce the amount of data scanned, making queries significantly faster.
**Example: Date-Based Partitioning (Best for Time-Series Data)**
```sql
CREATE TABLE transactions (
transaction_id INT,
transaction_date DATE,
customer_id INT,
amount DECIMAL
)
PARTITION BY RANGE(transaction_date);
```
**Impact:** Queries filtering on transaction\_date skip unnecessary partitions, improving query performance by **80%.**
**Example: Region-Based Partitioning (For Geo-Distributed Data)**
```sql
PARTITION BY LIST(region);
```
**Impact:** Reduces I/O operations by skipping irrelevant regions in analytical queries.
**Example: Clustering for Faster Joins**
```sql
ALTER TABLE transactions CLUSTER BY (customer_id, transaction_date);
```
**Impact:** Reduces shuffle and sort operations, significantly improving join performance.
### 2. Sparse Indexing for Large-Scale Queries [#2-sparse-indexing-for-large-scale-queries]
**Example: Creating an Index for Fast Customer Lookup**
```sql
CREATE INDEX idx_customer_id ON transactions (customer_id);
```
**Impact:** Accelerates JOIN performance on customer\_id, reducing lookup times by 70%.
### 3. Materialized Views for Precomputed Aggregations [#3-materialized-views-for-precomputed-aggregations]
**Example: Precomputing Monthly Revenue**
```sql
CREATE MATERIALIZED VIEW monthly_revenue AS
SELECT DATE_TRUNC('month', transaction_date) AS month, SUM(amount)
FROM transactions
GROUP BY month;
```
**Impact:** Cuts query latency from 30 seconds to milliseconds.
## Cloud Data Warehousing: Firebolt vs. Snowflake vs. Redshift vs. BigQuery [#cloud-data-warehousing-firebolt-vs-snowflake-vs-redshift-vs-bigquery]
| Feature | Firebolt | Snowflake | Redshift | BigQuery |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------- | ----------------- | ---------- |
| Query Latency | ⚡ Sub-second | Seconds | Seconds-minutes | Seconds |
| Storage Model | Decoupled compute & storage | Shared storage | Localized storage | Serverless |
| Indexing | Sparse indexing & data skipping | None | Zone maps | None |
| Concurrency | Thousands of queries per second | Medium | Medium | High |
| Cost Efficiency | Firebolt is: 7.5 to 8x better in price-performance compared to Snowflake; 7 to 11x better in price-performance compared to Redshift; 90x better in price-performance compared to BigQuery | | | |
**Key Takeaway:** Firebolt is optimized for sub-second analytics performance, offering better indexing, cost efficiency, and concurrency scaling compared to traditional cloud data warehouses.
## Data Warehouse vs. Data Lake vs. Lakehouse: Which One Do You Need? [#data-warehouse-vs-data-lake-vs-lakehouse-which-one-do-you-need]
| Feature | Data Warehouse | Data Lake | Lakehouse |
| ----------- | ------------------------------ | -------------------- | ---------------------- |
| Schema | Schema-on-write | Schema-on-read | Hybrid |
| Performance | High (optimized for analytics) | Low (raw storage) | Medium |
| Use Case | BI, reporting | Machine Learning, AI | Mixed workloads |
| Examples | Firebolt, Snowflake, Redshift | Hadoop, S3, ADLS | Databricks, Delta Lake |
**Key Takeaway:** A Lakehouse (e.g., Databricks) combines the best of data lakes and warehouses, enabling schema flexibility and analytics performance.
## Best Practices for Building a High-Performance Data Warehouse [#best-practices-for-building-a-high-performance-data-warehouse]
✔ Use ELT over ETL for faster cloud-based data ingestion.
✔ Leverage partition pruning to minimize data scans.
✔ Adopt columnar storage (Parquet/ORC/Delta Lake) for faster analytics.
✔ Implement sparse indexing to accelerate lookup queries.
✔ Use query caching & materialized views to reduce recomputation overhead.
✔ Optimize JOINs with broadcast vs. shuffle strategies.
✔ Automate workload management for query concurrency optimization.
✔ Monitor query performance and apply adaptive optimizations based on workload patterns.
## Conclusion: The Future of Data Warehousing [#conclusion-the-future-of-data-warehousing]
Data warehouses are powerful tools for **BI reporting and analytics, providing a centralized, accurate source of data for decision-making**. They consolidate structured and semi-structured data into a single repository, ensuring data integrity, accessibility, and performance at scale.
While challenges such as ETL complexity, schema evolution, and cost management exist, advancements in cloud-based, elastic architectures have significantly mitigated these issues. By leveraging columnar storage, distributed query execution, and automated indexing, organizations can achieve unmatched performance and scalability.
**Key Takeaways:**
* Optimize queries with partition pruning & indexing to reduce query execution times.
* Evaluate Firebolt for sub-second analytics performance with its decoupled storage & compute architecture.
* Leverage ELT workflows for modern cloud-native pipelines, enabling faster data processing and lower operational overhead.
* Adopt AI-driven query optimization techniques to ensure automated tuning and workload-aware performance improvements.
## Firebolt: The Fastest Cloud Data Warehouse For Data-intensive AI Applications [#firebolt-the-fastest-cloud-data-warehouse-for-data-intensive-ai-applications]
Firebolt provides the fastest cloud data warehouse with the performance to support ad hoc and high-performance analytics at scale, as well as semi-structured data analytics. Unlike traditional data warehouses, Firebolt is built for modern data and AI applications that require:
* **Sub-second query performance**
* **Massively parallel processing (MPP)**
* **Decoupled storage & compute for scalability**
* **Sparse indexing & optimized caching for faster analytics**
* **High user concurrency with real-time query execution**
**Firebolt's cloud-native design allows companies to:**
* **Run thousands of concurrent queries** without performance degradation.
* **Optimize costs with pay-as-you-go pricing**, eliminating over-provisioning.
* **Enable instant scaling** to handle unpredictable workloads.
🔗 **Explore Firebolt:** [Firebolt.io](https://www.firebolt.io)
# Data Warehousing - Definition (/glossary-items/data-warehousing)
Data warehousing is the process of developing, deploying, and managing a data warehouse. In order to perform data warehousing you also need to perform data integration, which is the process of integrating multiple data sources together for analytics, data migration, or operations. You will also have analytics or business intelligence processes.
They are all separate processes, though they should work well together so that you can build new reports quickly.
## What is data warehousing [#what-is-data-warehousing]
Data warehousing is the process of developing, deploying and managing a data warehouse. In order to perform data warehousing you also need to perform data integration, which is the process of integrating multiple data sources together for analytics, data migration or operations. You will also have analytics or business intelligence processes.
They are all separate processes, though they should work well together so that you can build new reports quickly.
## Challenges [#challenges]
Apart from the technical challenges a [cloud data warehouse](https://www.firebolt.io/the-cloud-data-warehousing-guide) or any [data warehouse](https://www.firebolt.io/glossary-items/data-warehouse) faces, data warehousing face a combination of challenges that force most companies to find their own balance between agility, consistency and governance.
* **Agility:** End users want and need new reports in days. But for many organizations, it can take 1-3 weeks to explain all the requirements for a new report, and 1-3 months to deliver it. Part of the reason is it can take 1-3 iterations to get the report right, and as many data warehouse administrators will tell you, if you touch a data warehouse it takes a month to test and roll out to production.
* **Consistency:** Data needs to be extracted from various applications and other sources, integrated, merged, cleansed and made consistent. This requires some combination of data integration, data quality, and other tools such as master data management.
* **Governance:** Beyond the tools and best practices required to ensure a level of data quality and consistency, there are also all kinds of internal business, customer, security, and regulatory requirements for different types of data.
Making data more accessible and shortening the time to new reports can come at the cost of data quality and consistency, and data governance.
## Benefits [#benefits]
The biggest benefit of data warehousing and its related discipline of data governance is visibility. The more data you can expose, and the more agile users can be, the more employees can improve business performance, and the happier customers will be. But bad, ungoverned or exposed data can lead to bad decisions, data breaches or regulatory fines.
Increasingly, data warehousing has incorporated the use of data lakes for holding raw data, which has improved both data access and agility. You end up with two processes. The more traditional data integration and data governance processes are focused on moving data from the data sources into the data lake. These processes can still take weeks to months.
But they do enable faster data access by data engineers. Now whenever people need new reports, data engineers can find the best data in the data lake and load it on their own using SQL-based [ELT (or ETL)](https://www.firebolt.io/glossary-items/etl-and-elt). This can shorten new report creation from weeks to hours.
## Firebolt and data warehousing [#firebolt-and-data-warehousing]
Firebolt is a cloud data warehouse that supports most traditional and newer data warehouse processes. Many companies use Firebolt with a data lake, and use Firebolt for SQL-based ELT, to help deliver new reports much faster as well as deliver more self-service analytics, including more interactive ad hoc, operational and customer-facing analytics.
# DDL and DML - Definition (/glossary-items/ddl-and-dml)
Databases are essential for storing and managing data in modern applications. But behind every well-organized database is a system that defines its structure and handles its data — this is where DDL and DML come in. Some industry analysts prefer the term Database Management System (DBMS) over "database" because, technically, a database is just the data store, while a DBMS combines the data store with the computational logic that defines, manages, and manipulates that data.
Most DBMSs operate using two core components: Data Definition Language (DDL) and Data Manipulation Language (DML). Together, DDL and DML handle the computation within a DBMS, while the database itself stores the data. In a Relational Database Management System (RDBMS), these operations are carried out using SQL.
DDL commands such as CREATE, ALTER, and DROP define and modify database structures, while DML commands like SELECT, INSERT, UPDATE, and DELETE manage the data within those structures.
Understanding the role of DDL and DML in database operations is essential for working with both transactional and analytical database systems. Read on to explore their roles and applications.
## What Is DDL (Data Definition Language)? [#what-is-ddl-data-definition-language]
Data Definition Language (DDL) consists of [SQL](https://www.firebolt.io/blog/simplifying-time-variance-in-a-sql-data-warehouse) commands that establish a database's framework. It is responsible for setting up structures such as tables and schemas, modifying their design, and removing them when they are no longer needed.
When setting up a database, DDL commands define tables, indexes, and relationships. These commands modify structures and remove database objects. Common examples include:
* **CREATE:** Defines new database objects like tables, indexes, or entire databases.
* **CREATE DATABASE my\_database**: Creates a new database.
* **CREATE TABLE users (id INT PRIMARY KEY, name VARCHAR(50), email VARCHAR(100))**: Defines a table with columns and data types.
* **CREATE INDEX idx\_email ON users(email)**: Generates an index for faster lookups.
* **ALTER:** Modifies existing database structures without affecting stored data.
* **ALTER TABLE users ADD COLUMN phone VARCHAR(20)**: Adds a new column to the users table.
* **ALTER TABLE users MODIFY COLUMN name VARCHAR(100)**: Changes a column's data type.
* **DROP:** Deletes database objects permanently.
* **DROP TABLE users**: Removes the users table and its data.
* **DROP INDEX idx\_email ON users**: Deletes an index. (SQL Server syntax: DROP INDEX users.idx\_email;)
* **TRUNCATE:** Clears all records from a table without removing its structure.
* TRUNCATE TABLE users: Empties the users table while keeping its schema intact.
### DDL Commands in SQL [#ddl-commands-in-sql]
Here are DDL statements following [MySQL and PostgreSQL standards](https://www.firebolt.io/blog/making-a-query-engine-postgres-compliant-part-i-functions):
```sql
CREATE TABLE users (
id INT PRIMARY KEY,
name VARCHAR(50),
email VARCHAR(100)
);
ALTER TABLE users ADD COLUMN phone VARCHAR(20);
DROP INDEX idx_email ON users; -- MySQL
-- DROP INDEX idx_email; -- PostgreSQL
TRUNCATE TABLE users;
```
For SQL Server, syntax varies slightly:
```sql
CREATE TABLE users (
id INT PRIMARY KEY,
name VARCHAR(50),
email VARCHAR(100)
);
ALTER TABLE users ADD phone VARCHAR(20); -- No "COLUMN" keyword in SQL Server
DROP INDEX users.idx_email; -- SQL Server requires table name prefix
```
In summary, DDL provides the foundation for database organization by defining structures, enforcing constraints, and allowing modifications as data needs evolve. By creating and managing tables, indexes, and relationships, it ensures databases remain structured, scalable, and efficient.
## What Is DML (Data Manipulation Language)? [#what-is-dml-data-manipulation-language]
Data Manipulation Language (DML) is a component of SQL used to insert, retrieve, modify, and delete records within a database structure. Unlike DDL, which defines database objects like [tables](https://www.firebolt.io/blog/simplicity-and-power-of-agg-indexes-at-scale-materialized-views-simplified) and indexes, DML manages data within those structures.
The most frequent examples of commands include:
**INSERT:** Adds new rows to a table.
**\*Example\*\***: The following statement inserts a new customer record into the customers table:\*
```sql
INSERT INTO customers (first_name, last_name, email)
VALUES ('Brian,' 'David,' 'briandavid@domain.com');
```
**UPDATE:** Modifies existing data within a table.
**\*Example\*\***: Updating a customer's email address based on their id:\*
```sql
UPDATE customers
SET email = 'briandavid@domain.com'
WHERE id = 123;
```
**DELETE:** Removes specific rows from a table.
**Example**: This statement deletes all customer records where the id is greater than 200:
```sql
DELETE FROM customers
WHERE id > 200;
```
**SELECT:** Retrieves data from one or more tables.
**Example:** This query selects all records from the customers table:
```sql
SELECT * FROM customers;
```
The above statements access and modify data across tables in an SQL database.
## DDL vs. DML: Key Differences [#ddl-vs-dml-key-differences]
Knowing the difference between DDL and DML allows you to write [well-structured queries](https://www.firebolt.io/blog/why-even-simple-queries-can-be-slow-but-not-with-firebolt), build more scalable schemas, and troubleshoot issues faster. Below is a side-by-side comparison of the two:
| Aspect | DDL | DML |
| -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------- |
| Purpose | Defines and manages the database structure, schema, and objects. | Works with the actual data stored in database tables. |
| Commands | CREATE, ALTER, DROP, TRUNCATE, COMMENT, RENAME | INSERT, UPDATE, DELETE, SELECT, MERGE |
| Impact | Modifies the database structure (e.g., adding or removing tables, columns, indexes). | Modifies or retrieves data within existing tables without changing the structure. |
| Examples | CREATE TABLE customers (id INT, name VARCHAR(50)); This command sets up the structure for storing customer data but does not insert any data into the table. | INSERT INTO customers (id, name) VALUES (1, 'William'); This code inserts a new row into the customers' table with the ID 1 and name 'William.' |
### How DDL and DML Work Together [#how-ddl-and-dml-work-together]
DDL establishes the structure and defines tables, constraints, and relationships. Once the schema is in place, DML allows operations on stored data and handles record insertion, retrieval, updates, and deletions.
Efficient database systems balance DDL for stability and DML for flexible data handling.
## Practical Examples of DDL and DML Usage [#practical-examples-of-ddl-and-dml-usage]
To illustrate the real-world usage of DDL vs DML, consider a simple database example that stores information about devices. The process involves two steps:
1. Defining the table structure using DDL
2. Adding and modifying data using DML
Here's how it works:
**Step 1: Creating a Table with DDL**
The following DDL statement defines the table structure in SQL:
```sql
CREATE TABLE devices (
device_id INT PRIMARY KEY,
device_name VARCHAR(50),
category VARCHAR(50),
price DECIMAL(10,2)
);
```
* **device\_id INT PRIMARY KEY**: Ensures unique device identification.
**VARCHAR(50)**: Stores variable-length text for names and categories.
**DECIMAL(10,2)**: Stores numeric values with two decimal places.
**Step 2: Adding & Modifying Data with DML**
Once the structure is defined, DML operations handle the data itself:
**Inserting Data (INSERT)**: Adds a new device record:
```sql
INSERT INTO devices
VALUES (1, 'Alpha,' 'Smartphone,' 299.99);
```
*Updating Data (UPDATE)*: Adjust the price of an existing device:
```sql
UPDATE devices
SET price= 279.99
WHERE device_id=1;
```
**Retrieving Data (SELECT)**: Retrieves all records:
```sql
SELECT * FROM devices;
```
### Common Usage Scenarios [#common-usage-scenarios]
DDL is used during database setup and maintenance, helping DBAs define tables, indexes, and schemas that keep data organized and reliable. In application development, software engineers use DDL to update and refine database models. DML supports daily operations like transactions, analytics, and data pipelines.
Data engineers rely on DML for building ETL and EL workflows, analysts use it to generate reports with SELECT queries, and developers integrate CRUD operations to manage data within applications.
## Importance of DDL and DML in SQL [#importance-of-ddl-and-dml-in-sql]
Without DDL and DML, databases would be unmanageable data dumps, making it impossible to organize and retrieve data. Let's explore why mastering these components is vital for efficient database management and optimized data operations:
* **Ensures Data Integrity:** DDL defines an optimized schema that models business concepts and enforces data accuracy through entities like primary keys, foreign keys, and constraints. DML supports standardized CRUD processes, ensuring accuracy as data is updated or modified.
* **Improves Security:** DDL also provides object-level security, granting and revoking user permissions on specific database objects. DML verifies access control lists during executions, ensuring users can only query and manipulate authorized data.
* **Simplifies Usage:** DDL and SML use simple, clear keywords and SQL syntax, making them easier to work with compared to proprietary coding languages. This streamlines development and analytics tasks for technical SQL users across organizations.
* **Enhances Flexibility:** Database schemas and data can be adjusted with incremental DDL and DML changes, avoiding the need for full recreations. As business needs evolve, the database model and contents can be modified iteratively.
* **Provides Insight into Database Usage and Performance:** DML supports query analytics, allowing business users to analyze database usage patterns and guide efficiency improvements. DBAs rely on DML procedures to monitor output and identify optimization opportunities.
## FAQs [#faqs]
Here are some questions to help you better understand DDL and DML. They clarify some additional concepts, commands, and how they are used in database management:
**Can DML Commands Modify Database Structures?**
No, DML commands cannot alter the structure or schema of a database. DML (Data Manipulation Language) focuses solely on managing data within existing tables or collections. Without DML, databases would be static, with no way to add, change, or remove records.
Commands like `INSERT`, `UPDATE`, and `DELETE` allow users to manipulate data but cannot modify the database structure itself. For structural changes, such as adding a new column, you need DDL (Data Definition Language) commands like `ALTER TABLE`.
For example, you would need to use the ALTER TABLE DDL statement to add a new column to an existing table. A DML command like UPDATE or INSERT will not change the table design.
**Are SELECT Queries Part of DML or DQL?**
SELECT queries belong to the Data Query Language (DQL), not DML. The `SELECT` command only reads data without altering it, while DML commands like `INSERT`, `UPDATE`, and `DELETE` change data.
While `SELECT` is categorized as DQL, most databases combine DQL and DML in the same structure. Complex `SELECT` queries often include conditional expressions found in DML.
**Is it Necessary to Know DDL and DML for Advanced SQL?**
Yes, understanding both DDL and DML is necessary for complex queries, [performance tuning](https://www.firebolt.io/blog/future-of-performance-is-not-about-performance), and database administration, as DDL and DML give full control over databases, making them vital for professionals like database developers, architects, and administrators. Junior analysts should also learn these languages to write better queries and advance toward engineering or architect roles.
# DWaaS - Definition (/glossary-items/dwaas)
Data Warehouse as a Service (DWaaS) is an outsourcing model where the service provider takes care of, configures, and manages the hardware and software resources required, providing the full-featured capabilities a company needs without any or minimal administrative overheads.
## What is DWaaS? [#what-is-dwaas]
DWaaS has become very popular recently due to the data-centric approaches most companies have picked up. It is not difficult to understand why data has become critical for companies to operate. It does it all—providing actionable analysis, giving fresh insights, and fueling business processes that are transformed digitally.
Unfortunately, on-premise [data warehouses](https://www.firebolt.io/blog/cloud-data-warehouse) can be quite large and extremely costly to build and maintain. As a result, not every organization can afford to own one. This is where DWaaS comes to address the challenge.
The core components are similar to on-premise warehouses, just that DWaaS is delivered as a cloud solution over the internet or as a solution over some private network. All the customer needs to do is provide data and pay for the managed service.
## Components related to DWaaS [#components-related-to-dwaas]
1. **Source System Integration**
A DWaaS involves a collection of connections and data feeds with source systems that allow a corporation to pull (or push) data into the warehouse from multiple end-users, apps, or any other media through which they generate their data. This also consists of a load manager that makes up the "front-component", performing all operations associated with loading all the received data into the warehouse to prepare it for the [ETL](https://www.firebolt.io/glossary-items/etl-and-elt) processes.
2. **Actual Warehouse Database/s**
After the data has been loaded to the warehouse, it needs to undergo processing to fit into a certain standardized form. This is because various applications or sources of data might possibly send varied forms or types of data. For example, a certain field might be numeric from one data source and a boolean from another.
Thus, extract-transform-load (ETL) is performed on the data before it is stored in the database—the gold mine of the entire philosophy being DWaaS. This component implements actions such as data analysis to verify consistency, index and view building, denormalization and aggregate generation, source data transformation and merging, and archive and backup creation.
The following image gives a clearer idea of the entire process. This database might or might not be the most simple-looking structure one can think of. Most companies use multiple databases to store the different types of data they come across: raw, freshly processed, or highly refined data.
3. **Administer Technical Configurations**
Once we have ensured consistency in the data, it requires technical configurations such as summarization and consolidation so it can be easily used and managed. Additionally, a set of tools, including metadata (both technical and business) generators and query managers that make up the "backend" component, are used for efficient scheduling of requests for data to deliver almost delay-less services.
4. **Analytics and services**
One of the most important parts of a DWaaS is the analytical capabilities it provides:
* business intelligence
* application development (interface supported)
* data mining
* statistical reporting
* query system
* ingestion
* modern technology like machine learning and AI
* other data integration tools
### Benefits of DWaaS [#benefits-of-dwaas]
There are many advantages to using data warehouses and DWaaS in particular.
* **Consolidation**: DWaaS consolidates data from multiple sources, acting as a single point of access for all data. It eliminates the need for users to connect to dozens, if not hundreds, of different systems, like marketing, sales, financing, and so on. The data warehouse delivers consistent data on a variety of cross-functional tasks. It also allows for ad-hoc reporting and querying. Additionally, the data warehouse assists in the integration of multiple data sources to alleviate stress on the production system.
* **High quality**: Data quality, accuracy, and consistency are all critical factors. A data warehouse applies a set of semantics to data, such as naming conventions, codes for different product types, languages, currencies, and so on.
* **Reduced supervision**: Moreover, a data warehouse aids in reducing user supervision. Your data warehouse makes it a point to show you discrepancies and rectify them before you enter data. This is especially useful for those who are careless or hurried when gathering data.
* **Analysis and reporting**: Utilizing a data warehouse can help speed up the analysis and reporting process. It is easier for the user to utilize data for reporting and analysis after restructuring and integrating it.
* **Time-saving**: DWaaS allows users to access crucial data from a variety of sources in one place. As a result, the user saves time by not having to retrieve data from several sources. Even major decisions like scaling up happen in seconds for an infrastructure that has fluctuating demands.
* **Long-lasting**: A vast volume of historical data is stored in a data warehouse. This allows users to compare and contrast different time periods and patterns to make predictions.
### Challenges of DWaaS [#challenges-of-dwaas]
There are also some disadvantages of using data warehouses that DWaaS is trying to address as it matures.
* **Third-Party Requirements**: Using data warehouses can require the help of a third-party business intelligence team, depending on the current system. Due to the complexities of operating systems, software, and programs, it may be difficult for a business owner to figure out how to properly use their data warehouse.
* **Security**: Association with third parties can lead to vulnerabilities as their policies may have exploitable points or loopholes. Attacks could expose sensitive and valuable user data. Any minor breach could result in loss of credibility and could lead to multiple lawsuit compensations depending on what was affected by the breach.
* **Performance**: Even though DWaaS separates analytical processing from traditional transactional databases, it is ultimately an online service. The service is dependant on stable connections so user experience can be highly degraded if a stable connection is not ensured.
* **High Costs**: According to [International Data Corporation](https://www.idc.com/) research, firms that invested in data warehousing produced more revenue and saved significant margins over any business model. Despite this, one should expect to spend more than their initial investment for the same. This includes regular system maintenance and upgrades, if one wants the latest technology at their fingertips for extracting the maximum value out of their data.
These are some apprehensions business owners still have about using data warehouses served to them by someone else. However as DWaaS matures like any other technology, these are likely to subside.
# ETL and ELT - Definition (/glossary-items/etl-and-elt)
ETL (Extract, Transform, Load) and ELT (Extract, Load, Transform) describe two common approaches for moving data into a [data warehouse](https://www.firebolt.io/blog/how-ai-is-transforming-etl-in-data-warehousing). In [ETL](https://www.firebolt.io/blog/etl-best-practices), data is transformed before it's loaded. With ELT, raw data lands in the warehouse first, and transformation happens there.
The shift to ELT reflects how modern [cloud-native](https://www.firebolt.io/blog/cloud-data-warehouse-vs-traditional-data-warehousing-why-the-shift) platforms work. They're faster, cheaper to scale, and flexible enough to process raw data on the fly. For teams dealing with siloed or high-volume datasets, these methods are essential to prep data for querying, modeling, or machine learning.
## Key Difference between ETL and ELT [#key-difference-between-etl-and-elt]
ETL transforms data before it enters the data warehouse. ELT loads raw data into the warehouse first, then transforms it using in-warehouse compute.
This affects:
* **Performance:** ELT leverages the warehouse's compute layer for transformation.
* **Cost:** ELT can reduce data movement and use lower-cost storage for raw data.
* **Governance:** ETL allows stricter control before ingestion.
* **Flexibility:** ELT simplifies iteration, especially with schema-on-read.
Firebolt handles both [ETL and ELT pipelines](https://www.firebolt.io/blog/etl-vs-elt-know-the-differences), as its high ingestion speeds, decoupled compute, and vectorized execution make it suited for ELT, where transformation runs inside the platform.
### ETL: Extract, Transform, Load [#etl-extract-transform-load]
Data is pulled from source systems. Transformation occurs outside the warehouse, using tools like Spark, dbt, or Python-based scripts. The final, cleaned dataset is then loaded into the target system.
ETL is common when transformation must happen before storage due to compliance, or when the destination lacks the processing power for complex transformations.
### ELT: Extract, Load, Transform [#elt-extract-load-transform]
Data is pulled and directly loaded into the warehouse in its raw format. All transformation takes place within the warehouse using SQL or integrated compute engines.
This method keeps raw data available for reprocessing, enables full audit trails, and handles evolving schemas. It's well-suited for modern platforms that can isolate and scale compute.
### When Each Makes Sense [#when-each-makes-sense]
Use ETL for strict data validation before ingestion or when dealing with legacy infrastructure.
Use ELT when performance, schema agility, and in-database processing are more important.
### Firebolt's Fit [#firebolts-fit]
Firebolt handles raw data ingestion at speeds up to 10TB per hour. It integrates with external transformation pipelines and provides in-warehouse transformation capabilities through SQL, vectorized execution, and indexed data structures.
This gives teams the option to process data before or after it lands based on volume, latency, and architectural requirements.
## Architecture and Performance: ETL vs ELT [#architecture-and-performance-etl-vs-elt]
How you design your data pipeline, ETL or ELT, impacts speed, cost, and infrastructure complexity.
### ETL Architecture [#etl-architecture]
ETL includes three parts: a source system, an external processing engine, and a data warehouse.
* Data must be moved, transformed, and then loaded before it can be queried.
* The transformation engine must be deployed, scaled, and monitored as a separate system.
* This adds more systems to manage and increases total processing time.
### ELT Architecture [#elt-architecture]
ELT moves data directly from the source into the warehouse.
* Transformations are run inside the warehouse using SQL and internal compute.
* This setup eliminates intermediate processing systems and shortens the time between extraction and analysis.
* It requires the warehouse to handle large query volumes and fast execution.
### Performance Comparison [#performance-comparison]
* **ETL:** Reduces demand on the warehouse but takes longer to make data available.
* **ELT:** Provides faster access to data for querying and modeling, but places higher demands on warehouse performance.
### Firebolt Context [#firebolt-context]
Firebolt separates storage, compute, and metadata into independent layers and makes it [scalable](https://www.firebolt.io/blog/high-volume-ingestion-scalable-and-cost-effective-data-loading).
* Data ingestion can run in parallel with querying and transformation.
* Query speed is maintained under high user load due to isolated processing and native indexing.
* This design handles high-frequency SQL transformations without delay.
## Common Tools and Use Cases [#common-tools-and-use-cases]
Tool selection depends on how data moves through your environment. Some setups rely on external processing before storage. Others push raw data into the warehouse and transform it there.
### ETL Tools [#etl-tools]
Examples: Apache NiFi, Talend, Informatica, SSIS
Best used when:
* Regulatory rules require pre-storage transformation
* Infrastructure includes on-prem systems or hybrid setups
* Schema is fixed and data arrives in regular batches
* Reducing warehouse compute usage is a priority
### ELT Tools [#elt-tools]
Examples: [dbt](https://www.firebolt.io/blog/elt-with-firebolt-using-dbt), Fivetran, Matillion, Airbyte, Rivery
Best used when:
* The stack is fully cloud-based
* Data includes nested or semi-structured formats
* Dashboards refresh frequently, or models are updated often
* Data is staged for machine learning or exploratory queries
### Firebolt Context [#firebolt-context-1]
Firebolt connects directly with ELT systems like dbt and Fivetran.
It ingests raw data at high speeds, which suits customer-facing dashboards and high-frequency query environments.
It also works with external systems for teams that still run transformations before loading.
## ETL vs ELT: Side-by-Side Comparison [#etl-vs-elt-side-by-side-comparison]
Here's a side-by-side comparison of ETL and ELT across critical technical categories. Each row breaks down a core aspect, including processing location, latency, scalability, and typical toolsets.
| Feature | Data Lake | Data Warehouse | Lakehouse |
| ------------------ | ------------------------ | ---------------------------- | ------------------------------------- |
| Storage | Object storage | Columnar storage engines | Object storage with format layer |
| Formats | Parquet, ORC, Avro | Proprietary (e.g., Redshift) | Open (Parquet, Delta, Iceberg) |
| Schema enforcement | Optional or late-binding | Strict | Configurable and flexible |
| Workloads | ML, batch processing | BI, dashboards | BI, ML, streaming, batch |
| Transactions | No | Yes | Yes (via Delta/Iceberg/Hudi) |
| Cost | Low | High | Low for storage, variable for compute |
| Governance | Limited, manual | Built-in | Varies, improving rapidly |
| Streaming support | Requires integration | Limited | Native in some engines |
| Tooling ecosystem | Fragmented | Mature | Growing and modular |
## Common Challenges [#common-challenges]
ETL and ELT each come with their own limitations. Understanding them upfront helps prevent downstream issues.
### ETL challenges [#etl-challenges]
* **Longer pipelines:** Data moves through multiple systems, each adding latency and increasing failure points.
* **Extra systems to manage:** Running and maintaining external transformation layers requires setup, scaling, and monitoring.
* **Rigid schemas:** Upstream schema changes often require manual pipeline rewrites and testing.
* **Slow time-to-analysis:** Data isn't available for querying until all transformation steps complete.
### ELT challenges [#elt-challenges]
* **Increased compute demand:** All transformations run inside the warehouse, which raises processing requirements.
* **Unexpected costs:** Poorly optimized queries or wide scans can drive up warehouse usage charges.
* **Governance pressure:** Storing raw data increases exposure to access control, data lineage, and retention risks.
* **Messy query logic:** Analysts often deal with overlapping raw and cleaned tables, increasing complexity in joins and filters.
### Firebolt context [#firebolt-context-2]
Firebolt isolates query execution, ingestion, and transformation, so one task doesn't delay another.
* It provides native controls to track and limit query resource usage.
* Indexing and aggressive filtering reduce the cost of working with large raw datasets.
* Access controls are granular and built for managing structured, semi-structured, or raw data at scale.
## Best Practices for ETL and ELT Implementation [#best-practices-for-etl-and-elt-implementation]
These practices help teams reduce errors, control cost, and stay flexible as pipelines evolve.
1. **Define data ownership early**
Assign clear responsibility for extraction, transformation logic, and data quality. This avoids overlap and gaps as systems grow.
2. **Optimize for cost and latency**
In ETL, decouple storage from compute so transformation systems can scale independently.
In ELT, track query usage by transformation stage to prevent excessive warehouse spending.
3. **Use modular SQL for ELT**
Break down logic into stages like staging, transformation, and marts. Modular SQL simplifies debugging and helps teams reuse code across projects.
4. **Build for change**
Use version-controlled tools like dbt. Choose flexible data formats and schema structures that support updates without downtime.
5. **Monitor and audit pipelines**
Track lineage, schema changes, and failures. In ELT environments, monitor query latency and compute load to avoid saturation.
## Firebolt and Modern Data Pipelines [#firebolt-and-modern-data-pipelines]
Firebolt is built to handle data at speed and scale, whether transformation happens inside or outside the warehouse.
1. **Handles ETL and ELT setups**
External tools can load into Firebolt after transformation. Firebolt can also ingest raw data and process it internally.
2. **Designed for ELT performance**
Ingests up to 10TB per hour and executes queries with sub-second latency. Compute runs independently from storage and metadata layers.
3. **Simplified design**
Built-in SQL engine with vectorized processing and indexing removes the need for OLAP cubes or manual aggregation.
4. **Developer-focused tools**
Works with Postgres-compatible SQL, integrates with dbt, and includes SDKs for custom applications. Easy to slot into existing systems.
## ETL vs ELT for Different Workloads [#etl-vs-elt-for-different-workloads]
| Workload | ETL Fit | ELT Fit |
| --------------------- | ---------------------------------------- | --------------------------------------------- |
| Business Intelligence | Stable, modeled data for dashboards | Dynamic reporting layers, agile model updates |
| Operational Analytics | Rare | Common (low-latency ops insights) |
| Machine Learning | Structured data prep with offline models | Feature pipelines using raw + enriched data |
| Customer-Facing Apps | Uncommon due to delay | Ideal for fast insights from fresh raw data |
| Data Archiving | Used for retention and schema curation | Used for reprocessing or backfill workflows |
### Can I use both ETL and ELT together? [#can-i-use-both-etl-and-elt-together]
Yes. Many teams extract and clean core data using ETL, then apply additional business logic with ELT in the warehouse.
### Does Firebolt require ELT? [#does-firebolt-require-elt]
No. Firebolt works with both approaches. It supports high-throughput ingestion and performs just as well with pre-modeled or raw data.
### Is ELT more expensive than ETL? [#is-elt-more-expensive-than-etl]
Not always. ELT can reduce infrastructure costs but may increase compute usage in your warehouse. Firebolt's workload isolation and cost control tools help manage this.
If you're running ETL or ELT at any serious scale [**Book a demo**](https://www.firebolt.io/book-a-demo) to see how Firebolt handles raw data, SQL transformations, and high-concurrency queries without slowing down.
# Materialized Views (/glossary-items/materialized-views)
A Materialized View (MV) is a database object that contains the result of a precomputed query. This object exists for a precise reason: to boost performance and efficiency and reduce network load. As database queries are usually the bottleneck in modern database-intensive applications, MVs are crucial to improving application performance.
Furthermore, as MVs contribute to better performance, they bring down costs, especially that of a [data warehouse](https://www.firebolt.io/glossary-items/data-warehouse). In those environments, one often finds its very frequent to find complex aggregation and intensive joins on huge tables containing billions of rows.
This is useful when, for example, business intelligence (BI) users have dashboards that need analytical data extraction ([OLAP](https://www.firebolt.io/glossary-items/olap)), one needs to execute internet of things (IoT) processing, or one is running an extract-transform-load process ([ETL](https://www.firebolt.io/glossary-items/etl-and-elt)). Another classical use case, combined with pre-joining collections and data aggregation, is using this kind of database object for data-filtering.
There is an important distinction between materialized views and views. The primary difference pertains to the physical allocation of data. An MV is actually a DB object and, like all tables, contains data. Conversely, a view is a virtual-only table, which does not contain anything. In fact, with views, data is created only when it gets questioned.
Generally speaking, we should create an MV when:
* **Query results are needed often:** Creating an MV means increasing storage cost together with further maintenance. This is not a worthwhile investment if the MV is not queried frequently. Better performance for queries that are executed usually impacts a lot on user experience. It is a precise cost/benefit balancing game to be played with full awareness.
* **Query results could be outdated:** We shouldn't rely on MVs for queries that need real-time data. If stringent consistency of the transaction is at play, any delay on the data update should be considered a big point of concern. Even if there are several mechanisms to automatically refresh MVs (especially in smart data warehouses), adding a level of caching on top of the only source of truth will always bring some data masking.
* **The query is expensive and uses lots of resources:** While it's true that it's possible to rely on MVs to execute simple data filtering or little aggregations, the real point of these objects' existence is to decrease costs. There is no point in increasing database storage occupation for operations that can be done in other ways.
In all the other cases, we should rely on **virtual views** or **database caching**. In some sense, database caching is similar to MVs. Query results are precomputed in both cases. But, as you can imagine, caching is a static process: data cannot be cached if input filters change dynamically, or you need to have lots of cached data.
In all modern data warehouses, it's possible to specify the refreshing mechanism when creating an MV. Usually, it can be automatic or manual. The automatic mode will check if any difference exists among the source tables of the warehouse and update the MV accordingly, reasoning on the delta of data. In some cases, a best-effort approach could be implemented, refreshing views only when really needed.
## Advantages [#advantages]
The main advantages of the MV are performance and efficiency. When using data warehouses, those objects are usually used for a repeatable and predictable workload. Thanks to them, it's possible to compute once and query many times. Let's enumerate the principal advantages of MVs:
1. **Performance**: If a standard query consumes a lot of resources (processing time and storage space), it's possible to cut costs and increase performance with MVs. It's also possible to increase performance even more by creating indexes on the MVs' relevant columns. It also reduces the execution time for simple queries which is great for situations where query computation cost is high but the resulting data set is small.
2. **Data Compactness**: Usually, relying on MVs simplifies query logic. Complex join and aggregation could become a single select with some static filters. This action will greatly reduce business logic errors and increase data compactness.
3. **Access**: MVs can be created in the same database where the base tables exist or in a different database as well. This is especially useful if the query is on a remote database. By replicating data, it's possible to provide local access to the target data.
4. **Caching**: It's possible to use caching on an MV to further optimize performance and costs. This can be especially useful when the query on the MV is executed with static input parameters.
## Challenges [#challenges]
Everything comes with a cost. Here, those are the main challenges that MVs bring along with their advantages:
1. **Restricted Syntax**: MVs use a restricted [SQL](https://www.firebolt.io/glossary-items/sql) syntax and a limited set of aggregation functions. Not every query is immediately convertible into a materialized view. All modern data warehouse and database vendors are trying to expand the syntax as much as possible to cover all use cases.
2. **Costs**: Data in MVs is not updated in real-time by default. Modern data warehouses can always show the latest state of the tables using a best-effort approach. This optimization, unfortunately, can increase costs and latency.
3. **Maintenance**: DML data editing on tables on which the materialized views exist has to be replicated onto the materialized view itself. This will always increase maintenance costs even if, with data warehouses, this task is automatically managed by the DBMS itself. It is usually part of the creation cost of MVs.
4. **Obsolete Data**: MVs can be refreshed manually or automatically in a scheduled fashion. In the latter case, it's difficult to choose an appropriate refresh schedule. If you refresh data too frequently, it may greatly increase the cost of the executions. Refreshing data with long delays can bring obsolete data to users or, if using a best-effort approach, long latencies, and increased costs.
Generally speaking, having two sources of information with the same result is against the standard third normal form of databases. This can also impact applications that use ORM. It will cause different objects to have the same semantic value. The same information could be retrieved from two different sources.
# OLAP - Definition (/glossary-items/olap)
Online Analytical Processing (OLAP) is the foundation for business intelligence tools – it is software for multidimensional analysis database queries to permit high speed processing on large volumes of data. They work with cloud data warehouses, data marts, and other centralized data stores and can be used for report views, predictive analysis, and other analytical calculations.
## What is OLAP? [#what-is-olap]
The multidimensional nature of OLAP is due to the fact that most businesses have multiple sections into which their data is broken down for different purposes. OLAP extracts data from multiple relational data sets to reogranize and recategorize it into different formats for better processing and analysis. This technology is used by solutions like Oracle's Essbase, Microsoft SWL Server Analysis Services, and Cognos.
OLAPs work based on **cubes** which are array-based multidimensional databases. They are more efficient than traditional relational databases which makes them a good option for cloud computing, too. Cubes extend the traditional relational database's single table concept and add additional dimensions to it to expand the hierarchy of data. The top layer of the cube would have one form of organization while additional layers would filter and sort in different ways. What makes OLAP cubes different is that they can potentially and theoretically contain infinite layers so you could have multiple permutations of tables and organizations of data. Simply: cubes allow you to slice and acquire the data you *really* need to understand your business analytics.
In addition to multidimensional OLAP cubes, there are also relational OLAP (ROLAP) and hybrid OLAP (HOLAP). ROLAP is multidimensional data analysis that takes place on data in unorganized **relational** tables – this is best when there's more emphases on working with massive amounts of data than performance or efficiency. HOLAP functions through the combination of multidimensional OLAP and ROLAP to divide processing load. It is ideal for data processing efficiency and high scalability for larger companies but won't be as fast as ROLAP. It's more expensive as well because of its complex architecture and frequent up-keep.
Nowadays, OLAP is used commonly with cloud computing as it's less expensive and easier to set up with so much cloud-based data. This is thanks to MPP – massively parallel processing – wherein OLAP can tap into high-level analytics on cloud-based data warehouses. For companies, this is extremely beneficial as it has the potential to maximise the usability of corporate data, leading to more thorough analysis on business insights. This can be especially useful for companies that would otherwise have physical data silos in different regions: the cloud optimizes access while OLAP makes it all more efficient.
All in all, OLAP is great for situations where you need to conduct complex analytics on large datasets to generate simply and readable reports quickly and consistently. OLAP systems, especially on the cloud (rather than legacy systems) are optimized for read-heavy situations and datasets and is ideal for trend-spotting and data exploration.
## OLAP Advantages [#olap-advantages]
Many organizations are switching to building Business Intelligence (BI) solutions using OLAP technology due to factors like speed, efficiency, and structural integrity.
* **Speed**: OLAP solutions are fast and help users perform data analysis and reporting on their own. They are best for data mining or discovering data items' relationships. With these tools, you do not have to perform calculations manually and compose complex reports anymore. OLAP tools help run queries in seconds.
* **Centralization**: They store all the transactional data, customer information, supplier information, data related to the company employees and their performance, etc., in a centralized location.
* **Data Organization**: OLAP tools follow a multidimensional approach where all the data is organized into various dimensions (sharing common characteristics) and later used for analysis. These dimensions include simple business categories that are easy to understand even for non-technical users.
* **Efficiency:** OLAP tools avoid the manipulation of database table fields. It is one of the best computing methods that help users extract and analyze information from a different point of view.
* **Ease of Use:** One can collect information through a [data warehouse](https://www.firebolt.io/blog/cloud-data-warehouse) and perform business analysis without any technical background. OLAP systems provide complete documentation, tutorials, and prompt technical support for users.
* **Response:** One can also practice "what-if" scenarios with the help of OLAP systems. If the OLAP cubes support write-back functions, you can replace values and look for other outcomes, foresee losses, and make strategies to prevent them. As a result, the tool helps organizations make better decisions to improve profit, sales, brand image, marketing strategies, and more.
* **All-Encompassing:** Furthermore, the system allows users to create reports like sales forecasting, management reporting, financial reporting, trends analysis, and more on their own. As a result, it helps reduce the demand for IT resources.
The OLAP system is highly beneficial as it helps aggregate and sum up data in cubes that are further rolled up, sliced, and diced to form the best solution. Once the cubes are made, teams can use existing business intelligence tools to instantly connect with the OLAP model and draw interactive real-time insights from their cloud data. The high-speed data processing features present in these tools enables users to run queries and prepare reports in just a few minutes.
## OLAP Challenges [#olap-challenges]
OLAP systems are a great investment for big companies as they can bring in more profit in the future. But, implementing OLAP systems in most cases can present some challenges.
* **Set-up**: Organizations find difficulties when structuring a database and creating a decision support architecture on their own. Thus, various organizations opt out and make no effort. The database must be structured properly to perform analysis on large amounts of datasets.
* **Structuring**: To use an OLAP system, one needs to define its structure in advance, i.e., pre-calculate each column and data type before creating a table. Similarly, having a good OLAP engine is essential. Without it, one cannot convert data into a pattern, and they will find it hard to operate directly. Overall, it may result in causing difficulties for quick results.
* **High-Level**: You will require IT pros at some stage, even if the OLAP system is meant for users to operate and run queries. Traditional OLAP tools follow a complex modeling procedure. Thus, to write codes/scripts/SQL, use [ETL](https://www.firebolt.io/glossary-items/etl-and-elt) tools for integration, or use the map dimensions of a model, you will require an IT expert and human resources. A non-technical user cannot handle this part by themself.
* **Power Requirement**: Some systems may lack computational power resulting in less flexibility of the OLAP tool. In the majority of cases, analyzers have to depend on a third party to perform calculations as they are restricted to a small area.
No doubt, OLAP tools are highly advantageous, but various challenges work as a setback. All the above-listed challenges without fixation can result in inappropriate or incomplete information to the users. This may further cause problems in decision-making.
The only method to combat these challenges is to strongly integrate the analysis tools and databases. Both the primary components of the OLAP system should not impose any restrictions on one other. It is essential to ensure that the ability to manage or analyze data will not be compromised in any way while structuring the components.
# Partitioning and Sharding - Definition (/glossary-items/partitioning-and-sharding)
Partitioning is a technique used in databases to break a single table into smaller chunks or partitions. This provides a foundation for faster queries and ingestion, while enabling ease of maintenance for extremely large tables. Having the data partitioned into smaller tables can help tailor the schema to be in line with the data access patterns. Vertical and Horizontal partitioning are two approaches used by customers depending on their needs.
## Vertical partitioning [#vertical-partitioning]
In the case of vertical partitioning, a table is split into two or more tables with different sets of columns. Each of the individual tables can be customized to include the columns that support specific query access patterns.

After vertical partitioning, above table is split into 2 separate tables each with its own set of columns.

## Horizontal partitioning [#horizontal-partitioning]
On the other hand, with horizontal partitioning a table is split into multiple tables where all tables have identical schemas but hold a smaller subset of the rows of the main table.

With horizontal partitioning, data is split into different partitions based on sales order date column. This is shown below.

## Sharding [#sharding]
Sharding typically references horizontal partitioning. In the case of sharding, the partitions themselves could be spread out on multiple machines enabling scale-out capabilities. This approach results in each individual machine processing only a portion of the data resulting in faster data processing. A shard key is used to split the data into smaller partitions.
## Clustering and partition pruning [#clustering-and-partition-pruning]
Partitioned or sharded data can then be efficiently organized using clustering to ensure data is sorted according to a key. This provides best options when sequential access to sorted data is needed. Partition pruning and clustering are techniques used by analytics platforms to avoid scanning or ordering data resulting in efficient query processing. Partition manipulation can also be used to archive data or to remove older data as conducting wholesale partition operations are efficient compared to deletes and updates. These techniques are leveraged in large scale analytics platforms.
# Query Federation - Definitions (/glossary-items/query-federation)
A federated query is defined as a feature that enables executing a single query across multiple, potentially heterogeneous data sources, such as operational databases, data warehouses, and data lakes, without the need to physically consolidate the data. This concept aligns with the broader idea of **database federation**, where a **federated database** system provides a virtualized view of distributed data, allowing users to interact with it as a single entity. The term "federated query" specifically refers to the query mechanism that facilitates this integration, often using SQL as the common language.
## Understanding federated queries [#understanding-federated-queries]
An analogy often used is that of a "universal translator" for SQL. Just as a universal translator enables communication between people speaking different languages, a federated query enables data retrieval and combination from diverse sources, such as MySQL, PostgreSQL, Amazon Redshift, Snowflake, and Amazon S3, using a standardized query interface. This is particularly relevant in modern data architectures where data is rarely centralized but distributed across various platforms, each with its own schema and access methods.
Research suggests that federated queries emerged from the need to handle distributed data systems, with early concepts dating back over 20 years, as noted in discussions around query federation for heterogeneous databases. The technology has gained traction with the rise of big data and multi-cloud environments, where organizations need to analyze data across disparate systems without incurring the overhead of data movement.
The primary purpose of federated queries is to enable seamless data integration and analysis without physically moving or replicating data across systems. This is achieved through **database federation**, which maintains the autonomy of each data source while providing a centralized query interface. The evidence leans toward this being a critical feature for distributed data environments, where data resides in operational databases (e.g., MySQL, PostgreSQL), data warehouses (e.g., Amazon Redshift, Snowflake), and data lakes (e.g., Amazon S3).
By allowing queries to span these diverse sources, federated queries help organizations gain insights from their entire data ecosystem without the need for complex extract, transform, load (ETL) pipelines. This is particularly valuable in scenarios where data latency is a concern, such as real-time analytics, or where data governance policies prohibit data replication. For instance, [Querying data with federated queries in Amazon Redshift](https://docs.aws.amazon.com/redshift/latest/dg/federated-overview.html) highlights that federated queries can incorporate live data from external databases into business intelligence (BI) and reporting applications, reducing the need for complex data ingestion processes.
Technical advantages include:
* **Reduced data movement:** Queries are executed at the source, minimizing network traffic and storage costs.
* **Performance optimization:** Many systems, such as Amazon Redshift, distribute computation to remote databases and leverage parallel processing.
* **Flexibility:** Supports heterogeneous data sources, accommodating different SQL dialects and data formats, as seen in Google BigQuery's support for AlloyDB, Spanner, and Cloud SQL ([Introduction to federated queries | BigQuery | Google Cloud](https://cloud.google.com/bigquery/docs/federated-queries-intro)).
However, there are limitations, such as read-only access in some implementations (e.g., BigQuery federated queries do not support DML/DDL), and performance may vary based on network latency and the proximity of data sources to the query engine.
### Practical examples and use cases [#practical-examples-and-use-cases]
Federated queries are particularly useful in real-world scenarios where data is distributed across multiple platforms. Below are detailed examples.
**Cross-platform querying for customer behavior analysis:** Imagine a scenario where a company needs to analyze customer behavior by combining data from its CRM system (stored in a PostgreSQL database), sales transactions (in an Amazon Redshift data warehouse), and user interaction logs (in an Amazon S3 data lake). With federated queries, a data engineer can write a single SQL statement that joins these datasets, such as:
```sql
SELECT c.customer_id, s.transaction_amount, l.log_date FROM postgresql.crm.customers c JOIN redshift.sales.transactions s ON c.customer_id = s.customer_id JOIN s3.logs.user_interactions l ON c.customer_id = l.customer_id;
```
This is enabled by *federated database* systems like Amazon Redshift, which support querying across RDS for PostgreSQL, Redshift, and S3.
**Real-time analytics with operational and analytical data:** In a retail scenario, a data architect might need up-to-the-minute sales data for reporting. Federated queries allow joining real-time transaction data from a MySQL database with historical sales data in Snowflake. For example:
```sql
SELECT m.transaction_id, m.amount, s.historical_trend FROM mysql.operational.transactions m JOIN snowflake.analytics.sales_trends s ON m.product_id = s.product_id;
```
This use case is supported by platforms like Azure Databricks, which enable querying across local and foreign catalogs, as noted in [Federated queries (Lakehouse Federation) - Azure Databricks](https://learn.microsoft.com/en-us/azure/databricks/sql/language-manual/sql-ref-federated-queries).
**Hybrid cloud and multi-cloud environments:** Organizations often operate in hybrid or multi-cloud setups, using services like AWS, Google Cloud, and Azure. Federated queries enable querying data residing in different clouds or on-premises databases from a single interface. For instance, [Google BigQuery's federated queries](https://cloud.google.com/bigquery/docs/cloud-sql-federated-queries) allow querying Cloud SQL (MySQL, PostgreSQL) alongside BigQuery storage. This simplifies data management and analysis, reducing the complexity of managing multiple data silos.
### Technical details and considerations [#technical-details-and-considerations]
For a deeper understanding, consider the following technical aspects, which are crucial for implementation:
* **Supported data sources:** Different platforms support varying external databases. For example, Amazon Redshift supports RDS for PostgreSQL, Aurora PostgreSQL, RDS for MySQL, and Aurora MySQL ([Querying data with federated queries in Amazon Redshift](https://docs.aws.amazon.com/redshift/latest/dg/federated-overview.html)), while BigQuery supports AlloyDB, Spanner, and Cloud SQL.
* **Query execution and performance:** Federated queries often optimize performance by delegating operations like filtering and column pruning to the external source, known as SQL pushdowns. For instance, BigQuery supports filter pushdowns for Cloud SQL MySQL.
* **Limitations and challenges:** Federated queries are typically read-only, with no support for data manipulation language (DML) or data definition language (DDL) in some systems (e.g., BigQuery). Performance may be slower than querying local storage due to network latency, and unsupported data types can cause query failures.
* **Pricing and quotas:** On-demand pricing is common, with charges based on bytes returned (e.g., BigQuery's on-demand pricing at [BigQuery Pricing](https://cloud.google.com/bigquery/pricing)). Quotas, such as a 1 TB cross-region query limit per project per day for BigQuery, can also affect scalability.
### A real-world example: querying multiple data sources with Firebolt [#a-real-world-example-querying-multiple-data-sources-with-firebolt]
Let's see **federated queries** in action with a practical example. Suppose you're a data engineer at an e-commerce company. You need to analyze customer behavior by combining:
* **Transactional sales data** from a PostgreSQL database.
* **Customer interaction logs** (e.g., website clicks) stored in a cloud storage bucket (e.g., Parquet files).
* **Historical demographic data** from an external analytical database.
With Firebolt's **federated query** capabilities, you can write a single SQL statement to join these sources without moving any data. Here's how it might look:
```sql
SELECT
t.transaction_id,
t.transaction_amount,
l.interaction_timestamp,
c.customer_name,
c.customer_age
FROM
external_postgres.sales_transactions t
INNER JOIN external_storage.customer_logs l
ON t.customer_id = l.customer_id
INNER JOIN external_analytics.customer_demographics c
ON t.customer_id = c.customer_id
WHERE
t.transaction_date >= '2025-01-01'
AND l.interaction_type = 'CLICK'
AND c.customer_age BETWEEN 25 AND 45;
```
**Breaking down the query**
* **PostgreSQL (`sales_transactions`):** An external table connected to a PostgreSQL database, containing sales data like `transaction_id` and `transaction_amount`.
* **Cloud storage (`customer_logs`):** An external table accessing Parquet files in a cloud storage bucket, with columns like `interaction_timestamp` and `interaction_type` (e.g., 'CLICK').
* **Analytical database (`customer_demographics`):** An external table from another analytical platform, providing `customer_name` and `customer_age`.
This query joins the datasets on `customer_id`, filters for transactions since January 1, 2025, click interactions, and customers aged 25–45. The result? A unified view of customer behavior, powered by Firebolt's blazing-fast **federated database** engine.
## Why federated queries are critical in modern data architectures [#why-federated-queries-are-critical-in-modern-data-architectures]
As organizations adopt multi-cloud, hybrid, and decentralized data strategies, federated queries have become a foundational capability for modern analytics. Traditional approaches that rely on centralized data warehouses and heavy ETL (Extract, Transform, Load) pipelines are increasingly ill-suited to today's data environments. That's where database federation and federated databases shine — they provide a flexible, scalable, and real-time alternative.
### Speed to insight without data duplication [#speed-to-insight-without-data-duplication]
In traditional data architectures, moving data between systems incurs high costs. You pay for every operation: extracting from the source, transforming data for compatibility, and storing redundant copies in your data warehouse. These costs scale quickly with data volume and pipeline complexity.
Federated queries reduce — and often eliminate — these costs. By enabling direct access to remote data sources, federated queries allow analytical workflows to avoid unnecessary storage and transformation overhead. This is especially impactful in multi-cloud or hybrid cloud deployments, where cross-region data transfers can become a major line item.
In a federated database environment, compute resources are used more efficiently because operations like filtering and aggregation can be pushed down to the source. Instead of pulling gigabytes of raw data into the warehouse, you retrieve only what's needed for the specific query. The result is lower compute usage, reduced egress charges, and a leaner, more cost-effective architecture.
This makes database federation not only a performance enhancer, but also a smart financial move — especially for organizations looking to scale analytics without scaling their infrastructure budget.
### Cost savings from reduced data movement [#cost-savings-from-reduced-data-movement]
In traditional data architectures, moving data between systems incurs high costs. You pay for every operation: extracting from the source, transforming data for compatibility, and storing redundant copies in your data warehouse. These costs scale quickly with data volume and pipeline complexity.
Federated queries reduce — and often eliminate — these costs. By enabling direct access to remote data sources, federated queries allow analytical workflows to avoid unnecessary storage and transformation overhead. This is especially impactful in multi-cloud or hybrid cloud deployments, where cross-region data transfers can become a major line item.
In a federated database environment, compute resources are used more efficiently because operations like filtering and aggregation can be pushed down to the source. Instead of pulling gigabytes of raw data into the warehouse, you retrieve only what's needed for the specific query. The result is lower compute usage, reduced egress charges, and a leaner, more cost-effective architecture.
This makes database federation not only a performance enhancer, but also a smart financial move — especially for organizations looking to scale analytics without scaling their infrastructure budget.
### Real-time analytics from operational systems [#real-time-analytics-from-operational-systems]
Businesses increasingly need real-time access to data — whether it's for fraud detection, personalized recommendations, or operational dashboards. Traditional ETL pipelines can't keep up with these demands because they introduce latency at every step: data must be extracted, cleaned, and loaded before it's ready for analysis.
Federated queries solve this by connecting directly to operational databases and other live data systems. For example, you can join real-time purchase data from a MySQL database with user demographic data from a cloud data warehouse — all in a single SQL query. There's no delay, no staging environment, and no waiting for a batch job to complete.
This is made possible by the architectural principles of database federation, where a federated database system acts as a virtualization layer across diverse data platforms. This design enables real-time access to disparate systems while preserving data integrity and control.
By querying data in place, organizations can build real-time analytics that respond to events as they happen — not hours or days later. This is critical for use cases like operational intelligence, time-sensitive decision-making, and AI-powered workflows that require access to the most current information possible.
### Common use cases [#common-use-cases]
Federated queries aren't just a theoretical advantage — they solve real-world problems across industries and roles. Here are a few of the most impactful use cases, each made easier, faster, and more scalable through the power of federated database systems.
**Customer 360 views across marketing, sales, and support systems** — Creating a unified customer profile is a top priority for marketing, sales, and support teams. However, customer data is often scattered across CRM platforms, transactional systems, support databases, and clickstream logs. Traditional ETL pipelines struggle to keep these systems in sync — especially when updates happen in real time.
With federated queries, data engineers can write SQL statements that join customer records across disparate systems without ingesting the data into a single store. For example, a federated query could join customer metadata from Salesforce (via PostgreSQL), purchase history from Redshift, and session logs stored as Parquet files in S3 — enabling a truly dynamic Customer 360 view. This approach leverages the flexibility of database federation to reduce duplication, eliminate latency, and simplify compliance by keeping sensitive data in its original location.
**IoT & log analysis directly from cloud storage or Kafka** — IoT devices and applications generate massive volumes of telemetry data that is typically stored in object storage systems like Amazon S3 or Azure Data Lake. Analyzing this data in combination with metadata or reference tables (e.g., device configurations or user profiles) requires joining structured and semi-structured data.
With federated queries, developers can seamlessly query log files or sensor data directly from cloud storage, join them with relational databases or warehouses, and extract insights without triggering massive ingestion jobs. This makes federated querying an ideal solution for real-time monitoring, anomaly detection, predictive maintenance, and AI model training pipelines that rely on raw sensor data.
**Real-time dashboards blending operational and historical data** — Business intelligence tools often rely on periodic data refreshes, meaning dashboards can lag behind real-time data by hours or even days. This is particularly problematic for decision-makers who need up-to-date information to drive strategy.
By powering dashboards with federated queries, developers and analysts can build live reports that pull data directly from source systems. For example, sales data from a PostgreSQL operational DB, marketing metrics from Snowflake, and customer support logs from S3 can all be queried in real time — and rendered in tools like Metabase or Tableau without delay. Federated database systems ensure that each component of the dashboard reflects the latest data, enabling more accurate decision-making across teams.
**Joining cloud data warehouse with lakehouse storage** — Many organizations maintain a lakehouse architecture, where structured warehouse data is complemented by unstructured or semi-structured files in object storage (e.g., Delta Lake or Apache Iceberg on S3). Traditionally, blending these datasets requires ingestion and transformation pipelines that can be brittle, slow, and expensive.
Federated queries provide a seamless alternative. With tools like Firebolt, developers can write SQL to join warehouse tables with external tables stored in Parquet, Delta, or Iceberg format — directly and efficiently. There's no need to load lake data into the warehouse or transform it in advance. This unlocks new possibilities for machine learning, behavioral analytics, and compliance workflows, all powered by live, federated access to the full spectrum of enterprise data.
## How federated queries work [#how-federated-queries-work]
At a high level, a federated query is a SQL statement that accesses and combines data from multiple systems as if they were part of a single database. It enables teams to join data from operational databases, cloud data warehouses, data lakes, and even SaaS APIs—without moving the data.
To make this possible, modern federated database systems rely on several architectural components that abstract away the complexity of distributed data environments. Let's explore how these systems work under the hood.
### SQL layer abstraction [#sql-layer-abstraction]
At a high level, a federated query is a SQL statement that accesses and combines data from multiple systems as if they were part of a single database. It enables teams to join data from operational databases, cloud data warehouses, data lakes, and even SaaS APIs—without moving the data.
To make this possible, modern federated database systems rely on several architectural components that abstract away the complexity of distributed data environments. Think of the SQL layer as the brain of the federated database — making intelligent decisions about how and where each part of the query should be executed.
### Source connectors and data virtualization [#source-connectors-and-data-virtualization]
At the core of any federated query engine is its ability to talk to external data sources. This is made possible through source connectors — specialized adapters that allow the federated system to connect with relational databases, object stores, and analytical engines.
Commonly supported connectors include:
* Relational databases (e.g., PostgreSQL, MySQL, SQL Server)
* Cloud data warehouses (e.g., Redshift, Snowflake, BigQuery)
* Data lakes (e.g., S3, Delta Lake, Iceberg, Parquet files)
* Streaming platforms (e.g., Kafka, Kinesis)
* SaaS APIs (e.g., Salesforce, HubSpot)
These connectors enable data virtualization, a key principle of database federation. Rather than ingesting or copying data, the federated system creates external tables that point to the remote data sources. These virtual tables behave like native tables in SQL, but they reference data that lives elsewhere.
> Data virtualization allows teams to join, filter, and aggregate remote data on the fly, without needing to move or transform it ahead of time.
This is especially valuable in enterprise environments where data is siloed across regions, clouds, or business units, and where replication is either too expensive or restricted by governance policies.
### Distributed query execution engine [#distributed-query-execution-engine]
Once the SQL abstraction layer has parsed the query and identified the relevant data sources, the federated query engine kicks in to optimize and execute the plan. Unlike traditional query engines that assume all data is local, federated query engines must operate across distributed systems — often with varying latency, data formats, and compute capabilities.
The process typically follows these steps:
1. Query decomposition: The SQL statement is broken down into logical subqueries targeted at each data source.
2. Pushdown logic: Where possible, filters, joins, and aggregations are pushed down to the remote system to minimize data transfer.
3. Parallel execution: Subqueries run in parallel across sources to reduce latency.
4. Intermediate result collection: Partial results are returned to the federated engine.
5. Final aggregation and assembly: The engine performs any final joins or aggregations, then returns the complete result to the user.
This approach allows federated queries to remain performant, even when querying petabyte-scale data across multiple systems. Modern engines like Firebolt enhance this flow by integrating smart indexing, vectorized execution, and query caching, resulting in sub-second latency even in federated scenarios. In contrast, legacy engines often suffer from poor concurrency and high overhead when stitching together results from multiple sources.
### Why does it matter? [#why-does-it-matter]
Understanding how federated queries work isn't just academic—it's essential for designing performant, scalable data architectures. By using a federated query engine with a robust SQL layer, flexible connector support, and distributed execution capabilities, developers and data engineers can unlock insights across silos—without ETL bottlenecks, data duplication, or pipeline sprawl.
As data continues to grow across environments, database federation provides a pragmatic and scalable alternative to monolithic data warehousing. Whether you're integrating a federated database into your stack or simply exploring how to reduce data movement, the benefits of this architecture are clear: speed, simplicity, and flexibility.
## Federated query vs. data federation vs. data virtualization [#federated-query-vs-data-federation-vs-data-virtualization]
In the world of distributed data systems, terms like federated query, data federation, and data virtualization are often used side by side. While related, each refers to a distinct concept in the data stack.
### Federated query: the action [#federated-query-the-action]
A federated query is the *act* of executing a single SQL query across multiple data sources — often in real time — without requiring data to be physically moved. It's a technical capability supported by modern query engines (like Firebolt, BigQuery, and Redshift) that enables direct access to data in systems like PostgreSQL, Snowflake, or S3.
Think of it as: query across systems without replicating data.
It typically supports:
* Standard SQL syntax
* Source-specific optimizations (e.g., pushdowns)
* Joins, filters, and aggregations across remote sources
### Data federation: the architecture [#data-federation-the-architecture]
Data federation refers to the *overall architecture* or strategy of presenting multiple heterogeneous data sources as if they were a single logical database. A federated database system supports this by managing metadata, access control, and connectors that power federated queries under the hood.
Think of it as: unify access to distributed data, abstracting away source details.
It often includes:
* Metadata management
* Schema mapping and normalization
* Access policies across data sources
* Governance and auditing across systems
### Data virtualization: the abstraction layer [#data-virtualization-the-abstraction-layer]
Data virtualization is a broader concept that emphasizes the *presentation layer* — allowing applications and users to interact with remote data as if it were local, without knowing where or how it's stored. It's often implemented via logical views or APIs.
Think of it as: pretend the data is local, even when it's remote.
It may include:
* Logical views with no physical data movement
* API or JDBC-based access to external sources
* Masking of schema differences and query complexity
## Firebolt's approach to federated query [#firebolts-approach-to-federated-query]
While many data platforms now support federated queries, few are optimized for the performance, concurrency, and scale demanded by modern data and AI applications. Firebolt takes a fundamentally different approach — built from the ground up for speed — by combining high-performance federated query execution with decoupled storage and compute architecture.
This means you can run fast, flexible queries across external tables, cloud data lakes, and remote databases — all while leveraging the full power of Firebolt's execution engine.
**High-performance federated queries with Firebolt's decoupled storage and compute**
At the heart of Firebolt's architecture is its decoupled storage and compute model, which separates where data lives from how it's processed. This allows Firebolt to scale each independently for maximum efficiency and flexibility — a major advantage over traditional federated database systems that suffer from tight coupling and resource contention.
For federated queries, this architecture is critical:
* Storage can reside in object stores (like Amazon S3 or Delta Lake), relational databases (like PostgreSQL or Snowflake), or other external sources.
* Compute resources can be spun up elastically to execute queries — even across federated sources — without moving the data.
This allows Firebolt to support federated query workloads at scale, while maintaining sub-second performance and cost efficiency.
**How Firebolt optimizes queries across sources**
Firebolt's federated query engine is designed to minimize latency and maximize performance when accessing and joining data from multiple systems. Firebolt can define external tables on top of file-based storage systems such as:
* Amazon S3 (Parquet, CSV, JSON)
* Delta Lake
* Apache Iceberg
* Custom lakehouse formats
Instead of ingesting the data, Firebolt reads directly from these sources using high-speed, vectorized I/O, and applies filter and projection pushdowns to minimize the amount of data scanned.
Benefits:
* No need for ingestion pipelines
* Schema-on-read with support for nested structures
* Near-instant query response times with indexing and caching
Whether you're blending clickstream logs from S3 with warehouse tables or analyzing machine-generated data at scale, Firebolt enables real-time querying of your lakehouse — no ETL required.
**Other databases (via connectors)**
In federated environments, organizations often need to join Firebolt-native tables with live data in operational databases like:
* PostgreSQL
* MySQL
* Snowflake
* SQL Server
Firebolt provides built-in connectors to these systems, allowing you to define external tables that point to remote databases. Firebolt pushes down as much logic as possible — including WHERE clauses, JOIN keys, and projections — to the source system.
This reduces:
* Data transfer costs
* Latency from remote sources
* Load on Firebolt's compute engine
In doing so, Firebolt turns even complex federated queries across distributed systems into efficient, low-latency analytics pipelines.
**Firebolt's query engine vs. traditional federated engines**
Traditional federated database systems — including older MPP engines or virtualization platforms — often struggle with performance when querying across sources. That's because they treat remote data sources as black boxes, requiring full table scans or complex joins in-memory. Firebolt's engine, by contrast, is designed for speed at every layer of federated query execution.
### Sub-second performance [#sub-second-performance]
Firebolt consistently delivers sub-second query latency, even when federating across cloud storage or remote databases. This is achieved through a combination of:
* Columnar storage format and vectorized execution
* Query plan optimization based on source capabilities
* Asynchronous prefetching and just-in-time execution
This makes it ideal for use cases like:
* Customer 360 dashboards
* AI feature stores
* Real-time personalization
* Time-sensitive operational reporting
### Smart indexing and pre-aggregation [#smart-indexing-and-pre-aggregation]
Most federated query engines can't leverage indexes or pre-aggregations across remote systems — which means every query can become a full scan.
Firebolt changes that by supporting:
* Aggregating indexes on local or external data
* Sparse indexes to skip irrelevant blocks
* Pre-aggregated materialized views on top of federated sources
This allows for blazing-fast filtering, grouping, and joining, especially when combined with external metadata caching and adaptive execution plans.
### Parallel distributed execution [#parallel-distributed-execution]
Firebolt's compute engine is inherently parallelized and distributed, meaning federated queries are broken into subtasks that execute concurrently across nodes.
Each subtask:
* Pulls from its assigned source(s)
* Executes partial aggregations or filters
* Streams results into Firebolt for final assembly
This distributed architecture makes it possible to handle high-concurrency workloads — such as dashboards, API queries, or AI pipelines — even when the data spans multiple sources and formats.
## Common challenges and considerations [#common-challenges-and-considerations]
While federated queries offer powerful advantages — like reduced data movement, real-time access, and cross-platform analytics — they aren't without trade-offs. Understanding the limitations and architectural considerations of federated database systems is key to building performant, secure, and cost-effective data solutions. Here are the most common challenges teams face when implementing federated queries and how to plan for them.
**Latency and performance bottlenecks**
One of the most significant hurdles in any federated query implementation is performance variability. Unlike centralized queries that run within a tightly optimized warehouse, federated queries span across remote systems, often introducing unpredictable latency due to:
* Network round trips to each data source
* Lack of indexes or caching on external systems
* Cold start penalties when accessing data lakes (e.g., S3 or Delta)
For example, if your query touches a low-latency database like PostgreSQL *and* a high-latency object store like S3, the slowest source becomes the bottleneck. This is why it's critical to choose a federated database engine — like Firebolt — that supports parallel distributed execution, aggressive pushdowns, and intermediate caching to minimize delays.
Best practices include:
* Limiting the volume of data retrieved from remote sources
* Indexing data when possible
* Using materialized views or pre-aggregations on federated inputs
**Security and access controls across sources**
With federated queries, you're accessing multiple independent systems — each with its own authentication, permissions, and governance rules. This introduces a new layer of security complexity:
* Ensuring consistent user access across all platforms
* Managing row-level and column-level security policies
* Preventing unauthorized data exposure through federated joins
In database federation architectures, it's vital to have a unified security layer that respects the permissions of each source system. Tools like fine-grained IAM, token-based authentication, and auditing help enforce proper access control. Firebolt, for instance, supports scoped credentials, secure external table definitions, and usage auditing — enabling secure cross-source federation without opening attack surfaces.
**Query complexity and optimization**
Not all federated queries are created equal. Complex joins, nested filters, and cross-system aggregations can lead to inefficient execution plans if not handled carefully.
In traditional ETL pipelines, data is pre-modeled and transformed. But in a federated database system, all of that complexity plays out at runtime. This means developers must be mindful of:
* Join order and cardinality
* Which operations are *pushed down* vs. executed locally
* Whether filters are applied early enough to reduce data transfer
Without proper tuning, a single poorly-optimized federated query could consume large amounts of compute or strain underlying source systems. That's why federated engines like Firebolt include cost-based optimizers, execution hints, and metadata-aware planning to guide developers toward efficient patterns.
**Cost implications of querying live sources**
While federated queries can reduce ETL costs, they also shift expenses in other directions — especially when querying live data from cloud-based systems or across regions.
Common cost drivers include:
* Egress charges for cross-cloud data access (e.g., GCP to AWS)
* On-demand compute spikes when federated sources scale up during query execution
* Query volume pricing, such as bytes scanned (e.g., in BigQuery or Snowflake)
Additionally, federated systems without caching or pushdowns may transfer large volumes of data unnecessarily, increasing both cost and latency.
To manage cost effectively:
* Use query filters to limit scan size
* Monitor bytes scanned vs. returned
* Pre-aggregate or cache high-traffic federated results
* Use tools like Firebolt's materialized views and external table caching to optimize access patterns
**Vendor lock-in and proprietary connectors**
Some federated query platforms rely on proprietary connectors that lock you into a specific vendor ecosystem. This can be problematic for enterprises adopting data mesh or multi-cloud strategies, where portability and openness are key.
Vendor lock-in typically shows up as:
* Custom formats that don't translate outside the platform
* Connectors that only work with specific query engines
* Lack of standards compliance (e.g., limited ANSI SQL support)
In contrast, Firebolt's federated query capabilities are designed to be open and extensible, supporting industry standards like Apache Iceberg, Delta Lake, and Parquet for lakehouse integration — while providing JDBC-compatible connectors for popular databases. This gives teams the flexibility to build open, federated architectures without long-term dependency risks.
## Best practices for federated query implementation [#best-practices-for-federated-query-implementation]
Adopting a federated query architecture unlocks powerful capabilities — but to get the best performance, cost efficiency, and maintainability, it's critical to follow well-established implementation strategies. Whether you're querying external files in a lakehouse or joining across cloud databases, these best practices will help you get the most out of your federated database stack.
**Use lazy evaluation when possible**
In the context of federated queries, lazy evaluation means delaying execution until the result is absolutely needed — especially for remote or large datasets. This ensures that unnecessary reads or joins aren't triggered prematurely, which is essential when accessing external sources like Amazon S3 or remote PostgreSQL databases.
By applying lazy evaluation strategies:
* Queries only pull data once downstream logic is finalized
* Intermediate steps that don't affect final output are skipped
* Data movement across networks is minimized
Firebolt's query planner supports deferred execution and intelligent optimization, helping avoid compute-intensive work when results aren't consumed immediately. For federated databases, this pattern is especially valuable when building exploratory queries, BI dashboards, or AI pipelines where not all data paths are accessed at once.
**Push down filters and aggregations**
Pushdown optimization is one of the most important levers in a federated query engine. Instead of retrieving all rows from remote sources and filtering them locally, pushdowns delegate operations like WHERE, JOIN, and GROUP BY directly to the source system — where compute is cheaper and data volumes are smaller.
This is a critical performance enhancer when federating:
* Data lakes (e.g., filtering Parquet files on S3)
* Operational databases (e.g., reducing query impact on PostgreSQL)
* Cloud warehouses (e.g., Snowflake joins and aggregations)
Firebolt's federated query engine automatically applies pushdowns when supported by the source. You can further optimize by:
* Writing explicit filters and predicates
* Avoiding overly nested subqueries
* Using source-compatible functions and data types
This best practice helps reduce scan size, improves federated query performance, and minimizes resource strain on both Firebolt and your external systems.
**Use metadata caching where supported**
Metadata caching refers to the ability to store schema information and statistics locally instead of fetching it on every query. In federated environments where source systems are remote — or have slow schema discovery — metadata caching can drastically reduce query planning time.
Benefits of metadata caching for federated queries include:
* Faster query compilation and optimization
* Reduced overhead on metadata-heavy systems (e.g., object stores)
* More responsive BI tools and ad-hoc exploration
In database federation architectures, this becomes especially important when connecting to systems like:
* S3/Delta/Iceberg with evolving schemas
* External catalogs or metastore APIs
* Cloud SQL sources with access throttling
Firebolt supports metadata introspection and caching for external tables, helping your queries remain fast, even in federated or semi-structured scenarios.
**Combine materialized views for hot data**
When certain federated queries are accessed repeatedly — for dashboards, machine learning models, or operational metrics — consider storing the results in materialized views. This strategy brings together the best of both worlds: federated flexibility and warehouse-level speed.
Materialized views in Firebolt:
* Persist results of complex joins or aggregations
* Can be scheduled to refresh incrementally
* Dramatically reduce response time for recurring queries
This is especially useful when federating across:
* Cold lake data (logs, events, IoT)
* Semi-structured datasets that require parsing
* Historical datasets joined with real-time data
Materializing these patterns ensures that frequent federated queries return sub-second results without hammering your source systems.
**Monitor source query cost and latency**
Just because you're not ingesting data doesn't mean you're not paying for it. Federated queries still consume compute, storage, and egress — often from multiple systems. That's why monitoring and optimizing source-level performance and costs is crucial for any federated database strategy.
Important metrics to track:
* Bytes scanned vs. returned from external sources
* Query latency per federated source
* Source-side CPU and IOPS impact
* Cross-region or cross-cloud egress charges
Firebolt provides built-in tools and query profiles that show exactly how federated sources are being accessed — so you can adjust filters, query frequency, or storage strategies accordingly.
## Getting started with federated query in Firebolt [#getting-started-with-federated-query-in-firebolt]
Firebolt makes it easy to get started with high-performance federated queries by combining a developer-friendly SQL interface with the scalability and speed of modern database federation. Whether you're joining Parquet files in S3 with real-time transactions in PostgreSQL, or blending cloud warehouse tables with lakehouse datasets — Firebolt gives you the tools to do it all, without the friction of traditional ETL.
**How to define external tables**
To run a federated query, you first need to tell Firebolt where your remote data lives. This is done by defining external tables, which are pointers to external storage systems or databases. These tables behave like native Firebolt tables in SQL, but they reference data outside of Firebolt.
Firebolt supports defining external tables for:
* Cloud object storage (Amazon S3, Delta Lake, Apache Iceberg)
* Parquet, CSV, JSON file formats
* External databases (e.g., PostgreSQL, MySQL, Snowflake) via connectors
Here's a basic example for a Parquet file in S3:
```sql
CREATE EXTERNAL TABLE s3_customer_logs (
customer_id VARCHAR,
interaction_type VARCHAR,
timestamp TIMESTAMP
)
STORED AS PARQUET
LOCATION = 's3://your-bucket/logs/';
```
For databases, you'll define a connection string and credentials to securely access the system — while keeping data in its source. In a federated database system, defining external tables is the foundation that enables real-time access to distributed data.
**Writing your first federated query**
Once your external tables are in place, writing your first federated query is no different from writing standard SQL. You can join external and internal tables, apply filters, aggregate results, and use all the capabilities of Firebolt's query engine.
Here's an example that joins three external sources:
```sql
SELECT
t.transaction_id,
t.transaction_amount,
l.timestamp,
c.customer_name
FROM
external_postgres.transactions t
JOIN
s3_customer_logs l
ON t.customer_id = l.customer_id
JOIN
external_snowflake.customer_profiles c
ON t.customer_id = c.customer_id
WHERE
l.interaction_type = 'CLICK'
AND t.transaction_date >= '2025-01-01';
```
This federated query accesses structured warehouse data, semi-structured log files, and real-time transactional data — all in one unified query interface.
**Connecting to other warehouses and lakes**
Firebolt offers robust connectors to a variety of external data systems, enabling cross-platform federation without complex middleware. Supported connections include:
* Cloud warehouses: Snowflake, Redshift
* Operational databases: PostgreSQL, MySQL, SQL Server, Aurora
* Data lakes: Amazon S3, Delta Lake, Apache Iceberg, Hudi
Each connection is defined once and then used to create one or more external tables. Firebolt supports secure credential storage, scoped access policies, and SSL encryption — ensuring that your federated database queries are both performant and compliant. This flexibility allows you to build lakehouse-style architectures, analyze streaming data, or enrich warehouse records with live operational context — all within Firebolt.
**Query optimization tips with Firebolt**
While Firebolt handles a lot of query planning behind the scenes, you can significantly improve federated query performance with a few best practices:
* Use pushdowns: Structure your SQL to push WHERE, JOIN, and GROUP BY logic down to the source. Firebolt automatically optimizes for this when possible.
* Index frequently filtered columns: Firebolt supports sparse and aggregating indexes even on federated data when materialized. These can drastically speed up filters and sorts.
* Materialize common joins: Use materialized views to cache the result of expensive cross-source joins, especially for dashboards or recurring reports.
* Monitor query plans: Use Firebolt's query profile and EXPLAIN plans to understand where data is pulled, filtered, and aggregated — and adjust accordingly.
* Filter early, scan less: Especially when working with large Parquet datasets in object storage, apply filters on partition columns to avoid unnecessary scanning.
By following these patterns, you'll ensure that your federated database implementation is not just functional, but blazing fast — with sub-second response times, even across distributed systems.
## Conclusion [#conclusion]
Federated queries have become an essential tool for modern data teams. Instead of spending time building and maintaining fragile ETL pipelines, developers and engineers can now query distributed data in place — gaining fast, actionable insights without the cost or complexity of data movement.
Whether you're joining structured and semi-structured data, querying across cloud platforms, or powering real-time dashboards, federated query technology enables a new level of agility and scalability. By abstracting the underlying data systems through a federated database architecture, teams can run cross-source analytics with standard SQL and zero duplication.
Firebolt takes this even further — delivering sub-second federated query performance thanks to its decoupled storage and compute, smart indexing, pushdown optimization, and massively parallel execution. It's not just federation; it's high-performance federation designed for today's most demanding analytics and AI workloads.
## FAQs [#faqs]
**Can a federated query join multiple cloud data warehouses?**
Yes — that's exactly what federated queries are designed to do. With Firebolt's federated capabilities, you can write a single SQL query that joins tables across multiple cloud data warehouses, including Snowflake, Redshift, BigQuery, and others, as well as data lakes like S3 and Delta Lake. Using Firebolt's external table definitions and connectors, your query engine treats these remote systems as part of a virtualized federated database, enabling seamless joins without the need to ingest or replicate the data first.
**Are federated queries slower than local queries?**
Not necessarily — but it depends on the query engine. In traditional systems, federated queries can suffer from latency due to network delays, lack of indexing, or inefficient scan operations. However, Firebolt changes the game with sub-second performance, even for federated workloads. Firebolt achieves this through:
* Smart pushdowns to minimize unnecessary data transfer
* Parallel query execution
* Metadata caching and smart indexing
* Columnar storage and vectorized processing
While querying remote sources will always involve more variables than querying local tables, Firebolt is purpose-built to ensure that even federated queries run fast at scale.
**How secure is data access in federated queries?**
Federated queries are only as secure as the data access model you implement — and Firebolt provides robust tools to ensure secure, governed, and compliant federated access across all your sources. Key security features include:
* Scoped credentials for external sources
* Token-based authentication and IAM integration
* Granular access controls at the table and column level
* Audit logging and query tracking
Because federated queries query data in place, you can maintain tight control over data residency, reduce unnecessary replication, and stay aligned with data governance policies.
**Do I need to replicate data into Firebolt to run federated queries?**
No — federated queries eliminate the need for data replication. With Firebolt, you define external tables that reference data in cloud warehouses, databases, or object stores like S3. These tables allow you to query remote data directly without ingesting it into Firebolt's storage. This approach is core to the database federation model, which lets you unify access to distributed data without physical consolidation. The result: faster time to insight, lower storage costs, and far less maintenance than traditional ETL-based pipelines.
# Reverse ETL - Definition (/glossary-items/reverse-etl)
Reverse ETL is the process of syncing data from your data warehouse into the SaaS tools used by sales, marketing, support, and operations teams. Unlike [ETL pipelines](https://www.firebolt.io/glossary-items/etl-and-elt) that centralize raw data for analysis, Reverse ETL operationalizes clean, transformed data, pushing it into CRM's, ad platforms, and support systems so teams can act on it without writing SQL or jumping into dashboards. It's a practical way to activate your warehouse as a source of truth across the business.
## Reverse ETL [#reverse-etl]
The process of moving processed data from a [data warehouse](https://www.firebolt.io/glossary-items/data-warehouse) into external operational systems such as CRMs, ad platforms, and support tools. It enables business teams to access actionable data inside the applications they already use, without requiring [SQL](https://www.firebolt.io/blog/postgresql-swiss-army-knife-and-the-analytics-workload) queries or dashboard access.
## How Reverse ETL Works [#how-reverse-etl-works]
* **Extract:** Queries are run against [warehouse](https://www.firebolt.io/blog/how-ai-is-transforming-etl-in-data-warehousing) tables or materialized views to pull records intended for operational use.
* **Transform (Optional):** Data may be formatted to meet destination tool requirements, including schema changes, typecasting, field enrichment, or filtering. Some workflows use pre-transformed tables to skip this step.
* **Load:** Data is delivered into external systems via APIs, SDKs, or direct database connections. Common targets include Salesforce (CRM data syncs), HubSpot (marketing automation updates), or Zendesk (ticket enrichment).
## Key Components [#key-components]
* **Data Mapping**: Establishes precise field-to-field relationships between warehouse data models and destination system schemas. Supports format validation and error handling for mismatched types or missing fields.
* **Scheduling**: Defines sync cadence, ranging from near-instant event-driven pushes to fixed intervals like every 15 minutes or hourly. Critical for balancing system load against operational freshness requirements.
* **Monitoring**: Tracks sync job success rates, latency, data volume discrepancies, and error rates. Alerts and audit logs are used to detect and resolve pipeline failures before they impact operations.
* **Version Control**: Manages change tracking for field mappings, transformation scripts, and destination schemas. Supports rollback to known good configurations in case of deployment errors or unexpected schema changes.
## Common Use Cases [#common-use-cases]
Reverse ETL supports business workflows by embedding warehouse data into operational systems.
* **Sales Support**: Sends product usage or engagement signals into CRM systems to drive lead scoring, pipeline prioritization, and sales enablement.
* **Marketing Segmentation**: Builds dynamic, high-resolution audiences based on behavioral or demographic warehouse data for targeted advertising and email campaigns.
* **Support Context**: Enriches support platforms with customer billing history, usage patterns, or subscription statuses to improve ticket triage and resolution speed.
* **Internal Ops**: Pushes alerts, operational KPIs, or workflow triggers into collaboration tools like Slack or Notion to keep teams updated without context switching.
## Supported Destinations [#supported-destinations]
Reverse ETL platforms offer integrations with a wide range of external business applications.

## Reverse ETL vs. ETL vs. ELT [#reverse-etl-vs-etl-vs-elt]
Each method manages a [different point](https://www.firebolt.io/blog/etl-vs-elt-know-the-differences) in the data flow lifecycle.
* **ETL**: [Extracts data](https://www.firebolt.io/blog/etl-best-practices) from source systems, transforms it for analysis, and then loads it into a data warehouse.
* **ELT**: Extracts and loads raw data into a warehouse first, performing transformations inside the warehouse using SQL or native functions.
* **Reverse ETL**: Extracts processed, trusted data from a warehouse and loads it into operational tools where non-technical teams can use it directly.
## Challenges [#challenges]
Reverse ETL introduces new technical and operational challenges that require careful management.
**Data Freshness**
* Sync frequency limitations (15 min/hourly) lead to lag between warehouse updates and destination tools
* Impacts customer-facing workflows relying on recent data (e.g. personalization, scoring)
* Solutions: Change data capture (CDC), tools with low-latency sync, lag monitoring
**Schema Drift**
* Changes in source schema (e.g. renamed columns, dropped fields) break syncs silently
* Operational systems may run on incorrect assumptions without immediate errors
* Solutions: Schema versioning, CI/CD validation checks, metadata monitoring
**Rate Limits**
* Destination APIs (Salesforce, HubSpot, etc.) throttle requests—can stall large sync jobs
* Backpressure affects critical updates (e.g. billing events, lifecycle changes)
* Solutions: Respect API limits, implement smart batching and retry logic
**Failure Recovery**
* Sync errors can corrupt downstream systems with partial or stale data
* Hard to detect without visibility into write results or failure types
* Solutions: Alerts on partial failures, idempotent writes, rollback plans, detailed observability
## Benefits [#benefits]
Reverse ETL maximizes the business value of centralized data infrastructure.
* **Puts Data to Work**: Operationalizes curated warehouse data across CRMs, marketing platforms, and support tools without additional engineering work.
* **Keeps Teams Aligned**: Ensures all departments access consistent, up-to-date data inside the applications they already use.
* **Cuts Down on Manual Tasks**: Eliminates the need for ad hoc scripts, CSV exports, or manual integrations.
* **Supports Targeted Communication**: Powers detailed, behavior-driven segmentation for ads, emails, and customer support interactions.
* **Makes Better Use of the Warehouse**: Activates data assets already modeled, cleaned, and governed, extending their value beyond reporting.
* **Speeds Up Execution**: Reduces the lag between analysis and action by automating data delivery into operational workflows.
* **Increases Data ROI**: Turns the warehouse into a live operational hub, not just a passive reporting layer.
## Tools That Support Reverse ETL [#tools-that-support-reverse-etl]
Several specialized platforms streamline Reverse ETL setup and management. These tools focus on bridging the gap between cloud data warehouses and operational systems by offering prebuilt connectors, scheduling, monitoring, and data transformation layers.
* **Hightouch** — One of the earliest platforms to define the Reverse ETL category. Hightouch supports a wide range of destinations, including CRMs, marketing tools, and ad platforms. It allows users to define sync logic with SQL or a visual editor, supports custom scheduling, and offers features like field-level mapping, sync logs, and versioning. Popular with teams that want fast time-to-value without building their own pipelines.
* **Census** — Focused heavily on data reliability and observability. Census lets teams sync data from cloud warehouses like Snowflake or BigQuery to tools like Salesforce, Marketo, and Zendesk. It includes features like row-level change detection and built-in testing. Integrates cleanly with dbt, making it a fit for teams already using modern data stacks.
* **RudderStack** — Started as a customer data platform (CDP), but now supports Reverse ETL as part of its broader offering. RudderStack appeals to engineering teams with its code-first approach and support for both event streaming and warehouse-to-tool syncs. It's a good fit when Reverse ETL is one part of a broader data movement strategy.
* **Polytomic** — Built with a visual-first design, Polytomic offers a no-code interface to model and sync data from warehouses to business tools. It's tailored to less technical users who still need fine-grained control over sync logic. Good for ops teams managing workflows in CRMs or support tools without relying on data engineering.
* **Omnata** — Designed for teams that work primarily in Salesforce. Omnata allows live querying of warehouse data directly inside Salesforce via connected apps and sync layers. Unlike batch-based tools, it focuses more on enabling embedded analytics and operational dashboards from warehouse data.
## When to Use Reverse ETL [#when-to-use-reverse-etl]
Reverse ETL makes sense when your organization is ready to operationalize warehouse data at scale.
* You maintain a centralized data warehouse like Snowflake, BigQuery, or Redshift.
* Your business teams actively work in external systems like Salesforce, Marketo, or Zendesk.
* You want to automate processes such as lead scoring, audience building, or support ticket enrichment directly from trusted warehouse data.
## Frequently Asked Questions [#frequently-asked-questions]
### What's the difference between Reverse ETL and ETL? [#whats-the-difference-between-reverse-etl-and-etl]
Their [difference](https://www.firebolt.io/blog/etl-vs-elt-know-the-differences) is that ETL moves raw data into a warehouse for analysis. Reverse ETL sends transformed data back out to tools like CRMs and ad platforms for operational use.
### Do I need Reverse ETL if I already use BI dashboards? [#do-i-need-reverse-etl-if-i-already-use-bi-dashboards]
Yes. BI tools are for analysis. Reverse ETL brings that insight into everyday tools where business teams can actually use it.
### What kind of data is typically synced with Reverse ETL? [#what-kind-of-data-is-typically-synced-with-reverse-etl]
Common examples include customer segments, product usage, support status, and revenue metrics. These are used by sales, marketing, and support teams.
### How often can data be synced? [#how-often-can-data-be-synced]
Sync frequency depends on the platform. Some tools offer batch updates hourly or daily. Others support near-instant updates through APIs.
### What tools does Reverse ETL integrate with? [#what-tools-does-reverse-etl-integrate-with]
Most platforms support CRMs like Salesforce, ad tools like Facebook and Google Ads, support systems like Zendesk, and internal tools like Slack.
# What is a Serverless Data Warehouse? (/glossary-items/serverless-data-warehouse)
Data warehouses without servers act like regular centralized repositories of information that can be accessed for analysis to aid decision making. Data and its analysis are now an indispensable tool for business competitiveness. Users often depend on reports, dashboards, and analytics tools to access valuable insights from the data they gather, so they can monitor business performance and derive important decisions from it. Data warehouses often support these tools. In this structure, data flows into the data warehouse with the help of ETL or ELT processes from various sources, such as transactional systems or relational databases. With (self-service) BI tools, spreadsheets, SQL, and other tools and services, business analysts, data engineers, and data scientists can access the data and work with it.
The advantage of a serverless data warehouse is that it is easier to operate compared to the classic on-prem [data warehouse](https://www.firebolt.io/glossary-items/data-warehouse). The provider is responsible for the infrastructure. So, there is no need to set up servers, firewalls, etc., but rather, you get IT on-demand, so to speak. There is also no need to worry about scaling resources. If more resources are needed, they are provided automatically.
## Advantages of a serverless data warehouse [#advantages-of-a-serverless-data-warehouse]
* **Accessibility and Ease of Use:** Unlike on-premise data warehouses, one can get access to [cloud data warehouses](https://www.firebolt.io/blog/cloud-data-warehouse) from anywhere in the world. Furthermore, today's data warehouses feature certain control factors that make sure that the data needed for business intelligence can only be seen by the affected people working in a company. It is interesting to note that data integrity is always ensured, even when the data within the data warehouse is accessed by multiple users simultaneously. Hence, companies don't have to fear that they will limit their data quality because of possible accessibility problems. In addition, modern data warehouses themselves or the associated BI layers are usually designed in a very user-friendly way, so that users from specialized fields can easily create and share their own reports.
* **Scalability and Elasticity:** The cloud architecture of SaaS-based data warehouses allows businesses to adjust their resource allocation to meet changing demands. With a cloud data warehouse, businesses with fluctuating requirements can pay only for the features and functionality they need. With on-premise data warehouses, you will have to buy more hardware if needed and still pay for it even when usage decreases. This automatic addition of resources (scaling) then works automatically and does not require any effort by your own IT staff.
* **Improved Performance:** SaaS data warehouses feature a distributed architecture where a number of servers carry the load together. The described servers make sure that the vast quantities of data are prepared at the same time. Cache and in-memory functions and more efficient architectures, such as column-based and [nested data structures](https://www.firebolt.io/glossary-items/data-nesting-and-data-unnesting), make these databases even faster than their predecessors. Also, the efficient and coordinated interfaces to the front-end make data analysis better and faster.
* **More and Cheap Data Storage:** One of the most important reasons for moving your data warehouse to the cloud is the expanded storage options. Cloud-based data warehousing solutions often offer a pay-as-you-go model for companies to create an individual storage space without any waste. The previously described approach can also be used with other features so that organizations are able to approach their [data warehousing](https://www.firebolt.io/glossary-items/data-warehousing) projects in different ways. With the right know-how, you can also save money, because cloud storage is often more cost effective.
* **Better Integration and More Data Formats:** Today's cloud data warehouse approaches are so created that they can include data from various sources such as cloud applications, databases, or file formats. Furthermore, even raw or semi-processed data can easily be filtered out and manifested within the structural architecture. With existing interfaces to other services within a cloud platform, the data can be easily linked to spreadsheet tools, self-service BI tools, or ML services, for example. This makes integrations and sharing of data within the company much easier and helps the company become more data driven. A common approach for this is using a [data lakehouse](https://www.firebolt.io/glossary-items/data-lakehouse), which combines the concepts of [data lakes](https://www.firebolt.io/glossary-items/data-lake) and data warehouses.
## Challenges [#challenges]
* **Data Governance and Security:** Data privacy laws and regulations have been a hot topic in the data world for several years. One example is the GDPR in Europe, which threatens stiff penalties for non-compliance. Organizations do have to comply with the strict governance regulations that handle everything that has anything to do with consumer data. A data warehouse should preferably feature data from various sources and make sure that the system is complete, without any gaps in between. Nonetheless, organizations can encounter difficulties if they omit any governance regulations or compliance requirements by mistake. Companies must therefore use additional encryption techniques for cloud data warehouses because the data is no longer stored in the company's own locked server room but is outsourced to the cloud provider.
* **Data Integration and Network:** This means that data no longer has to be distributed only within the company's own network, but also possibly across multiple cloud platforms. This means that companies have more work to do to create, secure, and monitor interfaces. Since the data is routed via the network, companies must of course also ensure an appropriate connection to the Internet, which is the only way to quickly transfer large volumes of data.
* **Know-How:** A problem for companies that is often underestimated is the lack of know-how. The cloud is a different world than the one of on-prem. In addition to the areas of network and company, a lot has happened in the area of data warehouses and data analysis. New architectures such as column-oriented or NoSQL databases have been developed and new services came on the market. However, the market is tight and companies are looking for employees in this area, but may not be able to get them. Companies must therefore look for suitable personnel at an early stage, be they data architects, data engineers, or scientists who are also familiar with modern cloud platforms and data warehouses.
# Slowly Changing Dimensions (SCDs) (/glossary-items/slowly-changing-dimensions-scds)
In the world of data engineering and database management, maintaining data's accuracy and historical integrity is paramount. This is where "Slowly Changing Dimensions" (SCDs) come into play. SCDs are a set of strategies and methods designed to handle changes in data over time while preserving historical records. As a data engineer, understanding SCDs and their types is crucial for effective data warehousing and business intelligence. In this article, we'll explore SCDs through examples and illustrations.
## Understanding SCD Types [#understanding-scd-types]
There are three common types of Slowly Changing Dimensions:
### SCD Type 1 (Overwrite) [#scd-type-1-overwrite]
SCD Type 1 is the most straightforward approach. When changes occur, it overwrites the existing data with the new values, effectively losing the historical information. Let's illustrate this with an example:
Example: Customer Address Change
Initially, the customer's data looks like this:

If John Doe's address changes to "456 Elm St," using SCD Type 1, the database will be updated like this:

This scenario's previous address is overwritten, and historical data is not retained.
### SCD Type 2 (Historical Tracking) [#scd-type-2-historical-tracking]
SCD Type 2 maintains a history of changes by creating new records when data changes occur. Each record has an associated timestamp or version number, allowing analysts to trace back to any point in time. Here's an example:
Example: Product Price Changes
Suppose you manage a product catalog. The initial record for a laptop might look like this:

When the laptop's price changes to $900, a new record is created:

SCD Type 2 preserves a complete historical view of price changes.
### SCD Type 3 (Limited History) [#scd-type-3-limited-history]
SCD Type 3 strikes a balance between Types 1 and 2. It maintains both the current and the previous values in separate columns, providing limited historical information. Here's an example:
Example: Employee Position Changes
Imagine you're managing employee records. Initially, Alina Smith was an Analyst:

When Alina gets promoted to a Manager, the database would store the change like this:

## Practical Applications for Data Engineers [#practical-applications-for-data-engineers]
As a data engineer, you'll encounter SCDs in various scenarios, such as customer data management, product catalogs, and employee records. Choosing the right SCD type depends on the specific business requirements and data characteristics.
Customer Data Management: SCDs help track changes in customer information, ensuring that historical data remains intact, which can be crucial for auditing and compliance.
Product Catalogs: When managing product data, SCDs assist in monitoring changes in attributes like price, category, or product names, allowing businesses to analyze price trends and product evolution over time.
Employee Records: SCDs play a significant role in tracking changes in employee positions, salaries, and roles, facilitating HR analytics and performance evaluation.
## Implementing SCDs as a Data Engineer [#implementing-scds-as-a-data-engineer]
To implement SCDs effectively, data engineers often use Extract, Transform, and Load (ETL) processes. These processes extract data from source systems, apply the necessary transformations to manage dimension changes, and load the data into a data warehouse or mart. ETL tools and data integration platforms automate these tasks, ensuring data accuracy and historical preservation.
In conclusion, Slowly Changing Dimensions is a fundamental concept in data engineering, enabling businesses to maintain historical data accuracy while accommodating changes over time. Whether you choose SCD Type 1 for simplicity, SCD Type 2 for comprehensive historical tracking, or SCD Type 3 for balancing the two, align your SCD strategy with your business needs to ensure that your data remains a valuable asset for decision-making, analysis, and reporting.
By mastering SCDs and their applications, data engineers contribute to the foundation of robust and insightful data-driven decision-making within organizations.
# What is a Spark Job? (/glossary-items/spark-job)
Apache Spark is an open-source unified analytics and data processing engine for big data. Its capabilities include near real-time or in-batch computations distributed across various clusters.
Simply put, a Spark Job is a single computation action that gets instantiated to complete a Spark Action. A Spark Action is a single computation action of a given Spark Driver. Finally, a Spark Driver is the complete application of data processing for a specific use case that orchestrates the processing and its distribution to clients. Each Job is divided into single "stages" of intermediate results. Finally, each stage is divided into one or more tasks.

It should be clear by now that a Spark Job is simply one of the single units of execution used to achieve the maximum possible configurability for cluster affinity and parallelization of resources.
By default, these Jobs are executed by Spark's scheduler in keeping with the FIFO (first in, first out) ordering—the first job gets priority on all cluster resources. In the latest versions of Spark, it is also possible to configure fair sharing, assigning tasks between jobs following the "round robin" algorithm, so that all jobs get an approximately equal share of cluster computing resources.
Spark distribution can work with different clustering technologies such as Mesos, Hadoop Yarn, or Kubernetes. Spark will take care of distributing the Job's workload among different cluster nodes through a cluster manager.
Most [cloud warehouse](https://www.firebolt.io/blog/cloud-data-warehouse) providers also offer Spark integration services called Spark Connectors. Amazon S3 has **"s3a connector"**, Azure Storage has **"wasb connector"**, and Google Cloud Storage has **"gs connector"**, making Spark an all-round software computational unit for modern distributed [data-warehousing](https://www.firebolt.io/glossary-items/data-warehousing) systems.
## Advantages [#advantages]
Spark and Spark Jobs, specifically, have many advantages, which are mostly related to speed and ease of use. Apart from those, using Spark Jobs provides the following additional advantages:
* **Lazy Evaluation**: Spark automatically triggers processing only when a specific Spark Action (or Job) is run. It's possible to classify Spark Jobs into "actions" and "transformations", which allows Spark to optimize the overall performance of the system by triggering transformations of data only when they're effectively needed and when all the actions needed have been executed correctly.
* **Easy Parallelism**: Spark's Jobs are easily configurable to be run in parallel. It's easy to split data into several partitions so that they can be processed in parallel and independently.
* **Caching of Intermediate Results**: Spark automatically manages to cache intermediate results. If Spark discovers that the results of a previous computation can be reused in a subsequent execution, it automatically reuses the previous data without re-executing the workload. It is still possible to configure peculiar types of caching among different jobs in different tasks.
* **Highly Configurable Computational Tasks**: We have seen that Jobs are executed by default in a FIFO fashion. If enabled, the fair scheduler also supports grouping jobs into *pools* by setting the weight for different scheduling options for each pool of jobs. This approach is modeled after the Hadoop Fair Scheduler. In this way, it's possible to create a pool of "high-priority" jobs whose computational analyses need to be prioritized. It's even possible to group the jobs of each user together.
* **Third-Layer Application Servers Available**: Despite the great parallelization available when handling workloads with Spark Jobs, Spark is not really well suited for concurrency. In fact, multiple concurrent requests for data analysis should be handled by interposing a third-layer application server like Apache Livy or Apache Hadoop between the clients and the Apache Spark Context Manage.
* **Resiliency**: The distribution of Jobs among different nodes relies on resilient distributed datasets (RDD), appropriately designed to handle the failure of any worker node in the cluster and make sure that data loss can be approximated to 0.
## Challenges [#challenges]
Adopting a distributed unified analytics and data processing engine comes with different challenges that need to be considered carefully.
* **Optimization:** Optimizing a Spark Job means taking care of the various aspects of processing, which can be difficult. Splitting datasets into multiple partitions so that the computation can be run in parallel is not always an easy task, as single data information could be more frequently than not heavily correlated and not atomisable. Caching intermediate results should be considered carefully, as handling a huge amount of data could result in heavy cluster resource consumption.
* **Harder to Read and Write**: When integrating Apache Stark with distributed data storage, it's important to remember that object stores do not function in the same way as filesystems; they are significantly different. Reading and writing data can be significantly slower than working with a normal filesystem.
* **Out Of Memory Issues**: Memory issues are one of the most frequent problems in designing Spark Applications. Driver memory can be affected by huge collectors or broadcasting. Executor memory issues can be caused by analyzing big partitions or by setting up very high concurrency.
* **Configuring the number of executors**: when designing an apache spark cluster, a particular focus should be put on the division of the job workload into different executors. Having too many executors with low memory can result easily in out-of-memory problems while having too few executors can result in an inability to parallelize workloads.
* **Consistency**: even if Amazon S3, Google Cloud, or Microsoft Azure object stores are consistent (meaning that it's possible to read a file immediately after it has been written) none of the store connectors provide any guarantees as to how their clients cope with overwritten objects while a stream is reading them. This could result in Spark having inconsistent results if the data cloud providers are overwriting objects.
With all the underlying and atomic division among computational units, a simple job execution could become really hard to configure. Spark Jobs are not suited for simple enough data-intensive computation, as all the configuration power that comes from the adoption of Spark Jobs may be excessive for the vast majority of use cases.
# Simple SQL Optimization Tricks (/glossary-items/sql-optimization-tricks)
Here are some SQL optimizations that are meant to be simple to implement and not compromise the overall code readability while giving you a performance boost.
## 1. UNION ALL vs. UNION [#1-union-all-vs-union]
Both operators are used to combine the result sets of two or more SELECT statements. The difference is that **UNION** removes the duplicate records by performing a DISTINCT operation on the result set, whereas **UNION ALL** does not remove duplicates and is, therefore, faster than **UNION**.
**Solution**: Use UNION ALL instead of UNION when unioned results are mutually exclusive or when duplicates are not a problem.
## 2. Using leading wildcards when filtering [#2-using-leading-wildcards-when-filtering]
When you use leading wildcards in a query, e.g., `WHERE name like '%Miles%'`, the query is *not sargable* – which is a fancy way of saying that a query cannot take advantage of an index to speed up its execution.
If, on the other hand, you replace the leading wildcard with a trailing wildcard like the following: `WHERE name like 'Miles%'`, your query will be able to leverage any of the indexes you have defined on the column when performing the search – the query will be sargable.
**Solution**: Replace `'%string%'` with `'string%'` where possible.
## 3. Using OR in the JOIN predicate vs. UNION ALL [#3-using-or-in-the-join-predicate-vs-union-all]
The OR operator is an expensive database used to chain multiple JOIN conditions—especially when used in the JOIN clause. One way to boost the query performance is to split the JOIN predicates into individual queries and use UNION to combine the result sets.
Using **OR**:
```sql
SELECT p.name
FROM dim_date d
INNER JOIN fact_salesorder f ON p.dim_product = f.dim_productid OR p.market = f.store_market
```
Using **UNION ALL**:
```sql
SELECT p.name
FROM dim_product p
INNER JOIN fact_salesorder f ON p.dim_productid = f.dim_productid
UNION ALL
SELECT p.name
FROM dim_product p
INNER JOIN fact_salesorder f ON p.market = f.store_market AND p.dim_productid != f.dim_productid
```
**Solution**: Split conditions chained with OR in the JOIN predicate into multiple queries combined with UNION ALL.
## 4. Using ANY\_VALUE(column) vs. GROUP BY column [#4-using-any_valuecolumn-vs-group-by-column]
When writing analytical queries, you will often use the GROUP BY clause to report metrics qualified by the selected attributes you want to report on.
But sometimes, it doesn't make sense to add all the attributes in the GROUP BY - they are not grouping data at a higher granularity but are only used for display purposes alongside the other attributes and metrics.
For example, if you want to see the total sales by customer name and id, we might write the following query:
```sql
SELECT c.dim_customerid, c.name, sum(f.sales_amount)
FROM dim_customer c
INNER JOIN fact_salesorder f ON c.dim_customerid = f.dim_customerid
GROUP BY c.dim_customerid, c.name;
```
Suppose we know that each dim\_customerid corresponds to one single customer name. In that case, we can safely remove the name attribute from the GROUP BY and instead select one single value from the group using the ANY\_VALUE aggregate function in the select clause:
```sql
SELECT c.dim_customerid, any(c.name), sum(f.sales_amount)
FROM dim_customer c
INNER JOIN fact_salesorder f ON c.dim_customerid = f.dim_customerid
GROUP BY c.dim_customerid;
```
## 5. Using indexes to speed up the query performance [#5-using-indexes-to-speed-up-the-query-performance]
Indexes are the primary way for users to accelerate query performance. While indexes are common with Online Transaction Processing applications, they are used sparingly with big data. If you want to find out more on how to use indexes with a data warehouse, check out our blog on [Indexes in Action](https://www.firebolt.io/blog/firebolt-indexes-in-action).
# What is SQL? (/glossary-items/sql)
SQL (Structured Query Language) is the [standard programming language](https://www.firebolt.io/the-cloud-data-warehousing-guide/data-warehousing-fundamentals) for managing and querying data in relational databases. It enables users to store, retrieve, and manipulate structured data.
SQL serves as the backbone of relational databases and is at the core of data-driven applications. It enables organizations to handle structured data, making it crucial for:
* **Business Intelligence and Reporting:** Extracting insights from large datasets.
* **Transactional Systems:** Managing data consistency and integrity in applications.
* **Cloud Data Warehouses:** Powering analytics with scalable, high-performance queries.
### Key Characteristics of SQL [#key-characteristics-of-sql]
In cloud analytics platforms like Firebolt, SQL remains the interface for high-speed query execution, now paired with vectorization and indexing under the hood.
* **Declarative vs. Procedural Nature:** SQL is [declarative](https://www.firebolt.io/blog/vector-databases-wont-replace-sql---andy-pavlo), meaning users specify what data they need rather than writing step-by-step instructions on how to retrieve it.
* **ANSI/ISO Standardization:** SQL follows standards set by [ANSI](https://www.firebolt.io/knowledge-center/faq) (American National Standards Institute) and ISO (International Organization for Standardization), ensuring compatibility with platforms like MySQL, PostgreSQL, SQL Server, and Oracle.
* **Designed for Structured Data:** SQL organizes data in tables with predefined columns and relationships. Each table stores specific types of information, and relationships between tables are maintained using primary and foreign keys.
### A Brief History of SQL [#a-brief-history-of-sql]
SQL began at IBM in the early 1970s as part of the System R project, originally named SEQUEL. It later became the standard for working with structured data. Today, SQL remains the foundation of most data systems, from legacy enterprise tools to modern cloud warehouses.
## Core Components and Architecture [#core-components-and-architecture]
SQL systems have key components that structure, store, and process data, ensuring consistency and reliability. Modern data warehouses like Firebolt still rely on these key SQL elements but with significant upgrades in execution, storage, and scalability
### Key Components of SQL Systems [#key-components-of-sql-systems]
*In Firebolt, the optimizer chooses the best execution plan, then leverages indexes and vectorized execution to reduce latency.*
* **Databases and Tables**: Data is stored in structured tables with primary keys for unique identification and foreign keys for relationships.
* **Queries and Statements**: SQL uses [commands](https://www.firebolt.io/blog/5-steps-to-debug-your-complex-sql-queries-in-firebolt) like SELECT, INSERT, UPDATE, and DELETE for data retrieval and modification.
* **Constraints and Stored Procedures**: Constraints enforce rules like NOT NULL and UNIQUE, while stored procedures automate tasks.
* **Indexes and Views**: Indexes speed up searches, and views provide pre-processed data access without modifying tables.
### SQL Execution Process [#sql-execution-process]
1. **Parsing**: The system checks syntax and rules before execution.
2. **Optimization**: The [query optimizer](https://www.firebolt.io/glossary-items/sql-optimization-tricks) selects the most efficient way to retrieve or modify data.
3. **Storage Engine Operations**: The [engine retrieves](https://www.firebolt.io/blog/engines-online-scaling-and-upgrades) or updates data while ensuring consistency and security.
## Essential SQL Commands and Operations [#essential-sql-commands-and-operations]
### SQL commands are divided into five categories [#sql-commands-are-divided-into-five-categories]
**DDL (Data Definition Language): Defines structures**
```sql
CREATE TABLE customers (id INT PRIMARY KEY, name VARCHAR(50));
```
```sql
ALTER TABLE customers ADD COLUMN email VARCHAR(100);
```
**DML (Data Manipulation Language): Modifies data**
```sql
INSERT INTO customers (id, name) VALUES (1, 'Alice');
```
```sql
UPDATE customers SET name = 'Bob' WHERE id = 1;
```
**DQL (Data Query Language): Retrieves data**
```sql
SELECT * FROM customers;
```
**DCL (Data Control Language): Manages access**
```sql
GRANT SELECT ON customers TO user1;
```
**TCL (Transaction Control Language): Handles transactions**
```sql
BEGIN TRANSACTION; COMMIT; ROLLBACK;
```
## ACID Properties in Transactions [#acid-properties-in-transactions]
ACID ensures reliable database transactions:
* **Atomicity**: Transactions are all-or-nothing.
* **Consistency**: The Database remains valid before and after a transaction.
* **Isolation**: Transactions do not interfere with each other.
* **Durability**: Committed changes are permanent.
Firebolt is ACID-compliant even at high concurrency, so complex analytics don't compromise data integrity.
## SQL Query Performance Best Practices [#sql-query-performance-best-practices]
Optimizing SQL queries improves speed and reduces resource usage. Data engineers should focus on execution plans, indexing, partitioning, and caching.
**Query Execution Plans:**
* Diagnose slow queries using EXPLAIN or EXPLAIN ANALYZE. In Firebolt, the EXPLAIN plan includes index usage, join type, and partition filters, all critical for fast queries.
* Avoid full table scans by using indexed columns.
* Optimize joins by selecting the right join strategy (nested loop, hash, merge). Firebolt supports hash joins, merge joins, and also allows pre-aggregated indexes to skip unnecessary joins altogether.
* Reduce sorting overhead with pre-sorted data or indexing.
**Indexing Strategies**
* Clustered indexes store data physically in sorted order.
* Non-clustered indexes create separate structures for faster lookups.
* Covering indexes include all required columns to avoid extra reads.
* Composite indexes combine multiple columns for better filtering.
**Partitioning and Denormalization Trade-Offs**
* Partitioning splits large tables to speed up queries.
* Range partitioning organizes data by values (e.g., date ranges).
* Hash partitioning distributes data evenly across partitions.
* List partitioning groups data into predefined categories.
* Denormalization reduces joins but increases storage and update costs.
* Firebolt supports range and hash partitioning on fact tables, enabling better pruning and reduced scan times for large-scale queries.
**Vectorized Execution and Caching**
* Vectorized execution processes multiple rows at once for faster performance.
* Result set caching stores previous query results to avoid recomputation.
* Metadata caching speeds up queries by skipping redundant schema lookups.
## SQL Security Considerations [#sql-security-considerations]
Protecting SQL databases from security threats is essential. Key challenges include injection attacks, access control, authentication, and encryption.
**Critical Security Challenges**
* **SQL Injection Attacks:** Attackers [manipulate queries](https://firebolt.io/faq/how-does-firebolt-mitigate-sql-injection-attacks) to gain unauthorized access or alter data.
* **Access Control:** [Poorly managed permissions](https://www.firebolt.io/resources/architecting-layered-security-in-a-cloud-data-warehouse) expose sensitive data.
* **Authentication:** Weak [login credentials](https://firebolt.io/faq/how-do-i-work-with-authentication-tokens-in-firebolt) increase the risk of breaches.
* **Data Encryption:** Unencrypted data is vulnerable to interception.
**Best Practices for Secure SQL Implementation**
* Use parameterized queries to prevent SQL injection.
* Limit privileges with role-based access control (RBAC).
* Enforce strong authentication with multi-factor authentication (MFA).
* Encrypt data at rest and in transit using AES and TLS.
* Monitor query logs to detect suspicious activity.
### How Firebolt Addresses SQL Security [#how-firebolt-addresses-sql-security]
Firebolt strengthens SQL security with [role-based access control](https://www.firebolt.io/data-security), encrypted data transmission, and secure authentication. It integrates with identity providers for centralized user management and ensures end-to-end encryption for data at rest and in transit. Query auditing and logging provide visibility into database activity, helping detect and prevent unauthorized access.
## Modern Applications and Use Cases of SQL [#modern-applications-and-use-cases-of-sql]
SQL powers data-driven applications across industries. It supports large-scale storage, analytics, and cloud-based processing.
### Common Applications of SQL [#common-applications-of-sql]
* **Data Warehousing**: Stores and organizes [massive datasets](https://www.firebolt.io/glossary-items/data-warehouse) for analysis and reporting.
* **Business Intelligence**: Powers dashboards and reports for [decision-making](https://www.firebolt.io/resources/guide-to-sub-second-analytics).
* **Real-Time Analytics**: Processes live data streams for immediate insights.
* **Cloud Computing**: Runs scalable, distributed databases called [cloud](https://www.firebolt.io/the-cloud-data-warehousing-guide/cloud-data-warehouse-features-and-benefits).
## Advanced SQL for Modern Data Warehouses [#advanced-sql-for-modern-data-warehouses]
SQL in cloud data warehouses operates differently from traditional relational databases. Modern architectures improve speed, scalability, and efficiency.
### Differences Between Traditional and Cloud Data Warehouse SQL [#differences-between-traditional-and-cloud-data-warehouse-sql]
* **Distributed Query Execution**: Traditional databases process queries on a single machine, while cloud data warehouses distribute workloads across multiple nodes. Firebolt optimizes this further with vectorized execution and indexing.
* **Columnar Storage & Vectorized Execution**: Unlike row-based storage, columnar formats reduce I/O and speed up analytical queries. Vectorized execution processes multiple rows in parallel for faster results.
* **Separation of Storage and Compute**: Traditional databases bundle storage and compute together, limiting flexibility. Cloud data warehouses separate them, allowing independent scaling based on workload demands.
### Advanced SQL for Modern Data Warehouses [#advanced-sql-for-modern-data-warehouses-1]
SQL in cloud data warehouses differs from traditional relational databases by improving speed, scalability, and efficiency.
* **Distributed Query Execution**: Traditional databases process queries on a single machine, while cloud warehouses distribute workloads across multiple nodes. Firebolt optimizes this further with indexing and vectorized execution.
* **Columnar Storage & Vectorized Execution**: Columnar formats reduce I/O, speeding up analytical queries. Vectorized execution processes multiple rows in parallel for faster results.
* **Separation of Storage and Compute**: Unlike traditional systems that bundle resources, cloud data warehouses scale storage and compute independently.
## Firebolt's SQL Implementation [#firebolts-sql-implementation]
Firebolt optimizes SQL for performance and scale in modern analytics.
* **High-performance analytics**: Processes large datasets with rapid execution.
* **Large-scale data processing**: It handles petabyte-scale workloads.
* **Complex query optimization**: Uses indexing and caching for faster queries.
Firebolt also integrates with modern data architectures:
* **Direct querying of data lakes**: Reads Parquet, ORC, and JSON files without preprocessing.
* **ELT capabilities**: [Loads raw data](https://www.firebolt.io/glossary-items/etl-and-elt) quickly and transforms it in place.
* **Performance optimizations**: Uses pushdown processing to reduce data movement.
### Firebolt vs. Traditional SQL Databases [#firebolt-vs-traditional-sql-databases]
| Feature | Firebolt | Legacy RDBMS (PostgreSQL, MySQL) | Cloud DWH (Snowflake, BigQuery) |
| -------------------------- | ----------------------- | -------------------------------- | ------------------------------------ |
| Query Execution | Distributed, vectorized | Single-node, row-based | Distributed, optimized for analytics |
| Storage Format | Columnar | Row-based | Columnar |
| Compute-Storage Separation | Yes | No | Yes |
| Query Performance | Sub-second | Slower for large datasets | Varies |
| Concurrency | High | Limited | High |
## SQL in Modern ELT Workflows [#sql-in-modern-elt-workflows]
Traditional ETL loads data **before** transformation, limiting flexibility. ELT loads raw data first, transforming it later for better scalability.
* **SQL-Based Data Loading**
* Supports direct data ingestion into cloud storage.
* Enables parallel loading and efficient processing.
* Works with multiple formats like Parquet, ORC, and JSON.
Firebolt supports **direct querying of data lake files**, enabling **high-speed ingestion** at 10TB/hour.
* **Pushdown Processing in Modern ELT**
* Moves computation closer to data to reduce movement.
* Transforms data at the source, improving efficiency.
### Traditional vs. Pushdown Processing [#traditional-vs-pushdown-processing]
| Approach | Data Movement | Compute Efficiency | Processing Speed |
| ----------- | ------------- | ------------------ | ---------------- |
| Traditional | High | Centralized | Slower |
| Pushdown | Low | Distributed | Faster |
Firebolt applies **SQL-based pushdown optimization** to accelerate queries, reduce costs, and improve performance.
### Why Firebolt's SQL is Built for High-Performance Analytics [#why-firebolts-sql-is-built-for-high-performance-analytics]
Firebolt's query engine processes SQL faster than traditional data warehouses, handling **high-throughput workloads** for analytics applications.
* **Use cases**: Customer-facing dashboards, real-time analytics, and machine learning pipelines.
* **Optimized SQL queries**: Firebolt benchmarks show significant speed improvements.
### SQL Best Practices for Firebolt [#sql-best-practices-for-firebolt]
* **Indexing**: Use primary and aggregating indexes to reduce scan time.
* **Partitions**: Organize large datasets efficiently for faster queries.
* **Aggregations**: Precompute results to avoid repetitive calculations.
# What is Stream Data Processing? (/glossary-items/stream-data-processing)
Stream data processing, or simply stream processing, is a Big data technology that goes by different names like real-time streaming analytics and event processing. It involves processing, storing, and analysing a constant data feed from various sources like servers, applications, payment processing networks, and security logs.
## Why is it necessary and what are its uses? [#why-is-it-necessary-and-what-are-its-uses]
Batch processing of data following a fixed schedule or by fixing a volume threshold (i.e., data was not immediately acted upon generation) is one of the most common processing methods in use. However, this method runs the risk of data becoming stale as data can lose its value when it is not realized at the moment. For example, customer buying patterns may indicate certain preferences specific to a sales event like Black Friday; when this data is processed later it may no longer be of relevance. Thus real-time data processing is essential. Moreover, data is generated more rapidly and data volumes have increased exponentially, so a spontaneous approach as stream processing is of significance today. As data is immediately acted upon in stream data processing, it not only meets the demands of this age but also provides valuable insights instantaneously. For this reason the technology has found uses in almost every industry. Some examples of utilizing stream data processing include real-time inventory management, quick customer/user personalisation (e.g., displaying products or ads similar to those viewed previously, movie and TV show recommendations on streaming platforms), social media feeds, and matching rider to driver on cab-hailing apps (i.e., connecting riders to drivers based on rider specifications, location, and driver availability).
Apart from its usage in predictive analytics, the technology is found valuable in detecting fraud and anomaly. For instance, stream processing enables quick and smooth transactions. So, credit-card processing delays that were typically seen in the traditional fraud detection methods are not an issue in the stream processing method. While a credit-card provider can benefit from the good client/customer experience that stream processing assures, a manufacturer can benefit from massive savings and prevent colossal wastage as the real-time process continuously spots errors in the production line that can be immediately rectified and increase production instead of identifying a whole batch as defective after completing production. Therefore, stream processing's usage extends to Internet of Things (IoT) edge analytics that companies can leverage.
### What are some related challenges? [#what-are-some-related-challenges]
Stream processing is incomplete without an effective stream processing application because data processing applications help convert raw data into use cases (e.g., location data, fraud detection, and excess inventory). So, companies can either utilize a stream processing software as Apache Kafka that is freely available or build one. Organizations choose to build their own stream processing system depending on the scalability, reliability, and fault tolerance that they desire. However, major challenges lie in building the application according to the company's requirements. For example, designing the application to accommodate a wealth of incoming data that may even be infrequent is a challenge. Another challenge is with handling data that is received from several sources, locations, and in different formats and volumes, as the application should be able to handle such variations and prevent disruptions.

Shown above is an example of a stream data processing architectural pattern that integrates various data sources into different destinations. Once the data is processed by the stream processing application, it can be persisted in data warehouses or data lakes. Cloud based data warehouses and data lakes are commonly used to store streaming data for data analyst and end user access. The ability to scale these backend data stores for capacity and performance is of importance depending on the volume, variety, and velocity of data. Data warehouse solutions such as Redshift, BigQuery, Firebolt, and Snowflake are widely used to address the challenges of scale effectively.
# What are Upserts? (/glossary-items/upserts)
An **upsert** is a database operation whose name combines "update" and "insert": it either updates an existing record if a matching one is found or inserts a new record if none exists. In a relational database management system (RDBMS), the in-built indexes make this straightforward, because they let the database quickly identify whether a specific record or entry already exists before deciding whether to update it or add a new row.
This article describes upsert functionality in general and is not specific to Firebolt.
## Understanding upserts [#understanding-upserts]
When dealing with large volumes of data, a relational database management system (RDBMS) helps users manage their database. It is this support system that offers key features to handle data. One feature of an RDBMS is the upsert. As the name suggests, the upsert command lets the user alter existing data by inserting a whole new row or just updating an existing one. A common scenario where an operation like upsert is useful is updating or expanding employee information in a database — for instance, adding a new row for an employee ID or changing employee contact details. Although the upsert is simply a command for an operation that must be executed, it is not universal: not all RDBMSs offer it as a command.
## Upsert examples [#upsert-examples]
### Updating an existing row [#updating-an-existing-row]
Here is an example of using upsert to update an existing table of employee details.
*SQL statement*
```sql
UPSERT INTO employees (id, name, email) VALUES (2, 'Shane', 'Shane@testxyzAsia.io');
```
*Current table*

*Table after upsert*

In the above example, the second row's old employee details were replaced with new employee information.
### Inserting a new row [#inserting-a-new-row]
The example below shows how upsert is used to insert a new row of employee details.
*SQL statement*
```sql
UPSERT INTO employees (id, name, email) VALUES (3, 'Raj', 'Raj@testxyzHR.corp');
```
*Current table*

*Table after upsert*

## Indexes, upserts, and the MERGE command [#indexes-upserts-and-the-merge-command]
The upsert is a simple function when performed on a database because the in-built indexes facilitate identification of a specific record or entry, enabling easy changes to an entry or entries. Without indexes, identifying specific records becomes difficult for operations like upsert. Although indexes have long been common in RDBMSs, they are not extensively found in data warehouses. Despite that, they are now considered essential in data warehouses because of the need to maintain a history of data with reference to the source data served into the ETL tool. A typical use case is maintaining slowly changing dimensions in a data warehouse: new records have to be inserted, data in the warehouse that is no longer in the source has to be flagged or removed, and updates made to the source must reflect in the warehouse. Upsert can help make some of these changes. However, doing the update and insert tasks separately for each instance can be tedious. Hence the MERGE command combines insert, update, and delete operations so these tasks are performed in one go without the need to write separate syntax for each. The MERGE command can be used to combine data from a source table into a target table; based on a specific condition, records can be updated, deleted, or inserted, simplifying the process of ingesting data.
# Firebolt Cookies Policy (/legal/cookies)
We use in our site Firebolt.io ("Site") cookies and similar files or technologies to automatically collect and store information about your computer, device, and Site usage, in order to improve their performance and enhance your user experience. We use the general term "cookies" in this policy to refer to these technologies and all such similar technologies that collect information automatically when you are using our Site where this policy is posted. You can find out more about cookies and how to control them in the information below.
If you do not accept the use of these cookies, please disable them using the instructions in this cookie policy or by changing your browser settings so that cookies from this Site cannot be placed on your computer or mobile device. Important: disabling cookies on this Site may seriously cripple the user experience and other features on the Site, to the point of rendering them useless.
In this Cookies Policy, we use the term Firebolt (and "**we**", "**us**" and "**our**") to refer to Firebolt Analytics Inc. Our Privacy Policy is available at Firebolt.io.
## What is a cookie? [#what-is-a-cookie]
Cookies are computer files containing small amounts of information which are downloaded to your computer or mobile device when you visit a website. Cookies can then be sent back to the originating website on each subsequent visit, or to another website that recognizes that cookie.
Cookies are widely used in order to make websites work, or to work more efficiently, as well as to provide information to the owners of the website. Cookies do lots of different jobs, like letting you navigate between pages efficiently, remembering your preferences, and generally improving the user experience. Cookies may tell us, for example, whether you have visited our Site before or whether you are a new visitor.There are two broad categories of cookies:
1. **First party cookies,** served directly by us to your computer or mobile device.
2. **Third party cookies,** which are served by a third party on our behalf. We use third party cookies for functionality, performance / analytics, marketing, unclassified and other technologies, and social media purposes.
Cookies can remain on your computer or mobile device for different periods of time. Some cookies are 'session cookies', meaning that they exist only while your browser is open. These are deleted automatically once you close your browser. Other cookies are 'permanent cookies', meaning that they survive after your browser is closed. They can be used by websites to recognize your computer when you open your browser and browse the Internet again.
## Web beacons [#web-beacons]
Cookies are not the only way to recognize or track visitors to a website. We may use other, similar technologies from time to time, like web beacons (sometimes called "tracking pixels" or "clear gifs"). These are small graphics files that contain a unique identifier that enable us to recognize when someone has visited our website. This allows us, for example, to monitor the traffic patterns of users from one page within our website to another, to deliver or communicate with cookies, to understand whether you have come to our website from an online advertisement displayed on a third party website, to improve website performance and to measure the success of email marketing campaigns. In most instances, these technologies are reliant on cookies to function, and therefore declining cookies prevents them from functioning.
If you don't want your cookie information to be associated with your visits to these pages, you can set your browser to turn off cookies as described further below. If you turn off cookies, web beacon and other technologies will still detect your visits to our Site; however, they will not be associated with information otherwise stored in cookies.
## Targeted advertising [#targeted-advertising]
Third parties may drop cookies on your computer or mobile device to serve advertising through our website. These companies may use information about your visits to this and other websites in order to provide relevant advertisements about goods and services that you may be interested in. They may also employ technology that is used to measure the effectiveness of advertisements. The information collected through this process does not enable us or them to identify your name, contact details or other personally identifying details unless you choose to provide these to us.
## How do we use cookies? [#how-do-we-use-cookies]
We use cookies to:
* Track traffic flow and patterns of travel in connection with our Site;
* Understand the total number of visitors to our Sites on an ongoing basis and the types of internet browsers (e.g. Chrome, Firefox, Safari, or Internet Explorer) and operating systems (e.g. Windows or Mac) used by our visitors;
* Monitor the performance of our Site and to continually improve it; and
* Customize and enhance your online experience.
## What types of cookies do we use? [#what-types-of-cookies-do-we-use]
The types of cookies used by us in connection with the Site can be considered 'essential website cookies', 'functionality cookies', 'analytics and performance cookies', 'marketing', 'unclassified', and 'other technologies'. We've set out some further information below, and the purposes of the cookies we set in the following table.
## 1. Cookies necessary for essential website purposes [#1-cookies-necessary-for-essential-website-purposes]
These cookies are essential to provide you with services available through this Site and to use some of its features, such as access to secure areas. Without these cookies we will not be able to provide services that you require, such as transactional pages and secure login accounts.
| Cookie name | Source | Expiry | Purpose |
| --------------------- | ----------------------------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| OptanonAlertBoxClosed | [www.firebolt.io](http://www.firebolt.io) | Session | Enabling the website not to show the message more than once to a user. |
| OptanonAlertBoxClosed | firebolt.io | 1 year | Enabling the website not to show the message more than once to a user. |
| OptanonConsent | [www.firebolt.io](http://www.firebolt.io) | Session | Enabling the website to remember the cookie preferences of returning visitors. |
| OptanonConsent | firebolt.io | 1 year | Enabling the website to remember the cookie preferences of returning visitors. |
| nlbi\_XXXXXXX | comeet.co | Session | Linking certain sessions to a specific visitor. |
| \_\_cf\_bm | hello.firebolt.io | Session | manage incoming traffic that matches criteria associated with bots. |
| \_cfuvid | hello.firebolt.io | Session | distinguish individual users who share the same IP address for the purpose of enforcing rate limiting rules. |
| merchant | stripe.com | 1 year | This domain is associated with Stripe, a company that provides payment processing software and application programming interfaces for e-commerce websites and mobile applications. |
## 2. Functionality Cookies [#2-functionality-cookies]
These cookies record information about choices you've made and allow us to tailor the website to you. These cookies allow us to provide you with our services in the way in which you have required, as you continue to use or come back to our Site. For example, these cookies allow us to:
* Save your location preference if you have set your location on the homepage in order to receive a localized information;
* Remember settings you have applied, such as layout, text size, preferences, and colors;
* Show you when you are logged in; and
* Store accessibility options.
| Cookie name | Source | Expiry | Purpose |
| ---------------------------- | ----------------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Hubspotutk | firebolt.io | 5 months | Keeping track of a visitor's identity. |
| Hubspotutk | [www.firebolt.io](http://www.firebolt.io) | Session | Keeping track of a visitor's identity. |
| messagesUtk | [www.firebolt.io](http://www.firebolt.io) | Session | Ensures continuity of the chat conversations and remembers the information of visitors who chat in multiple sessions. |
| \_\_cf\_bm | vimeo.com | Session | Supporting Cloudflare Bot Management. |
| \_\_cf\_bm | hubspotusercontent.com | Session | Supporting Cloudflare Bot Management. |
| \_\_cf\_bm | hs-analytics.net | Session | Supporting Cloudflare Bot Management. |
| \_\_cf\_bm | hsappstatic.net | Session | Supporting Cloudflare Bot Management. |
| \_\_cf\_bm | codepen.io | Session | Supporting Cloudflare Bot Management. |
| \_\_cf\_bm | usemessages.com | Session | Supporting Cloudflare Bot Management. |
| \_\_cf\_bm | hs-sites.com | Session | Supporting Cloudflare Bot Management. |
| \_\_cf\_bm | hs-scripts.com | Session | Supporting Cloudflare Bot Management. |
| AWSALBTG | api.typeform.com | 6 days | Used to attribute commission to affiliates when a visitor arrive at the website from an affiliate referral link. |
| AWSALBTG | form.typeform.com | 6 days | Used to attribute commission to affiliates when a visitor arrive at the website from an affiliate referral link. |
| SM | c.clarity.ms | Session | Used for synchronizing the MUID across Microsoft domains to facilitate consistent user tracking across different browser sessions |
| ANONCHK | c.clarity.ms | Session | Used to indicate whether a user is sampled for session recording; it ensures that the session data is tied to a unique ID for that specific visit. |
| visid\_incap\_xxxxxxx | comeet.co | 1 year | linking certain sessions to a specific visitor (visitor representing a specific computer). In order to identify clients that have already visited Incapsula. |
| \_cfuvid | vimeo.com | Session | A Cloudflare Rate Limiting cookie used to identify individual visitors sharing an IP address to ensure security and prevent DDoS attacks. |
| \_cfuvid | cdn.webflow\.com | Session | Used by Cloudflare to manage traffic spikes and rate-limiting policies, ensuring the CDN delivers content securely to the legitimate user. |
| incap\_ses\_xxxxxxxxxxxxxxxx | comeet.co | Session | Linking HTTP requests to a certain session (AKA visit) in order to maintain existing sessions (ie, session cookie) |
| MUID | clarity.ms | 1 year | Used as a unique user identifier to track visits across different Microsoft sites and help Clarity distinguish unique vs. returning users. |
| MR | c.clarity.ms | 6 days | Used to indicate whether to refresh the MUID; it helps maintain the accuracy of user tracking by checking if the identifier needs updating. |
| AWSALBTGCORS | form.typeform.com | 6 days | Used to attribute commission to affiliates when a visitor arrive at the website from an affiliate referral link. |
| AWSALBTGCORS | api.typeform.com | 6 days | Used to attribute commission to affiliates when a visitor arrive at the website from an affiliate referral link. |
| CLID | [www.clarity.ms](http://www.clarity.ms) | 1 year | Used to identify the first-time a user visited the site to provide long-term heatmaps and behavioral analytics for that specific browser. |
## 3. Performance / Analytics Cookies [#3-performance--analytics-cookies]
We use performance/analytics cookies to analyze how the website is accessed, used, or is performing. We do this in order to provide you with a better user experience and to maintain, operate and continually improve the website. For example, these cookies allow us to:
* Better understand our website visitors so that we can improve how we present our content;
* Test different design ideas for particular pages, such as our homepage;
* Collect information about Site visitors such as where they are located and what browsers they are using;
* Determine the number of unique users of the website;
* Improve the website by measuring any errors that occur; and
* Conduct research and diagnostics to improve product offerings.
| Cookie name | Source | Expiry | Purpose |
| --------------------------------- | ----------------------------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| \_\_hstc | firebolt.io | 5 months | Website analytics. |
| \_\_hstc | [www.firebolt.io](http://www.firebolt.io) | Session | Website analytics. |
| cb\_anonymous\_id | firebolt.io | 1 year | Attributing page views back to an anonymous individual and track the user's interactions and behavior on the website over time |
| cb\_anonymous\_id | [www.firebolt.io](http://www.firebolt.io) | Session | Attributing page views back to an anonymous individual and track the user's interactions and behavior on the website over time |
| \_hjTLDTest | firebolt.io | Session | Determining the most generic cookie path to use, instead of the page hostname. |
| cb\_user\_id | firebolt.io | 1 year | Keeps track of the visitor's identity by storing a unique user ID to attribute web activity to the individual |
| cb\_user\_id | [www.firebolt.io](http://www.firebolt.io) | Session | Keeps track of the visitor's identity by storing a unique user ID to attribute web activity to the individual |
| cb%3Atest | firebolt.io | 1 year | This cookie name is associated with analytics which is used to collect data on the user's visits to the website, such as the number of visits, average time spent on the website and what pages have been loaded with the purpose of generating reports for optimizing the website content. |
| \_\_hssrc | firebolt.io | Session | Website analytics. |
| \_\_hssrc | [www.firebolt.io](http://www.firebolt.io) | Session | Website analytics. |
| \_\_hssc | firebolt.io | Session | Website analytics. |
| \_\_hssc | [www.firebolt.io](http://www.firebolt.io) | Session | Website analytics. |
| rl\_anonymous\_id | firebolt.io | 1 year | Identifies a visitor to the website in an anonymous form and track user interactions with the site. |
| rl\_anonymous\_id | [www.firebolt.io](http://www.firebolt.io) | Session | Identifies a visitor to the website in an anonymous form and track user interactions with the site. |
| rl\_auth\_token | firebolt.io | Session | Stores the authentication token passed by the user. |
| rl\_group\_id | firebolt.io | Session | Stores the group ID set via the group API. |
| \_hjSessionUser\_xxxxxx | firebolt.io | 1 year | Persisting the Hotjar User ID, unique to that site on the browser; ensuring that behavior in subsequent visits to the same site will be attributed to the same user ID. |
| \_hjSessionUser\_xxxxxx | [www.firebolt.io](http://www.firebolt.io) | Session | Persisting the Hotjar User ID, unique to that site on the browser; ensuring that behavior in subsequent visits to the same site will be attributed to the same user ID. |
| \_ga | firebolt.io | 1 year | Registers a unique ID that is used to generate statistical data on how the visitor uses the web site |
| \_ga | [www.firebolt.io](http://www.firebolt.io) | Session | Registers a unique ID that is used to generate statistical data on how the visitor uses the web site |
| \_ga\_xxxxxxxxxx | firebolt.io | 1 year | Identify and track an individual session. |
| \_ga\_xxxxxxxxxx | [www.firebolt.io](http://www.firebolt.io) | Session | Identify and track an individual session. |
| rl\_group\_trait | firebolt.io | Session | Stores the group ID set via the group API. |
| rl\_page\_init\_referrer | firebolt.io | 1 year | Stores the initial referrer of the page when a user visits a site for the first time. |
| rl\_page\_init\_referrer | [www.firebolt.io](http://www.firebolt.io) | Session | Stores the initial referrer of the page when a user visits a site for the first time. |
| \_hjSession\_xxxxxx | firebolt.io | Session | Ensuring that subsequent requests within the session window will be attributed to the same Hotjar session. |
| \_hjSession\_xxxxxx | [www.firebolt.io](http://www.firebolt.io) | Session | Ensuring that subsequent requests within the session window will be attributed to the same Hotjar session. |
| rl\_page\_init\_referring\_domain | firebolt.io | Session | Stores the initial referring domain of the page when a user visits a site for the first time. |
| rl\_session | firebolt.io | 1 year | tracks a user's session, helping to provide insights into their behavior within a single visit to the website. |
| rl\_session | [www.firebolt.io](http://www.firebolt.io) | Session | tracks a user's session, helping to provide insights into their behavior within a single visit to the website. |
| \_uetvid | firebolt.io | 1 year | Engaging with a user that has previously visited our website. |
| \_uetvid | [www.firebolt.io](http://www.firebolt.io) | Session | Engaging with a user that has previously visited our website. |
| rl\_trait | firebolt.io | Session | Stores the user traits object set via the identify API. |
| rl\_user\_id | firebolt.io | Session | Stores the unique user ID for the purpose of Marketing/Tracking. |
| test\_rudder\_cookie | firebolt.io | 1 year | Checks whether the cookie storage of a browser is accessible or not. Once checked, the cookie is removed immediately. |
| rack.session | hubspot.clearbit.com | 6 days | stores session data for user sessions in Ruby on Rails web applications. |
| \_\_cf\_bm | hubspot.com | Session | Supporting Cloudflare Bot Management. |
## 4. Marketing/Advertising [#4-marketingadvertising]
We use marketing cookies to deliver many types of targeted digital marketing.. We do this in order to provide you with a better user experience and to maintain, operate and continually improve the website. The cookie store user data and behavior information, which allows advertising services to target audience according to variables. For example, these cookies allow us to:
* Observe the site performance and generate retargeting (Site retargeting, search retargeting, etc).
* Maintain and improve the website and our products
| Cookie name | Source | Expiry | Purpose |
| -------------------------- | ------------------------------------------- | ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| \_fbp | firebolt.io | 2 months | Used by Facebook to deliver a series of advertisement products such as real time bidding from third party advertisers. |
| \_fbp | [www.firebolt.io](http://www.firebolt.io) | Session | Used by Facebook to deliver a series of advertisement products such as real time bidding from third party advertisers. |
| cb\_group\_id | [www.firebolt.io](http://www.firebolt.io) | Session | This cookie is associated with Clearbit which is used to keep track of group information, like an account that a user belongs to or a company they work for, and to assign visitors into segments for website advertisement efficiency. |
| cb\_group\_id | firebolt.io | 1 year | This cookie is associated with Clearbit which is used to keep track of group information, like an account that a user belongs to or a company they work for, and to assign visitors into segments for website advertisement efficiency. |
| \_rdt\_uuid | [www.firebolt.io](http://www.firebolt.io) | \[Please confirm the expiry duration] | Tracking the views of embedded videos and performance of the advertisements. |
| \_gcl\_au | firebolt.io | 2 months | Used by Google AdSense for experimenting with advertisement efficiency across websites using their services |
| \_gcl\_au | [www.firebolt.io](http://www.firebolt.io) | Session | Used by Google AdSense for experimenting with advertisement efficiency across websites using their services |
| pfjs%3Acookies | firebolt.io | 1 year | This cookie name is associated with pfjs%3Acookies which is used to check if the user's browser supports cookies, collect data on visitor behavior from multiple websites to present more relevant advertisement, and set through the site by advertising partners to target the visitor with relevant advertisement. |
| \_uetsid | firebolt.io | 1 year | This cookie is used by Bing to determine what ads should be shown that may be relevant to the end user perusing the site. |
| \_uetsid | [www.firebolt.io](http://www.firebolt.io) | Session | This cookie is used by Bing to determine what ads should be shown that may be relevant to the end user perusing the site. |
| messagesUtk | firebolt.io | 1 year | Ensures continuity of the chat conversations and remembers the information of visitors who chat in multiple sessions. |
| SRM\_B | c.bing.com | 1 year | This domain is owned by Microsoft - it is the site for the search engine Bing. |
| \_\_cf\_bm | linkedin.com | 1 year | Supporting Cloudflare Bot Management. |
| \_\_cf\_bm | hsforms.com | Session | Supporting Cloudflare Bot Management. |
| \_cfuvid | hubspot.com | Session | This domain is owned by Hubspot. This company provides a range of online marketing and sales technology and services. |
| \_cfuvid | hsforms.com | Session | This domain is owned by Hubspot. This company provides a range of online marketing and sales technology and services. |
| \_\_Secure-xxxxxxx | youtube.com | Session | YouTube is a Google owned platform for hosting and sharing videos. YouTube collects user data through videos embedded in websites, which is aggregated with profile data from other Google services in order to display targeted advertising to web visitors across a broad range of their own and other websites. |
| VISITOR\_PRIVACY\_METADATA | youtube.com | 5 months | YouTube is a Google owned platform for hosting and sharing videos. YouTube collects user data through videos embedded in websites, which is aggregated with profile data from other Google services in order to display targeted advertising to web visitors across a broad range of their own and other websites. |
| bcookie | linkedin.com | 1 year | Used by the social networking service, LinkedIn, for tracking the use of embedded services. |
| bscookie | linkedin.com | 1 year | Used by the social networking service, LinkedIn, for tracking the use of embedded services |
| IDE | doubleclick.net | 1 year | Used by Google DoubleClick to register and report the website user's actions after viewing or clicking one of the advertiser's ads with the purpose of measuring the efficacy of an ad and to present targeted ads to the user. |
| lidc | linkedin.com | Session | Used by the social networking service, LinkedIn , for tracking the use of embedded services. |
| test\_cookie | doubleclick.net | Session | Used to check if the user's browser supports cookies. |
| VISITOR\_INFO1\_LIVE | youtube.com | 5 months | This cookie is used as a unique identifier to track viewing of videos |
| li\_gc | linkedin.com | 5 months | This domain is owned by LinkedIn, the business networking platform. It typically acts as a third party host where website owners have placed one of its content sharing buttons in their pages, although its content and services can be embedded in other ways. Although such buttons add functionality to the website they are on, cookies are set regardless of whether or not the visitor has an active Linkedin profile, or agreed to their terms and conditions. For this reason it is classified as a primarily tracking/targeting domain. |
| YSC | youtube.com | Session | YouTube is a Google owned platform for hosting and sharing videos. YouTube collects user data through videos embedded in websites, which is aggregated with profile data from other Google services in order to display targeted advertising to web visitors across a broad range of their own and other websites. |
| JSESSIONID | [www.linkedin.com](http://www.linkedin.com) | Session | This domain is owned by LinkedIn, the business networking platform. It typically acts as a third party host where website owners have placed one of its content sharing buttons in their pages, although its content and services can be embedded in other ways. Although such buttons add functionality to the website they are on, cookies are set regardless of whether or not the visitor has an active Linkedin profile, or agreed to their terms and conditions. For this reason it is classified as a primarily tracking/targeting domain. |
| lang | linkedin.com | Session | This domain is owned by LinkedIn, the business networking platform. It typically acts as a third party host where website owners have placed one of its content sharing buttons in their pages, although its content and services can be embedded in other ways. Although such buttons add functionality to the website they are on, cookies are set regardless of whether or not the visitor has an active Linkedin profile, or agreed to their terms and conditions. For this reason it is classified as a primarily tracking/targeting domain. |
| onfire\_id | onfire.ai | 12-month | To identify anonymous website visitors by matching browsing activity with third-party professional identity graphs for the purpose of personalized B2B marketing and sales outreach |
## How to control or delete cookies [#how-to-control-or-delete-cookies]
Most browsers allow you to change your cookie settings. These settings will typically be found in the "options" or "preferences" menu of your browser. In order to understand these settings and learn how to use them, please consult the "Help" function of your browser, or the documentation published online for your particular browser type and version.
However, please note that if you choose to refuse cookies you may not be able to use the full functionality of our Site.
The following pages have information on how to change your cookies settings for the different browsers:
* [Cookie settings in Chrome](https://support.google.com/chrome/answer/95647?hl=en\&ref_topic=14666)
* [Cookie settings in Firefox](https://support.mozilla.org/en-US/kb/cookies-information-websites-store-on-your-computer?redirectlocale=en-US\&redirectslug=Cookies)
* [Cookie settings in Internet Explorer](https://support.mozilla.org/en-US/kb/cookies-information-websites-store-on-your-computer?redirectlocale=en-US\&redirectslug=Cookies)
* [Cookie settings in Safari and iOS](https://support.mozilla.org/en-US/kb/cookies-information-websites-store-on-your-computer?redirectlocale=en-US\&redirectslug=Cookies)
## Third Party Websites' Cookies [#third-party-websites-cookies]
When using our website, you may be directed to other websites for such activities as surveys, to make payment in currency other than U.S. dollars, or for job applications. These websites may use their own cookies. We do not have control over the placement of cookies by other websites you visit, even if you are directed to them from our website.
If you use the buttons that allow you to share products and content with your friends via social networks like Google, Twitter and Facebook, these companies may set a cookie on your computer memory. Find out more about these here:
* [https://www.facebook.com/about/privacy](https://support.google.com/chrome/answer/95647?hl=en\&ref_topic=14666)
* [http://twitter.com/privacy](https://support.mozilla.org/en-US/kb/cookies-information-websites-store-on-your-computer?redirectlocale=en-US\&redirectslug=Cookies)
* [http://www.google.com/intl/en-GB/policies/privacy](https://support.mozilla.org/en-US/kb/cookies-information-websites-store-on-your-computer?redirectlocale=en-US\&redirectslug=Cookies)
## Need More Information? [#need-more-information]
If you would like to find out more about cookies and their use on the Internet, you may find the following link useful: [All About Cookies](http://www.allaboutcookies.org/)
## Cookies that have been set in the past [#cookies-that-have-been-set-in-the-past]
If you have disabled one or more Cookies, we may still use information collected from cookies prior to your disabled preference being set, however, we will stop using the disabled cookie to collect any further information.
## Contact us [#contact-us]
If you have any questions or comments about this cookies policy, or privacy matters generally, please contact us via email at [privacy@firebolt.io](mailto:privacy@firebolt.io).
# Firebolt Copyright policy (/legal/copyright)
It is the policy of FIREBOLT ANALYTICS INC. ("**Firebolt**") to respect the legitimate rights of copyright owners. Pursuant to the Digital Millennium Copyright Act, 17 U.S.C. Section 512 (the "**DMCA**"), Firebolt has designated an agent (specified below) to receive notifications of claimed copyright infringement on its sites. Please be advised that we enforce a policy that provides for the termination in appropriate circumstances of subscribers who are repeat infringers.
If you believe that your work has been copied in a way that constitutes copyright infringement, please provide Firebolt's "copyright agent" (identified below) with the following information in accordance with the [DMCA](https://dmca.copyright.gov/):
1. An electronic or physical signature of the person authorized to act on behalf of the owner of the copyright;
2. A description of the copyrighted work that you claim has been infringed;
3. A description of where the material that you claim is infringing is located on the Site, with enough detail that We may find it on our Site; providing URLs in the body of an email is the best way to help us locate content quickly;
4. Your address, telephone number, and email address;
5. A statement by you that you have a good faith belief that the disputed use is not authorized by the copyright owner, its agent, or the law; and
6. A statement by you, made under penalty of perjury, that the above information in your notice is accurate and that you are the copyright owner or authorized to act on the copyright owner's behalf.
Firebolt's agent for notice of claims of copyright infringement can be reached as follows:
**Copyright Claims**\
Meitar Law Offices\
16 Abba Hillel Rd.\
Ramat Gan, 5250608\
Israel\
Phone: 0097236103974\
Email: [copyright@meitar.com](mailto:copyright@meitar.com)
Please also note that under Section 512(f) any person who knowingly materially misrepresents that material or activity is infringing may be subject to liability.
## Counter-Notification [#counter-notification]
If you believe that the material you posted was removed by mistake, and that you have the right to post the material, you may elect to send us a counter notice. To be effective the counter-notification must be a written communication provided to our designated agent that includes substantially the following (please consult your legal counsel or see 17 U.S.C. Section 512(g)(3) to confirm these requirements):
1. A physical or electronic signature of the subscriber.
2. Identification of the material that has been removed or to which access has been disabled and the location at which the material appeared before it was removed or access to it was disabled. **Providing URLs in the body of an email is the best way to help us locate content quickly**.
3. A statement under penalty of perjury that the subscriber has a good faith belief that the material was removed or disabled as a result of mistake or misidentification of the material to be removed or disabled.
4. The subscriber's name, address, and telephone number, and a statement that the subscriber consents to the jurisdiction of Federal District Court for the judicial district in which the address is located, or if the subscriber's address is outside of the United States, for any judicial district in which the service provider may be found, and that the subscriber will accept service of process from the person who provided notification of infringement or an agent of such person.
Such written notice should be sent to our designated agent as follows:
**Copyright Claims**\
Meitar Law Offices\
16 Abba Hillel Rd.\
Ramat Gan, 5250608\
Israel\
Phone: 0097236103974\
Email: [copyright@meitar.com](mailto:copyright@meitar.com)
Please note that under Section 512(f) of the Copyright Act, any person who knowingly materially misrepresents that material or activity was removed or disabled by mistake or misidentification may be subject to liability.
# Master Subscription Agreement (/legal/master-subscription-agreement)
THIS MASTER SUBSCRIPTION AGREEMENT (THIS "**AGREEMENT**") IS ENTERED INTO BETWEEN FIREBOLT ANALYTICS INC. (OR IF YOU ARE LOCATED OUTSIDE THE UNITED STATES – FIREBOLT ANALYTICS IRELAND LTD.) ("**PROVIDER**") AND THE CUSTOMER LISTED ON THE ORDER TERMS ("**ORDER TERMS**") WHEREBY CUSTOMER PURCHASES A SUBSCRIPTION TO THE PROVIDER CLOUD-BASED DATA WAREHOUSE SERVICE WHICH ENABLES THE STORING, PROCESSING AND ANALYZING OF BUSINESS DATA RECEIVED FROM MULTIPLE SOURCES (THE "**SERVICE**").
CUSTOMER ACCEPTS AND AGREES TO BE BOUND BY THIS AGREEMENT BY ACKNOWLEDGING SUCH ACCEPTANCE DURING THE REGISTRATION PROCESS AND ALSO BY CONTINUING TO USE THE SERVICE. IF THE PERSON ENTERING INTO THIS AGREEMENT IS DOING SO ON BEHALF OF A COMPANY OR OTHER LEGAL ENTITY, SUCH PERSON REPRESENTS THAT HE/SHE HAS THE AUTHORITY TO BIND SUCH ENTITY TO THIS AGREEMENT.
The "**Effective Date**" of this Agreement is the date which is the earlier of Customer's initial access to the Service or the effective date of the first Order Terms or Reseller Order Terms (defined below), as applicable, referencing this Agreement. This Agreement will govern Customer's initial purchase on the Effective Date as well as any future purchases made by the Customer that reference this Agreement.
The parties hereby agree as follows:
1. **LICENSES**
* **Access Rights.** Provider hereby grants Customer, during the Term (defined below), a limited, non-transferable and non-exclusive license for Customer's employees and third-party consultants ("**Authorized Users**") to use and access the Services in accordance with the use parameters described in the Order Terms and Documentation, solely for Customer's internal business purposes consistent with the terms and conditions of this Agreement. "**Documentation**" shall mean the reference, administrative and user manuals, made available by Provider to Customer with the Service. Documentation shall not include marketing materials.
* **Administration.** Provider will issue to one Authorized User ("**Administrator**") an individual logon identifier and password ("**Administrator's Logon**") for purposes of administering the Services. Using the Administrator's Logon, the Administrator shall assign each remaining Authorized User a unique logon identifier and password and assign and manage the business rules that control each such Authorized User's access to the Services.
* **Customer Data.** "**Customer Data**" means any data or data files of any type that are uploaded and stored by or on behalf of Customer in a data repository that is within an account owned and controlled by Provider with a cloud service provider. Customer hereby grants Provider a worldwide, limited-term license to process the Customer Data via the Service in accordance with instructions provided by Customer. Subject to the limited license granted herein, Provider acquires no right, title or interest from Customer or Customer's licensors under this Agreement in or to the Customer Data.
* **Feedback.** Customer grants to Provider and its affiliates a worldwide, perpetual, irrevocable, royalty-free, transferrable license to use and incorporate into the Service any suggestion, enhancement request, recommendation, correction, or other feedback provided by Customer or Authorized Users relating to the operation of the Service.
* **Restrictions.** Customer and its Authorized Users shall be prohibited from and will not: (a) sell, lease, license or sublicense the Service, or include the Service in a service bureau or outsourcing offering; (b) modify, change, alter, translate, create derivative works from, reverse engineer, disassemble or decompile the Service or any software included in the Service; (c) provide, disclose, divulge or make available to, or permit use of the Service by, any third party (except as expressly provided for herein); (d) copy or reproduce all or any part of the Service (except as expressly provided for herein); (e) knowingly interfere, or attempt to interfere, with the Service in any way; (f) use the Service to engage in spamming, mailbombing, spoofing or any other fraudulent, illegal or unauthorized use of the Service; (g) knowingly introduce into or transmit through the Service any virus, worm, trap door, back door; (h) interfere with or disrupt the integrity or performance of the Service or third-party data contained therein, (i) remove, obscure or alter any copyright notice, trademarks or other proprietary rights notices affixed to or contained within the Service; (j) attempt to gain unauthorized access to the Service or its related systems or networks, or permit direct or indirect access to or use of the Service in a way that circumvents a contractual usage limit, or access the Services in order to build a competitive product or service; or (k) host, provide, or develop software to intercept, emulate or redirect the Service in any way, or create, use or maintain any unauthorized connections to the Service. In addition, Customer may not access the Services for purposes of monitoring their availability, performance or functionality.
* **User-Defined Function (UDF) Feature.** Customer acknowledges and agrees that the Service includes a feature enabling Customer and its Authorized Users, at no additional charge, to execute any Python code, owned or licensed by Customer, within a sandboxed, tenant-isolated server environment, at no additional charge, to execute any Python code, owned or licensed by Customer (the "**UDF Service**"). Customer shall bear sole and exclusive responsibility for ensuring that its utilization of any proprietary or third-party software or Customer Data within the UDF Service is in full compliance with all applicable third-party licenses, laws and regulations, including privacy laws and regulations. Without limiting the generality of the foregoing sentence, Provider shall have no liability or responsibility for Customer's and/or its Authorized Users' infringement of third-party intellectual property (including privacy) rights, misappropriation or misuse of third-party intellectual property, or violation of any applicable laws or regulations arising out of or in connection with Customer's use of the UDF Service.
Provider reserves the right to monitor Customer's usage of the UDF Service to verify compliance with the terms of this Agreement. Customer acknowledges and agrees that it is Provider's policy to respect the legitimate rights of copyright and other intellectual property owners, and that Provider will respond to clear notices of alleged copyright infringement in accordance with Provider's Copyright Policy, which may be viewed at [Copyright Policy](https://www.firebolt.io/copyright-policy).
2. **RESPONSIBILITIES**
* **Provision of Service.** Provider will (a) make the Service available to Customer pursuant to this Agreement and the Order Terms, (b) provide Provider standard Customer Support, as described on Schedule A, attached hereto) for the Service to Customer at no additional charge, and/or upgraded support (if made available by Provider and purchased by Customer), and (c) use commercially reasonable efforts to make the Service available 24 hours a day, 7 days a week, except for: (i) planned downtime (of which Provider shall give at least 8 hours electronic notice and which Provider shall schedule to the extent practicable during the weekend hours between 6:00 pm Friday and 8:00 am Monday UTC time), and (ii) any unavailability caused by circumstances beyond Provider's reasonable control, including, for example, an act of god, act of government, flood, general health crisis, fire, earthquake, civil unrest, act of terror, strike or other labor problem (other than one involving Provider's employees), Internet or cloud service provider failure or delay, non-Provider application, or denial of service attack (each a "**Force Majeure Event**").
* **Protection of Customer Data.** Provider will maintain administrative, physical, and technical safeguards for protection of the security, confidentiality, and integrity of the Service, as described in Provider's security documentation. Those safeguards will include, but will not be limited to, measures for preventing access, use, modification, or disclosure of Customer Data by Provider personnel except (a) to provide the Service and prevent or address service or technical problems, (b) as compelled by law, or (c) as Customer expressly permits in writing. However, Customer acknowledges that Customer and the applicable cloud service provider are responsible for the security of the Customer Data as such data resides in Customer's account with the applicable cloud service provider.
* **Professional Services.** Customer may order from Provider professional services that are beyond the scope of the Service, such as configuration, customization and data entry services, pursuant to the terms set forth in the Order Terms ("**Professional Services**"). No Service license purchases are contingent on any Professional Services. Customer shall have a license right to use any deliverables (including any documentation, code, software, training materials or other work product) delivered as part of the Professional Services ("**Deliverables**") solely in connection with Customer's licensed use of the software, subject to all the same terms and conditions as apply to customer's software license (including in Section 1.5 (Restrictions), the Provider Professional Services Agreement attached hereto as Schedule B, and subject to any additional terms and conditions provided with the Deliverables. Customer may order Professional Services under Order Terms or a mutually executed Statement of Work ("**Professional Services SOW**") describing the work to be performed, fees and any applicable milestones, dependencies and other technical specifications or related information. Customer will reimburse Provider for reasonable travel and lodging expenses as incurred. Professional Services shall be charged in accordance with the applicable Professional Services SOW. Each Professional Services SOW is hereby deemed incorporated into this Agreement by reference. To the extent of any conflict between the main body of this Agreement and a Professional Services SOW, the former shall prevail, unless and to the extent that the Professional Services SOW expressly states otherwise. The Professional Services will be performed by Provider and/or its affiliates. Provider may subcontract Professional Services (in whole or in part) to a third-party contractor, and Provider shall remain primarily responsible for such contractor's performance of the Professional Services.
* **Customer Responsibilities.** Customer will (a) be responsible for Authorized Users' compliance with this Agreement, (b) be responsible for the accuracy, quality and legality of Customer Data and the means by which Customer acquired the Customer Data, (c) use commercially reasonable efforts to prevent unauthorized access to or use of the Service, and notify Provider promptly of any such unauthorized access or use, (d) use the Service only in accordance with the Documentation and applicable laws and government regulations, and (e) obtain and maintain all equipment and components necessary for Customer's use of the Service.
3. **FEES; PAYMENT TERMS**
* **Fees.** In consideration of the rights to the Service granted in this Agreement, Customer shall pay the fees specified in the Order Terms. Customer will provide Provider with valid and updated credit card information, or with a valid purchase order or alternative document reasonably acceptable to Provider. If Customer provides credit card information to Provider, Customer authorizes Provider to charge such credit card for the Service as listed in the Order Terms for the initial subscription term and any renewal subscription term(s). Additionally, payments may be made via third-party platform or gateway providers. Certain such charges shall be made in advance, either annually or in accordance with any different billing frequency stated in the applicable Order Terms, and other usage fees will be charged in arrears. If the Order Terms specify that payment will be by a method other than a credit card, Provider will invoice Customer in advance and otherwise in accordance with the relevant Order Terms. Unless otherwise stated in the Order Terms, invoiced charges (as opposed to credit card charges) are due net 30 days from the invoice date. Customer is responsible for providing complete and accurate billing and contact information to Provider and notifying Provider of any changes to such information. Late payments will incur interest in an amount equal to the lesser of 1.5% per month or the maximum allowable under applicable law. Payment obligations are non-cancelable, and fees paid are non-refundable. All payments shall be in U.S. dollars. Customer shall reimburse Provider for any costs of collection, including reasonable attorneys' fee, incurred collecting from Customer overdue fees. The fees due for any pay-as-you-go model shall be based on the then current price list of the Company found at [https://www.firebolt.io/pricing](https://www.firebolt.io/pricing) as updated therein from time to time.
* **Taxes.** All fees quoted or specified on the Order Terms do not include, and Customer will pay or reimburse Provider for, any applicable sales tax, use tax, and value added taxes (VAT) or other taxes which are levied or imposed by reason of the performance by Provider under this Agreement, excluding income taxes. If Customer is a tax-exempt organization and is not obligated to pay taxes arising out of this Agreement, Customer will provide Provider with any required documentation to verify its tax-exempt status with the applicable taxing authorities. If Customer is required by any authority to withhold taxes, then, after such withholding, the amount invoiced shall be deemed increased to the original amount invoiced prior to the withholding.
* **Future Functionality.** Customer agrees that Customer purchase of the Service is not contingent on the delivery of any future functionality or features, or dependent on any oral or written public comments made by Provider regarding future functionality or features.
4. **LIMITED WARRANTIES**
* **Customer Warranty.** Customer represents, warrants, and covenants to Provider that: (a) it has the authority to enter into this Agreement and perform its obligations hereunder; and (b) it and its Authorized Users will only use the Service for lawful purposes and will not use the Services to violate any law of any country or the intellectual property rights of any third party.
* **Provider Warranty.** Provider warrants that: (a) Provider has the authority to enter into this Agreement and (b) the Service will substantially operate and conform to the Documentation.
* **Disclaimer.** EXCEPT AS SET FORTH IN SECTION 4.2, PROVIDER MAKES NO REPRESENTATIONS OR WARRANTIES, WHETHER EXPRESS OR IMPLIED REGARDING OR RELATING TO ANY PORTION OF THE SERVICE, THE UDF SERVICE, OR ANY OTHER MATTER COVERED BY THIS AGREEMENT. PROVIDER SPECIFICALLY DISCLAIMS ANY AND ALL IMPLIED WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NON-INFRINGEMENT. PROVIDER DOES NOT GUARANTEE THAT CUSTOMER'S ACCESS TO THE SERVICE, INCLUDING THE UDF SERVICE, WILL BE UNINTERRUPTED OR ERROR FREE OR THAT ALL ERRORS WILL BE CORRECTED. PROVIDER SHALL NOT BE LIABLE FOR DELAYS, INTERRUPTIONS, SERVICE FAILURES OR OTHER PROBLEMS INHERENT IN USE OF THE INTERNET AND ELECTRONIC COMMUNICATIONS OR FOR ISSUES RELATED TO THIRD-PARTY HOSTING PROVIDERS WITH WHOM CUSTOMER SEPARATELY CONTRACTS. PROVIDER DOES NOT MAKE ANY WARRANTIES AND SHALL HAVE NO OBLIGATIONS WITH RESPECT TO THIRD PARTY APPLICATIONS. CUSTOMER MAY HAVE OTHER STATUTORY RIGHTS, BUT THE DURATION OF STATUTORILY REQUIRED WARRANTIES, IF ANY, SHALL BE LIMITED TO THE SHORTEST PERIOD PERMITTED BY LAW.
5. **LIMITATION OF LIABILITY.** IN NO EVENT WILL PROVIDER OR ITS PARTNERS BE LIABLE FOR ANY LOSS OF PROFITS, LOSS OF USE, BUSINESS INTERRUPTION, LOSS OF OR DAMAGE TO ANY CONTENT OR DATA, COST OF COVER OR INDIRECT, SPECIAL, INCIDENTAL OR CONSEQUENTIAL DAMAGES OF ANY KIND, WHETHER ALLEGED AS A BREACH OF CONTRACT, TORT OR OTHER FORM OF ACTION, EVEN IF SUCH PARTY HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES. PROVIDER'S LIABILITY UNDER THIS AGREEMENT FOR ANY DIRECT DAMAGES OF ANY KIND WILL NOT EXCEED AN AMOUNT EQUAL TO THE FEES PAID BY CUSTOMER TO PROVIDER UNDER THIS AGREEMENT DURING THE 6 MONTHS PRECEDING THE DATE ON WHICH A CLAIM FIRST ACCRUES.
6. **CONFIDENTIAL INFORMATION; DATA PROTECTION**
* **Confidentiality.** "**Confidential Information**" shall mean all information that is identified as confidential at the time of disclosure by the disclosing party or should be reasonably known by the receiving party to be confidential or proprietary due to the nature of the information disclosed and the circumstances surrounding the disclosure. All Customer Data will be deemed Confidential Information of Customer without any marking or further designation. All Provider technology and the terms and conditions of this Agreement (including pricing) will be deemed Confidential Information of Provider without any marking or further designation. Confidential Information shall not include information that the receiving party can demonstrate: (i) was rightfully in its possession or known to it prior to receipt of the Confidential Information; (ii) is or has become public knowledge through no fault of the receiving party; (iii) is rightfully obtained by the receiving party from a third party without breach of any confidentiality obligation; or (iv) is independently developed by the receiving party without reliance on the disclosing party's Confidential Information. Each party will take all reasonable precautions necessary to safeguard the confidentiality of the other party's Confidential Information including, at a minimum, those precautions taken by a party to protect its own Confidential Information, which will in no event be less than a reasonable degree of care. Confidential Information may only be disclosed on a need-to-know basis to the receiving party's employees and financial and legal advisors that are bound by confidentiality restrictions no less protective that as provided under this Agreement. The receiving party may only use the disclosing party's Confidential Information for the purpose of the performance of this Agreement. Confidential Information will not include information that is required to be disclosed by order of a court or other governmental entity; provided no less than ten days' notice is given to the disclosing party so that such party may obtain a protective order or other equitable relief, subject to compliance with applicable law by the receiving party. Additionally, the Privacy Policy posted on the Provider website shall apply to information obtained by Provider during the performance of the Service. The receiving party acknowledges that disclosure or improper use of Confidential Information would cause substantial harm for which damages alone would not be a sufficient remedy, and therefore that upon any such disclosure or improper use by the receiving party, the disclosing party will be entitled to seek appropriate equitable relief in addition to whatever other remedies it might have at law.
* **HIPAA Data.** Customer agrees not to upload to the Service any HIPAA Data unless Customer has entered into BAA with Provider. Upon mutual execution of the BAA, the BAA is incorporated by reference into this Agreement and is subject to its terms. "**BAA**" means a business associate agreement governing the parties' respective obligations with respect to any HIPAA Data uploaded by Customer to the Service in accordance with the terms of this Agreement. "**HIPAA**" means the Health Insurance Portability and Accountability Act, as amended, and supplemented. "**HIPAA Data**" means any patient, medical or other protected health information regulated by HIPAA or any similar federal or state laws, rules or regulations.
* **Data Privacy.** Customer hereby warrants and represents that it will (i) provide all appropriate notices, (ii) obtain all required informed consents and/or have any and all ongoing legal bases, and (iii) comply at all times with any and all applicable privacy and data protection laws and regulations (including, without limitation, the EU General Data Protection Regulation ("**GDPR**")), for allowing Provider to use and process the data in accordance with this Agreement (including, without limitation, the provision of such data to Provider (or access thereto) and the transfer of such data by Provider to its affiliates, subsidiaries and subcontractors, including transfers outside of the European Economic Area), for the provision of the Service and the performance of this Agreement.
The parties shall comply with the DPA, which is incorporated herein by this reference and except as expressly stated therein, shall not be modified except by mutual written agreement of the parties. "**DPA**" means the Data Processing Addendum located at [https://www.firebolt.io/DPA](https://www.firebolt.io/DPA) on the Effective Date of this Agreement.
In the event Customer fails to comply with any data protection or privacy law or regulation, the GDPR and/or any provision of the DPA, and/or fails to return an executed version of the DPA to Provider, then: (a) to the maximum extent permitted by law, Customer shall be solely and fully responsible and liable for any such breach, violation, infringement and/or processing of personal data without a DPA by Provider and Provider's affiliates and subsidiaries (including, without limitation, their employees, officers, directors, subcontractors and agents); and (b) in the event of any claim of any kind related to any such breach, violation or infringement and/or any claim related to processing of personal data without a DPA, Customer shall defend, hold harmless and indemnify Provider and Provider's affiliates and subsidiaries (including, without limitation, their employees, officers, directors, subcontractors and agents) from and against any and all losses, penalties, fines, damages, liabilities, settlements, costs and expenses, including reasonable attorneys' fees.
* **Service Data.** Notwithstanding anything to the contrary in this Agreement, Provider may collect and use Service Data to develop, improve, support, and operate its products and services. Provider may not share any Service Data that includes Customer's Confidential Information with a third party except (i) in accordance with the confidentiality provisions of this Agreement, or (ii) to the extent the Service Data is aggregated and anonymized such that Customer and Customer's users cannot be identified. "**Service Data**" means query logs, and any data (other than Customer Data) relating to the operation, support and/or about Customer's use of the Service.
7. **PROPRIETARY RIGHTS.** Except for the license granted in Section 1, no right title or interest of intellectual property or other proprietary rights in and to the Service made available under this Agreement is transferred to Customer hereunder. Provider and its third-party licensors retain all right, title and interests, including, without limitation, all copyright, trademark, patent, and other proprietary rights in and to the Service and all, modifications, enhancements and derivatives thereof. Customer will retain all right, title and interest to the Customer Data and documents created by Customer using the Services.
8. **MUTUAL INDEMNIFICATIONS**
* **Indemnification by Provider.** Provider will defend Customer against any claim, demand, suit or proceeding made or brought against Customer by a third party alleging that the use of the Service in accordance with this Agreement infringes or misappropriates such third party's intellectual property rights (a "**Claim Against Customer**"), and will indemnify Customer from any damages, reasonable attorney fees and costs finally awarded against Customer as a result of, or for amounts paid by Customer under a court-approved settlement of, a Claim Against Customer, provided that Customer (a) promptly gives Provider written notice of the Claim Against Customer, (b) gives Provider sole control of the defense and settlement of the Claim Against Customer (except that Provider may not settle any Claim Against Customer unless it unconditionally releases Customer of all liability), and (c) gives Provider all reasonable assistance, at Provider's expense. If Provider receives information about an infringement or misappropriation claim related to the Service, Provider may in its discretion and at no cost to Customer (i) modify the Service so that it no longer infringes or misappropriates, without breaching the warranties under Section 4.2, (ii) obtain a license for Customer's continued use of the Service in accordance with this Agreement, or (iii) terminate Customer's right to use the Service upon 30 days' written notice and refund Customer any prepaid, unused fees covering the remainder of the Term of the terminated use. The above defense and indemnification obligations do not apply to the extent a Claim Against Customer arises from Customer's breach of this Agreement.
* **Indemnification by Customer.** Customer will defend Provider against any claim, demand, suit or proceeding made or brought against Provider by a third party (a) alleging that Customer Data, or Customer use of any Service in breach of this Agreement, infringes or misappropriates such third party's intellectual property rights or violates applicable law; or (b) arising out of Customer's use of the UDF Service, as described in Section 1.6 (collectively, a "**Claim Against Provider**"), and will indemnify Provider from any damages, reasonable attorney fees and costs finally awarded against Provider as a result of, or for any amounts paid by Provider under a court-approved settlement of, a Claim Against Provider, provided that Provider (i) promptly gives Customer written notice of the Claim Against Provider, (ii) gives Customer sole control of the defense and settlement of the Claim Against Provider (except that Customer may not settle any Claim Against Provider unless it unconditionally releases Provider of all liability), and (iii) gives Customer all reasonable assistance, at Customer's expense.
* **Exclusive Remedy.** This Section 8 states the indemnifying party's sole liability to, and the indemnified party's exclusive remedy against, the other party for any type of claim described in this Section 8.
9. **TERM AND TERMINATION**
* **Term.** The initial term of this Agreement shall be the term specified on the Order Terms. After expiration of the initial term specified on the Order Terms the Customer's subscription to the Services shall automatically renew for successive one-year periods (the initial term and each renewal term, a "**Term**") unless either party provides written notice of non-renewal at least 30 days prior to commencement of the applicable renewal term. To the extent Customer enters into a proof of concept ("**POC**") and/or any evaluation period with Provider, and Customer does not provide a written request to enter a paid subscription, the Agreement, including any Service, will terminate immediately upon termination and/or expiration of the POC/evaluation, without any obligation and/or liability on Provider.
* **Termination by Provider.** Provider shall have the right, upon notice to Customer, to suspend the Service and/or terminate this Agreement and/or the relevant Order Form if: (a) Customer fails to pay Provider any amount due hereunder and such failure to pay is not cured within 30 days following Provider's notice to Customer of such breach; (b) Customer materially breaches any term or condition of this Agreement, provided such breach is not cured by Customer within 30 days following Provider's notice to Customer of such breach; or (c) Customer (i) terminates or suspends its business activities; (ii) liquidates all or a substantial portion of its assets for the benefit of creditors, or becomes subject to direct control of a trustee, receiver or similar authority to effect such liquidation of assets; (iii) becomes subject to any bankruptcy or insolvency proceeding under federal or state statutes to effect such liquidation of assets, or (iv) has stopped using the EC2 usage for at least 30 days.
* **Termination by Customer.** Customer will have the right, upon notice to Provider, to terminate this Agreement if Provider is in material breach of this Agreement and Provider fails to remedy such material breach within 30 days of its receipt of such notice or Provider (i) terminates or suspends its business activities; (ii) liquidates all or a substantial portion of its assets for the benefit of creditors, or becomes subject to direct control of a trustee, receiver or similar authority to effect such liquidation of assets; or (iii) becomes subject to any bankruptcy or insolvency proceeding under federal or state statutes to effect such liquidation of assets.
* **Refund or Payment upon Termination.** If this Agreement is terminated by Customer in accordance with Section 9.3, Provider will refund Customer any prepaid, unused fees. If this Agreement is terminated by Provider in accordance with Section 9.2, Customer will pay any unpaid fees covering the Service. In no event will termination relieve Customer of Customer's obligation to pay any fees payable to Provider for the period prior to the effective date of termination.
* **Data Extraction.** Subject to Section 9.1, upon any termination and for a period of 30 days thereafter, Customer may request in writing (email shall suffice) and Provider shall provide Customer with account access so that Customer may download a copy of the data that have been uploaded or otherwise saved to the database provided as part of the Service subscription purchased by Customer under this Agreement. After 30 days from termination of the Agreement, Provider may delete all data/files. In addition, and without derogating from the DPA, should you seek to have us retain your data/files for a longer period following termination of this Agreement, you may request that we do so, prior to the termination or expiration of this Agreement, or within 30 days thereafter, and subject to your payment of an addition fee to be paid by you in advance, which fee shall be set by us in our sole discretion.
* **Survival.** Any provisions necessary to interpret the respective rights and obligations of the parties hereunder shall survive any termination or expiration of this Agreement, regardless of the cause of such termination or expiration.
10. **GOVERNING LAW; VENUE.** For US customers purchasing from Firebolt Analytics, Inc.: This Agreement will be governed by the laws of the State of California, excluding its rules regarding conflicts of law. Venue for any dispute hereunder shall be a court of competent jurisdiction located in San Francisco County, California, and the parties irrevocably submit to the exclusive jurisdiction of such courts. For NON-US customers purchasing from Firebolt Analytics Ireland Ltd.: This Agreement will be governed by the laws of England, excluding its rules regarding conflicts of law. Venue for any dispute hereunder shall be a court of competent jurisdiction located in London, England, and the parties irrevocably submit to the exclusive jurisdiction of such courts.
11. **FEDERAL GOVERNMENT END USER PROVISIONS.** Provider will provide the Service, including related software and technology, for ultimate federal government end use solely in accordance with the following: Government technical data and software rights related to the Service include only those rights customarily provided to the public as defined in this Agreement. This customary commercial license is provided in accordance with FAR 12.211 (Technical Data) and FAR 12.212 (Software) and, for Department of Defense transactions, DFAR 252.227-7015 (Technical Data – Commercial Items) and DFAR 227.7202-3 (Rights in Commercial Computer Software or Computer Software Documentation). If a government agency has a need for rights not granted under these terms, it must negotiate with Provider to determine if there are acceptable terms for granting those rights, and a mutually acceptable written addendum specifically granting those rights must be included in any applicable agreement.
12. **EXPORT COMPLIANCE.** The Service and other technology Provider makes available, and derivatives thereof may be subject to export laws and regulations of the United States and other jurisdictions. Each party represents that it is not named on any U.S. government denied-party list. Customer shall not permit Authorized Users to access or use the Service in a U.S.-embargoed country (currently Cuba, Iran, North Korea, Sudan or Syria) or in violation of any U.S. export law or regulation.
13. **ASSIGNMENT.** Neither party may assign any of its rights or obligations hereunder, whether by operation of law or otherwise, without the other party's prior written consent (not to be unreasonably withheld); provided, however, that Provider may assign this Agreement in its entirety, without Customer's consent to its affiliates or in connection with a merger, acquisition, corporate reorganization, or sale of all or substantially all of its equity or assets.
14. **RESELLER ORDERS.** Customer may procure the Service directly from a reseller pursuant to a separate agreement that includes the Reseller Order Form and other commercial terms (a "**Reseller Arrangement**"). Provider will be under no obligation to provide the Service to Customer under a Reseller Arrangement if it has not received a Reseller Order Form for Customer. A reseller is not authorized to make any changes to this Agreement or otherwise authorized to make any warranties, representations, promises or commitments on behalf of Provider or in any way concerning the Service. If Customer procured the Service through a Reseller Arrangement, then Customer agrees that Provider may share certain Service Data with reseller related to Customer consumption of the Service. "**Reseller Order Form**" means a duly signed ordering document between a reseller and Customer that references this Agreement and the Service being provided by Provider pursuant to this Agreement and pricing and payment terms determined by the reseller. For clarity, Provider is not a party to the Reseller Order Form.
15. **AGREEMENT CHANGES.** From time to time, Provider may modify this Agreement. Unless otherwise specified by Provider, changes become effective for Customer after the updated version of this Agreement goes into effect. Provider will use reasonable efforts to notify Customer of any material adverse changes through communications via Customer's account, email or other means.
16. **USE OF CUSTOMER NAME.** Customer agrees that Provider may use Customer's name and may disclose that Customer is a Customer of Provider products or services in advertising, press, promotion and similar public disclosures; provided, however, that such advertising, promotion or similar public disclosures shall not indicate that Customer in any way endorses any Provider products without prior written permission from Customer.
17. **GENERAL PROVISIONS.** Provider and Customer are independent contractors. Any notice required or permitted to be delivered pursuant to this Agreement shall be in writing. Excluding payment obligations, neither party shall have any liability to the other or to third parties for any failure or delay in performing any obligation under this Agreement due to a Force Majeure Event (defined in Section 2.1 hereof). The failure of either party to enforce, or the delay by either party in enforcing, any of its rights under this Agreement will not be deemed to be a waiver or modification by such party of any of its rights under this Agreement. If any provision of this Agreement is held to be unenforceable, in whole or in part, such holding will not affect the validity of the other provisions of this Agreement. This Agreement may be executed in counterparts, all of which shall be considered one and the same agreement. The headings used herein are for reference and convenience only and shall not enter into the interpretation hereof. No purchase order or any hand-written or typewritten text on a purchase order which purports to modify or supplement the printed text of this Agreement shall add to or vary the terms of this Agreement. All such proposed variations or additions (whether submitted by Provider or Customer) are objected to and shall have no force or effect. Nothing in this Agreement affects any statutory rights of consumers that cannot be waived or limited by contract. This Agreement will not create any right or cause of action for any third-party beneficiary or any other third party. This Agreement contains the entire agreement of the parties with respect to the subject matter of this Agreement and supersedes all previous communications, representations, understandings and agreements, either oral or written, between the parties with respect to said subject matter.
## Schedule A [#schedule-a]
Capitalized terms not otherwise defined in this Schedule shall have the meaning ascribed to them in the Agreement.
### Customer Support Terms [#customer-support-terms]
These Support Terms set forth the terms, conditions, and procedures under which maintenance and support ("**Support**") is offered for the Service during the Term of the Customer's subscription.
Support will consist of: (i) email and portal-based support; (ii) correction of errors to keep the Service in conformance with the Documentation; and (iii) updated versions of the Service provided by Provider to its general customer base of subscribers at no additional charge. Support will not include: (i) set-up, installation, or configuration of hardware and software required for Customer to access the Service; or (ii) consultation, error correction, or research with respect to Customer-created documents and information.
### Problem Resolution [#problem-resolution]
The severity level of the problems reported by Customer shall be reasonably determined by Provider. Provider will resolve each reported error or issue with the Service by using commercially reasonable efforts to provide: (i) a patch or fix as necessary; or (ii) a reasonable workaround for the error or issue; or, if either (i) or (ii) are not reasonably practicable, a specific action plan regarding how Provider intends to address the reported error or issue and an estimate on how long it may take to correct or workaround the error or issue. Customer agrees to use commercially reasonable efforts to assist and provide information to Provider as required to resolve errors or issues with the Services reported by Customer. If a permanent repair cannot be made, a temporary resolution (bypass and recovery) will be implemented to the extent possible.
Support covers any issue or problem that is the result of a verifiable, replicable error (Customer will use all reasonable means to verify and replicate) in the Service ("**Verifiable Provider Issue**"). An error will be a Verifiable Provider Issue if it constitutes a material failure by the Service to function in accordance with the Documentation included in the Service. If Technical Support reasonably determines that Customer's problem is not caused by Provider or its systems, equipment, or software, Provider is not obligated to provide support under this Agreement. Nevertheless, Provider will, if possible, offer suggestions as to how Customer can remedy the problem. If Provider determines that the issue was not the result of a Verifiable Provider Issue, Provider may offer to provide for out-of-scope professional services at Provider's then current rates upon its standard terms to address the issue.
### Additional Support [#additional-support]
Technical Support may also determine that Customer's request is a request for "**Additional Support**." Additional Support is any assistance not covered above. Examples of Additional Support include substantive questions regarding data or results, requests for Service customization, specialized training regarding use of the Service, custom documentation, and consulting. If Provider believes that it can appropriately and effectively provide the requested services, it will offer do so at its then-current rates upon its standard terms.
### Customer Responsibilities [#customer-responsibilities]
Customer's representative shall initiate all requests for Support. The representative must be trained, qualified, and authorized to communicate all necessary information, perform diagnostic testing under the direction of the Provider service representative and be available during the performance of any Support if required.
### Service Level Agreement [#service-level-agreement]
This Service Level Agreement ("**SLA**") shall apply to Provider's to the Service during the Term of the Customer's subscription. Capitalized terms not otherwise defined herein shall have the meaning ascribed to them in the Agreement.
**1. Availability.**
a. Formula. The Service will, subject to the exceptions listed below, be available 99.9% of the time during each calendar month from the time that the Services go-live in Customer's production environment (referred to herein as the "**Availability Commitment**"). The availability of the Service for a given month will be calculated according to the following formula (referred to herein as the "**Availability**"):
Where: Total minutes in the month = TMM\
Total minutes in the month the Service is unavailable = TMU\
And: ((TMM-TMU) X 100)/TMM
b. For purposes of this calculation, the Service will be deemed to be unavailable (referred to herein as "**Unavailable**") only (i) if the Service does not respond to HTTP requests issued by Provider's monitoring software, or (ii) for the duration of a Severity-1 Error. A "**Severity-1 Error**" shall mean that the Service suffers an error or issue in a production down situation which cannot be reasonably circumvented, and which so substantially impairs the performance of the Service or any components of the Service, which are critical to the Customer's business, as to effectively render them unusable. Further, the Service will not be deemed Unavailable for any downtime or outages excluded from such calculation by reason of the exceptions set forth in Section 2 of this SLA. Provider's records and data will be the basis for all SLA calculations and determinations.
c. Maintenance performed at Customer's request outside of the normally scheduled maintenance will not be considered an outage.
**2. Exceptions**
a. The Service will not be considered Unavailable for any outage that results from any planned maintenance performed by Provider during Provider's standard maintenance windows which occur between 6:00 pm Friday and 8:00 am Monday UTC time (referred to herein as "**Scheduled Maintenance**").
b. The Service will not be considered Unavailable for any outage unavailability of the Service due to (a) Customer's information content or application programming, acts or omissions of Customer or its agents; (b) a Force Majeure Event or other delays or failures due to circumstances beyond Provider's reasonable control that could not be avoided by its exercise of due care; or (c) failures of Internet backbone itself and the network by which Customer connects to the Internet backbone or any other network unavailability outside of the Provider network.
# Firebolt Privacy Policy (/legal/privacy)
In order to ensure transparency and give you more control over your Personal Data, this privacy policy ("**Privacy Policy**") governs how we, Firebolt Analytics Inc. (together, "**Firebolt**" "**we**", "**our**" or "**us**") use, collect and store personal data we collect or receive from or about you ("**you**") such as in the following use cases:
1. When you browse or visit our website, [https://www.firebolt.io/](https://www.firebolt.io/) ("Website");
2. When you browse or interact with our Website:
1. When you request a free trial or a product demo
2. When you subscribe to our distribution list(s) / newsletter(s) / blog(s)
3. When we process your job application
4. When you contact us (e.g. customer support, need help, submit a request, schedule a call)
5. When you react, participate to or comment on our forum, webinars and/or events
3. When you make use of, or interact with, our platform (the "**Platform**"):
1. When you create an account, log-in and make use of the Platform
2. When you react, participate to or comment on our forum
4. When you attend a marketing event and provide us with your personal data;
5. When we use the personal data of our customers (e.g. contact details)
6. When you interact with us on our social media profiles (e.g., Facebook, Instagram, Twitter, LinkedIn).
When you interact with us on our social media profiles (e.g., Facebook, Instagram, Twitter, LinkedIn). We greatly respect your privacy, which is why we make every effort to provide a platform that would live up to the highest of user privacy standards. Please read this Privacy Policy carefully, so you can fully understand our practices in relation to personal data. "**Personal data**" or "**personal information**" means any information that can be used, alone or together with other data, to uniquely identify any living human being. Important note: Nothing in this Privacy Policy is intended to limit in any way your statutory right, including your rights to a remedy or means of enforcement.
## Table of contents: [#table-of-contents]
1. What Information we Collect, Why we Collect it, and How it is Used
2. Period of Storage of Collected Personal Data
3. How we Protect your Personal Data
4. How we Share your Personal Data
5. Additional Information regarding Transfers of Personal Data
6. Your Privacy Rights; How to Delete your Account
7. Use by Children
8. Interaction with Third Party Products
9. Analytic tools
10. Contact Us
This Privacy Policy can be updated from time to time and, therefore, we ask you to check back periodically for the latest version of this Privacy Policy. If we implement significant changes to the use of your personal data in a manner different from that stated at the time of collection, we will notify you by posting a notice on our Website or by other means.
## 1. WHAT INFORMATION WE COLLECT, WHY WE COLLECT IT, AND HOW IT IS USED [#1-what-information-we-collect-why-we-collect-it-and-how-it-is-used]
### When you contact us (e.g. customer support, need help, submit a request) [#when-you-contact-us-eg-customer-support-need-help-submit-a-request]
| Specific personal data we collect | Why is the personal data collected and for what purposes? | Legal basis (GDPR only, if applicable) | Third parties with whom we share your personal data | Consequences of not providing the personal data |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Cookies, analytic tools and log files, including IP address, usage pattern and clickstream events. For more information, please read our [cookies policy](/firebolt-cookies-policy) | To operate, monitor and analyze the Website and provide you with certain features and functionalities on the Website To improve the user experience on our website For marketing purposes | We process professional identifiers to pursue our legitimate business interest in identifying and engaging with potential B2B customers. Consent: we rely on your explicit consent via our cookie banner before activating identity-resolution technologies. | For more information, please read our cookies policy | We will not be able to personalize the website experience to you. Certain website functionality may not be accessible Cannot provide customized marketing campaigns |
### When you request a free trial or a product demo or join a workshop or download whitepapers [#when-you-request-a-free-trial-or-a-product-demo-or-join-a-workshop-or-download-whitepapers]
| Specific personal data we collect | Why is the personal data collected and for what purposes? | Legal basis (GDPR only, if applicable) | Third parties with whom we share your personal data | Consequences of not providing the personal data |
| ------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Full name Email Address Title Company Name Business email address Name Phone number Country | To schedule a demo To open a trial account for you To allow you to try our service To allow you to participate in the workshop To communicate with you For marketing and lead generation purposes Managing interactions with potential customers | Processing is necessary for the performance of a contract to which the data subject is party or in order to take steps at the request of the data subject prior to entering into a contract; or Legitimate interest (e.g., to provide you with a demo) | Firebolt entities Salesforce.com RudderStack Onfire Calendly, LLC Cal.com Smartlead | We will not be able to schedule a demo We will not be able to open a trial account for you To allow you to try our service We will not be able to allow you to participate in the workshop We will not be able to communicate with you We will not be able to follow up with you for marketing and lead generation purposes. |
### When you subscribe to our distribution list(s) / newsletter(s) / blog(s) [#when-you-subscribe-to-our-distribution-lists--newsletters--blogs]
| Specific personal data we collect | Why is the personal data collected and for what purposes? | Legal basis (GDPR only, if applicable) | Third parties with whom we share your personal data | Consequences of not providing the personal data |
| --------------------------------- | ---------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| Email Address | To send you updates and marketing materials with interesting content, news about Firebolt and our services | Consent. Legitimate interest (to the extent that we have a relationship with you in connection with your interest in or use of our Website) | Zapier Onfire Smartlead | We will not be able to send you updates and marketing materials with interesting content, news about Firebolt and our services |
### When we process your job application [#when-we-process-your-job-application]
| Specific personal data we collect | Why is the personal data collected and for what purposes? | Legal basis (GDPR only, if applicable) | Third parties with whom we share your personal data | Consequences of not providing the personal data |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Full name Email Address Phone number LinkedIn Profile URL CV, cover letter and portfolio Details of your technical interview Any other information that you decide to provide us with | To review your application To assess your suitability for the role To schedule an interview To communicate with you (for candidacy-related matters) To pre-screen candidate's technical ability (e.g., via offline coding challenges) To provide a shared coding environment to conduct technical interviews | Processing is necessary to take steps at the request of the data subject before entering into a contract. Legitimate interest (e.g., in processing your application for the purpose of evaluating you for the applicable position) | Comeet CoderPad.io | We will not be able to review your job application. We will not be able to assess your suitability for the role. We will not be able to schedule an interview. We will not be able to communicate with you (for candidacy-related matters) |
### When you contact us (e.g. customer support, need help, submit a request) [#when-you-contact-us-eg-customer-support-need-help-submit-a-request-1]
| Specific personal data we collect | Why is the personal data collected and for what purposes? | Legal basis (GDPR only, if applicable) | Third parties with whom we share your personal data | Consequences of not providing the personal data |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| Full name Email Address Phone number Company name Communications with you (e.g., zoom meeting or phone calls) Any other information that you decide to provide us with | To respond to your inquiry To provide support (e.g., to solve problems, bugs or issues) and handle your complaints and/or feedback To customize your experience To record your meeting and calls (subject to your consent) To customize your experience Managing interactions with potential customers | Legitimate interest (e.g. respond to a query sent by you); or Processing is necessary for the performance of a contract to which the data subject is party or in order to take steps at the request of the data subject prior to entering into a contract Consent (e.g., when you provide consent to record the call/meeting) | Salesforce.com RudderStack Zapier Google (including Workspace and Gemini) Onfire Calendly, LLC Cal.com Smartlead | We will not be able to communicate with you and provide support We will not be able to customize your experience |
| Full name Email address Any other information you decide to provide us with | To react, participate to or comment on the forum To reply to your comments on the forum | Legitimate interest (e.g., to allow you to comment on the forum) | Discourse | We will not be able to react, participate to or comment on the forum We will not be able to reply to your comments |
### When you create an account, log-in and make use of the Platform [#when-you-create-an-account-log-in-and-make-use-of-the-platform]
| Specific personal data we collect | Why is the personal data collected and for what purposes? | Legal basis (GDPR only, if applicable) | Third parties with whom we share your personal data | Consequences of not providing the personal data |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Full name Email Address User ID and Password IP address Job position/title Any other information that you decide to provide/supply us with | To create your account and allow you to create the admin and regular users To allow you to log in/sign up to the Platform To authenticate you To fulfill your request for our services and related activities (e.g. account management, support) To perform/execute the relevant agreement To grant you access to the services (our Platform) To communicate with you for product-related matters To sign the contract with us | Processing in necessary for the performance of a contract to which the data subject is party or in the request of the data subject prior to entering into a contract Legitimate interest (e.g. to allow you to sign up to the Platform) | AWS Auth0 Docusign Slack Amberflo Pendo (only email address) Google (including Workspace and Gemini) Scarf | We will not be able to create your account and allow you to create the admin and regular users We will not be able to allow you to log in/sign up to the Platform We will not be able to authenticate you We will not be able to fulfill your requests for our services and related activities (e.g., account management, support) We will not be able to perform/execute the relevant agreement We could not grant you access to the services (our Platform) We will not be able to communicate with you for product-related matters |
| Metadata about your use of the Platform Usage pattern | To analyze your use of the Platform To improve our services | Processing is necessary for performance of contract to which the data subject is party or in order to take steps at the request of the data subject prior to entering into a contact. Legitimate interest (e.g., to improve our services) | OneTrust Pendo Google (including Workspace and Gemini) | We will not be able to analyze your use of the Platform We will not be able to improve our services |
| Email Address | To send you updates and marketing materials with interesting content, news about Firebolt and our services | Consent. Legitimate interest (e.g., to send you more information of similar or the same products and services that you are using) | Onfire Smartlead | We will not be able to send you updates and marketing materials with interesting content, news |
### When you attend a marketing event and provide us with your personal data [#when-you-attend-a-marketing-event-and-provide-us-with-your-personal-data]
| Specific personal data we collect | Why is the personal data collected and for what purposes? | Legal basis (GDPR only, if applicable) | Third parties with whom we share your personal data | Consequences of not providing the personal data |
| ------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| Full name Email Address Phone number Title Company name Any other information you decide to provide us with | To participate in a promotion To have a business communication | Depending on the case, Legitimate interest (e.g., in contacting you after you showed interest in our products and services), consent or processing is necessary in order to take steps at the request of the data subject prior to entering into a contract | Salesforce.com Clay Onfire Smartlead | We will not be able to include you in the promotion We will not be able to have business communication with you |
### When we use the personal data of our customers (e.g. contact details) [#when-we-use-the-personal-data-of-our-customers-eg-contact-details]
| Specific personal data we collect | Why is the personal data collected and for what purposes? | Legal basis (GDPR only, if applicable) | Third parties with whom we share your personal data | Consequences of not providing the personal data |
| ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
| Full name Company name Email address Phone number Payment information Any data available on Linkedin profiles Communications with you (e.g., zoom meeting or phone calls) Any other information you decide to provide us with | To provide our products and services To perform the applicable agreement To communicate with our customers (e.g., to send you contract-related communications) | Processing is necessary for the performance of a contract to which the data subject is party or in order to take steps at the request of the data subject prior to entering into a contract. Legitimate interest (e.g. to contact our service providers) Processing is necessary for compliance with a legal obligation to which the controller is subject. Legitimate interest (e.g. to contact our service providers). Consent (e.g., when you provide consent to record the call/meeting) | Salesforce.com Amberflo Slack RudderStack Onfire Smartlead | We could not contact our service providers/vendors. We could not perform the applicable agreement. We could not communicate with our customers |
### When we obtain Personal Data from third parties (e.g., lead generation companies) [#when-we-obtain-personal-data-from-third-parties-eg-lead-generation-companies]
| Specific personal data we collect | Why is the personal data collected and for what purposes? | Legal basis (GDPR only, if applicable) | Third parties with whom we share your personal data | Consequences of not providing the personal data |
| --------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
| Business contact details | To contact you for a business communication To create an AI-generate a personalized business connection and relationship | Depending on the case, legitimate interest (e.g., B2B marketing), consent or processing is necessary in order to take steps at the request of the data subject prior to entering into a contract | Salesforce.com Clay Industry Dive Smartlead Antword Inc (Spear) | We will not be able to contact you for a business communication We will not be able to generate a business connection and relationship |
### When you communicate with us on social media (e.g. LinkedIn, Twitter, Facebook, Instagram) [#when-you-communicate-with-us-on-social-media-eg-linkedin-twitter-facebook-instagram]
| Specific personal data we collect | Why is the personal data collected and for what purposes? | Legal basis (GDPR only, if applicable) | Third parties with whom we share your personal data | Consequences of not providing the personal data |
| ---------------------------------------------------------------------------- | -------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | --------------------------------------------------- | ----------------------------------------------- |
| Social media handle Any other information you decide to provide us with | To respond to your inquiry To initiate a business communication | Depending on the context, legitimate **interest** (e.g., in responding to your query), or consent | The various social media platforms | We will not be able to respond to your inquiry |
Finally, please note that some of the abovementioned personal data will be used for detecting, taking steps to prevent, and prosecuting fraud or other illegal activity, to identify and repair errors, to conduct audits, and for security purposes. Personal Data may also be used to comply with applicable laws, with investigations performed by the relevant authorities, for law enforcement purposes, and/or to exercise or defend legal claims. In certain cases, we may or will anonymize or de-identify your Personal Data and further use it for internal and external purposes, including, without limitation, to improve the services and for research purposes. "Anonymous Information" means information that does not enable identification of an individual user, such as aggregated information about the use of our services. We may use Anonymous Information and/or disclose it to third parties without restrictions (for example, in order to improve our services and enhance your experience with them).
## 2. PERIOD OF STORAGE OF COLLECTED PERSONAL DATA [#2-period-of-storage-of-collected-personal-data]
1. Retention Personal Data. Your Personal data will be stored until we: (i) no longer need the information and proactively delete it or (ii) you send a valid deletion request. Please note that we will retain it for a longer or shorter period in accordance with data retention laws.
2. Additional Information on Data Retention. Other circumstances in which we will retain your personal data for longer periods of time include: (i) where we are required to do so in accordance with legal, regulatory, tax or accounting requirements, or (ii) for us to have an accurate record of your dealings with us in the event of any complaints or challenges, or (iii) if we reasonably believe there is a prospect of litigation relating to your personal data or dealings. We have an internal data retention policy to ensure that we do not retain your personal data perpetually. Regarding retention of cookies, you can read more in our cookie policy (available in the link provided above).
## 3. HOW WE PROTECT YOUR PERSONAL DATA [#3-how-we-protect-your-personal-data]
We have implemented appropriate technical, organizational and security measures designed to protect your personal data. However, please note that we cannot guarantee that the information will not be compromised as a result of unauthorized penetration to our servers. As the security of information depends in part on the security of the computer, device or network you use to communicate with us and the security you use to protect your user IDs and passwords, please make sure to take appropriate measures to protect this information.
## 4. HOW WE SHARE YOUR PERSONAL DATA [#4-how-we-share-your-personal-data]
In addition to the recipients described above, we may share your personal data as follows:
1. With our business partners with whom we jointly offer products or services. We may also share Personal Data with our affiliated companies.
2. To the extent necessary, with regulators, courts or competent authorities, to comply with applicable laws, regulations and rules (including, without limitation, federal, state or local laws), and requests of law enforcement, regulatory and other governmental agencies or if required to do so by court order; We may use or disclose the information we collect in order to ensure that our users are complying with all applicable aspects of our policies.
3. We may disclose information with our lawyers, accountants, auditors and other professional advisors where necessary to obtain legal or other advice or otherwise protect and manage our business interests.
4. We may share hashed identifiers and browsing data with identity resolution partners to help us identify and contact business visitors
5. If, in the future, we sell or transfer, or we consider selling or transferring, some or all of our business, shares or assets to a third party, we will disclose your personal data to such third party (whether actual or potential) in connection with the foregoing events;
6. In the event that we are acquired by, or merged with, a third party entity, or in the event of bankruptcy or a comparable event, we reserve the right to transfer, disclose or assign your personal data in connection with the foregoing events; and/or
7. Where you have provided your consent to us sharing or transferring your personal data (e.g., where you provide us with marketing consents or opt-in to optional additional services or functionality).
If you want to receive the list of the current recipients of your Personal Data, please make your request by contacting us to [privacy@firebolt.io](mailto:privacy@firebolt.io).
## 5. ADDITIONAL INFORMATION REGARDING TRANSFERS OF PERSONAL DATA [#5-additional-information-regarding-transfers-of-personal-data]
1. Internal transfers: Transfers within the Firebolt group will be covered by an internal processing agreement entered into by members of the Firebolt group (an intra-group data processing agreement) which contractually obliges each member to ensure that Personal Data receives an adequate and consistent level of protection wherever it is transferred to.
2. External transfers: Where we transfer your Personal Data outside of EU/EEA (for example to third parties who provide us with services), we will generally obtain contractual commitments from them to protect your Personal Data.
3. Where we transfer your personal data outside of the EU/EEA, for example to third parties who help provide our products and services, we will obtain contractual commitments and or assurances from them to protect your Personal Data or rely on adequacy decisions issued by the European Commission.
## 6. YOUR PRIVACY RIGHTS; HOW TO DELETE YOUR ACCOUNT [#6-your-privacy-rights-how-to-delete-your-account]
1. Rights: The following rights (which may be subject to certain exemptions or derogations) shall apply to certain individuals (some of which only apply to individuals protected by the GDPR):
* You have a right to access personal data held about you. Your right of access may normally be exercised free of charge, however we reserve the right to charge an appropriate administrative fee where permitted by applicable law;
* You have the right to request that we rectify any personal data we hold that is inaccurate or misleading;
* You have the right to request the erasure/deletion of your personal data (e.g. from our records). Please note that there may be circumstances in which we are required to retain your personal data, for example for the establishment, exercise or defense of legal claims;
* You have the right to object, to or to request restriction, of the processing;
* You have the right to data portability. This means that you may have the right to receive your personal data in a structured, commonly used and machine-readable format, and that you have the right to transmit that data to another controller;
* You have the right to object to profiling;
* You have the right to withdraw your consent at any time. Please note that there may be circumstances in which we are entitled to continue processing your data, in particular if the processing is required to meet our legal and regulatory obligations. Also, please note that the withdrawal of consent shall not affect the lawfulness of processing based on consent before its withdrawal;
* You also have a right to request certain details of the basis on which your personal data is transferred outside the European Economic Area, but you acknowledge that data transfer agreements may need to be partially redacted for reasons of commercial confidentiality.
* You have a right to lodge a complaint with your local data protection supervisory authority (i.e., your place of habitual residence, place or work or place of alleged infringement) at any time or before the relevant institutions in your place of residence. We ask that you please attempt to resolve any issues with us before you contact your local supervisory authority and/or relevant institution.
You can exercise your rights by contacting us at [privacy@firebolt.io](mailto:privacy@firebolt.io). Subject to legal and other permissible considerations, we will make every reasonable effort to honor your request promptly in accordance with applicable law or inform you if we require further information in order to fulfil your request.
When processing your request, we may ask you for additional information to confirm or verify your identity and for security purposes, before processing and/or honoring your request. We reserve the right to charge a fee where permitted by law, for instance if your request is manifestly unfounded or excessive. In the event that your request would adversely affect the rights and freedoms of others (for example, would impact the duty of confidentiality we owe to others) or if we are legally entitled to deal with your request in a different way than initial requested, we will address your request to the maximum extent possible, all in accordance with applicable law.
2. Deleting your account: Should you ever decide to delete your account, you may do so by contacting Firebolt Support. If you terminate your account, any association between your account and personal data we store will no longer be accessible through your account. However, given the nature of sharing on the Services, any public or team activity on your Account prior to deletion will (to the extent permissible pursuant to data protection laws) remain stored on our servers and will remain accessible to the public or team. Please note that whilst you may be able to exercise your data protection rights under Section 6 in respect of such information, they may be subject to certain exemptions or derogations.
3. Marketing emails – opt-out: You may choose not to receive marketing email of this type by sending a single email with the subject "BLOCK" to [privacy@firebolt.io](mailto:privacy@firebolt.io). Please note that the email must come from the email account you wish to block OR if you receive an unwanted email from us, you can use the unsubscribe link found at the bottom of the email to opt out of receiving future emails, and we will process your request within a reasonable time after receipt.
## 7. USE BY CHILDREN [#7-use-by-children]
We do not offer our products or services for use by children and, therefore, we do not knowingly collect personal data from, and/or about children under the age of eighteen (18). If you are under the age of eighteen (18), do not provide any personal data to us without involvement of a parent or a guardian. In the event that we become aware that you provide personal data in violation of applicable privacy laws, we reserve the right to delete it. If you believe that we might have any such information, please contact us at [privacy@firebolt.io](mailto:privacy@firebolt.io).
## 8. INTERACTION WITH THIRD PARTY PRODUCTS [#8-interaction-with-third-party-products]
We enable you to interact with third party websites, mobile software applications and products or services that are not owned or controlled by us (each a "**Third Party Service**"). We are not responsible for the privacy practices or the content of such Third Party Services. Please be aware that Third Party Services can collect Personal Data from you. Accordingly, we encourage you to read the terms and conditions and privacy policies of each Third Party Service.
We enable you to interact with third party websites, mobile software applications and products or services that are not owned or controlled by us (each a "**Third Party Service**"). We are not responsible for the privacy practices or the content of such Third Party Services. Please be aware that Third Party Services can collect Personal Data from you. Accordingly, we encourage you to read the terms and conditions and privacy policies of each Third Party Service.
## 9. ANALYTIC TOOLS [#9-analytic-tools]
1. **Google Analytics**. The Website and/or platform uses a tool called "**Google Analytics**" to collect information about use of the Website. Google Analytics collects information such as how often users visit this Website, what pages they visit when they do so, and what other websites they used prior to coming to this Website. We use the information we get from Google Analytics to maintain and improve the Website and our products. We do not combine the information collected through the use of Google Analytics with personal information. Google's ability to use and share information collected by Google Analytics about your visits to this Website is restricted by the Google Analytics Terms of Service, available at [https://marketingplatform.google.com/about/analytics/terms/us/](https://marketingplatform.google.com/about/analytics/terms/us/), and the Google Privacy Policy, available at [http://www.google.com/policies/privacy/](http://www.google.com/policies/privacy/). You may learn more about how Google collects and processes data specifically in connection with Google Analytics at [http://www.google.com/policies/privacy/partners/](http://www.google.com/policies/privacy/partners/). You may prevent your data from being used by Google Analytics by downloading and installing the Google Analytics Opt-out Browser Add-on, available at [https://tools.google.com/dlpage/gaoptout/](https://tools.google.com/dlpage/gaoptout/).
2. **Microsoft Clarity.** We partner with Microsoft Clarity to capture how you use and interact with our website through behavioral metrics, heatmaps, and session replay to improve and market our products/services. Website usage data is captured using first and third-party cookies and other tracking technologies to determine the popularity of products/services and online activity. Additionally, we use this information for site optimization, fraud/security purposes, and advertising. For more information about how Microsoft collects and uses your data, visit the [Microsoft Privacy Statement](https://www.microsoft.com/he-il/privacy/privacystatement).
3. **Scheduling Tools.** We use Cal.com and Calendly to manage meetings. These tools may collect technical information (such as IP address and time zone) to ensure the scheduling interface functions correctly. You can view their privacy practices at Cal.com/privacy and Calendly.com/privacy.
4. **Do Not Track.** We do not currently respond to web browser "Do Not Track" signals and reserve the right to remove or add new analytic tools.
## 10. ADDITIONAL DOCUMENTS [#10-additional-documents]
1. Subscription agreement: [http://www.firebolt.io/subscription-agreement](https://d7umqicpi7263.cloudfront.net/eula/product/28a79f78-8008-4c02-a971-54ef6d8ed904/79c7d4d6-2f9e-4e72-92ff-543a371308a6.pdf)
2. DPA: [https://www.firebolt.io/dpa](https://www.firebolt.io/dpa)
## 11. CONTACT US [#11-contact-us]
If you have any questions, concerns or complaints regarding our compliance with this notice and the data protection laws, or if you wish to exercise your rights, we encourage you to first contact us at [privacy@firebolt.io](mailto:privacy@firebolt.io).
# Firebolt.io Terms of Use (/legal/terms)
Welcome to [www.firebolt.io](https://www.firebolt.io) (together with its subdomains, Content, Marks and services, the "**Site**"). Please read the following Terms of Use carefully before using this Site so that you are aware of your legal rights and obligations with respect to Firebolt Analytics Ltd. ("**Firebolt**", "**we**", "**our**" or "**us**"). By accessing or using the Site, you expressly acknowledge and agree that you are entering a legal agreement with us and have understood and agree to comply with, and be legally bound by, these Terms of Use, together with the Privacy Policy (collectively, the "**Terms**"). You hereby waive any applicable rights to require an original (non-electronic) signature or delivery or retention of non-electronic records, to the extent not prohibited under applicable law. If you do not agree to be bound by these Terms please do not access or use the Site.
1. **Background.** The Site is intended to provide visitors of the Site with information and Content about our products and services as well as general information and news related to the big data and analytics space.
2. **Modification.** We reserve the right, at our discretion, to change these Terms at any time. Such change will be effective ten (10) days following posting of the revised Terms on the Site, and your continued use of the Site thereafter constitutes your acceptance of those changes.
3. **Ability to Accept Terms.** The Site is only intended for individuals aged eighteen (18) years or older. If you are under eighteen (18) years please do not visit or use the Site.
4. **Site Access.** For such time as these Terms are in effect, we hereby grant you permission to visit and use the Site provided that you comply with these Terms and applicable law.
5. **Restrictions.** As a condition to your right to access and use the Site, you shall not (and shall not permit or encourage any third party to) do any of the following: (i) copy, distribute or modify any part of the Site without our prior written authorization; (ii) use, modify, create derivative works of, transfer (by sale, resale, license, sublicense, download or otherwise), reproduce, distribute, display or disclose Content (defined below), except as expressly authorized herein; (iii) disrupt servers or networks connected to the Site; (iv) use or launch any automated system (including without limitation, "robots" and "spiders") to access the Site; (v) circumvent, disable or otherwise interfere with security-related features of the Site or features that prevent or restrict use or copying of any Content or that enforce limitations on use of the Site; (vi) make a derivative work of the Site, or use the Site to develop any service or product that is the same as (or substantially similar to or competitive with) the Site; (vii) publish or transmit any robot, virus, malware, Trojan horse, spyware, or similar malicious item intended (or that has the potential) to damage or disrupt the Site; (viii) take any action that imposes or may impose (at our sole discretion) an unreasonable or disproportionately large load on the Site infrastructure, or otherwise interfere (or attempt to interfere) with the integrity or proper working of the Site; and/or (ix) use the Site to infringe, misappropriate or violate any third party's Intellectual Property Rights (as defined below), or any law.
6. **Account.** In order to use some of the services of the Site, you may have to create an account ("**Account**"). You agree not to create an Account for anyone else or use the account of another without their permission. When creating your Account, you must provide accurate and complete information. You are solely responsible for the activity that occurs in your Account, and you must keep your Account password secure. You must notify Firebolt immediately of any breach of security or unauthorized use of your Account. As between you and Firebolt, you are solely responsible and liable for the activity that occurs in connection with your Account. If you wish to delete your Account you may send an email request to Firebolt at [hello@firebolt.io](mailto:hello@firebolt.io)
7. **Payments to Firebolt.** Unless stated otherwise in the Terms, your general right to access and use the Site is free. However, Firebolt may in the future charge a fee for certain access or usage. You will not be charged for any such access or use of the Site unless you first agree to such charges, but please be aware that any failure to pay applicable charges may result in you not having access to some or all of the Site.
8. **Intellectual Property Rights.**
* Content and Marks. The (i) content on the Site, including without limitation, the text, documents, articles, brochures, descriptions, products, software, graphics, photos, sounds, videos, interactive features, and services (collectively, the "**Content**"), (ii) the trademarks, service marks and logos contained therein ("**Marks**"), are the property of Firebolt and/or its licensors and may be protected by applicable copyright or other intellectual property laws and treaties. "Firebolt", the Firebolt logo, and other marks are Marks of Firebolt or its affiliates. All other trademarks, service marks, and logos used on the Site are the trademarks, service marks, or logos of their respective owners. We reserve all rights not expressly granted in and to the Site and the Content.
* Use of Content. Content on the Site is provided to you for your information and personal use only and may not be used, modified, copied, distributed, transmitted, broadcast, displayed, sold, licensed, de-compiled, or otherwise exploited for any other purposes whatsoever without our prior written consent. If you download or print a copy of the Content you must retain all copyright and other proprietary notices contained therein. In any event you wish to use, publish, copy, distribute, transmit, broadcast, display or otherwise exploit such Content, please be in touch with us at [hello@firebolt.io](mailto:hello@firebolt.io) in order to receive our written consent.
9. **Information Description.** We attempt to be as accurate as possible. However, we cannot and do not warrant that the Content available on the Site is accurate, complete, reliable, current, or error-free. We reserve the right to make changes in or to the Content, or any part thereof, in our sole judgment, without the requirement of giving any notice prior to or after making such changes to the Content. Your use of the Content, or any part thereof, is made solely at your own risk and responsibility.
10. **Links.**
* The Site may contain links, and may enable you to post content, to third party websites that are not owned or controlled by Firebolt. We are not affiliated with, have no control over, and assume no responsibility for the content, privacy policies, or practices of, any third party websites. You: (i) are solely responsible and liable for your use of and linking to third party websites and any content that you may send or post to a third party website; and (ii) expressly release Firebolt from any and all liability arising from your use of any third party website. Accordingly, we encourage you to read the terms and conditions and privacy policy of each third party website that you may choose to visit.
* Firebolt permits you to link to the Site provided that: (i) you link to but do not replicate any page on this Site; (ii) the hyperlink text shall accurately describe the Content as it appears on the Site; (iii) you shall not misrepresent your relationship with Firebolt or present any false information about Firebolt and shall not imply in any way that we are endorsing any services or products, unless we have given you our express prior consent; (iv) you shall not link from a website ("**Third Party Website**") which prohibits linking to third parties; (v) such Third party Website does not contain content that: (a) is offensive or controversial (both at our discretion), or (b) infringes any intellectual property, privacy rights, or other rights of any person or entity; and/or (vi) you, and your website, comply with these Terms and applicable law.
11. **Privacy.** Our privacy policy which is available at [https://www.firebolt.io/privacy](https://www.firebolt.io/privacy). You agree that we may use personal information that you provide or make available to us in accordance with the Privacy Policy.
12. **Artificial Intelligence.** Firebolt utilizes and incorporates third-party artificial intelligence (AI) tools and services in various aspects of its operations. This includes, but is not limited to, the creation of content and the generation of podcasts that may be published on Firebolt's Site or other properties or channels. By using the Site, you acknowledge and agree that such third-party AI services may process data submitted by you for the purpose of providing enhanced features and services. Firebolt does not control these third-party AI tools and services and is not responsible for their operations. Your interaction with such third-party AI services is governed by their respective terms of use and privacy policies.
13. **Third Party Content.** The Site may present, or otherwise allow you to view, access, link to, and/or interact with, Content from third parties and other sources that are not owned or controlled by us (such Content, "**Third Party Content**"). The Site may also enable you to communicate with the related third parties. The display or communication to you of such Third Party Content does not (and shall not be construed to) in any way imply, suggest, or constitute any sponsorship, endorsement, or approval by us of such Third Party Content or third party, or by such third party of us, and nor any affiliation between us and such third party. We do not assume any responsibility or liability for Third Party Content, or any third party's terms of use, privacy policies, actions, omissions, or practices. Please read the terms of use and privacy policy of any third party that you interact with before you engage in any such activity.
14. **Warranty Disclaimers.**
* This section applies whether or not the services provided under the Site are for payment. Applicable law may not allow the exclusion of certain warranties, so to that extent certain exclusions set forth herein may not apply.
* THE SITE IS PROVIDED ON AN "AS IS" AND "AS AVAILABLE" BASIS, AND WITHOUT WARRANTIES OF ANY KIND EITHER EXPRESS OR IMPLIED. FIREBOLT HEREBY DISCLAIMS ALL WARRANTIES, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO IMPLIED WARRANTIES OF MERCHANTABILITY, TITLE, FITNESS FOR A PARTICULAR PURPOSE, NON-INFRINGEMENT, AND THOSE ARISING BY STATUTE OR FROM A COURSE OF DEALING OR USAGE OF TRADE. FIREBOLT DOES NOT GUARANTEE THAT THE SITE WILL BE FREE OF BUGS, SECURITY BREACHES, OR VIRUS ATTACKS. THE SITE MAY OCCASIONALLY BE UNAVAILABLE FOR ROUTINE MAINTENANCE, UPGRADING, OR OTHER REASONS. YOU AGREE THAT FIREBOLT WILL NOT BE HELD RESPONSIBLE FOR ANY CONSEQUENCES TO YOU OR ANY THIRD PARTY THAT MAY RESULT FROM TECHNICAL PROBLEMS OF THE INTERNET, SLOW CONNECTIONS, TRAFFIC CONGESTION OR OVERLOAD OF OUR OR OTHER SERVERS. WE DO NOT WARRANT, ENDORSE OR GUARANTEE ANY CONTENT, PRODUCT, OR SERVICE THAT IS FEATURED OR ADVERTISED ON THE SITE BY A THIRD PARTY.
* EXCEPT AS EXPRESSLY STATED IN OUR PRIVACY POLICY, FIREBOLT DOES NOT MAKE ANY REPRESENTATIONS, WARRANTIES OR CONDITIONS OF ANY KIND, EXPRESS OR IMPLIED, AS TO THE SECURITY OF ANY INFORMATION YOU MAY PROVIDE OR ACTIVITIES YOU ENGAGE IN DURING THE COURSE OF YOUR USE OF THE SITE.
15. **Limitation of Liability.**
* TO THE FULLEST EXTENT PERMISSIBLE BY LAW, FIREBOLT SHALL NOT BE LIABLE FOR ANY DIRECT, INDIRECT, EXEMPLARY, SPECIAL, CONSEQUENTIAL, OR INCIDENTAL DAMAGES OF ANY KIND, OR FOR ANY LOSS OF DATA, REVENUE, PROFITS OR REPUTATION, ARISING UNDER THESE TERMS OR OUT OF YOUR USE OF, OR INABILITY TO USE, THE SITE, EVEN IF FIREBOLT HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES OR LOSSES.
* IN NO EVENT SHALL THE AGGREGATE LIABILITY OF FIREBOLT FOR ANY DAMAGES ARISING UNDER THESE TERMS OR OUT OF YOUR USE OF, OR INABILITY TO USE THE SITE EXCEED: (A) THE TOTAL AMOUNT OF FEES, IF ANY, PAID BY YOU TO FIREBOLT FOR USING THE SITE DURING THE THREE (3) MONTHS PRIOR TO BRINGING THE CLAIM, OR (B) TEN US DOLLARS ($10), WHICHEVER IS HIGHER.
* Some jurisdictions do not allow the exclusion or limitation of incidental or consequential damages, or of other damages, and to the extent applicable to you, such exclusions and limitations shall not apply. Furthermore, nothing in this Agreement shall be deemed to exclude or limit liability for death or personal injury resulting from negligence, or for fraud or fraudulent misrepresentation
16. **Indemnity.** You agree to defend, indemnify and hold harmless Firebolt and our affiliates, and our respective officers, directors, employees and agents, from and against any and all claims, damages, obligations, losses, liabilities, costs and expenses (including but not limited to attorney's fees) arising from: (i) your use of, or inability to use, the Site; or (ii) your violation of these Terms.
17. **Term and Termination.** These Terms are effective until terminated by Firebolt or you. Firebolt, in its sole discretion, has the right to terminate these Terms and/or your access to the Site, or any part thereof, immediately at any time and with or without cause (including, without any limitation, for a breach of these Terms). Firebolt shall not be liable to you or any third party for termination of the Site, or any part thereof. If you object to any term or condition of these Terms, or any subsequent modifications thereto, or become dissatisfied with the Site in any way, your only recourse is to immediately discontinue use of the Site. Upon termination of these Terms, you shall cease all use of the Site. This Section 15 and Sections 8 (Intellectual Property Rights), 11 (Privacy), 13 (Warranty Disclaimers), 14 (Limitation of Liability), 15 (Indemnity), and 17 (Independent Contractors) to 20 (General) shall survive termination of these Terms.
18. **Independent Contractors.** You and Firebolt are independent contractors. Nothing in these Terms creates a partnership, joint venture, agency, or employment relationship between you and Firebolt. You must not under any circumstances make, or undertake, any warranties, representations, commitments or obligations on behalf of Firebolt.
19. **Assignment.** These Terms, and any rights and licenses granted hereunder, may not be transferred or assigned by you but may be assigned by Firebolt without restriction or notification to you. Any prohibited assignment shall be null and void.
20. **Governing Law.** Firebolt reserves the right to discontinue or modify any aspect of the Site at any time. These Terms and the relationship between you and Firebolt shall be governed by and construed in accordance with the laws of the State of Israel, without regard to its principles of conflict of laws. You agree to submit to the personal and exclusive jurisdiction of the courts located in Tel-Aviv and waive any jurisdictional, venue, or inconvenient forum objections to such courts, provided that Firebolt may seek injunctive relief in any court of competent jurisdiction
21. **General.** These Terms shall constitute the entire agreement between you and Firebolt concerning the Site. If any provision of these Terms is deemed invalid by a court of competent jurisdiction, the invalidity of such provision shall not affect the validity of the remaining provisions of these Terms, which shall remain in full force and effect. Except as may be expressly stated otherwise in this Agreement, no right or remedy conferred upon or reserved by any party under this Agreement is intended to be, or shall be deemed, exclusive of any other right or remedy under this Agreement, at law or in equity, but shall be cumulative of such other rights and remedies. No waiver of any term of these Terms shall be deemed a further or continuing waiver of such term or any other term, and a party's failure to assert any right or provision under these Terms shall not constitute a waiver of such right or provision. YOU AGREE THAT ANY CAUSE OF ACTION THAT YOU MAY HAVE ARISING OUT OF OR RELATED TO THE SITE MUST COMMENCE WITHIN ONE (1) YEAR AFTER THE CAUSE OF ACTION ACCRUES. OTHERWISE, SUCH CAUSE OF ACTION IS PERMANENTLY BARRED.
# Trial Subscription Agreement (/legal/trial-subscription-agreement)
THIS TRIAL SUBSCRIPTION AGREEMENT (THIS "**AGREEMENT**") IS ENTERED INTO BETWEEN FIREBOLT ANALYTICS INC. (OR IF YOU ARE LOCATED OUTSIDE THE UNITED STATES - FIREBOLT ANALYTICS IRELAND LTD.) ("**PROVIDER**") AND THE CUSTOMER ("**CUSTOMER**", "**YOU**") WHEREBY CUSTOMER GAINS ACCESS TO A SUBSCRIPTION TO THE PROVIDERS CLOUD-BASED DATA WAREHOUSE SERVICE WHICH ENABLES THE STORING, PROCESSING AND ANALYZING OF BUSINESS DATA RECEIVED FROM MULTIPLE SOURCES (THE "**SERVICE**").
CUSTOMER ACCEPTS AND AGREES TO BE BOUND BY THIS AGREEMENT BY ACKNOWLEDGING SUCH ACCEPTANCE DURING THE REGISTRATION PROCESS AND ALSO BY CONTINUING TO USE THE SERVICE. IF THE PERSON ENTERING INTO THIS AGREEMENT IS DOING SO ON BEHALF OF A COMPANY OR OTHER LEGAL ENTITY, SUCH PERSON REPRESENTS THAT HE/SHE HAS THE AUTHORITY TO BIND SUCH ENTITY TO THIS AGREEMENT.
The "**Effective Date**" of this Agreement is the date which is the earlier of Customer's initial access to the Service. This Agreement will govern Customer's initial purchase on the Effective Date as well as any future purchases made by the Customer that reference this Agreement.
The parties hereby agree as follows:
1. **LICENSES**
* **Access Rights.** Provider hereby grants Customer, during the Term (defined below), a limited, non-transferable and non-exclusive license for a single user who is Customer's employee or third-party consultant ("**Authorized User**") to use the Service in accordance with the use parameters described in this Agreement and the Documentation, solely for Customer's internal business purposes consistent with the terms and conditions of this Agreement. "**Documentation**" shall mean the reference, administrative and user manuals, made available by Provider to Customer with the Service. Documentation shall not include marketing materials.
* **Credits.** The Provider shall grant the Customer free of charge, the equivalent of $200 USD (or such larger amount as decided by Provider at its own discretion) in Service usage credits which may only be used to receive access to the Service for the Term (as defined below) of this Agreement ("**Credits**"). Credits expire at the end of the Term, meaning that any Credits which Customer does not use during the Term will not roll over into future subscriptions to the Service. Credits have no cash value or any other value outside of the Service and are not redeemable for cash. For the avoidance of doubt, Credits do not operate or serve as stored value facilities in any way. Customer may not transfer, trade, gift or otherwise exchange Credits.
* **Customer Data.** "**Customer Data**" means any data or data files of any type that are uploaded and stored by or on behalf of Customer in a data repository that is within an account owned and controlled by Provider with a cloud service provider. Customer hereby grants Provider a worldwide, limited-term license to process the Customer Data via the Service in accordance with instructions provided by Customer. Subject to the limited license granted herein, Provider acquires no right, title or interest from Customer or Customer's licensors under this Agreement in or to the Customer Data. Provider reserves the right to delete Customer Data if Provider, at its own discretion, sees no activity in the user account.
* **Feedback.** Customer grants to Provider and its affiliates a worldwide, perpetual, irrevocable, royalty-free, transferrable license to use and incorporate into the Service any suggestion, enhancement request, recommendation, correction or other feedback provided by Customer or Authorized User relating to the operation of the Service.
* **User-Defined Function (UDF) Service.** Customer acknowledges and agrees that the Service includes a feature enabling Customer and its Authorized Users, at no additional charge, to execute any Python code, owned or licensed by Customer, within a sandboxed, tenant-isolated server environment, at no additional charge, to execute any Python code, owned or licensed by Customer (the "**UDF Service**"). Customer shall bear sole and exclusive responsibility for ensuring that its utilization of any proprietary or third-party software or Customer Data within the UDF Service is in full compliance with all applicable third-party licenses, laws and regulations, including privacy laws and regulations. Without limiting the generality of the foregoing sentence, Provider shall have no liability or responsibility for Customer's and/or its Authorized Users' infringement of third-party intellectual property (including privacy) rights, misappropriation or misuse of third-party intellectual property, or violation of any applicable laws or regulations in connection with Customer's use of the UDF Service.
Provider reserves the right to monitor Customer's usage of the UDF Service to verify compliance with the terms of this Agreement. Customer acknowledges and agrees that it is Provider's policy to respect the legitimate rights of copyright and other intellectual property owners, and that Provider will respond to clear notices of alleged copyright infringement in accordance with Provider's Copyright Policy, which may be viewed at: [Copyright Policy](https://www.firebolt.io/copyright-policy).
2. **RESTRICTIONS.** Customer and its Authorized User shall be prohibited from and will not: (a) sell, lease, license or sublicense the Service, or include the Service in a service bureau or outsourcing offering; (b) modify, change, alter, translate, create derivative works from, reverse engineer, disassemble or decompile the Service or any software included in the Service; (c) provide, disclose, divulge or make available to, or permit use of the Service by, any third party (except as expressly provided for herein); (d) copy or reproduce all or any part of the Service (except as expressly provided for herein); (e) knowingly interfere, or attempt to interfere, with the Service in any way; (f) use the Service to engage in spamming, mailbombing, spoofing or any other fraudulent, illegal or unauthorized use of the Service; (g) knowingly introduce into or transmit through the Service any virus, worm, trap door, back door; (h) interfere with or disrupt the integrity or performance of the Service or third-party data contained therein, (i) remove, obscure or alter any copyright notice, trademarks or other proprietary rights notices affixed to or contained within the Service; (j) attempt to gain unauthorized access to the Service or its related systems or networks, or permit direct or indirect access to or use of the Service in a way that circumvents a contractual usage limit, or access the Service in order to build a competitive product or service; or (k) host, provide, or develop software to intercept, emulate or redirect the Service in any way, or create, use or maintain any unauthorized connections to the Service.
3. **RESPONSIBILITIES.** Customer will (a) be responsible for Authorized User' compliance with this Agreement, (b) be responsible for the accuracy, quality and legality of Customer Data and the means by which Customer acquired the Customer Data, (c) prevent unauthorized access to or use of the Service, and notify Provider promptly of any such unauthorized access or use, and (d) use the Service only in accordance with the Documentation and applicable laws and government regulations.
4. **LIMITED WARRANTIES.** Customer represents, warrants and covenants to Provider that: (a) it has the authority to enter into this Agreement and perform its obligations hereunder; and (b) it and its Authorized User will only use the Service for lawful purposes and will not use the Service to violate any law of any country or the intellectual property rights of any third party. Provider warrants that it has the authority to enter into this Agreement. THE SERVICE, INCLUDING THE UDF SERVICE (DEFINED BELOW), IS PROVIDED AND MADE AVAILABLE HEREUNDER ON AN "AS IS" AND "AS AVAILABLE" BASIS, PROVIDER MAKES NO REPRESENTATIONS OR WARRANTIES, WHETHER EXPRESS OR IMPLIED REGARDING OR RELATING TO ANY PORTION OF THE SERVICE OR ANY OTHER MATTER COVERED BY THIS AGREEMENT. PROVIDER SPECIFICALLY DISCLAIMS ANY AND ALL IMPLIED WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NON-INFRINGEMENT. PROVIDER DOES NOT GUARANTEE THAT CUSTOMER'S ACCESS TO THE SERVICE, INCLUDING THE UDF SERVICE, WILL BE UNINTERRUPTED OR ERROR FREE OR THAT ALL ERRORS WILL BE CORRECTED. PROVIDER SHALL NOT BE LIABLE FOR DELAYS, INTERRUPTIONS, SERVICE FAILURES OR OTHER PROBLEMS INHERENT IN USE OF THE INTERNET AND ELECTRONIC COMMUNICATIONS OR FOR ISSUES RELATED TO THIRD-PARTY HOSTING PROVIDERS WITH WHOM CUSTOMER SEPARATELY CONTRACTS. PROVIDER DOES NOT MAKE ANY WARRANTIES AND SHALL HAVE NO OBLIGATIONS WITH RESPECT TO THIRD PARTY APPLICATIONS. CUSTOMER MAY HAVE OTHER STATUTORY RIGHTS, BUT THE DURATION OF STATUTORILY REQUIRED WARRANTIES, IF ANY, SHALL BE LIMITED TO THE SHORTEST PERIOD PERMITTED BY LAW.
5. **LIMITATION OF LIABILITY.** IN NO EVENT WILL PROVIDER OR ITS PARTNERS BE LIABLE FOR ANY LOSS OF PROFITS, LOSS OF USE, BUSINESS INTERRUPTION, LOSS OF OR DAMAGE TO ANY CONTENT OR DATA, COST OF COVER OR INDIRECT, SPECIAL, INCIDENTAL OR CONSEQUENTIAL DAMAGES OF ANY KIND, WHETHER ALLEGED AS A BREACH OF CONTRACT, TORT OR OTHER FORM OF ACTION, EVEN IF SUCH PARTY HAS BEEN ADVISED OF THE POSSIBILITY OF SUCH DAMAGES. PROVIDER'S AGGREGATE LIABILITY UNDER THIS AGREEMENT FOR ANY DIRECT DAMAGES OF ANY KIND WILL NOT EXCEED TEN U.S. DOLLARS (US$ 10). THE FOREGOING EXCLUSIONS AND LIMITATION SHALL APPLY: (A) EVEN IF PROVIDER OR ONE OF ITS AFFILIATES HAS BEEN ADVISED, OR SHOULD HAVE BEEN AWARE, OF THE POSSIBILITY OF LOSSES, DAMAGES, OR COSTS; (B) EVEN IF ANY REMEDY IN THIS AGREEMENT FAILS OF ITS ESSENTIAL PURPOSE; AND (C) REGARDLESS OF THE THEORY OR BASIS OF LIABILITY (INCLUDING WITHOUT LIMITATION BREACH OF CONTRACT, TORT, NEGLIGENCE AND STRICT LIABILITY).
6. **CONFIDENTIAL INFORMATION; DATA PROTECTION**
* **Confidentiality.** Each party (the "**Recipient**") may have access to certain non-public or proprietary information and materials of the other party (the "**Discloser**"), whether in tangible or intangible form ("**Confidential Information**"). Confidential Information shall not include information and material which: (a) at the time of disclosure by Discloser to Recipient hereunder, is in the public domain; (b) after disclosure by Discloser to Recipient hereunder, becomes part of the public domain through no fault of the Recipient; (c) was rightfully in the Recipient's possession at the time of disclosure by the Discloser hereunder, and which is not subject to prior continuing obligations of confidentiality; (d) is rightfully disclosed to the Recipient by a third party having the lawful right to do so; or (e) independently developed by the Recipient without use of, or reliance upon, Confidential Information received from the Discloser. The Recipient shall not disclose or make available the Discloser's Confidential Information to any third party (including without limitation by way of publishing), except to its employees, advisers, agents and investors, subject to substantially similar written confidentiality undertakings). Recipient shall take commercially reasonable measures, at a level at least as protective as those taken to protect its own Confidential Information of like nature (but in no event less than a reasonable level), to protect the Discloser's Confidential Information within its possession or control, from disclosure to a third party. The Recipient shall use the Discloser's Confidential Information solely for the purposes expressly permitted under this Agreement. In the event that Recipient is required to disclose Confidential Information of the Discloser pursuant to any Law, regulation, or governmental or judicial order, the Recipient will (a) promptly notify Discloser in writing of such Law, regulation or order, (b) reasonably cooperate with Discloser in opposing such disclosure, (c) only disclose to the extent required by such law, regulation or order (as the case may be). Upon termination of this Agreement, or otherwise upon written request by the Discloser, the Recipient shall promptly return to Discloser its Confidential Information (or if embodied electronically, permanently erase it), and certify compliance writing.
* **HIPAA Data.** Customer agrees not to upload to the Service any HIPAA Data unless Customer has entered into BAA with Provider. Upon mutual execution of the BAA, the BAA is incorporated by reference into this Agreement and is subject to its terms. "**BAA**" means a business associate agreement governing the parties' respective obligations with respect to any HIPAA Data uploaded by Customer to the Service in accordance with the terms of this Agreement. "**HIPAA**" means the Health Insurance Portability and Accountability Act, as amended and supplemented. "**HIPAA Data**" means any patient, medical or other protected health information regulated by HIPAA or any similar federal or state laws, rules or regulations.
* **Data Privacy.** Customer hereby warrants and represents that it will (i) provide all appropriate notices, (ii) obtain all required informed consents and/or have any and all ongoing legal bases, and (iii) comply at all times with any and all applicable privacy and data protection laws and regulations (including, without limitation, the EU General Data Protection Regulation ("**GDPR**")), for allowing Provider to use and process the data in accordance with this Agreement (including, without limitation, the provision of such data to Provider (or access thereto) and the transfer of such data by Provider to its affiliates, subsidiaries and subcontractors, including transfers outside of the European Economic Area), for the provision of the Service and the performance of this Agreement.
The parties shall comply with the DPA, which is incorporated herein by this reference and except as expressly stated therein, shall not be modified except by mutual written agreement of the parties. "**DPA**" means the Data Processing Addendum located at [https://www.firebolt.io/DPA](https://www.firebolt.io/DPA) on the Effective Date of this Agreement.
In the event Customer fails to comply with any data protection or privacy law or regulation, the GDPR and/or any provision of the DPA, and/or fails to return an executed version of the DPA to Provider, then: (a) to the maximum extent permitted by law, Customer shall be solely and fully responsible and liable for any such breach, violation, infringement and/or processing of personal data without a DPA by Provider and Provider's affiliates and subsidiaries (including, without limitation, their employees, officers, directors, subcontractors and agents); and (b) in the event of any claim of any kind related to any such breach, violation or infringement and/or any claim related to processing of personal data without a DPA, Customer shall defend, hold harmless and indemnify Provider and Provider's affiliates and subsidiaries (including, without limitation, their employees, officers, directors, subcontractors and agents) from and against any and all losses, penalties, fines, damages, liabilities, settlements, costs and expenses, including reasonable attorneys' fees.
Personal information for which Provider is considered a 'data owner' or 'controller' is subject to Provider's privacy policy available at [https://www.firebolt.io/firebolt-privacy-policy](https://www.firebolt.io/firebolt-privacy-policy).
* **Service Data.** Notwithstanding anything to the contrary in this Agreement, Provider may collect and use Service Data to develop, improve, support, and operate its products and services. Provider may not share any Service Data that includes Customer's Confidential Information with a third party except (i) in accordance with the confidentiality provisions of this Agreement, or (ii) to the extent the Service Data is aggregated and anonymized such that Customer and Customer's users cannot be identified. "Service Data" means query logs, and any data (other than Customer Data) relating to the operation, support and/or about Customer's use of the Service.
7. **PROPRIETARY RIGHTS.** Except for the license granted in Section 1, no right title or interest of intellectual property or other proprietary rights in and to the Service made available under this Agreement is transferred to Customer hereunder. Provider and its third party licensors retain all right, title and interests, including, without limitation, all copyright, trademark, patent, and other proprietary rights in and to the Service and all, modifications, enhancements and derivatives thereof.
8. **INDEMNIFICATION.** Customer will defend Provider against any claim, demand, suit or proceeding made or brought against Provider by a third party (a) alleging that Customer Data or Customer's use of any part of the Service in breach of this Agreement, infringes or misappropriates such third party's intellectual property rights or violates applicable law; or (b) arising out of Customer's use of the UDF Service, as described in Section 1.5 (a "Claim Against Provider"), and will indemnify Provider from any damages, reasonable attorney fees and costs finally awarded against Provider as a result of, or for any amounts paid by Provider under a court-approved settlement of, a Claim Against Provider, provided Provider (i) promptly gives Customer written notice of the Claim Against Provider, (ii) gives Customer sole control of the defense and settlement of the Claim Against Provider (except that Customer may not settle any Claim Against Provider unless it unconditionally releases Provider of all liability), and (iii) gives Customer all reasonable assistance, at Customer's expense.
9. **TERM AND TERMINATION.** This Agreement commences on the Effective Date and will remain in full force and effect for a period of 3 months unless earlier terminated as provided herein (the "Term"). Provider may terminate this Agreement at any time, for any reason or no reason, upon at least five (5) days prior written notice (email acceptable) to Customer. Notwithstanding the foregoing, Provider may terminate this Agreement immediately upon written notice (email acceptable) to Customer if Customer breaches any provision of this Agreement. Upon termination of this Agreement, Customer shall immediately cease all access to and use of the Service. Any provision in this Agreement that is stated (or by its nature ought) to survive termination, shall survive, including, without limitation, Sections 1.5, 2, 4, 5, 6, 7, 8, 9 and 10. Any provisions necessary to interpret the respective rights and obligations of the parties hereunder shall survive any termination or expiration of this Agreement, regardless of the cause of such termination or expiration.
10. **GOVERNING LAW; VENUE.** For US customers purchasing from Firebolt Analytics, Inc.: This Agreement will be governed by the laws of the State of California, excluding its rules regarding conflicts of law. Venue for any dispute hereunder shall be a court of competent jurisdiction located in San Francisco County, California, and the parties irrevocably submit to the exclusive jurisdiction of such courts. EACH PARTY IRREVOCABLY WAIVES ITS RIGHT TO TRIAL OF ANY ISSUE BY JURY. EXCEPT TO PROTECT OR ENFORCE A PARTY'S INTELLECTUAL PROPERTY RIGHTS OR CONFIDENTIALITY OBLIGATIONS, NO ACTION, REGARDLESS OF FORM, UNDER THIS AGREEMENT MAY BE BROUGHT BY EITHER PARTY MORE THAN ONE (1) YEAR AFTER TERMINATION OF THE AGREEMENT. For NON-US customers purchasing from Firebolt Analytics Ireland Ltd.: This Agreement will be governed by the laws of England, excluding its rules regarding conflicts of law. Venue for any dispute hereunder shall be a court of competent jurisdiction located in London, England, and the parties irrevocably submit to the exclusive jurisdiction of such courts.
11. **FEDERAL GOVERNMENT END USER PROVISIONS.** Provider will provide the Service, including related software and technology, for ultimate federal government end use solely in accordance with the following: Government technical data and software rights related to the Service include only those rights customarily provided to the public as defined in this Agreement. This customary commercial license is provided in accordance with FAR 12.211 (Technical Data) and FAR 12.212 (Software) and, for Department of Defense transactions, DFAR 252.227-7015 (Technical Data - Commercial Items) and DFAR 227.7202-3 (Rights in Commercial Computer Software or Computer Software Documentation). If a government agency has a need for rights not granted under these terms, it must negotiate with Provider to determine if there are acceptable terms for granting those rights, and a mutually acceptable written addendum specifically granting those rights must be included in any applicable agreement.
12. **EXPORT COMPLIANCE.** The Service and other technology Provider makes available, and derivatives thereof may be subject to export laws and regulations of the United States and other jurisdictions. Each party represents that it is not named on any U.S. government denied-party list. Customer shall not permit Authorized User to access or use the Service in a U.S.-embargoed country (currently Cuba, Iran, North Korea, Sudan or Syria) or in violation of any U.S. export law or regulation.
13. **MISCELLANEOUS.** This Agreement represents the entire agreement of the parties with respect to the subject matter hereof, and supersedes and replaces all prior and contemporaneous oral or written understandings and statements by the parties with respect to such subject matter. In entering into this Agreement, neither party is relying on any representation not expressly specified in this Agreement. This Agreement may only be amended in writing signed by each party. This Agreement may be executed in two or more counterparts. Section headings herein are for convenience only. Provider may assign this Agreement (or any of its rights and obligations) without restriction or obligation. Customer may not assign this Agreement (or any of its rights and obligations) without Provider's prior express written consent. Any prohibited assignment shall be null and void. If any provision of this Agreement is held by a court of competent jurisdiction to be invalid, illegal, or unenforceable, then the remaining provisions of this Agreement shall remain in full force and effect. Rights and remedies herein are cumulative of all rights and remedies available at law or in equity. No failure or delay on the part of any party hereto in exercising any right or remedy under this Agreement shall operate as a waiver thereof, nor shall any single or partial exercise of any such right or remedy preclude any other or further exercise thereof or the exercise of any other right or remedy. Any waiver granted hereunder must be in writing and shall be valid only in the specific instance in which given. The relationship of the parties is solely that of independent contractors. Provider shall not be responsible for any failure to perform any obligation because of any cause beyond its reasonable control. Provider may use Customer's name and logo on Provider's website and in its promotional materials to state that Customer is a customer of Provider.