Episode Summary
Episode Show Notes & Transcript
About Ken
- OpsLevel: https://www.opslevel.com/
- LinkedIn: https://www.linkedin.com/company/opslevel/
- Twitter: https://twitter.com/OpsLevelHQ
Transcript
Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.
Corey: Welcome to Screaming in the Cloud, Iâm Corey Quinn, about, oh I donât know, two years ago and change, I wound up writing a blog post titled, âDeveloper Portals are An Anti Pattern,â and I havenât really spent a lot of time thinking about them since. This promoted guest episode is brought to us by our friends at OpsLevel, and they have sent their CTO and co-founder Ken Rose, presumably in an attempt to change my perspective on these things. Letâs find out. Ken, thank you for agreeing to, well, run the gauntlet, for lack of a better term.
Ken: Hey, Corey. Thanks again for having me. And Iâve heard, you know, heard and listened to your show a bunch, and really excited to be here today.
Corey: Letâs begin with defining our terms. Iâm curious to know what a developer portal is. âWhat would you say a developer portal means to you?â Like itâs a college entrance essay.
Ken: Right? Definitely. You know, so really, a developer portal is this consolidated place for developers to come to, especially in large organizations to be able to get their jobs done more easily, right? A large challenge that developers have in large organizations, thereâs just a lot to do and a lot to take care of. So, a developer portal is a place for developers to be able to better own, manage, and run the services, theyâre responsible for that run in production, and they can do that through access, easy access to self-service tooling.
Corey: I guess, on some level, this turns into one of those alignment charts of, like, what is a database and, like, how prescriptive you want to be. Itâs like, well is a senior engineer a database because you can query them and they have information? Would you consider, for example, Kubernetes be a developer platform, and/or would the AWS console?
Ken: Yeah, thatâs actually an interesting question, right? So, I think thereâs actually twoâweâre going to get really niggly hereâthereâs developer platform and developer portal, right? And the word portal for me is something that sits above a developer platform. I donât know if you remember, like, the late-90s, early-2000s, like, portals were all the rage.
Like, Yahoo and AltaVistas were like search portals, they were trying to, at the time, consolidate all this information on a much smaller internet to make it easy to access. A developer portal is sort of the same thing, but custom-built for developers and trying to consolidate a lot of the tooling that exists. Now, in terms of the AWS console? Yeah, maybe. Like, it has a suite of tools and suite of offerings. It doesnât do a lot on the well, how do I quickly find out whatâs running in production and who is responsible for it? I donât know, unless AWS shipped, like, their, you know, three-hundredth new offering in the last week that I havenât, you know, kept on top of.
But you know, thereâs definitely some spectrum in terms of what goes into a developer portal. For me, thereâs kind of three main things you need. You do need some kind of a catalog, like, whatâs out there who owns it; you need some kind of a way to measure, like, how good are those services, like, how well built are they; and then you need some access to self-service tooling. And that last part is where, like, the Kubernetes or AWS could be, you know, sort of a dev portal as well.
Corey: My experience with developer portalsâthere was a time when I loved it. RightScale was what I usedâat some depthâback in I want to say 2010, 2011 because the EC2 console was clearly not built or designed by anyone who had not built EC2 themselves with their bare hands and sweat of their brow. And in time, the EC2 console got better where it wasnât written in hieroglyphics, as best we could tell, and it became âclick button to launch instance.â And RightScale really didnât have a second act and they wound up getting acquired by our friends over at Flexera years later. And I havenât seen their developer portal in at least eight years as a direct result of this.
So, the problem, at least when I was viewing it purely in the context of AWS services, it feels like you are competing against AWS iterating forward on developer experience, which they iterate slowly, sometimes, and unevenly across their breadth of services, but it does feel like at some level by building an internal portal, you are, first, trying to out-innovate AWS, in some ways, and two, you are inherently making the trade-off of not using recent features and enhancements that have not themselves been incorporated into the portal. Thatâs where the, I guess the start, the genesis of my opposition to the developer portal approach comes from. Is that philosophy valid these days? Not as much. Because I can see an argument for it shifting.
Ken: Yeah, I think itâs slightly different. I think of a developer portal as again, itâs something that sort of sits on top of AWS or Google Cloud or whatever cloud provider use, right? You give an example for example with RightScale and EC2. So, provisioning instances is one part of the activity you have to do as a developer. Now, in most modern organizations, you have, like, your product developers that ship features. They donât actually care about provisioning instance themselves. There are another group called the platform engineers or platform group that are responsible for building automation and tooling to help spin up instances and create CI/CD pipelines and get everything you need set up.
And they might use AWS under the covers to do that, but the automation built on top and making that accessible to developers, thatâs really what a developer portal can provide. In addition, it also provides links to operational tooling that you need, technical documentation, itâs everything you need as a developer to do your job, in one place. And though AWs bills itself is that, I think of them as more, they have a lot of platform offerings, right, they have a lot of infra-offerings, but they still havenât been able to, I think, customize that, unless youâre an organization that buildsâthat has kind of gone in-all on AWS and doesnât build any of your own tooling, thatâs where a developer portal helps. It really helps by consolidating all that information in one place, by making that information discoverable for every developer so they have less⌠less cognitive load, right? Weâve asked developers to kind of do too much that we donât⌠weâve asked to shift left and well, how do we make that information more accessible?
Regarding the point of, you know, AWS adds new features or new capabilities all the time and, like, well you have this dev portal, thatâs sort of your interface for how to get things done. Like, how will you use those? Dev portal doesnât stop you from doing that, right? So, my mental model is, if Iâm a developer, and I want to spin up a new service, I can just press a button inside of my dev portal in my company and do that. And I have a service that is built according to the latest standards, it has a CI/CD pipeline, it already has aâyou know, itâs registered in PagerDuty, itâs registered in Datadog, it has all the various bits.
And then thereâs something else that I want to do that isnât really on the golden path because maybe this is some new service or some experiment, nothing stops us from doing that. Like, you still can use all those tools from AWS, you know, kind of raw. And if those prove to be valuable for the rest of the organization, great. They can make their way into the dev portal; they can actually become a source of leverage. But if theyâre not, then they can also just sit there on the vine. Like, not everything that eight of us ever produces will be used by every company.
Corey: Many years ago, I got a Cisco pair of certifications because recession was hitting and I needed to be better at networking. And taking those certifications, in those days before Cisco became the sad corporate dragon with no friends we all know today, they were highly germane and relevant. But I distinctly remember, even now, 15 years later, that there was this entire philosophy of pretend that the entire world is Cisco only, which in networking is absolutely never true. It feels like a lot of the AWS designs and patterns tend to assume, oh yeah, youâre going to use AWS services for everything. I have never yet found that to be true, other than when Iâm just trying to be obstinate.
And hell is interoperability between a bunch of different things. Yes, I may want to spin up an EC2 instance and an AWS load balancer and some S3 storage or whatnot, but Iâm also going to want to monitor it with PagerDuty, Iâm going to want to have a CDN that isnât CloudFront because most CDN these days donât hate you in quite the same economic ways and are simpler to work with, et cetera, et cetera, et cetera. So, thereâs definitely a story wherein Iâve found that thereâs anâthe interoperability of tying these things together is helpful. How do you avoid falling down the trap of oh, everyone should be multi-cloud, single pane of glass at cetera, et cetera? In practice that always seems to turn to custard.
Ken: Yeah, I think multi-cloud and single pane of glass are actually two different things. So multi-cloud, like, I agree with you to some sense. Like, pick a cloud and go with it, like, unless you have really good business reasons to go for multi-cloud. And sometimes you do, like, years ago, I worked at PagerDuty, they were multi-cloud for a reliability reason, that hey, if one cloud provider goes down, you donât want [crosstalk 00:08:40]â
Corey: They were an example I used all the time for that storyâ
Ken: Right.
Corey: âspecifically the thing woke you up was homed in a bunch of different places, whereas the marketing site, the onboarding flow, the periphery stuff around it was not because it didnât need to be.
Ken: Exactly.
Corey: Like, the core business need of wake you up was very much multi-cloud because once upon a time, it wasnât and it went down with the rest of us-east-1 and people werenât woken up to be told their site was on fire.
Ken: A hundred percent. And on the kind of like application side where, even then, pick a cloud and go with it, unless thereâs a really compelling business reason for your business to go multi-cloud. Maybe thereâs something credits or compliance or availability, right? There might be reasons, but you have to be articulate about whether theyâre right for you.
Now, single pane of glass, I think thatâs different, right? I do think thatâs something that, ultimately, is a net boon for developers. In any large organization, there is a myriad of internal tools that have been built. And itâs like, well, how do I provision a new topic in the Kafka cluster? How do I actually get access to the AWS console? How do I spin up a new service, right? How do I kind of do these things?
And if Iâm a developer, I just want to ship features. Like, thatâs what Iâm incented to do, thatâs what Iâm optimizing for. And all this other stuff I have to do as part of my job, but I donât want to have to become, like, a Kubernetes guru to be able to do it, right? So, what a developer portal is trying to do is be that single pane of glass, bringing all these common set of tools and responsibilities that you have as a developer in one place. Theyâre easy to search for, theyâre easy to find, theyâre easy to query, theyâre easy to use.
Corey: I should probably have asked this earlier on, but letâs disambiguate for a little bit here. Because when Iâm setting up to use a new service or product and kick the tires on it, no two explorations really look the same. Whereas at most responsible mature companies that are building products that areâservices that are going to production use, theyâve standardized around a number of different approaches. What does your target customer look like? Is there a certain point of scale, a certain level of complexity, a certain maturity of process?
Ken: Absolutely. So, a tool like OpsLevel or a developer portal really only makes sense when you hit some critical mass in terms of the number of services you have running in production, or the number of developers that you have. So, when you hit 20, 30, 50 developers or 20, 30, 50 services, an important part of a developer portal is this catalog of whatâs out there. Once you kind of hit the Dunbar number of services, like, when you have more than you keep in your head, thatâs when you start to need tooling like this. If you look at our customer base, theyâre all you know, kind of medium to large-sized companies. If youâre a startup with, like, ten people, OpsLevel is probably not right for you. We use all playable internally at OpsLevel, and you know, like, weâre still a small company. Itâs like, we make it work for us because we know how to get the most out of it, but like, itâs not the perfect fit because itâs not really meant for, you know, smaller companies.
Corey: Oh, I hear you. I think Iâm probably⌠I have a better AWS bill analytic system running internally here at The Duckbill Group than some banks do. So, I hear you on that front.
Ken: I believe it.
Corey: But also implies to me that thereâs no OpsLevel prospect or customer deployment that has ever been greenfield. Itâs always youâre building existing things, thereâs already infrastructure in place, vendors have been selected across the board. You arenâtâdonât to want to starting a company day one, theyâre going to all right, time to spin up our AWS account and weâre also going to wind up signing up for OpsLevel, from the sound of it.
Ken: Correctâ
Corey: Accurate? Inaccurate?
Ken: I think thatâs actually accurate. Like, a lot of the problems, we solve other problems that come as you start to scale both your product and your engineering team. And itâs the problems of complexity.
Corey: What do those painful problems look like? In other words, what is someone sitting at home right now listening to this, or driving to work debating whether want to ram a bridge abutment or go into the office depending on their mental state today, what painful problem did they have that OpsLevel is designed to fix?
Ken: Yeah, for sure. So, letâs help people self-select. So, hereâs my mental model for any [unintelligible 00:12:25]. There are product developers, platform developers, and engineering leaders. Product developers, if youâre asking questions like, âI just got paged for the service. I donât know what this does.â Or, âItâs upstream from here. Where do I find the technical documentation?â Or, âI think I have to do something with the payment service. Where do I find the API for that?â
You know, when you get to that scale, a developer portal can help you. If youâre a platform engineer and you have questions like, âOkay, we got to migrate. Weâre migrating, I donât know, from Datadog to Honeycomb, right? We got to get these fifty or a hundred or thousands of services and all these different owners to, like, switch to some new tool.â Or, âHey, weâve done all this work to ship the golden path. Like, how to actually measure the adoption of all this work that weâre doing and if itâs actually valuable?â Right?
Like, we want everybody to be on a certain set of CI tooling or a certain minimum version of some library or framework. How do we do that? How do we measure that? OpsLevel is for you, right? We have a whole bunch of stuff around maturity.
And if youâre engineering leader, ultimately, questions you care about, like, âHow fast are my developers working? I have this massive team, weâve made this massive investment in hiring all these humans to write software and bring value for our customers. How can we be more efficient as a business in terms of that value delivery?â And thatâs where OpsLevel can help as well.
Corey: Guardrails, whether they be economic, regulatory, or otherwise, have to make it easier than doing things incorrectly because one of the miracle aspects of cloud also turns into a bit of a problem, which is shadow IT is only ever a corporate credit card away. Make it too difficult to comply with corporate policies and people wonât. And theyâre good actors; theyâre trying to get work done. Theyâre not trying to make peopleâs lives harder, but they donât want to spend six weeks provisioning an EC2 cluster. So, thereâs always that weird trade-off.
Now, it feelsâand please correct me if Iâm wrongâonce someone has rolled out OpsLevel at their organization, where it really shines is spinning up a new service where okay, great, youâre going to spin up the automatic observability portion of it, youâre going to spin up the underlying infrastructure in certain ways that comply with our policies, itâs going to build the CI/CD pipelines around it, youâre going to wind up having the various cost instrumentation rolled out to it. But for services that are already excellent within the environment, is there an OpsLevel story for them?
Ken: Oh, absolutely. So, I look at it as, like, the first problem OpsLevel helps solve is the catalog and whatâs out there and who owns it. So, not even getting developers to spin up new services that are kind of on the golden path, but just understanding the taxonomy of what are the services we have? How do those services compose into higher-level things like systems or domains? Whatâs the whole set of infrastructure we have?
Like, I have 50 AWS accounts, maybe a handful of GCP ones, also, some Azure. I have all this infrastructure that, like, how do I start to get a handle on, like, whatâs out there in prod and whoâs responsible for it. And that helps you get in front of compliance risks, security risks. Thatâs really the starting point for OpsLevel building that catalog. And we have a bunch of integrations that kind of slurp all this data to automatically assemble that catalog, or YAML as well if thatâs your thing. But thatâs the starting point is building that catalog and figuring out this assignment of, like, okay, this service and this human, or thisâsorryâteam, like, theyâre paired together.
Corey: A number of offerings in this space, which honestly, my exposure to it is bounded simultaneously to things that are ten years old and no one uses anymore, or a bunch of things I found on GitHub. And the challenge that both of those products tend to have is that they assume certain things to be true about a given environment: that theyâre using Terraform to manage everything, or theyâre always going to be using CloudFormation, or everyone there knows Python or something else like that. What are the prerequisites to get started with OpsLevel?
Ken: Yeah, so we worked pretty hard to build just a ton of integrations. I would say integrations is are just continuing thing we have going on in the background. Like, when we started, like, we only supported a GitHub. Now, we support all the gits, you know, like GitHub, GitLab, Bitbucket, Azure DevOps, like, weâre building [unintelligible 00:16:19]. Thereâs just a whole, like, long tail of integrations.
The same with APM tooling. The same with vulnerability management tooling, right? And the reason we do that is because thereâs just this huge vendor footprint, and people, you know, want OpsLevel to work for them. Now, the other thing we try to do is we also build APIs. So, anything we have as, like, a core integration, we also have kind of like an underlying API for, so that thereâs, no matter what you have an escape hatch. If like, youâre using some tool that we donât support or you have some homegrown thing, thereâs always a way to try to be able to integrate that into OpsLevel.
Corey: When people think about developer portals, the most common one that pops to mind is Backstage, which Spotify wound up building, internally, championing, open-sourcing, and I believe, on some level, turned into a product because if thereâs one thing people want, itâs to have their podcast music company become a SaaS vendor, which is weird to me. But the criticisms that Iâve seen about and across the board have rung relatively true, including from people internal at Spotify who have used the thing, which is, well first is underestimating the amount of effort that is necessary to maintain Backstage itself, that the build versus buy discussion is always harder to buâengineers love to build, but they shouldnât be building things outside of their core competency half the time, and the other is driving adoption within the org where you can have the most amazing developer portal in the known universe, but if people donât use it, it may as well not exist and doing the carrot and stick approach often doesnât work. I think you have a pretty good answer that I need not even ask you to elaborate on, âWell, how do we avoid having to maintain this ourselves,â since you have a company that does this, but how do you find companies are driving adoption successfully once they have deployed OpsLevel?
Ken: Yeah, thatâs a great question. So, absolutely. Like, I think the biggest thing you need first, is kind of cultural buy-in and that this is a tool that we want to invest in, right? I think one of the reasons Spotify was successful with Backstage and I think it was System Z before that was that they had this kind of flywheel of, like, they saw that their developers were getting, you know better faster, working happier, by using this type of tooling, by reducing the cognitive load. The way that we approach it is sort of similar, right?
We want to make sure that there is executive buy-in that, like, everybody agrees this is, like, a problem thatâs worth solving. The first step we do is trying to build out that catalog again and helping assign ownership. And that helps people understand, like, hey, these are the services Iâm responsible for. Oh, look, and now hereâs this other context that I didnât have before. And then helping organizations, you know, whatâit depends on the problem weâre trying to solve, but whether itâs rolling out self-serve automation to help developers, like, reduce what was before a ton of cognitive load or if itâs helping platform teams define what good looks like so they can start to level up the overall health of whatâs running in production, we kind of work on different problems, but itâs picking one problem and then you know, kind of working with the customers and driving it forward.
Corey: On some level, I think that this is going to be looked down upon inherently just by automatic reflex of folks with infrastructure engineering backgrounds. Itâs taken me some time to learn to overcome my own negative reaction to it. Because itâs, Iâm here to build things and I want to build things out in such a way that itâs portable and reusable without having to be tied to a particular vendor and move on. And it took me a long time to realize that what that instinct was whispering in my ear was in fact, no, you should be your own cloud provider. If thatâs really what I want to do, I probably should just brush up on you know, computer science trivia from 20 years ago and then go see if I can pass Googleâs SRE interview.
Iâm not here to build the things that just provision infrastructure from scratch every company I wind up landing at. It feels like thereâs more important, impactful work that I can do. And letâs be clear, people are never going to follow guardrails themselves when they have to do a bunch of manual steps. It has to be something that is done for them. And I donât know how you necessarily get there without having some form of blueprint or something like that, provided for them with something that is self-service because otherwise, itâs not going to work.
Ken: I a hundred percent agree, by the way, Corey. Like, the take that, like, automation is the only way to drive a lot of this forward is true, right? If for every single thing youâre tryingâlike, we have a concept called a rubric and itâs basically how you measure the service health. And you canâitâs very customizable, you have different dimensions. But if, for any check thatâs on your rubric, it requires manual effort from all your developers, that is going to be harder than something you can just automate away.
So, vulnerability management is a great example. If you tell developers, âHey, you have to go upgrade this library,â okay, some percentage [unintelligible 00:20:47], if you give developers, âHereâs a pull request thatâs already been done and has a test passing and now you just need to merge it,â youâre going to have a much better adoption rate with that. Similarly with, like, applying templates being able to [up-level 00:20:57], you know, kind of apply the latest version of a template to an existing service, those types of capabilities, anything where you can automate what the fixes are, absolutely youâre going to get better adoption.
Corey: As you take a look at your existing reference customersâwhich is something I always look for on vendor websites because, like, oh, we have many customers who will absolutely not admit to being customers, itâs like, that sounds like something thatâs easy to sayâyou have actual names tied to these things. Not just companies, but also individuals. If you were to sit down and ask your existing customer base, âSo, why did you wind up implementing OpsLevel and what has the value thatâs delivered to you been since that implementation?â What do they say?
Ken: Definitely. I actually had to check our website because we, you know, land new customers and put new logos on it. I was like, âOh, I wonder what the current set is out right now?â
Corey: I have the exact same challenge. Like oh, we have some mutual customers. And itâs okay. I donât know if I can mention them by name because I havenât checked our own list of testimonials [unintelligible 00:21:51] lately because say the wrong thing and thatâs how you wind up being sued and not having a company anymore.
Ken: Yeah. So, I donâtâI definitely, you know, want to stay [on side 00:22:00] on that part, but in terms of, like, kind of sample reference customer, a lot of the folks that we initially worked with are the platform teams, right? Theyâre the teams that care about whatâs out there, and they need to know whoâs responsible for it because theyâre trying to drive some kind of cross-cutting change across the entire, you know, production footprint. And so, the first thing that generally people will say isâand I love this quote. This cameâI wonât name them, but like, itâs in one of our case studies.
It was like, âI had, like, 50 different attempts at making a spreadsheet and theyâre all, like, in the graveyard, like, to be able to capture whatâs out there and whoâs responsible for it.â And just OpsLevel helping automate that has been one of the biggest values that theyâve gotten. The second point, then is now be able to drive maturity and be able to measure how well those services are being built. And again, itâs sort of this interesting thing where we start with the platform teams. And then sometime later security teams find out about OpsLevel, and theyâre like, âOh, this is a tool I can use to, like, get developers to do stuff? Like, Iâve been trying to get developers to do stuff for the longest time.â
And theyâI file Jira tickets and they just sit there and nothing gets done. But when it becomes part of this, like, overall health score that youâre trying to increase a part of the across the board, yeah, itâs just a way to kind of drive action.
Corey: I think that thereâs a dichotomy of companies that emerge. And I tend to see the world through a lens of AWS bills, so letâs go down that path. I feel like there are some companies presumably like OpsLevel, whereas if Iâassuming youâre running on top of AWSâif I were to pull your AWS bill, I would see upwards of 80% of your spend is going to be on this application called OpsLevel, the service that you provide to people. As opposed to the other side of the world, which is large enterprises, where theyâre spending hundreds of millions of dollars a year, but the largest application they have is a million-and-a-half a year in spend because just, they have thousands of these things scattered everywhere. That latter case is where I tend to see more platform teams, where I start to see a lot of managing a whole bunch of relatively small workloads. And developer platforms really seem to be where a lot of solutions lead, whereas 80% of our workload is one application, we donât feel the need for that as much. Is that accurate? Am I misunderstanding some aspect of it?
Ken: No, a hundred percent youâd hit the nail on the head. Like, okay, think about the typical, like, microservices adoption journey. Like, you started with, you know, some small companyâlike usâyou started with a monolith. Ah, maybe you built out a second appâ
Corey: Then you read on Hacker News and realize, âOh, if we want to hire people, weâve got to be doing what all the cool kids are up to.â
Ken: Right. We got a microservice all the thingâbut thatâs actually you know, microservices should come later, right, as a response to you needing scale your org and scale yourâ
Corey: As someone who started building some application with microservices, I could not agree more.
Ken: A hundred percent. So, itâs as youâre starting to take steps to having just more moving parts in your production infrastructure, right? If you have one moving part, unless itâs like a really large moving part that you can internally break down, like, kind of this majestic monolith where you do have kind of like individual domains that are owned by different teams, but really the problem weâre trying to solve, itâs more about, like, who owns what. Now, if thatâs a single atomic unit, great, but can you decompose that? But if you just have, like, one small application, kind of like the whole team is owning everything, again, a developer portal is probably not the right tool for you. It really is a tool that you need as you start to scale your engineer work and as you start to scale the number of moving parts in your production infrastructure.
Corey: I tended to use to think of that in terms of boring companies versus innovative ones and I donât think thatâs accurate. I think it is the question of maturity and where companies lead to. On some level, of OpsLevel starts growing and becomes larger and larger in different ways and starts doing acquisitions and launching into other areas, at some point, you donât have just one product offering, you have a multitude of them. At which point having something like that is going to be critical. But I have to ask, given that you are sort of not exactly your target customer profile, what are the sharp edges been on using it for your use case?
Ken: Yeah. So, we actually have an internal Slack channel, we call OpsLevel on OpsLevel. And finding those sharp edges actually has been really useful for us. You know, all the good stuff, dogfooding and it makes your own product better. Okay, so we have our main app, we also do have a bunch of smaller things and itâs like, oh yeah, you know, we have, like, I donât know, various Hackaday things that go on, itâs important we kind of wind those down for, you know, compliance, we have our marketing site, we have, like, our Terraform.
Like, so thereâs, like, stuff. Itâs not, like, hundreds or thousands of things, but thereâs more than just the main app. The second though, is itâs really on the maturity piece that we really try to get a lot of value out of our own product, right? Helpingâwe have our own platform team. Theyâre also trying to drive certain initiatives with our product developers.
There is that usual tension of our, like, our own product developers are like, âI want to ship features.â Whatâs this security thing I have to go take care of right now? But OpsLevel itself, like, helps reflect that. We had an operational review today and it was like, âOh, this one service is actually nowââwe have platinum as a level. Itâs in gold instead of platinum. Itâs like, âWhy?â âOh, thereâs this thing that came up. We got to go fix that.â âGreat. Letâs go actually go fix that so weâre back into platinum.â
Corey: Do you find that thereâs often a choice you have to make internally, where you could make the product more effective for your specific use case, but that also diverges from where your typical customer needs or wants the product to go?
Ken: No, I think a lot of the things we find for our use case are, like, theyâre more small paper cuts, right? Theyâre just as weâre using it, itâs like, âHey, like, as Iâm using this, I want to see the report for this particular check. Why do I have to click six times to get?â You know, like, âWouldnât it be great if we had a button?â Right?
And so, itâs those type of, like, small innovations that kind of come up. And those ultimately lead to, you know, a better product for our customers. We also work really closely with our customers and developers are not shy about telling you what they donât like about your product. And I say this with love, like, a lot of our customers give us phenomenal feedback just on how our product can be better and we try to internalize that and you know, roll that feedback into the product.
Corey: You have a number of integrations of different SaaS providers, infrastructure providers, et cetera, that you wind up working with. I imagine that given your scale and scope and whatnot, those offerings are dictated by what customers say, âHey, weâre using this thing. Are you going to support that or are you not going to maintain our business?â Which is a great way to wind up financing a lot of product development and figuring out what matters to people. My question for you is, if you look across the totality of your user base, what are the most popularly used integrations, if you can say?
Ken: Yeah, for sure. I think right nowâI could actually dive in to pull the numbersâGitHub and GitLabâor⌠I think GitHub, like, has slightly more adoption across our customer base. At least with our customers, almost nobody uses Bitbucket. I mean, we have, like, a small number, but, like, itâs⌠I think, single-digit percentage. A lot of people use PagerDuty, which you know, hey, Iâm an ex-PagerDuty person [crosstalk 00:28:24] and Iâm glad to see that.
Corey: I have a free tier PagerDuty account that will automatically page me for my home automation stuff. Specifically, if you know, the fire alarm goes off. Like, yeah, okay, there are certain things I want to be woken up for, but itâs a very short list.
Ken: Yeah, itâs funny, the running default message when we use a test PagerDuty was, âThe server is on fire.â [unintelligible 00:28:44] be like, âThe house is on fire.â Like you know, go get that taken care of. Thereâs one other tool so thatâs used a lot. Datadog actually is used a ton by just across our entire customer base, despite its⌠weâre also Dataâweâre a Datadog partner, weâre a Datadog customer, you know? Itâs not cheap, but itâs a good product for, you know, monitoring and logs and there are [crosstalk 00:29:01]â
Corey: No other than cloud infrastructure providers, I get the number one most common source of inquiries is Datadog optimization. It has now risen to a board-level concern in many cases because observability is expensive. Thatâs a sign of success, on some level. Meanwhile, Iâm sitting here, like, Date-a-dog? Oh, my God, thatâs disgusting. Itâs like Tinder for Pets. Which it turns out is not at all what they do.
Ken: Nice.
Corey: Yeah.
[audio break 00:29:23]âoptimizing their Slack integrations, their GitHub integration, et cetera. Or are they starting with the spinning up the servers piece of it?
Ken: A lot of the timeâand again, that first problem theyâre trying to solve is just get me a handle on everything we have running in production. You know, if you have multiple AWS accounts, multiple Kubernetes clusters, dozens or even hundreds of teams, God help you if youâre going to try to, like, build a list manually to consolidate all that information. Thatâs really the first part is, like, integrate Kubernetes, integrate your CI/CD pipelines, integrate Git, integrate your Cloud account, like, will integrate with everything and will try to build that map of, like, hereâs everything thatâs out there, and start to try to assign it to, like, and hereâs people that we think might be responsible in terms of owning the software. Thatâs generally the starting point.
Corey: Which makes an awesome amount of sense. I think going at it from the infrastructure first perspective is where Iâve seen most developer platforms founder. And to be fair, the job is easier now than it was years ago because it used to be that you were being out-innovated by AWS constantly. Innovation has slow down there. And you know that because of how much they say the pace of innovation has only sped up.
And whenever AWS says something in a marketing context, theyâre insecure about it. Iâve learned this through the fullness of time observing that company. And these days, most customers do not use the majority of features available for any given service. They have solidified to a point where you can responsibly build on top of these things. Now, it seems that the problem is all the âyes, andâ stuff that gets built on top of it.
Ken: Yeah. Do you have an example, actually, like, one of the kinds of, like, âyes, andâ tools that youâre thinking about?
Corey: Oh, absolutely. We have a bunch of AWS environment stuff so we should configure CloudWatch to look at all these things from an observability perspective. No, you should not. You should set up Datadog. And the first time someone does that by hand, they enable all have the observability and the rest and suddenly get charged approximately the GDP of Guam.
And okay, maybe we shouldnât do that because then you have the downstream impact of that on your CloudWatch bill. So okay, how do we optimize this for the observability piece directly tied to that? How do we make sure that we get woken up when the site is down or preferably before that, but not every time basically, a EBS volume starts to get a little bit toasty? You have to start dialing this stuff in. And once youâve found a lot of those aspects, being able to templatize that and roll that out on an ongoing basis and having the integrations all work together feels like itâs the right problem to be solving.
Ken: Yeah, absolutely. And the group that I think is responsible for that kind ofâbecause itâs a set of problems you describedâis really, like, platform teams. Sometimes service owners for like, how should we get paged, but really, what youâre describing are these kind of cross-cutting engineering concerns that platform teams are uniquely poised to help solve in an [unintelligible 00:32:03] organization, right? I was thinking what you said earlier. Like, nobody just wants to rebuild the same info over and over, but itâs sort of like, itâs not just building an [unintelligible 00:32:09]; itâs kind of like solving this, like, how do we ship? Can we actually run stuff in prod? And not just run it but get observability and ensure that weâre woken up for it and, like, whatâs that total end-to-end look like from, like, developers writing code to running software in production thatâs serving traffic? And solving all the problems [unintelligible 00:32:24], thatâs what I think of was platform engineering.
Corey: So, my last question before we wind up wrapping this episode comes down to, I am very adept at two different programming languages, and those are brute force and enthusiasm. What implementation language is most of what you find yourself working with? And why is it in invariably going to be YAML?
Ken: Yeah, thatâs a great question. So, I think thereâs, in terms of implementing OpsLevel and implementing a service catalog, we support YAML. Like, you know, thereâs this very common workflow, you just drop a YAML spec, basically, in your repo, if youâre a service owner. And that, we can support that. I donât think thatâs a great take, though.
Like, we have other integrations. Again, if the problem youâre trying to solve is I want to build a catalog of everything thatâs out there, asking each of your developers hey, can you please all write YAML files that, like, describe the services you own and drop them into this repo? Youâve inverted this, like, database that essentially youâre trying to build, like, whatâs out there and stored it in Git, potentially across several hundreds or thousands of repos. You put a lot of toil now on individual product developers to go write and maintain these files. And if you ever had to, like, make a blanket update to these files, thereâs no atomic way to kind of do that, right?
So, I look at YAML as, like, I get it, you know? Like, we use the YAML for all the things in DevOps, so why not their service catalog as well, but I think itâs toil. Like, there are easier ways to build a catalog. By, kind of, just integrate. Like, hook up AWS, hook up GitHub, hook up Kubernetes, hook up your CI/CD pipeline, hook up all these different sources that have information about whatâs running in prod, and let the software, let the tool, automatically infer whatâs actually running as opposed to requiring humans to manually enter data.
Corey: I find that there are remarkably few technical holy wars that I cannot unify both sides on by nominating something far worse. Like, the VI versus Emacs stuff, the tabs versus spaces, and of course, the JSON versus YAML folks. My JSON versus YAML answer is XML: Godâs language. I find that as soon as you suggest that, people care a hell of a lot less about the differences between JSON and YAML because their job is to now kill the apostate, which is me.
Ken: Right. Yeah. I remember XML, like, oh, man, 2002. SOAP. I remember SOAP as a protocol. That was a thing.
Corey: Some of the earliest S3 API calls were done in SOAP, and I think they finally just used it to wash their mouths out when all was said and done.
Ken: Nice. Yeah.
Corey: I really want to thank you for taking the time to do your level best to attempt to convert me, and I would argue in many respects, you have succeeded. Iâm thinking about this differently than I did half an hour ago. If people want to learn more, whereâs the best place for them to find you?
Ken: Absolutely. So, you can always check out our website, opslevel.com. Weâre also fairly active on LinkedIn. If Twitter hasnât imploded by the time this episode becomes launched, then they can also check us out at twitter.com/OpsLevelHQ. Weâre always posting, just different content on, like, how to be successful with service maturity, DevOps, developer productivity, so that you know, ultimately, that you can ship out to customers faster.
Corey: And we will, of course, put links to that in the [show notes 00:35:23]. Thank you so much for taking the time, not just to speak with me, but also for sponsoring this episode. It is appreciated.
Ken: Cheers.
Corey: Ken Rose, CTO and co-founder at OpsLevel. Iâm Cloud Economist Corey Quinn and this has been a promoted guest episode of Screaming in the Cloud. If youâve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if youâve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment which, upon further reflection, you could have posted to all of the podcast platforms if only you had the right developer platform to pull it off.
Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

