---
title: "The Invisible Network Powering AWS with Matt Rehder"
id: "15436"
type: "podcast"
slug: "the-invisible-network-powering-aws-with-matt-rehder"
published_at: "2026-09-10T10:30:00+00:00"
modified_at: "2026-09-10T10:32:28+00:00"
url: "https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/the-invisible-network-powering-aws-with-matt-rehder/"
markdown_url: "https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/the-invisible-network-powering-aws-with-matt-rehder.md"
taxonomy_shows:
  - "Screaming in the Cloud"
---

About the Author Corey is the Chief Cloud Economist at Duckbill, where he specializes in helping companies improve their AWS bills by making them smaller and less horrifying. He also hosts the "Screaming in the Cloud" and "AWS Morning Brief" podcasts; and curates "Last Week in AWS," a weekly newsletter summarizing the latest in AWS news, blogs, and tools, sprinkled with snark and thoughtful analysis in roughly equal measure.

[https://podcasts.apple.com/us/podcast/screaming-in-the-cloud/id1361244178](https://podcasts.apple.com/us/podcast/screaming-in-the-cloud/id1361244178)

[https://overcast.fm/itunes1361244178/screaming-in-the-cloud](https://overcast.fm/itunes1361244178/screaming-in-the-cloud)

[https://pca.st/7l2e](https://pca.st/7l2e)

[https://open.spotify.com/show/3fBA9eNkGliCzp3Xuy1GVd](https://open.spotify.com/show/3fBA9eNkGliCzp3Xuy1GVd)

[https://feeds.transistor.fm/screaming-in-the-cloud](https://feeds.transistor.fm/screaming-in-the-cloud)

## Episode Summary

## Episode Video

## Episode Show Notes & Transcript

AWS VP of Global Networking Matt Rehder joins Corey Quinn to pull back the curtain on the massive network infrastructure behind AWS. They explore resiliency at scale, AWS’s move toward flatter networks, the advantages of building custom hardware, and why AI is making networking exciting again.

**Show Highlights:**

(01:08) Meet AWS Networking Lead

(02:01) Why AWS Avoids Global Outages

(05:23) RNG Flat Network Explained

(10:08) Overbuild Capacity And Custom Hardware

(17:37) VPC Virtual Network Origins

(19:50) Why TCP Still Wins

(20:38) SRD Inside AWS

(21:54) Opt In SRD Transport

(25:31) Networks As Utilities

(27:24) Learning And Growing Engineers

(29:22) Training Talent In House

(31:27) AI Rekindles Networking

(34:26) Where To Learn Networking

### Transcript

Matt: We build more network capacity than we need, both from a reliability perspective so that we have a lot of redundancy and resiliency. Many devices can fail, many links can fail, we still have sufficient capacity

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. Somehow, I have managed to breach containment and go talk to people at Amazon. My guest today is Matt Rehder, who is the VP of Global Networking. Matt, thank you for joining me.

Matt: Yeah, thanks for having me.

Corey: This episode is sponsored in part by my day job, Duckbill.

Do you have a horrifying AWS bill? That can mean a lot of things. Predicting what it's going to be, determining what it should be, negotiating your next long-term contract with AWS, or just figuring out why it increasingly resembles a phone number, but nobody seems to quite know why that is. To learn more, visit duckbillhq.com.

Remember, you can't duck the Duckbill bill, which my CEO reliably informs me is absolutely not our slogan.

So what is it you do exactly? Networking, what's the point of it all? "That feels like something old people care about," said the old person.

Matt: Yeah. Uh, I mean, it... The simplest answer is we make all the servers talk to each other.

Um, but, um- Otherwise,

Corey: they're just expensive space heaters.

Matt: Otherwise, they're very expensive space heaters. Yeah. So it's, it's really interconnecting all the servers, interconnecting all the data centers around the world, and then connecting them to the, all the people around the world on the internet, is, is the gist of it.

Corey: Yes, I... You were kind enough to recently give me a tour- Mm ... of the networking lab, one of the networking labs. I'm guessing there might be more than one. We have more

Matt: than one.

Corey: Because we're Amazon.

Matt: Yeah. Yeah. We didn't show you the super-secret one, but we can, yeah.

Corey: Oh, exactly. Yeah. The, the real secret one.

Yeah. That's where the cables are longer so we can fit more data into them before it comes back around. No, yeah, that's how networking works, right? My networking background myself- Mm ... is such that it's just enough to know what I don't know, and... but it does guide me toward asking some of the right questions.

Mm-hmm. There's a whole different sense of scale that I don't think most folk appreciate. Yeah. Uh, one of the enduring qualities of AWS that has also led to frustration, but I maintain is the right answer, is your networking is phenomenal. It i- You have a harsh separation between regions. You have, knock wood, uh, never yet seen a global networking outage.

There's never been a global rolling outage of things going down. Error, failures are region-bound in virtually every case I am aware of. And that was a choice, and it does mean that when you log into an AWS account, each region makes it look like there are now 31 AWS accounts. Go hunt and find it down. But the resilience and durability story are unparalleled by any other company.

I include every other hyperscaler in that list. Resilience is great, and it's been that way for a long time. Now it seems that- Systems have gotten good enough across the board, especially since modern AI systems are not that reliable. We're not seeing five nines of uptime on anything, but companies are putting it in their critical path to a point where it feels like resilience is something companies pay lip service to largely.

It is an afterthought more than a lot of other things. It just isn't an area of focus. My sense is that this is something that happens when you've been too long without a really good, really notable outage. It's sort of like a self-inflicted, we've been too good. People forget this stuff can break. How do you see it?

Matt: Well, first of all, thank you for the kind words. Uh, we care a lot about resiliency. That is our number one guiding principle, especially for our network, is massive reliability. And it, and obviously we wanna contain outages to within a region, but we never wanna have them span to a region. We go to the level of isolation, even with inside of individual data centers, we're, we're building isolation at every level to get to that.

In, in terms of resiliency, I, I've definitely noticed a trend. There's, there's increased risk appetite, I would say, from a lot of customers, who they wanna move faster, uh, and they're willing to, to trade off some of the multiple nines of availability that traditionally have been a primary focus for all customers.

Mm-hmm. Uh, but I think, I think that's a, a smaller number of customers. I think most customers still care deeply about resiliency. Uh, I talk to lots of them all the time- Yes ... and they're, they're very focused on resiliency. They're pushing harder and harder for higher availability, higher reliability on the network all of the time, which is great.

Uh, that's, that helps us all be better. Uh, I think for us it's a, it's a balancing act of how do we build great services that have all the level of resiliency that some customers want, uh, but don't slow ourselves down with all the resiliency, uh, in order to move as fast as we possibly can. So that's the, that's the challenge that we're facing right now.

I- I-

Corey: it's interesting. The, the resilience demands are much higher in cloud than they are in the era when most companies ran their own data centers- Yes ... just because of the correlation risk. Like, great, today my bank is down, tomorrow your grocery store is down. Well, you're bad at running websites, and so am I, the end.

When AWS has a problem, suddenly lots of companies are suddenly impacted, even those who have previously done a great job of planning for resilience. But, oh, no, we have a critical path dependency on a third-party vendor who did not realize they themselves had that ven- Yep ... dependency exposure.

Everything's interconnected. Some of these are circular. And the fact that this is not broadly appreciated, understood, et cetera, by even by many engineers and technologists, again, is testament to you folks getting this very right very early on. Yeah. Which brings us to, I guess, the showpiece thing that you wanted to show me in your networking lab, which is the, I guess, the embodiment of the jellyfish paper.

Matt: Mm-hmm.

Corey: Can you explain that in a way that... I, I mean, I could make an attempt at it- Yeah ... but I'm gonna sound way dumber than you will. Please take it away.

Matt: Yeah, no, a- as you said, uh, like, when AWS was building the cloud, we, we knew customers were trusting us with their business, and it's a, it's a tremendous responsibility.

And in order to convince them that this was a good idea, we knew we had to have nearly perfect availability. We had to be far superior to what they've been able to do. And again, that's always been a guiding principle, especially at our infrastructure and network layers. Network is underneath everything.

If you have a network issue, it affects many services, many customers. And we've had a very, very reliable network for many, many years. We're very proud of our network. But we saw an opportunity to make it even better, uh, and particularly in terms of its availability and resiliency. And so the new network is called Resilient Network Graph Resilience, or RNG for short.

Uh, and the idea is, is the way networks have been built for really the past 20 or 30 years, they're hierarchical, they're factory networks for people who know about networks, and it's really stacked layers of switches. And so there's a tier one, there's a tier two, there's a tier three. They're all interconnecting into each other, and that's how you get big scale networks.

Oh, and

Corey: three, you have aggregation switches that are on, on leaf. Yep. You have

Matt: edge switches. Distribution cores, edges. Yes. You

Corey: have a core switch. You have an out of band switch for the managing- All the things ... the servers that like to break, and

Matt: mm. Yes. All the things. Yeah. And so many layers of switches.

And for a long time, people have asked, "Why do you need all those layers in your network, Matt? Why, why can't you have a flatter network?" And, uh, mathematical theory will show you that as- is actually the most efficient way to build a network, both from a cost perspective, but also from a reliability perspective.

The problem with building a flat network is all of the interconnectedness that needs to happen to, to plug everything together. At minimal scale, you can take a, a switcher that might have, like, 32 ports on it, and if you have a small number of switches, you can connect them all together Easy, right? Uh, if you have many, many thousands of switches, you can't plug them all together.

You don't have enough ports to. So you need to effectively randomly interconnect them with some links, and then you have to figure out how to route through multiple different switch hops to get to your destination. There's, there's no structure like there is in a hierarchical network where everything is very, like, clean and, and orderly.

So Matt has taught us for years this is the best way to build a network. Many people have tried to build networks like this. No one, to my knowledge, has ever actually built a network like this at any sort of scale. We decided, uh, three or four years ago that we think we can actually pull this off. And so we started working down that path and doing a bunch of research, doing a bunch of simulation.

This is a very novel concept. Most of the network engineers on my team thought this was a really bad idea. "Don't do this, Matt. You know, this is not how we run networks. It's gonna break. It's gonna be weird." We, we eventually convinced ourselves through a lot of simulation that, no, like, this will work, uh, and it will work more reliably than the network that we have today.

And so we invested, built a- built an engineering team, ag- did a lot of work, actually started to build these things for real. Uh, learned a lot of lessons, and that's, that's one of the other keys to resiliency is learning a lot of lessons and actually taking the learnings from those lessons and then trying to, like, bake them into your product.

And so in order to build this for real, we had to build this in smaller scale, we had to test this out, we had to see how it was going to fail, the problems we were gonna have, and then iterate quickly into that before we could deploy at scale. And today we're in production in multiple data centers. Uh, and this is now the new default network for all core services for AWS.

For any new data center we're building, we're, we're-- basically, we've switched to this new RNG flat network from the traditional networks we used to build.

Corey: Yeah. Uh, my take on this is not that i- it's not gonna be possible for you to do... A real brave take on my part given you already done it and proven it.

Yeah, with the benefit of hindsight, I, I sound real awesome. No, my question is, why do it at all? Why

Matt: do it at all?

Corey: Exactly. Yeah. Because at some point it's no longer your bottleneck. Your, your network is already ridiculously reliable. So at some point, making it even more reliable- Yeah ... seems like that's no longer the bottleneck.

That is no longer the area of focus that- Yeah ... of contention that is customer exposed. Why spend the investment, research, energy, hardware time- Yeah ... et cetera, on that rather than other areas that are more, I guess, directly perceived as impactful to customer workloads?

Matt: Yeah. I, I see any outage in the network as a problem, and while our network is extremely durable and customer applications are highly reliable, we still see behind the curtain, and we still see the opportunities and the risks, and it's just we wanna make this better for our customers.

Like, to me, the network is still in the way in the sense of you know it's there, you feel it's there. Sometimes you still have network degradation. If we can make it even more reliable, it becomes more invisible, and it enables more customer workloads. And so it's part of our customer obsession, I guess, at the end of the day of like, it's just not good enough.

And I think it never will be good enough, really. It's just a constant, uh, work that we're doing to continually try and drive improvement in this. And not all customers will care, but some of them do and, and therefore, it's worth it for those customers.

Corey: Somewhere between 10 and 20 years ago. Yeah. I wound up d- helping with a cluster build-out for more or less a 256-node, uh, HPC cluster w- alike, where they wound up putting a whole bunch of servers into racks in a data center cage that they had.

Great, awesome. And the idea was that customers would then pay to basically load their data into this, do all kinds of number crunching. Yeah. And it was a multi-petabyte cluster, which at that time was reasonably impressive. Today, you're like, "Ha, that's cute, I have S3 buckets like that," which, different era.

And a question I had for them during the planning process was, okay, great, I see the aggregation switches at the top of each rack. I see the, uh, the core networking structure here. Great, awesome. So what... how are these petabytes of data getting here exactly? "Oh, customers will ship us a pile of disks."

Awesome, great. "And then we plug them in over there on the bench in the rack. See, we thought about this. Stop with your impertinent questions." Cool. Just one more, um, what is the sustained throughput rate between all of those switches? And that... and when you saturate, even assuming line rate, how long does it take to fill all of those servers throughout there?

And the answer distilled down to, "Oh, no." Because it's a narrow pipe problem. How do you do this? Uh, it is... and that was not well understood. These are smart people, I'm not trying to dunk on them. Oh, yeah. Yeah. It, it's the sort of thing that's really obvious the second time. The first time, you have questions on this.

And even now, companies that were born in AWS and are exploring, "Well, what if we move this workload to our own data center? Let's go ahead and build that out," they are misled by a very strange aspect of AWS that I confess I don't fully understand myself, and I'm hoping you can shed some light. Uh, in my experience- And you can do this concurrently with basically everything you imagine.

Pick any two points in your, uh, between, in the same availability zone, in the same cluster, et cetera. You can get damn near line rate network transfer between those points. Mm-hmm. And then when I have a bunch of things doing that simultaneously, I don't see a subsequent degradation. It is still right there- Yeah

at line rate. And the only answer I have on that is witchcraft, which, m- yeah, from a certain level of ignorance, everything seems like magic. Yep. I get it. How do you do that?

Matt: Yeah. We overbuild the network. I mean, it, the, it is that simple. We build more network capacity than we need, both from a reliability perspective so that we have a lot of redundancy and resiliency.

Many devices can fail, many links can fail. We still have sufficient capacity for all of the customer traffic. And so it's, it's really is that overbuilding and building as much capacity as we possibly can to create that illusion of elasticity or that illusion that basically there is no network there.

These servers are all just directly wired together. You can send as much data as you want between them all the time. Um, and that's really been our mission for the last 15 years is if you wanna do that and you wanna do that at massive scale, you have to do it reliably, but you also have to have a cost structure so that you can affordably do that.

A lot of the reasons that there are network problems, it comes down to constraints or lack of capacity, and then people get creative in that they wanna, like, maximize the usage of the network, and you get into things like quality of service or how do I do traffic engineering or how do I move this traffic around?

If you just had more network- You don't actually need that stuff, and your network runs much, much more reliably

Corey: if I

Matt: have a more-

Corey: Well, how do you even QoS? Well, I don't know. If you have the capacity to put it all through with the same, uh, latency target- You

Matt: don't- ... you don't need it ... you don't need it. Yeah, and that, and that's been our mission for many, many years, is we really try and keep our network very simple.

I mean, that's another secret to our resiliency. We try not to do creative or fun things in the network. Keep it as simple as absolutely possible, have a lot of capacity, have a lot g- more capacity than you actually need so things can fail and you're still fine, and running a giant network becomes something you can actually achieve.

It- but the, but the key there is it's the cost structure and the ability to have that much capacity. Like, how do you actually pull that off? Mm-hmm. Uh, for us, it's we started investing in building our own hardware. And so taking control of our own hardware, we did this about 15 years ago and started this journey, meant that we could be really prescriptive about just the things we wanted in the hardware and in the software on our devices, and not take along everything else that when you buy from a vendor is gonna come along.

It's like- Well,

Corey: you'd have that section of the firmware in case you wanted to plug it into a Mellanox thing later. Exactly. So yeah, we know we're not gonna do that. Yep. Why bother, uh, taking up the RAM?

Matt: Delete that code. Yeah. It's, it's one less thing that can fail. Again, uh, few- fewer features, fewer functionality, simple, tailored to your use case specifically, uh, is, is really the key there.

And so over the last 15 years, we've, again, iteratively matured. We keep making it better and better and better with every generation, and you get to this point where today 100% of the AWS network is built with devices that we, we design ourselves. And, and the other fun secret about the way we build our network, we actually use the same switch everywhere.

Uh, so again, most you get these hierarchical networks we talked about before with core and aggregation and edge and all these other layers, but also each one of those layers would use different type of router for, for different fit for purpose. And there's good advantages and reasons to do that, but then now you have more complexity, you have more device types to manage.

They all have their own little nuances. We said, "Well, why can't we just use the basically our top of rack switch, and why don't we just use it everywhere?" And we've achieved that at this stage. It took us about 10 years to fully get to the point where we could do that, but we're bull-headed, and we just kind of continued to push forward and get to that stage.

But so now, like our entire internet network, our backbone network, runs on the same top of rack switch that sits in the top of rack switch and connects to EC2 servers.

Corey: Well, I, I have memories, uh, uh, misspent youth of driving a van with a core switch in it that cost more than the van we had rented to do this.

Yep, yep. It's like, "Well, I'm pretty sure we're not insured for this, but that's the boss's problem," certainly not mine. It, uh... And we didn't hit anything, so fortunately it was no one's problem. Yep, yep. Yeah. It, it's the, the world has changed. The, the way you address these things has changed and, and it's, it's wild.

Uh, one caveat, I want- uh, that I know I'm gonna get comments on this, otherwise if I don't say this. Like, well, yeah, you talk about economies of scale, that's why data transfer is so expensive. I, I get it, but in the context of inside of a VPC, inside of a subnet, you're getting that full magic line rate between two endpoints, assume both EC2 instances, that is free.

There is no additional charge metered to customers- Correct ... until it starts crossing other boundaries. Correct. So it's, it's not just that you've overbuilt and, uh, made it super awesome. It, it is not charged explicitly for the most common use cases where those things matter. Correct. And I wanna make sure that nuance is clear.

Correct. Because otherwise it's, well, yeah, if I was charging X dollars per gigabyte, I too would invest in making sure you could shove as many things as possible through it, but that's not what's going on here.

Matt: No. That's, that's not charged, and as availability zones get larger, again, there's the magic of elasticity that comes from AWS.

Behind the scenes, those are actual data centers filled with devices that all have to be interconnected together as we scale out these availability zones, which means more and more and more network, and that's also a driver for us to drive cost efficiency. And so we can keep it free effectively, like we don't want to charge for that because we want to get the network out of the way and let customers just move their data at whatever speed they possibly can.

Corey: And that's valuable. Uh, as long as people can predict that it's not gonna cross those chargeable boundaries, which is where it's... It's not even that it's too expensive, it's that I thought it was free and it's not, that scares people. Yes. Uh, one other bit of magic in here that I talk to folks, even folks who have networking backgrounds there, right?

You look at the typical... You, you start inspecting the traffic that's going over the wire. You look and you see the physical, uh, the back address on the physical layer, Layer 2. You look at the IP logic around Layer 3. You look at TCP and all the stuff that's being built in as those things happen. Mm-hmm.

The entire network that you see as a customer is fake. Yes. It is an emulated, virtualized imagining of a network designed to look like networks looked 25 years ago, and it is entirely living on top of what the reality actually is. And that is wildly exciting. Was it always that way?

This episode is sponsored by my own company, Duckbill.

Having trouble with your AWS bill? Perhaps it's time to renegotiate a contract with them. Maybe you're just wondering how to predict what's going on in the wide world of AWS. Well, that's where Duckbill comes in to help. Remember, you can't duck the Duckbill bill, which I am reliably informed by my business partner is absolutely not our motto.

To learn more, visit duckbillhq.com.

Matt: Not in the very, very, very first days. It, it, when EC2 originally launched, you could see the actual network. Within about a year, we very quickly realized this was going to be a bad idea, uh, and that's when we invested in building VPC. Mm-hmm. And since, since it's really, I think, since about 2007 or 2008 when we introduced VPC, everything from that point was virtualized, and the whole idea was make a very simple virtual network for customers, abstract it and separate it from the physical network.

That way your physical network can change, and you can do whatever you want on the physical world behind the scenes, and you're not really disrupting customers. You're changing the way customers perceive their network to function, and it's been very powerful for us to have that separation. Like, I like to tell people I run the real network at AWS 'cause it's the physical network.

It's the network network.

Corey: I love the fact that it just sounds like you're talking smack. It's amazing. Like, well, you know, those fake network things. I don't know.

Matt: No, and we have, we have amazing teams who build all of the network services the customers use, uh, and it's super powerful. Like, I, I work with those teams, but my teams are not tightly coupled, right?

Like, we can build our network relatively separately from the services that are delivered on top of that network to customers, and that lets us all move faster.

Corey: One thing that I have always wondered about is the reason that the internet exists is the idea of interoperable standards. Mm-hmm. And the one

Like, there were a bunch of early things that came out, like, oh, AppleTalk. Sure, great. Uh, IP- ISP, uh, what is it, IPX? ISP- Yeah ... SPX, that. Yeah. Things that I don't even remember because that, they did not win. TCP on top of IP did. It's

Matt: very durable.

Corey: But if you look at the protocol definitions and see how they are structured, what they are built for, it's the reason the internet works.

It is designed for a wide variety of network environments with a wide variety of ever-changing conditions- Yep ... in those environments. Inside of a virtualized network like you have built, a lot of the things that those services and protocols have been built for historically will not happen. You can state that with a certainty.

Like, we didn't build the ability for the simulation to ask about the nature of itself or whatever the, uh- Mm-hmm ... the challenge is. Does that mean that there is a path to a more efficient protocol for some workloads as long as it doesn't have to start dealing with the rest of the broader internet? Yeah.

How do you think about that?

Matt: We already have that. Um, we're-

Corey: You do under the hood. I know that much- Yeah ... historically. It was SRD that you were talking about- SRD, yes ... a few years back. Yes. You announced this at re:Invent, and it was great. Like, we have this amazing TCP replacement protocol- Yep ... that we use at the time to power, I think it was EFA and a lot of the higher EBS, uh- Yep.

Matt: Everything, everything behind- ... resources ... EFA runs on SRD as well as all EBS traffic, uh, within AWS regions. Yeah, so all the storage traffic.

Corey: Yes. And, and there was like, "Wow, that's really interesting. Can I learn more about it?" And the answer was basically no. And cool, so why are you telling about it? Like, honestly, it's really neat, and we're proud of it, and it do- it's how we do these things.

But outside of the context of how we run, what we do, how we see it- Yeah ... it is not something a customer would even find useful, much less make sense to even go too far down the path. Yeah, yeah. I'm talking about the other side of it. When, once you're in- inside of the virtualized network, once you have... I have an EC2 instance that's talking to another, and I wanna send a bunch of this data over- Yeah

very quickly. Yeah. Like, as the connection stands up, you start seeing TCP window scaling. Yeah. Starts slow, speeds up, et cetera. Yeah. If I have certain guarantees around this, I could see, and these are famous disasters, I could write my own protocol and put traffic over that. I've been involved with a project once that did it, and I was there- It's hard

monitoring it, I said, "Turn it off." Yeah. It's really hard. But I could see you folks putting something like that out for-

Matt: Yeah. We're gonna... We, we have. So SRD, you can go to an EC2 instance right now- Mm-hmm ... on your, your ENI or, and you can enable, basically, SRD transport under the covers. It'll still look like TCP that you're communicating with, but behind the scenes, we wrap your TCP packets in SRD, which effectively makes them extremely reliable, lower latency, higher throughput, and send the TCP bits.

And so from a TCP perspective, it just kind of looks like you're on a magic network with infinite bandwidth and no packet loss. But you can... But that way, there's no- That one

Corey: snuck completely past me. Fantastic ...

Matt: but there's no change for the customer, and that's, that's the big thing. Like, you can adopt EFA, right?

But if you are adopting EFA, there's a lot of work for you to do in your application so that you can make it EFA capable, and a lot of customers just won't do that or can't do that, and we wanted to, like, how do we make the network better for all of them? And so right now this is an opt-in feature- Yeah

that you can go turn on. Again, we're busily working behind the scenes to be, 'How do we make this the default? Like, how is this just the way it works for everyone?' And there's a bunch of technical challenges, but we're, we're chipping away at that.

Corey: Today at least, what are the workloads for which that makes sense, and which is one of those-

Matt: Almost all workloads, honestly.

Corey: Okay. Okay. Is there any exception cases? Ooh, do not use it

Matt: for

Corey: that one

Matt: thing? Or- The, the, the exception case is there's, there is a tax of, uh, maximum packets per second. Mm. So in order to do that encapsulation on the hardware layer, the peak PPS you can get is reduced slightly, and that's the reason we haven't gone and turned it on by default- Okay

for people yet. Again, we're chipping away and r- and reducing and kind of removing that bottleneck. But for the vast majority of workloads, unless you're very, very sensitive to peak PPS, uh, enabling that functionality gets you more reliability in your connectivity. And, and the real way you see this is if you're looking at, like, tail latency for your application.

You're looking at your P99, P99.9 latency. That's gon- usually gonna be dominated by some type of packet loss or some other, uh, other issue. And if you enable this with SRD based transport, mostly that disappears. And that's, that's the reason we've done it for EBS. Tail latency matters a whole lot when you're talking about storage workloads.

And so we were on a mission to be like, 'How do we drive that tail latency down?' And that's where we came up with SRD and moved all, uh, all of the EBS traffic over to SRD.

Corey: Okay, I've seen it all now. All I have to do is talk to one more AWS customer, and I see things that are new. Uh, the challenge I always ask, okay, where is this not an appropriate fit for, is because very often I will talk to folks who hear about these things.

They're like, 'Great, we're gonna go and implement that everywhere.' Great. What does your workload look like? 'Oh, it calls out to, uh, an AI inference provider, and it waits for the response, and the response comes back, and we're getting fast speeds. We're getting almost 200 tokens a second on that. It's-' I promise the network is not your bottleneck right now.

No. It's fine. This is not the place to optimize. Yeah. There are, there are better paths forward for you. Yeah. I, I, I'm excited for the day where this does become the default, with the obvious caveat then that I can see a lot of workloads that people are gonna try deploying somewhere else get disparate results as a result of this once that becomes a default, and start to wonder.

I mean, sure, the, the blame is going to accrue in directions that are not aimed at you because you're the good example in this. How come it's so crappy in our data center? Probably our terrible networking team. Which is not true. It's just this weird- Yeah ... almost under- under the hood magic.

Matt: Yeah. And, and, and really, like something like SRD and our capability to do that, it's, it's similar to our network investment and we wanna own our own devices, we wanna own our own software.

And we've done on the, on the server side for years with Nitro. We have our own hardware. We, we abstract all of the customer bits with VPCs, which means we can do interesting thing behind the scenes and we don't have to change, you know, many, many years long standards like TCP. You wanna change TCP? Good luck.

And you wanna try and change everyone's applications. The way to make meaningful change is to leave TCP alone, but just basically magically make it better by doing something behind the scenes. Embrace, extend,

Corey: and yeah.

Matt: Yeah.

Corey: Yeah. So as you talk to people about what you've been doing for the last, well, 15 years at least- Yeah

uh, what do you find is the biggest misconception that customers have about AWS's network?

Matt: About AWS's network? Most people don't know it exists.

Corey: I agree wholeheartedly. Yeah. Yes, yes. It-

Matt: Yes. And so it's, it's just the lack of understanding of the sheer scale and all of the investment and all the work that goes on behind the scenes to just make all of this function and work for customers.

And, and friends and family and everyone else, they just see it as, well, there's just, you know, some magic happens to the internet, things happen, so.

Corey: Uh, a while back, uh, in one of his iterations, Peter DeSantis, before he transitioned to his Amazon role now, was the SVP of Utility Computing. Mm-hmm. And I love the term because it, it encapsulated an experience I think lots of folks have, even if they didn't notice it.

Uh, when something is a utility, uh, like, uh, electricity or water- Mm-hmm ... you don't turn on the switch or turn on the faucet and then wonder if it's gonna work this time. It is expected. Yep. And we've all gone through that as consumers. There was a time at some point where when I go to google.com in my web browser and it doesn't work, in the early days, Google must be having an issue.

And at some point, it switched for all of us, "Oh, my Wi-Fi must be having problems." Yes, it's my problem. Yeah. Exactly. To the point where on the rare occasion, uh, I can't think of one in the last 10 years, but toward the end of that period when Google would actually be down and have an issue serving, I didn't realize that was even a possibility.

How could that happen? It, uh... And I really think the network, adding AWS to the site, is in so many ways almost a victim of its own success.

Matt: The electricity example's exactly what we have in our head. Like, we want the network to be like the light switch. You flip it on and off, you expect it to always work, and it's extremely rare and surprising when it doesn't work, and that's, that's been our mission, is make the network invisible.

Get it out of the way.

Corey: A, a question I often get from folks is, "Great, how, how did you get started?" How did I become me? Yeah. And the answer, I started off as a Unix, followed shortly by Linux- Yeah ... systems administrator. And during the, uh, the global financial crisis in 2008, suddenly salary freeze, everyone hates their job, but no one's hiring.

What do I do for the next year? Well, I started learning networking, 'cause that had always been an area- Yeah ... I hand waved over. And I did dabble as a network engineer after that. But by and large, it made me a much better systems person. Yeah. Because once I understood what was on the wire, I could reason about it.

I could understand- Yeah ... what was going on. The challenge now is that we have complexity on top of complexity on top of complexity and so on and so forth. It is dizzyingly high now. And I worry on some level that networking as a whole is no longer as central and core to modern engineering understanding of this.

There are publicly traded companies built in the cloud that do very well, and their internal networking teams, their entire professional expertise is more or less around configuring things like Transit Gateway- Yeah, yeah ... and Cloud WAN, and these are all abstractions on top of it. They don't understand that even the old days of MDI-X autosense, which was a consumer feature then came to us, like, great, when I plug the ethernet cable in, it's not lighting up.

Why? Oh, that's a s- that's a rollover cable. Yeah, you got to cross that over, yeah. You need a straight through. It's... Wait, you mean there's different standards for the, for the wiring standard? I have a mug that just has the colors of the ethernet B spec. Mm-hmm. And every once in a while, someone sees it, they're like, "Yes!"

Like, okay, I found my people. It's great. But now I sound like an old man talking about the Great War- Yeah ... once upon a time. Where does the next generation of all of this come from?

Matt: Yeah, we're, uh... It's something we grapple with. And, like, again, we've been very successful such that most of our customers don't have to deal with the vagaries of networking, and you like it.

Yeah. Most people look at that, and they're like, "That's terrible. I don't wanna do that." I

Corey: like it, and I'm far enough away from it now that I can wax nostalgic about this. If I were on the phone this morning with Cisco TAC because of a weird routing issue- Man ... I would be considerably less charitable. Yes.

But that's 10 years in my past.

Matt: It... For, for us, it's really grow your own. Uh, and so, like, I mean, I joined Amazon as a very junior engineer back in 2008, and my initial job was kind of doing monitoring of the network and, you know, seeing things breaking and learning to fix it. And I've stuck around long enough and learned enough that I managed to grow into this position, uh, and it's kind of the same path we have with all of our new hires.

We're, we're hiring lots of junior engineers, either from... We hire lots of people from our data centers, actually, some amazingly talented people who were there. I was just talking to someone who was visiting one of our data centers, and it's someone who was working at a gas station three years ago, and they were just wowed with what they knew.

And I'm like, "We should go hire that guy for the networking team," 'cause they actually touch the physical equipment. They understand how this stuff works, and they, they often become our best engineers in the future. Um, but it is a deep investment in us in terms of, like, hiring younger talent, junior talent, and, like, going through the process and training and growing them and teaching them how we build networks 'cause there's...

At the end of the day, there's no accelerator for this. It is increasingly complicated, as you said. It's much, much more complicated than it was 15 years ago when I was starting. But you still need to know the same foundational things that I had to know 15 years ago, but you also have to know this increasing mountain of new stuff over the top, and that only comes from just- Learning the basics, learning the practical bits, and then slowly adding the knowledge on top with experience

Corey: I mean, 'cause you know that.

Customers don't have to. I, I wanna highlight just how rare that is. It, it seems that the idea of developing talent internally- Yeah ... has almost gone out of fashion in the industry. It's, "Oh, no, this person doesn't have experience. We can't hire them for the role 'cause it'll take six months to get them up to speed."

Yeah, I get it. Nine months later, the role's still open. Yeah. So what are we really doing here, friends? Yeah. The idea of teaching people and, and arming them to do this, 'cause there was a time where every company that wanted to be on the internet, which let's face it, is probably most of them- Mm-hmm ... needed to have a networking- Yeah

if not person, team. Team,

Matt: yeah.

Corey: Where it was you need people able to handle this stuff. It gets really weird really fast. Yeah. And now you don't need that. I mean, I, I went through a parallel evolution running email systems. Yep. Now most companies do not run their own mail servers. Nor

Matt: should

Corey: you. Thank God.

Yeah. Yeah, n- and networking is, is still continuing to evolve, and I'm not naive to th- enough to think that we are at the end of history, and, like, in our final form. What's the next step? What's the next phase? Where do you see this going?

Matt: Where do I see this going? Uh, it's- It's really more of the same. I mean, the, there's so much, there's been an acceleration in investment in networking because of AI, which is exciting as someone who does networking.

The, there's bandwidth demands are higher, there's more capacity to build, there's more things to connect. And so I've seen more energy go into networking in the last three years than I had in the previous 10 years. And so there's just accelerated in investments both at the silicon and the hardware level, but as well as the software, the protocols, everything else.

It, it feels like networking is exciting again i- in a way that maybe we made it fairly boring by the time you got to 2019. And that it's just working. I will say

Corey: when, in my misspent youth, there were times I wished it had been a lot less exciting- Yes ... some nights. Yes. But yes.

Matt: Yes. Um, but yes, uh, and so I, I think it's just gonna be more of that.

Like, we're gonna continue to find... The world has an insatiable appetite for bandwidth. Yes. A- and I think we're only scratching the surface of that. And if we were able to 100X the internet and network capacity, it wouldn't take long for applications to find uses for that. And I think there's still a lot of things that are not done, or they're done in ineffective ways because of either, uh, not enough capacity of network and/or, like, costs of networking.

And so the more you reduce that and eliminate that, I think you just open up new opportunities and innovation. And so I think it's just gonna be this constant march for more and more and more and more capacity. Well,

Corey: we, we've seen that in real time. The, the thing that drove that home to me more than almost anything else was the day that GDPR took effect because a lot of US-based websites, uh, suddenly were not fully compliant, so they turned off all the tracking and all the rest.

Yeah. And suddenly the internet was blindingly fast. Super

Matt: fast.

Corey: It was like, wow, you're loading an article, and, like, the total load was something like- Yeah ... 50K. Yeah. It was wild to me. And you look at the same article with all the other stuff turned on, it's why is loading that 50K article taking 25 megabytes- Yeah

of nonsense? It-- As soon as you have the capacity, people find ways to fill it. Or just- I'm not opining on whether it's all useful ... or

Matt: just look at streaming video, for example, right? Yeah. Like, I mean, you, you can get very good quality with the, the latest codecs, but it's not the best quality. Like real broadcast level quality video is, is like an order of magnitude or two higher bandwidth than what's streamed to you even if you're getting like the 4K Ultra whatever on your Fire TV.

And so that- the reason is bandwidth. There's always need for more bandwidth, and there always will be need for more bandwidth. And so the more we can do to- Mm ... to make it cheaper and more reliable and just more prevalent, we're only gonna be benefiting our customers.

Corey: No, like you have the, the news anchor trucks that show up at the scene of something going on.

Yeah. They have the big satellite dish on the roof. It's like, do you think that's because they haven't quite figured out that you can tether to your cell phone? There are some very dedicated, very concerning bandwidth requirements around this. Yes. And it's one of those areas, networking more so than many, but complexity passes everywhere, and it just slips below the baseline surface of awareness.

Thank you so much for being as generous with your time as you are and explaining how some of this, this magic all works. If someone's watching this or listening to this, depending upon their point of view, and they wanna learn more about how networks work and how infrastructure happens, where should they start?

Where should they begin learning about this, this secret thing that still drives our entire world?

Matt: Yeah. It, it's, it's again, it's going back to the, the raw material. And then there's a book written in I think the '70s or '80s. It's The TCP/IP Illustrated Book by Stevens. It's where I'd suggest you go. Like that's where I send anyone who's wants to get into networking and learn something new 'cause those same fundamentals like you talked about of all those standards They still matter, and the, the entire world is built on y- you know, versions of them that have, have evolved over time.

But that fundamental skill set will, will take you a long way, and you have to have it to really understand how these things work.

Corey: And we'll of course put a link to that. I love that book myself.

It's a great book. Though I will say the, some of the bandwidth references in there of like, "Oh, that's a really fast 10BASE-T connection there," that might not have aged super well.

Just

Matt: shows you how far we've come. But yeah, I, I just had someone, uh, actually one of our data techs this morning was messaging me asking me, "How do I learn about networking, Matt?" That's exactly what I sent him. Yeah. "Go get this on Amazon."

Corey: I have no idea if it's still up, and I'll check it. If not, my apologies, but there was a great site, WarriorsofTheDotNet.

Mm-hmm. Was always a video that explained in a like five or six minute cartoon approach of how packet switching routed networks work in an extraordinarily accessible way. I hope that's still out there. Yeah. That was always a fantastic primer. Like, yeah, some of their terminology choices you can argue with, and like, well, there's a lot more to it there.

Yeah, but you're going from zero to one, and you're giving people a chance to see what's next. Yes. And, huh, that's curious. I wanna dig deeper. The thing that actually tipped me over the edge was subnet masks. Okay, it's this weird number string, and all I know is that when I get it wrong, some things work and other things don't, and I don't understand it.

Maybe it's time I stop hand-waving over it.

Matt: Yes.

Corey: And for my sins, I learned how networks work. Yeah. 'Cause I didn't know how networks work, and I didn't believe it was possible. And now that I know better, I do not believe that networks work, and I'm astounded that it's possible. It feels like it works in spite of itself, but clearly it does.

It

Matt: is tremendously complicated, but it's, uh, it gives us a lot of pleasure to, to make these things simpler for everyone.

Corey: Thank you so much. Matt Rader, VP of Global Networking at AWS. I'm Corey Quinn. Stick around.

 View Full Transcript  Hide Full Transcript

## You might also like

[More Podcast Episodes](https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/)

### [AI Can Do the Work, But Should It? with Adam Larsen](https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/ai-can-do-the-work-but-should-it-with-adam-larsen/)

Screaming in the Cloud

08.27.2026

30 Minutes

[Play Episode](https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/ai-can-do-the-work-but-should-it-with-adam-larsen/)

### [When AI Starts Writing the Pull Requests with Madelyn Olson](https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/when-ai-starts-writing-the-pull-requests-with-madelyn-olson/)

Screaming in the Cloud

06.25.2026

30 Minutes

[Play Episode](https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/when-ai-starts-writing-the-pull-requests-with-madelyn-olson/)

### [The Appalachian Cloud Trail: Hiking, Cloud Economics, and Finding Perspective](https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/the-appalachian-cloud-trail-hiking-cloud-economics-and-finding-perspective/)

Screaming in the Cloud

06.11.2026

33 Minutes

[Play Episode](https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/the-appalachian-cloud-trail-hiking-cloud-economics-and-finding-perspective/)
