WEBVTT

00:00:00.000 --> 00:00:04.000
This week, containerd disclosed a stack of CRI

00:00:04.000 --> 00:00:06.820
plugin vulnerabilities in the runtime layer,

00:00:07.080 --> 00:00:10.800
a huge number of Kubernetes nodes trust to start

00:00:10.800 --> 00:00:14.759
your containers. Datadog ran a PostgreSQL

00:00:14.759 --> 00:00:17.660
gameday and learned their database could fail over

00:00:17.660 --> 00:00:21.300
just fine. It just couldn't do it safely. AWS

00:00:21.300 --> 00:00:25.039
DevOps Agent and Datadog's MCP Server are both

00:00:25.039 --> 00:00:28.660
now generally available. And the new AWS integration

00:00:28.660 --> 00:00:33.179
means AI incident response just graduated from

00:00:33.179 --> 00:00:37.500
demo to on-call rotation. And EKS will now route

00:00:37.500 --> 00:00:40.060
your Kubernetes control plane's outbound traffic

00:00:40.060 --> 00:00:44.020
through your own VPC, which is great, right up

00:00:44.020 --> 00:00:47.439
until a stale route table quietly kills your

00:00:47.439 --> 00:00:50.280
admission webhooks. Put those together and the

00:00:50.280 --> 00:00:53.259
shape of the episode is pretty clear. The control

00:00:53.259 --> 00:00:56.859
plane keeps getting wider. Runtimes. Databases.

00:00:57.079 --> 00:01:00.899
Incident agents. API-server egress. credentials,

00:01:01.039 --> 00:01:04.219
even the cloud console. One by one, they are

00:01:04.219 --> 00:01:07.140
all sliding into your production blast radius.

00:01:07.439 --> 00:01:10.159
And here's the part that matters. Your users

00:01:10.159 --> 00:01:13.700
don't care which control plane failed. They just

00:01:13.700 --> 00:01:16.799
feel the wait. I'm Brian Teller from Teller's

00:01:16.799 --> 00:01:36.859
Tech, and this is Ship It Weekly. Welcome back

00:01:36.859 --> 00:01:39.959
to Ship It Weekly, the show about the DevOps,

00:01:40.280 --> 00:01:45.000
SRE, cloud, platform, and security stories that

00:01:45.000 --> 00:01:47.980
actually matter when you are the person who has

00:01:47.980 --> 00:01:51.219
to keep the thing running at 3 a.m. If you are

00:01:51.219 --> 00:01:54.540
new here, follow or subscribe wherever you are

00:01:54.540 --> 00:01:57.519
watching or listening. And if you want the weekly

00:01:57.519 --> 00:02:01.219
story list and source links, check out OnCallBrief.com

00:02:01.219 --> 00:02:05.069
For past episodes, full show notes, and

00:02:05.069 --> 00:02:08.069
more from the show, head over to ShipItWeekly.fm

00:02:08.069 --> 00:02:12.469
We open with the containerd CRI plugin vulnerabilities,

00:02:13.009 --> 00:02:16.650
because your node runtime is the trust boundary

00:02:16.650 --> 00:02:20.530
underneath the trust boundary. Then, Datadog's

00:02:20.530 --> 00:02:24.430
PostgreSQL HA gameday, where the scary discovery

00:02:24.430 --> 00:02:28.530
wasn't that failover was hard, it was that failover

00:02:28.530 --> 00:02:32.800
was unsafe. After that, AWS DevOps Agent and

00:02:32.800 --> 00:02:36.639
Datadog MCP Server going GA. And what it means

00:02:36.639 --> 00:02:40.080
when an AI agent gets a seat near your control

00:02:40.080 --> 00:02:43.680
plane. Then, EKS customer-routed control-plane

00:02:43.680 --> 00:02:47.599
egress. Because your API server is now part of

00:02:47.599 --> 00:02:50.180
your network perimeter, whether you plan for

00:02:50.180 --> 00:02:53.680
it or not. In the lightning round, GitHub Credential

00:02:53.680 --> 00:02:57.840
Revocation. AWS Console Private Access. Vercel

00:02:57.840 --> 00:03:01.780
Connect, and S3 annotations. And we close with

00:03:01.780 --> 00:03:05.620
Marc Brooker on waiting, on why your customers

00:03:05.620 --> 00:03:08.860
live in the tail of your latency distribution,

00:03:09.240 --> 00:03:12.360
even when your dashboards swear everything's

00:03:12.360 --> 00:03:20.039
fine. Let's get into it. First up, containerd

00:03:20.039 --> 00:03:23.780
has a batch of CRI plugin vulnerabilities. And

00:03:23.780 --> 00:03:26.949
if you run Kubernetes, this one's yours. AWS

00:03:26.949 --> 00:03:30.030
published a security bulletin spanning

00:03:30.030 --> 00:03:35.030
containerd branches 1.7 through 2.3. And the list is

00:03:35.030 --> 00:03:38.490
not a fun read. Image cache poisoning through

00:03:38.490 --> 00:03:41.770
checkpoint image references. Host command execution.

00:03:42.490 --> 00:03:45.949
through unsanitized image labels, CDI annotation

00:03:45.949 --> 00:03:49.710
handling that can inject devices and host mounts,

00:03:50.030 --> 00:03:53.050
host file reads through symlinked container

00:03:53.050 --> 00:03:57.150
log paths during checkpoint restore, and a denial

00:03:57.150 --> 00:04:01.110
of service from crafted images that exhaust memory.

00:04:01.370 --> 00:04:05.319
So not exactly a relaxing Patch Tuesday. Here's

00:04:05.319 --> 00:04:08.840
why it matters. containerd sits underneath an

00:04:08.840 --> 00:04:12.099
enormous number of clusters, and we spend almost

00:04:12.099 --> 00:04:15.219
all of our security attention on the layers above

00:04:15.219 --> 00:04:18.800
it. Pod specs, admission control, image scanning,

00:04:19.040 --> 00:04:23.220
RBAC, network policy, runtime classes, all the

00:04:23.220 --> 00:04:26.180
familiar Kubernetes machinery. But eventually,

00:04:26.519 --> 00:04:29.680
something has to actually pull the image, unpack

00:04:29.680 --> 00:04:34.100
it, restore it, wire up devices. handle the logs,

00:04:34.180 --> 00:04:37.699
and start the container. That layer is a trust

00:04:37.699 --> 00:04:41.040
boundary too. And in some ways, it's the more

00:04:41.040 --> 00:04:44.019
dangerous one. Because by the time a workload

00:04:44.019 --> 00:04:47.399
reaches the runtime, the rest of the system has

00:04:47.399 --> 00:04:50.800
already decided this thing is allowed to exist.

00:04:51.160 --> 00:04:54.240
That's why the boring fields turn out to matter.

00:04:54.459 --> 00:04:57.939
Labels, annotations, checkpoint and restore paths,

00:04:58.319 --> 00:05:02.459
CDI, log paths, every field. that feels like

00:05:02.459 --> 00:05:05.819
plumbing can become an input to privileged behavior

00:05:05.819 --> 00:05:09.920
on the node. A malicious image isn't just application

00:05:09.920 --> 00:05:14.079
code. It's metadata, build time weirdness, and

00:05:14.079 --> 00:05:17.240
a set of assumptions the runtime makes about

00:05:17.240 --> 00:05:20.800
what it can trust. The takeaway is direct. Patch

00:05:20.800 --> 00:05:23.980
containerd. Check your managed node groups,

00:05:24.240 --> 00:05:28.519
your self-managed nodes, your AMIs, your Bottlerocket

00:05:28.519 --> 00:05:31.899
versions, your distro packages, anything

00:05:31.899 --> 00:05:35.319
that controls the runtime. If you lean on checkpoint

00:05:35.319 --> 00:05:39.899
restore, CDI devices, or GPU workloads, look

00:05:39.899 --> 00:05:43.079
harder. And if you don't use any of that, don't

00:05:43.079 --> 00:05:46.699
relax. At least one of these issues doesn't need

00:05:46.699 --> 00:05:49.620
checkpoint and restore turned on at all. Your

00:05:49.620 --> 00:05:53.379
node runtime is the trust boundary under the

00:05:53.379 --> 00:05:56.139
trust boundary. Stop treating it like invisible

00:05:56.139 --> 00:06:03.699
plumbing. Second story. Datadog published a genuinely

00:06:03.699 --> 00:06:07.160
good engineering write-up on running high availability

00:06:07.160 --> 00:06:11.500
PostgreSQL on Kubernetes. And it's one of those

00:06:11.500 --> 00:06:14.839
pieces that sounds boring until the real problem

00:06:14.839 --> 00:06:17.980
comes into focus. The problem wasn't that the

00:06:17.980 --> 00:06:20.779
database couldn't fail over. It was that it couldn't

00:06:20.779 --> 00:06:24.189
fail over safely. During a gameday, Datadog

00:06:24.189 --> 00:06:27.269
simulated a zonal failure. That added network

00:06:27.269 --> 00:06:30.930
latency, replication lag grew, and when the cluster

00:06:30.930 --> 00:06:34.209
needed a new primary, Patroni couldn't safely

00:06:34.209 --> 00:06:37.410
promote a standby without risking data loss.

00:06:37.730 --> 00:06:40.930
So the system got stuck in the worst possible

00:06:40.930 --> 00:06:44.889
spot. The old primary was unhealthy. The standbys

00:06:44.889 --> 00:06:47.550
weren't safe to promote, and the only correct

00:06:47.550 --> 00:06:50.199
move was to wait. That's the kind of failure

00:06:50.199 --> 00:06:53.680
mode that ages every SRE in the room about three

00:06:53.680 --> 00:06:56.680
years. Because on paper, you have everything.

00:06:56.939 --> 00:07:00.519
Multiple nodes, standbys, Kubernetes, automation,

00:07:01.040 --> 00:07:04.079
failover machinery. And then the actual failure

00:07:04.079 --> 00:07:07.699
arrives and the system says, yes, but not safely.

00:07:07.920 --> 00:07:10.839
Which, honestly, is the right answer. Promoting

00:07:10.839 --> 00:07:14.120
a stale standby might hand you a writable primary

00:07:14.120 --> 00:07:17.569
faster. But if it costs you data loss, split

00:07:17.569 --> 00:07:21.110
brain, or a broken consistency guarantee, you

00:07:21.110 --> 00:07:23.829
haven't fixed the outage. You've traded it for

00:07:23.829 --> 00:07:26.470
a corruption event. That's not an improvement.

00:07:26.769 --> 00:07:29.910
It's just a different postmortem. The real lesson

00:07:29.910 --> 00:07:33.009
is that HA isn't only about whether the service

00:07:33.009 --> 00:07:35.829
comes back. It's about whether the recovery path

00:07:35.829 --> 00:07:39.329
itself is safe. Can you fail over without losing

00:07:39.329 --> 00:07:42.509
writes? Can you prove which standby is safe to

00:07:42.509 --> 00:07:45.060
promote? Can your automation tell the difference

00:07:45.060 --> 00:07:47.959
between available and correct? And does your

00:07:47.959 --> 00:07:51.139
whole team agree on which one it should prefer

00:07:51.139 --> 00:07:54.899
before the incident call is on fire? Datadog's

00:07:54.899 --> 00:07:57.360
answer was to move toward synchronous replication

00:07:57.360 --> 00:08:01.300
and stronger Patroni guardrails. So a promoted

00:08:01.300 --> 00:08:05.019
standby is guaranteed to have the writes it needs.

00:08:05.300 --> 00:08:07.899
And that's the part that's worth copying. They

00:08:07.899 --> 00:08:11.040
didn't just ask how to recover faster. They asked

00:08:11.040 --> 00:08:14.480
how to recover safely. So test your database

00:08:14.480 --> 00:08:18.079
HA against real constraints, not the easy ones.

00:08:18.360 --> 00:08:22.019
Ask what happens under replication lag. Ask what

00:08:22.019 --> 00:08:25.259
happens during a zone failure. Ask what happens

00:08:25.259 --> 00:08:28.579
when the network is slow instead of cleanly dead.

00:08:28.860 --> 00:08:32.039
Ask what happens when every standby is behind.

00:08:32.399 --> 00:08:35.419
And ask whether your automation prefers safety

00:08:35.419 --> 00:08:38.600
or availability. And whether everyone actually

00:08:38.600 --> 00:08:41.779
agrees with that choice. Because failover is

00:08:41.779 --> 00:08:44.860
useless, if the only safe option is waiting.

00:08:45.039 --> 00:08:52.919
But unsafe failover can be a lot worse. Third

00:08:52.919 --> 00:08:56.679
story, AWS DevOps Agent is now generally available

00:08:56.679 --> 00:09:00.820
and Datadog's MCP Server is GA as a standard

00:09:00.820 --> 00:09:04.379
way for AI agents to reach Datadog monitoring

00:09:04.379 --> 00:09:07.379
data. This is one of those announcements. where

00:09:07.379 --> 00:09:10.019
the slide says autonomous incident resolution

00:09:10.019 --> 00:09:13.919
and the operator says, cool, but what exactly

00:09:13.919 --> 00:09:17.720
is it allowed to touch? The idea is solid. AWS

00:09:17.720 --> 00:09:21.820
DevOps Agent can work through Datadog MCP Server

00:09:21.820 --> 00:09:26.179
to investigate an incident across logs, metrics,

00:09:26.419 --> 00:09:30.340
traces, deployment events, and AWS infrastructure

00:09:30.340 --> 00:09:33.980
context. Instead of one engineer bouncing between

00:09:33.980 --> 00:09:38.679
CloudWatch, Datadog, deploy history, traces, dashboards,

00:09:38.679 --> 00:09:42.700
and Slack, the agent correlates the signals and

00:09:42.700 --> 00:09:45.779
helps push the incident forward and nobody wants

00:09:45.779 --> 00:09:48.620
to spend the first 30 minutes of an outage doing

00:09:48.620 --> 00:09:51.879
browser-tab archaeology if an agent can gather

00:09:51.879 --> 00:09:55.460
context, summarize what changed, flag a suspicious

00:09:55.460 --> 00:09:59.259
deploy and propose likely causes that's real

00:09:59.259 --> 00:10:02.820
time saved but this is also the moment AI incident

00:10:02.820 --> 00:10:06.549
response stops being a chatbot and becomes a

00:10:06.549 --> 00:10:10.289
production workflow. It's an agent reading operational

00:10:10.289 --> 00:10:13.549
telemetry, interpreting signals, recommending

00:10:13.549 --> 00:10:17.909
fixes, and potentially wired into Slack, PagerDuty,

00:10:18.049 --> 00:10:21.669
ServiceNow, your code, your deploys, and your

00:10:21.669 --> 00:10:24.549
runbooks. That puts it right next to the control

00:10:24.549 --> 00:10:27.870
plane. And once something sits next to the control

00:10:27.870 --> 00:10:31.029
plane, the question stops being, is it smart?

00:10:31.210 --> 00:10:34.840
And becomes, what authority does it have? Can

00:10:34.840 --> 00:10:38.240
it only read? Can it write? Can it open tickets?

00:10:38.539 --> 00:10:42.200
Trigger automation? Roll back a deploy? Restart

00:10:42.200 --> 00:10:46.679
a service? Change config? Page a human at 4 a.m.?

00:10:46.940 --> 00:10:51.120
Can it make things worse quickly and very confidently?

00:10:51.460 --> 00:10:54.919
That last one is the whole game. Incident response

00:10:54.919 --> 00:10:58.840
isn't about speed. It's about safe speed. So

00:10:58.840 --> 00:11:02.659
treat AI incident tooling like any other production

00:11:02.659 --> 00:11:05.759
automation. Give it the least privilege that

00:11:05.759 --> 00:11:09.059
still leaves it useful. Log what it sees and

00:11:09.059 --> 00:11:11.860
what it does. Make the human approval boundary

00:11:11.860 --> 00:11:15.879
impossible to miss. And draw a hard line between

00:11:15.879 --> 00:11:19.299
what it can recommend and what it can execute.

00:11:19.600 --> 00:11:23.279
Have rollback rules. Know what happens when it's

00:11:23.279 --> 00:11:26.419
wrong. And don't grade it only on time to answer.

00:11:26.580 --> 00:11:29.480
Grade it on whether the answer was safe, auditable,

00:11:29.620 --> 00:11:33.100
and actually useful under pressure. AI incident

00:11:33.100 --> 00:11:36.679
response is moving from demo to production. That's

00:11:36.679 --> 00:11:44.080
exciting. Production just needs guardrails. Fourth

00:11:44.080 --> 00:11:48.120
story. Amazon EKS now supports customer-routed

00:11:48.120 --> 00:11:51.940
control-plane egress. That's a very AWS phrase.

00:11:52.240 --> 00:11:55.000
So here's the human version. The Kubernetes API

00:11:55.000 --> 00:11:58.940
server sometimes needs to call outward to admission

00:11:58.940 --> 00:12:02.940
webhooks, OIDC providers, aggregated API servers,

00:12:03.159 --> 00:12:05.940
other endpoints that you control. Historically,

00:12:06.019 --> 00:12:09.639
that outbound traffic took AWS managed egress

00:12:09.639 --> 00:12:12.679
paths. Now you can route it through your own

00:12:12.679 --> 00:12:16.299
VPC, which hands platform teams control over

00:12:16.299 --> 00:12:19.700
routing, inspection, firewalls, NAT, private

00:12:19.700 --> 00:12:22.419
connectivity, and compliance boundaries. For

00:12:22.419 --> 00:12:25.519
regulated environments, that's a real win. It

00:12:25.519 --> 00:12:28.320
also makes the control plane feel a lot more

00:12:28.320 --> 00:12:30.899
like part of your network. which of course it

00:12:30.899 --> 00:12:34.000
always was. The difference is that now you own

00:12:34.000 --> 00:12:37.559
the outbound path and AWS is blunt about what

00:12:37.559 --> 00:12:40.379
that ownership means. In customer routed mode,

00:12:40.639 --> 00:12:43.639
you are responsible for making sure the control

00:12:43.639 --> 00:12:46.700
plane can reach the endpoints it needs. Wrong

00:12:46.700 --> 00:12:50.100
route table, too-tight security group, a NACL

00:12:50.100 --> 00:12:52.940
that blocks the wrong thing, a broken firewall

00:12:52.940 --> 00:12:55.980
hop, and control plane operations start failing.

00:12:56.120 --> 00:13:00.139
That includes admission webhook calls, and OIDC

00:13:00.139 --> 00:13:03.700
authentication. So yes, great feature. But it

00:13:03.700 --> 00:13:06.580
isn't a checkbox. It's a failure mode change.

00:13:06.919 --> 00:13:10.480
If your API server can't reach an admission webhook,

00:13:10.539 --> 00:13:14.460
do pod creates fail? Do deploys hang? Does authentication

00:13:14.460 --> 00:13:18.039
break? Does your incident response now depend

00:13:18.039 --> 00:13:21.779
on a firewall path some other team owns? And

00:13:21.779 --> 00:13:25.019
do you have a metric? a test, and a name on the

00:13:25.019 --> 00:13:28.120
pager for when it breaks? This is a feature you

00:13:28.120 --> 00:13:31.100
bring to a design review. Not because it's risky,

00:13:31.259 --> 00:13:34.580
but because it's powerful. Map the traffic. Map

00:13:34.580 --> 00:13:38.440
the dependencies. Test the webhooks. Test OIDC.

00:13:38.580 --> 00:13:41.679
Test the failure modes. Make the routing visible.

00:13:41.899 --> 00:13:44.679
And write the runbook before the control plane

00:13:44.679 --> 00:13:47.820
starts failing in creative ways. The Kubernetes

00:13:47.820 --> 00:13:50.659
control plane is becoming part of your network

00:13:50.659 --> 00:14:00.740
perimeter. Treat it like one. Quick lightning

00:14:00.740 --> 00:14:03.980
round. First, GitHub added self-service credential

00:14:03.980 --> 00:14:07.059
revocation for incident response. Enterprise

00:14:07.059 --> 00:14:11.159
owners now get a break-glass capability to revoke

00:14:11.159 --> 00:14:14.179
a compromised user's credentials in one move.

00:14:14.399 --> 00:14:17.200
This matters because credential cleanup should

00:14:17.200 --> 00:14:20.200
never be a scavenger hunt. You do not want to

00:14:20.200 --> 00:14:22.639
be hand-hunting through SSO authorizations,

00:14:22.899 --> 00:14:27.149
personal access tokens, SSH keys, and OAuth grants

00:14:27.149 --> 00:14:30.490
while everyone argues in Slack. Revocation is

00:14:30.490 --> 00:14:33.169
incident response infrastructure. Know who can

00:14:33.169 --> 00:14:35.889
trigger it, know what it kills, know what it

00:14:35.889 --> 00:14:39.169
logs, and put it in the compromised-account runbook.

00:14:39.370 --> 00:14:42.429
Second, AWS Management Console private access

00:14:42.429 --> 00:14:45.529
now works without internet connectivity. Console

00:14:45.529 --> 00:14:48.490
traffic for supported services can flow over

00:14:48.490 --> 00:14:51.549
VPC endpoints instead of the public internet.

00:14:51.899 --> 00:14:54.460
It's a strong story for regulated environments.

00:14:54.899 --> 00:14:57.399
Even the console is getting pulled behind private

00:14:57.399 --> 00:15:00.259
network boundaries. The lesson? Console access

00:15:00.259 --> 00:15:03.179
is part of your control plane too. And private

00:15:03.179 --> 00:15:06.779
link, endpoint policies, and known-account restrictions

00:15:06.779 --> 00:15:10.360
are becoming cloud operations, not just app networking.

00:15:10.879 --> 00:15:13.960
Third, Vercel shipped Vercel Connect. And the

00:15:13.960 --> 00:15:17.179
idea worth catching is runtime credential exchange.

00:15:17.659 --> 00:15:20.419
Instead of stashing a long-lived provider token

00:15:20.419 --> 00:15:23.799
for an agent, The app proves its identity and

00:15:23.799 --> 00:15:26.500
gets a short-lived task-scoped credential.

00:15:26.799 --> 00:15:28.980
That's the pattern that we've been tracking for

00:15:28.980 --> 00:15:31.720
weeks. Agent credentials moving from store this

00:15:31.720 --> 00:15:35.679
token forever to prove who you are and get scoped

00:15:35.679 --> 00:15:38.600
access when you need it. Short-lived credentials

00:15:38.600 --> 00:15:41.720
don't solve every agent security problem, but

00:15:41.720 --> 00:15:44.360
they beat long-lived secrets sitting around

00:15:44.360 --> 00:15:47.879
waiting to become next quarter's incident. Fourth,

00:15:48.039 --> 00:15:52.009
Amazon S3 annotations are here. mutable, queryable

00:15:52.009 --> 00:15:56.149
context attached directly to S3 objects. Sounds

00:15:56.149 --> 00:15:59.909
dull, but object metadata has driven a lot of

00:15:59.909 --> 00:16:02.850
awkward platform design over the years. Side

00:16:02.850 --> 00:16:06.470
tables, DynamoDB metadata stores, Lambda sync

00:16:06.470 --> 00:16:10.490
jobs, custom catalogs, and constant drift between

00:16:10.490 --> 00:16:14.029
the object and whatever's describing it. If annotations

00:16:14.029 --> 00:16:16.720
shrink that glue layer, That's worth watching.

00:16:16.820 --> 00:16:20.259
Object metadata is quietly becoming a first-class

00:16:20.259 --> 00:16:24.000
platform layer, especially for data, AI, search,

00:16:24.200 --> 00:16:27.799
and agent workflows that need to know what an

00:16:27.799 --> 00:16:38.779
object is, not just where it lives. The human

00:16:38.779 --> 00:16:42.340
closer this week comes from a Marc Brooker post

00:16:42.340 --> 00:16:47.370
about waiting, latency, MTTR, and why averages

00:16:47.370 --> 00:16:50.990
can lie. The point is that your users don't experience

00:16:50.990 --> 00:16:54.070
your averages the way your dashboards report

00:16:54.070 --> 00:16:57.570
them. You measure mean latency, mean time to

00:16:57.570 --> 00:17:00.529
recovery, average outage duration. But people

00:17:00.529 --> 00:17:03.590
are far more likely to land in the long waits

00:17:03.590 --> 00:17:07.289
simply because long waits take up more of the

00:17:07.289 --> 00:17:10.970
time. That's the inspection paradox. A 10-minute

00:17:10.970 --> 00:17:15.019
outage catches a few users. A 10-hour outage

00:17:15.019 --> 00:17:18.000
catches a lot of them. Your incident tracker

00:17:18.000 --> 00:17:22.000
counts both as one outage. Your dashboard says

00:17:22.000 --> 00:17:26.279
MTTR looks fine. Your users say they spent all

00:17:26.279 --> 00:17:29.140
morning waiting. Both are true. And that's the

00:17:29.140 --> 00:17:31.460
whole episode, really. When the system breaks,

00:17:31.740 --> 00:17:34.460
nobody experiences your architecture diagram.

00:17:34.759 --> 00:17:37.859
They experience waiting. Waiting for a request.

00:17:38.160 --> 00:17:40.890
Waiting for recovery. Waiting for a credential

00:17:40.890 --> 00:17:44.150
to get revoked. Waiting for a deploy to stop

00:17:44.150 --> 00:17:46.950
failing. Waiting for the control plane to come

00:17:46.950 --> 00:17:50.190
back. Waiting for someone to find the right context.

00:17:50.549 --> 00:17:53.930
So here's the takeaway. Don't only measure the

00:17:53.930 --> 00:17:56.990
system from the server side. Measure it from

00:17:56.990 --> 00:18:00.650
the waiting side. Because your users don't live

00:18:00.650 --> 00:18:04.309
in your average. They live in the tail. And the

00:18:04.309 --> 00:18:07.789
tail is usually where the real reliability story

00:18:07.789 --> 00:18:10.799
is hiding. That's it for this week of Ship It

00:18:10.799 --> 00:18:13.519
Weekly. We covered containerd runtime risk,

00:18:13.799 --> 00:18:17.000
Postgres failover safety, AI incident response,

00:18:17.460 --> 00:18:20.839
EKS control-plane egress, and why your users

00:18:20.839 --> 00:18:23.819
feel the wait more than your dashboards show.

00:18:24.039 --> 00:18:27.019
If this episode was useful, follow or subscribe

00:18:27.019 --> 00:18:29.960
wherever you are watching or listening. If you're

00:18:29.960 --> 00:18:32.940
on YouTube, hit subscribe. If you're in a podcast

00:18:32.940 --> 00:18:36.250
app, follow the show there. And if you know someone

00:18:36.250 --> 00:18:39.109
wrestling with Kubernetes runtime security, database

00:18:39.109 --> 00:18:42.750
failover, AI incident response, or platform control

00:18:42.750 --> 00:18:45.930
planes, send them this one. It genuinely helps

00:18:45.930 --> 00:18:48.549
the show grow, and it helps me keep making this

00:18:48.549 --> 00:18:51.549
for people who actually live with these systems.

00:18:51.849 --> 00:18:54.430
You can find the weekly brief at OnCallBrief.com

00:18:54.430 --> 00:18:57.549
and the full show notes, links, and past

00:18:57.549 --> 00:19:01.230
episodes at ShipItWeekly.fm. I'm Brian Teller

00:19:01.230 --> 00:19:03.670
from Teller's Tech. Thanks for listening. And

00:19:03.670 --> 00:19:06.140
remember, your dashboards measure the average.

00:19:06.339 --> 00:19:08.220
Your users feel the wait.
