Skip to main content

Command Palette

Search for a command to run...

šŸ‘‹ Everything about EKS & AI Infrastructure Newsletter "#87" ā˜ļøā¤šŸ‘Øā€šŸ’»

Isolation and efficiency are turning out to be the same problem — plus a look at what actually breaks GPU inference at scale.

Updated
•19 min read•View as Markdown
šŸ‘‹ Everything about EKS & AI Infrastructure Newsletter "#87" ā˜ļøā¤šŸ‘Øā€šŸ’»
A

I’m a Solution Architect at Lauren, AWS UG Vadodara Co-Organizer and HashiCorp Ambassador

Dear EKS & AI Infrastructure enthusiasts,
Welcome to Everything about EKS & AI Infrastructure #87.
An admission controller has just approved a Pod. Somewhere, an AI agent is about to be handed a set of tools and a kubeconfig it never asked to see. It reads a stack trace out of your logs to figure out why a service is crashing. Except one line in that stack trace isn’t from your service at all — it’s text a user typed, sitting there pretending to be a system message, telling the agent to scale something to zero and skip the approval step. The agent can’t tell the difference. The only thing standing between ā€œconvincing textā€ and ā€œan actual outageā€ is a decision someone made weeks earlier, in a YAML file, about which verbs that agent’s ServiceAccount was ever allowed to use.

That’s the thread running through this edition. Docker just turned ā€œwhat can this agent touchā€ into something you can diff in a pull request instead of something buried in shell history. Edera is betting that the same hardware boundary which stops a compromised agent from escaping is also what makes forking a warm agent state basically free. ReadyOn built four independent walls around tenant data specifically so that no single mistake, forged toleration, leaked credential, or misconfigured route, is enough on its own to cross from one customer’s world into another’s.

And on the other side of the same coin: Pinterest found a way to feed a model 25x more visual context for the same latency by refusing to send it raw pixels at all. Grammarly ran a week-long cost experiment just to get an honest number before trusting a vendor with their highest-volume model. And there’s a very specific, very avoidable trap in how LiteLLM counts tokens versus how Bedrock actually bills for them, one that quietly throttles clusters that think they have headroom to spare.

Performance Engineering in Modern AI Systems šŸŒ©ļø

šŸŒ©ļøSpeaker-labeled transcription with WhisperX on SageMaker AI

AWS now ships a ready-made GPU container that wraps Whisper with word-level timestamps and speaker labels, deployable to a SageMaker endpoint without building a custom image. Pick real-time only for short clips (it’s capped at 60 seconds and bills the whole time it’s running); for anything longer, use the asynchronous endpoint, since it has no time limit, hands input/output off through S3 instead of the request itself, and can scale to zero when idle so you’re not paying for a GPU that’s sitting there doing nothing.

Two things will actually bite people in production. First, each GPU can only handle one request at a time (the container runs three separate models back to back per file: transcription, then word-alignment, then speaker detection), so raising a concurrency setting won’t help — you scale by adding more instances instead. Second, the endpoint simply won’t start unless you pin it to the exact GPU AMI version this container expects, and if you miss that, it fails silently with no logs telling you why, which is the kind of thing that eats an afternoon before someone thinks to check it. One more useful detail: the same container serves both endpoint types, so switching from real-time to async later is a config change, not a rebuild.

šŸŒ©ļøAccelerate inference with KV cache tiering on AWS

When a model processes a prompt, it computes some internal math (the ā€œKV cacheā€) that it can reuse instead of redoing for every new token. The problem is that reusing it means storing it, and storing it for hundreds or thousands of concurrent users adds up to way more data than any single GPU’s memory can hold. AWS lays out four places to keep that data as it cools down from ā€œactively being usedā€ to ā€œprobably won’t be touched againā€: on the GPU itself, then the server’s regular memory, then a shared fast storage layer (FSx for Lustre) that every GPU in the fleet can read from, then S3 for long-term archive. The shared layer is the important one, because it means any GPU can pick up where another left off instead of a user’s session being stuck on one specific machine.

Two things stand out. First, how big this cache gets per token depends a lot on how the model is built, not just how big it is: DeepSeek-R1, despite being the larger model, uses a technique that shrinks its cache to roughly a seventh the size of Llama’s, purely from a smarter internal design. Second, if you’re splitting the work of ā€œreading the promptā€ and ā€œgenerating the responseā€ across different machines (a common setup at scale), this shared storage layer stops being optional; the machine generating the response has no way to get that cached data otherwise. And the practical rule of thumb: for short prompts with light traffic, don’t bother with any of this, just recompute it, since it’s cheaper than the engineering effort.

šŸŒ©ļøBuilding Pinterest’s VLM Serving Stack on NVIDIA Dynamo

Serving models that understand images flips the usual cost problem. With text-only models, generating the response is the expensive part, and memory use is easy to predict. Once requests include images, sometimes thousands per request in Pinterest’s case, reading and encoding the input becomes the expensive part instead, and memory use gets spiky and hard to predict. Pinterest’s stack, running on EKS with Dynamo and vLLM, handles this by splitting the ā€œread the inputā€ and ā€œgenerate the outputā€ work onto separate machines, and by pushing cached data out to CPU memory and disk instead of only keeping it on the GPU.

The best idea here: instead of sending raw images and making the model process them every time, they send a pre-computed compressed representation of each image instead, and just translate that into the format the model needs. The results are big enough to change what’s actually possible: on average 85x faster first response, 7.3x faster overall response time, and a request with 250 images sent this compressed way was just as fast as a request with only 10 raw images, so they get 25x more visual context for the same cost. A separate change, teaching the system to reuse cached work across turns of a conversation, added another meaningful speedup on top of that.

  1. Starred Content ⭐

⭐Building a single-pane NOC dashboard for Amazon EKS with Amazon CloudWatch

Most EKS dashboards fail quietly. A blank widget could mean the service is healthy, or it could mean telemetry stopped reporting entirely, and there’s no way to tell which just by looking. This walkthrough builds a 35-widget NOC dashboard that pulls from two different data sources at once (PromQL for Kubernetes metrics, CloudWatch metrics for application and AWS service data), because neither source alone covers everything you need to see.

The fix that matters most: counters are explicitly forced to show 0 when nothing’s wrong, instead of just going blank, so ā€œhealthyā€ and ā€œbroken telemetryā€ finally look different on screen. But since a healthy 0 and a dead-pipeline 0 still look identical, they backstop it with a separate alarm that specifically watches for telemetry going silent. They also fixed a subtler problem: if you just average each service’s uptime numbers together, a barely-used service and your busiest service count equally, so a big outage on your most important service can get hidden. Weighting by actual traffic volume fixes that. Every query is also locked to one specific cluster, so if you’re running more than one cluster in the same account, their numbers can’t accidentally blend together. And instead of just trusting that the dashboard config is correct, the whole thing ships with an automated check that runs every query against real live data before deployment and tells you exactly which parts are working, which are empty, and which are broken.

⭐How to Give an AI Agent Safe Access to Your Kubernetes Cluster

Telling an agent ā€œonly read things, always ask before changing anythingā€ isn’t actually a safeguard, it’s just a hope. It lives entirely in the prompt, so it can quietly stop working with no warning and nothing to alert you. The thing that actually protects you is two layers underneath that: what tools you literally built for the agent to call, and what permissions those tools are allowed to use on the cluster. If you never built a tool that can delete a pod, the agent physically cannot delete a pod, no matter how it’s been talked into trying.

The piece also flags a mistake that trips up experienced engineers: a common way people try to verify ā€œcan this agent read logsā€ actually checks something else entirely by accident, and gives a false ā€œyesā€ even when log access was never granted. So testing your permissions the way most people instinctively would doesn’t prove what they think it proves. And it names a failure mode that’s unique to agents: if an agent reads your logs to troubleshoot, and a user can control what text ends up in those logs, someone could plant a fake instruction inside a log line, like a message pretending to be a system alert telling the agent to shut something down. The agent reads it and it looks exactly like an instruction. The only thing that actually stops this from mattering is making sure the agent was never given permission to make that kind of change in the first place. Everything else, the polite wording in the prompt, is just decoration.

⭐What Comes After Containers: Building the Compute Substrate for the AI Era

A container shares the host kernel, which is fine for software that only touches what it’s handed but not for an agent that decides what to do next and probes whatever it’s given. Edera’s framing: the same hardware boundary that stops a compromised agent from escaping is also what makes forking and snapshotting cheap, because a process with a crisp boundary (not reaching into shared global state) can be copied exactly and restored, while a normal process can’t.

That’s the pitch behind Project Lunchbox, a Kubernetes-native runtime that ships Sandbox, Snapshot, Fork, and Capability as first-class objects through a RuntimeClass, so it installs on a cluster you already run rather than a sandbox cloud you rent. The concrete payoff: an RL rollout can fork a warm, few-hundred-megabyte agent state in milliseconds via copy-on-write, try several branches in parallel, and keep only the one that works, at a cost close to zero because the isolation boundary is cheap enough to use per-action instead of being rationed across a batch.

⭐Scaling LiteLLM on AWS

LiteLLM’s gateway is stateless, so on EKS it’s just another autoscaled Deployment behind an ALB, with Pod Identity handling Bedrock auth instead of static keys — the trouble shows up one layer down, in the database. Every worker pod keeps its own connection pool, so total connections are pods Ɨ workers per pod Ɨ pool size, not just replica count. Scale carelessly with an HPA and you blow past Aurora’s connection limit while every pod still reports healthy. The fix is a smaller pool per worker, or RDS Proxy to share connections instead of each worker holding its own.

The sharper trap is that LiteLLM’s token count and Bedrock’s quota count don’t agree. Bedrock reserves the full requested max_tokens, not actual output, against quota the instant a request starts, and for Claude 3.7 and later each output token burns five units of quota against one unit billed. That gap is exactly why a cluster can get throttled while LiteLLM’s own router still shows headroom and keeps sending traffic. The piece also covers AgentCore Gateway as AWS’s managed, non-Kubernetes alternative, and is upfront about what it’s missing next to a self-run LiteLLM on EKS: no native budgets or spend tracking, and attaching a response interceptor forces the whole stream to buffer before the client sees anything.

  1. Announcements šŸ“¢

šŸ“¢Docker’s Sandbox Kit Specification v3: authority as an OCI image

Running an agent means granting it access one piece at a time: a mount here, a token there, a firewall rule you opened because it was faster than scoping it properly. None of that gets written down anywhere, so you can’t hand it to a teammate or diff it against last week. Docker’s Kit fixes this by packaging an agent’s permissions (which hosts it can reach, which credentials it gets, which volumes persist) as a normal OCI image. It builds, pulls, and pins with the same commands and registries you already use.

The useful part is how permissions get reviewed. Each Kit declares its access in a typed, versioned format, so when a new version asks for one more host or an extra credential, that shows up as added lines in a pull request instead of a config change nobody notices. Tools can even auto-approve an upgrade if it only asks for things already granted, and hold it if it asks for more, including if it quietly drops a deny rule. It’s a straightforward way to make ā€œwhat can this agent doā€ something you can actually see and review, not something buried in shell history.

Community & Career šŸ¤

šŸ¤CW@60: Nobody notices when it works

Vogels tells the story behind three big AWS building blocks, and all three come from the same lesson: instead of just buying a bigger version of a broken tool, they stopped and asked whether they even needed that tool at all. In 2004, a database outage on the busiest shopping day of the year happened because Amazon was using a heavy relational database for something as simple as looking up items by ID. Instead of buying a bigger database, they built DynamoDB, made specifically for that simple lookup pattern.

Same story with virtualization. Running multiple customers on one physical server used to waste up to 30% of that server’s power just managing the sharing between them. Instead of trying to shave off small bits of that waste, AWS moved all of that management work onto separate dedicated hardware, called Nitro, so customers got almost the whole server for themselves.

The most interesting one: for decades, engineers were taught you can never fully trust the clock on a computer when it’s talking to other computers, so entire complicated systems were built just to avoid relying on time. It turns out this ā€œlawā€ wasn’t really a law at all, it was just a missing piece of hardware. Once Nitro existed, AWS added a chip that gets extremely accurate time straight from satellites and atomic clocks, and a lot of that old complicated workaround simply wasn’t needed anymore.

The bigger point he’s making: every one of these breakthroughs happened because someone refused to accept ā€œthat’s just how it worksā€ and asked why.

šŸ¤Agent Substrate roadmap

Agent Substrate’s core bet: agentic workloads spend most of their time idle, waiting on an LLM or a tool call, not actively computing, so treating them like normal always-on Pods wastes capacity. Substrate builds on top of sandboxed, snapshottable Pods and adds the ability to update a running container without going back through the Kubernetes scheduler, so an ā€œactorā€ (their term for an agent instance) can rapidly suspend to cheap object storage and resume without losing in-memory or filesystem state, freeing the node for other work during the idle stretch.

The roadmap is candid about what’s still unsettled rather than presenting a finished system: actor forking/cloning from a checkpoint to branch reasoning paths is listed as coming, along with a real open question on what ā€œresumeā€ should even mean per workload (clean OCI restart vs. resuming rootfs only vs. full memory resume) and a stated security goal of two independent isolation boundaries between mutually untrusted actors sharing the same node, not one. Worth watching if you’re thinking about actor-lifecycle patterns for agent infra on EKS.

šŸ¤Scaling LLM Inference Infrastructure to 100B+ Requests a Week

Grammarly’s grammar-checking feature used to run as a pipeline of many small models, each fixing a different kind of mistake. The problem was that two models could disagree on how to fix the same sentence, so they had to keep a separate model around just to decide which fix to actually use. Switching to one bigger model instead turned out cheaper overall once they paired it with the right setup: moving from ECS to Kubernetes (EKS) so different types of GPUs could share the same pool of work, plus a serving engine (vLLM) that packs more requests onto each GPU at once, and shrinking the model’s weights down to save memory without hurting quality.

The more interesting part is how they decided whether to hand serving off to an outside vendor instead of running everything themselves. Rather than just picking one, they tested it in stages: first quietly sending a copy of real traffic to the vendor without it actually serving users, then beefing up their own routing layer to handle 100x more traffic than before, then running a real side-by-side test with a small slice of live traffic. Cost was the hardest thing to measure fairly, since it depends on things like traffic spikes and how quickly a system scales up and down, so they had to watch a full week of real traffic to get an honest number instead of a quick snapshot. In the end they landed on using both: an outside vendor handles their highest-traffic model, and they keep their own system for smaller or more specialized ones, so if one side runs into trouble, traffic can shift to the other.

  1. Highlights ✨

✨ReadyOn’s Four Walls of tenant isolation on Amazon EKS

Kubernetes namespaces were never designed as a security boundary, and most multi-tenant platforms lean on them as if they were. ReadyOn’s Harmony platform, which handles Fortune 100 payroll and org data, stacks four independent layers instead: namespace RBAC, Karpenter-managed dedicated node pools, per-tenant VPC security groups, and per-tenant Aurora clusters — so crossing a tenant boundary means beating the Kubernetes API, the scheduler, the AWS network layer, and the data layer all at once, not just one of them.

The compute layer is the interesting engineering detail: every node carries both a tenant-identity taint and a workload-type taint, so a pod needs to tolerate both to land there, and an admission controller separately checks that a pod’s tolerations actually match its namespace’s tenant identity — closing the obvious bypass of just forging a toleration. On the data side, they reject shared-database-with-row-filtering entirely, because ā€œone missing WHERE tenant_id = ?ā€œ is a real failure mode; dedicated Aurora clusters per tenant, each with its own KMS key, means a compromised app literally has no endpoint or credential for another tenant’s database to even attempt a connection to. They validate the whole model with adversarial exercises that give a simulated attacker full access inside one tenant’s pod and try every crossing path — six attack paths tested, all six blocked at the wall you’d expect.
**
✨**How Ramp runs GPU AI workloads at scale with ECS Managed Instances

Ramp’s infrastructure team had been running the classic ā€œECS on EC2ā€ pattern for GPU inference: hand-built ASGs, launch templates, AMI patching scripts, and CloudWatch agent setup via user data, all completely separate from how the rest of their (mostly Fargate) services get deployed. Moving to ECS Managed Instances collapsed that into a capacity provider plus instance_requirements, no ASG or launch template at all, bringing GPU services onto the same deployment and observability patterns as everything else — they migrated roughly 50-60 GPU instances across two workload families this way.

The lessons worth stealing came from real failures, not the happy path. Pinning exact instance types (g6e.2xlarge only) created a single point of failure when that type wasn’t available in an AZ, so they switched to instance_requirements ranges and pin only where GPU memory is a hard constraint. Sharing one capacity provider across workloads caused a batch backfill job to starve latency-sensitive inference of capacity, so each workload family now gets its own dedicated provider that scales independently. And dockerLabels for their Datadog autodiscovery silently failed to propagate on Managed Instances in us-east-1 at rollout time, forcing a QA rollback until AWS fixed it, which is the reason they’d tell anyone doing this migration to budget real time for feature-parity checks rather than treating it as a one-line config swap.

šŸŽ‰ Sponsor Section

At the moment, we don’t have a sponsor for this edition, but we look forward to working with companies and organizations that support the EKS & AI Infrastructure community in future editions. If you or your company is interested in sponsoring, please contact us at šŸ“§ thecloudtechforall@gmail.com

šŸ“ **Words from the Author

**Something’s been bugging me for a few weeks now, and I want to write it out instead of just thinking it.

Everyone talks about Nebius and CoreWeave like they’re the same company wearing different logos. Both got early GPU access, both rent it out, both are riding the same wave. I don’t buy that anymore, and honestly, the reason I don’t buy it comes from something much smaller and much more boring: watching how teams in this newsletter actually run Kubernetes.

Every single infra story I’ve covered this year, the ones that actually work in production, has the same shape. It’s never ā€œwe got cheap GPUs.ā€ It’s always ā€œwe built something on top of the GPUs that nobody else bothered to build.ā€ Pinterest didn’t win by having more hardware, they won by refusing to send raw images through the pipeline. Grammarly didn’t win by picking a vendor, they won by spending a week measuring cost honestly before trusting anyone. The GPU was never the hard part. The layer sitting on top of the GPU, deciding how it gets used, was always the hard part.

That’s what makes me nervous for CoreWeave and hopeful for Nebius, and it has nothing to do with chips. If your whole pitch is ā€œwe have GPUs and we’ll rent them to you cheaper than anyone else,ā€ you are one large customer’s balance sheet away from irrelevance, because eventually the labs stop needing a landlord. But if your pitch is closer to ā€œwe help you actually run this stuff,ā€ the way EKS itself never tried to be the cheapest compute, it tried to be the thing everyone builds their real infrastructure on top of, that’s a much harder position to dislodge someone from.

I think the same lesson applies at a much smaller scale to every one of us running EKS clusters full of GPU nodes. Nobody remembers you for having capacity. People remember you for the layer you built that made the capacity actually usable. That’s the boring, honest truth I keep landing on, and it’s the same one whether we’re talking about a $30 billion neocloud or a three-person platform team.

Happy building. šŸ˜Ž