š Everything about EKS & AI Infrastructure Newsletter "#87" āļøā¤šØāš»
Isolation and efficiency are turning out to be the same problem ā plus a look at what actually breaks GPU inference at scale.

Dear EKS & AI Infrastructure enthusiasts,
Welcome to Everything about EKS & AI Infrastructure #87.
An admission controller has just approved a Pod. Somewhere, an AI agent is about to be handed a set of tools and a kubeconfig it never asked to see. It reads a stack trace out of your logs to figure out why a service is crashing. Except one line in that stack trace isnāt from your service at all ā itās text a user typed, sitting there pretending to be a system message, telling the agent to scale something to zero and skip the approval step. The agent canāt tell the difference. The only thing standing between āconvincing textā and āan actual outageā is a decision someone made weeks earlier, in a YAML file, about which verbs that agentās ServiceAccount was ever allowed to use.
Thatās the thread running through this edition. Docker just turned āwhat can this agent touchā into something you can diff in a pull request instead of something buried in shell history. Edera is betting that the same hardware boundary which stops a compromised agent from escaping is also what makes forking a warm agent state basically free. ReadyOn built four independent walls around tenant data specifically so that no single mistake, forged toleration, leaked credential, or misconfigured route, is enough on its own to cross from one customerās world into anotherās.
And on the other side of the same coin: Pinterest found a way to feed a model 25x more visual context for the same latency by refusing to send it raw pixels at all. Grammarly ran a week-long cost experiment just to get an honest number before trusting a vendor with their highest-volume model. And thereās a very specific, very avoidable trap in how LiteLLM counts tokens versus how Bedrock actually bills for them, one that quietly throttles clusters that think they have headroom to spare.
Performance Engineering in Modern AI Systems š©ļø
š©ļøSpeaker-labeled transcription with WhisperX on SageMaker AI
AWS now ships a ready-made GPU container that wraps Whisper with word-level timestamps and speaker labels, deployable to a SageMaker endpoint without building a custom image. Pick real-time only for short clips (itās capped at 60 seconds and bills the whole time itās running); for anything longer, use the asynchronous endpoint, since it has no time limit, hands input/output off through S3 instead of the request itself, and can scale to zero when idle so youāre not paying for a GPU thatās sitting there doing nothing.
Two things will actually bite people in production. First, each GPU can only handle one request at a time (the container runs three separate models back to back per file: transcription, then word-alignment, then speaker detection), so raising a concurrency setting wonāt help ā you scale by adding more instances instead. Second, the endpoint simply wonāt start unless you pin it to the exact GPU AMI version this container expects, and if you miss that, it fails silently with no logs telling you why, which is the kind of thing that eats an afternoon before someone thinks to check it. One more useful detail: the same container serves both endpoint types, so switching from real-time to async later is a config change, not a rebuild.
š©ļøAccelerate inference with KV cache tiering on AWS
When a model processes a prompt, it computes some internal math (the āKV cacheā) that it can reuse instead of redoing for every new token. The problem is that reusing it means storing it, and storing it for hundreds or thousands of concurrent users adds up to way more data than any single GPUās memory can hold. AWS lays out four places to keep that data as it cools down from āactively being usedā to āprobably wonāt be touched againā: on the GPU itself, then the serverās regular memory, then a shared fast storage layer (FSx for Lustre) that every GPU in the fleet can read from, then S3 for long-term archive. The shared layer is the important one, because it means any GPU can pick up where another left off instead of a userās session being stuck on one specific machine.
Two things stand out. First, how big this cache gets per token depends a lot on how the model is built, not just how big it is: DeepSeek-R1, despite being the larger model, uses a technique that shrinks its cache to roughly a seventh the size of Llamaās, purely from a smarter internal design. Second, if youāre splitting the work of āreading the promptā and āgenerating the responseā across different machines (a common setup at scale), this shared storage layer stops being optional; the machine generating the response has no way to get that cached data otherwise. And the practical rule of thumb: for short prompts with light traffic, donāt bother with any of this, just recompute it, since itās cheaper than the engineering effort.
š©ļøBuilding Pinterestās VLM Serving Stack on NVIDIA Dynamo
Serving models that understand images flips the usual cost problem. With text-only models, generating the response is the expensive part, and memory use is easy to predict. Once requests include images, sometimes thousands per request in Pinterestās case, reading and encoding the input becomes the expensive part instead, and memory use gets spiky and hard to predict. Pinterestās stack, running on EKS with Dynamo and vLLM, handles this by splitting the āread the inputā and āgenerate the outputā work onto separate machines, and by pushing cached data out to CPU memory and disk instead of only keeping it on the GPU.
The best idea here: instead of sending raw images and making the model process them every time, they send a pre-computed compressed representation of each image instead, and just translate that into the format the model needs. The results are big enough to change whatās actually possible: on average 85x faster first response, 7.3x faster overall response time, and a request with 250 images sent this compressed way was just as fast as a request with only 10 raw images, so they get 25x more visual context for the same cost. A separate change, teaching the system to reuse cached work across turns of a conversation, added another meaningful speedup on top of that.
- Starred Content ā
āBuilding a single-pane NOC dashboard for Amazon EKS with Amazon CloudWatch
Most EKS dashboards fail quietly. A blank widget could mean the service is healthy, or it could mean telemetry stopped reporting entirely, and thereās no way to tell which just by looking. This walkthrough builds a 35-widget NOC dashboard that pulls from two different data sources at once (PromQL for Kubernetes metrics, CloudWatch metrics for application and AWS service data), because neither source alone covers everything you need to see.
The fix that matters most: counters are explicitly forced to show 0 when nothingās wrong, instead of just going blank, so āhealthyā and ābroken telemetryā finally look different on screen. But since a healthy 0 and a dead-pipeline 0 still look identical, they backstop it with a separate alarm that specifically watches for telemetry going silent. They also fixed a subtler problem: if you just average each serviceās uptime numbers together, a barely-used service and your busiest service count equally, so a big outage on your most important service can get hidden. Weighting by actual traffic volume fixes that. Every query is also locked to one specific cluster, so if youāre running more than one cluster in the same account, their numbers canāt accidentally blend together. And instead of just trusting that the dashboard config is correct, the whole thing ships with an automated check that runs every query against real live data before deployment and tells you exactly which parts are working, which are empty, and which are broken.
āHow to Give an AI Agent Safe Access to Your Kubernetes Cluster
Telling an agent āonly read things, always ask before changing anythingā isnāt actually a safeguard, itās just a hope. It lives entirely in the prompt, so it can quietly stop working with no warning and nothing to alert you. The thing that actually protects you is two layers underneath that: what tools you literally built for the agent to call, and what permissions those tools are allowed to use on the cluster. If you never built a tool that can delete a pod, the agent physically cannot delete a pod, no matter how itās been talked into trying.
The piece also flags a mistake that trips up experienced engineers: a common way people try to verify ācan this agent read logsā actually checks something else entirely by accident, and gives a false āyesā even when log access was never granted. So testing your permissions the way most people instinctively would doesnāt prove what they think it proves. And it names a failure mode thatās unique to agents: if an agent reads your logs to troubleshoot, and a user can control what text ends up in those logs, someone could plant a fake instruction inside a log line, like a message pretending to be a system alert telling the agent to shut something down. The agent reads it and it looks exactly like an instruction. The only thing that actually stops this from mattering is making sure the agent was never given permission to make that kind of change in the first place. Everything else, the polite wording in the prompt, is just decoration.
āWhat Comes After Containers: Building the Compute Substrate for the AI Era
A container shares the host kernel, which is fine for software that only touches what itās handed but not for an agent that decides what to do next and probes whatever itās given. Ederaās framing: the same hardware boundary that stops a compromised agent from escaping is also what makes forking and snapshotting cheap, because a process with a crisp boundary (not reaching into shared global state) can be copied exactly and restored, while a normal process canāt.
Thatās the pitch behind Project Lunchbox, a Kubernetes-native runtime that ships Sandbox, Snapshot, Fork, and Capability as first-class objects through a RuntimeClass, so it installs on a cluster you already run rather than a sandbox cloud you rent. The concrete payoff: an RL rollout can fork a warm, few-hundred-megabyte agent state in milliseconds via copy-on-write, try several branches in parallel, and keep only the one that works, at a cost close to zero because the isolation boundary is cheap enough to use per-action instead of being rationed across a batch.
LiteLLMās gateway is stateless, so on EKS itās just another autoscaled Deployment behind an ALB, with Pod Identity handling Bedrock auth instead of static keys ā the trouble shows up one layer down, in the database. Every worker pod keeps its own connection pool, so total connections are pods Ć workers per pod Ć pool size, not just replica count. Scale carelessly with an HPA and you blow past Auroraās connection limit while every pod still reports healthy. The fix is a smaller pool per worker, or RDS Proxy to share connections instead of each worker holding its own.
The sharper trap is that LiteLLMās token count and Bedrockās quota count donāt agree. Bedrock reserves the full requested max_tokens, not actual output, against quota the instant a request starts, and for Claude 3.7 and later each output token burns five units of quota against one unit billed. That gap is exactly why a cluster can get throttled while LiteLLMās own router still shows headroom and keeps sending traffic. The piece also covers AgentCore Gateway as AWSās managed, non-Kubernetes alternative, and is upfront about what itās missing next to a self-run LiteLLM on EKS: no native budgets or spend tracking, and attaching a response interceptor forces the whole stream to buffer before the client sees anything.
- Announcements š¢
š¢Dockerās Sandbox Kit Specification v3: authority as an OCI image
Running an agent means granting it access one piece at a time: a mount here, a token there, a firewall rule you opened because it was faster than scoping it properly. None of that gets written down anywhere, so you canāt hand it to a teammate or diff it against last week. Dockerās Kit fixes this by packaging an agentās permissions (which hosts it can reach, which credentials it gets, which volumes persist) as a normal OCI image. It builds, pulls, and pins with the same commands and registries you already use.
The useful part is how permissions get reviewed. Each Kit declares its access in a typed, versioned format, so when a new version asks for one more host or an extra credential, that shows up as added lines in a pull request instead of a config change nobody notices. Tools can even auto-approve an upgrade if it only asks for things already granted, and hold it if it asks for more, including if it quietly drops a deny rule. Itās a straightforward way to make āwhat can this agent doā something you can actually see and review, not something buried in shell history.
Community & Career š¤
š¤CW@60: Nobody notices when it works
Vogels tells the story behind three big AWS building blocks, and all three come from the same lesson: instead of just buying a bigger version of a broken tool, they stopped and asked whether they even needed that tool at all. In 2004, a database outage on the busiest shopping day of the year happened because Amazon was using a heavy relational database for something as simple as looking up items by ID. Instead of buying a bigger database, they built DynamoDB, made specifically for that simple lookup pattern.
Same story with virtualization. Running multiple customers on one physical server used to waste up to 30% of that serverās power just managing the sharing between them. Instead of trying to shave off small bits of that waste, AWS moved all of that management work onto separate dedicated hardware, called Nitro, so customers got almost the whole server for themselves.
The most interesting one: for decades, engineers were taught you can never fully trust the clock on a computer when itās talking to other computers, so entire complicated systems were built just to avoid relying on time. It turns out this ālawā wasnāt really a law at all, it was just a missing piece of hardware. Once Nitro existed, AWS added a chip that gets extremely accurate time straight from satellites and atomic clocks, and a lot of that old complicated workaround simply wasnāt needed anymore.
The bigger point heās making: every one of these breakthroughs happened because someone refused to accept āthatās just how it worksā and asked why.
Agent Substrateās core bet: agentic workloads spend most of their time idle, waiting on an LLM or a tool call, not actively computing, so treating them like normal always-on Pods wastes capacity. Substrate builds on top of sandboxed, snapshottable Pods and adds the ability to update a running container without going back through the Kubernetes scheduler, so an āactorā (their term for an agent instance) can rapidly suspend to cheap object storage and resume without losing in-memory or filesystem state, freeing the node for other work during the idle stretch.
The roadmap is candid about whatās still unsettled rather than presenting a finished system: actor forking/cloning from a checkpoint to branch reasoning paths is listed as coming, along with a real open question on what āresumeā should even mean per workload (clean OCI restart vs. resuming rootfs only vs. full memory resume) and a stated security goal of two independent isolation boundaries between mutually untrusted actors sharing the same node, not one. Worth watching if youāre thinking about actor-lifecycle patterns for agent infra on EKS.
š¤Scaling LLM Inference Infrastructure to 100B+ Requests a Week
Grammarlyās grammar-checking feature used to run as a pipeline of many small models, each fixing a different kind of mistake. The problem was that two models could disagree on how to fix the same sentence, so they had to keep a separate model around just to decide which fix to actually use. Switching to one bigger model instead turned out cheaper overall once they paired it with the right setup: moving from ECS to Kubernetes (EKS) so different types of GPUs could share the same pool of work, plus a serving engine (vLLM) that packs more requests onto each GPU at once, and shrinking the modelās weights down to save memory without hurting quality.
The more interesting part is how they decided whether to hand serving off to an outside vendor instead of running everything themselves. Rather than just picking one, they tested it in stages: first quietly sending a copy of real traffic to the vendor without it actually serving users, then beefing up their own routing layer to handle 100x more traffic than before, then running a real side-by-side test with a small slice of live traffic. Cost was the hardest thing to measure fairly, since it depends on things like traffic spikes and how quickly a system scales up and down, so they had to watch a full week of real traffic to get an honest number instead of a quick snapshot. In the end they landed on using both: an outside vendor handles their highest-traffic model, and they keep their own system for smaller or more specialized ones, so if one side runs into trouble, traffic can shift to the other.
- Highlights āØ
āØReadyOnās Four Walls of tenant isolation on Amazon EKS
Kubernetes namespaces were never designed as a security boundary, and most multi-tenant platforms lean on them as if they were. ReadyOnās Harmony platform, which handles Fortune 100 payroll and org data, stacks four independent layers instead: namespace RBAC, Karpenter-managed dedicated node pools, per-tenant VPC security groups, and per-tenant Aurora clusters ā so crossing a tenant boundary means beating the Kubernetes API, the scheduler, the AWS network layer, and the data layer all at once, not just one of them.
The compute layer is the interesting engineering detail: every node carries both a tenant-identity taint and a workload-type taint, so a pod needs to tolerate both to land there, and an admission controller separately checks that a podās tolerations actually match its namespaceās tenant identity ā closing the obvious bypass of just forging a toleration. On the data side, they reject shared-database-with-row-filtering entirely, because āone missing WHERE tenant_id = ?ā is a real failure mode; dedicated Aurora clusters per tenant, each with its own KMS key, means a compromised app literally has no endpoint or credential for another tenantās database to even attempt a connection to. They validate the whole model with adversarial exercises that give a simulated attacker full access inside one tenantās pod and try every crossing path ā six attack paths tested, all six blocked at the wall youād expect.
**
āØ**How Ramp runs GPU AI workloads at scale with ECS Managed Instances
Rampās infrastructure team had been running the classic āECS on EC2ā pattern for GPU inference: hand-built ASGs, launch templates, AMI patching scripts, and CloudWatch agent setup via user data, all completely separate from how the rest of their (mostly Fargate) services get deployed. Moving to ECS Managed Instances collapsed that into a capacity provider plus instance_requirements, no ASG or launch template at all, bringing GPU services onto the same deployment and observability patterns as everything else ā they migrated roughly 50-60 GPU instances across two workload families this way.
The lessons worth stealing came from real failures, not the happy path. Pinning exact instance types (g6e.2xlarge only) created a single point of failure when that type wasnāt available in an AZ, so they switched to instance_requirements ranges and pin only where GPU memory is a hard constraint. Sharing one capacity provider across workloads caused a batch backfill job to starve latency-sensitive inference of capacity, so each workload family now gets its own dedicated provider that scales independently. And dockerLabels for their Datadog autodiscovery silently failed to propagate on Managed Instances in us-east-1 at rollout time, forcing a QA rollback until AWS fixed it, which is the reason theyād tell anyone doing this migration to budget real time for feature-parity checks rather than treating it as a one-line config swap.
š Sponsor Section
At the moment, we donāt have a sponsor for this edition, but we look forward to working with companies and organizations that support the EKS & AI Infrastructure community in future editions. If you or your company is interested in sponsoring, please contact us at š§ thecloudtechforall@gmail.com
š **Words from the Author
**Somethingās been bugging me for a few weeks now, and I want to write it out instead of just thinking it.
Everyone talks about Nebius and CoreWeave like theyāre the same company wearing different logos. Both got early GPU access, both rent it out, both are riding the same wave. I donāt buy that anymore, and honestly, the reason I donāt buy it comes from something much smaller and much more boring: watching how teams in this newsletter actually run Kubernetes.
Every single infra story Iāve covered this year, the ones that actually work in production, has the same shape. Itās never āwe got cheap GPUs.ā Itās always āwe built something on top of the GPUs that nobody else bothered to build.ā Pinterest didnāt win by having more hardware, they won by refusing to send raw images through the pipeline. Grammarly didnāt win by picking a vendor, they won by spending a week measuring cost honestly before trusting anyone. The GPU was never the hard part. The layer sitting on top of the GPU, deciding how it gets used, was always the hard part.
Thatās what makes me nervous for CoreWeave and hopeful for Nebius, and it has nothing to do with chips. If your whole pitch is āwe have GPUs and weāll rent them to you cheaper than anyone else,ā you are one large customerās balance sheet away from irrelevance, because eventually the labs stop needing a landlord. But if your pitch is closer to āwe help you actually run this stuff,ā the way EKS itself never tried to be the cheapest compute, it tried to be the thing everyone builds their real infrastructure on top of, thatās a much harder position to dislodge someone from.
I think the same lesson applies at a much smaller scale to every one of us running EKS clusters full of GPU nodes. Nobody remembers you for having capacity. People remember you for the layer you built that made the capacity actually usable. Thatās the boring, honest truth I keep landing on, and itās the same one whether weāre talking about a $30 billion neocloud or a three-person platform team.
Happy building. š



