Blog
Saving Money by Migrating EKS to Graviton with Karpenter
· 16 min read
To migrate EKS to Graviton with Karpenter, build every image for linux/arm64 and linux/amd64, find the images and dependencies that can't run on arm64, add a weighted arm64 NodePool with an x86 NodePool as fallback, move stateless services first, and measure cost per workload with split cost allocation data.
The rest of this post covers the details that decide whether it takes two weeks or two quarters, and whether the savings show up on your invoice.
At my last company I owned our EKS upgrades from 1.23 through 1.32. We ran cluster-autoscaler for years before moving to Karpenter, and Karpenter was better: it provisioned nodes faster, and declarative NodePools express capacity more clearly than a pile of node groups.
On Graviton work, the instance price is the easy part. Most of the calendar time goes to confirming that every workload runs on arm64.
Prices, discount rates, and dates below are accurate as of this writing. AWS changes pricing, regions, and instance availability over time, and your Region and account may show different numbers, so confirm current details on AWS's pricing pages before you plan a migration around them.
Which Graviton generations can you run on EKS in 2026?
Graviton4 is the mature, widely available generation: M8g, C8g, R8g, and their variants. AWS has kept adding regions for it, including Paris, Osaka, Canada Central, and Bahrain in late 2025. For most teams it's the default target.
Graviton5 is newer. AWS previewed M9g at re:Invent in December 2025 and made M9g and M9gd generally available on June 10, 2026. Compute-optimized C9g and C9gd followed on June 30, and memory-optimized R9g and R9gd on August 31. All three launched in US East (N. Virginia), US East (Ohio), US West (Oregon), and Europe (Frankfurt). On September 3, AWS added M9g and M9gd in Europe (Ireland), Singapore, Sydney, and Tokyo. AWS claims up to 25% better compute performance than the matching Graviton4 instances.
Don't build the migration around Graviton5. Build it around arm64 and let Karpenter choose generations. The 9th-gen instances list about 9% above Graviton4 on demand, and they're in eight regions so far. If you're in one of them, allow them in your NodePool. Don't make the project wait on them.
Is Graviton actually cheaper than x86?
Per hour, yes, in every pairing I checked for this post. How much cheaper depends on which generations you compare, and that's where the common advice falls apart.
The number you'll see everywhere is 20%. AWS's own Graviton pages say EKS workloads get up to 40% better performance at 20% lower cost than comparable x86 instances. That 20% roughly holds for Graviton3 against Intel's 7th generation. It doesn't hold for the newest instances.
| Instance (2 vCPU, 8 GiB) | Processor | Arch | On-demand, us-east-1 | Gap vs. same-gen Intel | Gap vs. m8i.large |
|---|---|---|---|---|---|
| m7i.large | Intel, 7th gen | x86_64 | $0.1008/hr | baseline | 4.7% cheaper |
| m7g.large | Graviton3 | arm64 | $0.0816/hr | 19.0% cheaper than m7i | 22.9% cheaper |
| m8i.large | Intel, 8th gen | x86_64 | $0.1058/hr | baseline | baseline |
| m8g.large | Graviton4 | arm64 | $0.08976/hr | 15.2% cheaper than m8i | 15.2% cheaper |
| m9g.large | Graviton5 | arm64 | $0.0978/hr | no 9th-gen Intel M family to compare | 7.6% cheaper |
Linux on-demand rates in us-east-1 as reported by public pricing trackers, checked September 26, 2026. Prices change, so confirm against the AWS pricing page before you budget.
The hourly gap shrinks as Graviton gets newer: about 19% at generation 7, 15% at generation 8, and under 8% if you compare m9g against m8i. The remaining case for Graviton5 is performance per vCPU, and performance only saves money when you act on it by lowering CPU requests or running fewer nodes. If you leave requests alone, you get the list-price difference and nothing else.
One reason performance often improves: on Intel instances a vCPU is a hyperthread, while on Graviton each vCPU is a full physical core. AWS's prescriptive guidance calls this out directly. It also means CPU utilization graphs aren't comparable across architectures, so don't benchmark by eyeballing dashboards.
What breaks when you move EKS workloads to arm64?
Kubernetes itself runs fine on arm64. What breaks is any container image without an arm64 variant, including the ones you didn't build.
Start with an inventory of every image running in the cluster, init containers included:
kubectl get pods -A -o jsonpath='{range .items[*]}{range .spec.containers[*]}{.image}{"\n"}{end}{range .spec.initContainers[*]}{.image}{"\n"}{end}{end}' \
| sort -u > images.txt
while read -r img; do
if docker buildx imagetools inspect "$img" 2>/dev/null | grep -q 'linux/arm64'; then
echo "OK $img"
else
echo "NO $img"
fi
done < images.txt
Log in to ECR and any private registries first, or everything private shows up as NO. Single-architecture images also show as NO, which is the conservative answer you want.
Then look at how your own images get built. Grep your Dockerfiles and CI config for hardcoded architectures:
grep -rnE 'amd64|x86_64' --include='Dockerfile*' --include='*.yml' --include='*.yaml' .
The usual culprits are a curl that downloads a linux-amd64 binary, a base image pinned by digest to a single-arch manifest, and old versions of native dependencies: Python packages without arm64 wheels, JNI libraries, Node native addons. Pure Python, Node, Ruby, and JVM code generally runs unchanged. The failures live in native code and downloaded binaries.
For source-level checks, AWS's open source Porting Advisor for Graviton scans a repo for known incompatible code patterns and outdated dependencies, and suggests minimum versions. It covers Python, Java, Go, C/C++, and Fortran, and has dependency scanners for npm and NuGet. It only reads source, not binaries, and it runs on x86 machines:
./porting-advisor-linux-x86_64 ~/src/billing-service --output report.html
Don't skip DaemonSets. Many cluster agents ship with a toleration that matches every taint, so the first arm64 node you launch gets every DaemonSet immediately, tainted or not. Your CNI, log shipper, metrics agent, and security agent all need arm64 images before that first node joins.
How do you build multi-arch images in CI?
The core change is one flag. With Docker Buildx you build both platforms and push a single tag that points to a manifest list:
docker buildx create --use --name multiarch
docker buildx build \
--platform linux/amd64,linux/arm64 \
-t 123456789012.dkr.ecr.us-east-1.amazonaws.com/billing:${GIT_SHA} \
--push .
Build time is the problem. On an x86 runner, the arm64 half runs under QEMU emulation, which is fine for interpreted languages and painfully slow for anything that compiles. You have two fixes. Run arm64 builds on native arm64 runners (CodeBuild offers an ARM container environment type, and the major CI vendors offer hosted arm64 Linux runners), or cross-compile where the toolchain supports it. For Go:
FROM --platform=$BUILDPLATFORM golang:1.23 AS build
ARG TARGETOS TARGETARCH
WORKDIR /src
COPY . .
RUN CGO_ENABLED=0 GOOS=$TARGETOS GOARCH=$TARGETARCH go build -o /out/app ./cmd/app
FROM gcr.io/distroless/static
COPY --from=build /out/app /app
ENTRYPOINT ["/app"]
Add one CI gate: after pushing, run docker buildx imagetools inspect on the tag and fail the pipeline if linux/arm64 is missing. That keeps a new service from shipping x86-only six months from now.
Build multi-arch, not arm64-only, for the entire migration and for a while after. Your x86 NodePool is only a fallback if every image can still run there.
How should you configure Karpenter NodePools for arm64 with x86 fallback?
Use two NodePools. The arm64 pool gets a higher weight and a taint. The x86 pool gets a lower weight and no taint. This uses the Karpenter v1 API (karpenter.sh/v1).
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: graviton
spec:
weight: 100
template:
spec:
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: default
requirements:
- key: kubernetes.io/arch
operator: In
values: ["arm64"]
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand"] # add "spot" after validation
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["c", "m", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gte
values: ["7"]
taints:
- key: example.com/arm64-ready
value: "true"
effect: NoSchedule
limits:
cpu: "200"
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 1m
---
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: x86
spec:
weight: 10
template:
spec:
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: default
requirements:
- key: kubernetes.io/arch
operator: In
values: ["amd64"]
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["c", "m", "r"]
Workloads opt in with a toleration, and only after their images pass the inventory:
tolerations:
- key: example.com/arm64-ready
operator: Equal
value: "true"
effect: NoSchedule
A tolerating pod matches both pools, and Karpenter prefers the one with the higher weight. A pod without the toleration can only land on x86.
Apply and test NodePool changes like these in a non-production cluster first, and check them against whatever manages your cluster config (Terraform, Argo CD, or similar) so a manual change doesn't drift or get silently reverted.
Common advice says to put both architectures in one NodePool and let Karpenter pick the cheapest. That's fine at the end of the migration. During it, it's how you get outages. Karpenter doesn't inspect image manifests, so a pod with an x86-only image and no architecture constraint can be placed on an arm64 node and crash with exec format error. The taint makes arm64 opt-in per workload, which is the control you want.
Check your EC2NodeClass too. If amiSelectorTerms uses an alias such as al2023@<version>, Karpenter resolves an AMI for each architecture. If you pinned AMI IDs, you need arm64 IDs as well, or arm64 NodeClaims will never launch.
Test the fallback on purpose. In staging, drop the graviton pool's limits.cpu low, scale up a tolerating deployment, and confirm the overflow lands on x86.
One more thing Karpenter won't do for you: it picks instances by price against the requests it sees. It doesn't know a Graviton5 core does more work than a Graviton4 core. If you want 9th-gen for a latency-sensitive service, give it its own pool or a nodeSelector on karpenter.k8s.aws/instance-generation, and set that service's requests from load-test results.
What order should you migrate workloads in?
- Make every DaemonSet and platform component multi-arch before the first arm64 node exists.
- Move stateless internal services with decent test coverage onto on-demand arm64 nodes. Watch error rates and p95/p99 latency for at least a week.
- Move customer-facing stateless APIs, one service at a time, using the same checks.
- Move async workers, queue consumers, and batch jobs. These are good Spot candidates, so this is when I add
spotto the graviton pool. - Leave in-cluster stateful workloads for last, or out of scope entirely for phase one.
- Once every image passes the inventory, remove the taint and let both architectures compete on price.
I keep the arm64 pool on-demand during steps 2 and 3 on purpose. Spot interruptions during validation muddy the signal, and you won't be able to tell an arm64 bug from a reclaimed node.
How do Graviton, Spot, and Savings Plans work together?
For Spot, Karpenter prioritizes reserved capacity, then Spot, then on-demand when a NodePool allows several capacity types, and it falls back quickly when capacity is unavailable. Keep instance requirements broad (several categories and generations) so Spot has deep pools to draw from, and make sure Karpenter's interruption queue is configured.
Savings Plans are where Graviton migrations lose money. A Compute Savings Plan applies across instance families and regions, plus Fargate and Lambda, so it follows you from Intel to Graviton. An EC2 Instance Savings Plan is locked to one family in one region, so an m6i or m7i commitment won't cover a single Graviton hour. Neither type applies to Spot.
The risk is with the flexible kind. Say you have a Compute Savings Plan committing $4.00/hr, and your EKS usage at Savings Plan rates is exactly $4.00/hr. You move to Graviton and that usage drops 15%, to $3.40/hr. You still pay $4.00/hr. The migration saves $0 until you grow back into the commitment or it expires. The same thing happens when you shift on-demand work to Spot.
So check commitment coverage before you promise the CFO a number. If you hold EC2 Instance Savings Plans on x86 families, one option is to leave that much x86 capacity running until the term ends and migrate the rest. I'd generally avoid buying new commitments mid-migration, and would let usage settle for a month on the new architecture before sizing them. I compare commitment types in more detail in Database Savings Plans vs. Reserved Instances.
If you're not sure how much of your EKS spend is already committed, or whether a migration would just leave a plan underused, that's the part of the analysis I do in the Cloudshipped Cost Audit for AWS. You can book a free discovery call here.
How do you prove the savings are real?
Measure cost per workload, before and after, using split cost allocation data. It's free apart from S3 storage for the reports.
- From the management account, enable split cost allocation data in the Billing and Cost Management console.
- In Data Exports, create a CUR 2.0 export with split cost allocation data included. Turn on resource IDs and hourly granularity.
- Label pods consistently, for example
app.kubernetes.io/nameandteam. Split cost allocation imports up to 50 labels per pod, sorted alphabetically, and drops the rest, so keep label counts sane. - Activate those labels as cost allocation tags in the management account. They appear in the CUR within 24 hours.
- Capture at least two weeks of baseline before moving a service, then query cost per app per day in Athena and divide by a unit of work from your metrics, such as requests served or jobs processed.
Cost per unit of work matters because split cost allocation builds on pod CPU and memory requests. If you move a service and leave its requests untouched, its allocated cost falls by the price ratio and no further, even if it's now running at half the CPU. The additional savings only show up after you lower requests and Karpenter consolidates onto fewer nodes.
What savings should you actually expect?
From price alone, roughly 8% to 19% off the on-demand rate of the nodes you move, depending on which generations you're moving between (see the table above). AWS publishes up-to-40%-better-performance claims for Graviton. Treat that as upside you earn with load tests, not as the plan.
The math, with assumptions stated: Take 20 m8i.2xlarge nodes on demand in us-east-1, running all month (730 hours), at $0.4234/hr each. That's $6,181.64/month. The same 20 nodes as m8g.2xlarge at $0.35904/hr cost $5,241.98, saving $939.66/month, or 15.2%.
If load tests justify lower CPU requests and Karpenter consolidates you to 16 nodes, the cost drops to $4,193.59, a 32.2% reduction from the original. That second number only exists if you do the request tuning.
If that fleet sits inside a $30,000/month AWS bill, the 15% case is about 3% of total spend. Worth doing, but it rarely beats fixing commitment coverage or idle capacity first, and it's the combination of all three that moves runway.
If you want those numbers for your own account, including what Graviton would save after your existing Savings Plans are accounted for, book a free discovery call for the Cloudshipped Cost Audit for AWS.
FAQ
Is Graviton5 (M9g) available for EKS?
Yes. M9g and M9gd became generally available June 10, 2026, followed by C9g/C9gd on June 30 and R9g/R9gd on August 31. Availability started in four US and EU regions, and M9g added Ireland, Singapore, Sydney, and Tokyo in September. Karpenter can launch them like any other arm64 instance.
How much does Graviton save on EKS?
On-demand list prices in us-east-1 put Graviton 7.6% to 19% below comparable Intel instances, depending on generation. AWS advertises 20% lower cost plus higher performance. Extra savings come only if you lower CPU requests or node counts after load testing. Measure cost per unit of work, not cost per node.
Will my Savings Plans cover Graviton instances?
Compute Savings Plans apply to any EC2 instance family and region, plus Fargate and Lambda, so they follow you to Graviton. EC2 Instance Savings Plans are locked to one family in one region, so x86 commitments won't cover Graviton. Neither applies to Spot. Cheaper usage can also leave a Compute Savings Plan underused.
How do I set up Karpenter to prefer Graviton with an x86 fallback?
Create two NodePools: an arm64 pool with a higher weight and a taint, and an amd64 pool with a lower weight. Add a matching toleration only to workloads whose images include linux/arm64. Karpenter prefers the higher-weight pool, and pods still match the x86 pool when the arm64 pool hits its limits.
Disclaimer
This article is general information, not financial, tax, or legal advice for any specific AWS account. AWS pricing, discount rates, instance availability, and Karpenter or Kubernetes behavior change over time and vary by Region and account, so confirm current details with AWS before acting. You're responsible for testing and validating any change, including NodePool and workload migrations, in non-production before applying it to production. Cloudshipped isn't liable for costs, outages, or losses from actions taken based on this article, and isn't affiliated with or endorsed by Amazon Web Services.
Sources
- AWS What's New, Amazon EC2 M9g and M9gd instances powered by Graviton5 now available (Jun 2026)
- AWS News Blog, Now available: Amazon EC2 M9g and M9gd instances powered by new AWS Graviton5 processors
- AWS What's New, Amazon EC2 M9g and M9gd instances now available in four more regions (Sep 2026)
- AWS What's New, Amazon EC2 C9g and C9gd instances powered by Graviton5 now available (Jun 2026)
- AWS What's New, Amazon EC2 R9g and R9gd memory-optimized instances now available (Aug 2026)
- AWS What's New, Amazon EC2 M8g instances in additional regions (Oct 2025)
- InfoQ, AWS Graviton5 reaches general availability (Jun 2026)
- AWS, Graviton Fast Start
- AWS, Getting started with AWS Graviton
- AWS Prescriptive Guidance, Running .NET on Graviton
- AWS, Compute Savings Plans pricing
- AWS Documentation, Split cost allocation data: Kubernetes labels
- AWS What's New, Split cost allocation data for Amazon EKS now supports Kubernetes labels (Oct 2025)
- AWS Cloud Financial Management Blog, Using Kubernetes labels to split and track application costs on Amazon EKS
- Karpenter Documentation, NodePools
- GitHub, aws/porting-advisor-for-graviton
- AWS What's New, Porting Advisor for Graviton (Jan 2023)
- Holori, m9g.large pricing, us-east-1
- Holori, m8i.large pricing, us-east-1
- Holori, m7g.large pricing, us-east-1
- Vantage, m8g.large
- DevZero, m7i.large
Free discovery call
Find the savings your AWS bill is hiding
Book a free discovery call to see how Cloudshipped can help save you money on AWS spend.
- Fixed price
- 1-week target
- Fee refunded if under 10%
Prefer email? Write to support@cloudshipped.co
