AWS EKS IRSA & Karpenter Production Autoscaling: Step-by-Step Hands-On Runbook
1. The Problem: Why Legacy Cluster Autoscaler Fails at Scale
In traditional AWS Elastic Kubernetes Service (EKS) setups, the standard Kubernetes Cluster Autoscaler interacts with EC2 Auto Scaling Groups (ASGs). When sudden traffic spikes hit your services, the workflow involves waiting for pods to fail scheduling, waiting for the ASG to request EC2 instances, waiting for Cloud-Init bootstrap, and hoping the fixed instance type is available in the current Availability Zone. Total time to scale: 4 to 8 minutes.
Karpenter bypasses EC2 Auto Scaling Groups completely. It directly interrogates the Kubernetes API for unschedulable pod resource requirements (CPU, memory, GPU, architecture, topology spread), invokes the AWS EC2 Fleet API, and provisions the most cost-effective Spot or On-Demand instance in under 45 seconds.
Kube-scheduler detects insufficient capacity on existing cluster nodes.
Evaluates required CPU, memory, and topology spread in milliseconds.
Calls AWS EC2 Fleet directly to provision right-sized Spot or On-Demand instance.
Node registers with EKS cluster; all pending pods scheduled immediately.
2. Prerequisites & Target Stack
- An existing AWS EKS cluster running version 1.28, 1.29, or 1.30+.
- AWS CLI v2 authenticated with sufficient administrative rights (
AdministratorAccessor targeted IAM permissions). kubectlconfigured to point to your EKS cluster context.- Helm v3 installed locally.
- Cluster name and AWS Region saved to shell variables:
export CLUSTER_NAME="production-eks-cluster"
export AWS_REGION="us-east-1"
export AWS_ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)
export KARPENTER_NAMESPACE="kube-system"
export KARPENTER_VERSION="1.0.8"
3. Step 1: Associate IAM OIDC Provider with EKS
IAM Roles for Service Accounts (IRSA) requires your EKS cluster to have an associated OpenID Connect (OIDC) identity provider in AWS IAM. This allows the pod's projected ServiceAccount token to be validated by AWS Security Token Service (STS).
# Check and attach OIDC provider via eksctl
eksctl utils associate-iam-oidc-provider \
--cluster "${CLUSTER_NAME}" \
--region "${AWS_REGION}" \
--approve
# Verify OIDC endpoint in AWS IAM
export OIDC_ENDPOINT=$(aws eks describe-cluster --name "${CLUSTER_NAME}" --region "${AWS_REGION}" --query "cluster.identity.oidc.issuer" --output text | sed -e "s/^https:\/\///")
echo "OIDC Issuer: ${OIDC_ENDPOINT}"
4. Step 2: Create IAM Roles & Trust Policies for Karpenter
Karpenter requires two IAM roles:
- KarpenterControllerRole: Assumed by the controller pod via IRSA to manage EC2 instances, launch templates, and tags.
- KarpenterNodeRole: Instance profile attached to the newly provisioned worker nodes so kubelet can communicate with EKS.
cat <<EOF > controller-trust-policy.json
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::${AWS_ACCOUNT_ID}:oidc-provider/${OIDC_ENDPOINT}"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"${OIDC_ENDPOINT}:aud": "sts.amazonaws.com",
"${OIDC_ENDPOINT}:sub": "system:serviceaccount:${KARPENTER_NAMESPACE}:karpenter"
}
}
}
]
}
EOF
# Create Controller IAM Role
aws iam create-role \
--role-name "KarpenterControllerRole-${CLUSTER_NAME}" \
--assume-role-policy-document file://controller-trust-policy.json
Now attach the Karpenter controller policy granting granular EC2 and Pricing permissions:
cat <<EOF > controller-policy.json
{
"Version": "2012-10-17",
"Statement": [
{
"Action": [
"ec2:CreateLaunchTemplate",
"ec2:CreateFleet",
"ec2:RunInstances",
"ec2:CreateTags",
"ec2:TerminateInstances",
"ec2:DescribeLaunchTemplates",
"ec2:DescribeInstances",
"ec2:DescribeSecurityGroups",
"ec2:DescribeSubnets",
"ec2:DescribeInstanceTypes",
"ec2:DescribeInstanceTypeOfferings",
"ec2:DescribeAvailabilityZones",
"ec2:DescribeSpotPriceHistory",
"pricing:GetProducts",
"ssm:GetParameter"
],
"Effect": "Allow",
"Resource": "*"
},
{
"Action": "iam:PassRole",
"Effect": "Allow",
"Resource": "arn:aws:iam::${AWS_ACCOUNT_ID}:role/KarpenterNodeRole-${CLUSTER_NAME}"
}
]
}
EOF
aws iam put-role-policy \
--role-name "KarpenterControllerRole-${CLUSTER_NAME}" \
--policy-name KarpenterControllerPolicy \
--policy-document file://controller-policy.json
5. Step 3: Install Karpenter v1 Controller via Helm
Next, deploy Karpenter into your cluster via the official OCI Helm repository. Note that Karpenter runs as a deployment in kube-system and must be scheduled on existing management node groups (e.g. Fargate or static baseline nodes).
# Install Karpenter via Helm using OCI image
helm upgrade --install karpenter oci://public.ecr.aws/karpenter/karpenter \
--version "${KARPENTER_VERSION}" \
--namespace "${KARPENTER_NAMESPACE}" \
--set "serviceAccount.annotations.eks\.amazonaws\.com/role-arn=arn:aws:iam::${AWS_ACCOUNT_ID}:role/KarpenterControllerRole-${CLUSTER_NAME}" \
--set "settings.clusterName=${CLUSTER_NAME}" \
--set "settings.clusterEndpoint=$(aws eks describe-cluster --name ${CLUSTER_NAME} --query cluster.endpoint --output text)" \
--set "settings.interruptionQueueName=${CLUSTER_NAME}" \
--wait
# Verify Karpenter controller status
kubectl get pods -n kube-system -l app.kubernetes.io/name=karpenter
6. Step 4: Define EC2NodeClass & NodePool Custom Resources
In Karpenter v1, provisioning configuration is split cleanly into:
- EC2NodeClass: Infrastructure layer (Subnets, Security Groups, AMI Family, Instance Profile).
- NodePool: Scheduling layer (Instance Types, Spot vs On-Demand, Architecture, Disruption Budget).
apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
metadata:
name: default
spec:
amiFamily: AL2023
role: "KarpenterNodeRole-production-eks-cluster"
subnetSelectorTerms:
- tags:
karpenter.sh/discovery: "production-eks-cluster"
securityGroupSelectorTerms:
- tags:
karpenter.sh/discovery: "production-eks-cluster"
---
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: general-compute
spec:
template:
spec:
nodeClassRef:
group: karpenter.k8s.aws
kind: EC2NodeClass
name: default
requirements:
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]
- key: kubernetes.io/arch
operator: In
values: ["amd64", "arm64"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["c", "m", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["5"]
limits:
cpu: "1000"
memory: "4000Gi"
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 1m
karpenter.sh/discovery: production-eks-cluster. If this tag is missing, Karpenter will not be able to locate subnets to launch EC2 nodes.
7. Step 5: Verification & Smoke Test Scale-Out
Deploy a scale-test workload using the inflate deployment. This deployment requests 20 replicas, each with 1 CPU core and 1.5 GiB memory, forcing immediate compute scale-out:
apiVersion: apps/v1
kind: Deployment
metadata:
name: inflate
spec:
replicas: 0
selector:
matchLabels:
app: inflate
template:
metadata:
labels:
app: inflate
spec:
containers:
- name: inflate
image: public.ecr.aws/eks-distro/kubernetes/pause:3.7
resources:
requests:
cpu: "1"
memory: "1.5Gi"
Scale the deployment to 20 replicas and watch Karpenter launch nodes in real-time:
# Scale up
kubectl scale deployment inflate --replicas=20
# Stream Karpenter controller decisions
kubectl logs -f -n kube-system -l app.kubernetes.io/name=karpenter -c controller
# Check nodes arriving within 40 seconds
kubectl get nodes -l karpenter.sh/nodepool=general-compute -w
8. Step 6: Production Troubleshooting Runbook
1. Karpenter Controller Fails to Assume IAM Role (AccessDenied / 403 Forbidden)
- Root Cause: The OIDC thumbprint or ServiceAccount subject (
sub) claim does not match the trust policy. - Fix: Run
kubectl get sa karpenter -n kube-system -o yamland verify the annotationeks.amazonaws.com/role-arnexactly matchesarn:aws:iam::ACCOUNT_ID:role/KarpenterControllerRole-CLUSTER_NAME.
2. Nodes Created in AWS but Never Join the EKS Cluster
- Root Cause: The instance profile role is missing EKS cluster join permissions in AWS Auth or EKS Access Entries.
- Fix: For EKS clusters with API Access Entries, ensure you created an access entry:
aws eks create-access-entry --cluster-name $CLUSTER_NAME --principal-arn arn:aws:iam::$AWS_ACCOUNT_ID:role/KarpenterNodeRole-$CLUSTER_NAME --type EC2_LINUX
3. no subnets found with tags Warning in Logs
- Root Cause: Karpenter EC2NodeClass uses
karpenter.sh/discovery: CLUSTER_NAME, but subnets lack this exact key-value pair. - Fix: Tag subnets via AWS CLI:
aws ec2 create-tags --resources subnet-1234 subnet-5678 --tags Key=karpenter.sh/discovery,Value=$CLUSTER_NAME
Deepen your infrastructure mastery across our unified platform:
9. Frequently Asked Questions
capacity-type: ["spot"] with diversified instance types (e.g. c5, c6i, c6a, m5, m6i) across multiple availability zones and enabling the AWS EventBridge SQS interruption queue, Karpenter captures 2-minute Spot termination notices and drains workloads smoothly before termination.