AWS Deep Dives

AWS / EKS Networking

VPC CNI: Prefix Delegation

// Interactive explainer — subnet fragmentation & pod IP capacity

What is the VPC CNI?

The AWS VPC CNI (Container Network Interface) gives every Kubernetes pod a real VPC IP address. Overlay CNIs invent a private network that only the cluster understands. The VPC CNI does not. A pod IP is an ordinary subnet IP, routable from anything else in the VPC.

Every pod consumes a real subnet IP — your subnet size directly limits your total pod count.

What Actually Determines Your Subnet IP Usage

There is no single bottleneck. Four things together decide whether your subnet can handle prefix delegation:

1

Instance type

Sets the ceiling: how many ENI slots exist, and so how many /28 blocks a node could hold. A t3.medium has 15 slots; an m5.4xlarge has 232. That is the hardware limit, not what actually gets allocated.

2

Actual pod count

Decides how many /28 blocks are actually allocated. The VPC CNI allocates on demand: ceil(pods / 16) + WARM_PREFIX_TARGET. A node running 30 pods takes 3 blocks, not 15.

3

Subnet size

Decides how many /28 blocks are available. A /24 has only 14 usable blocks. A /22 has 62. A /20 has 254. Smaller subnets hit the wall sooner.

4

Nodes per subnet

Multiplies the demand. Every node takes its own /28 blocks. 5 nodes × 3 blocks each = 15 blocks, which is more than a /24 subnet holds. Spreading nodes across AZs helps.

The formula

blocks_needed = nodes_per_subnet × (ceil(actual_pods / 16) + WARM_PREFIX_TARGET)
fits? = blocks_needed ≤ usable /28 blocks in subnet
Example A: 2 nodes × 30 pods × /24 → 2 × (2+1) = 6 blocks ≤ 14 ✓
Example B: 2 nodes × 110 pods × /24 → 2 × (7+1) = 16 blocks > 14 ✗
Example C: 5 nodes × 30 pods × /22 → 5 × (2+1) = 15 blocks ≤ 62 ✓

Concrete Example: 3 × t3.medium on /24

t3.medium has 3 ENIs with 6 IPs each, so 5 usable slots per ENI once the primary IP is taken. A /24 subnet has 256 IPs, 251 usable, and 14 allocatable /28 blocks. With 3 AZs, each subnet holds 1 node. This example uses max-pods=110, the value the EKS max-pods calculator recommends here, plus the shipped CNI default WARM_PREFIX_TARGET=1. Prefixes are dynamic: a node starts with fewer /28s and grows toward ceil(pods/16)+warm, which is 8 blocks at 110 pods.

Default Mode Prefix Delegation
Max pods per node 17 110 (max-pods cap)
Total pods (3 nodes) 51 330
Subnet IPs consumed per node 18 130 (8 × 16 + 2 ENI IPs)
Subnet utilization (per AZ) 7% (18 / 251) 52% (130 / 251)
/28 blocks used (per subnet) N/A 8 / 14 (57%)
Room for another node? Yes — 12+ more easily No — 6 blocks left, a node needs 8
The /24 squeeze with prefix delegation At 110 pods, one t3.medium needs 8 /28 blocks, or 128 IPs. A /24 has only 14 blocks to give. One node takes 8 of them, so a second node cannot fit. The same node in default mode uses 18 IPs.

The key insight: In default mode, 1 pod costs 1 subnet IP. With prefix delegation the CNI allocates /28 blocks (16 IPs each), and a block holding a single pod still reserves all 16 IPs. In this t3.medium /24 example you gain about 6.5× the pod capacity but reserve about 7.2× the subnet IPs at full load. On a /24 or smaller, that trade can exhaust your address space fast.

Use Prefix Delegation When

• You need >17 pods per node (or more than your instance's default limit)
• Your subnets are /22 or larger (62 usable /28 blocks)
• You run dense workloads — many small containers per node
• You can use a secondary CIDR (e.g. 100.64.0.0/16) for pod IPs

Don't Use Prefix Delegation When

• Your subnets are /24 or smaller — you'll exhaust /28 blocks instantly
• You run fewer than 20 pods per node — default mode is sufficient
• You value subnet IP efficiency over pod density
• You have many nodes sharing small subnets — fragmentation will bite

What is an Elastic Network Interface (ENI)?

An Elastic Network Interface (ENI) is a virtual network card you attach to an EC2 instance. Think of it as the NIC in a traditional server, virtualized and managed by AWS. Each ENI has its own private IP address, MAC address and security groups, and can also carry a public IP or Elastic IP.

Every EC2 instance launches with one primary ENI (eth0). You can attach more ENIs up to the instance type's limit. Each ENI can also hold secondary private IP addresses on top of its primary IP. Those secondary IPs are what the VPC CNI hands out to pods.

Why does AWS use ENIs?

ENIs give you VPC-native networking. Every IP on an ENI is a real, routable VPC address. No overlay, no NAT, no encapsulation. Pods reach RDS, ElastiCache and other VPC resources over plain VPC routing.

How does it work?

The VPC CNI plugin (the aws-node DaemonSet) runs on every node. It pre-allocates ENIs and secondary IPs from the subnet. When a pod starts, the CNI moves one of those IPs into the pod's network namespace using a veth pair and Linux routing rules.

Why is it useful?

Performance & simplicity. VPC-native IPs avoid the encapsulation overhead of VXLAN overlays such as flannel or Calico. VPC Flow Logs capture pod traffic. AWS load balancers target pods directly in IP mode. Pods share the node's security groups; per-pod groups need the security groups for Pods feature.

ENI Lifecycle on a Kubernetes Node

1. Node boots
The primary ENI (eth0) is attached at launch. The VPC CNI plugin starts and allocates secondary IPs on this ENI for pods.
2. Warm pool fills
The CNI keeps a "warm pool" of spare IPs. When it runs low, the CNI asks the EC2 API for more secondary IPs or prefixes, and attaches a new ENI if the current ones are full.
3. Pod scheduled
The container runtime calls the CNI plugin. It takes an IP from the warm pool, creates a veth pair, and sets up routing so the pod can send and receive traffic on that VPC IP.
4. Pod terminates
The IP returns to the warm pool after a 30 second cooldown. If enough IPs stay free, the CNI may detach an ENI and give its IPs back to the subnet (controlled by WARM_ENI_TARGET, default 1).
Key limitation: The number of ENIs and IPs per ENI is fixed by the EC2 instance type. A t3.medium gives 3 ENIs with 6 IPs each; an m5.xlarge gives 4 ENIs with 15 IPs each. You cannot raise these limits. Prefix delegation lives inside the same slot budget, but each slot holds a /28 block of 16 IPs instead of a single IP.

How ENIs Work on EC2

Each EC2 instance type caps how many ENIs it can attach and how many IPs each ENI can hold. Those two numbers set the node's pod capacity.

EC2 Node
t3.medium
ENI 0 (primary)
eth0 — node IP
ENI 1
eth1
ENI 2
eth2

Default: Individual IPs

ENI 0: 1 reserved + 5 pod IPs
ENI 1: 1 reserved + 5 pod IPs
ENI 2: 1 reserved + 5 pod IPs
15 pod IPs + 2 host-network pods = 17 pods max

Prefix Delegation: /28 Blocks

ENI 0: 1 reserved + 5 prefix slots
ENI 1: 1 reserved + 5 prefix slots
ENI 2: 1 reserved + 5 prefix slots
Each slot = 1 × /28 = 16 IPs
15 slots × 16 + 2 = 242 pods; capped at 110

Instance Type Comparison

Instance Max ENIs IPs/ENI Default Pods Prefix Pods /28s at full load
t3.small 3 4 11 110 8
t3.medium 3 6 17 110 8
m5.large 3 10 29 110 8
m6g.large 3 10 29 110 8
m5.xlarge 4 15 58 110 8
m5.2xlarge 4 15 58 110 8
c5.4xlarge 8 30 234 110 8
Key insight: EKS sets max-pods from the Default Pods column, so prefix delegation only pays off once you raise it yourself. The EKS max-pods calculator caps its answer at 110, or 250 on instances with more than 30 vCPUs, so every instance in this table lands on 110. The VPC CNI then allocates ceil(pods/16) + WARM_PREFIX_TARGET /28 blocks, not one per ENI slot. At 110 pods plus 1 warm prefix that is 8 /28 blocks (128 IPs) on every instance type.

What is a /28 Block?

A /28 fixes the first 28 bits of the address and leaves 4 bits for hosts. That is exactly 24 = 16 IP addresses.

With prefix delegation on, the VPC CNI asks the EC2 API for whole /28 blocks instead of single secondary IPs. It calls AssignPrivateIpAddresses with the Ipv4PrefixCount parameter. Each block must be contiguous and naturally aligned, so its first IP falls on a 16-IP boundary. Prefixes only work on Nitro-based instance types, including bare metal, and need VPC CNI 1.9.0 or later. IPv6 clusters always run in prefix mode.

Why /28 specifically?

/28 is the only IPv4 prefix length EC2 accepts; IPv6 prefixes are always /80. At 16 IPs it is small enough not to waste a whole subnet on a quiet node, and large enough to matter: a slot that held 1 IP now holds 16.

Alignment requirement

A /28 block must start at an IP whose last octet divides by 16 (0, 16, 32, 48...). AWS cannot carve one from an arbitrary starting IP. If the free IPs are not aligned into a contiguous run of 16, the allocation fails. That is what fragmentation means here.

Natural /28 boundaries in 10.0.1.0/24:

10.0.1.0–1510.0.1.16–3110.0.1.32–4710.0.1.48–63 10.0.1.64–7910.0.1.80–9510.0.1.96–11110.0.1.112–127 10.0.1.128–14310.0.1.144–15910.0.1.160–17510.0.1.176–191 10.0.1.192–20710.0.1.208–22310.0.1.224–23910.0.1.240–255

16 blocks × 16 IPs = 256. AWS reserves .0, .1, .2, .3 and .255, and they fall inside the first and the last block. Neither can be handed out as a prefix, so a /24 gives 14 allocatable /28 blocks.

AWS reserved IPs in every subnet

AWS reserves the first 4 and the last IP in every VPC subnet, whatever its size:

10.0.1.0 — Network address
10.0.1.1 — VPC router
10.0.1.2 — DNS server
10.0.1.3 — Reserved for future use
10.0.1.255 — Broadcast (not supported in VPC, but reserved)

10.0.1.0/24 layout — each cell = 1 IP:

AWS Reserved
Block A (Node 1, ENI 1)
Block B (Node 1, ENI 2)
Block C (Node 2)
Free

/28 Alignment Math

A valid /28 start IP must have its last octet divisible by 16. This is natural alignment: the address space fixes the boundaries, you do not get to pick them.

Formula: IP_last_octet mod 16 = 0 → valid /28 start 10.0.1.0 → 0 mod 16 = 0  ✓ valid
10.0.1.16 → 16 mod 16 = 0 ✓ valid
10.0.1.32 → 32 mod 16 = 0 ✓ valid
10.0.1.7 → 7 mod 16 = 7  ✗ NOT a valid /28 start
10.0.1.20 → 20 mod 16 = 4 ✗ NOT a valid /28 start

So you cannot shift a /28 block onto whichever IPs happen to be free. If 10.0.1.5–10.0.1.20 are free, that is 16 IPs, but they straddle two block boundaries (block 0 and block 1). AWS cannot build a /28 out of them. That is what makes fragmentation dangerous.

How Prefix Delegation Differs from Default Mode

Default mode Prefix delegation
Each ENI slot holds 1 secondary IP 1 × /28 prefix (16 IPs)
Subnet IPs per slot 1 16
IP allocation granularity Individual IPs 16-IP aligned blocks
Wasted IPs (1 pod on slot) 0 15
Fragmentation risk Low High
Pod density Low (limited by ENI count) High (16x more per slot)

AWS Documentation References

Increase the amount of available IP addresses for your Amazon EC2 nodes

Official EKS guide on enabling prefix delegation, configuration, and subnet sizing recommendations.

docs.aws.amazon.com/eks/latest/userguide/cni-increase-ip-addresses.html

Amazon VPC CNI plugin for Kubernetes

The plugin README, with the defaults for ENABLE_PREFIX_DELEGATION, WARM_PREFIX_TARGET, WARM_IP_TARGET, MINIMUM_IP_TARGET and WARM_ENI_TARGET.

github.com/aws/amazon-vpc-cni-k8s/blob/master/README.md

Elastic network interfaces (ENI) limits

Full table of ENI counts and IPv4 addresses per ENI for every EC2 instance type. Used to calculate max pod capacity.

docs.aws.amazon.com/AWSEC2/latest/UserGuide/AvailableIpPerENI.html

AssignPrivateIpAddresses API

The EC2 API call the CNI uses to assign /28 prefixes. Documents the Ipv4PrefixCount parameter and alignment constraints.

docs.aws.amazon.com/AWSEC2/latest/APIReference/API_AssignPrivateIpAddresses.html

VPC subnet sizing

Details on AWS-reserved IPs in each subnet and CIDR block sizing considerations for VPCs.

docs.aws.amazon.com/vpc/latest/userguide/subnet-sizing.html

What is Subnet Fragmentation?

Fragmentation is when a subnet has plenty of free IPs but they are scattered, so no aligned run of 16 is left. AWS then cannot allocate a new /28 prefix. It is disk fragmentation for IP space: the room exists, just not in usable chunks.

Why it happens

Nodes join and leave over time, taking and releasing /28 blocks at scattered positions. The freed IPs end up in gaps that are too small, or too misaligned, for a new /28.

Why it's dangerous

New nodes fail to start because the CNI cannot get a prefix. Monitoring still reports free IPs, while pods fail with IP exhaustion errors. That mismatch makes the fault hard to diagnose.

How to prevent it

Use larger subnets (/22 or bigger) so there are far more /28 boundaries to choose from. Better still, add a subnet CIDR reservation of type prefix; EC2 then draws prefixes from that reserved space.

Example (simplified demo): A /24 subnet has 251 usable IPs and 14 allocatable /28 blocks. 4 nodes taking 3 blocks each use 12, leaving 2. A 5th node asking for 3 blocks fails, even though 59 individual IPs are still free.

Interactive Demo

This simulates a 10.0.1.0/24 subnet: 256 IPs, 14 usable /28 blocks. To keep it simple, each node here claims a fixed 3 × /28 when it joins. Real clusters allocate prefixes on demand as pods arrive. Click "Add Node" to watch the subnet fill and then fragment. Hover a cell to see its IP address.

Press "Add Node" to start →

Step-by-Step: What Happens When a Node Joins

1
kubelet starts on new EC2 node

The aws-node DaemonSet (VPC CNI) starts. It reads instance metadata to learn the ENI capacity and its prefix delegation settings.

2
CNI requests /28 prefix assignments from the EC2 API

The CNI calls AssignPrivateIpAddresses with Ipv4PrefixCount=N. AWS looks for N aligned, contiguous 16-IP blocks in the subnet. If it cannot find them, the call fails.

3
AWS reserves all 16 IPs in each /28 block

Even with zero pods running, every IP in those /28 blocks counts as in use at the VPC level. No other ENI, on any instance, can take them. That is why prefix delegation exhausts subnets.

4
Pods schedule and get IPs from the warm pool

When a pod is scheduled, the CNI hands it an IP from a prefix it already holds. There is no EC2 API call per pod, so it is fast. WARM_PREFIX_TARGET sets how many spare prefixes to keep ready, unless WARM_IP_TARGET overrides it.

Fragmentation — new nodes fail to join

Nodes join and leave, and released /28 blocks leave gaps. The subnet may hold 50+ free IPs, but with no aligned run of 16 the EC2 API returns InsufficientCidrBlocks. New nodes cannot start, and pending pods sit in ContainerCreating.

How to Spot Fragmentation in Production

Common symptoms, and the commands that tell /28 fragmentation apart from plain IP exhaustion:

Symptom: Pods stuck in ContainerCreating

kubectl describe pod shows: failed to assign an IP address to container

kubectl get pods --field-selector=status.phase!=Running -A

Symptom: aws-node logs show allocation failure

The CNI daemonset logs will contain errors about prefix allocation:

kubectl logs -n kube-system -l k8s-app=aws-node --tail=50 | grep -i "failed\|error\|prefix"

Diagnosis: Check subnet available IPs

If AvailableIpAddressCount looks healthy but nodes can't allocate, it's fragmentation — free IPs exist but not in aligned /28 blocks:

aws ec2 describe-subnets --subnet-ids <id> \
  --query 'Subnets[].[SubnetId,AvailableIpAddressCount,CidrBlock]'

Diagnosis: View prefix assignments on a node

See which /28 prefixes are currently assigned to a node's ENIs:

aws ec2 describe-network-interfaces \
  --filters Name=attachment.instance-id,Values=<instance-id> \
  --query 'NetworkInterfaces[].[Ipv4Prefixes,PrivateIpAddresses]'

VPC CNI Capacity Calculator

Set your cluster parameters below. Results update as you type.

0 = show max only

Solutions & Mitigations

Expand your subnets Recommended

Move to /22 or larger: 62 usable /28 blocks instead of 14. Create new subnets and move node groups onto them.

Use a secondary CIDR block

Attach a secondary CIDR (for example 100.64.0.0/16) to your VPC. Put pod subnets there and leave your primary range for everything else.

Disable prefix delegation

If you do not need 100+ pods per node, go back to individual IPs. Create new node groups and drain the old ones; it is not a hot swap.

Use smaller instance types

Fewer pods per node means fewer /28 blocks per node. Every block one node takes is a block no other node can use.

Tune WARM_PREFIX_TARGET

WARM_PREFIX_TARGET ships at 1 and cannot go to 0 while prefix delegation is on. To waste less on quiet nodes, set WARM_IP_TARGET below 16 with MINIMUM_IP_TARGET; both override it.

Useful AWS CLI Commands

# Check if prefix delegation is enabled
kubectl describe daemonset aws-node -n kube-system | grep ENABLE_PREFIX_DELEGATION

# View available IPs in your subnet
aws ec2 describe-subnets --subnet-ids <subnet-id> --query 'Subnets[].AvailableIpAddressCount'

# Disable prefix delegation (drain and recycle nodes after!)
kubectl set env daemonset aws-node -n kube-system ENABLE_PREFIX_DELEGATION=false

# Check ENI assignments on a node
aws ec2 describe-network-interfaces --filters Name=attachment.instance-id,Values=<instance-id>