AWS / EKS Networking
// Interactive explainer — subnet fragmentation & pod IP capacity
The AWS VPC CNI (Container Network Interface) gives every Kubernetes pod a real VPC IP address. Overlay CNIs invent a private network that only the cluster understands. The VPC CNI does not. A pod IP is an ordinary subnet IP, routable from anything else in the VPC.
There is no single bottleneck. Four things together decide whether your subnet can handle prefix delegation:
Instance type
Sets the ceiling: how many ENI slots exist, and so how many /28 blocks a node could hold. A t3.medium has 15 slots; an m5.4xlarge has 232. That is the hardware limit, not what actually gets allocated.
Actual pod count
Decides how many /28 blocks are actually allocated.
The VPC CNI allocates on demand: ceil(pods / 16) + WARM_PREFIX_TARGET.
A node running 30 pods takes 3 blocks, not 15.
Subnet size
Decides how many /28 blocks are available. A /24 has only 14 usable blocks. A /22 has 62. A /20 has 254. Smaller subnets hit the wall sooner.
Nodes per subnet
Multiplies the demand. Every node takes its own /28 blocks. 5 nodes × 3 blocks each = 15 blocks, which is more than a /24 subnet holds. Spreading nodes across AZs helps.
The formula
t3.medium has 3 ENIs with 6 IPs each, so 5 usable slots per ENI
once the primary IP is taken.
A /24 subnet has 256 IPs, 251 usable, and 14 allocatable /28 blocks.
With 3 AZs, each subnet holds 1 node. This example uses max-pods=110, the value the EKS max-pods
calculator recommends here, plus the shipped CNI default WARM_PREFIX_TARGET=1.
Prefixes are dynamic: a node starts with fewer /28s and grows toward ceil(pods/16)+warm, which is 8 blocks at 110 pods.
| Default Mode | Prefix Delegation | |
|---|---|---|
| Max pods per node | 17 | 110 (max-pods cap) |
| Total pods (3 nodes) | 51 | 330 |
| Subnet IPs consumed per node | 18 | 130 (8 × 16 + 2 ENI IPs) |
| Subnet utilization (per AZ) | 7% (18 / 251) | 52% (130 / 251) |
| /28 blocks used (per subnet) | N/A | 8 / 14 (57%) |
| Room for another node? | Yes — 12+ more easily | No — 6 blocks left, a node needs 8 |
The key insight: In default mode, 1 pod costs 1 subnet IP. With prefix delegation the CNI allocates /28 blocks (16 IPs each), and a block holding a single pod still reserves all 16 IPs. In this t3.medium /24 example you gain about 6.5× the pod capacity but reserve about 7.2× the subnet IPs at full load. On a /24 or smaller, that trade can exhaust your address space fast.
An Elastic Network Interface (ENI) is a virtual network card you attach to an EC2 instance. Think of it as the NIC in a traditional server, virtualized and managed by AWS. Each ENI has its own private IP address, MAC address and security groups, and can also carry a public IP or Elastic IP.
Every EC2 instance launches with one primary
ENI (eth0).
You can attach more ENIs up to the instance type's limit. Each ENI can also hold secondary private IP addresses
on top of its primary IP. Those secondary IPs are what the VPC CNI hands out to pods.
ENIs give you VPC-native networking. Every IP on an ENI is a real, routable VPC address. No overlay, no NAT, no encapsulation. Pods reach RDS, ElastiCache and other VPC resources over plain VPC routing.
The VPC CNI plugin (the aws-node DaemonSet) runs on every
node. It pre-allocates ENIs and secondary IPs from the subnet.
When a pod starts, the CNI moves one of those IPs into the pod's network namespace
using a veth pair and Linux routing rules.
Performance & simplicity. VPC-native IPs avoid the encapsulation overhead of VXLAN overlays such as flannel or Calico. VPC Flow Logs capture pod traffic. AWS load balancers target pods directly in IP mode. Pods share the node's security groups; per-pod groups need the security groups for Pods feature.
eth0) is attached at launch. The
VPC CNI plugin starts and
allocates secondary IPs on this ENI for pods.
WARM_ENI_TARGET, default 1).
t3.medium gives 3 ENIs with 6 IPs each; an
m5.xlarge gives 4 ENIs with 15 IPs each.
You cannot raise these limits.
Prefix delegation lives inside the same slot budget, but each slot holds a /28 block of 16 IPs
instead of a single IP.
Each EC2 instance type caps how many ENIs it can attach and how many IPs each ENI can hold. Those two numbers set the node's pod capacity.
| Instance | Max ENIs | IPs/ENI | Default Pods | Prefix Pods | /28s at full load |
|---|---|---|---|---|---|
| t3.small | 3 | 4 | 11 | 110 | 8 |
| t3.medium | 3 | 6 | 17 | 110 | 8 |
| m5.large | 3 | 10 | 29 | 110 | 8 |
| m6g.large | 3 | 10 | 29 | 110 | 8 |
| m5.xlarge | 4 | 15 | 58 | 110 | 8 |
| m5.2xlarge | 4 | 15 | 58 | 110 | 8 |
| c5.4xlarge | 8 | 30 | 234 | 110 | 8 |
max-pods from the Default Pods column, so
prefix delegation only pays off once you raise it yourself.
The EKS max-pods calculator caps its answer at 110, or 250 on instances with more than 30 vCPUs,
so every instance in this table lands on 110.
The VPC CNI then allocates ceil(pods/16) +
WARM_PREFIX_TARGET /28 blocks, not one per ENI slot.
At 110 pods plus 1 warm prefix that is 8 /28 blocks (128 IPs) on every instance type.
A /28
fixes the first 28 bits of the address and leaves 4 bits for hosts. That is exactly
24 = 16 IP addresses.
With prefix delegation on, the VPC CNI asks the EC2 API for whole /28 blocks instead of single
secondary IPs. It calls AssignPrivateIpAddresses
with the Ipv4PrefixCount
parameter. Each block must be contiguous and naturally
aligned, so its first IP falls on a 16-IP boundary.
Prefixes only work on Nitro-based instance types, including bare metal, and need VPC CNI 1.9.0 or
later. IPv6 clusters always run in prefix mode.
Why /28 specifically?
/28 is the only IPv4 prefix length EC2 accepts; IPv6 prefixes are always /80. At 16 IPs it is small enough not to waste a whole subnet on a quiet node, and large enough to matter: a slot that held 1 IP now holds 16.
Alignment requirement
A /28 block must start at an IP whose last octet divides by 16 (0, 16, 32, 48...). AWS cannot carve one from an arbitrary starting IP. If the free IPs are not aligned into a contiguous run of 16, the allocation fails. That is what fragmentation means here.
Natural /28 boundaries in 10.0.1.0/24:
16 blocks × 16 IPs = 256. AWS reserves .0, .1, .2, .3 and .255, and they fall inside the first and the last block. Neither can be handed out as a prefix, so a /24 gives 14 allocatable /28 blocks.
AWS reserved IPs in every subnet
AWS reserves the first 4 and the last IP in every VPC subnet, whatever its size:
10.0.1.0/24 layout — each cell = 1 IP:
A valid /28 start IP must have its last octet divisible by 16. This is natural alignment: the address space fixes the boundaries, you do not get to pick them.
So you cannot shift a /28 block onto whichever IPs happen to be free. If 10.0.1.5–10.0.1.20 are free, that is 16 IPs, but they straddle two block boundaries (block 0 and block 1). AWS cannot build a /28 out of them. That is what makes fragmentation dangerous.
| Default mode | Prefix delegation | |
|---|---|---|
| Each ENI slot holds | 1 secondary IP | 1 × /28 prefix (16 IPs) |
| Subnet IPs per slot | 1 | 16 |
| IP allocation granularity | Individual IPs | 16-IP aligned blocks |
| Wasted IPs (1 pod on slot) | 0 | 15 |
| Fragmentation risk | Low | High |
| Pod density | Low (limited by ENI count) | High (16x more per slot) |
Increase the amount of available IP addresses for your Amazon EC2 nodes
Official EKS guide on enabling prefix delegation, configuration, and subnet sizing recommendations.
docs.aws.amazon.com/eks/latest/userguide/cni-increase-ip-addresses.htmlAmazon VPC CNI plugin for Kubernetes
The plugin README, with the defaults for ENABLE_PREFIX_DELEGATION, WARM_PREFIX_TARGET, WARM_IP_TARGET, MINIMUM_IP_TARGET and WARM_ENI_TARGET.
github.com/aws/amazon-vpc-cni-k8s/blob/master/README.mdElastic network interfaces (ENI) limits
Full table of ENI counts and IPv4 addresses per ENI for every EC2 instance type. Used to calculate max pod capacity.
docs.aws.amazon.com/AWSEC2/latest/UserGuide/AvailableIpPerENI.htmlAssignPrivateIpAddresses API
The EC2 API call the CNI uses to assign /28 prefixes. Documents the Ipv4PrefixCount parameter and alignment constraints.
docs.aws.amazon.com/AWSEC2/latest/APIReference/API_AssignPrivateIpAddresses.htmlVPC subnet sizing
Details on AWS-reserved IPs in each subnet and CIDR block sizing considerations for VPCs.
docs.aws.amazon.com/vpc/latest/userguide/subnet-sizing.htmlFragmentation is when a subnet has plenty of free IPs but they are scattered, so no aligned run of 16 is left. AWS then cannot allocate a new /28 prefix. It is disk fragmentation for IP space: the room exists, just not in usable chunks.
Why it happens
Nodes join and leave over time, taking and releasing /28 blocks at scattered positions. The freed IPs end up in gaps that are too small, or too misaligned, for a new /28.
Why it's dangerous
New nodes fail to start because the CNI cannot get a prefix. Monitoring still reports free IPs, while pods fail with IP exhaustion errors. That mismatch makes the fault hard to diagnose.
How to prevent it
Use larger subnets (/22 or bigger) so there are far more /28
boundaries to choose from. Better still, add a subnet CIDR reservation of type prefix; EC2 then draws prefixes from that
reserved space.
This simulates a 10.0.1.0/24
subnet: 256 IPs, 14 usable /28 blocks. To keep it simple, each node here claims a fixed 3
× /28 when it joins.
Real clusters allocate prefixes on demand as pods arrive. Click "Add
Node" to watch the subnet fill and then fragment.
Hover a cell to see its IP address.
The aws-node DaemonSet (VPC CNI) starts. It reads instance metadata to learn the ENI capacity and its prefix delegation settings.
The CNI calls AssignPrivateIpAddresses with Ipv4PrefixCount=N. AWS looks for N
aligned, contiguous 16-IP blocks in the subnet. If it cannot find them, the call fails.
Even with zero pods running, every IP in those /28 blocks counts as in use at the VPC level. No other ENI, on any instance, can take them. That is why prefix delegation exhausts subnets.
When a pod is scheduled, the CNI hands it an IP from a prefix it
already holds. There is no EC2 API call per pod, so it is fast. WARM_PREFIX_TARGET sets how many spare
prefixes to keep ready, unless WARM_IP_TARGET overrides it.
Nodes join and leave, and released /28 blocks leave gaps.
The subnet may hold 50+ free IPs, but with no aligned run of 16 the EC2 API returns InsufficientCidrBlocks. New nodes cannot
start, and pending pods sit in ContainerCreating.
Common symptoms, and the commands that tell /28 fragmentation apart from plain IP exhaustion:
Symptom: Pods stuck in ContainerCreating
kubectl describe pod shows: failed to assign an IP address to container
Symptom: aws-node logs show allocation failure
The CNI daemonset logs will contain errors about prefix allocation:
Diagnosis: Check subnet available IPs
If AvailableIpAddressCount looks healthy but nodes can't allocate, it's fragmentation — free IPs exist but not in aligned /28 blocks:
Diagnosis: View prefix assignments on a node
See which /28 prefixes are currently assigned to a node's ENIs:
Set your cluster parameters below. Results update as you type.
0 = show max only
Calculator cap: 110, or 250 above 30 vCPUs
Spare /28s kept ready. Default 1
Spare IPs. Overrides WARM_PREFIX_TARGET
Move to /22 or larger: 62 usable /28 blocks instead of 14. Create new subnets and move node groups onto them.
Attach a secondary CIDR (for example 100.64.0.0/16) to
your VPC. Put pod subnets there and leave your primary range for everything else.
If you do not need 100+ pods per node, go back to individual IPs. Create new node groups and drain the old ones; it is not a hot swap.
Fewer pods per node means fewer /28 blocks per node. Every block one node takes is a block no other node can use.
WARM_PREFIX_TARGET ships at
1 and cannot go to 0 while prefix delegation is on.
To waste less on quiet nodes, set WARM_IP_TARGET below 16 with
MINIMUM_IP_TARGET; both
override it.