- If there are unschedulable Pods in the cluster due to resource constraints, the cluster autoscaler adds new nodes to accommodate these Pods.
- If nodes in the cluster are underutilized, the cluster autoscaler removes these nodes in order to optimize resource usage and reduce costs.
Set up autoscaling for new node groups
You can set up autoscaling when creating a new node group:- Web console
- CLI
- Go SDK
- Python SDK
- JavaScript SDK
When creating a node group:
- Under Size, select Enable autoscaling.
- Specify the Min. nodes and Max. nodes numbers in the group.
Set up autoscaling for existing node groups
You cannot manage autoscaling for existing node groups in the web console. Use the CLI or an SDK instead.- CLI
- Go SDK
- Python SDK
- JavaScript SDK
To enable autoscaling for an existing node group, add the following parameters to the nebius mk8s node-group update command:For example, to set the autoscaling from 2 to 4 nodes, add
--autoscaling-min-node-count 2 --autoscaling-max-node-count 4 to the nebius mk8s node-group update command.Troubleshooting
More GPU nodes than required
- Issue: When a Managed Kubernetes cluster has the NVIDIA GPU Operator and the NVIDIA Network Operator installed, and workloads on a GPU node group are run with autoscaling, the cluster autoscaler can create more nodes in the group than the workloads require.
- Possible reason: A bug in Kubernetes Autoscaler that causes inconsistency in how nodes are considered ready or not ready for Pods. For more information, see Excess multiGPU nodes when using GPU + network operators in the Kubernetes Autoscaler repository on GitHub.
-
Solution:
- Uninstall the NVIDIA operators.
- Create a GPU node group and migrate your workloads to it. A node group created this way uses the GPU-adapted boot disk image offered by Managed Kubernetes, which solves the issue because the NVIDIA operators are no longer required.
See also
- Cluster autoscaler parameters in the official GitHub repository.