gpu-b300-sxm), available in uk-south1 and eu-west2*gpu-b200-sxm and gpu-b200-sxm-a), available in us-central1 and me-west1 respectively:
| Preset name | Number of GPUs | Number of vCPUs | RAM, GiB |
| --------------------- | -------------- | --------------- | -------- |
| `1gpu-20vcpu-224gb` | 1 | 20 | 224 |
| `8gpu-160vcpu-1792gb` | 8 | 160 | 1792 |
* NVIDIA® H200 NVLink with Intel Sapphire Rapids (gpu-h200-sxm), available in eu-north1, eu-north2*eu-west1 and us-central1:
| Preset name | Number of GPUs | Number of vCPUs | RAM, GiB | [Regions](https://docs.nebius.com/overview/regions.md) |
| --------------------- | -------------- | --------------- | -------- | ----------------------------------------------------------------------------------------------- |
| `1gpu-16vcpu-200gb` | 1 | 16 | 200 | eu-north1, eu-west1, us-central1, eu-north2 |
| `8gpu-128vcpu-1600gb` | 8 | 128 | 1600 | eu-north1, eu-west1, us-central1, eu-north2 |
* NVIDIA® H100 NVLink with Intel Sapphire Rapids (gpu-h100-sxm), available in eu-north1:
| Preset name | Number of GPUs | Number of vCPUs | RAM, GiB |
| --------------------- | -------------- | --------------- | -------- |
| `1gpu-16vcpu-200gb` | 1 | 16 | 200 |
| `8gpu-128vcpu-1600gb` | 8 | 128 | 1600 |
* NVIDIA® RTX PRO™ 6000 with Intel Granite Rapids (gpu-rtx6000), available in us-central1:
| Preset name | Number of GPUs | Number of vCPUs | RAM, GiB |
| --------------------- | -------------- | --------------- | -------- |
| `1gpu-24vcpu-218gb` | 1 | 24 | 218 |
| `8gpu-192vcpu-1744gb` | 8 | 192 | 1744 |
* NVIDIA® L40S PCIe with Intel Ice Lake (gpu-l40s-a), available in eu-north1:
| Preset name | Number of GPUs | Number of vCPUs | RAM, GiB |
| ------------------- | -------------- | --------------- | -------- |
| `1gpu-8vcpu-32gb` | 1 | 8 | 32 |
| `1gpu-16vcpu-64gb` | 1 | 16 | 64 |
| `1gpu-24vcpu-96gb` | 1 | 24 | 96 |
| `1gpu-32vcpu-128gb` | 1 | 32 | 128 |
| `1gpu-40vcpu-160gb` | 1 | 40 | 160 |
* NVIDIA® L40S PCIe with AMD EPYC Genoa (gpu-l40s-d), available in eu-north1:
| Preset name | Number of GPUs | Number of vCPUs | RAM, GiB |
| --------------------- | -------------- | --------------- | -------- |
| `1gpu-16vcpu-96gb` | 1 | 16 | 96 |
| `1gpu-32vcpu-192gb` | 1 | 32 | 192 |
| `1gpu-48vcpu-288gb` | 1 | 48 | 288 |
| `2gpu-64vcpu-384gb` | 2 | 64 | 384 |
| `2gpu-96vcpu-576gb` | 2 | 96 | 576 |
| `4gpu-128vcpu-768gb` | 4 | 128 | 768 |
| `4gpu-192vcpu-1152gb` | 4 | 192 | 1152 |
### Presets compatible with GPU clusters
If you are adding a VM to a [GPU cluster](https://docs.nebius.com/compute/clusters/gpu/index.md), select from the following platforms and presets:
| Platform | Presets | [Regions](https://docs.nebius.com/overview/regions.md) |
| ----------------------------------------------------------------------- | --------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| NVIDIA® B300 NVLink with Intel Granite Rapids cpu-d3):
| Preset name | Number of vCPUs | RAM, GiB | [Regions](https://docs.nebius.com/overview/regions.md) |
| ---------------- | --------------- | -------- | ----------------------------------------- |
| `4vcpu-16gb` | 4 | 16 | All regions |
| `8vcpu-32gb` | 8 | 32 | All regions |
| `16vcpu-64gb` | 16 | 64 | All regions |
| `32vcpu-128gb` | 32 | 128 | All regions |
| `48vcpu-192gb` | 48 | 192 | All regions |
| `64vcpu-256gb` | 64 | 256 | All regions |
| `96vcpu-384gb` | 96 | 384 | All regions |
| `128vcpu-512gb` | 128 | 512 | All regions |
| `160vcpu-640gb` | 160 | 640 | All regions except eu-north1 |
| `192vcpu-768gb` | 192 | 768 | All regions except eu-north1 |
| `224vcpu-896gb` | 224 | 896 | All regions except eu-north1 |
| `256vcpu-1024gb` | 256 | 1024 | All regions except eu-north1 |
* Non-GPU Intel Ice Lake (cpu-e2), available in eu-north1:
| Preset name | Number of vCPUs | RAM, GiB |
| -------------- | --------------- | -------- |
| `2vcpu-8gb` | 2 | 8 |
| `4vcpu-16gb` | 4 | 16 |
| `8vcpu-32gb` | 8 | 32 |
| `16vcpu-64gb` | 16 | 64 |
| `32vcpu-128gb` | 32 | 128 |
| `48vcpu-192gb` | 48 | 192 |
| `64vcpu-256gb` | 64 | 256 |
| `80vcpu-320gb` | 80 | 320 |
## Compatibility with boot disk images
Nebius AI Cloud provides boot disk images for GPU and non-GPU VMs. The image that you choose for a VM must be compatible with the VM's platform. For compatibility details, see [Boot disk images for Compute virtual machines](https://docs.nebius.com/compute/storage/boot-disk-images.md).
## Compatibility with VM types
All platforms support regular VMs.
All platforms with GPUs also support [preemptible VMs](https://docs.nebius.com/compute/virtual-machines/preemptible.md).
To get an up-to-date list of platforms, run the `nebius compute platform list` [command](https://docs.nebius.com/cli/reference/compute/platform/list). The platforms available for preemptible VMs are marked as `allowed_for_preemptibles: true`.
# How to find out platforms and presets available in a project
Source: https://docs.nebius.com/compute/virtual-machines/list-platforms.md
Different virtual machine platforms and presets are available for different regions and projects. All supported platforms and presets are listed in [Types of virtual machines and GPUs in Nebius AI Cloud](https://docs.nebius.com/compute/virtual-machines/types.md).
## Prerequisites
nebius compute instance create command, see the [Nebius AI Cloud CLI reference](https://docs.nebius.com/cli/reference/compute/instance/create).
nebius\_compute\_v1\_instance Terraform resource, see the [provider reference](https://docs.nebius.com/terraform-provider/reference/resources/compute_v1_instance).
gpu-h100-sxm) | eu-north1 |
| `fabric-3` | NVIDIA® H100 NVLink with Intel Sapphire Rapids (gpu-h100-sxm) | eu-north1 |
| `fabric-4` | NVIDIA® H100 NVLink with Intel Sapphire Rapids (gpu-h100-sxm) | eu-north1 |
| `fabric-5` | NVIDIA® H200 NVLink with Intel Sapphire Rapids (gpu-h200-sxm) | eu-west1 |
| `fabric-6` | NVIDIA® H100 NVLink with Intel Sapphire Rapids (gpu-h100-sxm) | eu-north1 |
| `fabric-7` | NVIDIA® H200 NVLink with Intel Sapphire Rapids (gpu-h200-sxm) | eu-north1 |
| eu-north2-a | NVIDIA® H200 NVLink with Intel Sapphire Rapids (gpu-h200-sxm) | eu-north2\* |
| eu-west2-a | NVIDIA® B300 NVLink with Intel Granite Rapids (gpu-b300-sxm) | eu-west2\* |
| me-west1-a | NVIDIA® B200 NVLink with Intel Emerald Rapids (gpu-b200-sxm-a) | me-west1 |
| uk-south1-a | NVIDIA® B300 NVLink with Intel Granite Rapids (gpu-b300-sxm) | uk-south1 |
| us-central1-a | NVIDIA® H200 NVLink with Intel Sapphire Rapids (gpu-h200-sxm) | us-central1 |
| us-central1-b | NVIDIA® B200 NVLink with Intel Emerald Rapids (gpu-b200-sxm) | us-central1 |
{role} role within your tenant
{defaultGroup ? <>; for example, the default {defaultGroup} group> : ""}.
You can check this in the Administration → IAM section of the web console.
>;
};
Nebius AI Cloud's *capacity advisor* provides insights into GPU capacity availability for launching virtual machines (VMs) with specific hardware presets. It helps you understand where you can launch VMs based on your [quotas](https://docs.nebius.com/compute/resources/quotas-limits.md) and the current physical capacity in Nebius AI Cloud [regions](https://docs.nebius.com/overview/regions.md).
## Scope
The capacity advisor provides data for the following virtual machines with GPUs:
* **By service**:
* VMs that you create directly in Compute
* [Managed Soperator](https://docs.nebius.com/slurm-soperator/index.md) nodes
* [Managed Kubernetes®](https://docs.nebius.com/kubernetes/index.md) nodes
* VMs launched by [Serverless AI](https://docs.nebius.com/serverless/index.md) for running jobs and endpoints
* **By type**:
* [Regular VMs](https://docs.nebius.com/compute/virtual-machines/manage.md)
* [Preemptible VMs](https://docs.nebius.com/compute/virtual-machines/preemptible.md)
* [VMs with reservations](https://docs.nebius.com/compute/virtual-machines/reservations.md)
* [Container VMs](https://docs.nebius.com/compute/virtual-machines/containers.md)
Data about computing resource availability of [standalone applications](https://docs.nebius.com/applications/types.md) and VMs without GPUs isn't provided.
| Component | `ubuntu24.04-cuda12` | `ubuntu24.04-cuda13.0` |
|---|---|---|
| Drivers preset | `cuda12.8` | `cuda13.0` |
| CUDA Toolkit | 12.8 ([release notes](https://docs.nvidia.com/cuda/archive/12.8.0/)) | 13.0 ([release notes](https://docs.nvidia.com/cuda/archive/13.0.0/)) |
| NVIDIA® Data Center GPU Driver | 570.x | 580.x |
| Linux kernel | 6.11, NVIDIA® HWE | 6.11, NVIDIA® HWE |
| Networking package | [NVIDIA® DOCA](https://developer.nvidia.com/networking/doca) 2.9.2 ([release notes](https://docs.nvidia.com/doca/archive/2-9-2-lts-ovs-update/doca+release+notes/index.html)) | NVIDIA® DOCA 3.1.0 ([release notes](https://docs.nvidia.com/doca/archive/3-1-0/doca+release+notes/index.html)) |
| Other components |
|
|
nebius compute instance list.
2. Create an image:
```bash
nebius compute image create \
--name nebius compute image list.
* To edit your image (for example, rename it), run:
```bash
nebius compute image update computeimage-e***
```
For more information about the command parameters, see the [command reference](https://docs.nebius.com/cli/reference/compute/image/update).
* To delete your image, run:
```bash
nebius compute image delete computeimage-e***
```
nebius compute image list.
2. [Create a VM](https://docs.nebius.com/compute/virtual-machines/manage.md#create-a-vm) with the new disk.
Use the ID of the boot disk created earlier with the `--boot-disk-existing-disk-id` parameter.
nebius compute filesystem list.
* `mount_tag`: Tag for [mounting a filesystem to a VM](https://docs.nebius.com/compute/storage/use.md#shared-filesystems).
Create your own tag, such as `my-filesystem`. Make sure that it is unique within a VM.
If you do not specify the tag, it takes the `filesystem-N` default value where `N` is an integer. For example, `filesystem-0`.
3. Create or update a VM with the volumes configured:
* To **create a VM**, run the following command:
```bash
nebius compute instance create \
--secondary-disks "$(cat secondary_disk.json)" \
--filesystems "$(cat filesystem.json)" \
nebius compute instance list.
The `--patch` parameter allows you to only update the specified parameters. Without it, the command resets the VM settings to their default values.
* To **attach shared filesystems to an existing VM**, stop the VM first:
```bash
nebius compute instance stop --id nebius compute instance list. To get the disk ID, run nebius compute disk list.
nebius compute instance list. To get the filesystem ID, run nebius compute filesystem list.
nebius compute disk list or nebius compute disk get \. If the disk is used on a VM, the output contains the VM's ID in the `.status.read_write_attachment` field:
```yaml
status:
read_write_attachment: computeinstance-e00***
```
For more details about the commands, see the references for [nebius compute disk list](https://docs.nebius.com/cli/reference/compute/disk/list) and [nebius compute disk get](https://docs.nebius.com/cli/reference/compute/disk/get).
* For a filesystem, run nebius compute filesystem list or nebius compute filesystem get \. If the filesystem is used on any VMs, the output lists their IDs in the `.status.read_write_attachments` field:
```yaml
status:
read_write_attachments:
- computeinstance-e00***
- computeinstance-e00***
```
For more details about the commands, see the references for [nebius compute filesystem list](https://docs.nebius.com/cli/reference/compute/filesystem/list) and [nebius compute filesystem get](https://docs.nebius.com/cli/reference/compute/filesystem/get).
nebius compute instance update with the `--network-interfaces` parameter to add a public address to the VM. For the command reference, see [nebius compute instance update](https://docs.nebius.com/cli/reference/compute/instance/update).
* Create another VM with a public address and add the volume to it. For details and examples, see [How to create a virtual machine in Nebius AI Cloud](https://docs.nebius.com/compute/virtual-machines/manage.md).
eu-west1 region, you can try NVIDIA® H100 NVLink with Intel Sapphire Rapids in eu-north1.
For VMs with GPUs, you can use the [capacity advisor](https://docs.nebius.com/compute/virtual-machines/capacity-advisor.md) to get information about the availability of computing resources based on your quotas and the current physical capacity.
| **Command parameter** | **Description** | **Environment variable** | **#SBATCH directive** | **Slurm default** |
| `--job-name= | The name of your job. | `SBATCH_JOB_NAME` | `#SBATCH --job-name= | `job-name= |
| `--nodes= | The number of worker nodes to allocate for the job. | *N/A* | `#SBATCH --nodes= | `nodes= |
`--nodelist=
| The list of specific worker nodes to allocate for the job. For example, `--nodelist="worker-0,worker-2"` or `--nodelist="worker-[0-2,3]"`. | *N/A* | `#SBATCH --nodelist=
| `nodelist=
|
`--exclude=
| The list of worker nodes to exclude from the job allocation. For example, `--exclude="worker-3"` or `--nodelist="worker-[4-7]"`. | *N/A* | `#SBATCH --exclude=
| `exclude=
|
| `--output= | The path to the file for the job's standard output. The path can contain special replacement symbols; for example, `%j` is replaced by the job ID. For more details, see the [Slurm documentation](https://slurm.schedmd.com/sbatch.html#SECTION_FILENAME-PATTERN). | `SBATCH_OUTPUT` | `#SBATCH --output= | `output= |
| `--error= | The path to the file for the job's error output. The path can contain the replacement symbols as described for the `output` setting. | `SBATCH_ERROR` | `#SBATCH --error= | `error= |
| `--time= | The time limit for the job. When the job reaches the time limit, all its tasks (processes) are terminated. For example, `01:00` limits the job to one hour, and `1-00` limits the job to one day (24 hours). | `SBATCH_TIMELIMIT` | `#SBATCH --time= | `time= |
| `--gpus-per-node= | The number of GPUs to allocate for the job on each worker node. | `SBATCH_GPUS_PER_NODE` | `#SBATCH --gpus-per-node= | `gpus-per-node= |
| `--ntasks-per-node= | The maximum number of tasks to run on each worker node. When you define resources for the job in per-task settings like `gpus-per-task`, `cpus-per-task`, etc., total resources in the job allocation are based on the value of `ntasks-per-node`. | *N/A* | `#SBATCH --ntasks-per-node= | `ntasks-per-node= |
| `--exclusive` | Allocates all CPUs on the allocated worker nodes to the job, preventing other jobs from using these nodes. This allows the total number of tasks of the job to be unlimited. | `SBATCH_EXCLUSIVE` | `#SBATCH --exclusive` | *N/A* |
| `--cpus-per-task= | The number of CPUs to allocate for the job per task. | *N/A* | `#SBATCH --cpus-per-task= | `cpus-per-task= |
| `--mem= | The RAM size to allocate for the job on each worker node. For example, `mem=4G`. To allocate all available RAM, specify `mem=0`. | `SBATCH_MEM_PER_NODE` | `#SBATCH --mem= | `mem= |
| `--partition= | The Slurm partition to allocate nodes from. | `SBATCH_PARTITION` | `#SBATCH --partition= | `partition= |
| `--account= | The Slurm account name. | `SBATCH_ACCOUNT` | `#SBATCH --account= | `account= |
| `--requeue` | Requeues the job automatically: restarts it (with the same ID) when its worker nodes fail or other, higher-priority jobs preempt them. | `SBATCH_REQUEUE` | `#SBATCH --requeue` | `requeue` |
| `--no-requeue` | Disables automatically requeuing the job (see `requeue`). | `SBATCH_NO_REQUEUE` | `#SBATCH --no-requeue` | `no-requeue` |
`--dependency=
| Dependencies of the job. For example:
| *N/A* | `#SBATCH --dependency=
| `dependency=
|
| `--parsable` | Changes the standard output of `sbatch` from `Submitted batch job | *N/A* | `#SBATCH --parsable` | `parsable` |
| `--verbose` or `-v` | Increases the verbosity of `sbatch`'s informational messages. For more verbosity, use the parameter, `#SBATCH` directive or Slurm default multiple times, or set the `SBATCH_DEBUG` environment variable to `2`, `3`, etc. | `SBATCH_DEBUG` | `#SBATCH --verbose` or `#SBATCH -v` | `verbose` or `v` |
cr.eu-north1.nebius.cloud: For the eu-north1 region.
* cr.eu-west1.nebius.cloud: For the eu-west1 region.
For more information, see [Container Registry documentation](https://docs.nebius.com/container-registry/registries/manage.md) and [Enroot documentation](https://github.com/NVIDIA/enroot/blob/master/doc/cmd/import.md#description).
## How to run a job for a local image
nebius mk8s cluster list.
2. Update the cluster settings:
```bash
nebius mk8s cluster update \
--id $K8S_CLUSTER_ID \
--labels Managed Kubernetes tries to launch a node group in any available and suitable reservation. If none is found, resources for all nodes in the node group are provided from the common pool, not from a reservation.
If a reservation doesn't have enough capacity for the whole node group, it uses all the resources available in this reservation and also takes resources from the common pool. In other words, resources for some nodes are provided from the reservation, and resources for the rest of the nodes are provided from the common pool.
| | | `auto` | Specified | Managed Kubernetes tries to launch a node group in one of the specified reservations. If none of them fit (for example, they are not currently active or there are not enough GPUs), the same logic of the `auto` policy applies. | | | `forbid` | Not specified | Node group resources are provided from the common pool. No reservations are used. | | | `forbid` | Specified | Not supported. If you apply this combination, it will result in a validation error. | | | `strict` | Not specified | Managed Kubernetes tries to launch a node group in any available and suitable reservation. If none is found, a request for creating or updating a node group fails. | | | `strict` | Specified | Managed Kubernetes tries to launch a node group in one of the specified reservations. If none of them fit, a request for creating or updating a node group fails. | |Managed Kubernetes tries to launch a node group in any available and suitable reservation. If none is found, resources for all nodes in the node group are provided from the common pool, not from a reservation.
If a reservation doesn't have enough capacity for the whole node group, it uses all the resources available in this reservation and also takes resources from the common pool. In other words, resources for some nodes are provided from the reservation, and resources for the rest of the nodes are provided from the common pool.
| | `AUTO` | Specified | Managed Kubernetes tries to launch a node group in one of the specified reservations. If none of them fit (for example, they are not currently active or there are not enough GPUs), the same logic of the `AUTO` policy applies. | | `FORBID` | Not specified | Node group resources are provided from the common pool. No reservations are used. | | `FORBID` | Specified | Not supported. If you apply this combination, it will result in a validation error. | | `STRICT` | Not specified | Managed Kubernetes tries to launch a node group in any available and suitable reservation. If none is found, a request for creating or updating a node group fails. | | `STRICT` | Specified | Managed Kubernetes tries to launch a node group in one of the specified reservations. If none of them fit, a request for creating or updating a node group fails. |nebius mk8s node-group create, see the [CLI reference](https://docs.nebius.com/cli/reference/mk8s/node-group/create).
nebius\_mk8s\_v1\_node\_group Terraform resource, see the [provider reference](https://docs.nebius.com/terraform-provider/reference/resources/mk8s_v1_node_group).
cr.eu-north1.nebius.cloud/\/nginx:mynginx (you can get the registry ID in the web console or with the [nebius registry list](https://docs.nebius.com/cli/reference/registry/list)) CLI command), here is how to refer to it in a deployment manifest:
>
> ```yaml
> apiVersion: apps/v1
> kind: Deployment
> metadata:
> name: nginx-deployment
> spec:
> replicas: 1
> selector:
> matchLabels:
> app: nginx
> template:
> metadata:
> labels:
> app: nginx
> spec:
> containers:
> - name: nginx
> image: cr.eu-north1.nebius.cloud/gpu-h100-sxm VM platform with the `8gpu-128vcpu-1600gb` preset.
* The nodes use a boot disk image offered by Managed Kubernetes that contains drivers and other components for GPUs. Without this image, you need to install the drivers and components manually. For more details, see [GPU drivers and other components](https://docs.nebius.com/kubernetes/gpu/set-up.md#gpu-drivers-and-other-components).
4. Generate a kubeconfig file with the cluster details for kubectl:
```bash
nebius mk8s cluster get-credentials \
--id $MK8S_CLUSTER_ID --external
```
To verify that kubectl is connected to the cluster, you can run `kubectl cluster-info`.
### Run the NCCL tests
1. Install the [Kubeflow Training Operator](https://www.kubeflow.org/docs/components/trainer/legacy-v1/overview/) (also known as Kubeflow Trainer).
```bash
kubectl apply --server-side -k "github.com/kubeflow/training-operator/manifests/overlays/standalone?ref=v1.9.3"
```
2. Create a namespace for the tests, named `nccl-test` in this tutorial:
```
kubectl create ns nccl-test
```
3. Create `nccl-test.yaml` with an `MPIJob` for your tests.