Serverless AI jobs run container images as one-off or scheduled batch workloads. They are suitable for training, fine-tuning and data processing where you want to use computing resources only to perform a task and stop when the task is done. Each job runs on a Compute container virtual machine (VM) managed by Serverless AI and billed only while the job is running.
Prerequisites
Make sure you are in a group that has at least the editor role within your tenant or project; for example, the default editors group. You can check this in the Administration → IAM section of the web console.
-
Install and configure the Nebius AI Cloud CLI.
Check that your project ID is saved in the Nebius AI Cloud CLI profile configuration:
-
Make sure you are in a group that has at least the
editor role within your tenant or project; for example, the default editors group. You can check this in the Administration → IAM section of the web console.
How to create a job
To run a container image as a batch workload for training, fine-tuning or data processing, create a job:
-
In the sidebar, go to
Serverless AI → Jobs.
-
Click
Create job.
-
In Configuration, select Custom.
-
Configure Job settings:
-
In Image path, set the path to the container image.
-
If you use a private registry, in Private registry, select an existing registry or click
Add, and provide the details for your registry.
-
(Optional) In Entrypoint command, specify the entrypoint command for the container.
If you need to pass container arguments, specify them in this field as well.
-
(Optional) In Environment variables, specify environment variables in key-value pairs.
-
(Optional) In Secret environment variables, click
Create secret to store a sensitive value in SecretStash and inject it as an environment variable. In the window that opens, specify the Key (environment variable name) and Value (secret data), then click Create.
-
(Optional) In Job timeout in hours, specify the number of hours after which the job will be canceled if not completed.
-
(Optional) Configure the Computing resources section:
-
Select whether the container VM should have GPUs.
-
Specify the VM type: regular or preemptible.
VMs without GPUs only support the regular type.
-
Select the platform and preset.
-
If you selected the preemptible VM type, in Max price, select or create a pricing policy, or choose Follow spot price.
-
(Optional) Configure Storage settings:
-
(Optional) In the Files section, add one or more small configuration files to inject into the container at launch:
- Select Upload files to upload files from your machine, or Create text files to enter the file content and click Create file.
- In Mount path, specify the absolute path to the file in the container, for example,
/mnt/files/config.yaml.
Injected files are read-only and limited to 64 KiB each.
-
(Optional) In the Access section, configure User name and SSH key to connect to the running workload for debugging.
You can create new credentials or select existing ones. If you decide to use an existing credential, make sure that the SSH key is stored for the nebius username.
-
Configure the Network section:
- Select a subnet or create a new one.
- Select the IP address type: Public static IP or Private IP. If you want to connect to the resource from the internet, select Public static IP.
-
Click Create job.
With the command below, you specify all values directly in the command. This is useful for scripts, configuration-driven workflows, continuous integration (CI) and agents. Alternatively, you can run nebius ai create to use the interactive flow where the CLI prompts you for parameter values step by step.
Run the following command:
In the command, specify the following parameters:
-
Job settings:
-
--image: Container image reference in the registry/path:tag or registry/path@digest format. Use an image from a public registry or your authenticated private registry.
-
--registry-username, --registry-password (optional): Credentials to authenticate if you pull an image from a private registry. Alternatively, use --registry-secret for credentials stored in SecretStash.
--registry-username: Username.
--registry-password: Personal access token, password or an API key. Depends on where your registry is hosted. It can be Docker Hub, Microsoft Azure, GitHub, NVIDIA or a custom registry.
If you pull an image from a public registry or from Container Registry in the same project, you don’t need to specify credentials.
-
--registry-secret (optional): SecretStash secret selector with registry_username and registry_password payload keys. You can specify a secret name, secret ID, version ID or a combined secret/version selector such as mbsec-e00***@mbsecver-e00***.
-
--container-command (optional): Entrypoint command for the container.
-
--args (optional): Arguments for docker run to pass to the entrypoint command.
-
--env (optional): Environment variables for the container. Set them in the key=value format where key is the environment variable and the value is the value of this variable. If you need to set several variables, list the key=value pairs separated by commas.
-
--env-secret (optional): Environment variables loaded from a SecretStash secret. Set them in the key=secret_selector format, where key is the environment variable name and must match a payload key in the secret, and secret_selector is a secret name, secret ID, version ID or a combined secret/version selector such as mbsec-e00***@mbsecver-e00***. If you need to set several variables, list the pairs separated by commas. You cannot use the same key in both --env and --env-secret.
-
--working-dir (optional): Working directory (absolute path).
-
--timeout (optional): Job timeout (for example, 2h30m10s, 24h). Minimum: 1h, maximum: 168h. Default: 24h.
-
--restart-policy (optional): Whether to automatically restart the job when its container exits with an error (--restart-policy on-failure) or to only run it once and never restart it (--restart-policy never). For more details, see Automatic job restarts.
-
--inject-file (optional): Mount a local file into the job container at launch. Use the format <local_path>:<absolute_container_path>. To inject multiple files, repeat the parameter. The mounted file is read-only and limited to 64 KiB.
-
--volume (optional): Bucket or shared filesystem to mount to the job container and to store the job results and checkpoints. Volumes persist if the job is recreated after a maintenance event.
Specify the value in either format:
source:container_path[:mode] for mounting Nebius shared filesystems and existing bucket or volume resources by ID or name.
s3://bucket:/container_path[:mode[:profile]] for mounting an Object Storage bucket with AWS profile credentials or S3 credentials stored in SecretStash. The profile is the AWS credentials profile to use. If you manage your credentials with SecretStash, use profile@<secret_selector>, where <secret_selector> is a secret name, secret ID, version ID or a combined secret/version selector such as mbsec-e00***@mbsecver-e00***
The supported modes are ro, read only, and rw, read-write (default). Repeat for multiple volumes. For example:
-
Underlying container VM characteristics:
-
--subnet-id: Subnet ID. Required if the project has multiple subnets.
-
--platform: VM platform. See available platforms in Types of virtual machines and GPUs in Nebius AI Cloud.
-
--preset: Number of GPUs, vCPUs and RAM allocated to the container. The preset must match the selected platform. See available presets in Presets for GPU platforms.
-
--disk-size: Disk size of the container VM. Specify the value such as 100Gi, 500Gi or 1Ti. The default value is 250Gi.
See how disk performance depends on disk size.
-
--shm-size (optional): Shared memory size of /dev/shm. Specify the value such as 64Mi, 128Mi or 1Gi. The default value is 16Gi.
-
--ssh-key (optional): SSH key to access the container VM by SSH. When you add an SSH key, a public dynamic IP address is assigned. Before you add the key, check the quota on the number of public IP addresses in the web console.
-
--preemptible (optional): Use a preemptible VM. Preemptible VMs can help reduce computing costs for workloads that tolerate interruptions, but they can be stopped by Compute at any time. Charges depend on the current spot price. Only GPU platforms offer preemptible VMs. If you omit this parameter, the container runs on a regular VM.
-
--spot-pricing-policy-id (optional): Pricing policy ID that sets the maximum price you agree to pay for the preemptible VM. Requires --preemptible. Mutually exclusive with --follows-spot-price and --on-demand.
-
--follows-spot-price (optional): Accept the current spot price for the preemptible VM, with no maximum price. Requires --preemptible. Mutually exclusive with --spot-pricing-policy-id and --on-demand.
-
--on-demand (optional): Run on a regular, non-preemptible VM. This is the default when you omit --preemptible. Cannot be combined with --preemptible, --spot-pricing-policy-id or --follows-spot-price.
For the full list of parameters, see the nebius ai job create CLI reference.Send the following request:
In the Authorization header, replace <access_token> with the access token that you got in the prerequisites.The request includes the following parameters:
-
Job settings:
-
metadata.parentId: Project ID.
-
metadata.name: Job name.
-
spec.image: Container image reference in the registry/path:tag or registry/path@digest format. Use an image from a public registry or your authenticated private registry.
-
spec.registryCredentials (optional): Credentials to authenticate if you pull an image from a private registry.
spec.registryCredentials.username: Username.
spec.registryCredentials.password: Personal access token, password or an API key. Depends on where your registry is hosted. It can be Docker Hub, Microsoft Azure, GitHub, NVIDIA or a custom registry.
If you pull an image from a public registry or from Container Registry in the same project, you don’t need to specify credentials.
-
spec.containerCommand (optional): Entrypoint command for the container.
-
spec.args (optional): Arguments for docker run to pass to the entrypoint command.
-
spec.environmentVariables (optional): Environment variables for the container.
spec.environmentVariables[].name: Environment variable name.
spec.environmentVariables[].value: Environment variable value.
-
spec.workingDir (optional): Working directory (absolute path).
-
spec.timeout (optional): Job timeout in seconds with the s suffix (for example, 9000s). Minimum: 3600s (1 hour), maximum: 604800s (168 hours). Default: 86400s (24 hours).
-
spec.volumes (optional): Buckets or shared filesystems to mount to the job container and to store the job results and checkpoints. Volumes persist if the job is recreated after a maintenance event.
spec.volumes[].source: Volume ID or name.
spec.volumes[].containerPath: Absolute path where the volume is mounted in the container.
spec.volumes[].mode: Mount mode: READ_ONLY or READ_WRITE.
-
spec.injectedFiles (optional): Files to mount into the job container at launch. To inject multiple files, add multiple items. Each mounted file is read-only and limited to 64 KiB.
spec.injectedFiles[].containerPath: Absolute path where the file is mounted in the container.
spec.injectedFiles[].content: Base64-encoded file content.
-
Underlying container VM characteristics:
spec.subnetId: Subnet ID. Required if the project has multiple subnets.
spec.platform: VM platform. See available platforms in Types of virtual machines and GPUs in Nebius AI Cloud.
spec.preset: Number of GPUs, vCPUs and RAM allocated to the container. The preset must match the selected platform. See available presets in Presets for GPU platforms.
spec.disk.type: Disk type for the container VM.
spec.disk.sizeBytes: Disk size of the container VM in bytes. The default value is 250 GiB. See how disk performance depends on disk size.
spec.shmSizeBytes (optional): Shared memory size of /dev/shm in bytes. The default value is 16 GiB.
spec.sshAuthorizedKeys (optional): SSH public keys to access the container VM by SSH. When you add an SSH key, a public dynamic IP address is assigned. Before you add the key, check the quota on the number of public IP addresses in the web console.
spec.preemptible (optional): Whether to use a preemptible VM. Preemptible VMs can help reduce computing costs for workloads that tolerate interruptions, but they can be stopped by Compute at any time. Charges depend on the current spot price. Only GPU platforms offer preemptible VMs. If omitted or set to false, the container runs on a regular VM.
spec.pricingModel (optional): Pricing model for the VM. It must match spec.preemptible: use onDemand for a regular VM, or followsSpotPrice or spotPricingPolicy for a preemptible VM. Specify one of:
spec.pricingModel.onDemand: Empty object {} for a regular, non-preemptible VM.
spec.pricingModel.followsSpotPrice: Empty object {} to accept the current spot price for a preemptible VM, with no maximum price.
spec.pricingModel.spotPricingPolicy.id: ID of the pricing policy that sets the maximum price for a preemptible VM.
The job creation usually takes a few minutes. Jobs run until the workload finishes.
When the job completes successfully or fails, the container VM is deleted automatically. If you mounted volumes, they will remain, and you should delete them manually.
How to check job logs
For this operation, it’s enough to be in a group that has the viewer role within your tenant; for example, the default viewers group. You can check this in the Administration → IAM section of the web console.
Checking job logs is available only in the web console and CLI.
- In the sidebar, go to
Serverless AI → Jobs.
- Next to the job, click View logs. Alternatively, select the job that you want to view the logs for and switch to the Logs tab.
You can use the period or log level filters to filter the logs. You can also use the LogQL query language.Run the following command:You can add the following parameters to control the output:
--follow or -f: Stream logs in real time.
--since <value>: Show logs starting from the specified time. For example, 1h (from 1 hour ago), 30m (from 30 minutes ago) or 2024-01-01 (from that date).
--tail <value>: Number of recent lines to show in the output.
--timestamps: Include timestamps in the output.
--until <value>: Show logs up to the specified time. For example, 1h (up to 1 hour ago), 30m (up to 30 minutes ago) or 2024-01-01 (up to that date).
How to cancel a job
If you don’t need a job to continue running, you can cancel it. The jobs that finish with COMPLETED status are canceled automatically.
To cancel a job:
- In the sidebar, go to
Serverless AI → Jobs.
- Find the job and then click
→ Cancel.
- In the window that opens, confirm canceling the job.
-
List jobs:
In the output, copy the ID of the required job.
-
To cancel a job, run:
-
List jobs:
In the request, specify:
- Your access token in the
Authorization header.
- The project ID in the
parentId query parameter.
In the response, copy the metadata.id of the required job.
-
Cancel the job:
In the request, specify:
- Your access token in the
Authorization header.
<job_ID>: Job ID that you copied.
Canceling a job immediately stops the container VM and deletes the container disk. The job remains in the list of jobs. Mounted volumes are retained. You can remove the mounted volumes manually. See the guides on deleting a filesystem and deleting a bucket.
If you need to remove any record about the job from the job list, delete the job instead of canceling it.
How to delete a job
To delete a job:
- In the sidebar, go to
Serverless AI → Jobs.
- Locate the job and then click
→ Delete.
- In the window that opens, confirm deleting the job.
-
List jobs:
In the output, copy the ID of the required job.
-
Delete the job:
-
List jobs:
In the request, specify:
- Your access token in the
Authorization header.
- The project ID in the
parentId query parameter.
In the response, copy the metadata.id of the required job.
-
Delete the job:
In the request, specify:
- Your access token in the
Authorization header.
- The job ID that you copied in the request path.
When the job is deleted, it disappears from the list of jobs. If a job is running, deleting cancels the job first.
If the job uses additional volumes, they are not deleted with it. You can remove the mounted volumes manually. See the guides on deleting a filesystem and deleting a bucket.