Skip to main content
Serverless AI is a Nebius AI Cloud service for running containerized AI workloads as interactive endpoints or non-interactive jobs. By deploying your workloads in Serverless AI, you can focus on them without worrying about the infrastructure: the service handles resource provisioning and lifecycle, and usage-based, per-second billing. The service is available in all Nebius AI Cloud regions except for the private eu-west2 region.

About Serverless AI

Read about how Serverless AI works and how to choose between endpoints and jobs

Getting started with jobs

Create your first job that runs nvidia-smi and prints information about GPUs in use

Getting started with endpoints

Launch a simple endpoint and send authenticated requests to it

Deploying an LLM

Deploy a large language model on an endpoint and chat with it

Fine-tuning an LLM

Run a job that fine-tunes a large language model by using Axolotl

Monitoring

Track resource utilization to schedule quota increases and to quickly identify anomalies