Devlabs, endpoints and jobs
Serverless AI supports the end-to-end AI/ML workflow through three deployment types. Choose the deployment type that best fits what you want to do:- Experiment and prototype: Write code, explore models and data in a browser-based IDE or notebook → Devlab.
- Run a task to completion: Train or fine-tune a model, preprocess data, run batch inference or evaluation → Job.
- Serve requests: Host a model or application at a URL for real-time inference → Endpoint.
Observability and debugging
Each Serverless AI deployment (a Devlab, an endpoint or a job) has a status that indicates the current stage in the lifecycle. You can view its status and logs on the resource page under Devlabs, Jobs or Endpoints in the web console. Deployments also report GPU and vCPU utilization metrics on the Metrics tab of this page. To investigate errors or unexpected outcomes, see Debugging failed deployments. For metric definitions and details, see Monitoring deployments.Pricing and quotas
Serverless AI follows Compute billing and quota rules. Billing depends on whether the deployment is running or stopped:- While a Devlab, job or endpoint is running, you are billed for its computing resources and storage.
- While a Devlab is stopped, you are billed for the entire container disk, but not for computing resources.
- While an endpoint is stopped, you are not billed for computing resources or storage.