status.state field returned by the CLI and REST API.
Devlab statuses
Endpoint statuses
Job statuses
Typical transitions
When a Devlab or endpoint starts, it moves through the following statuses:PROVISIONING → STARTING → IMAGE_PULLING → RUNNING
When you stop a Devlab and restart it, or stop an endpoint and start it again:
RUNNING → STOPPING → STOPPED → STARTING → IMAGE_PULLING → RUNNING
When a job runs to completion:
PROVISIONING → STARTING → IMAGE_PULLING → RUNNING → COMPLETED
When you cancel a running job:
RUNNING → CANCELLING → CANCELLED
IMAGE_PULLING can be quick enough that the status never surfaces, and a Devlab can return to STARTING after it. Read these flows as the statuses a workload can pass through rather than a fixed order.
Provisioning and capacity
When you create or start a Devlab, endpoint or job, it first entersPROVISIONING, during which Serverless AI allocates the resources the workload needs, including the requested GPU platform and preset, and prepares it to run. If the requested capacity is available, PROVISIONING completes and the workload moves on to STARTING and RUNNING.
If the requested capacity is not immediately available, Serverless AI keeps trying for up to 30 minutes; this window doesn’t depend on whether the workload uses regular or preemptible compute. If capacity still is not available within that window, the workload moves on to ERROR with the NotEnoughResources code (see FAILED vs. ERROR) and a message such as:
Due to high demand, we can’t start this workload. Please try to create another one with a different platform or preset.The workload is not retried automatically. To run it, you can:
- Try a different platform or preset: capacity varies by GPU type and size.
- Try a different region, if your use case allows it.
- Retry a little later, as availability changes over time.
- Check your quotas: a capacity error can also mean you have reached a limit.
Devlab specifics
- A Devlab’s managed HTTPS URL is available only in the
RUNNINGstate. If the managed HTTPS endpoint is enabled, the Devlab stays inSTARTINGuntil the endpoint is ready, and only then becomesRUNNING. - Unlike jobs and endpoints, Devlabs are not preemptible and have no preempted state.
STOPPEDis a normal, resumable state, not a failure: you can restart the Devlab from it.- If the virtual machine behind a running Devlab is lost, the Devlab reports
STOPPINGand thenSTOPPEDwith no status details and preserves the workspace disk. You can restart it normally.
FAILED vs. ERROR
The FAILED and ERROR statuses indicate different kinds of problems:
FAILEDapplies to jobs and Devlabs. For a job, it means the job did not complete successfully, for example, because of an error in your code, an incorrect entrypoint, a missing dependency in the image, or an exceeded timeout. For a Devlab, it means the Devlab failed to start.ERRORmeans Serverless AI could not run the workload. For jobs, this indicates an internal Serverless AI error rather than a problem with your workload. Endpoints have noFAILEDstatus, so theirERRORcovers workload problems as well, such as a container that could not start. For Devlabs,ERRORcovers capacity and infrastructure problems.
status.state_details field in the CLI, before you retry. Its code and message name the cause. The message wording can change; the code is the stable identifier.
For jobs and endpoints, the code can be:
StartFailed: A container that could not startContainerFailed: A workload that exited with a non-zero codeTimeoutExceeded: A job that ran out of timeNotEnoughResources: No capacity for the requested platform and preset
code can be:
When the status details are empty, the cause is internal to Serverless AI. Retry your workload, and if the status persists, contact support. For deeper investigation, see Debugging failed Devlabs, jobs and endpoints.