/
New feature Public Preview

AI Runtime (GPU serverless)

AI Runtime (GPU serverless) is a Databricks compute / bi capability, introduced June 2025.

GPU support for Databricks serverless compute - on-demand accelerators for deep-learning training and fine-tuning, single-node and multi-node, with no clusters to provision.

  • It puts on-demand H100s behind serverless compute, so a deep-learning job spins up GPUs without anyone sizing a cluster - single-node tasks are in Public Preview while the distributed multi-GPU training API is still catching up in Beta.
  • It first showed up in June 2025 under the plainer name 'Serverless GPU compute.'

Limitations: Only A10 and H100 accelerators; the maximum workload runtime is seven days; no FedRAMP High or DoD IL5 compliance. For scheduled jobs, Environments-panel dependencies and auto-recovery for incompatible packages aren't supported, and cross-region GPUs during peak demand can incur egress costs.

Open in REbricked →
Category
Compute / BI
Introduced
June 2025
Announced at
Debuted as Serverless GPU compute (Beta), Data + AI Summit June 2025
Also known as
AI Runtime, Serverless GPU compute
Verified
2026-07-23

Sources

Related in Compute / BI