Loading…
Introducing AI spend controls with Unity AI Gateway
Kevin Stumpf
- Source
- Databricks
- Published
- Added to Yomu
Summary
Unity AI Gateway now offers AI Spend Controls, extending existing cost visibility with proactive budget alerts across models and workloads. The feature supports budgets at user, use-case, workspace, and account levels, plus shared and per-user thresholds that can trigger email alerts or enforce hard caps by stopping requests after a limit is exceeded. Configuration starts in account settings under Usage and Budgets, where administrators select Unity AI Gateway, optionally scope workspaces and resource tags, and define monthly limits and recipients. Budget status and trends are available in the Cost section, while customizable Cost Analytics dashboards use Unity Catalog system tables to attribute DBU and model-provider costs by identity, workspace, endpoint, tags, model, provider, and request tags. The release positions Databricks budgets, Unity AI Gateway, and Unity Catalog as a combined governance layer for controlling AI access, usage, and spend.
Context
AI adoption spans dozens of teams, hundreds of users, and thousands of agents, while workloads can incur unpredictable costs through retries, accidental experiments, or excessive token usage. The release addresses the need for uniform spend controls across AI workloads and organizational levels.
Approach / What changed
Administrators configure Unity AI Gateway budgets with shared or per-user monthly thresholds, workspace and resource-tag filters, alert recipients, and optional hard caps. They can monitor budget trends in the Cost section and analyze detailed usage through customizable dashboards backed by Unity Catalog system tables.
Takeaways
- Budgets can be configured per user, use case, workspace, or account, allowing controls to match different organizational responsibilities.
- Hard spend caps automatically stop further requests after a budget is exceeded, until the limit is raised or the next billing period begins.
- Cost tracking includes DBU costs, provisioned throughput, uptime, pay-per-token usage, and external provider token costs, with attribution across identities, models, providers, workspaces, and tags.