Databricks as a Model Provider: The Safe Route into Enterprise Adoption of Open-Weight Models

There's a topic not getting enough attention: the role Databricks is taking on as an AI model provider. Not as a data platform that "also" serves models, but as an inference provider in its own right. And especially now, when everyone is talking about Chinese open-weight models.

Kimi K3 (Moonshot AI) and the GLM-5 family (Z.ai / Zhipu) represent today's state of the art open-weight models. They're reliable, fast, and cheap. Databricks has quietly become one of the best options for consuming them efficiently, quickly, and most importantly, securely.

What The Numbers Say

According to Artificial Analysis (the independent reference for inference provider benchmarks), Databricks is the second-fastest and second-lowest-latency provider for both Kimi K3 (max) and GLM-5.3 (max), among twelve and eight measured providers respectively, and in some cases it does so at the same price as the model maker.

ar chart comparing Kimi K3 max output speed by provider — Databricks ranks second at 117 tokens per second, behind Nebius at 129 t/s and ahead of Fireworks at 112 t/s, based on September 2026 data from artificialanalysis.ai

With 10,000 input tokens, Databricks delivers 117 tokens/second. Only Nebius is ahead (129 t/s), and it charges nearly twice as much.

Bar chart comparing Kimi K3 max time to first answer token by provider — Databricks ranks second fastest at 18.1 seconds, behind Nebius at 17.1s and ahead of Fireworks at 18.9s, based on September 2026 data from artificialanalysis.ai

When it comes to latency, the story is the same: 18.2 seconds to first answer token (including the model's reasoning time), versus 57 seconds on Moonshot's official API.

And on price, Databricks charges exactly what Moonshot charges: $2.31 per million blended tokens. No platform markup.

Bar chart comparing Kimi K3 max blended price per million tokens by provider — Databricks at $2.31 matches the model maker's price, sitting mid-range between Makora at $1.96 and Nebius at $4.20, based on September 2026 data from artificialanalysis.ai

With GLM-5.3 (max), the story repeats with even better numbers:

GLM-5.3 output speed and time to first token — Databricks ranks second across both metrics among top 5 providers, September 2026

Databricks serves GLM-5.3 (max) at 295.5 tokens/second with 7.6 seconds to first answer token; only Makora is ahead (340 t/s, 6.8 s). The third-place provider, Multiverse Computing, runs at less than half the speed (129 t/s). Z.ai's official API doesn't even make the top 5. The Databricks catalog also carries GLM-5.3 Flash (275 t/s) and GLM-5.2 (max) for cost-driven use cases.

One honest caveat: the Kimi K3 endpoint on Databricks exposes a 205k-token context window, versus the 1M advertised by other providers. For the vast majority of corporate processes, this is irrelevant. For extreme-context use cases, it's worth knowing.

But This Isn't About Being "The Best"

Rankings shift depending on how you measure and from one week to the next. What doesn't shift is that using Moonshot's or Zhipu's official APIs is not an option for a Western enterprise. And aggregators like OpenRouter, where you can pick from several providers, don't solve the problem either. Each provider would need to pass its own vendor approval process, and none offers the governance and sovereignty guarantees a security team demands.

The benefits of Databricks as a model provider include:

  • Self-hosted - Kimi K3 and the GLM models run on Databricks infrastructure; requests are not forwarded to the model maker's API. The weights are open. The model is Chinese, the inference is not.

  • Data residency - For workspaces in the EU Data Boundary, processing stays in European data centers. For US workspaces, processing is done in the US. Kimi K3 launched with US hosting and more locations are on the way.

  • Zero data retention - Databricks doesn't use prompts or responses to train models, and endpoints operate under ZDR.

  • Compliance - HIPAA across all regions along with PCI-DSS, FedRAMP, IRAP, CCCS, and UK Cyber Essentials Plus in supported regions. For banking and insurance, this isn't a nice-to-have but rather a requirement.

  • Unity AI Gateway - This includes guardrails, cost limits, rate limiting, payload logging, model fallback, and observability. The same controls for Kimi K3 as for Claude, GPT, or Gemini within the platform.

  • OpenAI-compatible API - Switching models means changing a string. That's what makes replacing a commercial model with an open one realistic without rewriting anything.

  • Availability - Databricks is an established company, with platform SLAs and provisioned throughput for production. For a process that needs to be available most of the time, that matters more than a 10 tokens/second difference.

These features are all in one place, with the same credentials, invoice, and governance catalog the data team already uses.

Diagram comparing Kimi K3 and GLM-5.3 deployment options — Databricks delivers enterprise-ready AI with zero retention, data residency, and compliance vs. unmanaged APIs and aggregators

Where This Is Going

I have no doubt enterprises will increasingly use open-weight models, in two layers:

  1. Local AI on employees' machines: models like Qwen 3.8 27B run remarkably well on any machine with 32 GB of unified memory or more.

  2. Provider-hosted models for productive work: process automation, document extraction, classification, and agents. This is where Databricks fits, on competitive and secure terms.

And a third, increasingly relevant reason: commercial frontier models charge separate credits, with quotas and plans that change every quarter. Having open alternatives of comparable quality inside the same platform means having negotiating leverage.

My Conclusion (After a Lot of Testing)

I've recently been running a multitude of tests and I'm convinced that if I had to automate a corporate process today that requires an LLM or SLM, I would seriously consider doing it with an open-weight model. They have comparable features, they're cheaper and more efficient, and everything points to that gap widening.

And within that option, I have no doubt that I'd use Databricks as the provider. It's top-tier in speed, latency, and availability (critical for keeping the process running); it's trustworthy and it guarantees governance and, above all, sovereignty over my data.

For any Western enterprise evaluating the replacement of a commercial model with an open-weight one for a specific process or use case, I can't think of a better option than starting with an endpoint on Databricks. Before considering buying hardware. Before approving yet another inference vendor. If the numbers later justify self-hosting, evaluate it then, with real consumption data in hand.

Taking all relevant factors into account, there is no better option. 

Sources: Artificial Analysis — Kimi K3 providers · Artificial Analysis — GLM-5.3 providers · Artificial Analysis — Databricks · Databricks Blog — Kimi K3 via Unity AI Gateway · Databricks Docs — Foundation Model APIs compliance · Databricks Designated Services

Next
Next

Genie ZeroOps: A Hands-On Preview