aiexpert
Home / News / Brief
Breaking · Aug 06, 2026, 03:05 PM · 3 sources

Baseten launches on Hugging Face Inference Providers; adds support for DeepSeek V4 Flash, GLM-5.2, Kimi K3

Baseten, an AI infrastructure platform for serverless inference and training, has launched as a supported Inference Provider on the Hugging Face Hub as of August 6, 2026. The integration enables developers to access and deploy Baseten-hosted models directly through Hugging Face's model pages, client SDKs (Python and JavaScript), and integrated Agent Harnesses (Pi, OpenCode, Hermes, OpenClaw). Initial support covers conversational and text-generation tasks with popular open-weight LLMs including DeepSeek V4 Flash, GLM-5.2, Kimi K3, and others, with support for additional tasks rolling out soon.

Baseten is available through two inference modes: custom API key (direct requests billed to the Baseten account) and routed by HF (requests routed through Hugging Face with charges applied to HF account, no provider token required). Users can set preferences in their HF account settings to order providers by choice and route requests through preferred infrastructure. Integration is live in the huggingface_hub Python SDK (>= 1.26.1) and @huggingface/inference JavaScript SDK, enabling seamless integration into existing Agent Harnesses without additional glue code. HF PRO users receive $2 monthly in Inference credits, usable across providers; free users get limited free quota with option to upgrade.

Baseten joins DeepInfra and Scaleway as supported Inference Providers on HF, expanding HF's roster of third-party inference options. This move mirrors the broader consolidation of AI infrastructure, where specialized inference platforms integrate upstream into model hubs and developer platforms to lower friction for deployment. HF continues to position itself as a central router and aggregator for model inference across multiple providers rather than a single monolithic inference service.

For architects: this is a practical consolidation point. Previously, developers choosing Baseten over HF's native Inference Endpoints required separate integrations and vendor-specific code. Now, model discovery, vendor preference, billing aggregation, and deployment all flow through a single UI/SDK interface. The routed-by-HF billing mode removes vendor lock-in friction. Expect other specialized inference providers (Fireworks, Together AI, Replicate) to follow. This establishes HF as the de facto inference routing layer for open models, reducing deployment friction and API lock-in for teams managing multi-provider stacks.

Sources

Everything this brief rests on
  1. 01 Primary source huggingface.co
  2. 02 Hugging Face Blog: Baseten on Inference Providers huggingface.co “Baseten is now a supported Inference Provider on the Hugging Face Hub”
  3. 03 Baseten org on Hugging Face Hub huggingface.co “Baseten catalog of models available on HF”