Loading…
Scaleway on Hugging Face Inference Providers 🔥
Hugging FaceGuillaume Noale, Franck Pagny, Fred Bardolle, Guillaume Calmettes, Constance Morales, Célina Hanouti, Julien Chaumond, Simon Brandeis, Lucain Pouget
Summary
Scaleway is integrated as a supported serverless Inference Provider on the Hugging Face Hub, expanding model deployment options across Hub model pages and official client SDKs for JavaScript and Python. Operating out of European data centers located in Paris, France, Scaleway Generative APIs host open-weight models including gpt-oss, Qwen3, DeepSeek R1, and Gemma 3 with structured outputs, function calling, multimodal processing, and sub-200ms first-token response times. Developers can route inference requests using their Hugging Face tokens or supply direct Scaleway API keys. Billing for routed requests charges standard provider rates starting at €0.20 per million tokens without markups, while direct requests bill to Scaleway accounts. Hugging Face PRO subscribers receive two dollars of monthly inference credits applicable across providers.
Context
Hugging Face required expanded serverless inference options and European data sovereignty for running open-weight models directly from Hub model pages and client SDKs.
Approach / What changed
Scaleway Generative APIs integrated as an Inference Provider on Hugging Face Hub, allowing developers to execute inference via Python and JavaScript SDKs using either direct Scaleway API keys or Hugging Face token routing.
Takeaways
- Scaleway Generative APIs run in Paris, France data centers to provide European data sovereignty, multimodal support, and sub-200ms first-token response latencies.
- Users can access Scaleway models through Hugging Face JavaScript and Python SDKs using either a direct Scaleway API key or automated routing with a Hugging Face token.
- Hugging Face applies pass-through billing for routed requests at standard provider rates starting at €0.20 per million tokens, with PRO users receiving $2 in monthly inference credits.
Related reading
Public AI on Hugging Face Inference Providers 🔥
Hugging Face has integrated Public AI as a supported Inference Provider on the Hugging Face Hub. Public AI operates as a nonprofit, open-source project providing access to sovereign and public models from institutions such as the Swiss AI Initiative and AI Singapore. Its distributed infrastructure combines a vLLM-powered backend serving OpenAI-compatible APIs across partner-donated clusters with a global load-balancing routing layer. Users can access these models through the Hugging Face web UI, Python client SDK, and JavaScript SDK using either direct provider API keys or routed Hugging Face tokens. At the time of announcement, inference through the Public AI provider is free of charge, supported by donated GPU time and advertising subsidies.
Joseph Low, Joshua Tan, Célina Hanouti, Julien Chaumond, Simon Brandeis, Lucain PougetOVHcloud on Hugging Face Inference Providers 🔥