Loading…
Building for an Open Future - our new partnership with Google Cloud
Hugging FaceJeff Boudier, Simon Pagezy
Summary
Hugging Face and Google Cloud announced an expanded strategic partnership designed to streamline the deployment and management of open models on Google Cloud infrastructure. Prompted by a tenfold increase in Hugging Face usage on Google Cloud over three years, the collaboration introduces a joint CDN Gateway using Hugging Face Xet technology to cache models and datasets directly on Google Cloud. This caching mechanism aims to shorten download times, strengthen model supply chain resilience, and accelerate time-to-first-token across Vertex AI, Google Kubernetes Engine, Cloud Run, and Compute Engine virtual machines. Hugging Face plans to integrate native library support for Google Cloud TPUs, lower instance prices on Inference Endpoints, and enhance Hub security scanning via VirusTotal, Google Threat Intelligence, and Mandiant.
Context
Usage of Hugging Face by Google Cloud customers grew tenfold over three years, resulting in tens of petabytes of monthly downloads across billions of requests and driving demand for lower latency, robust model supply chains, and tighter infrastructure integration.
Approach / What changed
Hugging Face and Google Cloud are deploying a dedicated CDN Gateway utilizing Hugging Face Xet storage technology and Google Cloud infrastructure to cache repositories directly on GCP. The initiative also adds native library support for Google TPUs, improves integration across Vertex AI, GKE, and Cloud Run GPUs, and integrates VirusTotal, Google Threat Intelligence, and Mandiant to secure Hugging Face assets.
Takeaways
- A dedicated CDN Gateway built with Hugging Face Xet technology and Google Cloud networking caches repositories directly on GCP to reduce download times and improve supply chain robustness.
- Hugging Face libraries are adding native support for seventh-generation Google TPUs to streamline model acceleration alongside GPUs.
- Security across Hugging Face models, datasets, and Spaces is being reinforced through integrations with VirusTotal, Google Threat Intelligence, and Mandiant.
Related reading
huggingface.co ·
Graphcore and Hugging Face Launch New Lineup of IPU-Ready Transformers
Graphcore and Hugging Face expanded the range of machine learning modalities and tasks available in Hugging Face Optimum. Developers can now access ten transformer models optimized for Graphcore IPUs across natural language processing, speech, and computer vision. The available architectures include BERT, ViT, GPT-2, RoBERTa, DeBERTa, BART, LXMERT, T5, HuBERT, and Wav2Vec2, complete with IPU configuration files and ready-to-use pre-trained or fine-tuned weights. The integration supports the Bow IPU processor, which uses 3D Wafer-on-Wafer stacking to achieve up to 350 teraFLOPS of AI compute. Optimum also integrates with the Poplar SDK 2.5, enabling compatibility with frameworks such as PyTorch, TensorFlow, Docker, and Kubernetes.
Sally Dohertyhuggingface.co ·
Habana Labs and Hugging Face Partner to Accelerate Transformer Model Training
Training transformer models across computer vision, speech, and natural language processing tasks at scale often demands heavy compute resources, incurring significant time and financial expense. To address these bottlenecks, Habana Labs and Hugging Face partnered to integrate the SynapseAI software suite into the Hugging Face Optimum open-source library. This integration allows machine learning practitioners to accelerate transformer training workflows on Habana Gaudi processors using minimal code adjustments. Habana Gaudi hardware, featured in Amazon EC2 DL1 instances and Supermicro X12 servers, incorporates ten 100 Gigabit Ethernet ports per processor to scale from single units to thousands of chips. The combined hardware and software architecture supports TensorFlow and PyTorch while delivering price and performance metrics up to 40% lower than comparable training alternatives.