# SAIR: Accelerating Pharma R&D with AI-Powered Structural Intelligence

[Hugging Face](https://yomu.fyi/company/hugging-face) · Arman Zaribafiyan, Georgia Channing, Rudi Plesch, Zane Beckwith · Sep 2, 2025

**Type:** Announcement

## Summary

SandboxAQ released the Structurally Augmented IC50 Repository (SAIR), an open-source dataset containing 5.24 million computationally co-folded 3D protein-ligand structures paired with empirical IC50 binding potency data. The repository addresses training data scarcity in structure-based drug discovery, where over 40 percent of the included target proteins lack experimental structures in the Protein Data Bank. To construct the dataset, engineers executed over 130,000 GPU hours of the Boltz1 co-folding model on 760 NVIDIA H100 processors hosted via NVIDIA DGX Cloud on Google Cloud Platform. Infrastructure optimizations maintained over 95 percent GPU compute utilization, compressing the generation timeline from three months to three weeks. Quality validation using PoseBusters confirmed that 97 percent of the predicted complexes met physical plausibility and chemical sanity standards.

## Context

AI-driven drug design is bottlenecked by the scarcity of structural training data linking 3D protein-ligand conformations to biological potency measurements such as IC50. Experimental methods like X-ray crystallography and cryo-EM are slow and expensive, leaving many disease-relevant targets in the dark proteome without structural representation in the Protein Data Bank.

## Approach / What changed

SandboxAQ created SAIR, an open-source dataset pairing 5.24 million computationally co-folded 3D protein-ligand structures across over 1 million pairs with curated IC50 labels from ChEMBL and BindingDB. The team utilized the Boltz1 model on a cluster of 760 NVIDIA H100 GPUs via NVIDIA DGX Cloud on Google Cloud Platform, optimizing workloads to exceed 95 percent GPU utilization and evaluating structure plausibility with PoseBusters.

## Takeaways

- SAIR contains 5.24 million 3D complexes spanning over 1 million unique protein-ligand pairs with experimental IC50 labels, with over 40 percent of proteins lacking prior structures in the Protein Data Bank.
- Workload and metric optimizations across 760 NVIDIA H100 GPUs on NVIDIA DGX Cloud achieved over 95 percent compute utilization, finishing 130,000 GPU hours of Boltz1 generation in three weeks.
- Quality checks with the PoseBusters benchmarking tool showed that 97 percent of SAIR's generated 3D structures passed chemical sanity and physical plausibility tests.

**Tags:** [Google Cloud](https://yomu.fyi/topic/gcp), [Machine Learning](https://yomu.fyi/topic/machine-learning), [Open Source](https://yomu.fyi/topic/open-source), [Python](https://yomu.fyi/topic/python)

[Read original post](https://huggingface.co/blog/SandboxAQ/sair-data-accelerating-drug-discovery-with-ai)
