Loading…
SAIR: Accelerating Pharma R&D with AI-Powered Structural Intelligence
Hugging FaceArman Zaribafiyan, Georgia Channing, Rudi Plesch, Zane Beckwith
Summary
SandboxAQ released the Structurally Augmented IC50 Repository (SAIR), an open-source dataset containing 5.24 million computationally co-folded 3D protein-ligand structures paired with empirical IC50 binding potency data. The repository addresses training data scarcity in structure-based drug discovery, where over 40 percent of the included target proteins lack experimental structures in the Protein Data Bank. To construct the dataset, engineers executed over 130,000 GPU hours of the Boltz1 co-folding model on 760 NVIDIA H100 processors hosted via NVIDIA DGX Cloud on Google Cloud Platform. Infrastructure optimizations maintained over 95 percent GPU compute utilization, compressing the generation timeline from three months to three weeks. Quality validation using PoseBusters confirmed that 97 percent of the predicted complexes met physical plausibility and chemical sanity standards.
Context
AI-driven drug design is bottlenecked by the scarcity of structural training data linking 3D protein-ligand conformations to biological potency measurements such as IC50. Experimental methods like X-ray crystallography and cryo-EM are slow and expensive, leaving many disease-relevant targets in the dark proteome without structural representation in the Protein Data Bank.
Approach / What changed
SandboxAQ created SAIR, an open-source dataset pairing 5.24 million computationally co-folded 3D protein-ligand structures across over 1 million pairs with curated IC50 labels from ChEMBL and BindingDB. The team utilized the Boltz1 model on a cluster of 760 NVIDIA H100 GPUs via NVIDIA DGX Cloud on Google Cloud Platform, optimizing workloads to exceed 95 percent GPU utilization and evaluating structure plausibility with PoseBusters.
Takeaways
- SAIR contains 5.24 million 3D complexes spanning over 1 million unique protein-ligand pairs with experimental IC50 labels, with over 40 percent of proteins lacking prior structures in the Protein Data Bank.
- Workload and metric optimizations across 760 NVIDIA H100 GPUs on NVIDIA DGX Cloud achieved over 95 percent compute utilization, finishing 130,000 GPU hours of Boltz1 generation in three weeks.
- Quality checks with the PoseBusters benchmarking tool showed that 97 percent of SAIR's generated 3D structures passed chemical sanity and physical plausibility tests.
Related reading
AI for Food Allergies
Global food allergies affect an estimated 220 million individuals, yet computational progress in biomedical discovery remains constrained by fragmented and inaccessible scientific data. The AI for Food Allergies initiative addresses this challenge by establishing an open research community and releasing the curated Awesome Food Allergy Datasets collection across multiple biological layers. Computational pipelines leverage deep learning models, such as AllergenAI and NetAllergen-1.0, which incorporate sequence motifs and computationally predicted MHC class II presentation propensities to evaluate allergenicity. Additionally, molecular property prediction benchmarks like QM9 provide high-accuracy quantum-mechanical properties for approximately 134,000 molecules, supporting generative and virtual screening workflows targeting IgE–FcεRI binding. These combined efforts systematically structure molecular, clinical, and chemogenomic resources to accelerate allergy diagnostics and therapeutic protein engineering.
Ludovico Comito, Antonis Vozikis, Vaibhav Pandey, Kisejjere Rashid