---
title: "Latest reads"
description: "The engineering internet, summarised so you can actually read it."
---

# Latest reads
> The engineering internet, summarised so you can actually read it.

## Articles

### [MongoDB.local NYC 2025: Definire il database ideale per l'era dell'AI](https://yomu.fyi/post/mongodb-local-nyc-2025-definire-il-database-ideale-per-l-era-dell-ai.md)
- Company: [MongoDB](https://yomu.fyi/company/mongodb.md)
- Author: Dev Ittycheria, President and CEO, MongoDB
- Published: Sep 18, 2025

MongoDB announced several product releases and platform updates aimed at supporting artificial intelligence and agentic workflows at MongoDB.local NYC. The release of MongoDB 8.2 arrives alongside integrations with Voyage AI embedding and reranker models designed to enhance data retrieval precision. In addition, MongoDB launched Search and Vector Search in public preview for both Community Edition and Enterprise Server deployments, bringing vector capabilities to self-managed environments. To assist organizations transitioning away from rigid legacy database systems, MongoDB also introduced the Application Modernization Platform, which combines AI-driven tooling and specialized migration expertise. Early benchmarks from the modernization platform demonstrate legacy migrations running two to three times faster while accelerating code rewriting tasks by an order of magnitude.


### [Celebrating Excellence: MongoDB Global Partner Awards 2025](https://yomu.fyi/post/celebrating-excellence-mongodb-global-partner-awards-2025.md)
- Company: [MongoDB](https://yomu.fyi/company/mongodb.md)
- Author: Olivier Zieleniecki
- Published: Sep 18, 2025

MongoDB announced the recipients of its 2025 Global Partner Awards, recognizing cloud providers, systems integrators, and technology vendors for driving enterprise modernization and artificial intelligence adoption. Microsoft earned Global Cloud Partner of the Year for joint integrations linking MongoDB Atlas on Azure with native Microsoft services across healthcare, telecommunications, and financial services. Amazon Web Services received the Global AI Cloud Partner award, highlighted by a joint deployment with Novo Nordisk that reduced a key workflow from twelve weeks to ten minutes using Amazon Bedrock and Atlas. Confluent was named Global Tech Partner of the Year with over 550 joint customer deployments focused on event-driven streaming and multi-agent systems alongside LangChain. Additional honorees included Google Cloud, Accenture, BigID, Pureinsights, gravity9, IBM, and Alibaba Cloud for their contributions across public sector solutions, database-as-a-service offerings, and generative AI frameworks.


### [庆祝卓越：MongoDB 全球合作伙伴奖 2025](https://yomu.fyi/post/mongodb-2025.md)
- Company: [MongoDB](https://yomu.fyi/company/mongodb.md)
- Author: Olivier Zieleniecki
- Published: Sep 18, 2025

MongoDB announced its 2025 Global Partner Award recipients to recognize organizations driving AI adoption, legacy system modernization, and collaborative market expansion. Microsoft earned Global Cloud Partner for Azure integrations, while Amazon Web Services took Global AI Cloud Partner after assisting Novo Nordisk in cutting a key workflow from twelve weeks to ten minutes with Amazon Bedrock. Google Cloud received the Global Cloud GTM Partner award, and Confluent was named Global Technology Partner with over 550 joint deployments. LangChain and Pureinsights received honors for their integrations supporting retrieval-augmented generation and search solutions. Additional awards celebrated Accenture, BigID, gravity9, IBM, and Alibaba Cloud for their enterprise and public-sector impact.


### [우수성을 기념하기: 2025년 MongoDB 글로벌 파트너 어워드](https://yomu.fyi/post/2025-mongodb.md)
- Company: [MongoDB](https://yomu.fyi/company/mongodb.md)
- Author: Olivier Zieleniecki
- Published: Sep 18, 2025

MongoDB announced the recipients of its 2025 Global Partner Awards to recognize partner contributions across cloud modernization, artificial intelligence, and enterprise integration. Microsoft received the Global Cloud Partner award for Azure integrations, while Amazon Web Services secured the Global AI Cloud Partner award following joint generative AI implementations such as reducing workflow durations for Novo Nordisk. Google Cloud earned the Global Cloud GTM Partner distinction through shared sales development programs, and Accenture took Global SI Partner honors after establishing a dedicated engineering Center of Excellence. Confluent was named Global Tech Partner with more than 550 joint streaming deployments, and LangChain was recognized as Global AI Tech Partner for vector search and agentic frameworks. The awards demonstrate partner-led technical solutions spanning distributed databases, data streaming pipelines, and scalable enterprise cloud architectures.


### [Celebrando la Excelencia: Premios Globales de Emparejar de MongoDB 2025](https://yomu.fyi/post/celebrando-la-excelencia-premios-globales-de-emparejar-de-mongodb-2025.md)
- Company: [MongoDB](https://yomu.fyi/company/mongodb.md)
- Author: Olivier Zieleniecki
- Published: Sep 18, 2025

MongoDB presented its 2025 Global Partner Awards to honor partner organizations across cloud infrastructure, systems integration, generative AI, and enterprise data management. Microsoft received the Global Cloud Partner award for joint go-to-market solutions integrating MongoDB Atlas with native Azure services across healthcare, telecommunications, and financial services. Amazon Web Services earned the Global AI Cloud Partner award after delivering an Amazon Bedrock and MongoDB Atlas workflow implementation for Novo Nordisk that reduced processing time from twelve weeks to ten minutes. Confluent was recognized as Global Technology Partner for exceeding 550 joint deployments leveraging Apache Kafka for real-time data streaming and event-driven AI architectures. Other honored partners include Google Cloud for joint sales development programs, Accenture for software engineering centers of excellence, LangChain for retrieval-augmented generation tooling, and IBM for Watsonx.ai integrations.


### [Hommage à l’excellence : MongoDB Global Partner Awards 2025](https://yomu.fyi/post/hommage-a-l-excellence-mongodb-global-partner-awards-2025.md)
- Company: [MongoDB](https://yomu.fyi/company/mongodb.md)
- Author: Olivier Zieleniecki
- Published: Sep 18, 2025

MongoDB announced the recipients of its 2025 Global Partner Awards, recognizing enterprise collaboration across cloud infrastructure, artificial intelligence, and systems modernization. Microsoft received the Global Cloud Partner award for joint integrations with Azure, while Amazon Web Services earned the Global AI Cloud Partner award following deployments combining Amazon Bedrock and MongoDB Atlas. Google Cloud was recognized for joint go-to-market initiatives, and Accenture received honours for establishing a dedicated MongoDB Center of Excellence within its software engineering division. Additional recognitions included Confluent for event-driven streaming deployments exceeding 550 customer implementations, BigID for data governance, gravity9 for modernization services, and Alibaba Cloud for managed database services. IBM and Accenture Federal Services were also acknowledged for strategic enterprise impact across mainframe architectures and public sector programs.


### [Democratizing AI Safety with RiskRubric.ai](https://yomu.fyi/post/democratizing-ai-safety-with-riskrubric-ai.md)
- Company: [Hugging Face](https://yomu.fyi/company/hugging-face.md)
- Author: Gal Moyal
- Published: Sep 18, 2025

Cloud Security Alliance and Noma Security introduced RiskRubric.ai to provide standardized, transparent risk assessments across the open AI model ecosystem. The framework evaluates AI models across six pillars—transparency, reliability, security, privacy, safety, and reputation—using over 1,000 reliability tests, 200 adversarial security probes, automated code scanning, and harmful content evaluations. Each model receives 0–100 scores and A–F letter grades, supplemented by specific vulnerability findings and recommended mitigation strategies to assist deployment filtering. Initial benchmark results across models showed composite scores ranging from 47 to 94 with a median of 81, revealing polarized safety distributions and indicating that security hardening directly correlates with reduced safety risks.


### [Public AI on Hugging Face Inference Providers 🔥](https://yomu.fyi/post/public-ai-on-hugging-face-inference-providers.md)
- Company: [Hugging Face](https://yomu.fyi/company/hugging-face.md)
- Author: Joseph Low, Joshua Tan, Célina Hanouti, Julien Chaumond, Simon Brandeis, Lucain Pouget
- Published: Sep 17, 2025

Hugging Face has integrated Public AI as a supported Inference Provider on the Hugging Face Hub. Public AI operates as a nonprofit, open-source project providing access to sovereign and public models from institutions such as the Swiss AI Initiative and AI Singapore. Its distributed infrastructure combines a vLLM-powered backend serving OpenAI-compatible APIs across partner-donated clusters with a global load-balancing routing layer. Users can access these models through the Hugging Face web UI, Python client SDK, and JavaScript SDK using either direct provider API keys or routed Hugging Face tokens. At the time of announcement, inference through the Public AI provider is free of charge, supported by donated GPU time and advertising subsidies.


### [Taming the monorepo beast: Our journey to a leaner, faster GitLab repo](https://yomu.fyi/post/taming-the-monorepo-beast-our-journey-to-a-leaner-faster-gitlab-repo.md)
- Company: [Grab](https://yomu.fyi/company/grab.md)
- Author: Nagendra Gangwar
- Published: Sep 16, 2025

Grab's decade-old Go monorepo grew to 12.7 million commits and 250GB of Git data, causing Gitaly replication delays of up to four minutes that routed all read traffic exclusively to the primary node and slowed developer operations. After staging tests proved that shallow history reduced replication lag from hundreds of seconds to under three seconds, standard rewriting tools like git filter-repo and git rebase failed due to complex merge histories and repository scale. To overcome runner memory limits and lengthy git garbage collection cycles, the engineering team implemented a custom two-phase migration script. The script selectively migrated 2,000+ critical dependency tags and one month of recent history, flattening merge commits, embedding legacy hashes for traceability, and reducing total commit volume by 99.9%.


### [Visible Watermarking with Gradio](https://yomu.fyi/post/visible-watermarking-with-gradio.md)
- Company: [Hugging Face](https://yomu.fyi/company/hugging-face.md)
- Author: Margaret Mitchell
- Published: Sep 15, 2025

Generative artificial intelligence tools produce images, video, audio, and text that have become increasingly difficult to distinguish from real-world media captures. To improve synthetic media transparency, Hugging Face introduced visible watermarking capabilities directly into the Gradio web application framework. Developers can now overlay visual watermarks on image and video outputs by specifying a single watermark parameter using file paths, open image objects, or NumPy arrays. The implementation also supports QR code watermarks, which can match the visual style of generated media while conveying detailed background information. Furthermore, the Gradio Chatbot component accepts a text watermark parameter that automatically attaches attribution metadata whenever users copy generated responses to their clipboards.


### [Introducing the Palmyra-mini family: Powerful, lightweight, and ready to reason!](https://yomu.fyi/post/introducing-the-palmyra-mini-family-powerful-lightweight-and-ready-to.md)
- Company: [Hugging Face](https://yomu.fyi/company/hugging-face.md)
- Author: Rakshith, Tom Peres
- Published: Sep 11, 2025

WRITER released three open models in the Palmyra-mini family ranging between 1.5B and 1.7B parameters, designed for efficient inference and specialized reasoning tasks. Built on the Qwen architecture, the release includes the standard palmyra-mini base model alongside two Chain of Thought variants, palmyra-mini-thinking-a and palmyra-mini-thinking-b. In benchmark evaluations, palmyra-mini achieved 52.6% on Big Bench Hard, while palmyra-mini-thinking-a reached 82.87% on GSM8K and palmyra-mini-thinking-b reached 92.5% on AMC23. For palmyra-mini-thinking-b, applying reinforcement learning fine-tuning to an OpenReasoning-Nemotron-1.5B base improved single-shot pass@1 accuracy while reducing sampling diversity and majority@64 performance. The models are available in GGUF and MLX-BF16 quantizations and support inference engines including vLLM, SGLang, TRTLLM, and TGI.


### [Tricks from OpenAI gpt-oss YOU 🫵 can use with transformers](https://yomu.fyi/post/tricks-from-openai-gpt-oss-you-can-use-with-transformers.md)
- Company: [Hugging Face](https://yomu.fyi/company/hugging-face.md)
- Author: Aritra Roy Gosthipaty, Sergio Paniego, Vaibhav Srivastav, Pedro Cuenca, Arthur Zucker, Nathan Habib, Cyril Vallez
- Published: Sep 11, 2025

To support OpenAI's GPT-OSS series of models, the transformers library introduced several performance upgrades that apply across supported architectures. The release integrates zero-build custom kernels downloadable directly from the Hub, reducing external dependency bloat and compilation friction for operations like Liger RMSNorm and MegaBlocks MoE. Native support for MXFP4 quantization groups vector elements into 32-value blocks with shared scales, allowing GPT-OSS 20B to fit in roughly 16 GB of VRAM and GPT-OSS 120B in roughly 80 GB. In addition, transformers incorporates Flash Attention 3 with attention sinks, continuous batching via the generate\_batch API for experimentation, and automatic memory pre-allocation to speed up model loading on GPUs.


### [Fine-tune Any LLM from the Hugging Face Hub with Together AI](https://yomu.fyi/post/fine-tune-any-llm-from-the-hugging-face-hub-with-together-ai.md)
- Company: [Hugging Face](https://yomu.fyi/company/hugging-face.md)
- Author: Zain Hasan, Artem Chumachenko, Egor Timofeev, Max Ryabinin
- Published: Sep 10, 2025

Together AI and Hugging Face have introduced an integration enabling developers to fine-tune compatible Hugging Face Hub models directly on Together AI's managed infrastructure. To launch a fine-tuning job via the Python SDK, users provide a base model from Together's catalog as a configuration template alongside the target Hugging Face repository identifier. This base model template dictates GPU allocation, memory configuration, training pipelines, and inference setup for custom models with matching architectures and sizes. The workflow operates bidirectionally, pulling from public or token-authenticated private repositories and optionally pushing completed checkpoints back to the Hub upon completion. Geared toward CausalLM models under 100 billion parameters, the capability enables faster iteration cycles and domain adaptation without custom DevOps infrastructure.


### [Jupyter Agents: training LLMs to reason with notebooks](https://yomu.fyi/post/jupyter-agents-training-llms-to-reason-with-notebooks.md)
- Company: [Hugging Face](https://yomu.fyi/company/hugging-face.md)
- Author: Baptiste Colle, Hanna Yukhymenko, Leandro von Werra
- Published: Sep 10, 2025

Small language models often struggle to compete with large frontier models on complex, agentic data science tasks. To improve notebook-based reasoning, researchers simplified agent scaffolding down to roughly two hundred lines of code with dedicated execution and final answer tools, boosting baseline easy accuracy on the DABStep benchmark from 44.4 percent to 59.7 percent. They constructed a curated training dataset by deduplicating two terabytes of Kaggle notebooks, automatically fetching five terabytes of linked datasets, and scoring educational value and relevance using Qwen3-32B. Question-answer pairs grounded in verified execution traces were generated to fine-tune compact Qwen3-4B thinking and instruct models. The team released the trained models alongside the dataset and execution sandboxes, establishing a foundation for reinforcement learning and distillation on notebook workflows.


### [mmBERT: ModernBERT goes Multilingual](https://yomu.fyi/post/mmbert-modernbert-goes-multilingual.md)
- Company: [Hugging Face](https://yomu.fyi/company/hugging-face.md)
- Author: Marc Marone, Orion Weller, William Fleshman, Eugene Yang, Dawn Lawrie, Ben Van Durme
- Published: Sep 9, 2025

mmBERT is a massively multilingual encoder model trained on more than 3T tokens across over 1,800 languages to improve upon existing multilingual architectures like XLM-R. Built upon ModernBERT with a Gemma 2 tokenizer, mmBERT employs a three-phase training curriculum consisting of pre-training on 60 languages, mid-training on 110 languages, and a final decay phase covering 1,833 languages. The training pipeline integrates an inverse mask ratio schedule, dynamic language temperature annealing, and TIES merging across three decay variants. Benchmark evaluations demonstrate strong natural language understanding on English GLUE and multilingual XTREME, as well as competitive retrieval performance on MTEB v2 and CoIR. The release includes standard base and small models alongside open training data and checkpoints.


### [Welcome EmbeddingGemma, Google's new efficient embedding model](https://yomu.fyi/post/welcome-embeddinggemma-google-s-new-efficient-embedding-model.md)
- Company: [Hugging Face](https://yomu.fyi/company/hugging-face.md)
- Author: Tom Aarsen, Joshua, Alvaro Bartolome, Aritra Roy Gosthipaty, Pedro Cuenca, Sergio Paniego
- Published: Sep 4, 2025

Google DeepMind released EmbeddingGemma, a multilingual embedding model with 308 million parameters and a 2048-token context window designed for on-device applications. Based on the Gemma3 transformer backbone, the architecture replaces causal attention with bidirectional attention to function as an encoder, followed by mean pooling and two dense layers producing 768-dimensional vectors. The model incorporates Matryoshka Representation Learning, allowing outputs to be truncated down to 512, 256, or 128 dimensions for reduced memory and storage footprints. Trained on approximately 320 billion multilingual tokens across more than 100 languages, the quantized model operates under 200 MB of RAM. In domain-specific evaluations on the MIRIAD dataset, fine-tuning increased NDCG@10 from 0.8340 to 0.8862, outperforming larger baselines.


### [SAIR: Accelerating Pharma R&D with AI-Powered Structural Intelligence](https://yomu.fyi/post/sair-accelerating-pharma-r-d-with-ai-powered-structural-intelligence.md)
- Company: [Hugging Face](https://yomu.fyi/company/hugging-face.md)
- Author: Arman Zaribafiyan, Georgia Channing, Rudi Plesch, Zane Beckwith
- Published: Sep 2, 2025

SandboxAQ released the Structurally Augmented IC50 Repository (SAIR), an open-source dataset containing 5.24 million computationally co-folded 3D protein-ligand structures paired with empirical IC50 binding potency data. The repository addresses training data scarcity in structure-based drug discovery, where over 40 percent of the included target proteins lack experimental structures in the Protein Data Bank. To construct the dataset, engineers executed over 130,000 GPU hours of the Boltz1 co-folding model on 760 NVIDIA H100 processors hosted via NVIDIA DGX Cloud on Google Cloud Platform. Infrastructure optimizations maintained over 95 percent GPU compute utilization, compressing the generation timeline from three months to three weeks. Quality validation using PoseBusters confirmed that 97 percent of the predicted complexes met physical plausibility and chemical sanity standards.


### [Data mesh at Grab part I: Building trust through certification](https://yomu.fyi/post/data-mesh-at-grab-part-i-building-trust-through-certification.md)
- Company: [Grab](https://yomu.fyi/company/grab.md)
- Author: Chun Rong Phang
- Published: Aug 19, 2025

Rapid business growth across multiple verticals led Grab's centralized data engineering model to become an unscalable bottleneck, resulting in duplicate pipelines, ambiguous ownership, and broken downstream dependencies. To resolve these issues, the organization initiated a data mesh journey called Signals Marketplace that decentralizes data management and treats data as a product. A central data certification system establishes formal data contracts covering schemas, SLAs, freshness, and retention, while assigning clear Business Data Owners and Technical Data Owners. Breaches in contract guarantees automatically generate Data Production Incident tickets to enforce accountability and root-cause fixes. Consequently, 75% of internal queries now target certified assets, redundant tables saw a 400% year-over-year deprecation increase, and the total number of top-used datasets dropped by over 58%.


### [From Intern Project to Production: How I Shipped the Draw Tool for Canva's Present Mode](https://yomu.fyi/post/from-intern-project-to-production-how-i-shipped-the-draw-tool-for-canv.md)
- Company: [Canva](https://yomu.fyi/company/canva.md)
- Author: Edwina Adisusila
- Published: Aug 6, 2025

Canva engineers developed and shipped a real-time drawing tool for presentation mode after user feedback highlighted it as a highly requested feature. Integrating the existing editor-bound draw functionality into presentations required resolving tight package coupling, dual-window scaling mismatches, and performance regressions. To overcome architectural boundaries, core draw logic was extracted into a shared common package using interface abstractions. Coordinate normalization resolved positioning and scaling differences across presenter and audience views, while code-splitting deferred loading the core engine until activation. As a result, presentation load time regression dropped from 7% to 0.24%, and the feature reached over 470,000 monthly active users in production.


### [The evolution of Grab's machine learning feature store](https://yomu.fyi/post/the-evolution-of-grab-s-machine-learning-feature-store.md)
- Company: [Grab](https://yomu.fyi/company/grab.md)
- Author: Daniel Tai
- Published: Jul 24, 2025

Grab redesigned its initial machine learning feature store, Amphawa, to address high-dimensional data, complex entity retrieval, and versioning challenges during feature updates. The new architecture adopts a feature-table model where data scientists output Parquet datasets to Amazon S3 using Spark, which are then atomically ingested into Amazon Aurora PostgreSQL via a reverse ETL workflow. To prevent noisy-neighbor contention and optimize infrastructure costs, the platform utilizes Aurora's distributed storage to separate reads from writes. Grab pairs Aurora Serverless on writer nodes to scale up during daily batch ingestion with Provisioned instances on read replicas for steady serving traffic.


[Newer posts](https://yomu.fyi/page/29.md) · [Older posts](https://yomu.fyi/page/31.md)
