# Latest reads
> The engineering internet, summarised so you can actually read it.

## Articles

### [Scaling out Distroless adoption With AI](https://yomu.fyi/post/scaling-out-distroless-adoption-with-ai.md)
- Company: [Grab](https://yomu.fyi/company/grab.md)
- Author: Jia Yee Chong
- Published: Jun 22, 2026

Grab is transitioning its microservices to Distroless base images to eliminate unnecessary binaries and reduce vulnerability risks, but the migration risks runtime failures from missing shared objects and system utilities. To safely validate container execution in continuous integration without staging dependencies, the team relied on medium tests that run containerized services alongside internal dependencies managed by Testcontainers. Because hundreds of services lacked this test harness, Grab implemented an agentic workflow using Claude Code and Model Context Protocol integrations to inspect repositories, generate test boilerplate, and resolve configuration errors. Once test baselines are established, an automated patch-test-compare pipeline updates Dockerfiles, constructs multi-stage builds for necessary dynamic libraries, and creates draft merge requests for human approval.


### [Palana (Part 2): Architecting isolation, identity, and auditability for AI agents](https://yomu.fyi/post/palana-part-2-architecting-isolation-identity-and-auditability-for-ai.md)
- Company: [Grab](https://yomu.fyi/company/grab.md)
- Author: Kevin Littlejohn
- Published: Jun 21, 2026

Grab's Palana platform provisions isolated, Kubernetes-native runtime environments for autonomous AI agents using dedicated per-agent namespaces and role-based access controls. The architecture separates network enforcement across layers, applying Layer 3 and Layer 4 containment with Cilium and NetworkPolicy alongside Layer 7 application filtering evaluated by Open Policy Agent. Agent interactions with large language models route through a LiteLLM proxy wrapper that retrieves credentials from HashiCorp Vault based on Kubernetes pod context rather than client headers. Secrets management is divided between directly readable agent paths and proxy-only placeholder paths that prevent raw tokens from residing in runtime filesystems. An automated reaper monitors multi-source activity signals to shut down idle compute resources while preserving persistent storage and configuration state.


### [The Data Canary: How Netflix Validates Catalog Metadata](https://yomu.fyi/post/the-data-canary-how-netflix-validates-catalog-metadata.md)
- Company: [Netflix](https://yomu.fyi/company/netflix.md)
- Author: Netflix Technology Blog
- Published: Jun 19, 2026

A manual mitigation action during an incident corrupted a data feed for a subset of titles, causing playback issues and catalog service failures that existing code canary systems failed to catch. To protect streaming reliability, Netflix built an automated data canary system that validates transformed catalog metadata prior to publication. The architecture utilizes a dedicated orchestrator alongside permanent baseline and canary service clusters to coordinate validation using real production traffic. By leveraging custom chaos experiment thresholds, sticky session affinity, and Starts Per Second playback metrics, the system detects regressions in under ten minutes and blocks publication automatically. Controlled failure injection experiments routing approximately 0.2% of global traffic confirmed that issues could be identified in 2.5 to 4 minutes.


### [Dispatches from O'Reilly: From capabilities to responsibilities](https://yomu.fyi/post/dispatches-from-o-reilly-from-capabilities-to-responsibilities.md)
- Company: [Stack Overflow](https://yomu.fyi/company/stack-overflow.md)
- Author: Artur Huk
- Published: Jun 19, 2026

High-stakes AI agents capable of mutating external state often face governance failures when relying on system prompts or manual Human-in-the-Loop approval queues that quickly degrade into alert fatigue. The Responsibility-Oriented Agent architecture addresses this operational bottleneck by shifting system design from open-ended capability framing to deterministic, contract-enforced responsibilities. Under this model, underlying orchestration frameworks like LangChain operate in User Space with their side-effecting tools removed, isolating the agent to epistemic reasoning. The agent expresses its intended action exclusively by emitting a structured policy proposal to a privileged Kernel Space runtime. The runtime deterministically evaluates the proposal against versioned YAML contracts registered in an agent registry, ensuring that only genuine policy exceptions are escalated to human supervisors.


### [Palana (Part 1): Why Grab built a secure platform for autonomous AI Agents](https://yomu.fyi/post/palana-part-1-why-grab-built-a-secure-platform-for-autonomous-ai-agent.md)
- Company: [Grab](https://yomu.fyi/company/grab.md)
- Author: Kevin Littlejohn
- Published: Jun 19, 2026

Autonomous AI agents introduce significant operational and security risks when granted network access, persistent state, and credentials. To address these concerns without impeding developer productivity, Grab created Palana, an in-house Kubernetes-native execution substrate. The platform isolates each agent workload within its own namespace, pairing it with dedicated storage, network policies, and role-based access control. Network egress is funneled through an Envoy and Open Policy Agent proxy layer that audits requests and injects credentials from HashiCorp Vault using placeholder tokens, keeping raw secrets outside the agent runtime. This design allows Grab to securely host hundreds of long-running workflows, remote coding environments, and automation bots.


### [The new bottleneck](https://yomu.fyi/post/the-new-bottleneck.md)
- Company: [Stack Overflow](https://yomu.fyi/company/stack-overflow.md)
- Author: Eira May
- Published: Jun 18, 2026

AI coding tools have significantly lowered the cost of generating software, yet many engineering organizations fail to realize overall delivery speed improvements. Applying the Theory of Constraints reveals that eliminating code production as a bottleneck shifts inventory directly into surrounding legacy processes that remain unadjusted. New friction points consistently emerge across underspecified requirements, prolonged design handoff gates, senior engineer review capacity, and external sign-offs from legal or security. Organizations can address these blockers by interrogating legacy agile ceremonies checkpoint by checkpoint to determine if their original constraints still exist. Practical remedies include adopting real-time co-development between product and engineering, treating initial designs as fluid starting points, and restructuring processes around running rapid, high-volume experiments.


### [AI agents are a confused deputy with the keys to your kingdom](https://yomu.fyi/post/ai-agents-are-a-confused-deputy-with-the-keys-to-your-kingdom.md)
- Company: [Stack Overflow](https://yomu.fyi/company/stack-overflow.md)
- Author: Fabio Salvadori
- Published: Jun 17, 2026

Attackers recently compromised over twenty thousand Instagram accounts by manipulating Meta's AI support assistant to rebind recovery email addresses without verifying account ownership. This incident illustrates the classic confused deputy security problem, where a privileged process is persuaded by an unprivileged user to perform unauthorized operations. Because large language model interfaces operate purely on natural language and cannot distinguish instructions from data, the model itself cannot serve as an authorization boundary. Securing AI agents requires verifying caller identity through external policy checks against authenticated sessions rather than relying on chat context or prompt engineering. Teams must enforce least privilege with short-lived scoped credentials, place irreversible actions behind hard policy gates or human approvals, and maintain audit trails of agent actions.


### [Building AI Agents for AR Glasses and XR Devices with NVIDIA XR AI](https://yomu.fyi/post/building-ai-agents-for-ar-glasses-and-xr-devices-with-nvidia-xr-ai.md)
- Company: [NVIDIA Developer Blog](https://yomu.fyi/company/nvidia-developer-blog.md)
- Author: Greg Barbone
- Published: Jun 16, 2026

Building AI experiences for augmented reality and wearable devices requires bridging hardware with live media streams, multimodal models, enterprise tools, and runtime infrastructure. NVIDIA XR AI provides an open-source, modular framework connecting extended reality headsets and smart glasses to GPU-accelerated services across cloud, edge, and workstations. In this architecture, camera frames and microphone audio ingest into an XR Media Hub that routes data while keeping raw video pixels in shared memory to minimize overhead. The ecosystem uses NVIDIA Cosmos models for vision-language grounding, NVIDIA Nemotron models for reasoning and tool invocation, and the Model Context Protocol for enterprise integrations. Optional agent orchestration via NVIDIA NeMo Agent Toolkit and spatial streaming through NVIDIA CloudXR support complex workflows across healthcare and manufacturing.


### [Build Your Own Transaction Foundation Model for Financial Intelligence](https://yomu.fyi/post/build-your-own-transaction-foundation-model-for-financial-intelligence.md)
- Company: [NVIDIA Developer Blog](https://yomu.fyi/company/nvidia-developer-blog.md)
- Author: Benjamin Wu
- Published: Jun 16, 2026

Production financial intelligence systems frequently rely on hand-engineered tabular features and rule sets that are brittle, expensive to maintain, and blind to historical sequential patterns. NVIDIA demonstrates an accelerated reference pipeline that replaces standard BPE tokenization with a GPU-based domain tokenizer, converting raw transactions into semantic tokens with an 8,192-token context window. Using the NeMo AutoModel library, a compact 29-million parameter decoder-only transformer is pretrained from scratch on unlabeled transaction sequences using causal language modeling. Learned sequence embeddings are extracted, compressed using PCA, and concatenated with raw tabular features to train a downstream GPU-accelerated XGBoost fraud detection model. On the IBM TabFormer benchmark dataset, this combined approach achieves a 41.76% lift in Average Precision over the baseline model relying solely on raw tabular features.


### [NVIDIA Blackwell Tops MLPerf Training 6.0 with Industry-Leading Scale and Performance](https://yomu.fyi/post/nvidia-blackwell-tops-mlperf-training-6-0-with-industry-leading-scale.md)
- Company: [NVIDIA Developer Blog](https://yomu.fyi/company/nvidia-developer-blog.md)
- Author: Farshad Ghodsian
- Published: Jun 16, 2026

NVIDIA submitted results across all MLPerf Training v6.0 benchmarks, setting performance records on workloads including the DeepSeek-V3 and GPT-OSS-20B Mixture of Experts models. To scale training up to 8,192 Blackwell GPUs, NVIDIA combined hardware cluster designs with scale-out networking via Spectrum-X Ethernet and Quantum InfiniBand. The engineering team eliminated CPU-GPU synchronization bottlenecks in token-dropless MoEs by implementing full-iteration CUDA graphs with synchronization-free operators and paged stashing. Further software optimizations included CuTe DSL kernel fusions, an MXFP8 attention block in the Transformer Engine, and transitioning router elementwise math to FP32. Across these benchmarks, the GB300 NVL72 platform achieved the fastest time to train at scale and the highest normalized per-accelerator performance.


### [Build On-Device AI Companions with the NVIDIA ACE Game Agent SDK and Unreal Engine 5 Plugins](https://yomu.fyi/post/build-on-device-ai-companions-with-the-nvidia-ace-game-agent-sdk-and-u.md)
- Company: [NVIDIA Developer Blog](https://yomu.fyi/company/nvidia-developer-blog.md)
- Author: Phillip Singh
- Published: Jun 16, 2026

Game developers integrating on-device artificial intelligence companions face challenges such as conversational latency, state synchronization, and loop prevention. NVIDIA announced the open-source ACE Game Agent SDK alongside a suite of Unreal Engine 5 plugins to streamline native, hardware-accelerated character pipelines. The SDK offers Agent, Chat, and RAG APIs, enabling characters to execute multi-step tool-assisted reasoning and query game databases. Unreal Engine 5 plugins introduce local runtime models including nemo-conformer-ctc-120m for speech recognition, Qwen 3.5 4B for language generation, and Chatterbox Turbo 350M for speech synthesis. Additional tooling expands to motion generation via Animotive Kimodo and rendering updates with the DLSS 4.5 Unreal Engine plugin.


### [How to Optimize Transformer-Based Models for Low-Precision Training](https://yomu.fyi/post/how-to-optimize-transformer-based-models-for-low-precision-training.md)
- Company: [NVIDIA Developer Blog](https://yomu.fyi/company/nvidia-developer-blog.md)
- Author: Jonathan Mitchell
- Published: Jun 16, 2026

Accelerating transformer training requires understanding how low-precision formats such as FP8 and NVFP4 impact specific general matrix multiplication (GEMM) workloads. Transformer configurations do not explicitly reveal these shapes, making microbenchmarks necessary before committing to full training runs. NVIDIA Transformer Engine enables quantization and kernel dispatch across precisions, which can be evaluated in realistic autocast mode or kernel-only prequantized mode. Profiling the ESM2-15B model on NVIDIA B300 GPUs demonstrated that NVFP4 achieved a 1.79x blended forward propagation speedup over MXFP8 and up to 4.01x over BF16 in prequantized execution. Although large GEMM dimensions successfully overcome quantization overheads, dynamic scaling, Hadamard transforms, and kernel selection asymmetries moderate real-world gains.


### [How Dropbox uses MCP and Dash to close the design-to-code security gap](https://yomu.fyi/post/how-dropbox-uses-mcp-and-dash-to-close-the-design-to-code-security-gap.md)
- Company: [Dropbox](https://yomu.fyi/company/dropbox.md)
- Author: Ilya Yakovlev,Andrew Cheung,Binoy Dash,Simran Jumani,Dmitriy Meyerzon,Mark Breitenbach,Ishan Mishra
- Published: Jun 12, 2026

Dropbox developed a system using the Model Context Protocol (MCP) and Dash's semantic search to bridge the gap between security threat models and code implementation. By retrieving original security documents during pull requests, an LLM agent automatically evaluates whether the proposed code adheres to previously agreed-upon security requirements. This approach surfaces design regressions and missing controls that traditional static analysis tools miss.


### [Better, faster, less wrong: Enhancing issue grouping](https://yomu.fyi/post/better-faster-less-wrong-enhancing-issue-grouping.md)
- Company: [Sentry](https://yomu.fyi/company/sentry.md)
- Author: Kush Dubey, Yuval Mandelboum
- Published: Jun 12, 2026

Sentry upgraded its AI-driven issue grouping system to better prevent duplicate issues without merging distinct application errors. The team trained lightonai/modernbert-embed-large using Matryoshka Representation Learning on hundreds of thousands of stacktrace pairs labeled by Claude Sonnet 4.5. To optimize the high-throughput ingestion pipeline, embeddings were truncated from 768 to 64 dimensions, combined with bfloat16 precision, PyTorch SDPA, and CUDA graph compilation. A phased live rollout used threshold-gated index backfilling and automatic fallback to prevent issue spikes during migration. In production, the v2 model cuts the overgrouping rate from 8% to 4%, increases prevented duplicate issues to 70%, and delivers 6x faster inference.


### [Production-Ready Agents Need A Production-Ready Data Platform](https://yomu.fyi/post/production-ready-agents-need-a-production-ready-data-platform.md)
- Company: [MongoDB](https://yomu.fyi/company/mongodb.md)
- Author: Pablo Stern
- Published: Jun 11, 2026

AI development teams face constant shifts in model providers and agent frameworks, demanding data platforms that provide scalable, real-time context management. Agentic workloads require blending unstructured enterprise data, short-term session state, and persistent long-term memory. MongoDB addresses these requirements through its native JSON document model and integrated retrieval capabilities, combining full-text search, vector search, and hybrid search directly over operational data. Customer implementations such as DevRev, ElevenLabs, and Adobe rely on Atlas to achieve sub-100 millisecond hybrid retrieval and handle billions of requests. Additionally, MongoDB is collaborating with LangChain and ecosystem partners to establish open reference architectures and shared interfaces for portable agent memory across frameworks.


### [Agentic Testing: Where Agents Fit in the E2E Testing Stack](https://yomu.fyi/post/agentic-testing-where-agents-fit-in-the-e2e-testing-stack.md)
- Company: [Slack](https://yomu.fyi/company/slack.md)
- Author: Sergii Gorbachov
- Published: Jun 11, 2026

Traditional end-to-end tests validate rigid user journeys, whereas agentic tests verify whether broad goals can be achieved by adapting actions dynamically. To evaluate agentic testing tradeoffs, researchers executed over 200 runs across Playwright Model Context Protocol (MCP), Playwright CLI, and agent-generated Playwright tests using Claude models. Playwright MCP demonstrated high reliability with failure rates of 0% on simple thread replies and approximately 12% on complex search discovery flows. Playwright CLI and generated code struggled more on complex workflows, exhibiting failure rates of approximately 20% and 48% respectively. Although generated tests were faster with average runtimes of roughly three minutes, agentic testing provides a distinct exploratory layer atop deterministic CI test suites.


### [Catch visual regressions with Snapshots, now in beta](https://yomu.fyi/post/catch-visual-regressions-with-snapshots-now-in-beta.md)
- Company: [Sentry](https://yomu.fyi/company/sentry.md)
- Author: Max Topolsky
- Published: Jun 11, 2026

Sentry has released the beta version of Sentry Snapshots to catch visual regressions on pull requests across any platform with a frontend. The tool captures screenshots of an application across various viewports, themes, and languages, compares them against a baseline, and blocks pull requests when diffs are detected. Teams can attach custom context metadata, including test file paths, to help developers and AI agents locate test sources and generate additional test cases. By leveraging the sentry-mcp integration, REST API, and sentry-cli, workflows can retrieve base images, diff masks, and run local diff verification before pushing commits. Originally created at Emerge Tools for mobile applications, this iteration expands automated visual testing capabilities to broader frontend codebases.


### [Metric Semantic Layer: How Lyft Governs and Scales Key Data Definitions](https://yomu.fyi/post/metric-semantic-layer-how-lyft-governs-and-scales-key-data-definitions.md)
- Company: [Lyft](https://yomu.fyi/company/lyft.md)
- Author: Iraklikhorguani
- Published: Jun 10, 2026

As Lyft scaled, different teams developed conflicting definitions for key business metrics due to the lack of centralized version control and shared standards. To resolve this, Lyft implemented an internal Metric Semantic Layer as a Python package that stores authoritative metric definitions in YAML files with Jinja SQL templates. The system restricts onboarding to Golden Metrics used across multiple applications, requiring team-based approval from both Business and Operational Owners for any definition changes. Standardized definitions are exposed through Python APIs, integrated into the Amundsen data catalog and self-service user interfaces, and surfaced to AI tools via a Model Context Protocol.


### [Graviton5&#8217;s improved design increases speed and energy efficiency &#8212; beyond Moore&#8217;s law](https://yomu.fyi/post/graviton5-8217-s-improved-design-increases-speed-and-energy-efficiency.md)
- Company: [Amazon](https://yomu.fyi/company/amazon.md)
- Author: Ali Saidi
- Published: Jun 10, 2026

Amazon announced the general availability of M9g and M9gd EC2 instances powered by the new Graviton5 processor. Built on a three-nanometer process, Graviton5 features 192 Neoverse V3 cores, DDR5-8800 memory support, PCIe gen6 interconnects, and 192 megabytes of level-three cache. The chip transitions from seven discrete dies to four unified chiplets connected at 420 gigabytes per second, eliminating dedicated I/O and memory controller dies. These architecture updates deliver up to 25% higher computational performance over Graviton4, with up to 35% faster performance for web applications and machine learning inference. The new instances also introduce the Nitro Isolation Engine, which uses formal verification to enforce hardware and software isolation between virtual machines.


### [From Chaos to Clarity: How We Built a Unified, Self-Routing Support Ops Ticketing System at Lyft](https://yomu.fyi/post/from-chaos-to-clarity-how-we-built-a-unified-self-routing-support-ops.md)
- Company: [Lyft](https://yomu.fyi/company/lyft.md)
- Author: Atulgupta
- Published: Jun 9, 2026

Lyft Urban Solutions' Support Ops team faced fragmented ticketing across four Jira Help Centers with duplicate intake forms, lack of routing logic, and no dashboards. Over five years, the team redesigned the intake architecture by replacing redundant forms with dynamic Jira Proforma forms featuring conditional branching. Automated Jira rules were introduced to assign tickets, apply a taxonomy of over 30 labels, and trigger PagerDuty alerts for P0 incidents. To unify operations across disparate backend projects, the system automatically clones and links tickets from a single portal to downstream project boards. Jira Structures and a Jira-to-Mode ETL pipeline provided cross-project rollups and business analytics before an upcoming Jira Cloud migration required dashboard rebuilds.


[Newer posts](https://yomu.fyi/page/12.md) · [Older posts](https://yomu.fyi/page/14.md)
