Loading…
Building a soccer coaching app on Databricks
Samwel Emmanuel, Sheridan Harris, Andrew Helmreich, Kush Patel, Nick Ragonese
- Source
- Databricks
- Published
- Added to Yomu
Summary
Coach’s Corner, also called La Pizarra, turns high-frequency soccer tracking data into a bench-side application for replay, tactical analysis, scouting, standings, and agent-generated dossiers. Built as a Databricks App, it ingests NDJSON feeds at 25 frames per second through Auto Loader and Spark Declarative Pipelines, enforcing 46 data-quality expectations across bronze, silver, and gold layers. Liquid clustering supports 1–3-second DBSQL queries, while Lakebase synchronizes gold data to Postgres for millisecond replay reads and separates sequential playback from exploratory analytics. The scouting layer grounds Genie, Vector Search, a Unity Catalog-registered xG model, and an Agent Bricks supervisor in governed data, with Claude calls routed through the Unity AI Gateway, MLflow tracing, and a deterministic fallback. Together, these components are presented as a way to deliver traceable insights within seconds without forcing coaches to interpret raw tables or analysts to relay every result.
Context
Soccer matches generate extremely large, high-frequency tracking datasets, including 51 million rows for one tournament, but coaches need usable insights within seconds rather than through batch workflows, offline dashboards, or analyst-mediated interpretation.
Approach / What changed
The application unifies ingestion, transformation, governance, serving, analytics, and AI on Databricks. It uses Auto Loader, Spark Declarative Pipelines, Unity Catalog, liquid clustering, Lakebase, DBSQL, Genie, Vector Search, Agent Bricks, Model Serving, the Unity AI Gateway, and MLflow, with separate serving paths for replay and analytical workloads.
Takeaways
- Spark Declarative Pipelines process bronze, silver, and gold data while enforcing 46 named data-quality expectations for replay, analytics, and AI consumers.
- Lakebase serves sequential replay reads from synchronized Postgres tables at millisecond-level latency, while the Statement Execution API routes heavier exploratory analytics to SQL warehouses.
- The scouting supervisor combines Genie, Vector Search, a registered xG model, and Claude; MLflow traces execution and a deterministic scripted fallback prevents the agent from dead-ending.