Loading…
Databricks partners with OpenAI on GPT-5.5
Hanlin Tang, Ahmed Bilal, Arnav Singhvi, Ivan Zhou, Harish Gaur
- Source
- Databricks
- Published
- Added to Yomu
Summary
Databricks announces a partnership with OpenAI around GPT-5.5, described as OpenAI’s strongest frontier model for enterprise agentic work, complex document reasoning, and long-horizon coding agents. The model powers Codex and is presented as able to research, analyze data, create documents and spreadsheets, operate software, use tools, check outputs, recover from ambiguity, and continue through multi-part tasks. Databricks evaluated it on OfficeQA, a benchmark built from 89,000 pages of U.S. Treasury Bulletins that tests document retrieval, table interpretation, and precise calculation. With retrieval handled, GPT-5.5 scored 64.66% versus GPT-5.4’s 57.14%; in the full-agent OfficeQA Pro Agent Harness, it scored 52.63% versus 36.10%, representing reported improvements of roughly 13% and a 46% reduction in errors.
Context
Databricks evaluated GPT-5.5 to understand how its improvements translate to real enterprise workloads, using document-heavy, multi-step analytical tasks represented by the OfficeQA benchmark.
Approach / What changed
The partnership brings GPT-5.5’s agentic capabilities to Databricks, while the evaluation uses both a retrieval-assisted OfficeQA setup and a full-agent Codex harness in which the model finds documents, parses them, and computes answers.
Takeaways
- OfficeQA is built from 89,000 pages of U.S. Treasury Bulletins and measures document retrieval, complex table interpretation, and precise calculations.
- With retrieval already handled, GPT-5.5 scored 64.66%, compared with GPT-5.4’s 57.14%, described as a roughly 13% improvement.
- In the full-agent OfficeQA Pro Agent Harness, GPT-5.5 scored 52.63% versus GPT-5.4’s 36.10%, with a reported 46% reduction in errors.