---
title: "Introducing OfficeQA Pro V2: A New Benchmark for Enterprise Grounded-Reasoning"
description: "OfficeQA Pro V2 benchmarks whether AI agents generalize grounded reasoning to unfamiliar enterprise document collections rather than specialize in one corpus or task distribution. It contains 90 questions grounded in roughly 120,000 pages of U.S. Treasury accounting records spanning 1793–2024, including dense tables, charts, revised values, and changing reporting conventions. Questions were generated with asynth from seeded inputs, corpus evidence searches, traceable reasoning, and quality gates, then evaluated with deterministic exact match at 0.0% tolerance. Out-of-the-box frontier agents averaged 37.5% in one matched comparison, while competition systems averaged 41.1% and the winning team reached 63.3%; Databricks Genie raised matched-model accuracy by 24.0 percentage points on average, although failures remained in parsing, temporal reconciliation, and entity or category interpretation."
---

# Introducing OfficeQA Pro V2: A New Benchmark for Enterprise Grounded-Reasoning

[Databricks](https://yomu.fyi/company/databricks) · Databricks AI Research Team · Aug 6, 2026

**Type:** Benchmark

## Summary

OfficeQA Pro V2 benchmarks whether AI agents generalize grounded reasoning to unfamiliar enterprise document collections rather than specialize in one corpus or task distribution. It contains 90 questions grounded in roughly 120,000 pages of U.S. Treasury accounting records spanning 1793–2024, including dense tables, charts, revised values, and changing reporting conventions. Questions were generated with asynth from seeded inputs, corpus evidence searches, traceable reasoning, and quality gates, then evaluated with deterministic exact match at 0.0% tolerance. Out-of-the-box frontier agents averaged 37.5% in one matched comparison, while competition systems averaged 41.1% and the winning team reached 63.3%; Databricks Genie raised matched-model accuracy by 24.0 percentage points on average, although failures remained in parsing, temporal reconciliation, and entity or category interpretation.

## Context

The benchmark addresses whether improvements on the original OfficeQA reflect broader grounded-reasoning progress or specialization to one corpus and task distribution. Enterprise agents must operate across unfamiliar, heterogeneous document collections with changing reporting conventions.

## Approach / What changed

OfficeQA Pro V2 uses 90 questions grounded in roughly 120,000 pages of U.S. Treasury accounting records. Its questions were created with the asynth synthetic-data pipeline using seeded inputs, corpus search, traceable reasoning, and quality gates; the benchmark also compares provider harnesses with Databricks Genie.

## Takeaways

- The corpus contains roughly 1,400 PDFs and 120,000 pages of U.S. accounting records spanning 1793 through 2024, where concepts can change in name, location, schema, unit, time basis, and aggregation level.
- Databricks Genie improved accuracy by an average of 24.0 percentage points over default model-provider harnesses; in one Claude Fable 5 comparison, it also reduced cost by approximately nine times.
- Systems continued to fail on parsing fidelity, temporal reconciliation as accounting conventions changed, and interpreting entity scope or category granularity.

**Tags:** [Data Pipelines](https://yomu.fyi/topic/data-pipelines), [LLMs](https://yomu.fyi/topic/llm), [Performance](https://yomu.fyi/topic/performance), [Testing](https://yomu.fyi/topic/testing)

- Source: [Databricks](https://www.databricks.com/blog/introducing-officeqa-pro-v2-new-benchmark-enterprise-grounded-reasoning)
- Source URL: https://www.databricks.com/blog/introducing-officeqa-pro-v2-new-benchmark-enterprise-grounded-reasoning
- Ingested by Yomu: 2026-08-30T16:52:06.445Z

[Read original post](https://www.databricks.com/blog/introducing-officeqa-pro-v2-new-benchmark-enterprise-grounded-reasoning)
