# OfficeQA Pro

> OfficeQA Pro, published by Databricks AI Research on 9 March 2026, found frontier models answer under 5% of questions correctly from parametric knowledge alone and average just 34.1% when given the 89,000-page source corpus.

Enterprise multi-document grounded reasoning over US Treasury bulletins.

- **Category:** General capability
- **2026 frontier:** Frontier models under 5% from parametric knowledge alone; average 34.1% when given the documents
- **Credibility:** A clean measurement of the difference between knowing and looking it up: 89,000 pages and 26M+ numeric values with 133 questions. Even with the documents in hand, frontier models average 34.1%, which is the number that matters for enterprise RAG planning.
- **Verification:** Verified
- **Source:** [Databricks AI Research (arXiv:2603.08655)](https://arxiv.org/abs/2603.08655)
- **Last updated:** 2026-10-04
