OfficeQA Pro
General capabilityVerifiedLast updated 2026-10-04
Enterprise multi-document grounded reasoning over US Treasury bulletins.
OfficeQA Pro, published by Databricks AI Research on 9 March 2026, found frontier models answer under 5% of questions correctly from parametric knowledge alone and average just 34.1% when given the 89,000-page source corpus.
Key facts
| 2026 frontier | Frontier models under 5% from parametric knowledge alone; average 34.1% when given the documents |
|---|---|
| Credibility | A clean measurement of the difference between knowing and looking it up: 89,000 pages and 26M+ numeric values with 133 questions. Even with the documents in hand, frontier models average 34.1%, which is the number that matters for enterprise RAG planning. |
| Verification | Verified |
| Source | Databricks AI Research (arXiv:2603.08655) |
FAQ
What does OfficeQA Pro measure?
Enterprise multi-document grounded reasoning over US Treasury bulletins.
What is the 2026 frontier for OfficeQA Pro?
Frontier models under 5% from parametric knowledge alone; average 34.1% when given the documents
Is OfficeQA Pro credible?
A clean measurement of the difference between knowing and looking it up: 89,000 pages and 26M+ numeric values with 133 questions. Even with the documents in hand, frontier models average 34.1%, which is the number that matters for enterprise RAG planning.