Skip to content
A

OfficeQA Pro

General capabilityVerifiedLast updated 2026-10-04

Enterprise multi-document grounded reasoning over US Treasury bulletins.

OfficeQA Pro, published by Databricks AI Research on 9 March 2026, found frontier models answer under 5% of questions correctly from parametric knowledge alone and average just 34.1% when given the 89,000-page source corpus.

Key facts

2026 frontierFrontier models under 5% from parametric knowledge alone; average 34.1% when given the documents
CredibilityA clean measurement of the difference between knowing and looking it up: 89,000 pages and 26M+ numeric values with 133 questions. Even with the documents in hand, frontier models average 34.1%, which is the number that matters for enterprise RAG planning.
VerificationVerified
SourceDatabricks AI Research (arXiv:2603.08655)

Markdown version (for LLMs)

FAQ

What does OfficeQA Pro measure?

Enterprise multi-document grounded reasoning over US Treasury bulletins.

What is the 2026 frontier for OfficeQA Pro?

Frontier models under 5% from parametric knowledge alone; average 34.1% when given the documents

Is OfficeQA Pro credible?

A clean measurement of the difference between knowing and looking it up: 89,000 pages and 26M+ numeric values with 133 questions. Even with the documents in hand, frontier models average 34.1%, which is the number that matters for enterprise RAG planning.