Skip to content
Evaluation sources

OfficeQA Pro

Databricks / Professional work · Retrieving, calculating and answering questions from large volumes of real financial documents.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

Pro is a fixed 133 questions, separate from the 246 questions of Full and the 90 questions of Pro V2; end-to-end retrieval and directly providing the page containing the answer are different conditions. Scores for Claude Agent SDK, Codex SDK and Antigravity CLI belong to their respective run systems and cannot be mixed as a unified base-model evaluation. It measures file understanding, not the output quality of office documents.

How this evidence is used

Under observation: the official side keeps updating, but it currently mainly compares different products' run systems, and the raw aggregation protocol still needs verification; kept as reference material for understanding, without attributing scores from different tool systems directly to the base model.

Limits and data attribution

Code and question-bank licences do not automatically cover all external scores; old and new question sets, and end-to-end versus answer-page direct-supply conditions, cannot be mixed, and scores are not currently collected automatically.

Data licence: Code under Apache 2.0; question data under CC BY-SA 4.0, now moved to a Hugging Face dataset requiring access approval.

Scores published by Databricks. Raw scores and the News consensus score use different scales and cannot be added directly.