OfficeQA Pro
Databricks / Professional work · Retrieving, calculating and answering questions from large volumes of real financial documents.
What it measures, and how
Pro is a fixed 133 questions, separate from the 246 questions of Full and the 90 questions of Pro V2; end-to-end retrieval and directly providing the page containing the answer are different conditions. Scores for Claude Agent SDK, Codex SDK and Antigravity CLI belong to their respective run systems and cannot be mixed as a unified base-model evaluation. It measures file understanding, not the output quality of office documents.
How this evidence is used
Under observation: the official side keeps updating, but it currently mainly compares different products' run systems, and the raw aggregation protocol still needs verification; kept as reference material for understanding, without attributing scores from different tool systems directly to the base model.
Limits and data attribution
Code and question-bank licences do not automatically cover all external scores; old and new question sets, and end-to-end versus answer-page direct-supply conditions, cannot be mixed, and scores are not currently collected automatically.
Data licence: Code under Apache 2.0; question data under CC BY-SA 4.0, now moved to a Hugging Face dataset requiring access approval.
Scores published by Databricks. Raw scores and the News consensus score use different scales and cannot be added directly.