Skip to content
Evaluation sources

FACTS Parametric

Google DeepMind / Kaggle / Knowledge · Answers world-fact questions without search tools, testing the accuracy of the model's own knowledge.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

Only the Score of the fixed task version is considered, retaining the public set, private set and confidence intervals for verification; the three related scores must not be counted twice. FACTS is not among the ten underlying evaluations of AA v4.3.

How this evidence is used

Under observation: the official protocol and licence are confirmed, but there is no completed-question count or completeness marker; some row intervals are notably wide, and the aggregation basis for the total score against the public and private sets also has points to confirm. Awaiting clarification from the evaluator, without deleting anomalous rows to force inclusion, and without using scoring budget.

Limits and data attribution

Both the judge and the evaluation initiator come from Google, so other organisations' evidence is needed. The export does not state whether some results completed the same problem set, so partial results cannot be treated as a full evaluation; sample sizes back-calculated from intervals are not officially confirmed facts.

Data licence: Apache 2.0 · stated explicitly on the official benchmark and Score task pages; Kaggle documentation provides a login-free JSON download of the public leaderboard.

Scores published by Google DeepMind / Kaggle. Raw scores and the News consensus score use different scales and cannot be added directly.