Skip to content
Evaluation sources

BullshitBench v2

Peter Gostev / BullshitBench / Knowledge · Identify false premises in a question and observe whether the model keeps fabricating along the error.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

Uses the official default share of clearly identified attempts among all attempts, and retains the average score of answered responses; the refusal rate is kept for audit. Neither raw metric enters the current overall or category vote weight.

How this evidence is used

Original score reference: automatically collected after checking the aggregate, question sets, configuration and run date against the same public data version. Anonymous models are not disclosed, and rolling interfaces whose fixed version cannot be confirmed are not treated as representative model scores; they are not currently counted towards overall or category scores.

Limits and data attribution

All are false-premise questions and cannot represent general knowledge breadth; the model judge and refusal handling affect results, and rolling aliases and anonymous models must first complete identity verification.

Data licence: MIT · Copyright (c) 2026 Peter Gostev; official results are released with the same repository, with no separate data licence exception.

Scores published by Peter Gostev / BullshitBench. Raw scores and the News consensus score use different scales and cannot be added directly.