BullshitBench v2
Peter Gostev / BullshitBench / Knowledge · Identify false premises in a question and observe whether the model keeps fabricating along the error.
What it measures, and how
Uses the official default share of clearly identified attempts among all attempts, and retains the average score of answered responses; the refusal rate is kept for audit. Neither raw metric enters the current overall or category vote weight.
How this evidence is used
Original score reference: automatically collected after checking the aggregate, question sets, configuration and run date against the same public data version. Anonymous models are not disclosed, and rolling interfaces whose fixed version cannot be confirmed are not treated as representative model scores; they are not currently counted towards overall or category scores.
Limits and data attribution
All are false-premise questions and cannot represent general knowledge breadth; the model judge and refusal handling affect results, and rolling aliases and anonymous models must first complete identity verification.
Data licence: MIT · Copyright (c) 2026 Peter Gostev; official results are released with the same repository, with no separate data licence exception.
Scores published by Peter Gostev / BullshitBench. Raw scores and the News consensus score use different scales and cannot be added directly.