Evaluation sourcesOfficial evaluation
WeirdML
WeirdML / Coding · Implement machine learning solutions on unfamiliar data
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet
What it measures, and how
Compares on the actual task list and fixed GPU time, keeping the five-round scoring protocol.
How this evidence is used
Under observation: the official CSV is readable, but the score licence and per-model run times have not been fully verified.
Limits and data attribution
Under observation: the official CSV is readable, but the score licence and per-model run times have not been fully verified.
Data licence: Leaderboard data usage boundaries to be confirmed
Scores published by WeirdML. Raw scores and the News consensus score use different scales and cannot be added directly.