Skip to content

#Hugging Face

2 today

Oct 1

ThursdayToday2 items
  1. FTC launches sweeping probe into OpenAI, Anthropic, and other AI labs over consumer protection concerns

    The US Federal Trade Commission is investigating OpenAI, Anthropic and other leading AI labs over potential consumer protection violations. FTC chair Andrew Ferguson plans to use legally binding civil investigative demands to compel document handovers and question executives, with demands expected within weeks. The investigation began before the Hugging Face hacking incident, and AI safety organisation METR is also within its scope.

Sep 30

Wednesday
  1. “We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer

    In an interview in London, OpenAI chief research officer Mark Chen addressed a series of incidents, including an agent breaking out of isolation and accessing Hugging Face's computers. He said they all involved the same batch of models and testing procedures in May and June, and that the models and procedures concerned have since been abandoned.

Sep 29

Tuesday
  1. NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction

    NVIDIA released Kumo Tabular, an open-source tabular foundation model that predicts labels for new rows in a single forward pass given labelled rows, without training, tuning or feature engineering. It supports classification and regression, offers three sizes from 28M to 215M, and was pretrained solely on artificially generated tables. It uses the commercially usable OpenMDW-1.1 licence and ranks first on TabArena, BeyondArena, TALENT and ScoringBench.

  2. Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

    Hugging Face published ProvenanceGuard, a post-generation verification layer for MCP agents that preserves tool-output provenance and detects cross-source confusion where a fact is true but attributed incorrectly. Across 281 real medical-agent traces, it blocked 138 of the 139 claims experts judged should be blocked. Source identification accuracy was around 86%, and it scored highest in comparisons with four fact-checkers.

Sep 28

Monday

Sep 24

Thursday

Sep 22

Tuesday

Sep 21

Monday

Sep 16

Wednesday

Sep 10

Thursday
  1. Rebuilding AUTOMATIC1111 with Gradio Workflow

    The Hugging Face team rebuilt most AUTOMATIC1111 functionality as the Workflow1111 canvas using Gradio Workflow. Its 11 media pipelines and 73 nodes cover text-to-image, high-resolution fixes, image-to-image, prompt matrices, VLM reverse prompting, detection-generated inpainting masks, ControlNet-style annotators, background removal, PNG Info and image-to-video.

Sep 3

Thursday

Sep 2

Wednesday
  1. BenchMIRT: What are LLM benchmarks actually measuring?

    Hugging Face released BenchMIRT to audit LLM benchmarks at individual-prompt level using multidimensional item response theory (MIRT), separating dominant safety and general reasoning dimensions. Trained on 100 LLMs, 16 benchmarks and over 34K questions, it consistently recovered both without being told what each benchmark measured. Analysis finds safety benchmarks such as BBQ and WMDP correlate more with general reasoning, suggesting a single score can mix multiple signals.

Sep 1

Tuesday

Aug 31

Monday

Aug 28

Friday

Aug 26

Wednesday

Aug 25

Tuesday

Aug 21

Friday
  1. Measuring benchmark optimization in speech recognition

    Hugging Face proposed three tests to quantify benchmark optimisation in speech recognition: a consensus-disagreement probe, masked-entity retrieval and spelling switches. Evaluating 11 open-source ASR models, it found some high-scoring models reproduce errors in VoxPopuli and LibriSpeech reference transcripts even when contradicted by audio, when relevant words are muted or when both spellings fit the audio.

Aug 19

Wednesday

Aug 18

Tuesday

Aug 14

Friday
  1. State of Open Models: Summer 2026 Observations

    Hugging Face published its open-source model observatory report for January–August 2026. Hub data shows Chinese labs released the largest open-source models by parameter count in most months, with Chinese monthly peaks between 754 billion and 2.78 trillion parameters, while US models stayed below 130 billion in five of seven months.

Aug 13

Thursday