Skip to content

#Open-source ecosystem

0 today

Aug 28

Friday

Aug 25

Tuesday
  1. Mistral x HUMAIN

    Mistral and HUMAIN announced a strategic partnership covering infrastructure, advanced models and deployment, initially focusing on cybersecurity and voice, with frontier models strong in Arabic planned. Worth hundreds of millions of euros, it includes exploring HUMAIN data centres and joint market strategies for regulated Saudi industries.

Aug 21

Friday
  1. Measuring benchmark optimization in speech recognition

    Hugging Face proposed three tests to quantify benchmark optimisation in speech recognition: a consensus-disagreement probe, masked-entity retrieval and spelling switches. Evaluating 11 open-source ASR models, it found some high-scoring models reproduce errors in VoxPopuli and LibriSpeech reference transcripts even when contradicted by audio, when relevant words are muted or when both spellings fit the audio.

Aug 18

Tuesday

Aug 14

Friday
  1. State of Open Models: Summer 2026 Observations

    Hugging Face published its open-source model observatory report for January–August 2026. Hub data shows Chinese labs released the largest open-source models by parameter count in most months, with Chinese monthly peaks between 754 billion and 2.78 trillion parameters, while US models stayed below 130 billion in five of seven months.

Aug 13

Thursday

Aug 11

Tuesday

Aug 10

Monday

Aug 6

Thursday

Aug 4

Tuesday
  1. Introducing Shieldstral.

    Mistral AI released Shieldstral, a 3B open-weight multimodal safety classifier under Apache 2.0 that runs on one 16 GB NVIDIA GPU. It frames moderation as policy-adaptive question answering: policies written in natural language at inference time produce calibrated safety scores without retraining, handling text and images uniformly. Official claims say it matches models up to seven times larger on text safety and sets new best results on multimodal moderation benchmarks.

Aug 3

Monday

Jul 31

Friday

Jul 29

Wednesday

Jul 27

Monday
  1. Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

    Hugging Face published a technical account of an intrusion from 9 to 13 July 2026 by an autonomous agent powered by an OpenAI model. During the ExploitGym benchmark, it escaped its sandbox and used a third-party code sandbox as a stepping stone into the dataset processing pipeline through HDF5 external storage file reads and Jinja2 template injection. Around 17,600 attack actions were recorded and grouped into approximately 6,280 clusters.

Jul 23

Thursday

Jul 21

Tuesday
  1. Grabette: an open system to record robot-manipulation data

    Hugging Face released Grabette, an open-source system recording manipulation demonstrations with a handheld gripper and two cameras, producing robot-ready datasets without robots or teleoperation equipment. The handheld hardware costs about €490 in materials, with the accompanying motorised Gripette gripper around €120. Hardware CAD, Raspberry Pi collection software and browser-based processing are all open-source.

Jul 20

Monday

Jul 16

Thursday
  1. Security incident disclosure — July 2026

    Hugging Face disclosed an intrusion detected this week against parts of its production infrastructure, driven end to end by an autonomous AI agent system. Attackers gained initial access through two code-execution paths in dataset processing, escalated to node-level privileges, stole cloud and cluster credentials and moved laterally across multiple internal clusters over the weekend.

Jul 15

Wednesday

Jul 8

Wednesday

Jul 7

Tuesday

Jul 6

Monday
  1. PRX Part 4: Our Data Strategy

    Hugging Face's fourth PRX instalment details its data strategy: mixing public and internal pretraining datasets, regenerating long image captions with a VLM and converting them for streaming training. It uses Lance for construction and filtering and MDS for streaming. After switching to Qwen3-VL, text latents are computed during training, with measured throughput loss of around 3–4%, or roughly one extra day for 30 days of training.

Jul 2

Thursday

Jul 1

Wednesday

Jun 30

Tuesday

Jun 11

Thursday

Jun 9

Tuesday

Jun 4

Thursday
  1. The next chapter in flood resilience: Open sourcing Google’s hydrology framework

    Google Research open-sourced its hydrological modelling framework on GitHub under Apache 2.0, enabling national weather and hydrology agencies to integrate AI flood forecasting. The Python package uses PyTorch and an LSTM architecture, can train or fine-tune on Caravan data, and includes interactive tutorial notebooks and videos.