Transformers now runs llama.cpp quants
Hugging Face added GGUF support to transformers. Users can choose a quantised checkpoint from the Hub and pass gguf_file to from_pretrained for local generation without additional configuration.
AI that runs on phones, computers and edge devices: small models, local inference and on-device chips.
Hugging Face added GGUF support to transformers. Users can choose a quantised checkpoint from the Hub and pass gguf_file to from_pretrained for local generation without additional configuration.
Google Research proposed PhotoScan, a deep learning framework that estimates body fat percentage, A/G ratio and V/S ratio from ordinary 2D phone photos to predict insulin resistance.
Google DeepMind released SL2T, a multilingual sign-language-to-text model, bringing the capability to consumer products for the first time. Gboard and Live Transcribe on Pixel 11 support ASL-to-English dictation, with more devices and languages to follow.
Berkeley AI Research and IBM Research extended the K-Search evolutionary kernel search framework to MLX, transferring existing CUDA kernel knowledge to Apple Silicon through a structured CUDA-to-MLX translation layer.
Google Research published an architecture attaching multi-token prediction (MTP) heads to frozen Gemini Nano v3 models, accelerating on-device inference without changing backbone weights. It has shipped with the Pixel 9 and 10 series.
Google DeepMind released Gemma 4 12B, a multimodal model for local laptop use between the edge-focused E4B and 26B MoE. Its unified encoder-free architecture feeds visual and audio inputs directly into the LLM backbone.
Google Research published a study in Nature introducing PHRM, which records video with a phone’s front camera in the seconds after face unlock and uses deep learning to estimate heart rate and resting heart rate in the background.
Google DeepMind released the open-source Gemma 4 family in four sizes—E2B, E4B, 26B MoE and 31B Dense—under Apache 2.0.
Mistral released the Mistral 3 family, comprising 14B, 8B and 3B small dense models and its strongest yet Mistral Large 3, a sparse MoE with 41B active and 675B total parameters. All are open-source under Apache 2.0.