Hugging Face Blog·· 10 d ago
Transformers now runs llama.cpp quants
Transformers now runs llama.cpp quants
AI summary
Hugging Face added GGUF support to transformers. Users can choose a quantised checkpoint from the Hub and pass gguf_file to from_pretrained for local generation without additional configuration.
Selection record
Threshold 60Official, first-handFirst 62Second 62
AdmittedSum of both 124 ≥ twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:Transformers支持GGUF本地推理,属AI开发工具
- Why it was chosen
- Provides official instructions for loading GGUF directly in transformers and compares local inference performance, supporting decisions on migrating existing local workflows.
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Hugging Face Blog · huggingface.co