Mistral AI·· 2025-08-01
Unlocking the potential of vision language models on satellite imagery through fine-tuning
Unlocking the potential of vision language models on satellite imagery through fine-tuning
AI summary
Mistral fine-tuned Pixtral-12B with LoRA, substantially outperforming the untuned baseline on Aerial Image Dataset (AID) satellite classification and reducing hallucinated invalid class names. Fine-tuning uses Mistral's API or LaPlateforme UI without extensive hyperparameter tuning, with 8,000 training and 2,000 test samples.
Selection record
Threshold 60Official, first-handFirst 38Second 38
Not admittedSum of both 76 < twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:微调Pixtral视觉语言模型技术内容
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Mistral AI · mistral.ai