MoonDPO
Memory-Efficient DPO for Hallucination Mitigation in Small VLMs
Applies LoRA-based DPO to Moondream2 (1.9B) using preference pairs automatically generated from COCO annotations, without human or API labels. The DPO objective is derived and implemented from scratch, and the full pipeline is reproducible on a 16 GB laptop.
Results placeholder: POPE-Adversarial F1: __ → __; CHAIRi: __ → __
Research Setup
- Base model
- Moondream2 (1.9B)
- Adaptation
- LoRA + DPO
- Preference data
- COCO-derived pairs
- Compute
- 16 GB laptop
- Evaluation
- POPE + CHAIR