AIdanMED
Label-Efficient VLM Post-Training for Faithful, Grounded Medical Reasoning
Built during my internship at AIdanBio, this project trains a VLM with GRPO to produce reasoning together with a bounding box that must remain faithful to the model’s own reasoning, targeting label-efficient and grounded medical reasoning.
Status: In progress. SFT cold-start complete; GRPO post-training is running on 8×A100 GPUs. No evaluation results reported yet.
Research Setup
- Setting
- Medical vision-language
- Objective
- Faithful grounded reasoning
- Output
- Reasoning + bounding box
- Post-training
- SFT cold-start + GRPO
- Compute
- 8× NVIDIA A100
- Current stage
- GRPO training