I’m an undergraduate researcher focused on post-training small Vision-Language Models (VLMs) to improve reasoning faithfulness and reduce hallucination. My current research explores reinforcement learning and preference-based post-training, including GRPO and DPO, with an emphasis on data-, label-, and memory-efficient methods.

Alongside my research, I’m completing a B.E. in Electrical Engineering at Chungnam National University and work as an AI/ML Engineer Intern at AIdanBio, where I build LLM/RAG systems, AI backends, and biomedical AI applications.

Research Interests

Vision-Language Models; Reinforcement Learning; Preference-Based Post-Training (GRPO, DPO); Hallucination Mitigation; Faithful and Grounded Multimodal Reasoning; Efficient Small VLMs.

Thura Win Kyaw

Thura Win Kyaw

AI/ML Research Candidate

Dec 2024 – Present

AI/ML Engineer Intern

AIdanBio · Daejeon, KR

  • Developed AI backend systems, APIs, model-serving infrastructure, and LLM-powered applications.
  • Built RAG pipelines using vector databases, semantic search, prompt engineering, and domain-specific knowledge bases.
  • Fine-tuned Vision-Language Models (VLMs) for medical AI using multimodal SFT, grounding, GRPO, and evaluation pipelines.
  • Worked on knowledge distillation and teacher–student training to transfer reasoning from large multimodal models to smaller models.
Oct 2025 – Nov 2025

AI Software Engineer Intern

GRINDA AI · Daejeon, KR

Built an AI-powered Slack bot that automatically converts issue reports into structured GitHub issues using Claude AI and FastAPI, with auto-labeling, translation, and monitoring.

B.E. in Electrical Engineering

Chungnam National University · Daejeon, KR

Sept 2022 – Present
VLM / Multimodal AI
Vision-Language Models, Hallucination Mitigation, Faithful and Grounded Multimodal Reasoning, Small VLMs
Post-Training
Supervised Fine-Tuning (SFT), Reinforcement Learning, GRPO, DPO, LoRA, Preference Data Generation, Knowledge Distillation
Evaluation
POPE, CHAIR, LLM/VLM Evaluation, Hallucination and Faithfulness Analysis
Applied LLM Systems
RAG Pipelines, LangChain Orchestration, Vector Databases, Prompt Engineering, Local LLM Deployment (Ollama)
Backend / Systems
FastAPI, API Development, AI Model Integration
Tools & Infrastructure
Git, GitHub, VS Code, Claude Code
Programming Languages
Python, C
Languages
Burmese (Native), English (Fluent), Korean (TOPIK 5)
Operating Systems
Windows, Linux, macOS