I build multimodal agents that read messy documents, research questions across many sources, and operate software through its screen. Each one checks its own work before it hands anything back.
VinSmart Future · 2026
Staff engineer
Agentic document parsing
- layout model
- VLM OCR
- visual reflection
- self-heal
- Multi-agent loop: agents compare their parse to the page image and repair what doesn’t match.
- Optimized the layout model and VLM behind OCR on hard, scanned PDFs.
WER 9% → 5%on a challenging scanned-PDF test set
VinSmart Future · 2026
Main researcher
Deep research agent
- explore
- discover
- cross-source fact-check
- self-reflect
- Agent framework for thorough, long-horizon exploration and knowledge discovery.
- Raised citation accuracy and cut hallucinated citations by checking claims across sources.
FPT Smart Cloud · 2024 – 2026
Main researcher
Computer-use agent for enterprise workflows
- perceive screen
- plan
- act
- Designed and optimized the agent workflow that automates business processes end to end.
- Benchmarked commercial and open-source LLMs for computer use.
- Fine-tuned self-hosted LLMs for screen perception and planning.
FPT Smart Cloud · 2024 – 2026
Lead researcher
Ad video generation for clothing
- business brief
- diffusion video model
- ad video
- Compared open-source diffusion models for video against the customer’s requirements.
- Designed a fully automated generation system and shipped it.
Deployedin production for the customer