arXiv · 2026
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
Improving temporally grounded video reasoning through reinforcement learning and event-graph-based thinking.
arXiv · 2026
Improving temporally grounded video reasoning through reinforcement learning and event-graph-based thinking.
Pattern Recognition · 2026
Improving multimodal chain-of-thought reasoning by selectively combining specialised experts.
TMLR 2025
Improving visual question answering through training-free reasoning supported by lightweight vision tools.
ICML 2026
Using language sparse autoencoders to interpret and intervene on visual concepts in a vision-language model.
ICLR 2026
Recovering pretrained knowledge after adaptation through continued fine-tuning of vision-language models.
Under review · 2024
Subspace prompting for adapting vision-language models to multimodal tasks.
Selected papers co-authored by Da Li. For the full publication list, visit Google Scholar.