I am currently an Agent Expert at Xiaohongshu, focusing on cutting-edge research in Multimodal Large Language Models (MLLMs), MLLM Agents, and Agentic RL.
I received my Ph.D. degree from the Multimedia and Human Understanding Group (MHUG) at the Department of Information Engineering and Computer Science, University of Trento, Italy, in 2022. I was supervised by Prof. Nicu Sebe and Dr. Bruno Lepri, with my thesis defense committee including Vittorio Murino, Zhengyou Zhang, and Elisa Ricci.
Before my doctoral studies, I earned my B.Eng. degree in Photogrammetry and Remote Sensing (2015) and M.Eng. degree in Pattern Recognition and Intelligent System (2018) from Wuhan University, China.
We are actively recruiting daily interns for long-term positions. Please feel free to submit your resume to my email for exciting research opportunities!
Email: yahui.cvrs AT gmail.com
Experience
- 04/2026 - Present
Agent Expert, Xiaohongshu, Hangzhou, China
Research focus: MLLM Agent, Agentic RL - 01/2025 - 03/2026
Researcher, Kuaishou Technology, Beijing, China
Research focus: MLLMs, Formal Theorem Proving and AI Agents - 08/2022 - 01/2025
Researcher, Huawei, Shenzhen, China
Research focus: Image Generation and Enhancing (GANs and Diffusion Models) - 2021 - 06/2022
Research Intern, Tencent AI Lab, Shenzhen, China
Mentors: Dr. Linchao Bao and Dr. Wei Bi
Research focus: GANs, Image Domain Translation - 12/2018 - 06/2022
PhD Student, FBK and MHUG, Trento, Italy
Mentors: Prof. Nicu Sebe and Dr. Bruno Lepri
Research focus: Deep learning, GANs, Cross-modal Representations, Image Domain Translation - 11/2017 - 09/2018
Research Intern, Tencent AI Lab, Shenzhen, China
Mentors: Dr. Wei Bi and Dr. Xiaojiang Liu
Research focus: Deep Learning, Neural Dialogue Generation - 03/2015 - 06/2018
Master Student, Computer Vision and Remote Sensing (CVRS) Lab, Wuhan, China
Mentor: Prof. Jian Yao
Research focus: Deep Learning, Remote Sensing
Publications (Google Scholar Profile)
(*: equal contribution, †: equal contribution, †: correspondence)
Conference Papers
- DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing [arxiv]
ACL 2026 (Oral) - All-in-One Slider for Attribute Manipulation in Diffusion Models [arxiv]
CVPR 2026 - Evaluating Text Creativity Across Diverse Domains: A Dataset and Large Language Model Evaluator [arxiv]
ICLR 2026 - Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning [arxiv]
AAAI 2026 - What Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought Reasoning [arxiv]
EMNLP 2025 - Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search [arxiv] [code]
ACL 2025 - Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision Transformers [arxiv] [code]
CVPR 2023 - Efficient Training of Visual Transformers with Small Datasets [arxiv] [poster] [code]
NeurIPS 2021 - Smoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image Translation [supp] [video] [code]
CVPR 2021 - Describe What to Change: A Text-guided Unsupervised Image-to-Image Translation Approach [arxiv] [code]
ACM MM 2020
Journal Articles
- SeqPE: Transformer with Sequential Position Encoding [arxiv]
IEEE TPAMI, 2026 - A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models [arxiv]
IEEE TPAMI, 2025 - Spatial Entropy as An Inductive Bias for Vision Transformers
Machine Learning, 2024 (IF: 5.8) - ISF-GAN: An Implicit Style Function for High-Resolution Image-to-Image Translation
IEEE TMM, 2022 (IF: 8.4) - DeepCrack: A Deep Hierarchical Feature Learning Architecture for Crack Segmentation [code]
Neurocomputing, 2019 (IF: 5.5) - RoadNet: Learning to Comprehensively Analyze Road Networks in Complex Urban Scenes From High-Resolution Remotely Sensed Images [code]
IEEE TGRS, 2019 (IF: 7.5)
Recent News
- May 2026 — SeqPE has been accepted to TPAMI.
- May 2026 — DPWriter has been notified of an oral presentation at ACL 2026.
- April 2026 — Joined Xiaohongshu as Agent Expert, focusing on MLLM Agent and Agentic RL.
- April 2026 — DPWriter was accepted to ACL 2026.
- February 2026 — All-in-One Slider was accepted to CVPR 2026.
- January 2026 — CrEval was accepted to ICLR 2026.
- November 2025 — One paper was accepted to TPAMI.
- October 2026 — One paper was accepted to AAAI 2026.
- July 2025 — We released Leanabell-Prover-V2 for verifier-integrated reasoning via RL.
- May 2025 — We released LCoT2Tree for uncovering structural patterns in Long CoT, accepted to EMNLP.
- May 2025 — We released the UNITE framework for Multimodal Information Retrieval.
- May 2025 — One paper accepted to ACL main conference: MCTS-VCB.
- April 2025 — We released Capybara-VL and Capybara-Omni at ICLR 2025 SCI-FM workshop — our efficient MLLMs.
- April 2025 — We released Leanabell-Prover achieving SOTA 59.8% pass@32 on MiniF2F-test.
Academic Services
Conference Reviews
- ICML 2026, 2025
- NeurIPS 2026, 2025, 2024, 2023, 2022
- ICLR 2027, 2026, 2025, 2024
- CVPR 2026, 2025, 2024, 2023, 2022, 2021
- ICCV 2025, 2023, 2021
- AAAI 2026, 2025
- ACM MM 2026, 2025, 2024, 2023, 2022, 2021, 2020
- ACL/EMNLP 2026, 2025, 2024
- ECCV 2026, 2024, 2022
- IJCAI 2022, 2021
Journal Reviews
- IEEE TPAMI
- International Journal of Computer Vision (IJCV)
- IEEE Transactions on Industrial Informatics (TII)
- IEEE GRSL
- IEEE J-STARS
- IEEE TNNLS
- Machine Vision and Applications (MVAP)
- IEEE TMM
- Pattern Recognition Letters (PRL)
- Information Fusion
Selected Awards
- Pengcheng Excellent Talents(鹏城优才), Shenzhen, China, 2024
- Top Minds (天才少年) (第一档), Huawei, China, 2022
- Technical Expert (技术大咖), Tencent AI Lab, China, 2021