Tao Jin  (金涛)


Research Interests: Multimedia Analysis, Computer Vision, Natural Language Learning, Transfer Learning,
Address: Hangzhou/Ningbo, Zhejiang Province
Email: jint_zju@zju.edu.cn

Education

`

Work Experiences

  • Research Intern at Taobao Research
    Taobao Lab
    June 2020 - Sep 2020       Hangzhou, China
  • Research Intern at Kuake Research
    Kuake Lab
    Nov 2019 - Feb 2020       Hangzhou, China

Supervised and Co-supervised Students

  • School of Software, Zhejiang University
    Wang Lin (2021, linwanglw@zju.edu.cn, National Scholarship, PHD of ZJU),
    Linjun Li (2021, lilinjun21@zju.edu.cn, National Scholarship, Beidou Plan of Meituan),
    Xize Cheng (2021, chengxize@zju.edu.cn, National Scholarship, PHD of ZJU),
    Ye Wang (2021, yew@zju.edu.cn, National Scholarship, Daka of Tecent&2-1 of Bytedance),
    Zirun Guo (2024),
    Weicai Yan (2024),
    Dongjie Fu (2024),
    Xiaoda Yang (2024),

Publications(* denotes equal contributions, & denotes corresponding author)

  1. Concept Preservation and Unbinding in Continual Diffusion Customization
    Zirun Guo, Tao Jin,
    CVPR, 2025

  2. Towards Transformer-Based Aligned Generation with Self-Coherence Guidance
    Shulei Wang, Wang Lin, Tao Jin, Zhou Zhao,
    CVPR, 2025

  3. Non-Natural Image Understanding with Advancing Frequency-based Vision Encoders
    Wang Lin, Tao Jin, Zhou Zhao, Jingyuan Chen,
    CVPR, 2025

  4. Efficient Prompting for Continual Adaptation to Missing Modalities
    Zirun Guo, Shulei Wang, Wang Lin, Weicai Yan, Yangyang Wu, Tao Jin,
    NAACL, 2025

  5. Omni-Chart-600K: A Comprehensive Dataset of Chart Types for Chart Understanding
    Shulei Wang, Shuai Yang, Wang Lin, Zirun Guo, Sihang Cai, Hai Huang, Ye Wang, Jingyuan Chen, Tao Jin&,
    NAACL, 2025

  6. Data-Efficiently Learn Large Language Model for Universal 3D Scene Perception
    Zehan Wang, Haifeng Huang, Yang Zhao, Ziang Zhang, Tao Jin, Zhou Zhao,
    NAACL, 2025

  7. Smoothing the Shift: Towards Stable Test-time Adaptation under Complex Multimodal Noises
    Zirun Guo, Tao Jin&,
    ICLR, 2025

  8. Diff-Prompt: Diffusion-driven Prompt Generator with Mask Supervision
    Weicai Yan, Wang Lin, Tao Jin&,
    ICLR, 2025

  9. OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
    Xize Cheng, Tao Jin, Zhou Zhao,
    ICLR, 2025

  10. Improving Multi-modal Representations via Binding Space in Scale
    Zehan Wang, Ziang Zhang, Minjie Hong, Tao Jin, Hengshuang Zhao, Zhou Zhao,
    ICLR, 2025

  11. VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?
    Xize Cheng, Tao Jin, Zhou Zhao,
    ICLR, 2025

  12. Curriculum Learning aided Audio-Visual Speech Recognition with Arbitrary Number of Overlapping Speakers
    Yuxiao Lin, Tao Jin&, Xize Cheng, Zhou Zhao, Fei Wu,
    ICASSP, 2025

  13. Bridging the Gap for Test-time Multimodal Sentiment Analysis
    Zirun Guo, Tao Jin&, Wenlong Xu, Wang Lin, Yangyang Wu,
    AAAI, 2025

  14. Low-rank Sequence Adapter for Efficient Multimodal Transfer Learning
    Zirun Guo, Xize Cheng, Yangyang Wu, Tao Jin&,
    AAAI, 2025

  15. Speech Watermarking with Discrete Intermediate Representations
    Shengpeng Ji, Ziyue Jiang, Jialong Zuo, Minghui Fang, Yifu Chen, Tao Jin, Zhou Zhao,
    AAAI, 2025

  16. Exploring Embodied Emotion Through A Large-Scale Egocentric Video Dataset
    Wang Lin, Tao Jin, Zhou Zhao, Chang Yao, Jingyuan Chen,
    NeurIPS, 2024

  17. Action Imitation in Common Action Space for Customized Action Image Synthesis
    Wang Lin, Jingyuan Chen, Zirun Guo, Tao Jin, Zhou Zhao,
    NeurIPS, 2024

  18. Balancing Multimodal Learning with Classifier-guided Gradient Modulation
    Zirun Guo, Tao Jin&,
    NeurIPS, 2024

  19. AudioVSR: Enhancing Video Speech Recognition with Audio Data
    Xiaoda Yang, Xize Cheng, Tao Jin&,
    EMNLP, 2024

  20. Calibrating Prompt from History for Continual Vision-Language Retrieval and Grounding
    Tao Jin, Weicai Yan, Ye Wang, Zhou Zhao,
    ACM MM, 2024

  21. Boosting Speech Recognition Robustness to Modality-Distortion with Contrast-Augmented Prompts
    Dongjie Fu, Xize Cheng, Xiaoda Yang, Tao Jin&, Zhou Zhao,
    ACM MM (ORAL), 2024

  22. SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning
    Xiaoda Yang, Xize Cheng, Dongjie Fu, Tao Jin&, Zhou Zhao,
    ACM MM, 2024

  23. Low-rank Prompt Interaction for Continual Vision-Language Retrieval
    Weicai Yan, Ye Wang, Tao Jin&, Zhou Zhao,
    ACM MM, 2024

  24. TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation
    Xize Cheng, Tao Jin, Zhou Zhao,
    ACL, 2024

  25. Uni-Dubbing: Zero-Shot Speech Synthesis from Visual Articulation
    Songju Lei, Xize Cheng&, Tao Jin, Zhou Zhao,
    ACL, 2024

  26. Rethinking the Multimodal Correlation of Multimodal Sequential Learning
    Tao Jin, Zhou Zhao,
    ACL, 2024

  27. Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
    Zirun Guo, Tao Jin&, Zhou Zhao,
    ACL, 2024

  28. Two-Stream Generative Recommender with Behavior-Semantic Collaboration
    Ye Wang, Jiahao Xun, Minjie Hong, Jieming Zhu, Tao Jin, Zhou Zhao,
    KDD, 2024

  29. Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
    Yongqi Wang, Ruofan Hu, Rongjie Huang, Zhiqing Hong, Ruiqi Li, Wenrui Liu, Fuming You, Tao Jin, Zhou Zhao,
    NAACL, 2024

  30. Non-confusing Generation of Customized Concepts in Diffusion Models
    Wang Lin, Jingyuan Chen, Tao Jin, Zhou Zhao,
    ICML, 2024

  31. Molecule-Space: Free Lunch in Unified Multimodal Space via Knowledge Fusion
    Zehan Wang, Ziang Zhang, Xize Cheng, Rongjie Huang, Luping Liu, Tao Jin, Zhou Zhao,
    ICML, 2024

  32. MPOD123: One Image to 3D Content Generation Using Mask-enhanced Progressive Optimization
    Jimin Xu*, Tianbao Wang*, Tao Jin&, Zhou Zhao,
    CVPR, 2024

  33. Rethinking Missing Modality Learning from a Decoding Perspective
    Tao Jin, Zhou Zhao,
    ACM MM, 2023

  34. Exploring Group-Based Video Captioning with Efficient Relational Approximation
    Wang Lin*, Tao Jin*, Ye Wang, Zhou Zhao,
    ICCV, 2023

  35. Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition
    Xize Cheng*, Tao Jin*, Linjun Li, Zhou Zhao,
    ICCV, 2023

  36. Multi-Granularity Relational Attention Network for Audio-Visual QA
    Linjun Li*, Tao Jin*, Wang Lin, Hao Jiang, Zhou Zhao,
    TCSVT, 2023

  37. OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment
    Xize Cheng*, Tao Jin*, Linjun Li, Wang Lin, Xinyu Duan,
    ACL (ORAL), 2023

  38. TAVT: Towards Transferable Audio-Visual Text Generation
    Wang Lin*, Tao Jin*, Ye Wang, Wenwen Pan, Xize Cheng, Linjun Li, Zhou Zhao
    ACL, 2023

  39. Semantic-Conditioned Dual Adaptation for Query-based Visual Segmentation
    Ye Wang*, Tao Jin*, Wang Lin, Xize Cheng, Linjun Li, Zhou Zhao
    ACL, 2023

  40. Weakly-Supervised Spoken Video Grounding via Semantic Interaction Learning
    Ye Wang, Wang Lin, Shengyu Zhang, Tao Jin, Zhou Zhao
    ACL (ORAL), 2023

  41. Contrastive Token-Wise Meta-Learning for Unseen Temporal-Aligned Translation
    Linjun Li*, Tao Jin*, Xize Cheng, Ye Wang, Wang Lin, Rongjie Huang, Zhou Zhao,
    ACL, 2023

  42. DATE: Domain Adaptive Product Seeker for E-commerce
    Haoyuan Li, Hao Jiang, Tao Jin, Mengyan Li, Yan Chen, Zhijie Lin, Yang Zhao, Zhou Zhao,
    CVPR, 2023

  43. Gloss Attention for Gloss-free Sign Language Translation
    Aoxiong Yin, Tianyun Zhong, Li Tang, Weike Jin, Tao Jin, Zhou Zhao,
    CVPR, 2023

  44. Interaction Augmented Transformer with Decoupled Decoding for Video Captioning
    Tao Jin, Zhou Zhao, Peng Wang, Jun Yu, Fei Wu,
    Neurocomputing, 2022

  45. MC-SLT: Towards Low-Resource Signer-Adaptive Sign Language Translation
    Tao Jin, Zhou Zhao, Meng Zhang, Xingshan Zeng,
    ACM MM, 2022

  46. Prior Knowledge and Memory Enriched Transformer for Sign Language Translation
    Tao Jin, Zhou Zhao, Meng Zhang, Xingshan Zeng,
    ACL, 2022

  47. Generalizable Multi-Linear Attention Network
    Tao Jin, Zhou Zhao,
    NeurIPS, 2021

  48. Contrastive Disentangled Meta-Learning for Signer-Independent Sign Language Translation
    Tao Jin, Zhou Zhao,
    ACM MM (ORAL), 2021

  49. Dual Low-Rank Multimodal Fusion
    Tao Jin*, Siyu Huang*, Yingming Li, Zhongfei Zhang
    EMNLP, 2020

  50. SBAT: Video Captioning with Sparse Boundary-Aware Transformer
    Tao Jin, Siyu Huang, Ming Chen, Yingming Li, Zhongfei Zhang
    IJCAI, 2020

  51. Low-Rank HOCA: Efficient High-Order Cross-Modal Attention for Video Captioning
    Tao Jin, Siyu Huang, Yingming Li, Zhongfei Zhang,
    EMNLP, 2019

  52. Recurrent Convolutional Video Captioning with Global and Local Attention
    Tao Jin, Yingming Li, Zhongfei Zhang,
    Neurocomputing, 2019


Contest


Web Site Hit Counter Since Nov, 2018

Proudly powered by Bootstrap