基本信息
任烁 中国科学院自动化研究所
电子邮件: shuo.ren@ia.ac.cn
通信地址: 北京海淀区中关村东路95号
邮政编码: 100080
电子邮件: shuo.ren@ia.ac.cn
通信地址: 北京海淀区中关村东路95号
邮政编码: 100080
研究领域
多模态大模型,空间推理/世界模型,科学智能体,语音/代码基础模型
招生信息
长期从事多模态大模型、基础模型表示学习和智能体系统研究,近期重点聚焦视觉语言模型的空间理解、世界模型增强和动态视觉推理,以及面向科学发现的大模型智能体系统。
招生专业
081104-模式识别与智能系统
招生方向
大模型智能体,视觉推理,AI for Science
教育背景
2016-09--2021-06 北京航空航天大学 博士2012-09--2016-06 北京航空航天大学 学士
学历
博士研究生
学位
博士
工作经历
工作简历
2024-01~现在, 中国科学院自动化研究所, 副研究员2021-07~2022-08,微软亚洲研究院, Researcher
出版信息
发表论文
(1) Look again, think slowly: Enhancing visual reflection in vision-language models, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, (2) Teaching vision-language models to ask: Resolving ambiguity in visual questions, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, (3) KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical Reasoning, Advances in Neural Information Processing Systems 2025, 2025, (4) Collaborative Beam Search: Enhancing LLM Reasoning via Collective Consensus, Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, (5) SpeechLM: Enhanced speech pre-training with unpaired textual data, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2024, (6) Codexglue: A machine learning benchmark dataset for code understanding and generation, Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022, (7) Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, (8) Optimizing alignment of speech and language latent spaces for end-to-end speech recognition and understanding, ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, (9) Speech pre-training with acoustic piece, InterSpeech 2022, 2022, (10) GraphCodeBERT: Pre-training code representations with data flow, ICLR 2021, 2021, (11) SemFace: Pre-training Encoder and Decoder with a Semantic Interface for Neural Machine Translation, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021, (12) WavLM: Large-scale self-supervised pre-training for full stack speech processing, IEEE Journal of Selected Topics in Signal Processing, 2021, (13) A graph-based coarse-to-fine method for unsupervised bilingual lexicon induction, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, (14) A retrieve-and-rewrite initialization method for unsupervised machine translation, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, (15) Unsupervised Neural Machine Translation with SMT as Posterior Regularization, AAAI 2019, 2019, (16) Explicit cross-lingual pre-training for unsupervised machine translation, Proceedings of the 2019 Conference on Empirical Methods in Natural Language, 2019, (17) Triangular architecture for rare language translation, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2018,