Peijie Dong

Peijie Dong 董佩杰

PhD student in DSA, The Hong Kong University of Science and Technology (Guangzhou)

I am Peijie Dong (董佩杰), a final-year Ph.D. candidate in the Data Science and Analytics Thrust at the Hong Kong University of Science and Technology (Guangzhou), advised by Prof. Xiaowen Chu and Prof. Junxian He. I am currently a research intern with the WorkBuddy/CodeBuddy Coding Agent team at Tencent CSIG, where I work on post-training and evaluation for long-horizon coding agents.

Research Interests

My research focuses on improving the ability of coding agents to solve long-horizon, repository-level software engineering tasks. I am particularly interested in transforming interaction trajectories and environment feedback into effective training signals. My current research interests include:

  • Coding Agent Post-Training: Developing data and training recipes for coding agents, including trajectory curation, supervised fine-tuning, reinforcement learning, and reward design.
  • Long-Horizon Agent Evaluation: Building benchmarks and agent harnesses to study planning, tool use, repository navigation, error recovery, and end-to-end task completion.
  • Agent Data and Training-Evaluation Loops: Diagnosing behavioral failures from agent trajectories and translating them into targeted data, objectives, and evaluation signals.
  • Efficient Large Language Models: Improving the efficiency of LLM training and inference through model compression, low-precision training, efficient architectures, and systems optimization.

My long-term goal is to build coding agents that can learn from complete interaction trajectories and reliably improve through real-world task feedback. I welcome discussions and collaborations on coding agents, post-training, evaluation, and efficient LLMs.

🔥 News

  • 🎉🎉 Our paper “CPA: Efficient and Stable FP4 RL Training via Cross-Precision Alignment” is accepted by NeurIPS 2026.

  • 🎉🎉 Our paper “Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning” is accepted by EMNLP 2026 Main Conference.

  • 🎉🎉 Our tech report “Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction” is released to Arxiv.

  • 🎉🎉 Our paper “An Empirical Study of Reasoning Degradation in Quantized Multimodal Large Language Models” is accepted by ACM MM 2026.

  • 🎉🎉 Our paper “GreenMoE: Exploiting Dynamic Load Imbalance for Energy-Efficient Long-Context MoE Training” is accepted by ICML 2026 AdaptFM Workshop.

Earlier news
  • 🎉🎉 Our paper “Parameters as Agentic Memory: Internalizing Long-Horizon Memories for Efficient LLM Agents” is accepted by ICML 2026 AIWILD Workshop.

  • 🎉🎉 Our paper “Enhancing Knowledge Injection with Surrounding Backgrounds in Continual Training LLMs” is accepted by ICML 2026 FoGen Workshop.

  • 🎉🎉 Our paper “Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression” is accepted by ICML 2026.

  • 🎉🎉 Our paper “VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing” is accepted by ICML 2026.

  • 🎉🎉 Our paper “Identifying and Mitigating Errors in Gradient Aggregation of Distributed Data Parallel Training” is accepted by ICML 2026.

  • 🎉🎉 Our Paper “Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Understanding” is accepted by ICLR2026.

  • 🎉🎉 Two years after graduation, I was selected as an outstanding master’s student at the NUDT in Hunan Province.

  • 🎉🎉 Our Paper “ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference” is accepted by NeurIPS 2025.

  • 🎉🎉 Our Paper “Perovskite-LLM: Knowledge-Enhanced Large Language Models for Perovskite Solar Cell Research” is accepted by EMNLP 2025 findings.

  • 🎉🎉 Our Paper “Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks” is released to arxiv.

  • 🎉🎉 Our tech report “Intern-S1: A Scientific Multimodal Foundation Model” is released to arxiv. Great work by Intern-S1 team.

  • 🎉🎉 Our Paper “Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compresssion” is accepted by ICML25. We are especially grateful to the reviewer who awarded us a ‘5 (Strong Accept)’.

  • 🎉🎉 I’ve been invited to be an Area Chair in NeurIPS 2025.

  • 🎉🎉 Congratulations to our team (lead by @Ruibo) to get “SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUs” accepted by EuroSys 2025 as Best Paper !!!

  • 🎉🎉 I am awarded the Excellent Research Prize for the 2024 DSA Excellent Research Award!!!

  • 🎉🎉 Our STBLLM is accepted by ICLR25. STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs, International Conference on Learning Representations, 2025.

  • 🎉🎉 Our Lottery LLM Hypothesis is accepted by ICLR25 Blogpost Oral. The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?, International Conference on Learning Representations Blog Track Oral, 2025.

  • 🎉🎉 Our ParZC is accepted by AAA25 (Oral). ParZC: Parametric Zero-Cost Proxies for Efficient NAS, Association for the Advancement of Artificial Intelligence, 2025.

  • 🎉🎉 I was invited to give a talk to PDL about “Introduction to LLM Compression and Beyond”.

  • 🎉🎉 FuseFL is accepted by NeurIPS 2024 (Spotlight). FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Layer Fusion, Neural Information Processing Systems (NeurIPS) Spotlight, 2024.

  • 🎉🎉 DSA is accepted by NeurIPS 2024, Discovering Sparsity Allocation for Layer-wise Pruning of Large Language Models, Neural Information Processing Systems (NeurIPS), 2024.

  • 🎉🎉 Our paper “Should we really edit language models? on the evaluation of edited language models” is accepted by NeurIPS 2024.

  • 🎉🎉 LPZero is accepted by EMNLP 2024. LPZero: Language Model Zero-cost Proxy Search from Zero, Empirical Methods in Natural Language Processing (EMNLP), 2024. (paper, code)

  • 🎉🎉 LongGenBench is accepted by EMNLP 2024. LongGenBench: Long-context Generation Benchmark, Empirical Methods in Natural Language Processing (EMNLP), 2024.

  • 🎉🎉 Pruner-Zero is accepted by ICML 2024. This work evolves symbolic pruning metrics from scratch for large language models. (paper, code)

  • 🎉🎉 VMRNN is available. This work proposes the VMRNN cell, a new recurrent unit that integrates the strengths of Vision Mamba blocks with LSTM. We construct a network centered on VMRNN cells to tackle spatiotemporal prediction tasks effectively. (paper, code)

  • 🎉🎉 KD-Zero is accepted by NeurIPS 2023. This work evolves knowledge distiller for any teacher-student pairs. (paper)

  • 🎉🎉 EMQ is accepted by ICCV 2023. This work evolves training-free proxies for automated mixed precision quantization. (paper, code)

  • 🎉🎉 AutoKD: Automated KD via MCTS is accepted by ICCV 2023. This work proposes automated knowledge distillation via Monte Carlo Tree Search. (paper)

  • 🎉🎉 DisWOT is accepted by CVPR 2023. This work proposes student architecture search for distillation without training. (paper, code)

  • 🎉🎉 Progressive Meta-Pooling Learning is accepted by ICASSP 2023. This work proposes a lightweight image classification model. (paper)

  • 🎉🎉 RD-NAS is accepted by ICASSP 2023. This work enhances one-shot supernet ranking ability via ranking distillation. (paper)

  • 🎉🎉 AutoRF is accepted by MMM 2022. This work proposes auto learning receptive fields with spatial pooling. (paper)

  • 🎉🎉 Prior-Guided One-shot NAS is accepted by CVPR Workshop 2022. This work proposes prior-guided one-shot neural architecture search. (paper)

📖 Educations

  • 2023.09 - now, The Hong Kong University of Science and Technology (Guangzhou), PhD Candidate in Computer Science

    • Supervisor: Prof. Xiaowen Chu
    • Research Interests: Large Language Models, Model Compression
  • 2020.09 - 2023.06, National University of Defence Technology, Master of Engineering

    • Supervisor: Prof. Xin Niu
    • Research Interests: AutoML, Neural Architecture Search
    • Achievement: Outstanding Graduate
  • 2016.09 - 2020.06, Northwest Agriculture & Forestry University, B.S. in Software Engineering

    • GPA: 3.78/4.0 (Ranked 1st out of 93)
    • Advisor: Prof. Hongming Zhang
    • Achievements: National Scholarship, Principal’s Scholarship, Outstanding Graduate
    • Research Interests: Object Detection, Multi-Object Tracking

💻 Internship

  • 06/2026-present: Research Intern, Tencent CSIG WorkBuddy/CodeBuddy - post-training and evaluation for long-horizon coding agents
  • 10/2025–02/2026: Intern, Alibaba – large-scale model training
  • 03/2025–08/2025: Intern, Shanghai AI Lab – AI infrastructure for Xtuner project
  • 05/2022–08/2022: Intern, Shanghai AI Lab – model compression with MMRazor

👔 Professional Activities

  • 2022: ICASSP
  • 2023: NeurIPS, ICASSP, CIM
  • 2024:
    • Conferences: NeurIPS, ICLR, CVPR, ECCV, ICASSP, ACL (ARR)
    • Journals: TPAMI, Neural Networks, Information Fusion, CIM
  • 2025:
    • Conferences: NeurIPS (AC), ICLR, CVPR, ECCV, ICASSP
    • Journals: IJCV, Neural Networks
  • 2026:
    • Conferences: AAAI (PC), WACV, NeurIPS, ICLR
    • Journals: Neural Networks

🎖 Honors and Awards

  • 2024, Best Speaker in DSA Salon 2024.
  • 2023, Outstanding Graduate at School Level, National University of Defense Technology.
  • 2022, 1st Place, BDCI Retail Product Recognition based on MindSpore (CCF Big Data & Computing Intelligence Contest).
  • 2022, 1st Place, DCIC Intelligent Ship Detection Competition (Digital China Innovation Contest).
  • 2022, 2nd Place, DCIC Intelligent Cattle Segmentation Competition (Digital China Innovation Contest).
  • 2022, 1st Place, Baidu AI Competition - Blurred Document Image Recovery.
  • 2022, 3rd Place, Computer Vision and Pattern Recognition (CVPR) Third Workshop on NAS.
  • 2021, Outstanding MindSpore Developer.
  • 2020, Outstanding Dissertation, Northwest A&F University.
  • 2020, Outstanding Graduate, Northwest A&F University.
  • 2017, President’s Scholarship, Northwest A&F University.
  • 2016, National Scholarship, Northwest A&F University.

📝 Publications

Conference papers, workshop papers, technical reports, and blog publications.

  • CPA: Efficient and Stable FP4 RL Training via Cross-Precision Alignment — paper figureCPA: Efficient and Stable FP4 RL Training via Cross-Precision AlignmentG. Gong, Y. Wei, Y. Tao, T. Wu, P. Dong, R. Fan, W. Hu, Y. Yu, J. Wang, W. Su, G. Yang, L. Zhang, W. Wang, X. Chu.NeurIPS 2026.

  • Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning — paper figureArchitecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math ReasoningK. Liu, P. Dong, X. Xie, J. Gao, Q. Guo, X. Chu, S. Zhang, K. Chen.In EMNLP 2026 Main Conference.

  • Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction — paper figureTencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task ConstructionTencent WorkBuddy Bench Team (including P. Dong).Technical report, 2026.

  • An Empirical Study of Reasoning Degradation in Quantized Multimodal Large Language Models — paper figureAn Empirical Study of Reasoning Degradation in Quantized Multimodal Large Language ModelsACM MM 2026.

  • GreenMoE: Exploiting Dynamic Load Imbalance for Energy-Efficient Long-Context MoE Training — paper figureGreenMoE: Exploiting Dynamic Load Imbalance for Energy-Efficient Long-Context MoE TrainingICML 2026 AdaptFM Workshop.

  • Parameters as Agentic Memory: Internalizing Long-Horizon Memories for Efficient LLM Agents — paper figureParameters as Agentic Memory: Internalizing Long-Horizon Memories for Efficient LLM AgentsZ. Tang, F. Wei, Z. Tang, P. Dong, X. Liu, Q. Wang, X. Chu, B. Li.ICML 2026 AIWILD Workshop.

  • Enhancing Knowledge Injection with Surrounding Backgrounds in Continual Training LLMs — paper figureEnhancing Knowledge Injection with Surrounding Backgrounds in Continual Training LLMsZ. Tang, Z. Tang, Y. Hou, P. Dong, X. Liu, S. Shi, X. Chu, B. Li.ICML 2026 FoGen Workshop.

  • Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression — paper figureSemantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache CompressionX. Liu, Z. Tang, H. Chen, P. Dong, Z. Li, X. Zhou, B. Li, X. Hu, X. Chu.In ICML 2026.

  • VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing — paper figureVCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and EditingX. Su, P. Dong, Z. Tang, S. Tang, Y. Zhai, K. Lin, L. Chen, Y. Gai, Y. Luo, Q. Wang, X. Chu.In ICML 2026.

  • Identifying and Mitigating Errors in Gradient Aggregation of Distributed Data Parallel Training — paper figureIdentifying and Mitigating Errors in Gradient Aggregation of Distributed Data Parallel TrainingZ. Tang, J. Huang, Z. Tang, X. Kang, Y. Wang, P. Dong, S. Shi, X. Chu, B. Li.In ICML 2026.

  • Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Understanding — paper figureSmooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context UnderstandingK. Liu, Z. Su, P. Dong, F. Mo, J. Gao, S. Zhang, K. Chen.ICLR 2026.

  • ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — paper figureChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM InferenceX. Liu, Z. Tang, P. Dong, Z. Li, Y. Liu, B. Li, X. Hu, X. Chu.NeurIPS 2025.

  • Perovskite-LLM: Knowledge-Enhanced Large Language Models for Perovskite Solar Cell Research — paper figurePerovskite-LLM: Knowledge-Enhanced Large Language Models for Perovskite Solar Cell ResearchX. Liu, P. Sun, S. Chen, L. Zhang, P. Dong, H. You, Y. Zhang, C. Yan, X. Chu, T.-Y. Zhang.EMNLP 2025 Findings.

  • Intern-S1: A Scientific Multimodal Foundation Model — paper figureIntern-S1: A Scientific Multimodal Foundation ModelIntern-S1 Team, Shanghai AI Laboratory (including P. Dong).Technical report, 2025.

  • Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression — paper figureCan Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM CompressionP. Dong, Z. Tang, X. Liu, L. Li, X. Chu, B. Li.In ICML2025.

  • SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUs — paper figureSpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUsR. Fan, X. Yu, P. Dong, Z. Li, G. Gong, Q. Wang, W. Wang, X. Chu.In EuroSys2025, Best Paper.

  • ParZC: Parametric Zero-Cost Proxies for Efficient NAS — paper figureParZC: Parametric Zero-Cost Proxies for Efficient NASP. Dong, L. Li, Z. Tang, X. Liu, Z. Wei, Q. Wang, X. Chu.In AAAI2025, Oral.

  • STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs — paper figureSTBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMsP. Dong, L. Li, Y. Zhong, D. Du, R. Fan, Y. Chen, Z. Tang, Q. Wang, W. Xue, Y. Guo, X. Chu.In ICLR2025.

  • The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve? — paper figureThe Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?Z. Tang, X. Liu, Q. Wang, P. Dong, B. He, X. Chu, B. Li.ICLR 2025 Blogpost, Oral.

  • Discovering Sparsity Allocation for Layer-wise Pruning of Large Language Models — paper figureDiscovering Sparsity Allocation for Layer-wise Pruning of Large Language ModelsL. Li, P. Dong, Z. Tang, X. Liu, X. Pan, X. Chu.In NeurIPS 2024.

  • VMRNN: Integrating Vision Mamba and LSTM for Efficient and Accurate Spatiotemporal Forecasting — paper figureVMRNN: Integrating Vision Mamba and LSTM for Efficient and Accurate Spatiotemporal ForecastingY. Tang, P. Dong, Z. Tang, X. Chu, J. Liang.Preprint, 2024.

  • Should We Really Edit Language Models? On the Evaluation of Edited Language Models — paper figureShould We Really Edit Language Models? On the Evaluation of Edited Language ModelsQ. Li, X. Liu, Z. Tang, P. Dong, Z. Li, X. Pan, X. Chu,In NeurIPS 2024.

  • Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language Models — paper figurePruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language ModelsP. Dong, L. Li, Z. Tang, X. Liu, X. Pan, Q. Wang, X. Chu.In ICML 2024.

  • LPZero: Language Model Zero-cost Proxy Search from Zero — paper figureLPZero: Language Model Zero-cost Proxy Search from ZeroP. Dong, L. Li, X. Liu, Z. Tang, X. Liu, Q. Wang, X. Chu.Empirical Methods in Natural Language Processing (EMNLP), 2024.

  • LongGenBench: Long-context Generation Benchmark — paper figureLongGenBench: Long-context Generation BenchmarkX. Liu, P. Dong, X. Hu, X. Chu.In EMNLP 2024.

  • FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Layer Fusion — paper figureFuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Layer FusionZ. Tang, Y. Zhang, P. Dong, Y. Cheung, A. C. Zhou, B. Han, X. Chu.In NeurIPS Spotlight 2024.

  • DisWOT: Student Architecture Search for Distillation without Training — paper figureDisWOT: Student Architecture Search for Distillation without TrainingP. Dong, L. Li, Z. Wei.In CVPR 2023.

  • EMQ: Evolving Training-free Proxies for Automated Mixed Precision Quantization — paper figureEMQ: Evolving Training-free Proxies for Automated Mixed Precision QuantizationP. Dong, L. Li, Z. Wei, X. Niu$^*$, Z. Tian, H. Pan.In ICCV 2023.

  • Kd-zero: Evolving knowledge distiller for any teacher-student pairs — paper figureKd-zero: Evolving knowledge distiller for any teacher-student pairsL. Li, P. Dong, A. Li, Z. Wei, Y. Yang.In NeurIPS 2023.

  • Progressive Meta-Pooling Learning for Lightweight Image Classification Model — paper figureProgressive Meta-Pooling Learning for Lightweight Image Classification ModelP. Dong, X. Niu, Z. Tian, et al.In ICASSP 2023.

  • RD-NAS: Enhancing One-shot Supernet Ranking Ability via Ranking Distillation — paper figureRD-NAS: Enhancing One-shot Supernet Ranking Ability via Ranking DistillationP. Dong, X. Niu, L. Li, et al.In ICASSP 2023.

  • AutoRF: Auto Learning Receptive Fields with Spatial Pooling — paper figureAutoRF: Auto Learning Receptive Fields with Spatial PoolingP. Dong, X. Niu, H. Pan, et al.In MMM 2023.

  • Prior-Guided One-shot Neural Architecture Search — paper figurePrior-Guided One-shot Neural Architecture SearchP. Dong, X. Niu, L. Li, et al.In CVPR Workshop 2022.

  • Automated Knowledge Distillation via Monte Carlo Tree Search — paper figureAutomated Knowledge Distillation via Monte Carlo Tree SearchL. Li, P. Dong, Z. Wei, Y. Ya.In ICCV 2023.