Chen Junyang陈骏扬

Multimodal AI · Agentic RL · Environment-Grounded Reasoning多模态智能 · 智能体强化学习 · 环境约束推理

Graduate student, PALM Lab, School of Computer Science and Engineering, Southeast University · Nanjing, China东南大学计算机科学与工程学院 · PALM 实验室 · 中国南京

chenjunyang0806@gmail.com · chenjunyang@seu.edu.cn · Citations 236

My research studies AI systems that learn from imperfect supervision, reason over multimodal evidence, and make decisions in constrained interactive environments. I am particularly interested in RL distillation, multimodal reasoning, open-vocabulary perception, semi-supervised learning, and environment-grounded agent behavior.

我的研究关注如何让智能系统从不完美监督中学习,结合多模态证据进行推理,并在受约束的交互环境中完成决策。近期兴趣包括强化学习蒸馏多模态推理开放词汇感知、半监督学习与环境约束下的智能体行为

RL Distillation Multimodal Reasoning Open-Vocabulary Perception Grounded Agent Learning
Education
Master's DegreeSoutheast University (SEU), 2024.09–present
Bachelor's DegreeSichuan Agricultural University (SiCAU), 2020.09–2024.06
2026.03–present

Research Intern on Agent Post-Training & Interactive Reinforcement Learning

Current Internship

Working on application-level post-training and Agentic RL for Qwen3.5-122B-A10B and Qwen3.5-35B-A3B, with a focus on multi-turn decision-making and tool-mediated interaction for consumer-facing AI agents across travel, mobility, local services, instant commerce, and other real business scenarios.

Executable environments Model user goals, business rules, environment states, Agent SOPs, and Tool Schemas as interactive training environments.
Interactive agent trajectories Build trajectories for state tracking, tool selection, parameter completion, constraint following, and exception recovery.
Evaluation & data flywheel Support SFT cold start, Agentic RL training, offline evaluation, process-level reward design, bad-case attribution, and iterative data refinement.
§

News近期动态

  • Paper

    Positive-Unlabeled Reinforcement Learning Distillation for On-Premise Small Models accepted to ICML 2026.

  • Paper

    LoGoSeg accepted to AAAI 2026.

  • Paper

    RankMatch accepted to NeurIPS 2025.

  • Open

    Maintaining Arxiv-tracker for lightweight paper monitoring, filtering, and research queue management.

  • Lead

    Led the InternLM/Lagent teaching module and delivered a public Bilibili lecture.

📚

Selected Publications代表性论文

Full list of publications in Google Scholar. * Equal contribution.

🧩

Open Source Leadership开源主导与教学

ShuSheng LLM Practical Camp banner

ShuSheng LLM Practical Camp · Agent Module Lead

InternLM/Tutorial · course design · public lecture

2k

Led the Agent/Lagent teaching module for the InternLM open-source camp: designed hands-on course materials, structured runnable multi-agent examples, and delivered a public Bilibili lecture with 3k+ views.

Arxiv-tracker web preview

ArXiv Daily Paper Tracker · Project Lead

colorfulandcjy0806/Arxiv-tracker · daily digest · web + email

50

Built and maintain an automated research workflow for keyword/category search, bilingual LLM summaries, email digests, static site generation, freshness-window deduplication, OpenAI-compatible APIs, and daily GitHub Action automation.

🏆

Reviewer Experience & Awards审稿经历与荣誉

Reviewer Experience

  • Reviewer for Neural Information Processing Systems (NeurIPS 2026)
  • Reviewer for ACM Multimedia Conference (ACM MM 2025–2026)
  • Reviewer for AAAI Conference on Artificial Intelligence (AAAI 2025–2027)
  • Reviewer for Frontiers in Plant Science
  • Reviewer for International Conference on Machine Learning (ICML 2026)

Awards

  • Excellent Postgraduate Student Cadre, Southeast University (东南大学优秀研究生干部, Top 5%), 2025.10
  • Outstanding Graduates of Sichuan Province (四川省优秀毕业生, Top 4%), China, 2024.06
  • Outstanding Student Model (优秀学生标兵, the highest honor of the university, only 10 students selected, sole recipient in the entire college), Sichuan Agricultural University, 2022–2023
  • China National Scholarship (国家奖学金, Top 0.2%), 2023
  • China National Scholarship (国家奖学金, Top 0.2%), 2022