01
Agentic RL
Training LLM agents with reinforcement learning over long horizons — credit assignment, reward design, and learning from experience.
Research Scientist @ ByteDance Seed · Singapore
I build language-model agents that reason, prove, and act — trained with reinforcement learning at scale.
Currently working on agentic RL and formal theorem proving in Lean 4 as part of the Seed-Prover team. Previously at Salesforce AI Research and ByteDance AI Lab; Ph.D. from SUTD.
A Lean 4 snippet proving the playful theorem allan_jie: there exists a language model that reasons, proves and acts — discharged by ReFT, Seed-Prover and agentic reinforcement learning.
01/research
01
Training LLM agents with reinforcement learning over long horizons — credit assignment, reward design, and learning from experience.
02
Lean 4 provers that search deep and broad at test time and improve through large-scale RL — the Seed-Prover line of work.
03
From deductive math-word-problem solvers to reinforced fine-tuning (ReFT): making multi-step reasoning more reliable and more explainable.
04
Cheaper long-context serving, e.g. query-driven pruning of the KV cache (ThinK).
02/highlights
ByteDance Seed · 2025
A Lean 4 prover trained with large-scale agentic RL that keeps accumulating experience from Lean and other tools, plus a test-time workflow bridging natural-language and formal proofs.
ByteDance Seed · IMO 2025
Lemma-style whole-proof reasoning that iteratively refines proofs from Lean feedback, with deep and broad test-time search for olympiad-level problems.
Salesforce AI Research · ICLR 2025
The KV cache is redundant along the channel dimension too. ThinK prunes the least important key channels per query, cutting long-context memory while keeping accuracy.
ByteDance Research · ACL 2024
Reinforced fine-tuning: warm up with SFT, then run PPO over many sampled chain-of-thought paths, rewarded by answer correctness — generalising better than SFT alone on math reasoning.
03/papers
04/log
Measure Twice, Locate Once — mitigating hallucinations in LLM agents for repository-scale fault localization — published in ACM TOSEM.
The Seed2.0 Model Card is out; I’m one of the contributors.
Seed-Prover 1.5: agentic RL + test-time scaling solves 88% of PutnamBench and 11/12 Putnam 2025 problems within 9 hours.
Seed-Prover takes part in IMO 2025: an IMO-certified 30/42, silver-medal level. Papers: Seed-Prover and Delta Prover.
Joined ByteDance Seed in Singapore to work on LLM agents and reinforcement learning.
ThinK (query-driven KV-cache pruning) is a Spotlight at ICLR 2025 in Singapore.
ReFT: Reasoning with Reinforced Fine-Tuning appears at ACL 2024, alongside a Findings paper on non-autoregressive MT as a constrained HMM.
Joined Salesforce AI Research, Singapore.
Leveraging Training Data in Few-Shot Prompting for Numerical Reasoning accepted to Findings of ACL 2023.
Serving as a Senior Program Committee member for AAAI 2023.
Talk at the workshop on “A Science of Certified AI” (slides).
Learning to Reason Deductively accepted to ACL 2022; talk on math word problem solving at SMT, SUTD (slides).
Ph.D. from the StatNLP group at SUTD — received the Best Thesis Award (thesis & slides).
05/trajectory
6b801fe (HEAD → main) 2025 — now
LLM agents and reinforcement learning. Core contributor to Seed-Prover — Lean 4 theorem-proving agents trained with large-scale agentic RL (IMO 2025 silver-level score; 88% of PutnamBench).
15cf2b7 2024 — 2025
Efficient long-context inference for LLMs — ThinK, query-driven KV-cache pruning (ICLR 2025 Spotlight).
63ee5f5 2020 — 2024
Reasoning with language models: ReFT (reinforced fine-tuning, ACL 2024), chain-of-thought design, few-shot prompting for numerical reasoning, and visual document understanding.
1f9fee5 2019 — 2020
Knowledge-graph-to-text generation (ENT-DESC, EMNLP 2020) and named entity recognition.
022f388 2019
Hosted by Pradeep Dasigi and Ana Marasović.
1e3af2c (tag: phd) 2016 — 2020
StatNLP group, advised by Prof. Wei Lu. Thesis: “Leveraging Dependency Trees for Structured Prediction”. ★ Best Thesis Award