Research Scientist @ ByteDance Seed · Singapore

Zhanming (Allan) Jie

I build language-model agents that reason, prove, and act — trained with reinforcement learning at scale.

Currently working on agentic RL and formal theorem proving in Lean 4 as part of the Seed-Prover team. Previously at Salesforce AI Research and ByteDance AI Lab; Ph.D. from SUTD.

citations
1.7k+
h-index
17
papers
25
since
2014
lean AllanJie.lean
Lean 4 · Mathlib Ln 10, Col 33

A Lean 4 snippet proving the playful theorem allan_jie: there exists a language model that reasons, proves and acts — discharged by ReFT, Seed-Prover and agentic reinforcement learning.

01/research

What I work on

01

Agentic RL

Training LLM agents with reinforcement learning over long horizons — credit assignment, reward design, and learning from experience.

long-horizon RLreward designexperience

02

Formal theorem proving

Lean 4 provers that search deep and broad at test time and improve through large-scale RL — the Seed-Prover line of work.

Lean 4test-time scalingSeed-Prover

03

Reasoning in language models

From deductive math-word-problem solvers to reinforced fine-tuning (ReFT): making multi-step reasoning more reliable and more explainable.

chain-of-thoughtReFTmath

04

Efficient inference

Cheaper long-context serving, e.g. query-driven pruning of the KV cache (ThinK).

KV cachelong context

02/highlights

Selected work

all publications →

ByteDance Seed · 2025

Seed-Prover 1.5

A Lean 4 prover trained with large-scale agentic RL that keeps accumulating experience from Lean and other tools, plus a test-time workflow bridging natural-language and formal proofs.

88%
PutnamBench
11/12
Putnam 2025, ≤ 9 h
80%
Fate-H (graduate)

ByteDance Seed · IMO 2025

Seed-Prover

Lemma-style whole-proof reasoning that iteratively refines proofs from Lean feedback, with deep and broad test-time search for olympiad-level problems.

30/42
IMO 2025 · silver level
78.1%
past IMO problems
99.6%
miniF2F-test

Salesforce AI Research · ICLR 2025

ThinK

The KV cache is redundant along the channel dimension too. ThinK prunes the least important key channels per query, cutting long-context memory while keeping accuracy.

>20%
KV-cache memory saved
Spotlight
ICLR 2025

ByteDance Research · ACL 2024

ReFT

Reinforced fine-tuning: warm up with SFT, then run PPO over many sampled chain-of-thought paths, rewarded by answer correctness — generalising better than SFT alone on math reasoning.

400+
citations
ACL
2024 main

03/papers

Recent papers

Tech report 2025 ★ IMO 2025 · silver-level

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving

Luoxin Chen, Jinming Gu, …, Allan Jie, …, Thomas Hanwen Zhu

Browse all 25 papers

04/log

News

  1. paper

    Measure Twice, Locate Once — mitigating hallucinations in LLM agents for repository-scale fault localization — published in ACM TOSEM.

  2. release

    The Seed2.0 Model Card is out; I’m one of the contributors.

  3. release

    Seed-Prover 1.5: agentic RL + test-time scaling solves 88% of PutnamBench and 11/12 Putnam 2025 problems within 9 hours.

  4. release

    Seed-Prover takes part in IMO 2025: an IMO-certified 30/42, silver-medal level. Papers: Seed-Prover and Delta Prover.

  5. career

    Joined ByteDance Seed in Singapore to work on LLM agents and reinforcement learning.

  6. paper

    ThinK (query-driven KV-cache pruning) is a Spotlight at ICLR 2025 in Singapore.

  7. paper

    ReFT: Reasoning with Reinforced Fine-Tuning appears at ACL 2024, alongside a Findings paper on non-autoregressive MT as a constrained HMM.

git log --all · 6 older entries
  1. career

    Joined Salesforce AI Research, Singapore.

  2. paper

    Leveraging Training Data in Few-Shot Prompting for Numerical Reasoning accepted to Findings of ACL 2023.

  3. service

    Serving as a Senior Program Committee member for AAAI 2023.

  4. talk

    Talk at the workshop on “A Science of Certified AI” (slides).

  5. paper

    Learning to Reason Deductively accepted to ACL 2022; talk on math word problem solving at SMT, SUTD (slides).

  6. award

    Ph.D. from the StatNLP group at SUTD — received the Best Thesis Award (thesis & slides).

05/trajectory

Where I’ve been

full CV →
  1. 6b801fe (HEAD → main) 2025 — now

    Research Scientist @ ByteDance Seed Singapore

    LLM agents and reinforcement learning. Core contributor to Seed-Prover — Lean 4 theorem-proving agents trained with large-scale agentic RL (IMO 2025 silver-level score; 88% of PutnamBench).

  2. 15cf2b7 2024 — 2025

    Research Scientist @ Salesforce AI Research Singapore

    Efficient long-context inference for LLMs — ThinK, query-driven KV-cache pruning (ICLR 2025 Spotlight).

  3. 63ee5f5 2020 — 2024

    NLP Scientist @ ByteDance AI Lab / ByteDance Research Singapore

    Reasoning with language models: ReFT (reinforced fine-tuning, ACL 2024), chain-of-thought design, few-shot prompting for numerical reasoning, and visual document understanding.

  4. 1f9fee5 2019 — 2020

    Research Intern @ Alibaba Singapore

    Knowledge-graph-to-text generation (ENT-DESC, EMNLP 2020) and named entity recognition.

  5. 022f388 2019

    Research Intern @ Allen Institute for AI (AI2) Seattle

    Hosted by Pradeep Dasigi and Ana Marasović.

  6. 1e3af2c (tag: phd) 2016 — 2020

    Ph.D., Computer Science @ Singapore University of Technology and Design (SUTD)

    StatNLP group, advised by Prof. Wei Lu. Thesis: “Leveraging Dependency Trees for Structured Prediction”. ★ Best Thesis Award

esc
↑↓ navigate ↵ open