Trajectory Unlearning

GiRPO: Trajectory Unlearning on LLM-based Agents

Forgetting what an agent does, not just what it says.

Yingdan Shi1 Ren Wang1,*
1Illinois Institute of Technology

* Corresponding Author

Trajectory-Level Unlearning Group-injected Relative Policy Optimization ALFWorld & WebShop Benchmarks

Motivation

From knowledge unlearning to trajectory unlearning

The difference between knowledge unlearning and trajectory unlearning. Existing knowledge unlearning methods can prevent LLMs or tool-augmented agents from producing undesirable knowledge, but do not necessarily prevent agents from reproducing the corresponding task trajectories. A single task may admit multiple trajectories, and a trajectory may be either a complete trajectory that successfully accomplishes the task or an incomplete trajectory that fails to do so.

Overview

Beyond knowledge: unlearning what an agent does

LLM unlearning has focused on suppressing facts an agent knows. As LLMs act as autonomous agents, they must also stop reproducing undesired behaviors through their action trajectories.

Abstract

Existing large language model (LLM) unlearning has focused primarily on removing specific knowledge, such as harmful facts, private data, or copyrighted content. However, as LLMs are increasingly deployed as autonomous agents, a fundamental yet overlooked problem emerges: beyond suppressing what an agent knows, an agent should not reproduce undesired behaviors through its action trajectories. In this work, we introduce trajectory-level unlearning, a new problem formulation that targets the removal of specific action trajectories in long-horizon agentic tasks, rather than factual knowledge.

We identify two fundamental challenges that distinguish trajectory unlearning from knowledge unlearning: (1) our unlearning target is what the agent does, not what it says; and (2) trajectories are sequentially dependent action sequences that cannot be decomposed into isolated prompt-response pairs without losing inter-step structure. To address these challenges, we propose Group-injected Relative Policy Optimization (GiRPO), which injects forget trajectories into the policy rollout group with penalized rewards and isolates the normalization statistics, yielding a stable and bounded unlearning signal that does not corrupt gradient updates for normal task trajectories.

We construct trajectory unlearning benchmarks from two application scenarios, household tasks (ALFWorld) and online shopping (WebShop), and design three complementary metrics for evaluating forgetting quality and model utility. Experiments on ALFWorld and WebShop demonstrate that GiRPO effectively unlearns target trajectories while preserving task success rates, outperforming existing knowledge-unlearning baselines on both forgetting quality and task utility.

①

From Responses to Behaviors

Knowledge unlearning focuses on what the model says, whereas trajectory unlearning concerns what the agent actually does. This difference calls for different optimization objectives.

②

Long-Horizon Dependencies

A trajectory consists of a sequence of actions that depend on the interaction history and on previous actions. In contrast, text-based unlearning methods typically treat trajectories as a collection of independent prompt-response pairs, thereby overlooking the dependencies between steps.

Method

GiRPO: Group-injected Relative Policy Optimization

Forget trajectories ride alongside real rollouts inside the same policy-gradient update, but with a penalized reward and normalization statistics isolated from the normal-task group.

  1. 1
    Pseudo RolloutInject the forget trajectory into the policy's rollout group as an additional sampled rollout.
  2. 2
    Reward PenaltyAssign the injected trajectory a penalized reward that discourages the policy from reproducing it.
  3. 3
    Isolated NormalizationCompute its advantage using group statistics kept separate from the normal-task rollouts.
  4. 4
    Lower ClippingClip the resulting advantage at a lower bound so the unlearning signal stays bounded and stable.
GiRPO at a glance. Overview of our trajectory unlearning framework. Left: Construction of the forget trajectories 𝒟f, where each target task has at least one forget trajectory. Middle: GiRPO training pipeline, which injects forget trajectories into the rollout group with penalized rewards and isolated advantage estimation. Right: Unlearning scenarios spanning household tasks (ALFWorld) and web shopping (WebShop), together with our evaluation framework based on three complementary metrics.

Results

Unlearning the trajectory, keeping the skill

On both ALFWorld and WebShop, GiRPO drives down exact-match and LLM-as-judge similarity to the forgotten trajectory while matching or exceeding the base model's success rate on target and untargeted tasks.

87.4% Target-task success (avg.) vs. 87.1% base model
−20.9 pp Exact match (avg.) similarity to the forgotten trajectory
−15.3 pp LLM-as-judge (avg.) judged reproduction of the forgotten trajectory
+0.3 pp Untargeted-task success (avg.) vs. base model, essentially unaffected

Averaged over the four ALFWorld unlearning settings (Clean, Heat, Cool, Mixed) reported below; WebShop results follow.

ALFWorld trajectory unlearning results. Across Clean, Heat, Cool, and Mixed settings, GiRPO consistently lowers exact match and LLM-as-judge on target tasks while keeping target- and untargeted-task success rate close to the base model, outperforming GA, DPO, NPO, GRPO, and NPO+GRPO baselines on this trade-off.
WebShop trajectory unlearning results. GiRPO raises target-task Score and Success Rate above the base model (81.4% and 76.2%, versus 76.9% and 68.1%) while cutting Exact Match and LLM-as-judge similarity to the forgotten trajectory (down to 3.9% and 79.5%), and also improves untargeted-task Score and Success Rate — outperforming GA, DPO, NPO, GRPO, and NPO+GRPO on every column.

Analysis

Effect of Forget Trajectory's Completeness

In the forget trajectory dataset for ‘Clean’ target tasks, there is one forget trajectory for each of the 650 tasks. Among these, 595 forget trajectories successfully complete the corresponding task (complete) and 55 do not (incomplete). We analyze the effect of trajectory completeness on unlearning quality and task utility separately.

Forgetting outcomes by trajectory type. Of 595 forget trajectories that originally completed the task, GiRPO forgets 586 (98.5%); all 55 originally-incomplete forget trajectories are forgotten as well.
Task success rate before/after unlearning. Success on tasks tied to previously-complete forget trajectories declines only modestly (100.0% → 92.9%), while success on tasks tied to previously-incomplete forget trajectories rises (0.0% → 40.0%).
A

Complete trajectories are harder to unlearn

Incomplete trajectories correspond to behaviors the model already struggles to reproduce, whereas complete trajectories represent well-learned, high-reward behavioral patterns that are deeply embedded in the policy.

B

Forgetting a failure mode can help

Unlearning incomplete trajectories improves task success rate from 0% to 40%, since suppressing a failed action sequence implicitly steers the model away from suboptimal behaviors and toward successful alternatives.

Citation

Resources and citation

The paper and code are available now.

BibTeX
@article{shi2026traj,
  title   = {Trajectory Unlearning on LLM-based Agents},
  author  = {Shi, Yingdan and Wang, Ren},
  year    = {2026},
  journal = {arXiv preprint arXiv:2609.33639}
}