The task

Reach the goal you are given, g, having learned only from a fixed dataset of past trajectories. Nothing new is collected.

The method being reproduced

HIQL splits it in two: a high-level policy picks a nearby subgoal z, k steps ahead; a low-level policy picks the action toward z. Both read one goal-conditioned value function.

Experiment 01 · how one trajectory becomes training data · k = 2

From s0: goals s1…s5 · subgoal s2 · action 1

Value · (s, g)
15 samples
action-free data: yes
High-level · (s, g, z)
10 samples
action-free data: yes
Low-level · (s, z, a)
4 samples
action-free data: no — it needs the actions
Counts printed by my own check script (HIQL/train.py) on a six-state toy trajectory. A check of the method, not a result: no benchmark numbers yet.
← All work

Research · in progress

Offline goal-conditioned RL

Learning to reach a goal from a fixed dataset, in NTU’s Embodied AI Lab.

My part
Research exploration: reproducing baselines
Time
2025 — present · In progress
Affiliation
Embodied AI Lab · National Taiwan University
Kind
Research · Offline RL
Tags
Offline RL · Goal-conditioned RL · HIQL · OGBench

My part

Research exploration at National Taiwan University: reproducing offline reinforcement learning baselines, with a focus on goal-conditioned RL and decision-making from fixed datasets.

  • Reproducing offline reinforcement learning baselines for goal-conditioned tasks
  • Reading and implementing components from offline RL and contrastive RL papers
  • Dataset construction, policy learning and evaluation pipelines
  • Working from the authors’ code for HIQL (Park et al., 2023) and OGBench (Park et al., 2024), and a check script of my own (the experiment above)

Reading alongside: IQL, contrastive RL, OPAL, C-Planning, and a tutorial on offline RL.

The setup

Can an agent learn to reach a goal it is given using only a fixed dataset of past experience, without collecting anything new?

Offline goal-conditioned RL, as a diagramA fixed dataset of past transitions feeds policy learning; the goal-conditioned policy is then evaluated on reaching a given goal, with no new data collected.(s, a, s′)01 · Fixed datasetdataset constructionπ(a | s, g)02 · Goal-conditioned policypolicy learningg03 · Evaluationreach the given goalNo new interaction: everything is learned from data collected beforehand.

Where it stands

In progress

Reproduction is under way. There are no benchmark results yet, so none are shown; they will be, with their figures, once there is something to report.

An illustration

Goal-conditioned, in one picture: name a goal and the arm carries a block there. It is a sketch modelled for this site, scripted rather than learned — not the lab’s robot and not a result.

Illustration: a small robot arm stacking wooden blocks in a dark lab.Illustration · not a result

Interactive 3D sketch · point to aim, click to pick up or put down · drag to look around