Complete reward definitions, observation processing, training settings, and baseline configurations.
This supplement documents the implementation details needed to reproduce the training setup: all 25 reward terms, observation processing, action scaling, curriculum settings, and the E2E-RL baseline configuration. The paper presents the method, GAC derivation, frame construction, symmetry losses, and curriculum schedule.
The following tables define the 25 reward terms listed in Table I of the paper, using the same names and order. The notation also follows the paper: \(i \in \{l, r\}\) indexes arms or feet; \(q\) denotes joint positions; \(v_{xy}^{P}\) and \(\omega_{z}^{P}\) denote planar velocity and yaw rate in the body frame; and \(c = (v_{xy,d}^{P}, \omega_{z,d}^{P})\) is the locomotion command. The mean policy action is \(\mu_t = \mu_\theta(o_t)\), and the per-leg gait phase is \(\phi \in [-\pi, \pi]^2\). We use \(\mathbb{1}[\cdot]\) for the indicator function and \((\cdot)_+\) for the positive part:
A penalty based on \((x)_+\) applies only when \(x > 0\). The joint-position limit penalty uses the linear excess beyond a soft bound. The foot contact force and pre-touchdown velocity penalties use squared excess, \((x)_+^2\), to penalize forces or descent speeds above their thresholds.
Orientation rewards and penalties use projected gravity: the world gravity direction expressed in a body's local frame \(\{X\}\):
The horizontal component \(g^{X}_{xy}\) is zero when the frame's \(z\)-axis is vertical. Its magnitude, \(\lVert g^{X}_{xy}\rVert\), provides a measure of tilt without a reference orientation or an orientation singularity. We use this measure for the pelvis \(g^{P}\), torso \(g^{T}\), and feet \(g^{F_i}\). For the locomotion penalties below, \(\hat c = v_{xy,d}^{P}/\lVert v_{xy,d}^{P}\rVert\) denotes the unit commanded heading, and \(\psi\) denotes base yaw.
| Term | \(w_i\) | Definition |
|---|---|---|
| EEF position | 4.0 | \(\exp\!\left(-\dfrac{e_{2,t}}{\sigma_2} - \dfrac{e_{1,t}}{\sqrt{3}\,\sigma_1}\right)\). Stages A–D use the \(L_2^2\)-only form \(\exp(-e_{2,t}/\sigma_P)\). The baseline disables the \(L_1\) term in stages E–F by setting \(\sigma_1 \le 0\) (§6). |
| EEF orientation | 3.0 | \(\exp\!\left(-\operatorname{mean}_i\operatorname{tr}\left(I_3 - R_{DE_i,t}\right)/\sigma_R\right)\). Both poses are taken in the manipulation frame; the target rotation is reconstructed from the 6D representation carried in the command. |
| Term | \(w_i\) | Definition |
|---|---|---|
| Linear velocity | 2.0 | \(\exp\!\left(-\lVert v_{xy,d}^{P} - v_{xy}^{P}\rVert^2/\sigma_v\right)\), with \(v_{xy}\) the root linear velocity rotated into the base frame. |
| Angular velocity | 1.5 | \(\exp\!\left(-(\omega_{z,d}^{P} - \omega_{z}^{P})^2/\sigma_\omega\right)\), \(\sigma_\omega = 0.25\) throughout the curriculum. |
| Swing-foot trajectory | 5.0 | \(\exp\!\left(-\left[(z_l - z_l^{*})^2 + (z_r - z_r^{*})^2\right]/\sigma\right)\), \(\sigma = 0.008\), foot heights measured relative to the terrain. The reference \(z^{*}\) is defined below. |
| Pelvis orientation | 1.0 | \(\exp\!\left(-\lVert g^{P}_{xy}\rVert^2/\sigma\right)\), the projected gravity of the pelvis frame. |
| Torso orientation | 1.0 | \(\exp\!\left(-\lVert g^{T}_{xy}\rVert^2/\sigma\right)\) — the torso's projected gravity measures waist lean separately from pelvis tilt. |
| Base height | 1.0 | \(\exp\!\left(-(100\,(z^P - z^{P}_d))^2/\sigma\right)\), \(\sigma = 3.0\). The height error is converted to centimeters before squaring, giving a kernel width of \(3\ \mathrm{cm}^2\). The commanded height \(z_d^P\) varies with the curriculum. |
| Term | \(w_i\) | Definition |
|---|---|---|
| Arm residual | −0.05 | \(\lVert a_{\mathrm{RL,arm},t}\rVert^2\) uses the buffered residual after clipping and scaling. In terms of the raw policy output, this is \(s^2\lVert\operatorname{clip}(a_{\mathrm{RL,arm}},\pm c_r)\rVert^2\) with \(s = 3\). Scaling therefore multiplies the squared clipped residual by 9. |
| EEF body twist | −0.05 | \(\sum_i \lVert V^{E}_{RE,i}\rVert^2\) sums the ground-truth body-twist penalty over both arms. The twist comes from the controller when available and from a rigid-body computation otherwise. |
| Arm action rate | −0.5 | \(\lVert \Delta a_t\rVert^2\) over the 14 arm channels of the raw action. |
| Non-arm action rate | −1.0 | The same action-rate penalty over the 15 non-arm channels: 12 leg joints and 3 waist joints. |
| Mean-action accel. | −0.10 | \(\lVert \Delta^2\mu_t\rVert^2 = \lVert \mu_t - 2\mu_{t-1} + \mu_{t-2}\rVert^2\). Applied as two equally weighted terms over the 14 arm and 15 non-arm channels. Using the deterministic policy mean \(\mu\) prevents exploration noise in sampled actions from dominating the penalty. |
| Arm joint velocity | −0.05 | \(\sum_{j\in\text{arm}} \dot q_j^2\), using measured joint velocities. |
| Pose deviation | −1.0 | \(\sum_j w_j (q_j - q_j^{\text{def}})^2\). Only leg and waist joints contribute; arm weights are zero. Per-joint weights: hip pitch 0.01, hip roll 1.0, hip yaw 5.0, knee 0.01, ankle pitch and roll 5.0 (each leg); waist yaw, roll, pitch 50.0. |
| Joint-position limit | −1.0 | \(\sum_j \left[(q^{\text{lo}}_j - q_j)_+ + (q_j - q^{\text{hi}}_j)_+\right]\) penalizes the linear excess beyond the soft limits. These limits shrink the hardware range about its midpoint, \(q^{\text{lo/hi}} = m \mp 0.5\,\rho\,r\), with \(\rho = 0.95\). |
| Term | \(w_i\) | Definition |
|---|---|---|
| Foot proximity | −10 | The indicator \(\mathbb{1}[d < 0.15\ \mathrm{m}]\), where \(d = \lvert\cos\psi\,\Delta y - \sin\psi\,\Delta x\rvert\) is the foot separation projected onto the base lateral axis. The penalty is binary and depends on lateral separation, so it discourages foot crossing. |
| Foot orientation | −5 | \(\sum_i \lVert g^{F_i}_{xy}\rVert\) sums the projected-gravity norms of the two feet. Using the norm keeps the penalty sensitive to small tilts. |
| Alive / termination | 10 / −200 | The alive term is a constant \(+1\) per step. The termination term is \(\mathbb{1}[\text{reset} \wedge \neg\text{timeout}]\), applied once on a failure reset. A 1000-step timeout incurs no termination penalty. |
These penalties are active only for the commands specified in the Mask column. The masks use the commanded planar velocity \(v_{xy,d}^{P} = (v_{x,d}^{P}, v_{y,d}^{P})\) and yaw rate \(\omega_{z,d}^{P}\) from \(c_t\). The penalties use measured velocity \(v_{xy}^{P}\) and yaw rate \(\omega_{z}^{P}\). Velocity and yaw-rate thresholds are in m/s and rad/s, respectively.
| Term | \(w_i\) | Mask | Penalized quantity |
|---|---|---|---|
| Heading drift | −10 | \(\lVert v_{xy,d}^{P}\rVert > 0.05 \;\wedge\; \lvert\omega_{z,d}^{P}\rvert < 0.03\) | \(\left(\bar\omega_{z}^{P}\right)^2 + 0.15\left(\omega_{z}^{P}\right)^2\) |
| Lateral drift | −2 | \(\lVert v_{xy,d}^{P}\rVert > 0.03\) | \(\left(\bar v_\perp\right)^2 + 0.15\left(v_\perp\right)^2\), with the signed \(v_\perp = \hat c_x v_{y}^{P} - \hat c_y v_{x}^{P}\) |
| Turn-in-place drift | −2.5 | \(\lVert v_{xy,d}^{P}\rVert < 0.03 \;\wedge\; \lvert\omega_{z,d}^{P}\rvert > 0.1\) | \(\lVert v_{xy}^{P}\rVert^2\) |
| Standing yaw rate | −2.5 | \(\lVert v_{xy,d}^{P}\rVert < 0.03 \;\wedge\; \lvert\omega_{z,d}^{P}\rvert < 0.03\) | \(\left(\omega_{z}^{P}\right)^2\) |
Gait-cycle averaging. An overbar \(\bar\cdot\) denotes the running mean \(\bar e_k\) within the current gait cycle. A command-manager term updates this mean while the gait phase advances; standing adds no samples. The accumulator resets when the left-leg phase crosses zero in the positive direction. During the first 5 samples of each cycle, the previous cycle's mean is used instead. The mean penalizes net drift, while the instantaneous term, weighted by \(0.15\times\), provides a reward signal at each step.
| Term | \(w_i\) | Definition |
|---|---|---|
| Foot contact force | −2 | \(\sum_i \left(\dfrac{(F^{\text{peak}}_{z,i} - F_{\text{th}})_+}{F_{\text{th}}}\right)^2\), \(F_{\text{th}} = 400\ \mathrm{N}\) is a force threshold. \(F^{\text{peak}}_z\) is the maximum over all physics substeps in the control interval. A contact-force history buffer captures these values, including brief impacts between control ticks. The implementation raises an error if the buffer is shorter than the control decimation. |
| Pre-touchdown velocity | −2 | \(\sum_i \mathbb{1}[\text{active}_i]\left((-\dot z_i) - 0.15\right)_+^2\), where \(\text{active}_i\) is true when all three conditions hold: the foot is airborne (\(F_z < 5\ \mathrm{N}\)), the foot is near the ground (clearance \(< 0.05\ \mathrm{m}\)), and the velocity command is non-trivial (\(> 0.03\ \mathrm{m/s}\)). The penalty does not apply during mid-swing descent. |
Swing-foot reference. The target height \(z^{*}\) follows a two-piece cubic Bézier curve over normalized phase \(x = (\phi + \pi)/2\pi\). With \(b(x) = x^3 + 3x^2(1-x)\), the foot rises from \(0 \to h\) using \(b(2x)\) for \(x \le 0.5\), then falls from \(h \to 0\) using \(b(2x-1)\). The peak height is \(h = 0.09\ \mathrm{m}\). The curve has zero slope at liftoff and touchdown, giving zero commanded vertical velocity at both endpoints.
Observation processing follows this order: compute each term in physical units, add noise where enabled, apply the term's scale factor, concatenate and clip the terms, then normalize the resulting vector. Before normalization, each term has the form
dof_vel uses noise of ±0.05 rad/s and a scale of 0.05,
giving a noise contribution of ±0.0025 before normalization.enable_noise = False
for the entire group, including terms with a configured noise level.obs_normalization.history_length = 1.
The actor receives temporal information through two previous-action inputs and gait phase;
the critic also receives the gait-average state listed in §2.3.The actor receives 198 inputs across 16 configured terms. The Symbol column maps each term to the paper's observation vector \(o_t\). The first four terms form the task command \(c_t\); the remaining terms provide tracking errors, robot state, action history, and gait phase. Noise magnitudes are specified before scaling.
| Term | Symbol | Dim | Scale | Noise | Definition |
|---|---|---|---|---|---|
arm_command | \(g_{RD}\) in \(c_t\) | 18 | 1.0 | 0 | Desired end-effector pose per arm in \(\{R\}\): 3 position + 6D rotation. |
command_lin_vel | \(v_{xy,d}^{P}\) in \(c_t\) | 2 | 1.0 | 0 | Commanded planar velocity \((v_{x,d}^{P}, v_{y,d}^{P})\). |
command_ang_vel | \(\omega_{z,d}^{P}\) in \(c_t\) | 1 | 1.0 | 0 | Commanded yaw rate. |
base_height_command | \(z^{P}_{d}\) in \(c_t\) | 1 | 1.0 | 0 | Commanded pelvis height, randomized by the curriculum. |
ee_error_pose | \(g_{DE,t}\) | 18 | 1.0 | 0 | Per arm: \(R_{RD}^{\top}(p_{RE} - p_{RD})\) and the 6D form of \(R_{RD}^{\top}R_{RE}\). Both errors are expressed in the desired end-effector frame. |
ee_error_pose_tanh | \(p^{\mathrm{sat}}_{DE,t}\) | 6 | 1.0 | 0 | \(\tanh(p_{DE,t}/\tau)\) per component with \(\tau = 0.01\), i.e. \(\tanh(100\,p_{DE,t})\). Amplifies small position errors and approaches saturation beyond about 2 cm. This channel contains only position errors. |
controller_arm_joint_command | \(a_{\mathrm{GAC},t}\) | 14 | 1.0 | 0 | The cached nominal controller command to which the policy adds its residual. |
ee_body_twist | \(V^{E}_{RE,t}\) | 12 | 0.05 | 0.02 | Body twists for both arms, computed without base linear velocity because that quantity is unavailable at deployment. |
dof_pos_relative | \(q_t\) | 29 | 1.0 | 0.01 | Supplied as \(q_t - q^{\mathrm{def}}\), the offset from the default pose. |
dof_vel | \(\dot q_t\) | 29 | 0.05 | 0.05 | Measured joint velocity. The noise level reflects uncertainty in hardware velocity estimates. |
projected_gravity | \(g^{\mathrm{proj}}_{t}\) | 3 | 1.0 | 0.001 | The pelvis projected gravity \(g^{P}\) of §1. |
base_ang_vel | \(\omega_t\) | 3 | 0.25 | 0.01 | Base angular velocity expressed in the pelvis frame and obtained from the IMU at deployment. |
last_action | \(a_{\mathrm{RL},t-1}\) | 29 | 1.0 | 0 | Reads last_residual_policy_action: 15 non-arm targets concatenated with 14 scaled arm residuals. This stored vector has the same width as the raw action but contains processed values. |
prev_prev_action | \(a_{\mathrm{RL},t-2}\) | 29 | 1.0 | 0 | The raw 29-channel action two steps back, required by the mean-action acceleration penalty. |
sin_phase, cos_phase | \(\phi_t\) | 2 + 2 | 1.0 | 0 | \(\sin\phi_t\), \(\cos\phi_t\) per leg — encoding each leg's gait phase. |
| Term | Dim | Scale | Definition and reason for critic-only use |
|---|---|---|---|
base_lin_vel | 3 | 1.0 | Root linear velocity in the base frame. Not reliably estimated on the robot. |
base_height_error | 1 | 1.0 | \(z_P - z_d^P\). Requires absolute pelvis height, which needs external tracking. |
ee_body_twist_full | 12 | 0.05 | Ground-truth body twist including base linear velocity, which is omitted from the actor's version. |
gait_avg_velocity_state | 5 | 1.0 | \(\bar\omega_z\), \(\bar v_\perp\), both previous-cycle means, and the in-cycle sample count divided by 50. Provides the history used to compute the gait-averaged penalties, allowing the critic to account for their dependence on earlier motion. |
The critic receives 219 inputs: the actor's 198 inputs plus the 21 inputs
listed above. It includes both ee_body_twist and
ee_body_twist_full. Noise is disabled for the entire critic group, so the
configured dof_vel noise level of 0.02, compared with the actor's 0.05,
has no effect.
The policy outputs 29 action channels: 14 for the arms and 15 for the legs and waist. We write this partition as \(a_{\mathrm{RL},t} = [a_{\mathrm{RL,arm},t}^{\top},\, a_{\mathrm{RL,other},t}^{\top}]^{\top}\). The non-arm group contains the 12 leg joints and 3 waist joints. The policy directly controls this group, and the non-arm action-rate penalty applies to all 15 channels.
For the arms, the policy supplies a residual correction to the nominal controller command. Both outputs are normalized joint-position actions. The residual is clipped to \(\pm c\), multiplied by \(s\), and added to the controller command. The low-level controller then applies the per-joint scale \(\alpha\) and adds the default joint position:
The effective residual scale is therefore \(s_r = s\,\alpha = 0.75\), which limits the correction to \(\pm 0.75\) rad per arm joint. Leg channels use the same \(\alpha = 0.25\) without the residual pre-scale. The configuration stores the pre-scale of 3.0 and action scale of 0.25 separately; both factors are needed to determine the correction in radians.
The arrows in Table II denote linear ramps within stage D. Each parameter changes over the window listed below and remains constant outside that window. These gradual changes avoid abrupt shifts in reward scale while the off-policy replay buffer still contains transitions scored with earlier kernel settings.
| Term | Field | Start → End | Ramp window |
|---|---|---|---|
| EEF position | tracking_sigma | 0.005 → 0.001 | 20% – 70% |
| EEF orientation | tracking_sigma | 0.10 → 0.02 | 20% – 70% |
| Linear velocity | tracking_sigma | 0.25 → 0.10 | 20% – 80% |
| Arm residual | weight | −0.10 → −0.05 | 0% – 50% |
\(\sigma_v\) is the linear kernel; the angular kernel holds at 0.25 throughout and does not anneal.
Stages A–C may finish before their iteration budgets, as indicated by \(\le\) in Table II. Early stopping requires a mean episode length of at least 990 out of 1000 steps, sustained for a specified number of consecutive evaluations after a minimum iteration count. Stage A uses a patience of 1000 evaluations and a minimum of 5k iterations. Stage B uses 2000 and 5k; stage C uses 2000 and 10k. Later stages run their full budgets.
The reported run used 201,911 of the nominal 210,000 iterations. Early stopping therefore saved about 8k iterations across the first three stages.
The table below lists stage-dependent initialization ranges, height commands, command sampling probabilities, and workspace-bound weights. These settings complement the command ranges in Table II.
| A | B | C | D | E | F | |
|---|---|---|---|---|---|---|
| Initial joint scale | [.9, 1.1] | [.85, 1.15] | [.85, 1.15] | [.7, 1.3] | [.7, 1.3] | [.7, 1.3] |
| Commanded height [m] | 0.73 | [.68, .78] | [.68, .78] | [.55, .80] | [.55, .80] | [.55, .80] |
| Stand probability | 1.0 | 1.0 | 0.3 | 0.2 | 0.2 | 0.2 |
| Turn-in-place prob. | — | — | 0.05 | 0.05 | 0.10 | 0.10 |
| Workspace bound weight | — | 0.3 | — | 0.7 | 0.7 | 0.7 |
At reset, each default joint position is multiplied by a sampled initial joint scale. From stage D onward, this permits a deviation of up to 30% from the default value for each joint. Initial joint velocities are always zero.
The following randomization settings are shared by all six stages. Startup parameters are sampled once when the scene is created. Reset parameters are sampled at the start of each episode. Sampling occurs per environment and, where specified, per link or joint.
| When | Quantity | Range | Granularity |
|---|---|---|---|
| Startup | Link mass | ×[0.90, 1.10] | per link, over pelvis, hip yaw/roll/pitch, and knee links |
| Added base mass | [−0.5, +1.5] kg | per environment | |
| Ground friction | [0.50, 1.25] | per environment | |
| Base center-of-mass offset | ±0.05 m in \(x, y, z\) | per environment | |
| Reset | PD stiffness scale \(k_p\) | ×[0.90, 1.10] | per environment and per joint |
| PD damping scale \(k_d\) | ×[0.90, 1.10] | per environment and per joint | |
| Torque-RFI limit scale | ×[0.50, 1.50] | per environment and per joint | |
| Push schedule | resampled | per environment | |
| Step | Base velocity push | up to 0.5 m/s in \(x\) and \(y\), every 5–10 s | per environment |
Random force injection (RFI) adds torque noise proportional to each joint's effort limit:
The torque noise is resampled every control step, and the per-joint scale \(s_j\) is resampled at each reset. Together, the nominal 10% factor and the sampled scale produce a noise amplitude of 5–15% of each joint's effort limit. At the nominal scale, this corresponds to \(\pm 13.9\) N·m at a knee and \(\pm 0.5\) N·m at a wrist. Scaling by \(\tau^{\max}_j\) keeps the perturbation proportional to actuator capacity.
Observation noise is applied only to actor inputs. Its per-term settings are listed in §2.2.
| Parameter | Value | Parameter | Value |
|---|---|---|---|
| Discount \(\gamma\) | 0.97 | Target smoothing \(\tau\) | 0.125 |
| Batch size | 8192 | Buffer size (per env) | 1024 |
| Updates per step | 8 | Policy delay | 4 |
| Actor LR | 3×10−4 | Critic LR | 3×10−4 |
| Entropy LR | 3×10−4 | \(\alpha\) init / autotune | 0.001 / on |
| Critic ensemble | 2 | Distributional atoms | 101 |
| Value support | [−20, 20] | Weight decay | 0.001 |
| Actor hidden dim | 512 | Critic hidden dim | 768 |
| \(\log\sigma\) range | [−5, 0] | Grad-norm clip | disabled |
| Layer norm | on | Observation normalization | on |
| Symmetry coef. (actor) | 5.0 | Symmetry coef. (critic) | 0.5 |
| Mixed precision | bf16 | torch.compile | on |
| Parallel environments | 4096 | Learning starts | 10 |
The actor and critic are multilayer perceptrons with three hidden layers of widths \(H\), \(H/2\), and \(H/4\). Each hidden layer applies Linear → LayerNorm → SiLU. Neither network uses convolutions or recurrence, and observations are not stacked across time. The policy receives temporal information through gait phase and the two previous-action inputs.
| Actor | Critic (per Q-network) | |
|---|---|---|
| Input | 198 | 248 = 219 obs + 29 action |
| Hidden widths | 512 → 256 → 128 | 768 → 384 → 192 |
| Normalization | LayerNorm on every hidden layer | LayerNorm on every hidden layer |
| Activation | SiLU | SiLU |
| Output head | two \(\times\) 29: mean and log-std | 101 atoms (distributional) |
| Output squashing | \(\tanh\), log-std clamped to [−5, 0] | softmax over atoms on [−20, 20] |
| Parameters | 275,386 | 582,629 |
The two Q-networks contain 1,165,258 parameters. Together with the actor, the trained networks contain 1,440,644 parameters. Two additional target Q-networks are updated by Polyak averaging with \(\tau = 0.125\), bringing the total resident count to 2,605,902. Only the actor is deployed on the robot.
Replay capacity is specified per environment, whereas batch size is specified across all environments:
| Condition | Threshold |
|---|---|
| Base height | < 0.35 m |
| Base orientation (projected gravity \(x\) or \(y\)) | > 0.85 |
| Joint position limits | at 100% of range |
| Timeout | 1000 steps — not scored as failure |
Physics runs at 200 Hz with a control decimation of 4, giving a 50 Hz control rate. The simulator uses 1 substep and PhysX with 8 position and 4 velocity solver iterations. A contact-force history of length 4 captures every physics step in a control interval for the peak-force penalty. The scene uses plane terrain with nominal static and dynamic friction of 1.0 and environment spacing of 20 m. Episodes last up to 20 s.
E2E-RL removes the nominal controller and learns arm joint targets directly. The following tables describe its differences from ResGAC in action scaling, rewards, training settings, and observations. Both methods use the FastSAC optimizer settings listed in §5.1; differences in symmetry coefficients are listed below.
| ResGAC | E2E-RL | |
|---|---|---|
| Arm action | residual on \(a_{\text{GAC}}\) | direct joint target |
| Arm pre-scale | 3.0 | 7.0 |
| Arm clip | 1.0 | 1.0 |
| Action scale \(\alpha\) | 0.25 | 0.25 |
| Effective arm range | ±0.75 rad about \(a_{\text{GAC}}\) | ±1.75 rad about \(q^{\text{def}}\) |
| Leg action | direct, \(\alpha = 0.25\) | direct, \(\alpha = 0.25\) |
The baseline has 2.3× the arm action range because it must generate the full arm pose. Its ±1.75 rad range allows larger departures from the default pose than ResGAC's ±0.75 rad residual range, which corrects an existing controller command.
The baseline uses the same overall reward structure, with changes to tracking kernel widths, the arm-action penalty, and the pre-touchdown penalty. The table lists these differences alongside changes to command range and critic symmetry regularization. The baseline's tracking kernels and arm-action penalty were tuned for direct arm control.
| Final-stage setting | ResGAC | E2E-RL |
|---|---|---|
| EEF position kernel | \(\exp\!\left(-\frac{e_2}{\sigma_2} - \frac{e_1}{\sqrt{3}\sigma_1}\right)\) | \(\exp\!\left(-\frac{e_2}{\sigma_P}\right)\), \(L_1\) off |
| \(\sigma_2\) / \(\sigma_P\) | 0.001 | 0.005 |
| \(\sigma_1\) | 0.01 | 0 (disabled) |
| EEF orientation \(\sigma_R\) | 0.02 | 0.10 |
| Arm-action penalty, quantity | post-scale residual | raw arm action |
| weight | −0.05 | −0.5 |
| per-DOF weighting | uniform | shoulders/elbow 1.0, wrists 0.5 |
| Command range \((v_x, v_y, \omega_z)\) | ±0.7 | ±1.0 |
| Pre-touchdown velocity | −2 | not enabled |
| Symmetry coef. (critic) | 0.5 | 1.0 |
For an exponential kernel \(\exp(-e/\sigma)\), reward sensitivity depends on the error relative to \(\sigma\). When \(e \gg \sigma\), the reward is near zero and changes little as the error decreases. The kernel width therefore determines which error range provides a useful learning signal.
With a zero residual, ResGAC begins from the nominal controller's tracking behavior. Our interpretation is that this starting point makes a tighter kernel useful: the policy can refine an existing tracking solution. This motivated \(\sigma_2 = 0.001\) and the additional \(L_1\) term.
The baseline must learn arm tracking directly, and its tracking errors remained larger during tuning. A wider kernel can provide a useful reward signal over this larger error range. We therefore use \(\sigma_P = 0.005\) and \(\sigma_R = 0.10\) for the baseline, compared with 0.001 and 0.02 for ResGAC.
The \(L_1\) term adds \(e_1/(\sqrt{3}\,\sigma_1)\) to the same exponent as the squared-error term. At small errors, this linear contribution remains sensitive as the squared contribution diminishes. At larger errors, it can instead push the exponential reward close to zero. This may explain why the baseline trained better with the \(L_2^2\) kernel alone: a weak tracking signal could allow other reward terms to dominate.
This interpretation suggests that a nominal controller may expand the range of reward kernels that support learning, in addition to improving final tracking accuracy. Testing that hypothesis would require a separate \(\sigma\) sweep for each architecture. The present comparison establishes that the tuned configurations differ; it does not establish that their trainable ranges differ.
All kernel widths were hand-tuned. We do not claim that 0.001 and 0.005 are optimal, that their ratio has a theoretical basis, or that reward sensitivity is the only reason the baseline benefits from disabling the \(L_1\) term.
Table III therefore compares methods with separately tuned tracking rewards. The configuration differences above should be considered when interpreting the results.
The baseline omits the controller joint command,
controller_arm_joint_command (14 dimensions), and the controller-related
twist terms, ee_body_twist / ee_body_twist_full (12 dimensions
each), from the corresponding observation groups. The actor consequently uses 14 terms
instead of 16. Observation scales, noise levels, clipping, and normalization are unchanged.
The baseline uses a separate pipeline module. It does not receive the controller damping-ratio or angular-velocity-sigma flags, which apply only when a controller is present. Stage structure, iteration counts, trajectory datasets, and seeds are shared.
Reward weights were chosen in three broad groups. Tracking rewards have weights of 1–5 and unweighted values in \([0, 1]\). End-effector position and orientation receive weights of 4.0 and 3.0, compared with 2.0 and 1.5 for locomotion velocity tracking, reflecting the emphasis on manipulation accuracy. Shaping penalties have weight magnitudes of 0.05–2.5. Their unweighted values are unbounded, so these weights are kept small to limit their contribution during normal operation. Constraint penalties use larger magnitudes: 5–10 for foot proximity, foot orientation, and heading drift, and 200 for termination. These discourage violations strongly. A failure penalty of −200 equals 20 steps of the +10 survival bonus.
The survival bonus is 10.0, compared with 2.5 elsewhere in the codebase. This larger constant reward was chosen to support the scale of early value estimates in off-policy FastSAC training, when tracking rewards are still near zero.
For a tracking kernel \(\exp(-e/\sigma)\), the reward falls to \(1/e\) when the error equals \(\sigma\). The width thus specifies an error tolerance. The curriculum reduces \(\sigma_P\) from 0.02 to 0.001 after the policy learns to walk and reach, shifting reward sensitivity toward smaller errors. The combined \(L_2^2 + L_1\) kernel adds sensitivity near the target, where the squared-error contribution becomes small. Section 6.2 discusses why the two architectures use different kernel widths and the limits of that explanation.
Command masks activate penalties only when the associated motion is undesirable. For example, the lateral-drift penalty measures velocity perpendicular to the commanded direction, allowing intentional strafing. Gait-cycle averaging reduces sensitivity to the yaw and sway that occur within a stride and emphasizes net drift. The instantaneous term, weighted by 0.15×, retains a reward signal at each control step.
Observation scales bring channels to roughly unit range before normalization. Joint velocities and end-effector twists use 0.05, base angular velocity uses 0.25, and channels already treated as \(O(1)\), including commands, gravity, phases, and pose errors, use 1.0. Noise levels reflect the reliability of hardware estimates: 0.001 for projected gravity, 0.01 for joint positions and base angular velocity, and 0.05 for joint velocities. The saturated position channel \(\tanh(100\,p)\) amplifies millimeter-scale errors, which are of order \(10^{-3}\) in meters, while bounding the contribution of large errors.
These values were selected through iterative tuning on the reported task. No systematic search or ablation isolates the effect of individual weights. They describe one working configuration, supported by the design rationale above.