← Back to the project page

Additional Training Details

Complete reward definitions, observation processing, training settings, and baseline configurations.

Anonymous submission. The framework name and repository are withheld during review.

This supplement documents the implementation details needed to reproduce the training setup: all 25 reward terms, observation processing, action scaling, curriculum settings, and the E2E-RL baseline configuration. The paper presents the method, GAC derivation, frame construction, symmetry losses, and curriculum schedule.

Contents

  1. Reward definitions and weights
  2. Observation specification and pre-processing
  3. Action scaling
  4. Additional curriculum settings
  5. Training and simulation settings
  6. E2E-RL baseline configuration
  7. Parameter selection rationale

1. Reward definitions and weights

The following tables define the 25 reward terms listed in Table I of the paper, using the same names and order. The notation also follows the paper: \(i \in \{l, r\}\) indexes arms or feet; \(q\) denotes joint positions; \(v_{xy}^{P}\) and \(\omega_{z}^{P}\) denote planar velocity and yaw rate in the body frame; and \(c = (v_{xy,d}^{P}, \omega_{z,d}^{P})\) is the locomotion command. The mean policy action is \(\mu_t = \mu_\theta(o_t)\), and the per-leg gait phase is \(\phi \in [-\pi, \pi]^2\). We use \(\mathbb{1}[\cdot]\) for the indicator function and \((\cdot)_+\) for the positive part:

$$ (x)_+ \;\triangleq\; \max(x,\, 0) $$

A penalty based on \((x)_+\) applies only when \(x > 0\). The joint-position limit penalty uses the linear excess beyond a soft bound. The foot contact force and pre-touchdown velocity penalties use squared excess, \((x)_+^2\), to penalize forces or descent speeds above their thresholds.

Orientation rewards and penalties use projected gravity: the world gravity direction expressed in a body's local frame \(\{X\}\):

$$ g^{X} \;\triangleq\; R_{WX}^{\top}\,(0,\,0,\,-1)^{\top} \;\in\; \mathbb{R}^3 $$

The horizontal component \(g^{X}_{xy}\) is zero when the frame's \(z\)-axis is vertical. Its magnitude, \(\lVert g^{X}_{xy}\rVert\), provides a measure of tilt without a reference orientation or an orientation singularity. We use this measure for the pelvis \(g^{P}\), torso \(g^{T}\), and feet \(g^{F_i}\). For the locomotion penalties below, \(\hat c = v_{xy,d}^{P}/\lVert v_{xy,d}^{P}\rVert\) denotes the unit commanded heading, and \(\psi\) denotes base yaw.

1.1 Manipulation

Term\(w_i\)Definition
EEF position4.0 \(\exp\!\left(-\dfrac{e_{2,t}}{\sigma_2} - \dfrac{e_{1,t}}{\sqrt{3}\,\sigma_1}\right)\). Stages A–D use the \(L_2^2\)-only form \(\exp(-e_{2,t}/\sigma_P)\). The baseline disables the \(L_1\) term in stages E–F by setting \(\sigma_1 \le 0\) (§6).
EEF orientation3.0 \(\exp\!\left(-\operatorname{mean}_i\operatorname{tr}\left(I_3 - R_{DE_i,t}\right)/\sigma_R\right)\). Both poses are taken in the manipulation frame; the target rotation is reconstructed from the 6D representation carried in the command.

1.2 Locomotion

Term\(w_i\)Definition
Linear velocity2.0 \(\exp\!\left(-\lVert v_{xy,d}^{P} - v_{xy}^{P}\rVert^2/\sigma_v\right)\), with \(v_{xy}\) the root linear velocity rotated into the base frame.
Angular velocity1.5 \(\exp\!\left(-(\omega_{z,d}^{P} - \omega_{z}^{P})^2/\sigma_\omega\right)\), \(\sigma_\omega = 0.25\) throughout the curriculum.
Swing-foot trajectory5.0 \(\exp\!\left(-\left[(z_l - z_l^{*})^2 + (z_r - z_r^{*})^2\right]/\sigma\right)\), \(\sigma = 0.008\), foot heights measured relative to the terrain. The reference \(z^{*}\) is defined below.
Pelvis orientation1.0 \(\exp\!\left(-\lVert g^{P}_{xy}\rVert^2/\sigma\right)\), the projected gravity of the pelvis frame.
Torso orientation1.0 \(\exp\!\left(-\lVert g^{T}_{xy}\rVert^2/\sigma\right)\) — the torso's projected gravity measures waist lean separately from pelvis tilt.
Base height1.0 \(\exp\!\left(-(100\,(z^P - z^{P}_d))^2/\sigma\right)\), \(\sigma = 3.0\). The height error is converted to centimeters before squaring, giving a kernel width of \(3\ \mathrm{cm}^2\). The commanded height \(z_d^P\) varies with the curriculum.

1.3 Action and posture penalties

Term\(w_i\)Definition
Arm residual−0.05 \(\lVert a_{\mathrm{RL,arm},t}\rVert^2\) uses the buffered residual after clipping and scaling. In terms of the raw policy output, this is \(s^2\lVert\operatorname{clip}(a_{\mathrm{RL,arm}},\pm c_r)\rVert^2\) with \(s = 3\). Scaling therefore multiplies the squared clipped residual by 9.
EEF body twist−0.05 \(\sum_i \lVert V^{E}_{RE,i}\rVert^2\) sums the ground-truth body-twist penalty over both arms. The twist comes from the controller when available and from a rigid-body computation otherwise.
Arm action rate−0.5 \(\lVert \Delta a_t\rVert^2\) over the 14 arm channels of the raw action.
Non-arm action rate−1.0 The same action-rate penalty over the 15 non-arm channels: 12 leg joints and 3 waist joints.
Mean-action accel.−0.10 \(\lVert \Delta^2\mu_t\rVert^2 = \lVert \mu_t - 2\mu_{t-1} + \mu_{t-2}\rVert^2\). Applied as two equally weighted terms over the 14 arm and 15 non-arm channels. Using the deterministic policy mean \(\mu\) prevents exploration noise in sampled actions from dominating the penalty.
Arm joint velocity−0.05 \(\sum_{j\in\text{arm}} \dot q_j^2\), using measured joint velocities.
Pose deviation−1.0 \(\sum_j w_j (q_j - q_j^{\text{def}})^2\). Only leg and waist joints contribute; arm weights are zero. Per-joint weights: hip pitch 0.01, hip roll 1.0, hip yaw 5.0, knee 0.01, ankle pitch and roll 5.0 (each leg); waist yaw, roll, pitch 50.0.
Joint-position limit−1.0 \(\sum_j \left[(q^{\text{lo}}_j - q_j)_+ + (q_j - q^{\text{hi}}_j)_+\right]\) penalizes the linear excess beyond the soft limits. These limits shrink the hardware range about its midpoint, \(q^{\text{lo/hi}} = m \mp 0.5\,\rho\,r\), with \(\rho = 0.95\).

1.4 Foot penalties and survival rewards

Term\(w_i\)Definition
Foot proximity−10 The indicator \(\mathbb{1}[d < 0.15\ \mathrm{m}]\), where \(d = \lvert\cos\psi\,\Delta y - \sin\psi\,\Delta x\rvert\) is the foot separation projected onto the base lateral axis. The penalty is binary and depends on lateral separation, so it discourages foot crossing.
Foot orientation−5 \(\sum_i \lVert g^{F_i}_{xy}\rVert\) sums the projected-gravity norms of the two feet. Using the norm keeps the penalty sensitive to small tilts.
Alive / termination10 / −200 The alive term is a constant \(+1\) per step. The termination term is \(\mathbb{1}[\text{reset} \wedge \neg\text{timeout}]\), applied once on a failure reset. A 1000-step timeout incurs no termination penalty.

1.5 Command-dependent locomotion penalties

These penalties are active only for the commands specified in the Mask column. The masks use the commanded planar velocity \(v_{xy,d}^{P} = (v_{x,d}^{P}, v_{y,d}^{P})\) and yaw rate \(\omega_{z,d}^{P}\) from \(c_t\). The penalties use measured velocity \(v_{xy}^{P}\) and yaw rate \(\omega_{z}^{P}\). Velocity and yaw-rate thresholds are in m/s and rad/s, respectively.

Term\(w_i\)MaskPenalized quantity
Heading drift−10 \(\lVert v_{xy,d}^{P}\rVert > 0.05 \;\wedge\; \lvert\omega_{z,d}^{P}\rvert < 0.03\) \(\left(\bar\omega_{z}^{P}\right)^2 + 0.15\left(\omega_{z}^{P}\right)^2\)
Lateral drift−2 \(\lVert v_{xy,d}^{P}\rVert > 0.03\) \(\left(\bar v_\perp\right)^2 + 0.15\left(v_\perp\right)^2\), with the signed \(v_\perp = \hat c_x v_{y}^{P} - \hat c_y v_{x}^{P}\)
Turn-in-place drift−2.5 \(\lVert v_{xy,d}^{P}\rVert < 0.03 \;\wedge\; \lvert\omega_{z,d}^{P}\rvert > 0.1\) \(\lVert v_{xy}^{P}\rVert^2\)
Standing yaw rate−2.5 \(\lVert v_{xy,d}^{P}\rVert < 0.03 \;\wedge\; \lvert\omega_{z,d}^{P}\rvert < 0.03\) \(\left(\omega_{z}^{P}\right)^2\)

Gait-cycle averaging. An overbar \(\bar\cdot\) denotes the running mean \(\bar e_k\) within the current gait cycle. A command-manager term updates this mean while the gait phase advances; standing adds no samples. The accumulator resets when the left-leg phase crosses zero in the positive direction. During the first 5 samples of each cycle, the previous cycle's mean is used instead. The mean penalizes net drift, while the instantaneous term, weighted by \(0.15\times\), provides a reward signal at each step.

1.6 Impact handling

Term\(w_i\)Definition
Foot contact force−2 \(\sum_i \left(\dfrac{(F^{\text{peak}}_{z,i} - F_{\text{th}})_+}{F_{\text{th}}}\right)^2\), \(F_{\text{th}} = 400\ \mathrm{N}\) is a force threshold. \(F^{\text{peak}}_z\) is the maximum over all physics substeps in the control interval. A contact-force history buffer captures these values, including brief impacts between control ticks. The implementation raises an error if the buffer is shorter than the control decimation.
Pre-touchdown velocity−2 \(\sum_i \mathbb{1}[\text{active}_i]\left((-\dot z_i) - 0.15\right)_+^2\), where \(\text{active}_i\) is true when all three conditions hold: the foot is airborne (\(F_z < 5\ \mathrm{N}\)), the foot is near the ground (clearance \(< 0.05\ \mathrm{m}\)), and the velocity command is non-trivial (\(> 0.03\ \mathrm{m/s}\)). The penalty does not apply during mid-swing descent.

Swing-foot reference. The target height \(z^{*}\) follows a two-piece cubic Bézier curve over normalized phase \(x = (\phi + \pi)/2\pi\). With \(b(x) = x^3 + 3x^2(1-x)\), the foot rises from \(0 \to h\) using \(b(2x)\) for \(x \le 0.5\), then falls from \(h \to 0\) using \(b(2x-1)\). The peak height is \(h = 0.09\ \mathrm{m}\). The curve has zero slope at liftoff and touchdown, giving zero commanded vertical velocity at both endpoints.

2. Observation specification and pre-processing

2.1 Observation processing

Observation processing follows this order: compute each term in physical units, add noise where enabled, apply the term's scale factor, concatenate and clip the terms, then normalize the resulting vector. Before normalization, each term has the form

$$ o \;=\; \operatorname{clip}\Big(\; \alpha_{\text{term}} \cdot \big( f_{\text{term}}(\text{state}) + \eta \big) ,\; \pm 100 \Big), \qquad \eta \sim \mathcal{U}(-n_{\text{term}},\, +n_{\text{term}}) $$
  1. Noise is added before scaling. Noise magnitudes are specified in physical units. For example, dof_vel uses noise of ±0.05 rad/s and a scale of 0.05, giving a noise contribution of ±0.0025 before normalization.
  2. Noise is uniform. Each element receives an independent sample \(\eta \sim \mathcal{U}(-n, +n)\) at every step.
  3. Only actor inputs receive noise. The critic uses enable_noise = False for the entire group, including terms with a configured noise level.
  4. Clipping precedes normalization. The concatenated vector is clipped to \(\pm 100\). The agent then applies running mean/variance normalization through obs_normalization.
  5. Observations are not stacked. Both groups use history_length = 1. The actor receives temporal information through two previous-action inputs and gait phase; the critic also receives the gait-average state listed in §2.3.

2.2 Actor observations (16 terms)

The actor receives 198 inputs across 16 configured terms. The Symbol column maps each term to the paper's observation vector \(o_t\). The first four terms form the task command \(c_t\); the remaining terms provide tracking errors, robot state, action history, and gait phase. Noise magnitudes are specified before scaling.

TermSymbolDimScaleNoiseDefinition
arm_command\(g_{RD}\) in \(c_t\)181.00 Desired end-effector pose per arm in \(\{R\}\): 3 position + 6D rotation.
command_lin_vel\(v_{xy,d}^{P}\) in \(c_t\)21.00 Commanded planar velocity \((v_{x,d}^{P}, v_{y,d}^{P})\).
command_ang_vel\(\omega_{z,d}^{P}\) in \(c_t\)11.00 Commanded yaw rate.
base_height_command\(z^{P}_{d}\) in \(c_t\)11.00 Commanded pelvis height, randomized by the curriculum.
ee_error_pose\(g_{DE,t}\)181.00 Per arm: \(R_{RD}^{\top}(p_{RE} - p_{RD})\) and the 6D form of \(R_{RD}^{\top}R_{RE}\). Both errors are expressed in the desired end-effector frame.
ee_error_pose_tanh\(p^{\mathrm{sat}}_{DE,t}\)61.00 \(\tanh(p_{DE,t}/\tau)\) per component with \(\tau = 0.01\), i.e. \(\tanh(100\,p_{DE,t})\). Amplifies small position errors and approaches saturation beyond about 2 cm. This channel contains only position errors.
controller_arm_joint_command\(a_{\mathrm{GAC},t}\)141.00 The cached nominal controller command to which the policy adds its residual.
ee_body_twist\(V^{E}_{RE,t}\)120.050.02 Body twists for both arms, computed without base linear velocity because that quantity is unavailable at deployment.
dof_pos_relative\(q_t\)291.00.01 Supplied as \(q_t - q^{\mathrm{def}}\), the offset from the default pose.
dof_vel\(\dot q_t\)290.050.05 Measured joint velocity. The noise level reflects uncertainty in hardware velocity estimates.
projected_gravity\(g^{\mathrm{proj}}_{t}\)31.00.001 The pelvis projected gravity \(g^{P}\) of §1.
base_ang_vel\(\omega_t\)30.250.01 Base angular velocity expressed in the pelvis frame and obtained from the IMU at deployment.
last_action\(a_{\mathrm{RL},t-1}\)291.00 Reads last_residual_policy_action: 15 non-arm targets concatenated with 14 scaled arm residuals. This stored vector has the same width as the raw action but contains processed values.
prev_prev_action\(a_{\mathrm{RL},t-2}\)291.00 The raw 29-channel action two steps back, required by the mean-action acceleration penalty.
sin_phase, cos_phase\(\phi_t\)2 + 21.00 \(\sin\phi_t\), \(\cos\phi_t\) per leg — encoding each leg's gait phase.

2.3 Critic-only additions (4 terms)

TermDimScaleDefinition and reason for critic-only use
base_lin_vel31.0Root linear velocity in the base frame. Not reliably estimated on the robot.
base_height_error11.0\(z_P - z_d^P\). Requires absolute pelvis height, which needs external tracking.
ee_body_twist_full120.05Ground-truth body twist including base linear velocity, which is omitted from the actor's version.
gait_avg_velocity_state51.0\(\bar\omega_z\), \(\bar v_\perp\), both previous-cycle means, and the in-cycle sample count divided by 50. Provides the history used to compute the gait-averaged penalties, allowing the critic to account for their dependence on earlier motion.

The critic receives 219 inputs: the actor's 198 inputs plus the 21 inputs listed above. It includes both ee_body_twist and ee_body_twist_full. Noise is disabled for the entire critic group, so the configured dof_vel noise level of 0.02, compared with the actor's 0.05, has no effect.

3. Action scaling

The policy outputs 29 action channels: 14 for the arms and 15 for the legs and waist. We write this partition as \(a_{\mathrm{RL},t} = [a_{\mathrm{RL,arm},t}^{\top},\, a_{\mathrm{RL,other},t}^{\top}]^{\top}\). The non-arm group contains the 12 leg joints and 3 waist joints. The policy directly controls this group, and the non-arm action-rate penalty applies to all 15 channels.

For the arms, the policy supplies a residual correction to the nominal controller command. Both outputs are normalized joint-position actions. The residual is clipped to \(\pm c\), multiplied by \(s\), and added to the controller command. The low-level controller then applies the per-joint scale \(\alpha\) and adds the default joint position:

$$ q^{\text{target}} \;=\; q^{\text{default}} \;+\; \alpha \left( a_{\text{GAC}} + s\,\operatorname{clip}(a_{\text{RL,arm}}, \pm c) \right), \qquad \alpha = 0.25,\; s = 3.0,\; c = 1.0 $$

The effective residual scale is therefore \(s_r = s\,\alpha = 0.75\), which limits the correction to \(\pm 0.75\) rad per arm joint. Leg channels use the same \(\alpha = 0.25\) without the residual pre-scale. The configuration stores the pre-scale of 3.0 and action scale of 0.25 separately; both factors are needed to determine the correction in radians.

4. Additional curriculum settings

4.1 Parameter annealing within stage D

The arrows in Table II denote linear ramps within stage D. Each parameter changes over the window listed below and remains constant outside that window. These gradual changes avoid abrupt shifts in reward scale while the off-policy replay buffer still contains transitions scored with earlier kernel settings.

TermFieldStart → EndRamp window
EEF positiontracking_sigma0.005 → 0.00120% – 70%
EEF orientationtracking_sigma0.10 → 0.0220% – 70%
Linear velocitytracking_sigma0.25 → 0.1020% – 80%
Arm residualweight−0.10 → −0.050% – 50%

\(\sigma_v\) is the linear kernel; the angular kernel holds at 0.25 throughout and does not anneal.

4.2 Early stopping in stages A–C

Stages A–C may finish before their iteration budgets, as indicated by \(\le\) in Table II. Early stopping requires a mean episode length of at least 990 out of 1000 steps, sustained for a specified number of consecutive evaluations after a minimum iteration count. Stage A uses a patience of 1000 evaluations and a minimum of 5k iterations. Stage B uses 2000 and 5k; stage C uses 2000 and 10k. Later stages run their full budgets.

The reported run used 201,911 of the nominal 210,000 iterations. Early stopping therefore saved about 8k iterations across the first three stages.

4.3 Settings that vary by stage

The table below lists stage-dependent initialization ranges, height commands, command sampling probabilities, and workspace-bound weights. These settings complement the command ranges in Table II.

ABCDEF
Initial joint scale[.9, 1.1][.85, 1.15][.85, 1.15][.7, 1.3][.7, 1.3][.7, 1.3]
Commanded height [m]0.73[.68, .78][.68, .78][.55, .80][.55, .80][.55, .80]
Stand probability1.01.00.30.20.20.2
Turn-in-place prob.——0.050.050.100.10
Workspace bound weight—0.3—0.70.70.7

At reset, each default joint position is multiplied by a sampled initial joint scale. From stage D onward, this permits a deviation of up to 30% from the default value for each joint. Initial joint velocities are always zero.

4.4 Randomization shared by all stages

The following randomization settings are shared by all six stages. Startup parameters are sampled once when the scene is created. Reset parameters are sampled at the start of each episode. Sampling occurs per environment and, where specified, per link or joint.

WhenQuantityRangeGranularity
StartupLink mass×[0.90, 1.10]per link, over pelvis, hip yaw/roll/pitch, and knee links
Added base mass[−0.5, +1.5] kgper environment
Ground friction[0.50, 1.25]per environment
Base center-of-mass offset±0.05 m in \(x, y, z\)per environment
ResetPD stiffness scale \(k_p\)×[0.90, 1.10]per environment and per joint
PD damping scale \(k_d\)×[0.90, 1.10]per environment and per joint
Torque-RFI limit scale×[0.50, 1.50]per environment and per joint
Push scheduleresampledper environment
StepBase velocity pushup to 0.5 m/s in \(x\) and \(y\), every 5–10 sper environment

Torque RFI

Random force injection (RFI) adds torque noise proportional to each joint's effort limit:

$$ \tau \;\leftarrow\; \tau \;+\; \mathcal{U}(-1, 1)\;\cdot\; \lambda_{\mathrm{RFI}}\;\cdot\; s_j \;\cdot\; \tau^{\max}_j, \qquad \lambda_{\mathrm{RFI}} = 0.10,\quad s_j \sim \mathcal{U}(0.5,\, 1.5) $$

The torque noise is resampled every control step, and the per-joint scale \(s_j\) is resampled at each reset. Together, the nominal 10% factor and the sampled scale produce a noise amplitude of 5–15% of each joint's effort limit. At the nominal scale, this corresponds to \(\pm 13.9\) N·m at a knee and \(\pm 0.5\) N·m at a wrist. Scaling by \(\tau^{\max}_j\) keeps the perturbation proportional to actuator capacity.

Observation noise is applied only to actor inputs. Its per-term settings are listed in §2.2.

5. Training and simulation settings

5.1 FastSAC

ParameterValueParameterValue
Discount \(\gamma\)0.97Target smoothing \(\tau\)0.125
Batch size8192Buffer size (per env)1024
Updates per step8Policy delay4
Actor LR3×10−4Critic LR3×10−4
Entropy LR3×10−4\(\alpha\) init / autotune0.001 / on
Critic ensemble2Distributional atoms101
Value support[−20, 20]Weight decay0.001
Actor hidden dim512Critic hidden dim768
\(\log\sigma\) range[−5, 0]Grad-norm clipdisabled
Layer normonObservation normalizationon
Symmetry coef. (actor)5.0Symmetry coef. (critic)0.5
Mixed precisionbf16torch.compileon
Parallel environments4096Learning starts10

Architecture

The actor and critic are multilayer perceptrons with three hidden layers of widths \(H\), \(H/2\), and \(H/4\). Each hidden layer applies Linear → LayerNorm → SiLU. Neither network uses convolutions or recurrence, and observations are not stacked across time. The policy receives temporal information through gait phase and the two previous-action inputs.

ActorCritic (per Q-network)
Input198248 = 219 obs + 29 action
Hidden widths512 → 256 → 128768 → 384 → 192
NormalizationLayerNorm on every hidden layerLayerNorm on every hidden layer
ActivationSiLUSiLU
Output headtwo \(\times\) 29: mean and log-std101 atoms (distributional)
Output squashing\(\tanh\), log-std clamped to [−5, 0]softmax over atoms on [−20, 20]
Parameters275,386582,629

The two Q-networks contain 1,165,258 parameters. Together with the actor, the trained networks contain 1,440,644 parameters. Two additional target Q-networks are updated by Polyak averaging with \(\tau = 0.125\), bringing the total resident count to 2,605,902. Only the actor is deployed on the robot.

Replay capacity, batch size, and symmetry augmentation

Replay capacity is specified per environment, whereas batch size is specified across all environments:

5.2 Termination

ConditionThreshold
Base height< 0.35 m
Base orientation (projected gravity \(x\) or \(y\))> 0.85
Joint position limitsat 100% of range
Timeout1000 steps — not scored as failure

5.3 Simulation

Physics runs at 200 Hz with a control decimation of 4, giving a 50 Hz control rate. The simulator uses 1 substep and PhysX with 8 position and 4 velocity solver iterations. A contact-force history of length 4 captures every physics step in a control interval for the peak-force penalty. The scene uses plane terrain with nominal static and dynamic friction of 1.0 and environment spacing of 20 m. Episodes last up to 20 s.

6. E2E-RL baseline configuration

E2E-RL removes the nominal controller and learns arm joint targets directly. The following tables describe its differences from ResGAC in action scaling, rewards, training settings, and observations. Both methods use the FastSAC optimizer settings listed in §5.1; differences in symmetry coefficients are listed below.

6.1 Action space

ResGACE2E-RL
Arm actionresidual on \(a_{\text{GAC}}\)direct joint target
Arm pre-scale3.07.0
Arm clip1.01.0
Action scale \(\alpha\)0.250.25
Effective arm range±0.75 rad about \(a_{\text{GAC}}\)±1.75 rad about \(q^{\text{def}}\)
Leg actiondirect, \(\alpha = 0.25\)direct, \(\alpha = 0.25\)

The baseline has 2.3× the arm action range because it must generate the full arm pose. Its ±1.75 rad range allows larger departures from the default pose than ResGAC's ±0.75 rad residual range, which corrects an existing controller command.

6.2 Rewards and training settings

The baseline uses the same overall reward structure, with changes to tracking kernel widths, the arm-action penalty, and the pre-touchdown penalty. The table lists these differences alongside changes to command range and critic symmetry regularization. The baseline's tracking kernels and arm-action penalty were tuned for direct arm control.

Final-stage settingResGACE2E-RL
EEF position kernel\(\exp\!\left(-\frac{e_2}{\sigma_2} - \frac{e_1}{\sqrt{3}\sigma_1}\right)\)\(\exp\!\left(-\frac{e_2}{\sigma_P}\right)\), \(L_1\) off
  \(\sigma_2\) / \(\sigma_P\)0.0010.005
  \(\sigma_1\)0.010 (disabled)
EEF orientation \(\sigma_R\)0.020.10
Arm-action penalty, quantitypost-scale residualraw arm action
  weight−0.05−0.5
  per-DOF weightinguniformshoulders/elbow 1.0, wrists 0.5
Command range \((v_x, v_y, \omega_z)\)±0.7±1.0
Pre-touchdown velocity−2not enabled
Symmetry coef. (critic)0.51.0

Rationale for the tracking kernel widths

Tuning rationale. The explanation below is a hypothesis based on observations during tuning. It has not been isolated experimentally and is not a result claimed in the paper.

For an exponential kernel \(\exp(-e/\sigma)\), reward sensitivity depends on the error relative to \(\sigma\). When \(e \gg \sigma\), the reward is near zero and changes little as the error decreases. The kernel width therefore determines which error range provides a useful learning signal.

With a zero residual, ResGAC begins from the nominal controller's tracking behavior. Our interpretation is that this starting point makes a tighter kernel useful: the policy can refine an existing tracking solution. This motivated \(\sigma_2 = 0.001\) and the additional \(L_1\) term.

The baseline must learn arm tracking directly, and its tracking errors remained larger during tuning. A wider kernel can provide a useful reward signal over this larger error range. We therefore use \(\sigma_P = 0.005\) and \(\sigma_R = 0.10\) for the baseline, compared with 0.001 and 0.02 for ResGAC.

The \(L_1\) term adds \(e_1/(\sqrt{3}\,\sigma_1)\) to the same exponent as the squared-error term. At small errors, this linear contribution remains sensitive as the squared contribution diminishes. At larger errors, it can instead push the exponential reward close to zero. This may explain why the baseline trained better with the \(L_2^2\) kernel alone: a weak tracking signal could allow other reward terms to dominate.

This interpretation suggests that a nominal controller may expand the range of reward kernels that support learning, in addition to improving final tracking accuracy. Testing that hypothesis would require a separate \(\sigma\) sweep for each architecture. The present comparison establishes that the tuned configurations differ; it does not establish that their trainable ranges differ.

All kernel widths were hand-tuned. We do not claim that 0.001 and 0.005 are optimal, that their ratio has a theoretical basis, or that reward sensitivity is the only reason the baseline benefits from disabling the \(L_1\) term.

Table III therefore compares methods with separately tuned tracking rewards. The configuration differences above should be considered when interpreting the results.

6.3 Observations

The baseline omits the controller joint command, controller_arm_joint_command (14 dimensions), and the controller-related twist terms, ee_body_twist / ee_body_twist_full (12 dimensions each), from the corresponding observation groups. The actor consequently uses 14 terms instead of 16. Observation scales, noise levels, clipping, and normalization are unchanged.

The baseline uses a separate pipeline module. It does not receive the controller damping-ratio or angular-velocity-sigma flags, which apply only when a controller is present. Stage structure, iteration counts, trajectory datasets, and seeds are shared.

7. Parameter selection rationale

Reward magnitudes

Reward weights were chosen in three broad groups. Tracking rewards have weights of 1–5 and unweighted values in \([0, 1]\). End-effector position and orientation receive weights of 4.0 and 3.0, compared with 2.0 and 1.5 for locomotion velocity tracking, reflecting the emphasis on manipulation accuracy. Shaping penalties have weight magnitudes of 0.05–2.5. Their unweighted values are unbounded, so these weights are kept small to limit their contribution during normal operation. Constraint penalties use larger magnitudes: 5–10 for foot proximity, foot orientation, and heading drift, and 200 for termination. These discourage violations strongly. A failure penalty of −200 equals 20 steps of the +10 survival bonus.

The survival bonus is 10.0, compared with 2.5 elsewhere in the codebase. This larger constant reward was chosen to support the scale of early value estimates in off-policy FastSAC training, when tracking rewards are still near zero.

Kernel widths

For a tracking kernel \(\exp(-e/\sigma)\), the reward falls to \(1/e\) when the error equals \(\sigma\). The width thus specifies an error tolerance. The curriculum reduces \(\sigma_P\) from 0.02 to 0.001 after the policy learns to walk and reach, shifting reward sensitivity toward smaller errors. The combined \(L_2^2 + L_1\) kernel adds sensitivity near the target, where the squared-error contribution becomes small. Section 6.2 discusses why the two architectures use different kernel widths and the limits of that explanation.

Masks and gait averaging

Command masks activate penalties only when the associated motion is undesirable. For example, the lateral-drift penalty measures velocity perpendicular to the commanded direction, allowing intentional strafing. Gait-cycle averaging reduces sensitivity to the yaw and sway that occur within a stride and emphasizes net drift. The instantaneous term, weighted by 0.15×, retains a reward signal at each control step.

Observation scaling and noise

Observation scales bring channels to roughly unit range before normalization. Joint velocities and end-effector twists use 0.05, base angular velocity uses 0.25, and channels already treated as \(O(1)\), including commands, gravity, phases, and pose errors, use 1.0. Noise levels reflect the reliability of hardware estimates: 0.001 for projected gravity, 0.01 for joint positions and base angular velocity, and 0.05 for joint velocities. The saturated position channel \(\tanh(100\,p)\) amplifies millimeter-scale errors, which are of order \(10^{-3}\) in meters, while bounding the contribution of large errors.

Tuning procedure

These values were selected through iterative tuning on the reported task. No systematic search or ablation isolates the effect of individual weights. They describe one working configuration, supported by the design rationale above.