Reinforcement Learning
MyoAssist’s reinforcement learning (RL) pipeline uses Stable-Baselines3 (SB3) PPO and custom MuJoCo environments. These environments simulate human–exoskeleton interaction. This page gives an overview of how the parts fit together and where to find more information.
Reinforcement learning (RL) is a machine learning method. An agent learns to make decisions. It interacts with an environment and receives rewards as feedback. In MyoAssist, RL trains control policies for human–exoskeleton systems in MuJoCo simulation environments.
Observation Space:
In our environments, the agent receives observations that include:
- Joint angles
- Joint velocities
- Muscle activations
- Sensory data (such as ground contact, force sensors, etc.)
- Target velocity (the last observation component)
Action Space:
The agent outputs actions that control:
- Muscle activations (for the human actor network)
- Exoskeleton control values (for the exoskeleton actor network)
Training Workflow
- Define a config: start from an existing JSON preset, or create one from scratch.
- Launch training
python rl_train/run_train.py --config_file_path rl_train/train/train_configs/my_config.json - Monitor progress: logs and results go to
rl_train/results/train_session_*. - Evaluate policy:
python rl_train/run_policy_eval.py rl_train/results/train_session_<timestamp> - Analyze results: see Evaluation for the shared eval outputs.
Key Features
- Multi-Actor Support: Separate networks for human muscles and exoskeleton actuators (see Network Index Handler).
- Variable Terrain: Train on flat, sloped, rough, or tiled terrain, defined by Terrains.
- Reference Motion Imitation: Optional imitation reward using ground-truth gait trajectories.
- Realtime Evaluation: Run policies in realtime with
--flag_realtime_evaluate.
Getting Started
This guide shows you the fastest way to test the RL system and run training in the MyoAssist RL system.
RL Training Entry Points
Here is a quick overview of the main entry point scripts in the rl_train folder:
| File | Purpose |
|---|---|
run_sim_minimal.py | The simplest way to create and test a MyoAssist RL environment. No training, just environment creation and random actions. |
run_train.py | Main entry point for running RL training sessions. Loads configuration, sets up environments, and starts training. |
run_policy_eval.py | Entry point for evaluating and analyzing trained policies. Useful for testing policy performance and generating analysis results. |
Quick Test Commands
1. Environment Creation Example
See how to create a simulation environment and run for 150 frames(5sec):
python rl_train/run_sim_minimal.py
- mac:
mjpython rl_train/run_sim_minimal.pyNote: If you need MuJoCo visualizer in mac os, simply use
mjpythoninstead ofpythonto run your script.
You do not need to install anything extra. Just change the command:
Note:
If you see the error messageModuleNotFoundError: No module named 'flatten_dict', simply run the command again. This will usually resolve the problem automatically.
What this does:
- Shows an example of creating a Gym wrapped MuJoCo simulation environment
- No actual training - just environment creation example
Terminated vs Truncated In-depth explanation of the terminated and truncated values in Gymnasium’s Env.step API
2. Quick Training Test
Run a minimal training session to verify everything works:
python rl_train/run_train.py --config_file_path rl_train/train/train_configs/test.json --flag_rendering
What this does:
- Runs actual reinforcement learning training
- Training for only a few short timesteps
- Uses 1 environment (minimal resource usage)
- Enables rendering to see the simulation
- Logs results after every rollout (256 steps,
test.json) for immediate feedback
3. Check Results
After training, check the results folder:
# Results location
rl_train/results/train_session_[date-time]/
Concurrent runs never share a directory. Training claims train_session_[date-time], and steps to train_session_[date-time]_1, _2, and so on when the second-resolution timestamp is already taken.
What you’ll find:
analyze_results_[timesteps]_[evaluate_number]: analysis written during training, by the analyzer that the learning callback runs everylogger_params.evaluate_frequencyrolloutssession_config.json: Configuration used for this trainingtrain_log.json: Training log datatrained_models/: Trained models(.zip) saved at each log interval - can be used for evaluation or transfer learning
run_policy_eval.py writes analyze_results_[NN] instead, without the timestep prefix.
Full Training (When Ready)
Turn the model cache on first. Each of the num_envs workers composes its own model, so training without the cache is much slower:
export MYOASSIST_CACHE_DIR=~/.cache/myoassist
The variable covers RL and controller optimization. See Caching.
Once you’ve verified everything works, run full training:
python rl_train/run_train.py --config_file_path rl_train/train/train_configs/imitation_tutorial_22_separated_net_partial_obs.json
This file is the default example configuration we provide.
For more details, see the RL Configuration section.
Note:
The provided config setsnum_envsto 32.
Depending on your PC’s capability, try lowering this to 4, 8, or 16.
You should also adjustn_stepsaccordingly.
For example, if you usenum_envs=16(half of 32), you should doublen_stepsto keep the total batch size the same.
Policy Evaluation
Test a trained model:
python rl_train/run_policy_eval.py [path/to/trainsession/folder]
Point
run_policy_eval.pyat anytrain_session_*directory that you produced.
| Flag | Meaning |
|---|---|
--steps N | Override num_timesteps for every rollout. The configs ship 200 steps, about 5 strides. Works on already-trained sessions. |
--regen | Regenerate the evaluated gait data even if it already exists. |
--no-show | Skip the pop-out composite window. |
--varying | Replace evaluate_param_list with a single SINUSOIDAL 0.8-1.4 m/s rollout and emit the speed-tracking composite. |
--cmap {rainbow,teal,bluered} | Speed color map for varying-speed composites. |
--legacy-plots | Also write the legacy per-panel PNGs. |
After training, an analyze_results folder will be created inside your train_session directory.
This folder contains various plots and videos that visualize your agent’s performance.
- Where to find:
rl_train/results/train_session_[date-time]/analyze_results_[NN]/ - What’s inside:
composite.png(and.svg),replay.mp4, andgait_evaluated_data.json- See Evaluation for the full description of these outputs.
The parameters used for evaluation and analysis (such as which plots/videos are generated) are controlled by the evaluate_param_list in your session_config.json file.
For more details on how to customize these parameters, see the RL Configuration section.
Transfer Learning

python rl_train/run_train.py --config_file_path [path/to/transfer_learning/config.json] --config.env_params.prev_trained_policy_path [path/to/pretrained_model]
or you can specify the env_params.prev_trained_policy_path in config(.json) file
Note: The
[path/to/pretrained_model]should point to a.zipfile, but do not include the.zipextension in the path.
Realtime Policy Running
You can run a trained policy in realtime simulation:
- windows:
python rl_train/run_train.py --config_file_path [path/to/config.json] --config.env_params.prev_trained_policy_path [path/to/model_file] --flag_realtime_evaluate - mac:
mjpython rl_train/run_train.py --config_file_path [path/to/config.json] --config.env_params.prev_trained_policy_path [path/to/model_file] --flag_realtime_evaluate
Parameters:
[path/to/config.json]: Path to the JSON file in the train_session folder[path/to/model_file]: Path to the model file (.zip) without extension. It is located in the train_models folder
Use a
session_config.jsonand amodel_<steps>file from atrain_session_*directory that you produced.