Train 代码测试
- 作者仓库星标 1,155
- 许可证 BSD-2-Clause
- 作者仓库 LearningHumanoidWalking
/train — Launch a PPO Training Run
Parse the user's request from $ARGUMENTS and construct a training command.
Command Template
RAY_ADDRESS= uv run python run_experiment.py train --env <ENV> --logdir <LOGDIR> [OPTIONS...]
Available Environments
| Name | Description |
|---|---|
cartpole |
Cartpole swing-up (simplest, good for testing) |
h1 |
Unitree H1 standing task |
jvrc_walk |
JVRC humanoid basic walking |
jvrc_step |
JVRC humanoid stepping with planned footsteps |
Hyperparameters (defaults)
| Flag | Default | Description |
|---|---|---|
--n-itr |
20000 | Training iterations |
--lr |
1e-4 | Learning rate |
--gamma |
0.99 | Discount factor |
--std-dev |
0.223 | Action noise |
--learn-std |
off | Learn action noise (flag) |
--entropy-coeff |
0.0 | Entropy regularization |
--clip |
0.2 | PPO clipping |
--minibatch-size |
64 | Minibatch size |
--epochs |
3 | Optimization epochs per update |
--num-procs |
12 | Parallel workers |
--num-envs-per-worker |
1 | Vectorized envs per worker |
--max-grad-norm |
0.05 | Gradient clipping |
--max-traj-len |
400 | Episode horizon |
--eval-freq |
100 | Eval every N iterations |
--seed |
None | Random seed |
--device |
auto | Training device (auto/cpu/cuda) |
--no-mirror |
off | Disable symmetry wrapper (flag) |
--recurrent |
off | Use LSTM policy (flag) |
--continued |
None | Path to pretrained weights |
Instructions
- Determine the environment name from the user's request. If ambiguous, ask.
- Use
--logdir /tmp/training_runsunless the user specifies a different path. - Only include flags that differ from defaults — keep the command clean.
- Show the user the full command you're about to run.
- Run the command in the background using
run_in_background: trueon the Bash tool. Set a generous timeout (600000ms). - After launching, tell the user the logdir path and how to check progress (you can tail the output using the task ID).
- If the user asks to check on training, use
TaskOutputwithblock: falseto check the latest output.
Cartpole-Specific Defaults
For cartpole, these settings are known to work well with the current defaults
(--lr 3e-4 --max-grad-norm 0.5 --lam 0.95 --gamma 0.99):
--minibatch-size 256--std-dev 0.15 --learn-std --entropy-coeff 0.01--max-traj-len 500 --n-itr 500 --num-procs 12--no-mirror(cartpole has no body symmetry)
Suggest these defaults when the user trains cartpole, but let them override.
- 流狐分类
- AI 智能
- 作者声明 Agent
- 未找到明确声明;不据此推断已兼容或已测试
- 静态检查
- 94 / 100 · 启发式扫描,不代表运行安全
- 作者 / 版本 / 许可
- @rohanpsingh · BSD-2-Clause
- 流狐 Token 估算
- 低消耗
- 流狐接入估算
- 即装即用
- 是否需要外部 API Key
- 未发现要求
- 检测到的系统要求
- 未声明
- 底层运行要求
- Python
- 检测到的文件与系统行为
-
- 只读
- 允许写入 / 修改
- 检测到的网络行为
- 仅限本地
- 安装命令数
- 无(仅作为资料)
档案由构建时根据 SKILL.md 与安装命令自动衍生,可能与作者实际意图存在差异。
需要注意: 未限定 allowed-tools,默认拥有全部工具权限。
作者没有在当前 SKILL.md 中定义固定输出样例。 Command Template
Name · Description cartpole · Cartpole swing-up (simplest, good for testing) h1 · Unitree H1 standing task
Flag · Default · Description --n-itr · 20000 · Training iterations --lr · 1e-4 · Learning rate
Determine the environment name from the user's request. If ambiguous, ask. Use --logdir /tmp/trainingruns unless the user specifies a different path. Only include flags that differ from defaults — keep the command clean.
For cartpole, these settings are known to work well with the current defaults (--lr 3e-4 --max-grad-norm 0.5 --lam 0.95 --gamma 0.99): --minibatch-size 256
# /train — Launch a PPO Training Run
Parse the user's request from `$ARGUMENTS` and construct a training command.
## Command Template
```
RAY_ADDRESS= uv run python run_experiment.py train --env <ENV> --logdir <LOGDIR> [OPTIONS...]
```
## Available Environments
| Name | Description |
|------|-------------|
| `cartpole` | Cartpole swing-up (simplest, good for testing) |
| `h1` | Unitree H1 standing task |
| `jvrc_walk` | JVRC humanoid basic walking |
| `jvrc_step` | JVRC humanoid stepping with planned footsteps |
## Hyperparameters (defaults)
| Flag | Default | Description |
|------|---------|-------------|
| `--n-itr` | 20000 | Training iterations |
| `--lr` | 1e-4 | Learning rate |
| `--gamma` | 0.99 | Discount factor |
| `--std-dev` | 0.223 | Action noise |
| `--learn-std` | off | Learn action noise (flag) |
| `--entropy-coeff` | 0.0 | Entropy regularization |
| `--clip` | 0.2 | PPO clipping |
| `--minibatch-size` | 64 | Minibatch size |
| `--epochs` | 3 | Optimization epochs per update |
| `--num-procs` | 12 | Parallel workers |
| `--num-envs-per-worker` | 1 | Vectorized envs per worker |
| `--max-grad-norm` | 0.05 | Gradient clipping |
| `--max-traj-len` | 400 | Episode horizon |
| `--eval-freq` | 100 | Eval every N iterations |
| `--seed` | None | Random seed |
| `--device` | auto | Training device (auto/cpu/cuda) |
| `--no-mirror` | off | Disable symmetry wrapper (flag) |
| `--recurrent` | off | Use LSTM policy (flag) |
| `--continued` | None | Path to pretrained weights |
## Instructions
1. Determine the environment name from the user's request. If ambiguous, ask.
2. Use `--logdir /tmp/training_runs` unless the user specifies a different path.
3. Only include flags that differ from defaults — keep the command clean.
… 作者原文负责流程事实;流狐只索引当前章节、要点、文件与命令。
章节 -> Command Template → Available Environments → Hyperparameters (defaults) → Instructions → Cartpole-Specific Defaults
要点 -> Parse the user's request from $ARGUMENTS and construct a training command. · 1. Determine the environment name from the user's request. · Suggest these defaults when the user trains cartpole, but let them override.
文件/命令 -> $ARGUMENTS · cartpole · jvrcwalk · jvrcstep · --n-itr · --lr · --gamma · --std-dev
内容 SHA-256 -> 67a84e6e2fb7
原文结构
适用与边界
原文中的明确线索
$ARGUMENTS、cartpole、jvrcwalk、jvrcstep、--n-itr、--lr、--gamma、--std-dev