verl backend setup
The direct verl backend integrates AgentCore rollouts as a custom
AgentCoreAgentLoop through verl’s v1 agent-loop API. verl’s built-in
trainer.v1.trainer_mode=sync works when every rollout is guaranteed to produce
exactly one training row. Use agentcore_sync when a rollout may produce
multiple training rows; it preserves the configured optimizer-step schedule
while training every row. The in-repo rollout gateway captures token IDs, log
probabilities, and loss masks from multi-turn agent calls and converts each
trajectory-tree leaf into a verl training row.
The example training scripts have been run end to end with:
- FSDP full fine-tuning of Qwen3-4B on GSM8K.
- Megatron + LoRA fine-tuning of Qwen3-Coder-30B-A3B on MigrationBench.
Prerequisites
Section titled “Prerequisites”- A GPU cluster supported by the pinned CUDA 13 stack.
- AWS credentials that can invoke the AgentCore Runtime and read/write the result S3 bucket.
- A deployed
AgentCoreRLAppwhose OpenAI-compatible model client forwardspayload["_rollout"]["api_key"]. - Network routing from the AgentCore containers to the trainer nodes’ rollout gateway ports.
The backend is pinned to a tested verl commit through tool.uv.sources, so install
it from a checkout of this repository:
uv sync --extra verlThe full stack uses CUDA 13 wheels and requires driver 580.65.06 or newer and a GPU with compute capability 7.5 or newer.
For Megatron, use Python 3.12 and install the additional dependency group:
uv venv --python 3.12uv sync --extra verl --group verl-megatronAdapt the agent
Section titled “Adapt the agent”The trainer generates one capture-session key per rollout and supplies it as
_rollout.api_key. Pass that value to the model client:
@app.rollout_entrypointdef invoke_agent(payload: dict, context): rollout_config = payload["_rollout"] api_key = rollout_config.get("api_key") or "EMPTY" model = OpenAIModel( client_args={ "api_key": api_key, "base_url": rollout_config["base_url"], }, model_id=rollout_config["model_id"], params=rollout_config.get("sampling_params", {}), ) # Run the agent and return {"rewards": score}.The fallback "EMPTY" keeps local evaluation and unauthenticated inference
endpoints working.
Prepare data
Section titled “Prepare data”Each parquet row must contain a payload column holding the exact JSON object the
agent expects:
{"payload": {"prompt": "Natalia sold clips to...", "answer": "72"}}Configure verl to use PayloadDataset:
data: custom_cls: path: pkg://agentcore_rl_toolkit.backends.verl.dataset name: PayloadDatasetPayloadDataset synthesizes the chat-format prompt column required by verl from
payload["prompt"]. If the agent uses another field name, set
+data.payload_prompt_field=<field>. An explicit chat-format prompt column takes
precedence when present.
Run the GSM8K recipe
Section titled “Run the GSM8K recipe”cd src/agentcore_rl_toolkit/backends/verl/examples/math_agentexport AGENT_RUNTIME_ARN=arn:aws:bedrock-agentcore:...:runtime/...export ACR_S3_BUCKET=your-results-bucketpython preprocess_gsm8k.py --output-dir gsm8k./fsdp_fft_sync_grpo.shThe shell script configures verl and accepts additional Hydra overrides through its trailing arguments:
./fsdp_fft_sync_grpo.sh trainer.logger='["console","wandb"]'The accompanying agentcore_agent.yaml contains the AgentCore runtime ARN, result
bucket, per-turn token limit, timeout, and gateway settings. Values that vary by run
use OmegaConf environment interpolation.
Advanced example: MigrationBench with Megatron
Section titled “Advanced example: MigrationBench with Megatron”MigrationBench adds Megatron + LoRA, a Python 3.12 environment, and a two-stage
data-preparation flow that packages the source repositories in S3. Follow the
migration_agent example
for the complete setup and training commands.
Configuration constraints
Section titled “Configuration constraints”trainer.use_v1=trueis required becauseAgentCoreAgentLoop.run()returns a list with one row per trajectory-tree leaf. Usetrainer.v1.trainer_mode=agentcore_syncwhenever a rollout may emit multiple leaves.agentcore_syncsupports actor-only synchronous training with distillation disabled,parameter_sync_step=1, andactor_rollout_ref.actor.loss_agg_mode=seq-mean-token-sum.actor_rollout_ref.rollout.max_model_lenis the inference model’s context capacity and must be set explicitly.actor_rollout_ref.rollout.response_lengthsets the cumulative trajectory budget, whilemax_tokens_per_turninagentcore_agent.yamllimits each model call.- The only supported reward mode is agent-side scoring: the app returns
{"rewards": score}. Trainer-side reward functions are not yet supported.
For the full token-budget model, failure behavior, gateway networking options, and
troubleshooting notes, see the
backends/verl README.