Blog
Notes, deep dives, and experimental results from training agents with reinforcement learning on Bedrock AgentCore Runtime.
- Training multi-step agents on AgentCore Runtime with reinforcement learning — an end-to-end system for online RL on multi-step agents, with results from training on two long-horizon benchmarks: MigrationBench (Java code-migration) and OfficeBench (office automation).
- Growing verifiable training data for customer-service agents from the environment alone — a data synthesis recipe that grows a curriculum of verifiable RL training tasks from an agent’s environment alone, with results on the three τ²-Bench customer-service domains and CAR-bench.
- Stable Updates, Fewer Sequences: Better and Faster RL in ART — stable optimizer-step batching in verl and linear-history healing in the rollout gateway, with quality and efficiency results on MigrationBench.