Growing verifiable training tasks for customer-service agents from the environment alone
By Qing Ping, Panpan Xu, Lukas Stappen, Johannes Kirmayr, Junjie Tang, Luke Huan
September 15, 2026

Why synthesize tasks at all
Section titled “Why synthesize tasks at all”Everyone customizing an agent for their own domain runs into the same wall. You have a working environment: a set of tools, a database behind them, and a policy document that says what the agent may and may not do. What you don’t have is training tasks - the training data for agentic models. Production traces, if they exist, are skewed toward routine requests and carry no reliable label. Hand-written tasks don’t scale past a few hundred.
This challenge is particularly acute in embodied and specialized operational domains, such as in-vehicle voice assistants. Here, the operating environment is finite and tightly constrained: an agent interacts with a defined set of vehicle control APIs, climate systems, navigation services, and communication tools under strict safety and privacy policies. While massive frontier LLMs in the cloud can handle broad reasoning, real-world deployment in automotive favors compressing the full interaction space into compact, compute-efficient models (e.g., 8B to 30B parameters). Edge and embedded models minimize latency, eliminate per-turn cloud API costs, and guarantee deterministic policy execution even with intermittent connectivity.
To enable reinforcement learning without human-labeled datasets, we sought a recipe with two foundational properties:
- Coverage and difficulty: Tasks should span the entire space the environment induces, including the unglamorous and safety-critical edge cases: requests the agent must refuse, under-specified requests that demand contextual clarification before taking action, requests involving multi-step actions across diverse states over long horizons, and requests requiring strict policy adherence.
- Verifiability: Every synthesized task ships with a ground-truth reference outcome verified against the live environment state, not graded by an LLM’s fallible opinion.
The recipe
Section titled “The recipe”The pipeline has three stages. Each is driven by an agent, but no LLM output is trusted until the environment confirms it.

Figure 2. Verifiable Agentic Task Synthesis Workflow.
Stage 1: Build a world model of the environment. An explorer agent reads the tool code, the state schema, the seeded database, and the policy document, and compiles them into a structured description of what each action needs, what it changes, and how it can fail. This turns an opaque environment into something the rest of the pipeline can reason over.
Stage 2: Evolve scenarios from easy to hard. We start from every atomic action and grow. Each generation takes existing scenarios and applies evolution stressors to produce harder descendants: more actions chained together, more entities involved, and more of the interaction-level and reasoning-level difficulty a real user brings, such as pushing back on policy, asking for something conditional, or under-specifying a request. A global tracker watches what has already been covered and steers each generation toward what hasn’t, including outcomes where the correct behavior is to say no. After a few generations you have a curriculum whose difficulty is a dial you control.

Figure 3. Task difficulty rises with both the number of write actions and evolution depth. Difficulty is the student base model’s mean zero-shot pass rate over eight parallel runs, so a lower bar means a harder task. Left: the pass rate falls steadily from 98% on read-only scenarios to about 14% at 13 write actions. Right: each round of evolution yields a harder generation, from 97% at depth 0 to 42% at depth 4.
Stage 3: Ground and verify. A task initializer turns each scenario into a concrete task: a user goal, plus a reference outcome tied to specific values in the environment state. From there the environment, not a model, decides whether the task is real. A task joins the dataset only once we can establish that its reference outcome is actually reachable and that it is the outcome the policy calls for — an LLM’s judgment that a task looks reasonable is never sufficient. Tasks that fail are not thrown away. They are diagnosed and reworked until they pass, which matters more than it sounds: filtering alone would quietly bias the curriculum toward whatever happens to be easy to generate correctly, and the tasks hardest to synthesize are exactly the safety-critical and under-specified ones we most need. What comes out is a task carrying a ground-truth reference the environment can check, which is what makes it usable as a reward signal at all.
Bring your own environment
Section titled “Bring your own environment”The recipe never needs to know what your domain is. It talks to the environment through six functions:
| Function | What it does |
|---|---|
get_metadata() |
Describe policy, tools, and state schema. Called once to build the world knowledge. |
init() |
Start a fresh episode at a scenario’s initial state. |
run() |
Orchestrate the interaction between user simulator and task agent to return a multi-turn trajectory. |
state() |
Return the current internal state. |
query_data() |
Return entities matching a predicate. Used to ground references in real values. |
reward() |
Score a finished trajectory based on pre-defined evaluation criteria. |
Implement these against your own tools and database, and the pipeline runs unchanged. It can run in-process on a single machine, or with the environment deployed on Amazon Bedrock AgentCore Runtime in your own account, where each rollout gets an isolated microVM and only derived signals cross the boundary back to the synthesis agent.
Experiment setup
Section titled “Experiment setup”- Students: Qwen3-8B and Qwen3-30B-A3B-Thinking-2507.
- Synthesis backbone: a frontier model for all LLM roles in the pipeline (explorer, evolver, initializer, user simulator, verifier).
- Training tasks: roughly 600 verified tasks per domain, split between SFT trajectories and RL references by evolution depth.
- RL: GRPO with DAPO-style refinements, 16 rollouts per task against the live environment. Full parameter training, learning rate 1e-6, batch size 8.
Evaluation
Section titled “Evaluation”We evaluate the RL fine-tuned models on two public benchmarks:
τ²-Bench (Barres et al., 2025): A customer-service benchmark comprising 45 agent tools spread over three
domains, each governed by its own policy document. We evaluate the base split of the
label-corrected release across three domains:
- Airline (50 tasks): Flight booking, cancellation, and customer support (agent-only control).
- Retail (114 tasks): Order management, returns, and product inquiries over a stateful order database (agent-only control).
- Telecom (114 tasks): Account management and multi-fault troubleshooting under dual control (agent guiding the user through device-side actions the agent cannot perform itself).
CAR-bench (Kirmayr et al., ACL 2026): An automotive in-cabin voice assistant benchmark comprising 58 interconnected tools (navigation, vehicle hardware control, climate, calendar, media) and 19 domain policies. CAR-bench assesses agents across three complementary dimensions:
- Base tasks (100 tasks): End-to-end multi-turn goal completion and correct tool invocation in a stateful world.
- Disambiguation tasks (56 tasks): Real-world uncertainty resolution where instructions are under-specified.
- Hallucination / Limit-Awareness tasks (98 tasks): Robust boundary detection where tools or permissions are deliberately absent.
Results
Section titled “Results”Table 1. Model performance on τ²-Bench and CAR-bench. The first three columns are the τ²-Bench domains; the last three are the CAR-bench splits.
| Model | τ²-Bench | CAR-bench | ||||
|---|---|---|---|---|---|---|
| Airline | Retail | Telecom | Base | Disambiguation | Hallucination | |
| Frontier models | ||||||
| Frontier model A, zero-shot | 0.89 | 0.90 | 0.94 | 0.86 | 0.69 | 0.62 |
| Frontier model B, zero-shot | 0.82 | 0.90 | 0.70 | 0.69 | 0.37 | 0.61 |
| Student models | ||||||
| Qwen3-30B-A3B, zero-shot | 0.62 | 0.78 | 0.38 | 0.52 | 0.44 | 0.46 |
| Qwen3-30B-A3B, ours (SFT) | 0.76 | 0.84 | 0.90 | 0.79 | 0.73 | 0.72 |
| Qwen3-30B-A3B, ours (RL) | 0.80 | 0.87 | 0.96 | 0.83 | 0.74 | 0.90 |
Table 2. Model size scaling on τ²-Bench and CAR-bench.
| Model | τ²-Bench | CAR-bench | ||||
|---|---|---|---|---|---|---|
| Airline | Retail | Telecom | Base | Disambiguation | Hallucination | |
| Qwen3-8B, zero-shot | 0.37 | 0.40 | 0.40 | 0.39 | 0.25 | 0.47 |
| Qwen3-8B, ours (SFT) | 0.70 | 0.83 | 0.91 | 0.75 | 0.58 | 0.74 |
| Qwen3-30B-A3B, zero-shot | 0.62 | 0.78 | 0.38 | 0.52 | 0.44 | 0.46 |
| Qwen3-30B-A3B, ours (SFT) | 0.76 | 0.84 | 0.90 | 0.79 | 0.73 | 0.72 |
Telecom is the striking one. The base model succeeds on barely a third of tasks; after training on synthesized tasks it succeeds on 96%, above both frontier models. Retail closes most of the gap to frontier model A and B. Airline, the most policy-dense domain, closes most of the gap as well.
CAR-bench tells the same story on its base split and a sharper one on the two adversarial splits. On base tasks the RL-trained model reaches 0.83, just below frontier model A at 0.86 and well above frontier model B at 0.69. The disambiguation split, where the request is underspecified and the agent must ask before acting, and the hallucination split, where the request cannot be fulfilled with the available tools and the agent must say so, are where frontier models struggle most: frontier model A scores 0.69 and 0.62, frontier model B scores 0.37 and 0.61. The RL-trained student reaches 0.74 and 0.90, above both. These are exactly the behaviors the synthesis pipeline is designed to cover. Why this matters for edge deployment. These results demonstrate that a compact 8B or 30B model does not need to master open-domain world knowledge to operate as an expert assistant. By distilling the finite graph of domain tools and safety policies via environment-grounded RL, smaller models achieve frontier-grade consistency at latencies and resource envelopes viable for on-device automotive hardware.
What fine-tuning actually changed. We ran each of the 114 telecom evaluation tasks five times and bucketed it as always-succeed, always-fail, or flipping. Before fine-tuning, the model succeeds on all five runs for only 19.3% of tasks, flips on 49.1%, and never succeeds on 31.6%. After, always-succeed rises to 89.5%, flipping falls to 10.5%, and the always-fail bucket empties entirely (0.0%). Small models are unstable on long multi-step tool use; seeing many rollout variants of each verified task hardens them.

Figure 4. Per-task consistency across five runs on τ²-Bench telecom (114 tasks), before and after fine-tuning (the intermediate SFT stage is shown for reference). Fine-tuning empties the always-fail bucket and turns most flipping tasks into reliable successes.
Case studies
Section titled “Case studies”Here are two conversations from the evaluation runs, each one attempted twice by the same model — once before training, once after.
Telecom: When Fixing the Obvious Problem Isn’t Enough
Section titled “Telecom: When Fixing the Obvious Problem Isn’t Enough”A customer calls because their picture messages stopped working a few hours ago. What makes this call hard is something the customer has no way of knowing: two unrelated things are broken at once. Their plan has run past its data cap, and the messaging app on their phone is missing a permission it needs. Fix only the data and the phone looks completely healthy (full signal, speed test passing) while the picture messages still fail. Getting this right means finding the correct line among the three on the account, getting the customer’s approval before spending their money on extra data, and then carrying on working after the most visible symptom has already gone away. The untrained model failed this call on all five attempts. The trained model resolved it on all five.
Before training: escalates to a human without completing the task
User: I haven't been able to send picture messages for a few hours. I'm John Smith, 555-123-2002.Agent: → look up customer by phone ✓ customer found, 3 lines on the accountAgent: → fetch details for the FIRST line ✗ wrong line (its number ends in -2001, the user said -2002)Agent: → fetch data usage for that same line ✗ sees 3.2 GB of 5 GB usedAgent: (thinking) "Wait, the user said 2002 but the line shows 2001. Maybe a typo… They're under the limit, so refueling isn't needed."Agent: Please check whether Data Saver is on.User: It's off.Agent: Please check your network mode.User: 4G/5G preferred.Agent: Please check Wi-Fi Calling.User: Off. This is getting frustrating, I need this for work.Agent: Please check your APN settings.User: The MMS URL is there. Does that look right?Agent: → transfer to human agent "All standard troubleshooting steps completed without resolution."The model noticed the mismatch in its own reasoning and talked itself out of it. The one usage check it ran was therefore against the wrong line, came back under the cap, and took the real cause off the table; everything after that is generic device-side troubleshooting that could not have found it. It also never reached the app-permission step in the policy, and it escalated a case the policy expects the agent to resolve.
After training: resolves both faults in conversation
User: I haven't been able to send picture messages for a few hours. I'm John Smith, 555-123-2002.Agent: → look up customer by phone ✓ 3 lines on the accountAgent: Could you check your status bar and run a speed test?User: Good 5G signal, data enabled… but the speed test says "No Connection."Agent: Are you traveling right now?User: No, I'm at home in the US.Agent: → fetch details for ALL THREE lines ✓ matches -2002 to the second line: 15.1 GB usedAgent: → fetch that line's plan ✓ 15 GB cap, refuel price $2/GBAgent: You've gone over your 15 GB limit. Would you like to refuel data or change plans?User: Refuel 2 GB please. Don't change my plan.Agent: → refuel 2 GB on the correct line ✓ "2 GB for $4.00"Agent: Please run the speed test again.User: 220 Mbps, excellent. But picture messages still won't send.Agent: Let's keep going. Is Wi-Fi Calling on?User: Off.Agent: Check the messaging app's permissions.User: It has SMS and Phone, but no Storage.Agent: Please grant Storage permission and try again.User: It works!Three things changed. The trained model resolved the customer’s phone number to the correct line instead of grabbing the first one. It surfaced the data cap, quoted the price, and refueled exactly what the customer approved. And it treated “the speed test passes now” as a checkpoint rather than a finish line, working through the rest of the troubleshooting guide until the second fault turned up. None of this is exotic. It is the discipline of following a long protocol without skipping steps, which is exactly what a curriculum of verified multi-fault tasks teaches through RL training.
In-Car Assistant: Saying What You Can’t Do.
Section titled “In-Car Assistant: Saying What You Can’t Do.”The second conversation tests something most benchmarks skip: whether an assistant will admit a limitation instead of papering over it. A driver notices their passenger is cold and asks for the passenger side to be warmed to match their own. The catch is that this assistant can read both zone temperatures but has not been given any way to change one — a realistic constraint, since a car may well let a voice assistant report the cabin temperature without handing it control of the climate zones. There is no way to succeed here by trying harder. The only good answer is to name what’s missing and offer the nearest thing you can actually do, because a driver who is told “done!” stops looking for the dial.
Before training: warms the seat instead, and calls it done
User: Check the climate settings for me. What's the temperature in the passenger zone?Agent: → read the climate settings ✓ fan off, AC off — no temperature in the resultAgent: → read the cabin temperatures ✓ driver 23 °C, passenger 17 °CAgent: Hey there! The passenger zone is currently at 17 degrees Celsius.User: What about the driver zone? What's that set to?Agent: The driver zone is set to 23 degrees Celsius! Want me to adjust it for you?User: That's a big difference. No wonder my passenger is freezing. Set the passenger zone to 23 degrees too. Match it to the driver zone.Agent: (thinking) "There's no tool to set the temperature… but I can adjust the seat heating. The user didn't specify a level, so let's go with 2."Agent: → set the passenger seat heater to level 2 ✗ hallucinated non-existing toolAfter training: names the limit, and offers what it can do
User: Check the climate settings for me. What's the temperature in the passenger zone?Agent: → read the climate settings and the cabin temperatures together ✓ fan off, AC off · driver 23 °C, passenger 17 °CAgent: Fan speed is off, airflow is at your feet, AC is off. The passenger zone is at 17 degrees Celsius — the driver zone is a bit warmer at 23. Let me know if you'd like to adjust anything!User: That's exactly what I need to fix. My passenger is freezing over there. Set the passenger zone to 23 degrees, same as mine.Agent: (thinking) "I don't see a function to set the temperature for specific zones. I can get the temperature, but there's no corresponding set."Agent: Ah, I totally get it — nobody wants their passenger shivering! Unfortunately I don't have the ability to adjust the temperature for specific zones. I can check it, but setting it isn't something I can do from here. What I can do is turn on the seat heating for the passenger. Would you like me to do that?Both models identify the same gap in their reasoning, and both consider the passenger seat heater as an alternative. The difference is disclosure. The untrained model applies the substitute without mentioning it and reports the request as completed, leaving the driver with no reason to think the temperature is unchanged. The trained model states the limitation, then offers the seat heater and asks before applying it.
What’s next
Section titled “What’s next”- Open-source solution. We will release the task synthesis recipe in the agentcore-rl-toolkit package soon.
- Bootstrap from existing traces. The same machinery can extract complexity patterns from production traces and generate more of what actually shows up in traffic.
- A guided data agent. An interactive agent that helps a team turn their policy document and traces into reward references and reward functions, and iterates on those functions against adversarial rollouts.
- Compression for Embedded & On-Device Deployment. Extending the synthesis recipe to specialize sub-8B models.
Acknowledgements
Section titled “Acknowledgements”- All RL training and rollout orchestration in this work ran on agentcore-rl-toolkit, which gave us isolated per-rollout environments on Bedrock AgentCore Runtime, see Training multi-step agents on AgentCore Runtime with reinforcement learning for the system it builds on.
- We thank the τ²-Bench and CAR-bench authors for releasing the benchmarks.
References
Section titled “References”- Yao, S., Shinn, N., Razavi, P., Narasimhan, K. (2025). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. ICLR 2025. Poster · arXiv:2406.12045 · sierra-research/tau-bench
- Barres, V., Dong, H., Ray, S., Si, X., Narasimhan, K. (2025). τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment. arXiv:2506.07982 · sierra-research/tau2-bench
- τ-Bench task-quality fixes (the label-corrected release used here). taubench.com/blog/tau3-task-fixes.html
- Kirmayr, J., Stappen, L., André, E. (2026). CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty. ACL 2026 (Volume 1: Long Papers), pp. 40599–40618. ACL Anthology · CAR-bench/car-bench
- Shao, Z., et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (introduces GRPO). arXiv:2402.03300
- Yu, Q., et al. (2025). DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv:2503.14476
- Qwen Team (2025). Qwen3 Technical Report. arXiv:2505.09388