Skip to content

Growing verifiable training tasks for customer-service agents from the environment alone

By Qing Ping, Panpan Xu, Lukas Stappen, Johannes Kirmayr, Junjie Tang, Luke Huan

September 15, 2026

Figure 1 — Cost vs. Performance on the three τ²-Bench domains and CAR-bench.
Figure 1. Cost vs. Performance on the τ²-Bench and CAR-bench. RL-trained compact models match or outperform proprietary frontier models at a fraction of inference cost and latency.

Everyone customizing an agent for their own domain runs into the same wall. You have a working environment: a set of tools, a database behind them, and a policy document that says what the agent may and may not do. What you don’t have is training tasks - the training data for agentic models. Production traces, if they exist, are skewed toward routine requests and carry no reliable label. Hand-written tasks don’t scale past a few hundred.

This challenge is particularly acute in embodied and specialized operational domains, such as in-vehicle voice assistants. Here, the operating environment is finite and tightly constrained: an agent interacts with a defined set of vehicle control APIs, climate systems, navigation services, and communication tools under strict safety and privacy policies. While massive frontier LLMs in the cloud can handle broad reasoning, real-world deployment in automotive favors compressing the full interaction space into compact, compute-efficient models (e.g., 8B to 30B parameters). Edge and embedded models minimize latency, eliminate per-turn cloud API costs, and guarantee deterministic policy execution even with intermittent connectivity.

To enable reinforcement learning without human-labeled datasets, we sought a recipe with two foundational properties:

  • Coverage and difficulty: Tasks should span the entire space the environment induces, including the unglamorous and safety-critical edge cases: requests the agent must refuse, under-specified requests that demand contextual clarification before taking action, requests involving multi-step actions across diverse states over long horizons, and requests requiring strict policy adherence.
  • Verifiability: Every synthesized task ships with a ground-truth reference outcome verified against the live environment state, not graded by an LLM’s fallible opinion.

The pipeline has three stages. Each is driven by an agent, but no LLM output is trusted until the environment confirms it.

System workflow.

Figure 2. Verifiable Agentic Task Synthesis Workflow.

Stage 1: Build a world model of the environment. An explorer agent reads the tool code, the state schema, the seeded database, and the policy document, and compiles them into a structured description of what each action needs, what it changes, and how it can fail. This turns an opaque environment into something the rest of the pipeline can reason over.

Stage 2: Evolve scenarios from easy to hard. We start from every atomic action and grow. Each generation takes existing scenarios and applies evolution stressors to produce harder descendants: more actions chained together, more entities involved, and more of the interaction-level and reasoning-level difficulty a real user brings, such as pushing back on policy, asking for something conditional, or under-specifying a request. A global tracker watches what has already been covered and steers each generation toward what hasn’t, including outcomes where the correct behavior is to say no. After a few generations you have a curriculum whose difficulty is a dial you control.

Correlations between Task Difficulty versus Number of Write Actions and Evolution Depth.

Figure 3. Task difficulty rises with both the number of write actions and evolution depth. Difficulty is the student base model’s mean zero-shot pass rate over eight parallel runs, so a lower bar means a harder task. Left: the pass rate falls steadily from 98% on read-only scenarios to about 14% at 13 write actions. Right: each round of evolution yields a harder generation, from 97% at depth 0 to 42% at depth 4.

Stage 3: Ground and verify. A task initializer turns each scenario into a concrete task: a user goal, plus a reference outcome tied to specific values in the environment state. From there the environment, not a model, decides whether the task is real. A task joins the dataset only once we can establish that its reference outcome is actually reachable and that it is the outcome the policy calls for — an LLM’s judgment that a task looks reasonable is never sufficient. Tasks that fail are not thrown away. They are diagnosed and reworked until they pass, which matters more than it sounds: filtering alone would quietly bias the curriculum toward whatever happens to be easy to generate correctly, and the tasks hardest to synthesize are exactly the safety-critical and under-specified ones we most need. What comes out is a task carrying a ground-truth reference the environment can check, which is what makes it usable as a reward signal at all.

The recipe never needs to know what your domain is. It talks to the environment through six functions:

Function What it does
get_metadata() Describe policy, tools, and state schema. Called once to build the world knowledge.
init() Start a fresh episode at a scenario’s initial state.
run() Orchestrate the interaction between user simulator and task agent to return a multi-turn trajectory.
state() Return the current internal state.
query_data() Return entities matching a predicate. Used to ground references in real values.
reward() Score a finished trajectory based on pre-defined evaluation criteria.

Implement these against your own tools and database, and the pipeline runs unchanged. It can run in-process on a single machine, or with the environment deployed on Amazon Bedrock AgentCore Runtime in your own account, where each rollout gets an isolated microVM and only derived signals cross the boundary back to the synthesis agent.

  • Students: Qwen3-8B and Qwen3-30B-A3B-Thinking-2507.
  • Synthesis backbone: a frontier model for all LLM roles in the pipeline (explorer, evolver, initializer, user simulator, verifier).
  • Training tasks: roughly 600 verified tasks per domain, split between SFT trajectories and RL references by evolution depth.
  • RL: GRPO with DAPO-style refinements, 16 rollouts per task against the live environment. Full parameter training, learning rate 1e-6, batch size 8.

We evaluate the RL fine-tuned models on two public benchmarks:

τ²-Bench (Barres et al., 2025): A customer-service benchmark comprising 45 agent tools spread over three domains, each governed by its own policy document. We evaluate the base split of the label-corrected release across three domains:

  • Airline (50 tasks): Flight booking, cancellation, and customer support (agent-only control).
  • Retail (114 tasks): Order management, returns, and product inquiries over a stateful order database (agent-only control).
  • Telecom (114 tasks): Account management and multi-fault troubleshooting under dual control (agent guiding the user through device-side actions the agent cannot perform itself).

CAR-bench (Kirmayr et al., ACL 2026): An automotive in-cabin voice assistant benchmark comprising 58 interconnected tools (navigation, vehicle hardware control, climate, calendar, media) and 19 domain policies. CAR-bench assesses agents across three complementary dimensions:

  • Base tasks (100 tasks): End-to-end multi-turn goal completion and correct tool invocation in a stateful world.
  • Disambiguation tasks (56 tasks): Real-world uncertainty resolution where instructions are under-specified.
  • Hallucination / Limit-Awareness tasks (98 tasks): Robust boundary detection where tools or permissions are deliberately absent.

Table 1. Model performance on τ²-Bench and CAR-bench. The first three columns are the τ²-Bench domains; the last three are the CAR-bench splits.

Modelτ²-BenchCAR-bench
AirlineRetailTelecomBaseDisambiguationHallucination
Frontier models
Frontier model A, zero-shot0.890.900.940.860.690.62
Frontier model B, zero-shot0.820.900.700.690.370.61
Student models
Qwen3-30B-A3B, zero-shot0.620.780.380.520.440.46
Qwen3-30B-A3B, ours (SFT)0.760.840.900.790.730.72
Qwen3-30B-A3B, ours (RL)0.800.870.960.830.740.90

Table 2. Model size scaling on τ²-Bench and CAR-bench.

Modelτ²-BenchCAR-bench
AirlineRetailTelecomBaseDisambiguationHallucination
Qwen3-8B, zero-shot0.370.400.400.390.250.47
Qwen3-8B, ours (SFT)0.700.830.910.750.580.74
Qwen3-30B-A3B, zero-shot0.620.780.380.520.440.46
Qwen3-30B-A3B, ours (SFT)0.760.840.900.790.730.72

Telecom is the striking one. The base model succeeds on barely a third of tasks; after training on synthesized tasks it succeeds on 96%, above both frontier models. Retail closes most of the gap to frontier model A and B. Airline, the most policy-dense domain, closes most of the gap as well.

CAR-bench tells the same story on its base split and a sharper one on the two adversarial splits. On base tasks the RL-trained model reaches 0.83, just below frontier model A at 0.86 and well above frontier model B at 0.69. The disambiguation split, where the request is underspecified and the agent must ask before acting, and the hallucination split, where the request cannot be fulfilled with the available tools and the agent must say so, are where frontier models struggle most: frontier model A scores 0.69 and 0.62, frontier model B scores 0.37 and 0.61. The RL-trained student reaches 0.74 and 0.90, above both. These are exactly the behaviors the synthesis pipeline is designed to cover. Why this matters for edge deployment. These results demonstrate that a compact 8B or 30B model does not need to master open-domain world knowledge to operate as an expert assistant. By distilling the finite graph of domain tools and safety policies via environment-grounded RL, smaller models achieve frontier-grade consistency at latencies and resource envelopes viable for on-device automotive hardware.

What fine-tuning actually changed. We ran each of the 114 telecom evaluation tasks five times and bucketed it as always-succeed, always-fail, or flipping. Before fine-tuning, the model succeeds on all five runs for only 19.3% of tasks, flips on 49.1%, and never succeeds on 31.6%. After, always-succeed rises to 89.5%, flipping falls to 10.5%, and the always-fail bucket empties entirely (0.0%). Small models are unstable on long multi-step tool use; seeing many rollout variants of each verified task hardens them.

Per-task consistency across 5 runs on τ²-Bench telecom for zero-shot, SFT, and RL: always-succeed rises 19.3% → 75.4% → 89.5%, flipping falls 49.1% → 23.7% → 10.5%, and always-fail falls 31.6% → 0.9% → 0.0%.

Figure 4. Per-task consistency across five runs on τ²-Bench telecom (114 tasks), before and after fine-tuning (the intermediate SFT stage is shown for reference). Fine-tuning empties the always-fail bucket and turns most flipping tasks into reliable successes.

Here are two conversations from the evaluation runs, each one attempted twice by the same model — once before training, once after.

Telecom: When Fixing the Obvious Problem Isn’t Enough

Section titled “Telecom: When Fixing the Obvious Problem Isn’t Enough”

A customer calls because their picture messages stopped working a few hours ago. What makes this call hard is something the customer has no way of knowing: two unrelated things are broken at once. Their plan has run past its data cap, and the messaging app on their phone is missing a permission it needs. Fix only the data and the phone looks completely healthy (full signal, speed test passing) while the picture messages still fail. Getting this right means finding the correct line among the three on the account, getting the customer’s approval before spending their money on extra data, and then carrying on working after the most visible symptom has already gone away. The untrained model failed this call on all five attempts. The trained model resolved it on all five.

Before training: escalates to a human without completing the task

User: I haven't been able to send picture messages for a few hours. I'm John Smith, 555-123-2002.
Agent: → look up customer by phone ✓ customer found, 3 lines on the account
Agent: → fetch details for the FIRST line ✗ wrong line (its number ends in -2001, the user said -2002)
Agent: → fetch data usage for that same line ✗ sees 3.2 GB of 5 GB used
Agent: (thinking) "Wait, the user said 2002 but the line shows 2001. Maybe a typo…
They're under the limit, so refueling isn't needed."
Agent: Please check whether Data Saver is on.
User: It's off.
Agent: Please check your network mode.
User: 4G/5G preferred.
Agent: Please check Wi-Fi Calling.
User: Off. This is getting frustrating, I need this for work.
Agent: Please check your APN settings.
User: The MMS URL is there. Does that look right?
Agent: → transfer to human agent
"All standard troubleshooting steps completed without resolution."

The model noticed the mismatch in its own reasoning and talked itself out of it. The one usage check it ran was therefore against the wrong line, came back under the cap, and took the real cause off the table; everything after that is generic device-side troubleshooting that could not have found it. It also never reached the app-permission step in the policy, and it escalated a case the policy expects the agent to resolve.

After training: resolves both faults in conversation

User: I haven't been able to send picture messages for a few hours. I'm John Smith, 555-123-2002.
Agent: → look up customer by phone ✓ 3 lines on the account
Agent: Could you check your status bar and run a speed test?
User: Good 5G signal, data enabled… but the speed test says "No Connection."
Agent: Are you traveling right now?
User: No, I'm at home in the US.
Agent: → fetch details for ALL THREE lines ✓ matches -2002 to the second line: 15.1 GB used
Agent: → fetch that line's plan ✓ 15 GB cap, refuel price $2/GB
Agent: You've gone over your 15 GB limit. Would you like to refuel data or change plans?
User: Refuel 2 GB please. Don't change my plan.
Agent: → refuel 2 GB on the correct line ✓ "2 GB for $4.00"
Agent: Please run the speed test again.
User: 220 Mbps, excellent. But picture messages still won't send.
Agent: Let's keep going. Is Wi-Fi Calling on?
User: Off.
Agent: Check the messaging app's permissions.
User: It has SMS and Phone, but no Storage.
Agent: Please grant Storage permission and try again.
User: It works!

Three things changed. The trained model resolved the customer’s phone number to the correct line instead of grabbing the first one. It surfaced the data cap, quoted the price, and refueled exactly what the customer approved. And it treated “the speed test passes now” as a checkpoint rather than a finish line, working through the rest of the troubleshooting guide until the second fault turned up. None of this is exotic. It is the discipline of following a long protocol without skipping steps, which is exactly what a curriculum of verified multi-fault tasks teaches through RL training.

In-Car Assistant: Saying What You Can’t Do.

Section titled “In-Car Assistant: Saying What You Can’t Do.”

The second conversation tests something most benchmarks skip: whether an assistant will admit a limitation instead of papering over it. A driver notices their passenger is cold and asks for the passenger side to be warmed to match their own. The catch is that this assistant can read both zone temperatures but has not been given any way to change one — a realistic constraint, since a car may well let a voice assistant report the cabin temperature without handing it control of the climate zones. There is no way to succeed here by trying harder. The only good answer is to name what’s missing and offer the nearest thing you can actually do, because a driver who is told “done!” stops looking for the dial.

Before training: warms the seat instead, and calls it done

User: Check the climate settings for me. What's the temperature in the passenger zone?
Agent: → read the climate settings ✓ fan off, AC off — no temperature in the result
Agent: → read the cabin temperatures ✓ driver 23 °C, passenger 17 °C
Agent: Hey there! The passenger zone is currently at 17 degrees Celsius.
User: What about the driver zone? What's that set to?
Agent: The driver zone is set to 23 degrees Celsius! Want me to adjust it for you?
User: That's a big difference. No wonder my passenger is freezing. Set the passenger
zone to 23 degrees too. Match it to the driver zone.
Agent: (thinking) "There's no tool to set the temperature… but I can adjust the seat
heating. The user didn't specify a level, so let's go with 2."
Agent: → set the passenger seat heater to level 2 ✗ hallucinated non-existing tool

After training: names the limit, and offers what it can do

User: Check the climate settings for me. What's the temperature in the passenger zone?
Agent: → read the climate settings and the cabin temperatures together
✓ fan off, AC off · driver 23 °C, passenger 17 °C
Agent: Fan speed is off, airflow is at your feet, AC is off. The passenger zone is at
17 degrees Celsius — the driver zone is a bit warmer at 23. Let me know if you'd
like to adjust anything!
User: That's exactly what I need to fix. My passenger is freezing over there. Set the
passenger zone to 23 degrees, same as mine.
Agent: (thinking) "I don't see a function to set the temperature for specific zones. I
can get the temperature, but there's no corresponding set."
Agent: Ah, I totally get it — nobody wants their passenger shivering! Unfortunately I
don't have the ability to adjust the temperature for specific zones. I can check
it, but setting it isn't something I can do from here. What I can do is turn on
the seat heating for the passenger. Would you like me to do that?

Both models identify the same gap in their reasoning, and both consider the passenger seat heater as an alternative. The difference is disclosure. The untrained model applies the substitute without mentioning it and reports the request as completed, leaving the driver with no reason to think the temperature is unchanged. The trained model states the limitation, then offers the seat heater and asks before applying it.

  1. Open-source solution. We will release the task synthesis recipe in the agentcore-rl-toolkit package soon.
  2. Bootstrap from existing traces. The same machinery can extract complexity patterns from production traces and generate more of what actually shows up in traffic.
  3. A guided data agent. An interactive agent that helps a team turn their policy document and traces into reward references and reward functions, and iterates on those functions against adversarial rollouts.
  4. Compression for Embedded & On-Device Deployment. Extending the synthesis recipe to specialize sub-8B models.
  1. Yao, S., Shinn, N., Razavi, P., Narasimhan, K. (2025). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. ICLR 2025. Poster · arXiv:2406.12045 · sierra-research/tau-bench
  2. Barres, V., Dong, H., Ray, S., Si, X., Narasimhan, K. (2025). τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment. arXiv:2506.07982 · sierra-research/tau2-bench
  3. τ-Bench task-quality fixes (the label-corrected release used here). taubench.com/blog/tau3-task-fixes.html
  4. Kirmayr, J., Stappen, L., André, E. (2026). CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty. ACL 2026 (Volume 1: Long Papers), pp. 40599–40618. ACL Anthology · CAR-bench/car-bench
  5. Shao, Z., et al. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (introduces GRPO). arXiv:2402.03300
  6. Yu, Q., et al. (2025). DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv:2503.14476
  7. Qwen Team (2025). Qwen3 Technical Report. arXiv:2505.09388