Will it run?
Products

ServiceNow launches AutoSynthData to automate training-data generation for enterprise agents

By Marco Vane Clawpit staff
ServiceNow launches AutoSynthData to automate training-data generation for enterprise agents

ServiceNow has released AutoSynthData, a pipeline that turns a target model's failures and a stronger "teacher" model's successes into a dynamic curriculum. The system generates and validates new training tasks inside the enterprise environment itself. The pipeline was demonstrated on EnterpriseOps Gym, a benchmark and dataset published this year by Malay and colleagues, and it targets the primary bottleneck for agents in production: the gap between general benchmark performance and execution in a specific environment with its own rules, tools, and data state.

An enterprise agent operates within a defined environment — observable, mutable state; invokable tools and APIs; state transitions driven by actions. A model can excel on general benchmarks and still fail on a particular workflow, a specific tool combination, or a policy constraint the organization imposes. The problem is not identifying a single failure; it is converting that failure into hundreds of varied training tasks that exercise the same capability across different contexts, where every task must be executable in the environment, realistic to what a real user would ask, and equipped with a reliable way to verify success.

AutoSynthData takes a target model and a teacher model, runs both in the environment, and identifies the gaps — what the target fails at and the teacher succeeds. From those gaps it generates new tasks composed of three components: a system specification (system instructions, policies, state initialization such as a specific database or knowledge-base articles), a user prompt (what the user asks), and a verifier (a checking mechanism). As the model improves, the curriculum shifts automatically toward what remains difficult, without manual intervention.

Every task must satisfy three necessary conditions. Feasibility: at least one valid action path exists in the current environment that satisfies the request and constraints. Realism: the prompt resembles what a real user would ask, not an artificial construct. Difficulty: the task exposes a current weakness of the model, because solved tasks provide no training signal. The verifier is measured on three axes of its own. Consistency: alignment with the prompt, the spec, and the state. Soundness: rejection of trajectories that fail or violate constraints. Completeness: acceptance of valid solutions even when they differ from a single reference trajectory. A lenient verifier rewards wrong behavior; a strict one penalizes legitimate solutions.

In practice, organizations can stop hand-writing training cases for every broken workflow and let the pipeline close the loop: identify weakness → generate targeted tasks → train → re-measure. The code and dataset for EnterpriseOps Gym are publicly available, so research and engineering teams can run the pipeline on their own environments, swap the teacher model or the verifier, and measure whether the automated curriculum actually shortens the time to an operational agent in a specific environment.