Fireworks launches FireRouter with Opus, smart routing that cuts encoding costs 57%

Fireworks has launched FireRouter with Opus, a first-of-its-kind routing model that decides in real time where to send each user request. Routine tasks go to leading open models; more complex ones go to Claude Opus. The mechanism weighs not only model-to-task fit but also prompt-cache cost and the price of losing a cache hit when switching models — a consideration that translates directly into money and speed.
In an internal A/B test running more than a month, the same workloads and user cohorts were split between FireRouter with Opus and Opus alone. Cost per coding session fell 57 percent (error range ±19 percentage points), from $15.36 to $6.63. Accuracy was measured on three parameters: whether the agent completed the task, whether the answer was correct, and whether the user had to correct it in the next turn. FireRouter scored 78.7 percent ranked turns versus 80.2 percent for Opus, or 98.1 percent of total accuracy. Cache-hit rate came in at 94.2 percent against 97.8 percent for Opus alone, a deliberate trade-off that enables the larger saving.
The model is available now in the Fireworks CLI and as a serverless endpoint that any account can target like any other model. It is supported in common development environments — Claude Code, Codex, Cursor IDE and others — so engineering teams can adopt it without changing their existing harness. The current model set includes Claude Opus 5.5, GLM 5.3 and GLM 5.3 Flash, and is expected to refresh as new versions ship.