openJiuwen Releases X-Router, Saving More Than Half of Routing Tokens
openJiuwen recently released its self-evolving model routing technology, X-Router. In the Terminal-Bench and LLMRouterBench tests, the cost-priority mode reduced costs by 51.4% compared with top-tier models while achieving 95.2% of their success rate; the quality-priority mode achieved a 98.6% success rate with a 15.7% reduction in costs.
openJiuwen Releases X-Router, Targeting Agent Inference Costs
openJiuwen recently released X-Router, a self-evolving model routing technology designed to address a recurring cost in Agent deployment: assigning every task to the most powerful model regardless of difficulty. The team says that in Terminal-Bench and LLMRouterBench tests, X-Router reduced invocation costs by more than 50% while achieving task success rates close to those of top-tier models.
This is not a new model, but a decision-making layer that sits in front of model calls. It determines what capabilities a task requires and then selects an appropriate model from different tiers. As real-world feedback accumulates, the routing strategy can continue to adjust. For Agent developers, the key change is not simply that they can “switch to cheaper models,” but that model selection can move from rules hard-coded into business logic to a configurable, measurable runtime strategy.
A Knob for Balancing Cost Savings and Quality
According to supplementary materials, openJiuwen uses Opus 4.8, Qwen3.5-122B, and Qwen3.5-35B to form a three-tier model pool, conducting single-turn and multi-turn routing tests on Terminal-Bench and LLMRouterBench. In cost-priority mode, costs fell by 51.4%, while the task success rate reached 95.2% of the Opus 4.8 baseline. In quality-priority mode, costs fell by 15.7%, while the success rate reached 98.6%.
The value of these figures is that they do not present “saving money” as a free lunch. Cost-priority mode explicitly accepts some loss in success rate, while quality-priority mode trades a smaller cost reduction for performance closer to that of a top-tier model. Developers can choose according to their product scenarios: workloads such as internal batch processing and information extraction may be more tolerant of a cost-priority strategy, while high-value user-facing tasks may call for more conservative thresholds.
However, these results were obtained with a specific model pool and benchmark set, and do not mean that every application can reliably cut costs in half. The gains in a real system depend on task distribution, model pricing, retry mechanisms, context length, and evaluation criteria. In particular, for long-chain Agents, a single routing error may trigger retries or additional tool calls. Savings on individual tokens do not necessarily translate into a proportional reduction in the final bill. To determine whether X-Router is suitable for production, developers still need to replay their own traffic while monitoring success rate, latency, and total cost.
How It Decides Which Model Should Handle a Task
X-Router’s basic approach is to estimate task complexity first and then match it to an appropriate capability tier. Simple questions are routed to lightweight models, while coding, research, and complex reasoning tasks are progressively escalated. The supplementary materials state that its complexity classifier can use a locally deployed Qwen3-0.6B and run in-process without requiring an additional standalone service.
This is equivalent to placing a triage desk in front of the model pool: there is no need to consume an expensive reasoning model to answer “What does HTTP 429 mean?” A data-processing script can be assigned to a model with stronger general capabilities, while multi-document research or mathematical proofs can be escalated as needed. The goal is not to always choose the lowest-priced option, but to prevent expensive models from handling work that does not require them while preserving the ability to escalate when capabilities are insufficient.
A local classifier also introduces engineering trade-offs. It reduces additional network calls and service dependencies, but its own decisions affect downstream quality: underestimating difficulty may lead to failed responses, while overestimating it may simply spend the saved cost again. As a result, the router needs to be calibrated against real tasks. Classification accuracy alone is not enough; teams must also examine end-to-end results after routing.
From Model Selection to Orchestrating Execution
openJiuwen positions X-Router as more than a model-selection system. The materials also describe two types of extended capabilities: multi-model collaboration and policy self-orchestration.
Multi-model collaboration targets problems that a single model cannot answer reliably. The routing layer decides whether to call multiple models in parallel and which models to call. Their outputs are then semantically deduplicated, screened for quality, checked for conflicts, and condensed. This differs from simply sending the same question to multiple models and stacking the answers together. The key is whether the system can identify duplicated results and contradictory judgments. Multi-model collaboration increases invocation costs, so it is worthwhile only when a task genuinely requires cross-validation.
Policy self-orchestration expands the scheduling targets to models, Skills, Tools, and SubAgents. One step may have a model perform reasoning, another may call a tool for execution, and a SubAgent may then handle verification. In other words, routing is no longer limited to answering “Which model should be used?” It is beginning to decide “What capability should perform this step?” This aligns with the evolution of Agent frameworks, but it also increases complexity: teams need to record the inputs and outputs, invocation costs, and failure reasons for every step in order to determine whether orchestration is more effective than a fixed workflow.
The Key to Self-Evolution Is Feedback, Not Automatic Intelligence
X-Router emphasizes “self-evolution.” Its core promise is to feed production feedback into subsequent routing decisions. This direction has practical significance: task types change, and model capabilities and prices change as well, so static rules can quickly become outdated. If a system can revise its strategy based on success, failure, and cost data, it may gradually reduce unnecessary high-priced calls.
However, “self-evolution” does not mean that a router can modify its strategy online without constraints. Production environments need to define what the feedback signals are: whether a task succeeded, whether the user accepted the result, whether a retry was triggered, or whether the output received a quality label after human review. Different signals may conflict with one another, and short-term cost reductions may conceal a decline in response quality. A more prudent deployment approach is to evaluate new strategies through offline replays and shadow traffic first, then gradually increase exposure while preserving rollback capabilities.
The team also evaluated all 147 tasks in WorkSwarm and PinchBench, covering 11 categories including log analysis, data analysis, coding, and research. This setup is closer to Agent workloads than a single question-answering test, but the publicly available materials still do not provide enough information to fully assess the test details, such as the success criteria for each task, the number of repeated experiments, the pricing assumptions, and the comparison conditions for fixed-model strategies. For production teams, benchmark results are a starting point for screening tools, not a purchasing conclusion.
What This Means for Developers
The problem X-Router addresses is real. Once an Agent enters multi-turn execution, model calls are amplified layer by layer by planning, tool use, reflection, and retries. Simply specifying one model for every request can lead to excessive spending on simple tasks, while using only inexpensive models may convert savings into more failures and retries. The value of model routing lies in providing a tunable strategy layer across cost, quality, and latency.
It is also not a component every team needs to adopt immediately. If invocation volume is low and task types are fixed, a few explicit business rules may be easier to maintain. If workloads vary significantly, the model pool includes multiple price and capability tiers, and the team already has reliable task evaluations, intelligent routing is more likely to generate net benefits. At a minimum, evaluations should compare three options: a fixed powerful model, a fixed lightweight model, and a routing strategy. Results should also be broken down by task category rather than judged solely by an aggregate success rate.
From an industry perspective, X-Router reflects how model competition is expanding from “whose model is stronger” to “who can use models more effectively.” As the number of capability tiers grows and price differences persist, the routing layer will become important infrastructure for Agent systems. However, the router itself must also prove that it is not consuming the tokens it saves on classification, parallel calls, and failure recovery. openJiuwen has presented test results worth watching; the more important question now is whether developers can reproduce those gains across their own task distributions and clearly identify the true boundaries between quality and cost.



