Saluki 27B compressed to 7.89 GB

Underdog AI recently released Saluki 27B, a 2-bit quantized version of Qwen3.8-27B. The model files are only 7.89 GB, with a focus on optimizing local agents’ tool selection and parallel invocation capabilities.
7.89GB, 27B Models Begin Entering the Usable Range for Local Agents
Underdog AI released Saluki 27B on October 7. As of today (October 10), the most notable thing about this model is no longer that it is “another quantized version of Qwen,” but that it compresses a 27B-scale model down to 7.89GB while focusing its optimization on tool calling, one of the capabilities local Agents need most.
Saluki 27B is fine-tuned with 2-bit quantization based on Qwen3.8-27B and retains text input and output capabilities only. According to the official specifications, the model file is 7.89GB, while the unquantized version requires approximately 54GB of storage. This represents an approximately 85% reduction in size, leaving the file at roughly one-seventh the size of the original model.

This is not simply a matter of making the model “smaller.” 2-bit quantization essentially fits a larger model into limited memory and VRAM at the cost of sacrificing some weight precision. During quantization, model parameters are converted from higher-precision data types into lower-bit representations, after which scaling factors and other techniques are used to restore the original distribution as closely as possible. The lower the bit width, the lower the storage and bandwidth requirements, but the more likely the model is to lose its ability to handle details, complex reasoning, and long-chain tasks.
Saluki’s trade-offs are clear: it does not attempt to preserve Qwen3.8-27B’s performance across every benchmark. Instead, it prioritizes its limited precision budget for everyday reasoning, tool selection, and parallel calls to multiple tools. Underdog AI claims that Saluki even surpasses the full-size Qwen3.8-27B in tool selection and parallel multi-tool calling, although it performs relatively worse on mathematical tasks.
This indicates that its goal is not to become a local chat model that can “do everything,” but rather to serve as an Agent foundation that can run continuously on a computer and call local tools and external services.
For Local Agents, Tool Calling Matters More Than Question-Answering Scores
Traditional language-model evaluations typically focus on knowledge question answering, mathematics, code generation, and general reasoning. In Agent scenarios, however, the model’s actual job is often not to provide an answer directly, but first to determine which tools to call, then generate correctly structured parameters, wait for the tools to return results, and finally continue with the next round of actions.
For example, suppose a user says: “Check today’s high-speed trains from Shanghai to Hangzhou, filter for departures in the afternoon, and put the three cheapest options into a table.”
An ordinary chat model might fabricate train schedules from memory or provide only a natural-language suggestion. An Agent model, by contrast, needs to complete a sequence of actions:
- Identify the need to call a transportation query tool;
- Extract the departure location, destination, date, and time range;
- Generate parameters that conform to the tool’s schema;
- Process the multiple results returned by the tool;
- Further sort and filter the results, then output them as a table;
- If the tool call fails, decide whether to retry or use another tool.
In this process, selecting the right tool, providing complete parameters, and launching multiple tools in parallel usually have a more direct impact on the product experience than gaining a few extra points on a mathematics problem.
These are precisely the areas Saluki has chosen to strengthen, and they are areas that general-purpose leaderboards often overlook. It is more like the control center of a local Agent: responsible for understanding tasks, planning actions, assigning tools, and reorganizing the results.
The ability to make “parallel calls to multiple tools” is particularly important here. Suppose an Agent needs to query the weather, calendar, and flight information at the same time. If the model can initiate independent calls in a single operation, the overall response time does not need to equal the combined duration of three sequential calls. For desktop automation, coding assistants, personal knowledge bases, and local workflows, this capability can directly determine whether the interaction feels smooth.
However, it is important to note that the official information currently discloses positioning and performance conclusions, but does not provide sufficiently complete and reproducible details of a unified evaluation. The claim that Saluki “surpasses the full-size Qwen3.8-27B” needs to be understood in the context of the specific tool-calling dataset, prompt template, inference parameters, and quantization implementation. It should not be interpreted simply to mean that Saluki is stronger than the original model on every task.
7.89GB Does Not Mean It Will Run Easily with 8GB of VRAM
A model file of only 7.89GB does make Saluki appear to be within reach of consumer-grade hardware, but actual operation also needs to account for the context cache, runtime overhead, and the continuously growing conversation history generated during tool calls.
When loading a model in llama.cpp, VRAM or memory consumption usually includes more than the weight file itself:
- KV Cache, used to store context state;
- Computation buffers and temporary tensors;
- The overhead of the model loader and runtime itself;
- Additional memory requirements from longer contexts;
- Historical messages, tool schemas, and returned results generated by multiple Agent calls.
Therefore, 8GB of VRAM does not mean the entire Saluki 27B model can be placed in VRAM unconditionally. It may run with a short context, a lower batch size, and partial CPU offloading, but speed and stability will depend on the hardware. An ordinary desktop computer with more than 16GB of system memory is more likely to start the model in pure CPU or hybrid mode. If better interactive speed is desired, a more powerful GPU or higher memory bandwidth is still necessary.
In addition, Agent task contexts tend to grow more quickly than those of ordinary question answering. Tool definitions, call histories, web content, code snippets, and error logs are all added back into the context. If the goal is merely to verify whether the model can start, 7.89GB is a very attractive figure. If the goal is to run a local assistant that calls a dozen tools over an extended period, memory requirements need to be reassessed based on the actual context length.
In other words, 7.89GB lowers the entry barrier; it does not turn the hardware requirement into “8GB of VRAM is enough for full-speed operation.”
Standard llama.cpp Format Reduces Deployment Friction
Saluki uses the standard llama.cpp format and can be loaded directly in standard llama.cpp and applications built on it, without custom compilation. This may seem less eye-catching than its quantization precision, but it is highly significant for the practical adoption of local models.
Local model deployment often gets stuck in the final mile: the model itself has already been released, but users still need a specific branch, a special conversion script, custom operators, or an adaptation for a particular inference framework. For developers, these extra steps mean higher maintenance costs. For ordinary users, they mean that even after the model has finished downloading, it may not be ready to run immediately.
The value of a standard format is that Saluki can integrate relatively naturally into the existing llama.cpp ecosystem, including local chat interfaces, desktop model-management tools, command-line inference services, and some Agent frameworks. Developers can first use llama.cpp to verify the model’s behavior, then decide whether to integrate it into a more complete application layer based on project requirements.
For local Agents, this compatibility also means that the tool layer can remain independent. The model generates tool calls, llama.cpp handles inference, and external programs execute file operations, command-line tasks, browser actions, or API requests. As long as the calling format and schema conventions are clear, the cost of replacing the model will not be too high.
Of course, compatibility with llama.cpp does not mean that every tool-calling framework will work out of the box. Different applications do not handle function-calling formats, special tokens, stop conditions, and JSON constraints in exactly the same way. Developers still need to verify whether the model output strictly follows the requirements of the target framework, paying particular attention to whether:
- Tool names always come from the allowed list;
- Parameter types conform to the schema;
- The model proactively asks follow-up questions when required parameters are missing;
- It can recover correctly when a tool returns an error;
- Repeated execution or circular calls occur after multiple rounds of calling.
What It Is and Is Not Suited For
From a product-positioning perspective, Saluki is better suited to the following scenarios:
- Local coding assistants: Reading project files, searching code, running tests, and continuing to make modifications based on the results.
- Desktop automation: Combining file-system, terminal, calendar, note-taking, and local-search tools to complete everyday operations.
- Personal knowledge-base Agents: Searching and organizing information in local documents, exported email files, or note repositories.
- Offline or weak-connectivity environments: Workflows with high data-privacy requirements that do not want all context sent to the cloud.
- Agent prototype development: Testing tool orchestration, task planning, and error-recovery logic at a lower hardware cost.
It may not be well suited to mathematical reasoning, serious scientific computing, or tasks that require high-precision, long-chain derivations. 2-bit quantization inherently compresses the model’s expressive capabilities, while Saluki has also deliberately focused its training and fine-tuning on tool use. For scenarios involving complex formula derivation, precise numerical calculations, and verifiable proofs, it is better to have the model call a calculator, Python runtime, or specialized solver rather than making it handle all calculations on its own.
This is also an important principle in Agent design: the model does not need to know every answer, but it must know when to delegate a task to a tool. Saluki’s value lies precisely in making this division of labor smoother while consuming fewer local resources.
The Real Value of 2-Bit Quantization Is Redefining the Competitive Dimensions of Local Models
In the past, developers discussing local models often focused first on parameter count, context length, and general evaluation scores. Models like Saluki push the competitive dimension one step further: within the same hardware budget, which model can complete real-world workflows more reliably?
If a 27B model requires tens of gigabytes of VRAM, many individual developers can only run it in the cloud. If it can be compressed to 7.89GB while retaining a substantial portion of its tool-calling capabilities, it has the potential to become a practical component of local Agents. Even if its speed is ultimately lower than that of large cloud models, its privacy, cost, and low latency may still make it more valuable in specific scenarios.
But low-bit quantization is not free. Developers need to accept several realities:
- Complex reasoning capabilities may decline;
- The ability to preserve details in long contexts may weaken;
- Stability on specialized tasks may be worse than that of higher-precision versions;
- Speed can vary significantly across different hardware and inference backends;
- Tool-calling success rates need to be tested separately in real workflows.
Therefore, Saluki should not be evaluated solely by whether the model can load or by running a few simple question-answering tests. A more effective method is to build a small task set covering file search, code modification, command execution, web-information extraction, and parallel multi-tool calls, then measure tool-selection accuracy, parameter-validity rate, task-completion rate, average number of calls, and error-recovery success rate.
If a model appears intelligent in a single-turn conversation but frequently calls the wrong tool, generates invalid parameters, or repeats the same action after a failure, it is not suitable as an Agent foundation. Conversely, a model that produces less polished answers but completes tasks reliably may be more suitable for local automation products.
Conclusion: It Is Not a Smaller Chat Model, but a Cheaper Testing Ground for Agents
The value of Saluki 27B’s release does not lie in repackaging the “27B” parameter scale. Rather, it demonstrates a highly practical direction: through extreme quantization and task-oriented fine-tuning, a model that originally required substantial hardware resources can be compressed into a range that more developers can experiment with.
The 7.89GB size makes local deployment easier, the standard llama.cpp format lowers integration costs, and the optimizations for tool selection and parallel calling shift its target from “local chat” toward “local task execution.” It is the combination of these three points that makes Saluki genuinely worth watching.
Of course, its ultimate performance will still depend on the real operating environment. The model file size is not a complete accounting of hardware requirements, and the claimed advantage in tool calling still requires further independent evaluation. For developers, the most reasonable way to use it is not to treat it as a comprehensive replacement for a flagship cloud model, but to place it into a local Agent workflow and test whether it can reliably complete tasks with acceptable latency and resource consumption.
If the test results hold up, Saluki may become representative of a distinctive class of local models: models that do not aim to win on every leaderboard, but instead spend their limited parameter precision on the capabilities that truly determine whether an Agent can work: “selecting tools, calling tools, and continuing to act.” For developers who want to run Agents on personal computers, workstations, or edge devices, this trade-off is more practically meaningful than yet another general-purpose chat model.
References
- IT Home: Underdog AI Releases the Saluki 27B Model — Introduces the model’s release date, 2-bit quantization, 7.89GB file size, and primary capability positioning.
- Hugging Face: Underdog-Saluki-27B-1.0 — Provides the model weights and related deployment information for further testing.



