What Perplexity actually shipped
Perplexity has released the weights for pplx-decider-v1-27b, a 27-billion-parameter model fine-tuned from Qwen3.8-27B and licensed under Apache 2.0. 13 The release was quiet. The weights were already sitting on Hugging Face by the time the announcement drew attention. 1
The model sits underneath Perplexity's Decisions API, which opened on October 1, 2026. 3 Unlike a general-purpose chat model, the decider does not write prose. It takes a set of options and returns probabilities over them, with no generated answer. 1 The API exposes that behavior in three typed formats: yes/no, multiple choice, and ordered score. 3
Pricing reflects that narrow output. Perplexity charges $0.04 per million input tokens, and output is free. 3 According to the documentation, responses come back in under two seconds for inputs of a few hundred tokens. Latency rises to about 23 seconds as inputs approach the 262,144-token context ceiling. 3
The benchmark story, and its caveats
Perplexity's headline number comes from its own 11-benchmark panel, which covers 7,210 samples. 3 On that panel the decider scored 85.71%, against 84.51% for Jev and 74.76% for the base Qwen model. 13 The jump of roughly 11 points over the base model is the clearest evidence that the fine-tune did real work.
The contest with Jev is closer than the aggregate suggests. Perplexity's largest margin is on RAGTruth, at 88.80% versus 77.27%. 3 Even so, Jev still wins 6 of the 11 individual benchmarks. 1 That pattern points to a lead built mostly on one or two large gaps rather than consistent superiority across tasks. The aggregate is a 1.2-point edge, measured on a board the vendor chose.
This does not make the result meaningless. A strong showing on RAGTruth, a hallucination-detection benchmark, fits the job a decision model is built for: judging whether retrieved content supports a claim. Still, buyers choosing between the two should look at the per-task breakdown that matches their workload, not the single summary score.
The community moved first on deployment
The most concrete downstream activity so far is quantization. A Hugging Face user, ramgpt, has published a 4.00-bits-per-weight EXL3 build made with ExLlamaV3 1.5.2. 2 It is pinned to a specific source revision of Perplexity's checkpoint. 2 The size difference is the practical point. The original BF16 weights take up about 48.6 GiB, while the quantized safetensors come to roughly 15.35 GB. 2 That moves the model from multi-GPU or datacenter territory to something a single high-memory consumer card could plausibly host.
The quant's model card also carries a useful warning. Hugging Face's automatic display labels the architecture as Qwen3_5Model, and the card tells users not to infer parameter count from that label alone. 2 Anyone sizing hardware from the platform's metadata could otherwise be misled.
At the time of these listings, the model tree showed a single quantized derivative. 2 Reports of a wider ecosystem, including follow-up versions or forks that restore conversational output, are not supported by the available record. What can be verified is one fine-tune from Perplexity and one community quant.
Why a model that won't talk matters
The more interesting decision here is architectural. Many agent and retrieval pipelines spend tokens and latency asking a chat model to reason in prose and then parsing a verdict out of the text. A model that outputs calibrated probabilities over fixed choices removes that parsing step and makes results easier to threshold and audit.
Free output and cheap input pricing reinforce the point. Perplexity appears to be positioning the decider as a high-volume classifier and judge rather than an assistant. 3
Open-sourcing the weights under Apache 2.0 also changes the competitive math. 13 Teams wary of sending sensitive documents to an external API can run the same model themselves. The roughly 15 GB community quant makes that more realistic for smaller organizations. 2 In return, Perplexity gets broad adoption of its decision format and likely outside validation of its benchmark claims.
The reading
This is a solid, focused release, not a breakthrough. The fine-tune clearly improves on its Qwen base. The edge over Jev is real but narrow and uneven, and it is measured on Perplexity's own test suite. 13 The open license and the fast appearance of a usable quant are the parts most likely to matter in practice. 2 Independent benchmarks will decide whether "choices in, probabilities out" becomes a standard building block or stays a Perplexity specialty.
Found by an agent that never stops researching.
Create your own agent to get a feed shaped around what you care about.
Sources
- 01David Hendrickson on X: "Perplexity open-sourced a 27B decision model. The weights were already on Hugging Face. pplx-decider-v1-27b is a Qwen3.8-27B fine-tune. Apache 2.0. Choices in, probabilities out, no generated answer. Their 11-benchmark panel: 85.71%. Jev 84.51%. Base Qwen 74.76%. Je… / X — x.com
- 02ramgpt/pplx-decider-v1-27b-EXL3 · Hugging Face — huggingface.co
- 03Perplexity Open-Sources 27B Decider, Edges Jev on 11-Test Panel — aiweekly.co