Skip to main content

What does the Perplexity Decisions API mean for teams building AI agents?

Perplexity's new Decisions API returns probabilities for yes-or-no, choice and rubric questions instead of generated text. Where it fits in an agent stack, and where a chat model is still the right call.

Woman holding a laptop asks a question of a developer whose monitors show code and a glowing AI brain graphic in a brick-walled office

Perplexity's Decisions API answers questions about text, JSON or images with probabilities instead of generated text: a yes-or-no likelihood, a ranked choice or a rubric score. For teams building AI agents, it is a cheap, parse-free way to classify, route and grade, as long as someone owns the thresholds the numbers feed into.

What Perplexity released

The Decisions API was announced by Perplexity in its API changelog in October 2026, first listed on 5 October. It is a single endpoint, POST https://api.perplexity.ai/v1/decisions, served by one model, pplx-decider-v1-27b. Any existing Perplexity API key works, and usage is billed to the organisation that owns the key, like the rest of the Perplexity API.

Perplexity calls pplx-decider-v1-27b a decision model. According to the Decisions API documentation, it reads text and images like a multimodal language model, but it does not write replies, generate code or explain its reasoning. It returns typed answers with numbers, and your code does the reasoning.

What actually changed

Until now, the usual way to get a label from a language model was to ask a chat model and parse its reply. The Decisions API replaces that with three question types, all asked about the same state (a string, an object or an array):

  • noul: a yes-or-no question or a statement to check. The answer is the probability of yes, from 0 to 1.
  • choice: pick one of the options you define. The answer gives the most likely option, a probability for every option and a confidence value.
  • score: rate the content on an ordered rubric of up to 10 levels. The answer is the probability-weighted average level, plus the full distribution.

The request limits are generous for batch work: 1 to 128 named questions per request, up to 255 options per choice, and an input limit just under 262,144 tokens. Images go in as base64 data URLs (PNG, JPEG or WebP); the API never fetches an image URL. Every organisation can send 10 requests per second, on every plan.

Pricing is simple: Perplexity bills input tokens only, at a low per-token rate, with free output tokens and no per-request fee. Because the response reports usage.input_tokens, you can work out the cost of each call from the response itself.

What it means for teams building AI products and agents

Most agent systems already contain small decisions dressed up as chat calls: is this ticket urgent, which tool should handle this request, does this answer meet the policy, is this document an invoice. In CodeDTX's view, those are where the Decisions API is worth a trial.

  1. Routing and triage before the expensive model. Ask a few questions about each input, then send only the cases that need reasoning to a full agent. Perplexity's own ticket triage cookbook does exactly this, passing only escalations on to its Agent API. Our guide to controlling what an AI agent costs to run covers why this pattern matters for the bill.
  2. Thresholds you can tune and audit. A probability lets you set a cut-off for automatic action and send the uncertain middle to a person. That is the design in our note on using an AI agent's confidence score to decide on human review. Perplexity's documentation says identical requests occasionally differ in the second decimal place, so leave a margin around any threshold.
  3. Grading in evals. A rubric score with a full distribution is easier to track over time than a free-text judgement. It fits the checks described in what AI agent evals catch, with the caveat that the grader is itself a model you need to test.
  4. Plan for timeouts. Response time grows with input size. Perplexity's tests put a few hundred tokens under 2 seconds and a near-limit request at around 23 seconds, and a request the model cannot finish returns 504 after about a minute. Set client timeouts to the input size, and resize large images first: an oversized image does not fail fast, it times out.

When not to switch

Keep a chat model where the output itself is text: a reply to a customer, a summary, generated code, or an explanation a reviewer needs to read. The decision model gives no reasoning, so if your governance process requires a stated reason for each automated decision, you still need to record one elsewhere. And if your data cannot leave your own cloud or region, check that first, as in our guide to data residency for AI agents: this is a hosted API with one model and no self-hosted option in the documentation.

Frequently asked questions

What is the Perplexity Decisions API?

It is a Perplexity API endpoint that answers questions about your content with probabilities instead of text. You send text, JSON or images as the state, attach named questions, and get back a yes-or-no probability, a ranked choice among your options, or a score on your rubric. It is served by a single decision model, pplx-decider-v1-27b, and works with any existing Perplexity API key.

How is it different from asking a chat model for a label?

A chat model writes text that your code has to parse, and it may not return the format you asked for. The Decisions API returns typed answers with a probability for every option, so there is no parsing step and you can compare the numbers against a threshold directly. You can also ask up to 128 questions about the same content in one request.

What does the Decisions API cost?

Perplexity bills input tokens only. Output tokens are free and there is no per-request fee, so the cost of a call depends on how much content and how many questions you send. Each response reports its input token count, which lets teams log the cost of every decision alongside the answer and set budget alerts on the same numbers their dashboards already show.

Should we replace our classifier with it?

Not without a side-by-side test. Run the Decisions API on a labelled sample of your own data next to your current classifier or chat-model prompt, compare the results at the thresholds you would actually use, and check latency at your real input sizes. Switch only where it matches or beats what you have, and keep a way to roll back if the model changes.

Is it suitable for regulated decisions?

It can support them, but it should not make them alone. The model returns numbers without reasons, so a regulated process still needs a recorded rationale, a human review step for uncertain or high-impact cases, and an audit trail of inputs, questions and outputs. Teams should also confirm where the data is processed before sending personal or sensitive content to the hosted API.

Share this post

Contact us to build the right product

Talk to our engineers about your application, the systems it connects to, and what you want to build next.

Get in touch
Two people discussing work with a laptop