Decisions API - Perplexity
Fetch the complete documentation index at: /llms.txt Use this file to discover all available pages before exploring further. The Decisions API is billed at $0.04 per million input tokens.

- Fetch the complete documentation index at: /llms.txt Use this file to discover all available pages before exploring further.
- The Decisions API is billed at $0.04 per million input tokens.
- pplx-decider-v1-27b is a decision model, and the Decisions API is how you call it.
Fetch the complete documentation index at: /llms.txt Use this file to discover all available pages before exploring further. The Decisions API is billed at $0.04 per million input tokens. Output tokens are free. See Pricing . The Decisions API answers questions with probabilities. You send the content as state , which can be text, JSON, or images, attach one or more named questions, and get one answer per question: Question type You ask You get back noul A yes/no question, or a statement to check The probability of yes, from 0 to 1 choice Pick one of the options you define A probability for every option, plus the most likely one score Rate the content on an ordered rubric A probability for every level, plus the expected score A decision model is a class of model built to make fast, structured decisions that software can use directly. It reads natural-language text and images the way a multimodal language model does, but instead of writing text it returns typed answers with probabilities: yes or no, one of your options, or a level on your rubric. It does not write replies, generate code, or explain its reasoning; your code does the reasoning with the numbers it returns. pplx-decider-v1-27b is a decision model, and the Decisions API is how you call it. Use it where you would otherwise ask a chat model for a label and parse the reply: classifying, routing, grading against a rubric, or any decision you want to threshold. You get numbers you can compare against a cutoff, no output parsing, and as many questions as you need about the same content in one request. Any Perplexity API key works; create one at console.perplexity.ai if you need one. macOS/Linux export PERPLEXITY_API_KEY = "your_api_key_here" setx PERPLEXITY_API_KEY "your_api_key_here" # setx applies to new terminals; open a new one before you continue The Decisions API is a single JSON endpoint, so any HTTP client works. The Python example uses httpx . The TypeScript example uses the fetch built into Node.js 18 and later and has no dependencies; save it as decisions.mts and run npx tsx decisions.mts (npx downloads tsx on first use), or save it as decisions.mjs and run node decisions.mjs . Send a POST to https://api.perplexity.ai/v1/decisions with the content as state , the model pplx-decider-v1-27b , and your questions. Python "https://api.perplexity.ai/v1/decisions" , headers = { "Authorization" : f "Bearer { os.environ[ 'PERPLEXITY_API_KEY' ] } " }, "title" : "Battery died after two weeks" , "review" : "The headphones sound great, but the battery stopped charging after two weeks." , "instructions" : "Does the review report a product defect?" , "instructions" : "What is the overall sentiment of the review?" , "mixed" : "Praise and complaints in one review" , "instructions" : "How severe is the reported problem?" , "criteria" : [ "Cosmetic" , "Inconvenient" , "Product unusable" ], print (answers[ "defect" ][ "noul" ]) # probability of yes, 0 to 1 print (answers[ "sentiment" ][ "choice" ]) # the most likely option print (answers[ "severity" ][ "score" ]) # expected level, 0 to 2 const response = await fetch ( "https://api.perplexity.ai/v1/decisions" , { Authorization: Bearer ${ process . env . PERPLEXITY_API_KEY } , review: "The headphones sound great, but the battery stopped charging after two weeks." , instructions: "Does the review report a product defect?" , instructions: "What is the overall sentiment of the review?" , mixed: "Praise and complaints in one review" , instructions: "How severe is the reported problem?" , criteria: [ "Cosmetic" , "Inconvenient" , "Product unusable" ], signal: AbortSignal . timeout ( 30_000 ), throw new Error ( ${ response . status } ${ await response . text () } ); const { answers } = await response . json (); console . log ( answers . defect . noul ); // probability of yes, 0 to 1 console . log ( answers . sentiment . choice ); // the most likely option console . log ( answers . severity . score ); // expected level, 0 to 2 curl -X POST https://api.perplexity.ai/v1/decisions \ -H "Authorization: Bearer $PERPLEXITY_API_KEY " \ "review": "The headphones sound great, but the battery stopped charging after two weeks." "instructions": "Does the review report a product defect?" "instructions": "What is the overall sentiment of the review?", "mixed": "Praise and complaints in one review", "instructions": "How severe is the reported problem?", "criteria": ["Cosmetic", "Inconvenient", "Product unusable"] The response has one answer per question, under the name you gave it. This is the response the request above returned, with the values as the API sent them: defect : a 94% probability that the review reports a defect. Compare noul against a threshold you choose; values near 0.5 mean the model is unsure. sentiment : mixed is the option with the highest probability (95%). probabilities covers every option you defined and sums to about 1, so you can see how close the runner-up came. severity : score is the probability-weighted average of the level indices, so 1.78 sits between Inconvenient (1) and Product unusable (2), closer to 2. legend maps each index back to your rubric, and probabilities shows the full distribution over levels. confidence on choice and score answers is the model’s own certainty estimate, from 0 to 1. It is not the top probability: in the example above, sentiment has a top probability of 0.95 and a confidence of 0.93. It drops when the runner-up is close. Identical requests usually return identical numbers. Occasionally they differ in the second decimal place, so set thresholds with some margin. model echoes the model name you sent. Every question in a request refers to the same state , which can be a string, an object, or an array. You name each question, and the response uses the same names. Each question has a type , instructions (what to decide), and, depending on the type, criteria . Ask a question or state something to check. Give instructions , criteria , or both; criteria defines what counts as yes and what counts as no. A noul with neither returns 400 . "instructions" : "Does the review report a product defect?" , "true" : "Something is broken or not working." , "false" : "Normal wear or personal preference." The answer is noul , the probability of yes or true, from 0 to 1. criteria maps each option name to a description of when it applies. Use null as the description to let the name speak for itself. A question accepts 1 to 255 options. "instructions" : "Which team should handle this ticket?" , "billing" : "Charges, refunds, and invoices" , "shipping" : "Delivery status and lost packages" , The answer has choice (the option with the highest probability), probabilities (one value per option, summing to about 1), and confidence . criteria is an ordered array of level descriptions. The index in the array is the level’s score, starting at 0. A question accepts up to 10 levels. Use at least two: with a single level there is nothing to decide, so the answer is always a score of 0 with probability 1. "instructions" : "How severe is the reported problem?" , "criteria" : [ "Cosmetic" , "Inconvenient" , "Product unusable" ] The answer has score (the probability-weighted average of the level indices, which can fall between two levels), legend (each index, as a string, mapped back to your rubric entry), probabilities (one value per level, keyed like legend ), and confidence . state can carry images next to text. Pass state as an array and put each image in an OpenAI-style image part with a base64 data URL: { "type" : "image_url" , "image_url" : { "url" : "data:image/png;base64,iVBORw0KGgo..." }} "color" : { "type" : "choice" , "instructions" : "What color is the square?" , "criteria" : { "red" : null , "blue" : null , "green" : null }} PNG, JPEG, and WebP data URLs are accepted. The API never fetches a URL: an http or https image URL returns 400 . An image can also be the whole state , with no text. The API reads images in 32 × 32 pixel tiles. Keep each image at or under 2,048 tiles: round the width and the height to the nearest multiple of 32 and keep (width / 32) × (height / 32) at or under 2,048. 1440 × 1440 and 2048 × 1024 fit; 1600 × 1310 does not. A larger image does not return 400 : the request waits about a minute and then returns 504 . Resize before you send. Image tokens count toward usage.input_tokens and the input limit like text. In our tests an image cost about 1,000 input tokens per megapixel. Limit Value Questions per request 1 to 128, each with a non-empty name Options per choice 1 to 255 Levels per score 1 to 10 Input tokens per request Under 262,144, counting state , images, and every question Request body 32 MiB Image size 2,048 tiles of 32 × 32 pixels per image, for example 1440 × 1440 or 2048 × 1024 state must be a string, an object, or an array; null returns 400 . score levels follow the same rule. choice descriptions can also be null . An unknown top-level field returns 400 . One model serves the Decisions API: pplx-decider-v1-27b . Set it on every request. The response model field echoes the name you sent. { "error" : { "code" : null , "message" : "Invalid model 'pplx-decider-v1-27b-latest'. Permitted models can be found in the documentation at https://docs.perplexity.ai/docs/getting-started/models." , "param" : null , "type" : "invalid_request_error" }} Operation Endpoint Answer questions POST https://api.perplexity.ai/v1/decisions Send the key as Authorization: Bearer <PERPLEXITY_API_KEY> and the body as JSON. A key in an x-api-key header is not read, so the request returns 401 . Another method on the endpoint returns 405 with Allow: POST , and any other path, including a trailing slash, returns 404 . Responses carry an x-request-id header with a UUID, on success and on most errors. Log it with your results and quote it in support requests. A 401 , a 404 , and a 504 carry no request id. Every organization can send 10 requests per second to the Decisions API, on every plan. A token limit also applies to large bursts. Successful responses carry x-ratelimit-limit , x-ratelimit-remaining , x-ratelimit-used , and x-ratelimit-reset (Unix seconds). A request over a limit returns 429 with a Retry-After header in seconds. Wait that long, then retry. Most errors return a JSON body with an error object. Read error.message for the reason and error.type for the category; don’t branch on error.code , which is a string, a number, or null depending on the error. A 404 or 405 has an empty body, and a 504 can return an HTML page, so check the status before you parse the body. { "error" : { "message" : "Noul question must have criteria or instructions" , "type" : "invalid_request" , "code" : "400" }} Status Why What to do 400 The body is not valid JSON, model is missing or unknown, a field is unknown or the wrong type, a limit in Request limits is exceeded, or an image is not a supported data URL Fix the request; error.message names the problem 401 The API key is missing or invalid, or it was sent in x-api-key Send Authorization: Bearer <PERPLEXITY_API_KEY> and check that the key is active 404 Wrong path Use POST https://api.perplexity.ai/v1/decisions , with no trailing slash 405 Wrong method Use POST 413 The request body is over 32 MiB Send less state , or split the content across requests 429 Over the request or token limit; see Rate limits Wait Retry-After seconds, then retry 5xx The model did not answer in time ( 504 , after about a minute) or the service failed Retry with backoff; for large inputs, see Timeouts Response time grows with input size. In our tests on September 30, 2026, a request with a few hundred input tokens answered in under 2 seconds, about 90,000 tokens took 5 seconds, about 190,000 tokens took 14 seconds, and just under the input limit took 23 seconds. If the model does not answer in time, the request returns 504 ; in our tests that took about a minute. Set your client timeout to fit the input. The examples on this page use 30 seconds, which covers any request under the input limit; for small inputs, 10 seconds is plenty. The Decisions API costs $0.04 per million input tokens. Output tokens are free, and there is no per-request fee. Input tokens are the usage.input_tokens value in each response, so you can compute the cost of a request from the response you already receive. Usage is billed to the organization that owns the API key, like every other Perplexity API. Full request and response schema for POST /v1/decisions . Route twelve tickets with three questions per ticket, then send only the escalations to the Agent API. Web-grounded answers with citations, tools, and structured output. Direct access to open-weight models through OpenAI- and Anthropic-compatible endpoints. Need help? Check out our community for support. Responses are generated using AI and may contain mistakes. What is the Agent API? Which Agent API model should I use? How do I create an API key?
Sources
Related stories

Models & Pricing - DeepSeek
The prices listed below are in units of per 1M tokens. A token, the smallest unit of text that the model recognizes, can be a word, a number, or even a punctuation mark.

Building and Evaluating a QA System with LlamaIndex - LlamaIndex
LlamaIndex (GPT Index) offers an interface to connect your Large Language Models (LLMs) with external data. LlamaIndex provides various data structures to index your data, such as the list index, vector index, keyword index, and tree index.

How NVIDIA GPUs Help Accelerate OpenAI’s GPT-6 Astra Ultrafast
GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users.

Gemini 4 Argon: our next era of frontier intelligence
Gemini 4 Argon: our next era of frontier intelligence Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. SVP, Google DeepMind and Chief AI Architect, Google Google’s new Gemini 4 Argon model brings advanced reasoning to complex, long-horizon professional tasks.