AERIOXFLUX
← AI Tools
AI Tools · infra apis

OpenAI's Decisions API Bills Only Input, and Its Choices Run Hot

The new GPT-6 Luna endpoint charges $0.10 per million input tokens and nothing for output, but a forum tester found its choice format called a 70% coin heads 98% of the time.

Flux Desk·2026-10-09·5 min read

Give OpenAI's new Decisions API a coin that lands heads 70% of the time and ask it, in the right format, how likely heads is. It says 70%. Ask the same question as a multiple-choice pick and it says 98%. That result comes from a developer posting as platypus on OpenAI's own community forum, one day after the endpoint opened to everyone, and it is the most useful thing anyone has published about the product so far.

OpenAI put the Decisions API into public beta on October 6. It is a narrow tool with an unusual price. It does not write prose. It reads text or images and returns typed answers to questions the developer defines, and it bills only for what goes in.

A model that answers, not one that talks

The endpoint is POST /v1/decisions, and according to OpenAI's documentation, gpt-6-luna is the only supported model. A request carries an input and a list of named questions in one of three types. A predicate checks whether a condition is true and returns a probability from 0 to 1. A choice picks one option from a set the developer supplies and returns the pick, a probability for each option and a separate confidence value. A score rates the input against ordered levels and returns a probability-weighted average of the level indices, which can land between levels.

OpenAI says this produces answers about 10 times faster than the Responses API. It has not published a benchmark for that claim or an accuracy comparison against the same Luna model called the normal way, a gap aggregator AI Weekly also flagged. Images must be sent as inline base64 data; hosted URLs and uploaded file IDs are refused. Zero Data Retention and HIPAA use are available to eligible customers, with data residency in the US and Europe. General availability is expected "in the coming weeks," per the docs.

The pricing is the hook. Input costs $0.10 per million tokens. Output tokens are free, and there are no cache charges, though regional processing premiums and long-context multipliers may apply. For a product whose outputs are a handful of numbers, charging for output would be almost beside the point. The design makes cost a function of how much you ask the model to read, which is exactly the bill a content-moderation queue or a support-ticket router wants.

The coin test

Platypus ran batches of 1,000 requests and wrote them up in a separate forum thread on October 7. With predicates, the API reproduced the defined probabilities closely. A single flip came back 70% heads and 30% tails. Three-flip questions landed within a fraction of a point: all heads at 34.00% against an exact 34.30%, at least one head at 97.00% against 97.30%.

The choice format behaved differently. It returned 98% heads and 2% tails, and selected heads in all 1,000 responses. A jar of marbles split 50% red, 30% blue and 20% white produced red at 86% as a choice question, again selected every time. Reversing the order of the options, so red came last, dropped red's estimate to 73% and lifted white from 4% to 19%. The predicate estimates did not move at all. Platypus also found the score type unreliable, reporting probabilities that summed above 100%.

Another forum member, VeitB, offered the generous reading: a choice optimized for accuracy should pick heads every time, because heads is the better bet. Platypus agreed the format is "not estimating the true posterior at all," and concluded, for now, "don't use choice questions." A separate commenter argued the results reflect small-model unreliability and shift with prompt wording.

Both readings lead to the same practical rule. The choice format's selection is a decision. Its per-option probabilities are not a calibrated estimate, and they are sensitive to the order you list things in. Developer blog Classmethod saw a related pattern on a smaller scale: across five runs, OpenAI's own sample choice question came back with confidence pegged at 1.0, where the documentation's example shows 0.93. OpenAI's guide, to its credit, tells developers to set thresholds against labeled examples from their own application rather than trusting the raw numbers.

Cheap, but not the cheapest

OpenAI is not first to this shape of product. Classmethod lists two rivals with similar interfaces: TypeSafe's Jev, a text and JSON decision model at $0.042 per million input tokens, and Cloudflare's Clef on Workers AI, which takes images and costs $0.24, or $0.09 for its flash variant. The formats are not compatible, so switching means rewriting calls.

On the forum, the early comparisons were not flattering. One tester, jefff, ran a few hundred questions through both and reported that Jev was more accurate on nuanced judgment calls, roughly equal on simple yes-or-no checks, faster, and cheaper by two to three times. His bottom line was that Jev is the better default. OpenAI's advantage, for now, is images. Another commenter noted Jev does not yet handle image understanding and suggested Decisions for checks like whether an upload fits site guidelines.

Vercel is already routing traffic to the endpoint through its AI Gateway, according to AI Weekly. That is the realistic near-term use: a fast, cheap filter at the front of a pipeline, asked yes-or-no questions. The predicate format appears to earn that job. The choice format, on the evidence of one forum user's thousand coin flips, still has to prove its numbers mean what they say.

#openai#decisions-api#gpt-6-luna#calibration#api-pricing

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.