1

Let your app choose the right model, tool, or action in near real-time with Decisions API, now available to all developers in public beta.

The Decisions API makes decisions up to 10x faster than GPT-6 Luna through the Responses API.

Powered by GPT-6 Luna, it accepts text and image inputs and supports 3 kinds of outputs:

  • Predicates: Estimate the probability that a statement is true.
  • Choices: Select from predefined options, with confidence scores.
  • Scores: Evaluate an input against a numeric range.

Pricing

With gpt-6-luna, input costs $0.10 per 1M tokens. You pay only for input tokens: there are no cache-read, cache-write, or output-token charges.

You can now try the Decisions API in public beta:

Ah now Santa really did come early this year! Pumped to give this a try!

4

Let’s see who will be the first to benchmark Decisions against Jev!

Now this is what I was focused on from DevDay. OpenAI, really hoping you guys publish a blog explaining this and the model architecture to some extent like decider-2b.

6

I ran some quick tests vs jev for my use cases, didn’t see an improvement with Luna :frowning: Looking forward to seeing Luna on some real benchmarks though - is S1MB on huggingface the best place to look?

  • Accuracy: Jev was more accurate on nuanced judgment calls, such as whether two words are meaningfully related or what a person in a conversation is asking for. Luna made about three times as many confident wrong answers on those.
  • Simple factual questions: on narrow yes/no checks like “is this a real English word?”, the two were about equal. Luna was slightly more often right and slightly less well calibrated, and the difference wasn’t large enough to matter in practice.
  • Speed: Jev’s typical response was faster. In multi-step conversational use, its median turn took about 0.3 s against Luna’s 1.6 s. Luna was more consistent per call, while Jev had occasional slow outliers.
  • Cost: Luna cost about 2–3× as much for the same questions. Both cost fractions of a cent per call.

Bottom line: Jev is the better default. It’s cheaper, faster and more accurate on judgment-heavy questions. Luna didn’t show a clear advantage anywhere we tested.

This is based on a few hundred test questions from word games and one conversational game, so treat it as a strong early signal rather than a general benchmark.

Image understanding is something Jev does not support yet, I can see decisions API quite useful for those cases, eg:

  • Is image appropriate for sites guidelines?
  • Does image contain source code?

This should really give Jev it run for the money.

9

It seems that the decisions api has no cached input tokens. Is that expected?

10

Yes, there’s currently no caching available for the Decisions API. The Decisions API is still in beta, so I wouldn’t be surprised if caching is introduced in the future.

With gpt-6-luna, input costs $0.10 per 1M tokens. You pay only for input tokens: there are no cache-read, cache-write, or output-token charges.

11

Thank you, that is unfortunate as it makes classification tasks economically unattractive to switch from luna responses api to decisions api. I hope caching is enabled soon :crossed_fingers:

Has anyone been able to test this ?

13

Yes, I’m implementing a small test project where GPT models play a real-time 3D game.

On a slow internet connection of around 3 Mbps, the Decisions API responds to image inputs in about 0.8 seconds. There’s no caching of input tokens.

Is there anything specific you would like to know?

14

I’ve been playing around with it, to see if it’s susceptible to the same issues I found with Jev. One big finding is that there is quite a difference between predicate and choice based formats. The predicate seems to follow more closely expected results, while choice seems to put too much probability mass on the most likely outcome.

As an example, I ran a classic loaded coin test, where the coin was biased to show heads 70% of the time. Using predicate setup, across 1000 trials, I get heads 70% of the time - as expected.

When I setup the question as a choice, I would expect heads 70% of the time, but instead I got it 98% of the time.

15

So it’s optimizing for the most likely outcome and then it should actually be predicting Heads 100% of the time?
And it’s clearly not estimating the probability of the next flip being heads at ~70%.

16

Exactly, it’s not estimating the true posterior at all.

17

So in the game demo Romain showed, it doesnt matter, because you want it optimise for the most likely choice (the “opening” on the road). But let’s say you make the road wider, and you give it multiple openings, which one will it choose? And let’s say only one of those openings will lead to an optimal next-step outcome, and lets say Luna/Decisions can “see” that, will it choose the right one?

EDIT (can’t do more than 2 consecutive replies):folded_hands:

The plot thickens - choice order changes the probabilities :scream:

Should I make a dedicated post regarding the results?

I mean to me this is unreliable, but Jev isn’t any better.

18

Please go ahead.
I’m also looking into the best approach here to explore the probability distribution apart from predicting the most likely outcome.

19

My verdict right now is don’t use choice questions.

20

When it will be available on Azure or AWS?

21

I think that’s normally the question for Microsoft and AWS, if you have reps/account people there you can ask them.