simonw
9 hours ago
curl https://api.openai.com/v1/decisions \
-H "Authorization: Bearer $(llm keys get openai)" \
-H "Content-Type: application/json" \
--data '
{
"model": "gpt-6-luna",
"input": [{
"role": "user",
"content": [
{"type": "input_text", "text": "I am angry about the new product feature"}
]
}],
"questions": [{
"type": "predicate",
"name": "complaint",
"instructions": "Is this a complaint?"
}, {
"type": "predicate",
"name": "compliment",
"instructions": "Is this a compliment?"
}]
}'
Returned: {
"model": "gpt-6-luna",
"answers": [
{
"type": "predicate",
"name": "complaint",
"probability": 0.91
},
{
"type": "predicate",
"name": "compliment",
"probability": 0.06
}
],
"usage": {
"input_tokens": 310,
"input_tokens_details": {
"cached_tokens": 0,
"cache_write_tokens": 0
},
"output_tokens": 0,
"output_tokens_details": {
"reasoning_tokens": 0
},
"total_tokens": 310
}
}
That https://api.openai.com/v1/decisions endpoint is notable because usually when OpenAI define an endpoint like that it ends up as a defecto standard for other providers.(I turned this all into a new llm plugin: https://github.com/simonw/llm-openai-decisions)
Bassilisk
2 hours ago
>defecto standard for other providers.
I know it's just a little typo but it made my morning :)
chupchap
6 hours ago
How is this different from the categorisation models from ML era?
mogili
4 hours ago
These are essentially zero-shot classifiers; they don't need to be trained for a specific classification task. You could include some natural language context on the rules for classification and it should get good enough accuracy.
sethaurus
5 hours ago
The pitch is that it's a fully-general model, so you can skip training/tuning/selecting a particular categorisation model for each task.
chupchap
5 hours ago
That's great! So someone finally built the zero-shot model from the sales decks of 2015 =D
ehe78qhe
2 hours ago
Specifically, it happened a few weeks ago when Typesafe released Jev; this is OpenAI's competitor to Typesafe.
weird-eye-issue
an hour ago
This isn't really anything new it just seems like a new API but you could do the exact same thing with just a little bit of prompt engineering all the way back when GPT-3 was first released. Am I missing something?
sweetjuly
an hour ago
No amount of prompt engineering will give you the true probabilities for the model producing a certain response; this is something you can only get by inspecting the internal state at inference time.
weird-eye-issue
11 minutes ago
For most use cases is that actually needed though? Just having it choose between predefined responses seems like enough but I'm curious about specific use cases because I do feel like I'm missing something
ehe78qhe
3 minutes ago
This is useful for classification problems; any time you need to write software that looks at some fuzzy data and needs to make a probabilistic decision. It's far more cost-efficient and performant to use this type of model instead of an LLM.
Before now you had to train a model on your specific classification problem, now these new models don't require any specific training at all to do pretty well on novel problems.
Closi
an hour ago
It's much faster and cheaper (an order of magnitude).
And theoretically will give you better answers statistically as it's calibrated.