llama.cpp adds support for decision models that score options instead of writing text
The llama.cpp server now supports decision models through a new endpoint. These models read an input once and return a probability for each option you give them.

The ggml-org team announced on October 2 that the llama.cpp server supports decision models through a /v1/systemone endpoint. Users send a state, such as text, JSON or a screenshot, plus typed questions, and the model returns a probability for each option in one forward pass.
The post explains that the API follows the System One format introduced with TypeSafe's Jev model. It lists supported open models ranging from 144 million to 27 billion parameters, and suggests uses like routing requests, moderating content and checking an agent's steps.