Ollaya 发布,本地运行开源决策模型并兼容 TypeSafe API
Ollaya 发布,一个开源的本地决策模型运行工具,对文本或 JSON 的类型化问题返回毫秒级校准答案。支持 laya、decider、nli、gliclass 等开放权重模型,兼容 TypeSafe 的 /v1/systemone 和 /v1/models 接口,官方 TypeSafe Python SDK 0.7.1 可直接使用。
Ollaya and Ollama
Inspired by Ollama, measured side by side.
Ollaya borrows Ollama's design: one binary, pull and run, a local API. Ollama 0.35 now serves two decision models too, Nimble and Tev1, so we ran both on the same RTX 5090 over Bespoke Labs' public benchmark: 3,880 human-labeled questions from 13 datasets, scored with Bespoke's own code. Ollaya leads on accuracy, calibration and the models it runs; on the same Nimble weights, Ollama is faster.
- Ollaya 0.8.0
- Ollama 0.35.0
- winnow:12bon Ollaya0.7730.14160 ms
- decider:4bon Ollaya0.7560.043188 ms
- kev:9bon Ollaya0.7530.058229 ms
- nimbleon Ollama 0.350.7490.122210 ms
- nimble:9bon Ollaya0.7480.022310 ms
- tev1:4bon Ollama 0.350.7470.075406 ms
- winnow:e4bon Ollaya0.7340.07143 ms
- decider:2bon Ollaya0.7030.08798 ms
- tev1:0.8bon Ollama 0.350.6390.12961 ms
- laya:multilingualon Ollaya0.5790.15514 ms
| Ollaya | Ollama 0.35 | |
|---|---|---|
| Decision models | Ollaya16 families: encoders (laya, nli, gliclass, von) and decoders (winnow, clef, kev, decider, nimble, jeb, jeeves, cygnet and more) | Ollama 0.35Nimble and Tev1, decoders only |
| Small encoders (milliseconds, CPU-friendly) | Ollayalaya, nli, gliclass, von | Ollama 0.35None |
| Probabilities | OllayaCalibrated with each author's fitted temperature, refittable in a Modelfile | Ollama 0.35Raw softmax; documented as uncalibrated |
| Options per question | OllayaUp to 255, as TypeSafe | Ollama 0.35Up to 26 |
| Questions per request | OllayaUp to 256, as TypeSafe | Ollama 0.35Up to 64 |
| TypeSafe endpoints | Ollaya/v1/systemone, /v1/decisions, /v1/models | Ollama 0.35/v1/systemone |
| Language routing | Ollayalaya picks English or multilingual per request | Ollama 0.35None |
| Weights | OllayaThe author's files, pinned by commit and sha256 | Ollama 0.35Converted to GGUF and re-hosted |
Bespoke Labs' public benchmark, one request at a time with Ollaya 0.8.0 and Ollama 0.35.0. Accuracy is the mean over the 13 datasets; ECE is the calibration error of the top probability; latency is the median request, HTTP included. Ollama runs Nimble as Q8_0 on llama.cpp, Ollaya in fp32 with the author's temperature.
All results, every machine, with the raw data
Fast and accurate
Close to Jev's accuracy, in under 100 ms.
A decision model answers in a single forward pass, with no token-by-token generation. On an RTX 4090, winnow:e4b answers a five-question request in 89 ms end to end, and scores 0.722 on typed decisions against 0.738 for TypeSafe's hosted Jev. Smaller models such as laya answer in about 10 ms, and run well on a CPU.
- Ollaya
- TypeSafe's hosted Jev
ModelAccuracyLatency
- TypeSafe Jevhosted API0.738236–276 ms
- winnow:e4b0.72289 ms
- kev:9b0.722498 ms
- clef:flash0.703532 ms
- winnow:12b0.702131 ms
- cygnet:12b0.683202 ms
- decider:4b0.680520 ms
- jeeves:9b0.680838 ms
- kev:4b0.669354 ms
- nimble:9b0.6652297 ms
- jevk5:4b0.625105 ms
- decider:2b0.591190 ms
- nli0.54820 ms
- decider:0.8b0.506155 ms
- gliclass0.47715 ms
- kev:0.8b0.460128 ms
- von0.44723 ms
- laya:en0.36110 ms
- clm:8bquestions cached0.357149 ms
Accuracy: the typed-decisions test split (400 states, 2,000 questions) against the majority label; Jev's from Winnow's report on the same questions. laya:typed-decisions and jeb were trained on this dataset, so they are left out. Latency: the median five-question request through the HTTP API on an RTX 4090 (clm with its questions cached); Jev's is the hosted API in third-party benchmarks (AbdelStark/jev-benchmarks, nibzard/decision-model-benchmark), network included, so compare orders of magnitude.
Drop-in compatible
Speaks TypeSafe's API.
Ollaya serves /v1/systemone and /v1/models with TypeSafe's request and response shapes. The official TypeSafe Python SDK 0.7.1 works unchanged against a local server.
Request
# Point the TypeSafe SDK at Ollaya
export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local # any value works
export TYPESAFE_DEFAULT_MODEL=winnow:e4b
# …or call the compatible endpoint directly
curl http://localhost:11435/v1/systemone -d '{
"model": "winnow:e4b",
"state": "Can I get an invoice for last month?",
"questions": {
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": {
"invoice": "Needs an invoice or receipt",
"refund": "Wants money back",
"other": "Anything else"
}
}
}
}'Response
{
"model": "winnow:e4b",
"answers": {
"intent": {
"type": "choice",
"choice": "invoice",
"confidence": 0.9801,
"probabilities": {
"invoice": 0.9868,
"refund": 0.0026,
"other": 0.0106
}
}
},
"usage": {
"input_tokens": 120,
"output_tokens": 0
}
}Open models
Open weights, ready to pull.
Pick by what you need: winnow:e4b balances accuracy and speed best, laya is the fastest and runs well on a CPU, kev and decider scale up to 9B and 4B, von reads up to 8,192 tokens, and qwen3guard screens text for safety. The models page shows each one’s accuracy and speed.
- winnowDecision models by EldanRing, fine-tuned from Google's Gemma 4 and published as GGUF. Winnow reads the answer labels' logits after its own prompt; Ollaya runs the author's file on llama.cpp, on NVIDIA GPUs, Apple silicon or the CPU.7.5b · 12b
- layaOpen decision models from Convai Innovations. Typed, calibrated answers to choice, score and yes/no questions in a single forward pass, in English and 100+ languages.322m · 421m
- deciderDecoder decision models by Mapika on Qwen3.5: the answer is read from option-letter logits in one forward pass. decider:4b scores 0.680 on typed decisions, decider:2b 0.591.0.75b · 1.9b · 2.2b · 4.2b
- kevDecision models by Jared Palmer: a LoRA on a Qwen3.5 base plus a pointer head that scores every option at its own span, in one forward pass per question. Calibrated with Kev's own temperature.0.76b · 4.2b · 7.9b
- nliZero-shot classifiers by Moritz Laurer: every option becomes a hypothesis scored for entailment. The most accurate encoder model on typed decisions in our tests.396m · 435m
- gliclassInstruction-following zero-shot classifier by Knowledgator: all options of a question are scored in one pass, so cost barely grows with the number of options.439m
- qwen3guardSafety guard by the Qwen team: is a text safe, controversial or unsafe, and which unsafe category? It answers its own built-in questions, in 119 languages, in one forward pass.0.6b
- decisionDecision models by the vLLM Semantic Router contributors: a fully fine-tuned Qwen3.5 backbone plus an endpoint head that scores every option at its own last token against the question, in one forward pass per question. 16k-token rows.0.75b
Your data stays yours
Private by default.
Tickets, emails and user messages are often the most sensitive data you have. With Ollaya they are scored where they already live.
Platforms
Runs where you work.
A desktop app and a command line for macOS, Windows and Linux, and a Docker image for servers. Every model runs on the CPU; an NVIDIA GPU on Linux, Windows, WSL 2 or Docker takes a request down to milliseconds.
| Platform | Desktop app | Command line | GPU |
|---|---|---|---|
| macOSApple silicon, macOS 14+ | Desktop appMenu bar app.dmg | Command lineInstall script | GPUApple GPULaya and NLI on MLX |
| Windows10 and 11, x64 | Desktop appDesktop app.exe or .msi | Command linePowerShell script | GPUNVIDIA, CUDA 13 or 12Command line |
| Linuxx86-64 | Desktop appDesktop appAppImage, .deb, .rpm | Command lineInstall scriptsystemd service | GPUNVIDIA, CUDA 13 or 12 |
| LinuxARM64 | Desktop appNot available | Command lineInstall scriptsystemd service | GPUCPU only |
| WSL 2Linux on Windows | Desktop appNot available | Command lineInstall scriptSame as Linux | GPUNVIDIA, CUDA 13 or 12 |
| Dockeramd64 and arm64 | Desktop appNot available | Command lineImage on GHCR | GPUNVIDIA, CUDA 13 or 12:cuda and :cuda12, amd64 |
NVIDIA GPUs need driver R525 or newer; the install scripts fetch the CUDA libraries only when they find one. On a Mac, laya and nli run on the Apple GPU through MLX; other models, AMD and Intel GPUs, and the Windows and Linux desktop apps without the command line installed use the CPU.
来源:Hacker News 热门(buzzing.cc 中文翻译) · ollaya.dev