跳到正文
原文
Hacker News 热门(buzzing.cc 中文翻译)· Hacker News 热门(buzzing.cc 中文翻译)·· 10 天前AI 评分55

Ollaya 发布,本地运行开源决策模型并兼容 TypeSafe API

AI 导读

Ollaya 发布,一个开源的本地决策模型运行工具,对文本或 JSON 的类型化问题返回毫秒级校准答案。支持 laya、decider、nli、gliclass 等开放权重模型,兼容 TypeSafe 的 /v1/systemone 和 /v1/models 接口,官方 TypeSafe Python SDK 0.7.1 可直接使用。

正文

Ollaya and Ollama

Inspired by Ollama, measured side by side.

Ollaya borrows Ollama's design: one binary, pull and run, a local API. Ollama 0.35 now serves two decision models too, Nimble and Tev1, so we ran both on the same RTX 5090 over Bespoke Labs' public benchmark: 3,880 human-labeled questions from 13 datasets, scored with Bespoke's own code. Ollaya leads on accuracy, calibration and the models it runs; on the same Nimble weights, Ollama is faster.

Same GPU, same 3,880 human-labeled questions · accuracy, higher is better; calibration error and latency, lower is better
  • Ollaya 0.8.0
  • Ollama 0.35.0
  • winnow:12bon Ollaya0.7730.14160 ms
  • decider:4bon Ollaya0.7560.043188 ms
  • kev:9bon Ollaya0.7530.058229 ms
  • nimbleon Ollama 0.350.7490.122210 ms
  • nimble:9bon Ollaya0.7480.022310 ms
  • tev1:4bon Ollama 0.350.7470.075406 ms
  • winnow:e4bon Ollaya0.7340.07143 ms
  • decider:2bon Ollaya0.7030.08798 ms
  • tev1:0.8bon Ollama 0.350.6390.12961 ms
  • laya:multilingualon Ollaya0.5790.15514 ms
OllayaOllama 0.35
Decision modelsOllaya16 families: encoders (laya, nli, gliclass, von) and decoders (winnow, clef, kev, decider, nimble, jeb, jeeves, cygnet and more)Ollama 0.35Nimble and Tev1, decoders only
Small encoders (milliseconds, CPU-friendly)Ollayalaya, nli, gliclass, vonOllama 0.35None
ProbabilitiesOllayaCalibrated with each author's fitted temperature, refittable in a ModelfileOllama 0.35Raw softmax; documented as uncalibrated
Options per questionOllayaUp to 255, as TypeSafeOllama 0.35Up to 26
Questions per requestOllayaUp to 256, as TypeSafeOllama 0.35Up to 64
TypeSafe endpointsOllaya/v1/systemone, /v1/decisions, /v1/modelsOllama 0.35/v1/systemone
Language routingOllayalaya picks English or multilingual per requestOllama 0.35None
WeightsOllayaThe author's files, pinned by commit and sha256Ollama 0.35Converted to GGUF and re-hosted

Bespoke Labs' public benchmark, one request at a time with Ollaya 0.8.0 and Ollama 0.35.0. Accuracy is the mean over the 13 datasets; ECE is the calibration error of the top probability; latency is the median request, HTTP included. Ollama runs Nimble as Q8_0 on llama.cpp, Ollaya in fp32 with the author's temperature.

All results, every machine, with the raw data

Fast and accurate

Close to Jev's accuracy, in under 100 ms.

A decision model answers in a single forward pass, with no token-by-token generation. On an RTX 4090, winnow:e4b answers a five-question request in 89 ms end to end, and scores 0.722 on typed decisions against 0.738 for TypeSafe's hosted Jev. Smaller models such as laya answer in about 10 ms, and run well on a CPU.

Accuracy and speed of every model · typed-decisions accuracy, higher is better; latency, lower is better
  • Ollaya
  • TypeSafe's hosted Jev

ModelAccuracyLatency

  • TypeSafe Jevhosted API0.738236–276 ms
  • winnow:e4b0.72289 ms
  • kev:9b0.722498 ms
  • clef:flash0.703532 ms
  • winnow:12b0.702131 ms
  • cygnet:12b0.683202 ms
  • decider:4b0.680520 ms
  • jeeves:9b0.680838 ms
  • kev:4b0.669354 ms
  • nimble:9b0.6652297 ms
  • jevk5:4b0.625105 ms
  • decider:2b0.591190 ms
  • nli0.54820 ms
  • decider:0.8b0.506155 ms
  • gliclass0.47715 ms
  • kev:0.8b0.460128 ms
  • von0.44723 ms
  • laya:en0.36110 ms
  • clm:8bquestions cached0.357149 ms

Accuracy: the typed-decisions test split (400 states, 2,000 questions) against the majority label; Jev's from Winnow's report on the same questions. laya:typed-decisions and jeb were trained on this dataset, so they are left out. Latency: the median five-question request through the HTTP API on an RTX 4090 (clm with its questions cached); Jev's is the hosted API in third-party benchmarks (AbdelStark/jev-benchmarks, nibzard/decision-model-benchmark), network included, so compare orders of magnitude.

Drop-in compatible

Speaks TypeSafe's API.

Ollaya serves /v1/systemone and /v1/models with TypeSafe's request and response shapes. The official TypeSafe Python SDK 0.7.1 works unchanged against a local server.

Request

# Point the TypeSafe SDK at Ollaya
export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local        # any value works
export TYPESAFE_DEFAULT_MODEL=winnow:e4b

# …or call the compatible endpoint directly
curl http://localhost:11435/v1/systemone -d '{
    "model": "winnow:e4b",
    "state": "Can I get an invoice for last month?",
    "questions": {
      "intent": {
        "type": "choice",
        "instructions": "What does the customer want?",
        "criteria": {
          "invoice": "Needs an invoice or receipt",
          "refund": "Wants money back",
          "other": "Anything else"
        }
      }
    }
  }'

Response

{
  "model": "winnow:e4b",
  "answers": {
    "intent": {
      "type": "choice",
      "choice": "invoice",
      "confidence": 0.9801,
      "probabilities": {
        "invoice": 0.9868,
        "refund": 0.0026,
        "other": 0.0106
      }
    }
  },
  "usage": {
    "input_tokens": 120,
    "output_tokens": 0
  }
}

TypeSafe compatibility guide

Open models

Open weights, ready to pull.

Pick by what you need: winnow:e4b balances accuracy and speed best, laya is the fastest and runs well on a CPU, kev and decider scale up to 9B and 4B, von reads up to 8,192 tokens, and qwen3guard screens text for safety. The models page shows each one’s accuracy and speed.

Browse all models

Your data stays yours

Private by default.

Tickets, emails and user messages are often the most sensitive data you have. With Ollaya they are scored where they already live.

Platforms

Runs where you work.

A desktop app and a command line for macOS, Windows and Linux, and a Docker image for servers. Every model runs on the CPU; an NVIDIA GPU on Linux, Windows, WSL 2 or Docker takes a request down to milliseconds.

PlatformDesktop appCommand lineGPU
macOSApple silicon, macOS 14+Desktop appMenu bar app.dmgCommand lineInstall scriptGPUApple GPULaya and NLI on MLX
Windows10 and 11, x64Desktop appDesktop app.exe or .msiCommand linePowerShell scriptGPUNVIDIA, CUDA 13 or 12Command line
Linuxx86-64Desktop appDesktop appAppImage, .deb, .rpmCommand lineInstall scriptsystemd serviceGPUNVIDIA, CUDA 13 or 12
LinuxARM64Desktop appNot availableCommand lineInstall scriptsystemd serviceGPUCPU only
WSL 2Linux on WindowsDesktop appNot availableCommand lineInstall scriptSame as LinuxGPUNVIDIA, CUDA 13 or 12
Dockeramd64 and arm64Desktop appNot availableCommand lineImage on GHCRGPUNVIDIA, CUDA 13 or 12:cuda and :cuda12, amd64

Install for your platform

NVIDIA GPUs need driver R525 or newer; the install scripts fetch the CUDA libraries only when they find one. On a Mac, laya and nli run on the Apple GPU through MLX; other models, AMD and Intel GPUs, and the Windows and Linux desktop apps without the command line installed use the CPU.

来源:Hacker News 热门(buzzing.cc 中文翻译) · ollaya.dev