我们在RTX 3090上运行Qwen3.5-27B,获得了207 tok/s的性能
开发者在单张RTX 3090显卡上成功运行Qwen3.5-27B模型,实现了每秒207个token的生成速度。该项目已在GitHub平台开源,展示了消费级硬件运行270亿参数大语言模型的高性能潜力。相关成果在Hacker News获得105个点赞,引发技术社区对本地大模型部署效率与优化方案的关注。
Lucebox 把 Qwen3.5-27B 在 3090 上推到了 207 tok/s,靠的是定制投机解码和 KV 压缩,做本地推理的开发者值得照着配一遍,虽然调试门槛不低。
Speculative inference for heterogeneous machines and consumer GPUs.
Custom kernels, speculative prefill and decoding, tuned for each model and hardware target.
Inference Engine Optimizations
| Optimization | Measured setup | Result |
|---|---|---|
| DFlash2 | Qwen 3.8 27B on one R9700 | 208.1 tok/s average, 227.8 tok/s peak |
| DSpark | DeepSeek V4 on Strix Halo, native top-6 | 32.7 tok/s high-acceptance median; 27.9 tok/s mixed-eval average |
| PFlash + KVFlash | Laguna XS 2.1 33B at 256K on RTX 3090 | 6.1× prefill, 411 s to 67.3 s |
| Luce Spark | Laguna XS.2 33B on RTX 3090 | ~100 tok/s in 14.6 GiB |
| KVFlash | Laguna XS 2.1 33B at 256K on RTX 3090 | 152.3 tok/s with an 8K pool |
| Heterogeneous execution | DeepSeek V4 on R9700 + Strix Halo | 86 tok/s decode; 788 tok/s prefill at 2K |
| Paged attention + continuous batching | Qwen 3.8 27B + DFlash2 on R9700; DeepSeek V4 Flash AR on Strix Halo | 300.9 tok/s total at 5 clients (Qwen); 48.4 tok/s output-window at 4 clients (DeepSeek) |
| Three-tier experts | DeepSeek V4.1 Flash on R9700 + Strix Halo + SSD | 25-26 tok/s decode on code; 118-121 tok/s prefill at 4K, at the profile's 128K context |
| Megakernel | Qwen 3.5 0.8B on RTX 3090 | 413 tok/s, 1.87 tok/J |
| Vision (image input) | Qwen 3.8 27B vision on R9700 + DeepSeek V4 Flash Vision on Strix Halo, together | 3.2× the image-question throughput of a DGX Spark (58 vs 18 a minute); 2.4× faster at 8 users with the same model file |
Supported Models and Drafters
Each badge opens the exact weights used by the measured setup. Drafter badges open the published quant, or the source checkpoint when conversion is required.
| Model | 🤗 Weights | Optimization (blog) | Phase | Result |
|---|---|---|---|---|
| Qwen 3.5 0.8B | Prefill + decode | 1.9× prefill; 1.55× decode | ||
| Qwen 3.8 27B on R9700 |
Decode | 6.4× vs Lucebox AR; 3.8× vs llama.cpp with the same drafter | ||
| Laguna XS 2.1 33B | Prefill | 6.1×, 411 s to 67.3 s at 256K | ||
| Laguna XS 2.1 33B | Decode | 1.7× at 256K | ||
| Gemma 4 26B-A4B | Decode | 1.31× | ||
| Gemma 4 31B IT | Decode | 3.2× | ||
| DeepSeek V4 Flash on Strix Halo |
Decode | 42 tok/s at 8K and 39 tok/s on code and math with the plain launch (PR #729) | ||
| DeepSeek V4.1 Flash R9700 + Strix Halo + SSD, --profile ds41-lucebox |
Decode | 25.4-25.8 tok/s on code, 16.8-17.5 tok/s after a 4K prompt, at the profile's 128K context (guide) | ||
| Ling 3.0 Flash 124B-A5.1B on DGX Spark |
Autoregressive | Decode | 34.6 tok/s median AR | |
| Qwen 3.8 27B Vision on R9700 |
Image questions | 2.4× vs llama.cpp on DGX Spark at 8 users (10.5 s vs 25.4 s); 2.1× for one user | ||
| DeepSeek V4 Flash Vision encoder on R9700, 4 users at once |
Image questions | 1.5× sooner first token with 16 images when the R9700 encodes |
Tested Machines (GPU/APU)
The engine is not tied to one reference card. NVIDIA architectures are selected by CMake; HIP builds should target the device's exact gfx architecture.
| Architecture | Hardware | Runtime | Details | |
|---|---|---|---|---|
![]() |
RDNA4 gfx1201 |
Radeon AI PRO R9700 | ROCm 7.2 | Qwen 3.8 R9700 quick start |
![]() |
RDNA3.5 gfx1151 |
Ryzen AI MAX+ 395 / Strix Halo | ROCm 7.2 | DeepSeek V4 Strix profile |
![]() |
RDNA3 gfx1100 |
Radeon RX 7900 XT / XTX | ROCm 6+ | DeepSeek V4 dual AMD profile |
![]() |
Ampere sm_86 |
RTX 3090 | CUDA 12+ | Qwen 3.8 NVLink result and Megakernel results |
![]() |
Blackwell sm_120 |
RTX 5090 | CUDA 12.8+ | Qwen 3.8 single-GPU result |
![]() |
Blackwell sm_121 |
DGX Spark / GB10 | CUDA 12.9 | Qwen 3.5 NVFP4 results |
![]() |
Ada sm_89 |
RTX 4090 | CUDA 12+ | Linux and WSL2 community runs |
![]() |
Turing sm_75 |
RTX 2080 Ti | CUDA 12.0 | DFlash results |
![]() |
Volta sm_70, Pascal sm_61 |
V100, P40 | CUDA 12.0 | CUDA quick start |
| Not pictured | Blackwell sm_110 |
Jetson AGX Thor | CUDA 13.0 | Thor quick start |
Single-device results
| Hardware | Model | Measured result |
|---|---|---|
| R9700 | Qwen 3.8 27B UD-IQ4_XS + DFlash2 source | 208.1 tok/s HumanEval average; 227.8 tok/s best request |
| Strix Halo | DeepSeek V4 ROCmFPX MIX Strix + DSpark Q4RMFP4 | 42 tok/s decode and 320 tok/s prefill at 8K, 36 tok/s at 123K, 39 tok/s on code and math, 25 tok/s on prose, all six routed experts, plain launch (PR #729) |
| RTX 5090 | Qwen 3.8 27B | 110.6 tok/s for a 26,758-token prompt and 1,024-token continuation (PR #637) |
Heterogeneous and parallel results
| Hardware | Configuration | Measured result |
|---|---|---|
| 2x RTX 3090 + NVLink | Qwen 3.8 target tensor parallel + DFlash2 | 79.7 tok/s, 2.16× autoregressive decode (PR #637) |
| RX 7900 XT + Strix Halo | DeepSeek V4 with all six experts + DSpark verification width 4 | 45.0 to 47.7 tok/s decode; 111.2 tok/s prefill at 132,981 tokens (PR #604) |
| R9700 + Strix Halo | DeepSeek V4 across both AMD devices | 86 tok/s decode; 788 tok/s prefill at 2K |
These runs use different prompts, quantizations, and inference policies. They show which configurations work; they are not a cross-hardware ranking.
Recommended Setups
See Recommended server setups for the model and hardware matrix, including single-GPU and mixed-GPU profiles.
The DS4 guide also documents the Strix long-context sparse-verifier profile and Qwen3-0.6B PFlash integration. PFlash is lossy prompt compression; keep it off for exact-retrieval and matched true-context benchmarks.
Client Harnesses
harness/ runs Lucebox through popular coding clients and checks server compatibility.
|
Set the server binary and model paths, then run a launcher:
LUCE_SERVER_BIN=server/build/luce_server \ LUCE_TARGET=server/models/Qwen3.8-27B-UD-IQ4_XS.gguf \ LUCE_DRAFT=server/models/draft/qwen38-dflash2-q8_0.gguf \ MAX_CTX=32768 \ harness/clients/run_codex.sh
See the harness guide for setup, no-draft targets, and benchmarks.
Quick Start With Docker
Prebuilt images on GHCR track main. Mount the weights and serve the OpenAI-compatible API on :8000.
Put the target in |
|
Run the image for your GPU:
# NVIDIA docker run --rm --gpus all -p 8000:8080 \ -v "$PWD/server/models:/opt/lucebox-hub/server/models" \ ghcr.io/luce-org/lucebox-hub:cuda12 # AMD docker run --rm --device /dev/kfd --device /dev/dri \ --group-add video --group-add render --security-opt seccomp=unconfined \ -p 8000:8080 -v "$PWD/server/models:/opt/lucebox-hub/server/models" \ ghcr.io/luce-org/lucebox-hub:rocm
The container picks the GPU the model fits on (a discrete card before an integrated one, else the largest) and sizes the context from that GPU's memory. serve is the default command, so luce_server flags can follow the image name directly (or serve); they replace the values the container would pass. --target-device, --max-ctx and --profile all work:
# Show the GPUs, the model, and the device auto placement would use docker run --rm <gpu flags> -v ... ghcr.io/luce-org/lucebox-hub:rocm devices # DeepSeek V4 on Strix Halo with its qualified profile (DSpark drafter in models/draft/) docker run --rm <gpu flags> -p 8000:8080 -v ... ghcr.io/luce-org/lucebox-hub:rocm --profile ds4-strix # Pin a device docker run --rm <gpu flags> -p 8000:8080 -v ... ghcr.io/luce-org/lucebox-hub:rocm --target-device hip:1
Environment variables such as LUCE_TARGET, LUCE_TARGET_DEVICE, LUCE_MAX_CTX and LUCE_ARGS cover the same settings for compose files; see the header of server/scripts/entrypoint.sh.
Run the Server
This quick start runs the R9700 profile above. The complete flag reference is in the server guide.
# build (ROCm 7.2+, RDNA4) git clone --recurse-submodules https://github.com/Luce-Org/lucebox.git cd lucebox cmake -S server -B server/build-hip -G Ninja \ -DCMAKE_BUILD_TYPE=Release \ -DCMAKE_HIP_COMPILER=/opt/rocm/lib/llvm/bin/clang++ \ -DLUCE_GPU_BACKEND=hip \ -DLUCE_HIP_ARCHITECTURES=gfx1201 \ -DGGML_HIP_MMQ_MFMA=ON \ -DGGML_HIP_NO_VMM=ON cmake --build server/build-hip --target luce_server -j"$(nproc)" # target and DFlash2 drafter mkdir -p models huggingface-cli download unsloth/Qwen3.8-27B-GGUF \ Qwen3.8-27B-UD-IQ4_XS.gguf --local-dir models huggingface-cli download incoai/Qwen3.8-27B-DFlash2 --local-dir models/dflash2 python server/scripts/convert_dflash_to_gguf.py \ models/dflash2/model.safetensors models/qwen38-dflash2-f16.gguf python server/scripts/quantize_dflash_draft.py \ models/qwen38-dflash2-f16.gguf models/qwen38-dflash2-q8_0.gguf --scheme q8_0 # launch the measured profile ./server/build-hip/luce_server models/Qwen3.8-27B-UD-IQ4_XS.gguf \ --draft models/qwen38-dflash2-q8_0.gguf \ --draft-block-size 16 --max-ctx 131072 \ --cache-type-k q8_0 --cache-type-v q8_0 \ --port 8216 curl -s http://127.0.0.1:8216/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{"messages":[{"role":"user","content":"Write a Python LRU cache."}], "max_tokens":256,"temperature":0}'
To serve up to N concurrent requests, use this launch command with the same Qwen target and DFlash2 drafter. Set N to the desired concurrency (5 below). Qwen automatically sizes the shared KV pool from available GPU memory.
N=5
./server/build-hip/luce_server models/Qwen3.8-27B-UD-IQ4_XS.gguf \
--draft models/qwen38-dflash2-q8_0.gguf \
--draft-block-size 16 --max-ctx 16384 \
--paged-attention --max-concurrency "$N" \
--cache-type-k q8_0 --cache-type-v q8_0 \
--port 8216See Continuous batching in Lucebox for Qwen and DeepSeek V4 Flash results, latency measurements, and launch settings.
Documentation
| Topic | Guide |
|---|---|
| Recommended model and hardware profiles | Recommended setups |
| Runtime parameters | Server parameter reference |
| OpenAI Chat Completions, Responses, and Anthropic Messages | API reference |
| CUDA, HIP, and mixed-device placement | Mixed-backend guide |
| DeepSeek V4 single-device and heterogeneous profiles | DeepSeek V4 guide |
| DeepSeek V4.1 Flash on R9700 + Strix Halo + SSD | DeepSeek V4.1 guide |
| Image input (Qwen3.8, DeepSeek V4 Flash Vision) | Image input guide |
| Environment variables | Environment reference |
| Server internals | Architecture |
| Client integration and qualification | Harness guide |
| Server engine components | Engine components |
Benchmarks stay with each implementation: DFlash, PFlash, Spark, KVFlash, and Megakernel.
Tutorials
Video tutorials for each optimization and the harness setup.
| Luce Spark ▶ YouTube |
Luce DFlash ▶ YouTube |
Luce Turboquant ▶ YouTube |
| OpenClaw harness setup ▶ YouTube |
Luce PFlash ▶ YouTube |
Luce Megakernel ▶ YouTube |
| Luce KVFlash ▶ YouTube |
The Lucebox Machine
Local AI should be the default, not a privilege. Private data, no per-token bill, no vendor lock-in. Lucebox pairs the R9700 with Strix Halo and ships this open engine ready to run.
See the hardware and current benchmarks at lucebox.com.
Request for Contributions
We welcome focused contributions to CUDA and HIP kernels, speculative inference, support for more consumer GPUs and APUs, performance benchmarks, and client harnesses.
Citation
@software{lucebox_2026, title = {Lucebox: Speculative inference for heterogeneous consumer hardware}, author = {Lucebox}, url = {https://github.com/Luce-Org/lucebox}, year = {2026} }
Community
- Discord: discord.gg/yHfswqZmJQ
- Website: lucebox.com
- Issues: github.com/Luce-Org/lucebox/issues
- Blog: lucebox.com/blog
来源:Hacker News 热门(buzzing.cc 中文翻译) · github.com











