
GPU 服务器模型部署与技术总结
项目地址:https://github.com/Niko1221/Strata
采集时间:2026-07(实时 SSH 采集)
服务器配置:Debian GNU/Linux 13(2× RTX 4090 24GB,CUDA 13.0)
CPU:Montage Jintide(R) C5218R
内存:64GB
硬盘:HDD:100GB SSD:500GB
虚拟化:KVM
GPU0:SGLang 0.5.20(lmsysorg/sglang:latest),AWQ W4A16 量化
GPU1:Strata 推理引擎(strata:local),IQ2_XS 超压缩量化
一、服务器硬件
| 项目 | GPU 0 | GPU 1 |
|---|---|---|
| 型号 | NVIDIA GeForce RTX 4090 | NVIDIA GeForce RTX 4090 |
| 显存 | 49140 MiB (48GB) | 49140 MiB (48GB) |
| 当前已用 | 40629 MiB | 39982 MiB |
| 当前可用 | 7962 MiB | 8526 MiB |
| GPU 利用率 | 0%(空闲) | 0%(空闲) |
| 温度 | 48°C | 34°C |
| SM 频率 | 2520 MHz / 3105 MHz | 210 MHz / 3105 MHz |
| 功耗 | 67.40 W / 450 W | 28.56 W / 450 W |
| Persistence Mode | Enabled | Enabled |
| 运行进程 | sglang(lmsysorg/sglang) | strata(/opt/strata/engine/strata) |
两卡均处于空闲态(无活跃推理请求),模型权重常驻显存。
二、GPU1 — Strata 推理引擎(当前)
2.1 启动项目
services: strata: build: context: . dockerfile: Dockerfile args: CUDA_ARCHITECTURES: "89" BUILD_VISION: "1"
image: strata:local container_name: strata
restart: unless-stopped
ports: - "8080:8080"
environment: FAMILY: "qwen" MODEL: "IQ2_XS" CONTEXT: "262144" VISION: "no" KV: "int8"
# Strata 内部使用 GPU1 GPU: "0"
HOST: "0.0.0.0" PORT: "8080"
API_KEY: "sk-strata-a8f93d7c1e4b5f8a9c2d7e6f" LOW_RAM: "auto"
volumes: - ./data:/data
deploy: resources: reservations: devices: - driver: nvidia device_ids: - "1" capabilities: - gpu
ulimits: memlock: soft: -1 hard: -12.2 容器信息
| 项目 | 值 |
|---|---|
| 容器名 | strata |
| 镜像 | strata:local(本地构建,2026-10-01 21:41) |
| 镜像 ID | sha256:e31595f882648325a17c0efa578c751f133c650ab5acc33b08e073425c90e29a |
| 状态 | Up(healthy) |
| 端口映射 | 0.0.0.0:8080 → 8080/tcp(bridge 网络) |
| 数据卷 | /mnt/data/Strata/data → /data(rw) |
| 容器内进程 | /opt/strata/.venv/bin/python /opt/strata/serve/server.py |
| GPU 绑定 | 容器内 --gpu 0 = 宿主机 GPU1 |
2.3 模型与量化
| 项目 | 值 |
|---|---|
| 模型 | Qwen3.8-Flash-Next |
| 量化 | IQ2_XS(~2-bit 超压缩,GGUF 格式) |
| 模型文件 | Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00001-of-00002.gguf + 00002-of-00002.gguf |
| 模型标识 | qwen3.8-flash-next-iq2_xs |
| 视觉能力 | 无(VISION=no,images: false) |
| 显存占用 | 39982 MiB(~39 GB) |
IQ2_XS 是 2-bit 超压缩量化(类似 llama.cpp 的 IQ 系列),比 AWQ W4A16(4-bit)激进一倍。这是 27B 模型能在 24GB 4090 上跑起来的關鍵——2-bit 权重 + int8 KV cache 把总显存压到 ~40GB 以内。
2.4 引擎配置(strata-iq2_xs.json)
{ "exe": "/opt/strata/engine/strata", "args": [ "--pack", "/data/packs/iq2_xs", "--native", "/data/models/IQ2_XS/Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00001-of-00002.gguf", "--ple-gguf", "/data/models/IQ2_XS/Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00002-of-00002.gguf", "--expert-profile", "/opt/strata/data/expert-profile.bin", "--expert-cache", "auto", "--prefill", "auto", "--spec", "4", "--spec-min-p", "0.5", "--mtp", "/data/mtp/rt", "--max-context", "262144", "--kv", "int8", "--kv-resident", "32768" ], "tokenizer": "/data/packs/iq2_xs/tokenizer", "model_name": "qwen3.8-flash-next-iq2_xs", "port": 8080, "gpu": 0, "host": "0.0.0.0", "api_key": "sk-strata-a8f93d7c1e4b5f8a9c2d7e6f"}2.5 关键配置解读
| 参数 | 值 | 技术含义 |
|---|---|---|
--max-context | 262144(256K) | 最大上下文窗口,与 SGLang 部署一致 |
--kv | int8 | KV cache 用 INT8 量化(比 FP16 减半,比 FP8 e5m2 精度略低但更通用) |
--kv-resident | 32768 | 32K tokens 的 KV cache 常驻 GPU,超出部分可换出到 CPU |
--spec | 4 | 投机解码:draft 4 tokens(每次验证 4 个候选 token) |
--spec-min-p | 0.5 | 投机解码最小接受概率阈值 0.5 |
--mtp | /data/mtp/rt | MTP(Multi-Token Prediction)模块,路径 /data/mtp/rt |
--expert-profile | /opt/strata/data/expert-profile.bin | 专家路由配置文件(MoE 模型用) |
--expert-cache | auto | 专家缓存自动管理 |
--prefill | auto | Prefill 批处理大小自动决定 |
--pack | /data/packs/iq2_xs | 量化 pack 目录(含 tokenizer) |
--native | GGUF part 1 | 模型权重文件(第 1 片) |
--ple-gguf | GGUF part 2 | 模型权重文件(第 2 片,PLE = Position-Linear-Expert?) |
2.6 Strata 的投机解码 + MTP
这是 Strata 相比 SGLang 最大的技术差异。 Strata 同时启用了:
- 投机解码(
--spec 4):每次 decode 步骤 draft 4 个 token,然后一次验证 - MTP 模块(
--mtp /data/mtp/rt):模型自带的 Multi-Token Prediction 头
重要发现:之前分析 SGLang 文档时判断”Qwen3.8 不支持 MTP”——那是针对 AWQ W4A16 版本的 Qwen3.8-27B。但 Strata 跑的是 Qwen3.8-Flash-Next(不同变体),这个版本自带 MTP 头(
/data/mtp/rt目录)。所以:
Ar4ikov/Qwen3.8-27B-AWQ-W4A16-ASYM(SGLang 用)→ 无 MTP 头Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS(Strata 用)→ 有 MTP 头同一个 Qwen3.8 家族,不同 checkpoint 的 MTP 支持不同。
2.7 Strata API
| 端点 | 方法 | 说明 |
|---|---|---|
/ | GET | Web UI(内置聊天界面,HTML + CSS + JS) |
/health | GET | 健康检查(无需鉴权) |
/v1/models | GET | OpenAI 兼容模型列表(需 API key) |
/v1/chat/completions | POST | OpenAI 兼容 chat API(需 API key) |
/health 返回(实时):
{ "status": "ok", "max_context": 262144, "model": "qwen3.8-flash-next-iq2_xs", "images": false, "api_key": true, "loaded": true}API Key:sk-strata-a8f93d7c1e4b5f8a9c2d7e6f(配置在 strata-iq2_xs.json 和容器环境变量 API_KEY 中)
使用示例:
# 健康检查(无需 key)curl http://127.0.0.1:8080/health
# 模型列表(需要 key)curl http://127.0.0.1:8080/v1/models \ -H "Authorization: Bearer sk-strata-a8f93d7c1e4b5f8a9c2d7e6f"
# Chat completion(需要 key)curl http://127.0.0.1:8080/v1/chat/completions \ -H "Authorization: Bearer sk-strata-a8f93d7c1e4b5f8a9c2d7e6f" \ -H "Content-Type: application/json" \ -d '{ "model": "qwen3.8-flash-next-iq2_xs", "messages": [{"role": "user", "content": "你好"}], "max_tokens": 256 }'2.8 Strata 架构特点
从容器内文件结构看,Strata 是一个完整的推理引擎 + 服务框架:
/opt/strata/├── engine/│ ├── strata # 推理引擎二进制(38MB)│ └── strata-vision # 视觉推理引擎(73MB,当前未启用)├── serve/│ ├── server.py # HTTP 服务(120KB,FastAPI/Starlette)│ ├── frontend.py # Web UI 后端│ ├── mcp.py # MCP(Model Context Protocol)支持│ ├── chat_template.jinja # 聊天模板│ ├── telemetry.py # 遥测/监控│ └── web/ # 前端静态文件├── src/ # C++ 引擎源码├── include/ # C++ 头文件├── third_party/ # 第三方依赖├── bench/ # 基准测试├── tests/ # 测试└── tools/ └── vision/ # 视觉工具Strata 的技术栈:
- 引擎:C++ 编写(CMake 构建),直接调用 CUDA
- 服务层:Python(FastAPI),包装 C++ 引擎
- 量化:支持 GGUF 格式(IQ2_XS 等 llama.cpp 系列量化)
- 投机解码:引擎内置(
--spec)+ MTP 模块 - MoE 支持:expert profile + expert cache
- MCP 支持:
mcp.py+mcp_fake_server.py(Model Context Protocol,AI agent 工具调用标准) - Web UI:内置聊天界面(
frontend.py+web/)
三、GPU0 — SGLang(当前)
3.1 容器信息
| 项目 | 值 |
|---|---|
| 容器名 | sglang |
| 镜像 | lmsysorg/sglang:latest(= v0.5.20) |
| 状态 | Up 2 days(healthy) |
| 网络 | host 模式(无端口映射,直接用宿主机 30000) |
| 模型 | Ar4ikov/Qwen3.8-27B-AWQ-W4A16-ASYM |
| 量化 | AWQ W4A16(17.75GB) |
| 显存占用 | 40629 MiB |
3.2 启动命令(/proc/1/cmdline 实时读取)
python3 -m sglang.launch_server \ --model-path Ar4ikov/Qwen3.8-27B-AWQ-W4A16-ASYM \ --host 0.0.0.0 \ --port 30000 \ --trust-remote-code \ --enable-multimodal \ --mm-feature-transport cuda_ipc \ --keep-mm-feature-on-device \ --reasoning-parser qwen3 \ --default-chat-template-kwargs '{"enable_thinking": false}' \ --tool-call-parser qwen3_coder \ --context-length 262144 \ --max-total-tokens 229376 \ --kv-cache-dtype fp8_e5m2 \ --mem-fraction-static 0.86 \ --enable-metrics \ --enable-cache-report3.3 SGLang 完整配置(/get_server_info 实时返回)
| 配置项 | 值 | 说明 |
|---|---|---|
model_path | Ar4ikov/Qwen3.8-27B-AWQ-W4A16-ASYM | AWQ W4A16 量化 |
context_length | 262144 | 最大上下文 256K |
max_total_tokens | 229376 | KV cache 总容量上限 |
kv_cache_dtype | fp8_e5m2 | FP8 KV cache(动态范围 ±57344) |
mem_fraction_static | 0.86 | 86% 显存给 weights+KV |
enable_multimodal | true | 支持图像/视频 |
dtype | auto | 自动(AWQ 模型自动选) |
quantization | null | 未显式指定(从模型路径自动识别 AWQ) |
tp_size | 1 | 单卡 |
pp_size | 1 | 无流水线并行 |
dp_size | 1 | 无数据并行 |
chunked_prefill_size | 4096 | Prefill 分块 4096 tokens |
max_prefill_tokens | 16384 | 单次 prefill 最大 16K tokens |
schedule_policy | fcfs | 先进先出调度 |
schedule_conservativeness | 1.0 | 最保守(不抢占) |
retraction_policy | length | 按长度回退 |
page_size | 1 | token 级 KV page(prefix cache 精确匹配) |
c128_page_size | 16 | C128 模式 page size |
swa_full_tokens_ratio | 0.8 | 滑动窗口注意力 ratio |
radix_eviction_policy | lru | 前缀缓存 LRU 驱逐 |
max_running_requests | null(auto=13) | 自动推导最大并发 13 |
max_queued_requests | null | 无排队上限 |
enable_mixed_chunk | false | 不混合 chunked prefill |
disable_overlap_schedule | false | 启用 overlap scheduler |
num_continuous_decode_steps | 1 | 每步 1 个 decode |
watchdog_timeout | 300s | 看门狗超时 |
sleep_on_idle | false | 空闲时不睡眠 |
enable_dp_attention | false | 无 DP attention |
enable_priority_scheduling | false | 无优先级调度 |
3.4 SGLang 架构特性
Qwen3.8-27B 是混合架构模型:
- GDN(Gated Delta Network):48 层线性注意力(
mamba_backend: triton) - Full Attention:16 层标准注意力(
attention_backend: flashinfer) - 总计 64 层
在 sm_89(RTX 4090)上的 kernel 选择:
- GDN → triton(唯一可用的 GDN kernel)
- Attention → flashinfer(唯一支持 FP8 KV 的 Ada 后端;FA3 是 Hopper-only,FA4 不支持 FP8 KV)
3.5 SGLang 性能(之前实测)
| 并发 | 单请求速度 | 总吞吐 | 效率 |
|---|---|---|---|
| 1 | 54.6 tok/s | 54.5 tok/s | 1.0× |
| 2 | 42.1 tok/s | 84.2 tok/s | 1.54× |
| 3 | ~34 tok/s | ~101 tok/s | ~1.85× |
- 单流天花板 ~55 tok/s(17.75GB AWQ / ~1TB/s 带宽,memory-bound)
- CUDA graph decode bs:
[1,2,4,8,12,16,24,32] - 无 retract、无 OOM、无排队
3.6 SGLang 监控
SGLang 暴露 Prometheus 指标(--enable-metrics),端点 http://127.0.0.1:30000/metrics(host 网络)。
核心指标(前缀 sglang:):
sglang:num_running_reqs、sglang:num_queue_reqssglang:token_usage、sglang:cache_hit_ratesglang:gen_throughputsglang:time_to_first_token_seconds(TTFT)sglang:time_per_output_token_seconds(TPOT)sglang:e2e_request_latency_seconds
四、GPU1 历史配置 — SGLang(sglang2,已替换)
GPU1 之前运行 SGLang(容器名 sglang2),现已替换为 Strata。compose 文件保留在 /mnt/data/sglang-1/docker-compose.yml:
services: sglang: image: lmsysorg/sglang:v0.5.20 container_name: sglang2 volumes: - /mnt/data/sglang/cache/huggingface:/root/.cache/huggingface restart: always ports: - "8080:30000" deploy: resources: reservations: devices: - driver: nvidia device_ids: ["1"] capabilities: [gpu] environment: HF_TOKEN: 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 NVIDIA_DRIVER_CAPABILITIES: compute,utility entrypoint: python3 -m sglang.launch_server command: > --model-path Ar4ikov/Qwen3.8-27B-AWQ-W4A16-ASYM --host 0.0.0.0 --port 30000 --trust-remote-code --enable-multimodal --mm-feature-transport cuda_ipc --keep-mm-feature-on-device --reasoning-parser qwen3 --default-chat-template-kwargs '{"enable_thinking": false}' --tool-call-parser qwen3_coder --context-length 262144 --max-total-tokens 229376 --kv-cache-dtype fp8_e5m2 --mem-fraction-static 0.86 --enable-metrics --enable-cache-report ulimits: memlock: -1 stack: 67108864 ipc: host healthcheck: test: ["CMD-SHELL", "curl -f http://localhost:30000/health || exit 1"] interval: 30s timeout: 10s retries: 5与 GPU0 的 SGLang 配置完全相同(同样的模型、同样的参数),只是 GPU 绑定不同(GPU1 vs GPU0)和网络模式不同(bridge 8080→30000 vs host 30000)。
五、Strata vs SGLang 技术对比
| 维度 | SGLang(GPU0) | Strata(GPU1) |
|---|---|---|
| 引擎语言 | Python + Triton/FlashInfer kernels | C++(CMake)+ Python 服务层 |
| 量化 | AWQ W4A16(4-bit,17.75GB) | IQ2_XS(2-bit,~9GB 权重) |
| 最大上下文 | 262144(256K) | 262144(256K) |
| KV Cache | FP8 e5m2(动态范围 ±57344) | INT8(更通用) |
| KV 常驻 | 全部在 GPU(229376 tokens) | 32768 tokens 常驻,超出可换出 CPU |
| 投机解码 | 未启用 | 启用(--spec 4,draft 4 tokens) |
| MTP | 不支持(AWQ 版无 MTP 头) | 支持(--mtp /data/mtp/rt,Flash-Next 版有 MTP 头) |
| 多模态 | 支持(--enable-multimodal) | 不支持(VISION=no) |
| 视觉引擎 | 无独立视觉引擎 | strata-vision(73MB,未启用) |
| MoE 支持 | 无(稠密模型) | 有(expert profile + cache) |
| MCP 支持 | 无 | 有(mcp.py) |
| Web UI | 无 | 有(内置聊天界面) |
| API 鉴权 | 无(开放) | 有(API key) |
| Metrics | Prometheus /metrics | 无标准 metrics |
| Chat 模板 | qwen3 reasoning parser | chat_template.jinja(内置) |
| Tool Call | qwen3_coder parser | 通过 MCP |
| CUDA 版本 | 12.x(sglang 镜像) | 13.0(strata 镜像) |
| 显存占用 | 40629 MiB | 39982 MiB |
核心技术差异分析
-
量化策略不同:
- SGLang 用 AWQ W4A16(4-bit 权重 + 16-bit 激活),精度较高,权重 17.75GB
- Strata 用 IQ2_XS(2-bit 超压缩),精度较低,但权重只有 ~9GB,省下的空间给 KV cache 和 MTP
-
投机解码 + MTP:
- SGLang 的 Qwen3.8-27B(AWQ 版)没有 MTP 头,无法用 MTP 投机解码
- Strata 的 Qwen3.8-Flash-Next(IQ2_XS 版)有 MTP 头,
--spec 4+--mtp组合可以实现 4-token 投机解码 - 这是 Strata 在单流 decode 速度上可能超过 SGLang 的关键
-
KV cache 策略:
- SGLang:全部 KV 在 GPU(229376 tokens × FP8),适合高并发长上下文
- Strata:32768 tokens 常驻 GPU,超出换出 CPU,适合单请求超长上下文(256K)但并发能力弱
-
功能丰富度:
- Strata 多了:MCP、Web UI、视觉引擎(未启用)、MoE expert 管理
- SGLang 多了:Prometheus metrics、多模态、成熟的调度系统
六、监控与可观测性
6.1 监控容器
sglang-monitor:latest(容器名 sglang-monitor,端口 8189→8080,Up 5 days)
推测是 Prometheus + Grafana 或 DCGM exporter,用于监控两卡的 GPU 指标和 SGLang 的 Prometheus metrics。
6.2 各服务监控端点
| 服务 | 端点 | 鉴权 |
|---|---|---|
| SGLang(GPU0) | http://127.0.0.1:30000/metrics | 无 |
| SGLang(GPU0) | http://127.0.0.1:30000/health | 无 |
| Strata(GPU1) | http://127.0.0.1:8080/health | 无 |
| Strata(GPU1) | http://127.0.0.1:8080/v1/* | 需 API key |
6.3 SGLang 核心监控指标
# 实时并发sglang:num_running_reqssglang:num_queue_reqs
# KV cache 使用率sglang:token_usagesglang:full_token_usagesglang:swa_token_usage
# 吞吐sglang:gen_throughput
# 缓存命中率sglang:cache_hit_rate
# 延迟(histogram,需 rate() 计算 P50/P95/P99)sglang:time_to_first_token_secondssglang:time_per_output_token_secondssglang:e2e_request_latency_seconds
# 累计sglang:prompt_tokens_totalsglang:generation_tokens_totalsglang:num_requests_total七、部署架构总览
┌──────────────────────────────────────────────────────────────────┐│ 127.0.0.1 (2× RTX 4090 24GB, CUDA 13.0) ││ ││ GPU0 ──→ sglang (lmsysorg/sglang:latest = v0.5.20) ││ 网络: host (直接用宿主机 30000) ││ 模型: Ar4ikov/Qwen3.8-27B-AWQ-W4A16-ASYM ││ 量化: AWQ W4A16 (17.75GB) ││ 上下文: 262144 (256K) ││ KV: FP8 e5m2, 229376 tokens ││ 多模态: ✅ 投机解码: ❌ MTP: ❌ ││ API: OpenAI + SGLang 原生, 无鉴权 ││ Metrics: Prometheus /metrics ││ 显存: 40629 MiB ││ ││ GPU1 ──→ strata (strata:local) ││ 网络: bridge (8080→8080) ││ 模型: Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS ││ 量化: IQ2_XS (~2-bit, ~9GB) ││ 上下文: 262144 (256K) ││ KV: INT8, 32768 tokens 常驻 ││ 多模态: ❌ 投机解码: ✅ (--spec 4) MTP: ✅ ││ API: OpenAI 兼容, 需 API key ││ Web UI: ✅ MCP: ✅ MoE: ✅ ││ Metrics: 无标准端点 ││ 显存: 39982 MiB ││ ││ 宿主机 ──→ sglang-monitor (8189) ││ Prometheus/Grafana 监控栈 │└──────────────────────────────────────────────────────────────────┘八、遗留问题与待确认项
| # | 问题 | 影响 |
|---|---|---|
| 1 | Strata 的 IQ2_XS 2-bit 量化精度损失多大? | 需要 benchmark 对比 AWQ W4A16 的输出质量 |
| 2 | Strata --spec 4 + MTP 的实际加速比是多少? | 需要 benchmark 对比 SGLang 单流 55 tok/s |
| 3 | Strata 的 kv-resident 32768 在 256K 上下文时的换出性能? | 长上下文 decode 速度可能受 CPU 带宽限制 |
| 4 | Strata 的 API key 管理是否安全? | key 明文写在 config 和环境变量里 |
| 5 | Strata 的 core dump(40GB core.87) | 容器内有一个 40GB 的 core dump 文件,说明引擎曾经崩溃过 |
| 6 | GPU0 SGLang 为什么用 host 网络而非 bridge? | host 模式下 30000 直接暴露在宿主机,无隔离 |
| 7 | sglang-monitor 具体监控什么? | 需要查 8189 端口返回的内容 |
| 8 | Strata 的 --ple-gguf 参数含义? | PLE 可能是 Position-Linear-Expert 或 Parallel-Linear-Expert |
| 9 | 两卡跑不同变体的 Qwen3.8(27B AWQ vs Flash-Next IQ2_XS) | 输出可能不一致,需要确认是否 intentional |
本地大模型部署系列
连载中分享文章
生成精美分享图或复制链接,与更多人分享本文。
继续阅读
换条路线
从其他文章中稳定抽取
最后更新于 ,距今已过 0 天
部分内容可能已过时