DGX Spark GB10 + Qwen3.8-27B + DFlash2 lookup — 39 tok/s chat, 117 tok/s context replay ด้วย reproducible Spark CLI bundle
สารบัญ
- TL;DR
- Hardware: NVIDIA GB10 — ไม่ใช่ DGX Spark ที่คนคิด
- Setup ขั้นสุดท้าย (OP's โปรไฟล์ 57k/C4)
- Custom patches (2 only)
- Chat Decode Benchmark
- Why k15+lookup when k9 faster in chat?
- Context Replay Benchmark (23,386-token Markdown)
- Long-context TTFT Optimization
- Concurrent 4 Long Requests
- FP8 vs NVFP4 Comparison
- fastsafetensors: ทดลองแล้วปฏิเสธ
- Reproducibility:
spark run qwen38-dflash2-lookup - Lessons ที่ได้จากโพสต์นี้
- 1. Reproducible > one-off benchmark
- 2. Caveat ที่ดีทำให้ benchmark น่าเชื่อ
- 3. Profile trade-offs ชัดเจน
- 4. GB10 sweet spot ขึ้นกับ hardware
- 5. FP8 vs NVFP4 บน GB10
- สรุป
- Cross-references
- อ้างอิง
บันทึก — 25 สิงหาคม 2569 — เช้ามืด เจอกระทู้ reproducible benchmark ที่น่าสนใจ
คนใน r/LocalLLM โพสต์ benchmark ละเอียดมาก — Qwen3.8-27B-NVFP4 บน DGX Spark / ASUS Ascent GX10 เครื่องเดียว ใช้ DFlash2 (W4A16) + lookup speculation ได้:
- ~39 tok/s ใน chat ทั่วไป (consolidated profile)
- ~42.49 tok/s เมื่อ optimize สำหรับ chat ล้วง (k9 profile)
- ~117 tok/s สำหรับ context replay (RAG, code edits, document conversion)
- Cold TTFT @ 50k ลดจาก 29.94 → 23.46s ด้วย chunk 4,096 + O2 interactivity
ที่สำคัญที่สุด: reproducible ผ่าน spark run qwen38-dflash2-lookup — ไม่ใช่แค่บอก "ผมได้ X tok/s" แต่ bundle config ทั้งหมด (model, drafter, patched vLLM image, revisions ที่ pin, runtime arguments) ใน 1 command
TL;DR
- Hardware: NVIDIA GB10 — DGX Spark / ASUS Ascent GX10 เครื่องเดียว
- Model:
sakamakismile/Qwen3.8-27B-MTP-NVFP4+ Draftersyvai/Qwen3.8-27B-DFlash2-W4A16 - Stack: vLLM + 2 custom patches (W4A16 loading + DFlash2+lookup)
- Chat decode: NVFP4 baseline ~20 tok/s → +DFlash2 W4A16 k7 = 39.41 → k9 = 42.49 → k15+lookup = ~39 (consolidated)
- Context replay: 23,386-token Markdown → 117 tok/s decode (vs 71 tok/s control)
- TTFT @ 50k: Chunk 4,096 + O2 → 23.46s (vs 29.94s with chunk 8,192) — -21.6%
- Concurrent 4 × 49k requests: 135.38s aggregate, 7.56 tok/s (+9.4%)
- FP8 vs NVFP4: NVFP4 เร็วกว่าใน prose — 38.1 vs 23.07-25.99 tok/s (NVFP4 ลด weight bandwidth ผ่าน shared memory)
- Reproducible:
spark run qwen38-dflash2-lookupผ่าน github.com/massimo92/spark
Hardware: NVIDIA GB10 — ไม่ใช่ DGX Spark ที่คนคิด
โพสต์นี้เริ่มจากการแยกความสับสนเรื่องชื่อ:
"DGX Spark หมายถึงคอมพิวเตอร์ GB10 ของ NVIDIA เสมอ Spark หรือ spark หมายถึง CLI และโปรเจกต์ GitHub ของฉันเสมอ"
GB10 = NVIDIA Grace Blackwell (Grace CPU + Blackwell GPU) SoC — เดียวกับที่อยู่ใน DGX Spark และ ASUS Ascent GX10
Setup ขั้นสุดท้าย (OP's โปรไฟล์ 57k/C4)
| Component | Value |
|---|---|
| Target | sakamakismile/Qwen3.8-27B-MTP-NVFP4 |
| Drafter | syvai/Qwen3.8-27B-DFlash2-W4A16 |
| Speculation | DFlash2 draft 7 tokens; lookup extends verification to k15 |
| Attention | FlashAttention 2 |
| KV cache | auto |
| Mamba/DeltaNet state | BF16 |
| Prefill | Chunked, 4,096-token chunks |
| Runtime | O2, interactivity, synchronous scheduling |
| Caching | Prefix caching enabled |
| Context | 57,344 tokens |
| Concurrency limit | 4 sequences |
| Sampling | Normal FlashInfer sampler |
| Split-KV | OFF |
Custom patches (2 only)
ภาพขั้นสุดท้ายมีแค่ 2 custom patches (ไม่ใช่ fork vLLM ทั้ง project):
- W4A16 loading — โหลด QKV W4A16 weights ที่ pack โดย DFlash2 drafter
- DFlash2 + lookup — รักษา 7-token drafts trained แยกจาก k15 verification block, lookup เติม positions ที่เหลือด้วย matching จาก current request
ลบ patches ที่ไม่จำเป็น: custom sampler, split-KV, speculative INT8 KV, hybrid KV grouping, recurrent-state bounds
Chat Decode Benchmark
| Setup | Chat tok/s | Note |
|---|---|---|
| NVFP4 ไม่มี speculation | ~20.0 | baseline |
| NVFP4 + DFlash2 BF16 k7 | 36.15 | +speculation ช่วยมาก |
| NVFP4 + DFlash2 W4A16 k7 | 39.41 | W4A16 ดีกว่า BF16 |
| NVFP4 + DFlash2 W4A16 k9 | 42.49 | chat profile เร็วสุด |
| NVFP4 + W4A16 k15 + lookup | ~39.0 | consolidated profile |
Why k15+lookup when k9 faster in chat?
OP อธิบาย:
"เพราะฉันต้องการเซิร์ฟเวอร์โมเดลเดียวสำหรับการสนทนา, RAG และเอเจนต์การเขียนโค้ด ไม่ต้องการเปลี่ยนโปรไฟล์ตามคำขอถัดไป
ค่าการสนทนาทั่วไปประมาณ 8% แต่แลกกับ lookup ที่สามารถสร้างการเพิ่มขึ้นที่มากขึ้นเมื่อคำตอบมีอยู่แล้วใน prompt"
Insight: trade-off ~8% chat speed เพื่อ gain ใน RAG/code-editing เป็น choice ที่ดี — เพราะ single-server use case สำคัญกว่า peak chat speed
Context Replay Benchmark (23,386-token Markdown)
OP ทำ benchmark พิเศษ — ให้ model produce หรือแก้ไขเนื้อหาจาก prompt (typical RAG, code edit, document conversion)
| Setup | Decode tok/s |
|---|---|
| DFlash2 k7 control | 71.26 |
| Minimal k15+lookup bundle | 117.08 |
| Full experimental patch set | 128.62 |
OP caveat ที่ดี:
"นี่ไม่ใช่การอ้างสิทธิ์ 100+ tok/s ทั่วไป. Lookup ช่วยเมื่อผลลัพธ์สามารถคัดลอก, อ้างอิงหรือแก้ไขบริบทที่มีอยู่ ตัวอย่างที่ดีคือ RAG answers, code edits, document conversion. มันมีประโยชน์น้อยสำหรับ unstructured prose"
Insight ที่ดีมาก: benchmark ตัวเลข 117 tok/s มี context — ถ้า use case ของคุณไม่ใช่ "modify existing context" ก็จะไม่ได้ตัวเลขนี้
Long-context TTFT Optimization
OP ทดสอบ chunk size สำหรับ cold TTFT @ 50k tokens (เวลาที่ใช้ประมวลผล prompt แรกของ request ใหม่):
| Runtime | Decode tok/s | Cold TTFT @ 50k |
|---|---|---|
| Chunk 8,192, O2 balanced | 37.96 | 29.94s |
| Chunk 4,096, O2 interactivity | 38.12 | 23.46s (-21.6%) |
OP note: "2,048 และ 16,384 แย่กว่านี้บน GB10 นี้"
Insight: sweet spot ของ chunk size ขึ้นกับ hardware — GB10 ใช้ 4,096 ดีที่สุด ลด TTFT 21.6% โดยไม่เสีย decode speed
Concurrent 4 Long Requests
OP ส่ง 4 cold requests (~49k tokens each) พร้อมกัน แต่ละ request generate 256 tokens:
| Setup | Total time | Aggregate end-to-end tok/s |
|---|---|---|
| Default ก่อนหน้า | 149.48s | 6.85 |
| 4,096 + O2 interactivity | 135.38s | 7.56 (+9.4%) |
OP caveat:
"ตัวเลข 7.56 tok/s รวมถึงการเติมข้อมูลเย็นสี่ครั้ง มันคือ aggregate end-to-end output throughput, ไม่ใช่ per-session decode speed"
Insight: "aggregate end-to-end" ≠ "per-session decode speed" — ถ้าคนอ่านตัวเลขนี้โดยไม่เข้าใจ caveat อาจสับสน
FP8 vs NVFP4 Comparison
OP ทำ comparison เพิ่มเติม — เปลี่ยน quantization จาก NVFP4 → FP8:
| Setup | Decode tok/s |
|---|---|
| FP8 (averaged across tasks) | ~32 (median) |
| NVFP4 (prose-focused) | ~38.1 |
| FP8 (prose-focused) | 23.07-25.99 |
OP สังเกต:
"ในการเปรียบเทียบที่เน้นร้อยแก้ว, FP8 ทำได้ 23.07-25.99 tok/s NVFP4 ทำได้ประมาณ 38.1 tok/s โดยมีการยอมรับ speculation ที่คล้ายกัน คำอธิบายที่น่าจะเป็นคือ แบนด์วิดธ์น้ำหนักเป้าหมาย: GB10 ต้องอ่านเป้าหมาย FP8 ขนาดใหญ่ผ่านอินเตอร์เฟซหน่วยความจำที่แชร์"
Insight ทาง hardware: NVFP4 มี weight เล็กกว่า → อ่านจาก memory น้อยกว่า → GB10 (shared memory SoC) ชนะ — FP8 อาจจะดีกว่าใน discrete GPU ที่ไม่ bandwidth-limited
fastsafetensors: ทดลองแล้วปฏิเสธ
OP ลอง --load-format fastsafetensors:
| Setup | Result |
|---|---|
| NVFP4 init time | 19.7s (เร็วขึ้น) |
| Decode tok/s | 37.91 (เท่าเดิม) |
| Available KV | 24.94 GiB (vs 60.64 GiB) → 4.10x |
| Decision | ปฏิเสธ — KV น้อยเกินไปสำหรับ C4 |
Insight ที่ดี: fast load ไม่ใช่ always-better — ต้องดู side effect (KV headroom)
Reproducibility: spark run qwen38-dflash2-lookup
# Repo
https://github.com/massimo92/spark/
# Reproduce all benchmarks
spark run qwen38-dflash2-lookup
OP อธิบาย:
"CLI จะเก็บการตั้งค่าที่สามารถทำซ้ำได้ในรูปแบบ bundles แพ็คเกจจะกำหนดโมเดลเป้าหมาย, drafter, patched vLLM image, revisions ที่ถูก pin และ runtime arguments"
Bundle รวม:
- Model + Drafter versions (pinned)
- Patched vLLM image (2 patches)
- Runtime arguments
- Reproducible test prompts
OP ขอ feedback:
"อยากเห็นผลลัพธ์จากระบบ GB10 อื่นๆ ที่ใช้ prompt และ metrics เดียวกัน — โดยเฉพาะการวัด C4 decode เฉพาะอุ่นๆ เท่านั้น"
Lessons ที่ได้จากโพสต์นี้
1. Reproducible > one-off benchmark
OP ไม่ได้แค่บอก "ผมได้ 39 tok/s" — เขา pin versions + bundle config + แจก public repo
ผลคือ: คนอื่น verify ได้ → benchmark มีค่า
2. Caveat ที่ดีทำให้ benchmark น่าเชื่อ
OP ใส่ caveats ทุกที่:
- "นี่ไม่ใช่การอ้างสิทธิ์ 100+ tok/s ทั่วไป"
- "aggregate end-to-end ไม่ใช่ per-session decode"
- "lookup มีประโยชน์น้อยสำหรับ unstructured prose"
ถ้าไม่มี caveats → ตัวเลขจะ misleading
3. Profile trade-offs ชัดเจน
OP ไม่ได้ claim "profile ของผมเร็วที่สุด" — เขาบอกชัดว่า k9 เร็วกว่า k15+lookup ใน chat แต่ k15+lookup เร็วกว่าใน RAG/code editing แล้วเลือก consolidated profile เพราะ "single model server for chat, RAG, code editing"
Trade-off คือ feature ไม่ใช่ bug
4. GB10 sweet spot ขึ้นกับ hardware
- Chunk 4,096 → optimal (GB10)
- 2,048 และ 16,384 → แย่กว่า
คนอื่นใช้ hardware อื่นอาจได้ sweet spot อื่น — ต้องวัดเอง
5. FP8 vs NVFP4 บน GB10
NVFP4 ชนะบน GB10 (shared memory bandwidth-limited) แต่อาจแพ้บน discrete GPU
สรุป
ถ้ามี GB10 (DGX Spark / ASUS Ascent GX10) แล้วอยากได้ performance สูงสุดสำหรับ Qwen3.8-27B — ใช้ DFlash2 (W4A16) + lookup speculation + chunk 4,096 prefill + O2 interactivity ได้ ~39 tok/s chat และ ~117 tok/s context replay
Benchmark นี้น่าเชื่อถือเพราะ OP pin versions ทุกอย่าง + bundle config เป็น spark run qwen38-dflash2-lookup reproducible ได้ — ตัวเลข caveats ทุกที่ (lookup ใช้ได้เฉพาะ "modify existing context" ไม่ใช่ prose ทั่วไป)
Cross-references
- $10K budget — ซื้อ 2× DGX Spark เลย หรือรอ? — market perspective
- ซื้อ DGX Spark เครื่องเดียว คุ้มไหม — awkward spot
- Tesla V100 32GB + Qwen3.8-27B — second-hand hardware alternative
อ้างอิง
- Qwen3.8-27B at 39 tok/s chat and 117 tok/s context replay — r/LocalLLM — โพสต์ต้นเรื่อง
- spark CLI — github.com/massimo92/spark — reproducible bundle
- sakamakismile/Qwen3.8-27B-MTP-NVFP4 — HuggingFace — NVFP4 target
- syvai/Qwen3.8-27B-DFlash2-W4A16 — HuggingFace — W4A16 drafter
- syv-ai/qwen38-27b-rtx3090 — Reddit — original DFlash2 work (u/iamMess)
- Inco AI — DFlash2 — DFlash2 algorithm origin
- NVIDIA DGX Spark — official — GB10 SoC spec
- ASUS Ascent GX10 — same GB10 platform
- vLLM — official — patched inference engine
- FlashInfer — GitHub — sampler used
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee