Skip to main content

DGX Spark GB10 + Qwen3.8-27B + DFlash2 lookup — 39 tok/s chat, 117 tok/s context replay ด้วย reproducible Spark CLI bundle

· 8 min read

บันทึก — 25 สิงหาคม 2569 — เช้ามืด เจอกระทู้ reproducible benchmark ที่น่าสนใจ

คนใน r/LocalLLM โพสต์ benchmark ละเอียดมาก — Qwen3.8-27B-NVFP4 บน DGX Spark / ASUS Ascent GX10 เครื่องเดียว ใช้ DFlash2 (W4A16) + lookup speculation ได้:

  • ~39 tok/s ใน chat ทั่วไป (consolidated profile)
  • ~42.49 tok/s เมื่อ optimize สำหรับ chat ล้วง (k9 profile)
  • ~117 tok/s สำหรับ context replay (RAG, code edits, document conversion)
  • Cold TTFT @ 50k ลดจาก 29.94 → 23.46s ด้วย chunk 4,096 + O2 interactivity

ที่สำคัญที่สุด: reproducible ผ่าน spark run qwen38-dflash2-lookup — ไม่ใช่แค่บอก "ผมได้ X tok/s" แต่ bundle config ทั้งหมด (model, drafter, patched vLLM image, revisions ที่ pin, runtime arguments) ใน 1 command

TL;DR​

  • Hardware: NVIDIA GB10 — DGX Spark / ASUS Ascent GX10 เครื่องเดียว
  • Model: sakamakismile/Qwen3.8-27B-MTP-NVFP4 + Drafter syvai/Qwen3.8-27B-DFlash2-W4A16
  • Stack: vLLM + 2 custom patches (W4A16 loading + DFlash2+lookup)
  • Chat decode: NVFP4 baseline ~20 tok/s → +DFlash2 W4A16 k7 = 39.41 → k9 = 42.49 → k15+lookup = ~39 (consolidated)
  • Context replay: 23,386-token Markdown → 117 tok/s decode (vs 71 tok/s control)
  • TTFT @ 50k: Chunk 4,096 + O2 → 23.46s (vs 29.94s with chunk 8,192) — -21.6%
  • Concurrent 4 × 49k requests: 135.38s aggregate, 7.56 tok/s (+9.4%)
  • FP8 vs NVFP4: NVFP4 เร็วกว่าใน prose — 38.1 vs 23.07-25.99 tok/s (NVFP4 ลด weight bandwidth ผ่าน shared memory)
  • Reproducible: spark run qwen38-dflash2-lookup ผ่าน github.com/massimo92/spark

Hardware: NVIDIA GB10 — ไม่ใช่ DGX Spark ที่คนคิด​

โพสต์นี้เริ่มจากการแยกความสับสนเรื่องชื่อ:

"DGX Spark หมายถึงคอมพิวเตอร์ GB10 ของ NVIDIA เสมอ Spark หรือ spark หมายถึง CLI และโปรเจกต์ GitHub ของฉันเสมอ"

GB10 = NVIDIA Grace Blackwell (Grace CPU + Blackwell GPU) SoC — เดียวกับที่อยู่ใน DGX Spark และ ASUS Ascent GX10

Setup ขั้นสุดท้าย (OP's โปรไฟล์ 57k/C4)​

ComponentValue
Targetsakamakismile/Qwen3.8-27B-MTP-NVFP4
Draftersyvai/Qwen3.8-27B-DFlash2-W4A16
SpeculationDFlash2 draft 7 tokens; lookup extends verification to k15
AttentionFlashAttention 2
KV cacheauto
Mamba/DeltaNet stateBF16
PrefillChunked, 4,096-token chunks
RuntimeO2, interactivity, synchronous scheduling
CachingPrefix caching enabled
Context57,344 tokens
Concurrency limit4 sequences
SamplingNormal FlashInfer sampler
Split-KVOFF

Custom patches (2 only)​

ภาพขั้นสุดท้ายมีแค่ 2 custom patches (ไม่ใช่ fork vLLM ทั้ง project):

  1. W4A16 loading — โหลด QKV W4A16 weights ที่ pack โดย DFlash2 drafter
  2. DFlash2 + lookup — รักษา 7-token drafts trained แยกจาก k15 verification block, lookup เติม positions ที่เหลือด้วย matching จาก current request

ลบ patches ที่ไม่จำเป็น: custom sampler, split-KV, speculative INT8 KV, hybrid KV grouping, recurrent-state bounds

Chat Decode Benchmark​

SetupChat tok/sNote
NVFP4 ไม่มี speculation~20.0baseline
NVFP4 + DFlash2 BF16 k736.15+speculation ช่วยมาก
NVFP4 + DFlash2 W4A16 k739.41W4A16 ดีกว่า BF16
NVFP4 + DFlash2 W4A16 k942.49chat profile เร็วสุด
NVFP4 + W4A16 k15 + lookup~39.0consolidated profile

Why k15+lookup when k9 faster in chat?​

OP อธิบาย:

"เพราะฉันต้องการเซิร์ฟเวอร์โมเดลเดียวสำหรับการสนทนา, RAG และเอเจนต์การเขียนโค้ด ไม่ต้องการเปลี่ยนโปรไฟล์ตามคำขอถัดไป

ค่าการสนทนาทั่วไปประมาณ 8% แต่แลกกับ lookup ที่สามารถสร้างการเพิ่มขึ้นที่มากขึ้นเมื่อคำตอบมีอยู่แล้วใน prompt"

Insight: trade-off ~8% chat speed เพื่อ gain ใน RAG/code-editing เป็น choice ที่ดี — เพราะ single-server use case สำคัญกว่า peak chat speed

Context Replay Benchmark (23,386-token Markdown)​

OP ทำ benchmark พิเศษ — ให้ model produce หรือแก้ไขเนื้อหาจาก prompt (typical RAG, code edit, document conversion)

SetupDecode tok/s
DFlash2 k7 control71.26
Minimal k15+lookup bundle117.08
Full experimental patch set128.62

OP caveat ที่ดี:

"นี่ไม่ใช่การอ้างสิทธิ์ 100+ tok/s ทั่วไป. Lookup ช่วยเมื่อผลลัพธ์สามารถคัดลอก, อ้างอิงหรือแก้ไขบริบทที่มีอยู่ ตัวอย่างที่ดีคือ RAG answers, code edits, document conversion. มันมีประโยชน์น้อยสำหรับ unstructured prose"

Insight ที่ดีมาก: benchmark ตัวเลข 117 tok/s มี context — ถ้า use case ของคุณไม่ใช่ "modify existing context" ก็จะไม่ได้ตัวเลขนี้

Long-context TTFT Optimization​

OP ทดสอบ chunk size สำหรับ cold TTFT @ 50k tokens (เวลาที่ใช้ประมวลผล prompt แรกของ request ใหม่):

RuntimeDecode tok/sCold TTFT @ 50k
Chunk 8,192, O2 balanced37.9629.94s
Chunk 4,096, O2 interactivity38.1223.46s (-21.6%)

OP note: "2,048 และ 16,384 แย่กว่านี้บน GB10 นี้"

Insight: sweet spot ของ chunk size ขึ้นกับ hardware — GB10 ใช้ 4,096 ดีที่สุด ลด TTFT 21.6% โดยไม่เสีย decode speed

Concurrent 4 Long Requests​

OP ส่ง 4 cold requests (~49k tokens each) พร้อมกัน แต่ละ request generate 256 tokens:

SetupTotal timeAggregate end-to-end tok/s
Default ก่อนหน้า149.48s6.85
4,096 + O2 interactivity135.38s7.56 (+9.4%)

OP caveat:

"ตัวเลข 7.56 tok/s รวมถึงการเติมข้อมูลเย็นสี่ครั้ง มันคือ aggregate end-to-end output throughput, ไม่ใช่ per-session decode speed"

Insight: "aggregate end-to-end" ≠ "per-session decode speed" — ถ้าคนอ่านตัวเลขนี้โดยไม่เข้าใจ caveat อาจสับสน

FP8 vs NVFP4 Comparison​

OP ทำ comparison เพิ่มเติม — เปลี่ยน quantization จาก NVFP4 → FP8:

SetupDecode tok/s
FP8 (averaged across tasks)~32 (median)
NVFP4 (prose-focused)~38.1
FP8 (prose-focused)23.07-25.99

OP สังเกต:

"ในการเปรียบเทียบที่เน้นร้อยแก้ว, FP8 ทำได้ 23.07-25.99 tok/s NVFP4 ทำได้ประมาณ 38.1 tok/s โดยมีการยอมรับ speculation ที่คล้ายกัน คำอธิบายที่น่าจะเป็นคือ แบนด์วิดธ์น้ำหนักเป้าหมาย: GB10 ต้องอ่านเป้าหมาย FP8 ขนาดใหญ่ผ่านอินเตอร์เฟซหน่วยความจำที่แชร์"

Insight ทาง hardware: NVFP4 มี weight เล็กกว่า → อ่านจาก memory น้อยกว่า → GB10 (shared memory SoC) ชนะ — FP8 อาจจะดีกว่าใน discrete GPU ที่ไม่ bandwidth-limited

fastsafetensors: ทดลองแล้วปฏิเสธ​

OP ลอง --load-format fastsafetensors:

SetupResult
NVFP4 init time19.7s (เร็วขึ้น)
Decode tok/s37.91 (เท่าเดิม)
Available KV24.94 GiB (vs 60.64 GiB) → 4.10x
Decisionปฏิเสธ — KV น้อยเกินไปสำหรับ C4

Insight ที่ดี: fast load ไม่ใช่ always-better — ต้องดู side effect (KV headroom)

Reproducibility: spark run qwen38-dflash2-lookup​

# Repo
https://github.com/massimo92/spark/

# Reproduce all benchmarks
spark run qwen38-dflash2-lookup

OP อธิบาย:

"CLI จะเก็บการตั้งค่าที่สามารถทำซ้ำได้ในรูปแบบ bundles แพ็คเกจจะกำหนดโมเดลเป้าหมาย, drafter, patched vLLM image, revisions ที่ถูก pin และ runtime arguments"

Bundle รวม:

  • Model + Drafter versions (pinned)
  • Patched vLLM image (2 patches)
  • Runtime arguments
  • Reproducible test prompts

OP ขอ feedback:

"อยากเห็นผลลัพธ์จากระบบ GB10 อื่นๆ ที่ใช้ prompt และ metrics เดียวกัน — โดยเฉพาะการวัด C4 decode เฉพาะอุ่นๆ เท่านั้น"

Lessons ที่ได้จากโพสต์นี้​

1. Reproducible > one-off benchmark​

OP ไม่ได้แค่บอก "ผมได้ 39 tok/s" — เขา pin versions + bundle config + แจก public repo

ผลคือ: คนอื่น verify ได้ → benchmark มีค่า

2. Caveat ที่ดีทำให้ benchmark น่าเชื่อ​

OP ใส่ caveats ทุกที่:

  • "นี่ไม่ใช่การอ้างสิทธิ์ 100+ tok/s ทั่วไป"
  • "aggregate end-to-end ไม่ใช่ per-session decode"
  • "lookup มีประโยชน์น้อยสำหรับ unstructured prose"

ถ้าไม่มี caveats → ตัวเลขจะ misleading

3. Profile trade-offs ชัดเจน​

OP ไม่ได้ claim "profile ของผมเร็วที่สุด" — เขาบอกชัดว่า k9 เร็วกว่า k15+lookup ใน chat แต่ k15+lookup เร็วกว่าใน RAG/code editing แล้วเลือก consolidated profile เพราะ "single model server for chat, RAG, code editing"

Trade-off คือ feature ไม่ใช่ bug

4. GB10 sweet spot ขึ้นกับ hardware​

  • Chunk 4,096 → optimal (GB10)
  • 2,048 และ 16,384 → แย่กว่า

คนอื่นใช้ hardware อื่นอาจได้ sweet spot อื่น — ต้องวัดเอง

5. FP8 vs NVFP4 บน GB10​

NVFP4 ชนะบน GB10 (shared memory bandwidth-limited) แต่อาจแพ้บน discrete GPU

สรุป​

ถ้ามี GB10 (DGX Spark / ASUS Ascent GX10) แล้วอยากได้ performance สูงสุดสำหรับ Qwen3.8-27B — ใช้ DFlash2 (W4A16) + lookup speculation + chunk 4,096 prefill + O2 interactivity ได้ ~39 tok/s chat และ ~117 tok/s context replay

Benchmark นี้น่าเชื่อถือเพราะ OP pin versions ทุกอย่าง + bundle config เป็น spark run qwen38-dflash2-lookup reproducible ได้ — ตัวเลข caveats ทุกที่ (lookup ใช้ได้เฉพาะ "modify existing context" ไม่ใช่ prose ทั่วไป)

Cross-references​

อ้างอิง​

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...