Benchmark vLLM บน DGX Spark: Qwen3.6-35B-A3B-NVFP4 ที่ 1-12 Concurrent Requests
สารบัญ
- TL;DR
- Overviews
- เครื่องมือที่ใช้
- ตรวจสอบ Model บน vLLM Server
- ตรวจสอบ sparkrun status
- Benchmark Script
- วิธีรัน
- ผล Benchmark
- วิเคราะห์ผล
- Aggregate Throughput vs Concurrency
- Per-Request Throughput vs Concurrency
- Latency Analysis
- Sweet Spot
- รายละเอียดแต่ละ Concurrency Level
- Concurrency = 1
- Concurrency = 2
- Concurrency = 4
- Concurrency = 6
- Concurrency = 12
- Conclusion
- References
บันทึก 27 มิถุนายน 2569 — อยากรู้ว่า DGX Spark รัน vLLM กับโมเดล 35B NVFP4 แล้วรับโหลดหลาย concurrent request ได้แค่ไหน ก็เลยเขียน script วัดง่าย ๆ ด้วย Python แล้วรันผ่าน SSH
ผมมี vLLM server รันอยู่บน DGX Spark ตั้งแต่เมื่อวาน แต่ไม่เคยได้วัดจริง ๆ ว่ามันรับโหลดได้แค่ไหน ส่วนใหญ่ก็แค่เรียกใช้ทีละ request ก็จบ วันนี้เลยเขียน benchmark script ง่าย ๆ ด้วย Python concurrent.futures + requests แล้วยิงที่ 1, 2, 4, 6, 12 concurrent requests เพื่อดู throughput และ latency
TL;DR
- Aggregate output throughput เพิ่มจาก 104 → 352 TPS เมื่อเพิ่ม concurrency จาก 1 → 12 (3.4x)
- Per-request TPS ลดลงจาก 104 → 41.5 TPS เพราะ GPU ต้องแชร์กันหลาย request
- ที่ 12 concurrent, latency เฉลี่ย 12.6s (เทียบกับ 4.9s ที่ 1 request) — แต่ throughput รวมยังสูงสุด
- โมเดล:
nvidia/Qwen3.6-35B-A3B-NVFP4รันบน vLLM 0.23.1rc1 บน NVIDIA DGX Spark (GB10, aarch64) - Script เก็บไว้ใช้ต่อได้ รองรับ custom prompt, concurrency levels, และ warmup
Overviews
เครื่องมือที่ใช้
| ส่วนประกอบ | รายละเอียด |
|---|---|
| Hardware | NVIDIA DGX Spark (GB10, aarch64) |
| vLLM | 0.23.1rc1.dev480+gd980a3cc6.d20260626 |
| Model | nvidia/Qwen3.6-35B-A3B-NVFP4 (NVFP4 quantization) |
| Max Context | 262,144 tokens (256K) |
| Deployment | sparkrun — recipe Qwen3.6-35B-A3B-NVFP4-MTP-v9.yaml (tp=1) |
| Server | http://10.0.0.246:8000 (OpenAI-compatible API) |
ตรวจสอบ Model บน vLLM Server
ก่อน benchmark ตรวจสอบก่อนว่าโมเดลอะไรรันอยู่:
curl -s http://10.0.0.246:8000/v1/models | python3 -m json.tool
{
"object": "list",
"data": [
{
"id": "qwen3.6-35b-nvfp4",
"object": "model",
"created": 1782559139,
"owned_by": "vllm",
"root": "nvidia/Qwen3.6-35B-A3B-NVFP4",
"parent": null,
"max_model_len": 262144,
"permission": [
{
"id": "modelperm-88be6b784ee41606",
"object": "model_permission",
"created": 1782559139,
"allow_create_engine": false,
"allow_sampling": true,
"allow_logprobs": true,
"allow_search_indices": false,
"allow_view": true,
"allow_fine_tuning": false,
"organization": "*",
"group": null,
"is_blocking": false
}
]
}
]
}
สรุปสั้น ๆ:
- Model ID:
qwen3.6-35b-nvfp4— ชื่อที่ใช้ในmodelfield ตอนเรียก API - Root Model:
nvidia/Qwen3.6-35B-A3B-NVFP4— โมเดลต้นทางจาก HuggingFace - Max Context Length: 262,144 tokens (~256K)
- vLLM version:
0.23.1rc1.dev480+gd980a3cc6.d20260626
ตรวจสอบ sparkrun status
DGX Spark ใช้ sparkrun เป็น deployment tool สำหรับรัน vLLM ใน container ตรวจสอบสถานะได้ด้วย:
sparkrun status
Job: /home/kongvut/spark-vllm-docker/recipes/Qwen3.6-35B-A3B-NVFP4-MTP-v9.yaml (tp=1) [5fb580f55fbd] (1 container(s))
solo 127.0.0.1 Up 14 hours sparkrun-eugr-vllm
logs: sparkrun logs 5fb580f55fbd
stop: sparkrun stop 5fb580f55fbd
Total: 1 container(s) across 1 host(s)
จะเห็นว่า vLLM รันมา 14 ชั่วโมงแล้ว ใน container sparkrun-eugr-vllm โดยใช้ recipe Qwen3.6-35B-A3B-NVFP4-MTP-v9.yaml ที่ tp=1 (tensor parallel = 1 — ใช้ GPU 1 ตัว)
Note: sparkrun คืออะไร? — เป็น tool สำหรับ DGX Spark ที่จัดการ container lifecycle ผ่าน recipe YAML files คล้าย docker compose แต่เฉพาะสำหรับ NVIDIA DGX Spark ecosystem ติดตั้งผ่าน
uv tool install sparkrun
Benchmark Script
เขียน script ด้วย Python ใช้ concurrent.futures.ThreadPoolExecutor สำหรับ concurrency และ requests สำหรับ HTTP calls — ไม่ต้องลง dependency เพิ่มเพราะ requests มีอยู่แล้วบนเครื่อง
#!/usr/bin/env python3
"""
vLLM Benchmark Script — วัด throughput และ latency ที่ various concurrency levels.
Usage:
python3 vllm_benchmark.py --url http://10.0.0.246:8000 --model qwen3.6-35b-nvfp4
python3 vllm_benchmark.py --concurrency 1 2 4 6 12 --prompt "Tell me about AI"
python3 vllm_benchmark.py --max-tokens 512 --warmup
Saves results to benchmark_results_<timestamp>.json
"""
import argparse
import concurrent.futures
import json
import statistics
import sys
import time
from datetime import datetime
import requests
def send_request(url, model, prompt, max_tokens, timeout=120):
"""Send a single chat completion request and return timing metrics."""
payload = {
"model": model,
"messages": [{"role": "user", "content": prompt}],
"max_tokens": max_tokens,
"temperature": 0.0,
"seed": 42,
}
start = time.perf_counter()
try:
resp = requests.post(
f"{url}/v1/chat/completions",
json=payload,
timeout=timeout,
)
elapsed = time.perf_counter() - start
if resp.status_code != 200:
return {
"success": False,
"error": f"HTTP {resp.status_code}: {resp.text[:200]}",
"elapsed": elapsed,
}
data = resp.json()
usage = data.get("usage", {})
completion_tokens = usage.get("completion_tokens", 0)
prompt_tokens = usage.get("prompt_tokens", 0)
return {
"success": True,
"elapsed": elapsed,
"prompt_tokens": prompt_tokens,
"completion_tokens": completion_tokens,
"total_tokens": usage.get("total_tokens", prompt_tokens + completion_tokens),
"tps": completion_tokens / elapsed if elapsed > 0 else 0,
}
except Exception as e:
elapsed = time.perf_counter() - start
return {"success": False, "error": str(e), "elapsed": elapsed}
def run_concurrency_test(url, model, prompt, concurrency, max_tokens, timeout):
"""Run benchmark at a specific concurrency level."""
print(f"\n Running concurrency={concurrency} ...", flush=True)
start = time.perf_counter()
with concurrent.futures.ThreadPoolExecutor(max_workers=concurrency) as pool:
futures = [
pool.submit(send_request, url, model, prompt, max_tokens, timeout)
for _ in range(concurrency)
]
results = [f.result() for f in concurrent.futures.as_completed(futures)]
wall_time = time.perf_counter() - start
successful = [r for r in results if r.get("success")]
failed = [r for r in results if not r.get("success")]
if not successful:
return {
"concurrency": concurrency,
"success": False,
"error": failed[0].get("error", "unknown") if failed else "no results",
}
latencies = [r["elapsed"] for r in successful]
tps_values = [r["tps"] for r in successful]
total_output_tokens = sum(r["completion_tokens"] for r in successful)
total_input_tokens = sum(r["prompt_tokens"] for r in successful)
return {
"concurrency": concurrency,
"success": True,
"num_requests": concurrency,
"num_successful": len(successful),
"num_failed": len(failed),
"wall_time_s": round(wall_time, 3),
"latency_avg_s": round(statistics.mean(latencies), 3),
"latency_min_s": round(min(latencies), 3),
"latency_max_s": round(max(latencies), 3),
"latency_p50_s": round(statistics.median(latencies), 3),
"total_output_tokens": total_output_tokens,
"total_input_tokens": total_input_tokens,
"aggregate_output_tps": round(total_output_tokens / wall_time, 1) if wall_time > 0 else 0,
"per_request_tps_avg": round(statistics.mean(tps_values), 1),
"per_request_tps_min": round(min(tps_values), 1),
"per_request_tps_max": round(max(tps_values), 1),
"individual_results": results,
}
Script เต็ม ๆ อยู่ใน scripts/vllm_benchmark.py
วิธีรัน
รันจากเครื่องที่ SSH เข้า DGX Spark ได้ (หรือรันบนเครื่อง DGX Spark เอง):
python3 vllm_benchmark.py \
--url http://10.0.0.246:8000 \
--model qwen3.6-35b-nvfp4 \
--concurrency 1 2 4 6 12 \
--max-tokens 512 \
--warmup \
--timeout 300
Parameters ที่ใช้:
--warmup— ส่ง request เล็ก ๆ ก่อนเริ่ม benchmark เพื่อให้ vLLM warm up KV cache--max-tokens 512— จำกัด output 512 tokens ต่อ request (เพียงพอสำหรับวัด throughput)--temperature 0.0— ปิด randomness เพื่อให้ผลคงที่ (fixed ใน script)--seed 42— reproducibility--timeout 300— รอ request นานสุด 5 นาที
ผล Benchmark
======================================================================
vLLM Concurrency Benchmark
======================================================================
Server: http://10.0.0.246:8000
Model: qwen3.6-35b-nvfp4
Prompt: Write a short essay about the future of artificial intelligence in healthcare.
Max out: 512 tokens
Levels: [1, 2, 4, 6, 12]
Available models: ['qwen3.6-35b-nvfp4']
Warming up ...
Warmup done: 0.34s, 16 tokens
SUMMARY
======================================================================
Conc | Wall(s) | AvgLat | P50Lat | MaxLat | OutTPS | PerReqTPS | OK/Total
----------------------------------------------------------------------
1 | 4.92 | 4.92 | 4.92 | 4.92 | 104.1 | 104.2 | 1/1
2 | 6.31 | 6.17 | 6.17 | 6.31 | 162.3 | 83.0 | 2/2
4 | 7.74 | 7.70 | 7.68 | 7.74 | 264.5 | 66.5 | 4/4
6 | 9.74 | 9.52 | 9.46 | 9.74 | 315.4 | 53.8 | 6/6
12 | 17.45 | 12.62 | 11.77 | 17.44 | 352.1 | 41.5 | 12/12
Results saved to: benchmark_results_20260627_182207.json
วิเคราะห์ผล
Aggregate Throughput vs Concurrency
| Concurrency | Wall Time (s) | Total Output Tokens | Aggregate TPS | Scaling vs 1-req |
|---|---|---|---|---|
| 1 | 4.92 | 512 | 104.1 | 1.0x |
| 2 | 6.31 | 1,024 | 162.3 | 1.6x |
| 4 | 7.74 | 2,048 | 264.5 | 2.5x |
| 6 | 9.74 | 3,072 | 315.4 | 3.0x |
| 12 | 17.45 | 6,144 | 352.1 | 3.4x |
Aggregate throughput เพิ่มขึ้นตาม concurrency แต่ไม่ได้ scale เป็นเส้นตรง — จาก 1 → 12 concurrent (12x โหลด) throughput เพิ่มแค่ 3.4x แสดงว่า GPU เริ่มเต็ม capacity ประมาณ 6-12 concurrent
Per-Request Throughput vs Concurrency
| Concurrency | Per-Req TPS (avg) | Per-Req TPS (min) | Per-Req TPS (max) |
|---|---|---|---|
| 1 | 104.2 | 104.2 | 104.2 |
| 2 | 83.0 | 81.2 | 84.9 |
| 4 | 66.5 | 66.1 | 66.6 |
| 6 | 53.8 | 52.6 | 54.4 |
| 12 | 41.5 | 29.4 | 46.0 |
Per-request TPS ลดลงเรื่อย ๆ เมื่อ concurrency เพิ่ม — จาก 104 TPS (1 request) เหลือ 41.5 TPS (12 concurrent) แสดงว่าแต่ละ request ได้ GPU time น้อยลงเมื่อต้องแชร์กันหลายตัว
ที่ 12 concurrent สังเกตได้ว่า min ตกลงไปที่ 29.4 TPS ในขณะที่ max ยังอยู่ที่ 46.0 TPS — แสดงว่ามี request บางตัวที่รอนานกว่าคนอื่น (tail latency) ซึ่งเป็นพฤติกรรมปกติของ continuous batching เมื่อ batch เต็ม
Latency Analysis
| Concurrency | Avg Latency (s) | P50 Latency (s) | Max Latency (s) |
|---|---|---|---|
| 1 | 4.92 | 4.92 | 4.92 |
| 2 | 6.17 | 6.17 | 6.31 |
| 4 | 7.70 | 7.68 | 7.74 |
| 6 | 9.52 | 9.46 | 9.74 |
| 12 | 12.62 | 11.77 | 17.44 |
- ที่ 1-4 concurrent, latency กระจายน้อย (P50 ≈ Avg ≈ Max) — vLLM จัดการได้ดี
- ที่ 12 concurrent, Max latency กระโดดไป 17.4s เทียบกับ P50 ที่ 11.8s — มี request ที่รอนานกว่าคนอื่น ~50% เพราะ batch เต็มและต้องรอรอบถัดไป
Sweet Spot
ดูจากข้อมูลแล้ว 4-6 concurrent เป็น sweet spot สำหรับโมเดลนี้บน DGX Spark:
- ที่ 4 concurrent: throughput 264.5 TPS, latency 7.7s — ยังคง latency ต่ำ
- ที่ 6 concurrent: throughput 315.4 TPS, latency 9.5s — throughput เพิ่ม 19% แต่ latency เพิ่ม 23%
- ที่ 12 concurrent: throughput 352.1 TPS, latency 12.6s — throughput เพิ่มแค่ 12% จาก 6 แต่ latency เพิ่ม 32%
ถ้าแอปพลิเคชันสนใจ latency (เช่น chatbot) — 4 concurrent เหมาะที่สุด ถ้าสนใจ throughput (เช่น batch processing) — 6-12 concurrent ให้ผลรวมสูงกว่า
รายละเอียดแต่ละ Concurrency Level
Concurrency = 1
ที่ single request, vLLM ใช้เวลา 4.92s สร้าง 512 tokens — คิดเป็น 104.2 TPS โดยไม่มี contention เลย นี่คือ baseline throughput ของโมเดลบน GPU ตัวนี้
Concurrency = 2
Wall time เพิ่มจาก 4.92s → 6.31s (28%) แต่ throughput รวมเพิ่มจาก 104 → 162 TPS (56%) — vLLM ใช้ continuous batching ทำให้รับ 2 request พร้อมกันได้โดยใช้เวลาเพิ่มไม่มาก
Concurrency = 4
Wall time 7.74s, throughput 264.5 TPS — เพิ่ม 2.5x จาก baseline นี่คือจุดที่ GPU เริ่มทำงานหนักขึ้นแต่ยังจัดการได้ดี latency กระจายน้อยมาก (7.68s - 7.74s)
Concurrency = 6
Throughput 315.4 TPS — เพิ่ม 19% จาก 4 concurrent แต่ latency เริ่มเพิ่มชัดเจน (7.7s → 9.5s) GPU เริ่มเข้าใกล้ capacity
Concurrency = 12
Throughput 352.1 TPS — สูงสุด แต่ per-request TPS ตกลงไปเหลือ 41.5 (เฉลี่ย) บาง request ได้แค่ 29.4 TPS ส่วนที่เร็วได้ 46.0 TPS ความต่างระหว่าง min/max latency กว้างขึ้นมาก (11.1s - 17.4s) แสดงว่า batch เต็มและมี request ต้องรอคิว
Conclusion
การ benchmark vLLM บน DGX Spark กับ Qwen3.6-35B-A3B-NVFP4 แสดงให้เห็นว่า:
- Single-request throughput อยู่ที่ ~104 TPS — ใช้งานได้สบายสำหรับ interactive use
- Aggregate throughput เพิ่มได้ถึง 352 TPS ที่ 12 concurrent — เหมาะสำหรับ batch processing
- Sweet spot อยู่ที่ 4-6 concurrent สำหรับงานที่สมดุลทั้ง latency และ throughput
- vLLM continuous batching ทำงานได้ดี — throughput เพิ่ม 3.4x เมื่อโหลดเพิ่ม 12x
Script ที่เขียนเก็บไว้ใช้ต่อได้ ปรับ concurrency levels, prompt, max_tokens ได้ตามต้องการ และผลลัพธ์เก็บเป็น JSON สำหรับนำไปวิเคราะห์ต่อ
References
- vLLM Documentation — Continuous batching และ performance tuning
- nvidia/Qwen3.6-35B-A3B-NVFP4 (Hugging Face) — Model card และ NVFP4 quantization
- NVIDIA DGX Spark — Hardware specifications
- Benchmark Script (GitHub) — Script ที่ใช้ในบทความนี้
- Benchmark Results JSON (GitHub) — ผลลัพธ์ดิบจากการทดสอบ
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee