Skip to main content

Benchmark vLLM บน DGX Spark: Qwen3.6-35B-A3B-NVFP4 ที่ 1-12 Concurrent Requests

· 10 min read

บันทึก 27 มิถุนายน 2569 — อยากรู้ว่า DGX Spark รัน vLLM กับโมเดล 35B NVFP4 แล้วรับโหลดหลาย concurrent request ได้แค่ไหน ก็เลยเขียน script วัดง่าย ๆ ด้วย Python แล้วรันผ่าน SSH

ผมมี vLLM server รันอยู่บน DGX Spark ตั้งแต่เมื่อวาน แต่ไม่เคยได้วัดจริง ๆ ว่ามันรับโหลดได้แค่ไหน ส่วนใหญ่ก็แค่เรียกใช้ทีละ request ก็จบ วันนี้เลยเขียน benchmark script ง่าย ๆ ด้วย Python concurrent.futures + requests แล้วยิงที่ 1, 2, 4, 6, 12 concurrent requests เพื่อดู throughput และ latency

TL;DR​

  • Aggregate output throughput เพิ่มจาก 104 → 352 TPS เมื่อเพิ่ม concurrency จาก 1 → 12 (3.4x)
  • Per-request TPS ลดลงจาก 104 → 41.5 TPS เพราะ GPU ต้องแชร์กันหลาย request
  • ที่ 12 concurrent, latency เฉลี่ย 12.6s (เทียบกับ 4.9s ที่ 1 request) — แต่ throughput รวมยังสูงสุด
  • โมเดล: nvidia/Qwen3.6-35B-A3B-NVFP4 รันบน vLLM 0.23.1rc1 บน NVIDIA DGX Spark (GB10, aarch64)
  • Script เก็บไว้ใช้ต่อได้ รองรับ custom prompt, concurrency levels, และ warmup

Overviews​

เครื่องมือที่ใช้​

ส่วนประกอบรายละเอียด
HardwareNVIDIA DGX Spark (GB10, aarch64)
vLLM0.23.1rc1.dev480+gd980a3cc6.d20260626
Modelnvidia/Qwen3.6-35B-A3B-NVFP4 (NVFP4 quantization)
Max Context262,144 tokens (256K)
Deploymentsparkrun — recipe Qwen3.6-35B-A3B-NVFP4-MTP-v9.yaml (tp=1)
Serverhttp://10.0.0.246:8000 (OpenAI-compatible API)

ตรวจสอบ Model บน vLLM Server​

ก่อน benchmark ตรวจสอบก่อนว่าโมเดลอะไรรันอยู่:

curl -s http://10.0.0.246:8000/v1/models | python3 -m json.tool
{
"object": "list",
"data": [
{
"id": "qwen3.6-35b-nvfp4",
"object": "model",
"created": 1782559139,
"owned_by": "vllm",
"root": "nvidia/Qwen3.6-35B-A3B-NVFP4",
"parent": null,
"max_model_len": 262144,
"permission": [
{
"id": "modelperm-88be6b784ee41606",
"object": "model_permission",
"created": 1782559139,
"allow_create_engine": false,
"allow_sampling": true,
"allow_logprobs": true,
"allow_search_indices": false,
"allow_view": true,
"allow_fine_tuning": false,
"organization": "*",
"group": null,
"is_blocking": false
}
]
}
]
}

สรุปสั้น ๆ:

  • Model ID: qwen3.6-35b-nvfp4 — ชื่อที่ใช้ใน model field ตอนเรียก API
  • Root Model: nvidia/Qwen3.6-35B-A3B-NVFP4 — โมเดลต้นทางจาก HuggingFace
  • Max Context Length: 262,144 tokens (~256K)
  • vLLM version: 0.23.1rc1.dev480+gd980a3cc6.d20260626

ตรวจสอบ sparkrun status​

DGX Spark ใช้ sparkrun เป็น deployment tool สำหรับรัน vLLM ใน container ตรวจสอบสถานะได้ด้วย:

sparkrun status
Job: /home/kongvut/spark-vllm-docker/recipes/Qwen3.6-35B-A3B-NVFP4-MTP-v9.yaml (tp=1) [5fb580f55fbd] (1 container(s))
solo 127.0.0.1 Up 14 hours sparkrun-eugr-vllm
logs: sparkrun logs 5fb580f55fbd
stop: sparkrun stop 5fb580f55fbd

Total: 1 container(s) across 1 host(s)

จะเห็นว่า vLLM รันมา 14 ชั่วโมงแล้ว ใน container sparkrun-eugr-vllm โดยใช้ recipe Qwen3.6-35B-A3B-NVFP4-MTP-v9.yaml ที่ tp=1 (tensor parallel = 1 — ใช้ GPU 1 ตัว)

Note: sparkrun คืออะไร? — เป็น tool สำหรับ DGX Spark ที่จัดการ container lifecycle ผ่าน recipe YAML files คล้าย docker compose แต่เฉพาะสำหรับ NVIDIA DGX Spark ecosystem ติดตั้งผ่าน uv tool install sparkrun

Benchmark Script​

เขียน script ด้วย Python ใช้ concurrent.futures.ThreadPoolExecutor สำหรับ concurrency และ requests สำหรับ HTTP calls — ไม่ต้องลง dependency เพิ่มเพราะ requests มีอยู่แล้วบนเครื่อง

#!/usr/bin/env python3
"""
vLLM Benchmark Script — วัด throughput และ latency ที่ various concurrency levels.

Usage:
python3 vllm_benchmark.py --url http://10.0.0.246:8000 --model qwen3.6-35b-nvfp4
python3 vllm_benchmark.py --concurrency 1 2 4 6 12 --prompt "Tell me about AI"
python3 vllm_benchmark.py --max-tokens 512 --warmup

Saves results to benchmark_results_<timestamp>.json
"""

import argparse
import concurrent.futures
import json
import statistics
import sys
import time
from datetime import datetime

import requests


def send_request(url, model, prompt, max_tokens, timeout=120):
"""Send a single chat completion request and return timing metrics."""
payload = {
"model": model,
"messages": [{"role": "user", "content": prompt}],
"max_tokens": max_tokens,
"temperature": 0.0,
"seed": 42,
}
start = time.perf_counter()
try:
resp = requests.post(
f"{url}/v1/chat/completions",
json=payload,
timeout=timeout,
)
elapsed = time.perf_counter() - start
if resp.status_code != 200:
return {
"success": False,
"error": f"HTTP {resp.status_code}: {resp.text[:200]}",
"elapsed": elapsed,
}
data = resp.json()
usage = data.get("usage", {})
completion_tokens = usage.get("completion_tokens", 0)
prompt_tokens = usage.get("prompt_tokens", 0)
return {
"success": True,
"elapsed": elapsed,
"prompt_tokens": prompt_tokens,
"completion_tokens": completion_tokens,
"total_tokens": usage.get("total_tokens", prompt_tokens + completion_tokens),
"tps": completion_tokens / elapsed if elapsed > 0 else 0,
}
except Exception as e:
elapsed = time.perf_counter() - start
return {"success": False, "error": str(e), "elapsed": elapsed}


def run_concurrency_test(url, model, prompt, concurrency, max_tokens, timeout):
"""Run benchmark at a specific concurrency level."""
print(f"\n Running concurrency={concurrency} ...", flush=True)
start = time.perf_counter()
with concurrent.futures.ThreadPoolExecutor(max_workers=concurrency) as pool:
futures = [
pool.submit(send_request, url, model, prompt, max_tokens, timeout)
for _ in range(concurrency)
]
results = [f.result() for f in concurrent.futures.as_completed(futures)]
wall_time = time.perf_counter() - start

successful = [r for r in results if r.get("success")]
failed = [r for r in results if not r.get("success")]

if not successful:
return {
"concurrency": concurrency,
"success": False,
"error": failed[0].get("error", "unknown") if failed else "no results",
}

latencies = [r["elapsed"] for r in successful]
tps_values = [r["tps"] for r in successful]
total_output_tokens = sum(r["completion_tokens"] for r in successful)
total_input_tokens = sum(r["prompt_tokens"] for r in successful)

return {
"concurrency": concurrency,
"success": True,
"num_requests": concurrency,
"num_successful": len(successful),
"num_failed": len(failed),
"wall_time_s": round(wall_time, 3),
"latency_avg_s": round(statistics.mean(latencies), 3),
"latency_min_s": round(min(latencies), 3),
"latency_max_s": round(max(latencies), 3),
"latency_p50_s": round(statistics.median(latencies), 3),
"total_output_tokens": total_output_tokens,
"total_input_tokens": total_input_tokens,
"aggregate_output_tps": round(total_output_tokens / wall_time, 1) if wall_time > 0 else 0,
"per_request_tps_avg": round(statistics.mean(tps_values), 1),
"per_request_tps_min": round(min(tps_values), 1),
"per_request_tps_max": round(max(tps_values), 1),
"individual_results": results,
}

Script เต็ม ๆ อยู่ใน scripts/vllm_benchmark.py

วิธีรัน​

รันจากเครื่องที่ SSH เข้า DGX Spark ได้ (หรือรันบนเครื่อง DGX Spark เอง):

python3 vllm_benchmark.py \
--url http://10.0.0.246:8000 \
--model qwen3.6-35b-nvfp4 \
--concurrency 1 2 4 6 12 \
--max-tokens 512 \
--warmup \
--timeout 300

Parameters ที่ใช้:

  • --warmup — ส่ง request เล็ก ๆ ก่อนเริ่ม benchmark เพื่อให้ vLLM warm up KV cache
  • --max-tokens 512 — จำกัด output 512 tokens ต่อ request (เพียงพอสำหรับวัด throughput)
  • --temperature 0.0 — ปิด randomness เพื่อให้ผลคงที่ (fixed ใน script)
  • --seed 42 — reproducibility
  • --timeout 300 — รอ request นานสุด 5 นาที

ผล Benchmark​

======================================================================
vLLM Concurrency Benchmark
======================================================================
Server: http://10.0.0.246:8000
Model: qwen3.6-35b-nvfp4
Prompt: Write a short essay about the future of artificial intelligence in healthcare.
Max out: 512 tokens
Levels: [1, 2, 4, 6, 12]
Available models: ['qwen3.6-35b-nvfp4']

Warming up ...
Warmup done: 0.34s, 16 tokens

SUMMARY
======================================================================
Conc | Wall(s) | AvgLat | P50Lat | MaxLat | OutTPS | PerReqTPS | OK/Total
----------------------------------------------------------------------
1 | 4.92 | 4.92 | 4.92 | 4.92 | 104.1 | 104.2 | 1/1
2 | 6.31 | 6.17 | 6.17 | 6.31 | 162.3 | 83.0 | 2/2
4 | 7.74 | 7.70 | 7.68 | 7.74 | 264.5 | 66.5 | 4/4
6 | 9.74 | 9.52 | 9.46 | 9.74 | 315.4 | 53.8 | 6/6
12 | 17.45 | 12.62 | 11.77 | 17.44 | 352.1 | 41.5 | 12/12

Results saved to: benchmark_results_20260627_182207.json

วิเคราะห์ผล​

Aggregate Throughput vs Concurrency​

ConcurrencyWall Time (s)Total Output TokensAggregate TPSScaling vs 1-req
14.92512104.11.0x
26.311,024162.31.6x
47.742,048264.52.5x
69.743,072315.43.0x
1217.456,144352.13.4x

Aggregate throughput เพิ่มขึ้นตาม concurrency แต่ไม่ได้ scale เป็นเส้นตรง — จาก 1 → 12 concurrent (12x โหลด) throughput เพิ่มแค่ 3.4x แสดงว่า GPU เริ่มเต็ม capacity ประมาณ 6-12 concurrent

Per-Request Throughput vs Concurrency​

ConcurrencyPer-Req TPS (avg)Per-Req TPS (min)Per-Req TPS (max)
1104.2104.2104.2
283.081.284.9
466.566.166.6
653.852.654.4
1241.529.446.0

Per-request TPS ลดลงเรื่อย ๆ เมื่อ concurrency เพิ่ม — จาก 104 TPS (1 request) เหลือ 41.5 TPS (12 concurrent) แสดงว่าแต่ละ request ได้ GPU time น้อยลงเมื่อต้องแชร์กันหลายตัว

ที่ 12 concurrent สังเกตได้ว่า min ตกลงไปที่ 29.4 TPS ในขณะที่ max ยังอยู่ที่ 46.0 TPS — แสดงว่ามี request บางตัวที่รอนานกว่าคนอื่น (tail latency) ซึ่งเป็นพฤติกรรมปกติของ continuous batching เมื่อ batch เต็ม

Latency Analysis​

ConcurrencyAvg Latency (s)P50 Latency (s)Max Latency (s)
14.924.924.92
26.176.176.31
47.707.687.74
69.529.469.74
1212.6211.7717.44
  • ที่ 1-4 concurrent, latency กระจายน้อย (P50 ≈ Avg ≈ Max) — vLLM จัดการได้ดี
  • ที่ 12 concurrent, Max latency กระโดดไป 17.4s เทียบกับ P50 ที่ 11.8s — มี request ที่รอนานกว่าคนอื่น ~50% เพราะ batch เต็มและต้องรอรอบถัดไป

Sweet Spot​

ดูจากข้อมูลแล้ว 4-6 concurrent เป็น sweet spot สำหรับโมเดลนี้บน DGX Spark:

  • ที่ 4 concurrent: throughput 264.5 TPS, latency 7.7s — ยังคง latency ต่ำ
  • ที่ 6 concurrent: throughput 315.4 TPS, latency 9.5s — throughput เพิ่ม 19% แต่ latency เพิ่ม 23%
  • ที่ 12 concurrent: throughput 352.1 TPS, latency 12.6s — throughput เพิ่มแค่ 12% จาก 6 แต่ latency เพิ่ม 32%

ถ้าแอปพลิเคชันสนใจ latency (เช่น chatbot) — 4 concurrent เหมาะที่สุด ถ้าสนใจ throughput (เช่น batch processing) — 6-12 concurrent ให้ผลรวมสูงกว่า

รายละเอียดแต่ละ Concurrency Level​

Concurrency = 1​

ที่ single request, vLLM ใช้เวลา 4.92s สร้าง 512 tokens — คิดเป็น 104.2 TPS โดยไม่มี contention เลย นี่คือ baseline throughput ของโมเดลบน GPU ตัวนี้

Concurrency = 2​

Wall time เพิ่มจาก 4.92s → 6.31s (28%) แต่ throughput รวมเพิ่มจาก 104 → 162 TPS (56%) — vLLM ใช้ continuous batching ทำให้รับ 2 request พร้อมกันได้โดยใช้เวลาเพิ่มไม่มาก

Concurrency = 4​

Wall time 7.74s, throughput 264.5 TPS — เพิ่ม 2.5x จาก baseline นี่คือจุดที่ GPU เริ่มทำงานหนักขึ้นแต่ยังจัดการได้ดี latency กระจายน้อยมาก (7.68s - 7.74s)

Concurrency = 6​

Throughput 315.4 TPS — เพิ่ม 19% จาก 4 concurrent แต่ latency เริ่มเพิ่มชัดเจน (7.7s → 9.5s) GPU เริ่มเข้าใกล้ capacity

Concurrency = 12​

Throughput 352.1 TPS — สูงสุด แต่ per-request TPS ตกลงไปเหลือ 41.5 (เฉลี่ย) บาง request ได้แค่ 29.4 TPS ส่วนที่เร็วได้ 46.0 TPS ความต่างระหว่าง min/max latency กว้างขึ้นมาก (11.1s - 17.4s) แสดงว่า batch เต็มและมี request ต้องรอคิว

Conclusion​

การ benchmark vLLM บน DGX Spark กับ Qwen3.6-35B-A3B-NVFP4 แสดงให้เห็นว่า:

  1. Single-request throughput อยู่ที่ ~104 TPS — ใช้งานได้สบายสำหรับ interactive use
  2. Aggregate throughput เพิ่มได้ถึง 352 TPS ที่ 12 concurrent — เหมาะสำหรับ batch processing
  3. Sweet spot อยู่ที่ 4-6 concurrent สำหรับงานที่สมดุลทั้ง latency และ throughput
  4. vLLM continuous batching ทำงานได้ดี — throughput เพิ่ม 3.4x เมื่อโหลดเพิ่ม 12x

Script ที่เขียนเก็บไว้ใช้ต่อได้ ปรับ concurrency levels, prompt, max_tokens ได้ตามต้องการ และผลลัพธ์เก็บเป็น JSON สำหรับนำไปวิเคราะห์ต่อ

References​

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...