Skip to main content

Tesla V100 32GB + Qwen3.8-27B ที่ 23.6 tok/s และ 256K context — เมื่อ hardware อายุ 8 ปียังคุ้มค่า + power tuning + honest Velcro story

· 11 min read

บันทึก — 25 สิงหาคม 2569 — เช้ามืด เจอกระทู้ที่อ่านแล้วยิ้ม

คนใน r/LocalLLM โพสต์เรื่องการเอา Tesla V100 PCIe 32GB HBM2 ECC (การ์ด datacenter อายุ 8 ปี) มารัน Qwen3.8-27B Q3_K_M ที่ 23.59 tok/s ด้วย 150W + 256K native context — แถมใช้ พัดลมเซนตริฟูกัลติดด้วย Velcro tape แก้ปัญหา passive cooling

โพสต์นี้น่าสนใจเพราะตรงข้ามกับ discourse เรื่อง "ต้องซื้อ hardware ใหม่":

  • ราคา V100 32GB ตอนนี้ $150-600 (ราคา second-hand)
  • Performance ต่อวัตต์ดีกว่าที่คาด (sweet spot ที่ 150W ไม่ใช่ 200W)
  • ใช้ fp16 KV cache + flash-attn + draft-mtp → 160K+ context จริง
  • แต่ก็มี honest critique จาก ketosoy (Top 1%) ว่า 100W benchmark อาจมีปัญหา

TL;DR​

  • Setup: Ryzen 9 9950X + 48GB DDR5 + HPE/NVIDIA Tesla V100 PCIe 32GB HBM2 ECC + Ubuntu 24.04 + llama.cpp build 10499
  • Model: Qwen3.8-27B Q3_K_M, 12.86 GiB, 66/66 layers offloaded to V100, Q8 KV cache, native 262,144 context
  • Benchmark (4,096 token generation): 100W=12.91, 150W=23.59, 175W=25.58, 200W=26.96 tok/s
  • Sweet spot ที่ 150W — 83% more performance สำหรับ 50% พลังงานเพิ่ม; 175W = แค่ 8% gain, 200W = 14% gain + ความร้อนเพิ่ม
  • Cooling: 3D-printed duct + centrifugal fan + Velcro tape (ไม่ใช่ intended use) — 2700 RPM (ครึ่ง rated speed) เสียงไม่ดัง
  • Stability: 5 รอบติดที่ 150W, ไม่มี ECC errors, ไม่มี Xid errors, GPU 64°C / HBM2 67°C คงที่
  • Honest critique: ketosoy (Top 1%) ชี้ว่า 100W=12.91 น่าจะผิดเพราะ V100 อยู่ใต้ boost point — OP ตอบรับว่าจะทดสอบใหม่

Hardware: V100 ยังไม่ตาย​

Tesla V100 ออกปี 2017 (Volta architecture, sm_70) — แต่ spec ยังไม่เลว:

SpecV100 32GB PCIe (2017)
ArchitectureVolta (sm_70)
VRAM32 GB HBM2
Memory bandwidth900 GB/s
FP16 TFLOPs~28
TDP250W
ราคาตลาดปัจจุบัน$150-600 (second-hand)

ในขณะที่ DGX Spark (2026) มี:

  • 128GB unified memory แต่ bandwidth ต่ำกว่า
  • CUDA cores ใหม่ แต่ราคา $3K-$5K
  • การเปรียบเทียบ: V100 32GB × 2 ($300-1,200) vs DGX Spark × 1 ($3K-$5K) — สำหรับงาน 27B class ที่ context ยาว V100 อาจคุ้มกว่า

SlayBentos — ninfer-v100: 81 tok/sec บน V100 sm_70​

ก่อนเข้าเรื่อง OP ขอแนะนำตัวเลขที่ "ดึงดูด" จาก community — ถ้าใช้ custom port ที่ optimize สำหรับ V100 sm_70 ทำได้ 81 tok/sec (พร้อม prefill 1200 tokens) เทียบกับ stock vLLM ของ OP ที่ ~26 tok/s

Source: https://github.com/geoffwatts/ninfer-v100

# Custom port for V100 sm_70
# https://github.com/geoffwatts/ninfer-v100
# Result: 81 tok/sec with prefill 1200

SlayBentos: "I'm shocked by the performance I'm getting from Volta sm_70... I had to buy more because I'm afraid they'll sell out!"

Insights:

  • V100 sm_70 มี headroom มากกว่าที่ OP benchmark แสดง — stock vLLM ไม่ได้ optimize สำหรับ Volta
  • Custom port ได้ 3x throughput (81 vs 26 tok/s)
  • ราคา second-hand + custom port อาจคุ้มกว่า modern GPU สำหรับ budget-conscious

ตอนนี้เข้าเรื่อง OP จริงๆ — ตัวเลขของ OP เป็น stock vLLM ไม่ใช่ optimized port เลย

Setup ของ OP​

# System
Ryzen 9 9950X
48 GB DDR5
HPE/NVIDIA Tesla V100 PCIe 32 GB HBM2 ECC
Ubuntu 24.04
llama.cpp build 10499

# Model
Qwen3.8-27B Q3_K_M, 12.86 GiB
66/66 layers offloaded
Q8 KV cache
native context 262,144

# Use case
Full agent mode (read/write/edit files)
AMD box อีกเครื่องรัน ComfyUI พร้อมกัน

Power Tuning: 100W → 200W​

ผล benchmark 4,096-token generation:

PowerTok/sΔ vs 100WΔ EfficiencyNotes
100W12.91 ± 0.56baseline—"ประมาณการเต็มด้วย cooling ชั่วคราว 40mm"
150W23.59 ± 0.25+83%+1.66 tok/s per W5 รอบ, GPU 64°C / HBM2 67°C คงที่
175W25.58+98%+0.69 tok/s per Wผ่านเดียว, 69°C / 71°C
200W26.96+109%+0.46 tok/s per Wผ่านเดียว, 71°C / 73°C

Sweet spot ที่ 150W — ได้ 83% more performance สำหรับ 50% more power (linear-ish); 175W ได้แค่ 8% เพิ่ม; 200W ได้ 14% แต่ความร้อนเพิ่ม "อย่างเห็นได้ชัด"

Insight: สำหรับ daily production OP เลือก 150W — ไม่ใช่ max power คือ sweet spot เสมอ

Honest Critique: 100W Benchmark น่าจะมีปัญหา​

ketosoy (Top 1% Commenter) ชี้ประเด็นสำคัญ:

"V100 มี power floor ที่ 100W ตอน 100W การ์ดอาจอยู่ใต้ boost point — board components และ HBM2 power คงที่ เลยทำให้ที่ 100W มีพลังงานน้อยเกินไปสำหรับ compute cores

แต่ cooling อย่างเดียวไม่อธิบาย gap 13 vs 23 ได้ — ที่ 67°C ไม่ควร throttle

จาก curve ของคุณ 100W ที่ cooling เต็มที่ควรได้ ~21 tok/s ไม่ใช่ 12 (เว้นแต่พิมพ์ผิด)"

OP ตอบกลับดีมาก:

"เก็บข้อมูลได้ดีมาก ผล 12.91 tok/s ไม่ใช่พิมพ์ผิด แต่เห็นว่าค่าแบบนี้ก็น่าเช็คใหม่กับ blower ตัวใหม่

ผลนั้นมาจากการสร้างสามรอบที่ใช้ 4,096-token ที่ 100 W การ์ด GPU โชว์อุณหภูมิที่ระดับ 67°C แบบคงที่ ไม่มีปัญหาเรื่องความร้อนเลย

ตอนนี้คิดว่า 100 W ซึ่งเป็นขีดจำกัดพลังงานขั้นต่ำของการ์ดนั้น มันทำให้ V100 อยู่ใต้จุดที่ควรจะเพิ่มความเร็ว/ระดับพลังงาน

ผมไม่ได้บันทึกการทำงานของคอร์และหน่วยความจำในรอบ 100 W แรก และยังไม่ได้ทำซ้ำกับ blower ตัวใหม่เลย จะทำการ benchmark เดิมที่ 100 W กับการระบายความร้อนปัจจุบัน ถ้ายังอยู่ราวๆ 13 tok/s ก็คือ พลังงานที่ไม่เป็นเชิงเส้นนั้นมันของจริง ถ้าอยู่ใกล้ 20 ก็แสดงว่าอะไรบางอย่างในรอบแรกมันไม่สามารถเปรียบเทียวได้"

ผมชอบ response นี้มาก เพราะ:

  1. ยอมรับว่าอาจมีปัญหา (ไม่ defensive)
  2. ระบุ root cause hypothesis (under-boost point)
  3. วางแผน retest ที่ชัดเจน

นี่คือ scientific mindset ที่ดี — OP ไม่ได้ claim "ผมวัดได้ 12.91" แบบ final แต่ treat มันเป็น data point ที่ต้อง verify

Cooling Solution: 3D-Printed Duct + Velcro​

V100 PCIe เป็น passive cooling — ต้องการ airflow จาก server chassis โดยตรง

OP แก้ปัญหา:

  • 3D-printed duct
  • Centrifugal blower fan
  • Velcro tape (สองด้าน) — ไม่มีที่ยึด, ไม่มี zip ties, ไม่มี side panel
# Fan specs
Blower: 2700 RPM (ครึ่ง rated speed)
Noise: ไม่ดัง (<2000 RPM ตอน idle)
Control: BIOS-controlled (ไม่ตอบสนอง GPU load)

Saleen1310 ถามเรื่อง STL files → OP ไม่ได้ share แต่แนะนำ:

OrnateTech (ดีไซเนอร์ของ fan adapter) ตอบ:

"I'm the designer of the fan adapter... nice to see it actually working. This is the first time I've heard of someone using Velcro for this, but it's actually not a bad idea."

Techngro joke:

"Are you going to secure the fan to the GPU?" "I'll make it work how I want it to, honey..." "Heard of o-rings?" "I prefer zip ties for this kind of work, but double-sided tape worked great this time."

V100 vs Modern Alternatives: Real Comparisons จาก Community​

หลังจากเห็น 81 tok/s ของ SlayBentos แล้ว มาดูว่า community อื่นๆ ทำอะไรกับ V100 กันบ้าง — ทั้งหมดเป็น stock vLLM ไม่ใช่ custom port

Embarrassed_Adagio28 — V100 32GB + V100 16GB​

# Setup: 2 V100s with MTP + tensor parallelism
# Result: 77 tps กับ prompt เขียนโค้ดสั้นๆ
# ลงเหลือ 57 tps ในงาน 80k tokens (context ยาว)

rcriot25 — Proxmox + LXC + dual fan​

Setup:

  • AMD 2920X Threadripper, 128GB DDR4
  • V100 32GB PCIe + dual 40mm fans + 3D printed adapter
  • Proxmox 9.1.1 + LXC (Ubuntu 24.04)
  • Nvidia driver 580.159.04

Config:

./build/bin/llama-server \
--model Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-IQ4_XS.gguf \
--n-gpu-layers 99 \
--cache-ram 32768 \
--ctx-size 160144 \
--parallel 2 \
--kv-unified \
--cont-batching \
--flash-attn on \
--batch-size 4096 \
--ubatch-size 512 \
--cache-type-k f16 \
--cache-type-v f16 \
--spec-type ngram-mod,draft-mtp \
--spec-ngram-mod-n-match 24 \
--spec-ngram-mod-n-max 64 \
--threads 8 \
--split-mode layer

Result: TG เฉลี่ย ~45 t/s (32-65 range), PP 500-700 Log ตัวอย่าง:

slot print_timing: id 0 | task 906 | n_gen = 22169, tg = 47.78 t/s, tg_3s = 99.27 t/s
slot print_timing: id 0 | task 906 | n_gen = 23223, tg = 48.41 t/s, tg_3s = 65.95 t/s

illcuontheotherside: "ตั้งค่า cache fp16 เก่งจริงๆ" → fp16 KV cache ทำให้ context 160k+ ทำได้จริง

TheWolfOfWalmart — V100 vs RADEON PRO V620​

"ถ้าเรามองหาอุปกรณ์เก่าๆ ก็สามารถซื้อ V620 สองตัวในราคาใกล้เคียงกัน ตอนที่ลอง config นั้นได้ประมาณ 40 tok/s โดยใช้ 3.8 27B Q8_0 กับ Q8 KV cache และ config นี้ยังมี VRAM แค่ครึ่งนึงของ V100 ตัวเดียว"

V620 stack:

  • 8 ตัวที่ $350 ตัว (เมื่อก่อน; ตอนนี้ $600)
  • Best point: 180W
  • Q3.8 27B Q8_0 + Q8 KV = ~40 tok/s ต่อการ์ด

fallingdowndizzyvr — V100 Market Reality Check​

"ไม่ใช่ของดีราคาถูกเหมือนก่อนแล้ว

ยังหา V100 16GB ได้ที่ราคา $150 (PCIe 1x) แต่ V100 32GB PCIe ขั้นต่ำตอนนี้ $600"

Insight: second-hand enterprise hardware ขึ้นราคาแล้วเพราะ local AI community เริ่มใช้จริง

เทียบ DGX Spark vs V100 สำหรับ Qwen3.8-27B class​

จากข้อมูลใน community:

V100 32GB (2017)DGX Spark (2026)
ราคา second-hand$150-600$3,000-5,000 (ใหม่)
VRAM32 GB HBM2128 GB unified
Memory bandwidth900 GB/sต่ำกว่า (unified)
Sweet spot power150WN/A
27B Q3_K_M tok/s23.6 @ 150Wยังไม่ชัด
Native context262K128K (limit by model)
FP16 KV cacheทำได้ (V100 รองรับ)ทำได้
Flash attentionทำได้ (Volta patch)ทำได้
MTP (multi-token prediction)ใช้ได้ (rcriot25)ใช้ได้

Insight: สำหรับ 27B class model ที่ context ≤ 256K:

  • V100 32GB × 2 ($300-1,200): คุ้มกว่ามาก, แต่ต้อง deal with cooling
  • DGX Spark × 1 ($3K-$5K): จ่ายแพงกว่า 5-10x แต่ unified memory ดีกว่า + support ecosystem ดีกว่า

ถ้า use case ไม่ต้องการ unified memory 128GB — V100 ยังเป็น choice ที่คุ้มค่า

Lessons ที่ได้จากโพสต์นี้​

1. Power tuning > max power​

150W ให้ 83% performance gain vs 100W ด้วย 50% power increase — linear region

175W ให้ 8% gain (diminishing returns เริ่มต้น) 200W ให้ 14% gain + ความร้อนเพิ่ม

Max power ไม่ใช่ optimum เสมอ — ต้องวัด curve เอง

2. fp16 KV cache + flash-attn ปลดล็อก long context​

V100 รองรับ fp16 KV cache (Volta มี tensor cores FP16) + flash-attn (community patch) → 160K-262K context ทำได้จริง

ถ้าใช้ Q8 KV cache → context สั้นกว่า แต่ inference เร็วกว่า ถ้าใช้ fp16 KV cache → context ยาวกว่า แต่ inference ช้าลงเล็กน้อย

Trade-off ขึ้นกับ use case

3. Second-hand enterprise hardware ยังคุ้ม​

V100 32GB ราคา $150-600 vs DGX Spark ราคา $3K-$5K

สำหรับ budget-conscious dev + small-scale inference — V100 ยังเป็น choice ที่ดี ข้อเสีย: cooling ต้อง DIY, driver version sensitivity, no warranty

4. Honest data collection > defended benchmarks​

OP ตอบ ketosoy แบบยอมรับว่าอาจมีปัญหา — จะทดสอบใหม่

นี่คือ scientific mindset ที่ community ต้องการ — ไม่ใช่ "ผมวัดได้ 12.91, จบ"

5. Velcro tape สำหรับ cooling fix ใน residential setting​

Centrifugal blower + 3D-printed duct + Velcro tape (ไม่ใช่ intended use) — works

ไม่ต้องการ server chassis ถ้า airflow เพียงพอ + RPM control ผ่าน BIOS

สรุป​

Tesla V100 32GB ที่ second-hand ราคา $150-600 + power tuning ที่ 150W (sweet spot) + fp16 KV cache + flash-attn + Velcro-taped cooling → 23.59 tok/s บน Qwen3.8-27B Q3_K_M + native 256K context ได้จริง

OP ตอบ critique จาก ketosoy แบบ scientific mindset — ยอมรับว่า 100W benchmark อาจมีปัญหา (under-boost point hypothesis) และวางแผน retest นี่คือตัวอย่างที่ดีของ "honest data collection" ต่างจาก "defended benchmarks"

ถ้ามี V100 อยู่แล้วและไม่ต้องการ unified memory 128GB — V100 ยังเป็น choice ที่คุ้มค่าสำหรับ 27B class model ที่ context ≤ 256K

Cross-references​

ถ้าสนใจ local LLM infrastructure ลองอ่าน:

อ้างอิง​

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...