Tesla V100 32GB + Qwen3.8-27B ที่ 23.6 tok/s และ 256K context — เมื่อ hardware อายุ 8 ปียังคุ้มค่า + power tuning + honest Velcro story
สารบัญ
- TL;DR
- Hardware: V100 ยังไม่ตาย
- SlayBentos — ninfer-v100: 81 tok/sec บน V100 sm_70
- Setup ของ OP
- Power Tuning: 100W → 200W
- Honest Critique: 100W Benchmark น่าจะมีปัญหา
- Cooling Solution: 3D-Printed Duct + Velcro
- V100 vs Modern Alternatives: Real Comparisons จาก Community
- Embarrassed_Adagio28 — V100 32GB + V100 16GB
- rcriot25 — Proxmox + LXC + dual fan
- TheWolfOfWalmart — V100 vs RADEON PRO V620
- fallingdowndizzyvr — V100 Market Reality Check
- เทียบ DGX Spark vs V100 สำหรับ Qwen3.8-27B class
- Lessons ที่ได้จากโพสต์นี้
- 1. Power tuning > max power
- 2. fp16 KV cache + flash-attn ปลดล็อก long context
- 3. Second-hand enterprise hardware ยังคุ้ม
- 4. Honest data collection > defended benchmarks
- 5. Velcro tape สำหรับ cooling fix ใน residential setting
- สรุป
- Cross-references
- อ้างอิง
บันทึก — 25 สิงหาคม 2569 — เช้ามืด เจอกระทู้ที่อ่านแล้วยิ้ม
คนใน r/LocalLLM โพสต์เรื่องการเอา Tesla V100 PCIe 32GB HBM2 ECC (การ์ด datacenter อายุ 8 ปี) มารัน Qwen3.8-27B Q3_K_M ที่ 23.59 tok/s ด้วย 150W + 256K native context — แถมใช้ พัดลมเซนตริฟูกัลติดด้วย Velcro tape แก้ปัญหา passive cooling
โพสต์นี้น่าสนใจเพราะตรงข้ามกับ discourse เรื่อง "ต้องซื้อ hardware ใหม่":
- ราคา V100 32GB ตอนนี้ $150-600 (ราคา second-hand)
- Performance ต่อวัตต์ดีกว่าที่คาด (sweet spot ที่ 150W ไม่ใช่ 200W)
- ใช้ fp16 KV cache + flash-attn + draft-mtp → 160K+ context จริง
- แต่ก็มี honest critique จาก ketosoy (Top 1%) ว่า 100W benchmark อาจมีปัญหา
TL;DR
- Setup: Ryzen 9 9950X + 48GB DDR5 + HPE/NVIDIA Tesla V100 PCIe 32GB HBM2 ECC + Ubuntu 24.04 + llama.cpp build 10499
- Model: Qwen3.8-27B Q3_K_M, 12.86 GiB, 66/66 layers offloaded to V100, Q8 KV cache, native 262,144 context
- Benchmark (4,096 token generation): 100W=12.91, 150W=23.59, 175W=25.58, 200W=26.96 tok/s
- Sweet spot ที่ 150W — 83% more performance สำหรับ 50% พลังงานเพิ่ม; 175W = แค่ 8% gain, 200W = 14% gain + ความร้อนเพิ่ม
- Cooling: 3D-printed duct + centrifugal fan + Velcro tape (ไม่ใช่ intended use) — 2700 RPM (ครึ่ง rated speed) เสียงไม่ดัง
- Stability: 5 รอบติดที่ 150W, ไม่มี ECC errors, ไม่มี Xid errors, GPU 64°C / HBM2 67°C คงที่
- Honest critique: ketosoy (Top 1%) ชี้ว่า 100W=12.91 น่าจะผิดเพราะ V100 อยู่ใต้ boost point — OP ตอบรับว่าจะทดสอบใหม่
Hardware: V100 ยังไม่ตาย
Tesla V100 ออกปี 2017 (Volta architecture, sm_70) — แต่ spec ยังไม่เลว:
| Spec | V100 32GB PCIe (2017) |
|---|---|
| Architecture | Volta (sm_70) |
| VRAM | 32 GB HBM2 |
| Memory bandwidth | 900 GB/s |
| FP16 TFLOPs | ~28 |
| TDP | 250W |
| ราคาตลาดปัจจุบัน | $150-600 (second-hand) |
ในขณะที่ DGX Spark (2026) มี:
- 128GB unified memory แต่ bandwidth ต่ำกว่า
- CUDA cores ใหม่ แต่ราคา $3K-$5K
- การเปรียบเทียบ: V100 32GB × 2 ($300-1,200) vs DGX Spark × 1 ($3K-$5K) — สำหรับงาน 27B class ที่ context ยาว V100 อาจคุ้มกว่า
SlayBentos — ninfer-v100: 81 tok/sec บน V100 sm_70
ก่อนเข้าเรื่อง OP ขอแนะนำตัวเลขที่ "ดึงดูด" จาก community — ถ้าใช้ custom port ที่ optimize สำหรับ V100 sm_70 ทำได้ 81 tok/sec (พร้อม prefill 1200 tokens) เทียบกับ stock vLLM ของ OP ที่ ~26 tok/s
Source: https://github.com/geoffwatts/ninfer-v100
# Custom port for V100 sm_70
# https://github.com/geoffwatts/ninfer-v100
# Result: 81 tok/sec with prefill 1200
SlayBentos: "I'm shocked by the performance I'm getting from Volta sm_70... I had to buy more because I'm afraid they'll sell out!"
Insights:
- V100 sm_70 มี headroom มากกว่าที่ OP benchmark แสดง — stock vLLM ไม่ได้ optimize สำหรับ Volta
- Custom port ได้ 3x throughput (81 vs 26 tok/s)
- ราคา second-hand + custom port อาจคุ้มกว่า modern GPU สำหรับ budget-conscious
ตอนนี้เข้าเรื่อง OP จริงๆ — ตัวเลขของ OP เป็น stock vLLM ไม่ใช่ optimized port เลย
Setup ของ OP
# System
Ryzen 9 9950X
48 GB DDR5
HPE/NVIDIA Tesla V100 PCIe 32 GB HBM2 ECC
Ubuntu 24.04
llama.cpp build 10499
# Model
Qwen3.8-27B Q3_K_M, 12.86 GiB
66/66 layers offloaded
Q8 KV cache
native context 262,144
# Use case
Full agent mode (read/write/edit files)
AMD box อีกเครื่องรัน ComfyUI พร้อมกัน
Power Tuning: 100W → 200W
ผล benchmark 4,096-token generation:
| Power | Tok/s | Δ vs 100W | Δ Efficiency | Notes |
|---|---|---|---|---|
| 100W | 12.91 ± 0.56 | baseline | — | "ประมาณการเต็มด้วย cooling ชั่วคราว 40mm" |
| 150W | 23.59 ± 0.25 | +83% | +1.66 tok/s per W | 5 รอบ, GPU 64°C / HBM2 67°C คงที่ |
| 175W | 25.58 | +98% | +0.69 tok/s per W | ผ่านเดียว, 69°C / 71°C |
| 200W | 26.96 | +109% | +0.46 tok/s per W | ผ่านเดียว, 71°C / 73°C |
Sweet spot ที่ 150W — ได้ 83% more performance สำหรับ 50% more power (linear-ish); 175W ได้แค่ 8% เพิ่ม; 200W ได้ 14% แต่ความร้อนเพิ่ม "อย่างเห็นได้ชัด"
Insight: สำหรับ daily production OP เลือก 150W — ไม่ใช่ max power คือ sweet spot เสมอ
Honest Critique: 100W Benchmark น่าจะมีปัญหา
ketosoy (Top 1% Commenter) ชี้ประเด็นสำคัญ:
"V100 มี power floor ที่ 100W ตอน 100W การ์ดอาจอยู่ใต้ boost point — board components และ HBM2 power คงที่ เลยทำให้ที่ 100W มีพลังงานน้อยเกินไปสำหรับ compute cores
แต่ cooling อย่างเดียวไม่อธิบาย gap 13 vs 23 ได้ — ที่ 67°C ไม่ควร throttle
จาก curve ของคุณ 100W ที่ cooling เต็มที่ควรได้ ~21 tok/s ไม่ใช่ 12 (เว้นแต่พิมพ์ผิด)"
OP ตอบกลับดีมาก:
"เก็บข้อมูลได้ดีมาก ผล 12.91 tok/s ไม่ใช่พิมพ์ผิด แต่เห็นว่าค่าแบบนี้ก็น่าเช็คใหม่กับ blower ตัวใหม่
ผลนั้นมาจากการสร้างสามรอบที่ใช้ 4,096-token ที่ 100 W การ์ด GPU โชว์อุณหภูมิที่ระดับ 67°C แบบคงที่ ไม่มีปัญหาเรื่องความร้อนเลย
ตอนนี้คิดว่า 100 W ซึ่งเป็นขีดจำกัดพลังงานขั้นต่ำของการ์ดนั้น มันทำให้ V100 อยู่ใต้จุดที่ควรจะเพิ่มความเร็ว/ระดับพลังงาน
ผมไม่ได้บันทึกการทำงานของคอร์และหน่วยความจำในรอบ 100 W แรก และยังไม่ได้ทำซ้ำกับ blower ตัวใหม่เลย จะทำการ benchmark เดิมที่ 100 W กับการระบายความร้อนปัจจุบัน ถ้ายังอยู่ราวๆ 13 tok/s ก็คือ พลังงานที่ไม่เป็นเชิงเส้นนั้นมันของจริง ถ้าอยู่ใกล้ 20 ก็แสดงว่าอะไรบางอย่างในรอบแรกมันไม่สามารถเปรียบเทียวได้"
ผมชอบ response นี้มาก เพราะ:
- ยอมรับว่าอาจมีปัญหา (ไม่ defensive)
- ระบุ root cause hypothesis (under-boost point)
- วางแผน retest ที่ชัดเจน
นี่คือ scientific mindset ที่ดี — OP ไม่ได้ claim "ผมวัดได้ 12.91" แบบ final แต่ treat มันเป็น data point ที่ต้อง verify
Cooling Solution: 3D-Printed Duct + Velcro
V100 PCIe เป็น passive cooling — ต้องการ airflow จาก server chassis โดยตรง
OP แก้ปัญหา:
- 3D-printed duct
- Centrifugal blower fan
- Velcro tape (สองด้าน) — ไม่มีที่ยึด, ไม่มี zip ties, ไม่มี side panel
# Fan specs
Blower: 2700 RPM (ครึ่ง rated speed)
Noise: ไม่ดัง (<2000 RPM ตอน idle)
Control: BIOS-controlled (ไม่ตอบสนอง GPU load)
Saleen1310 ถามเรื่อง STL files → OP ไม่ได้ share แต่แนะนำ:
- Wathai Brushless Centrifugal Inflatable
- Amazon ASIN B0DDJM7X4R (อีกตัว)
OrnateTech (ดีไซเนอร์ของ fan adapter) ตอบ:
"I'm the designer of the fan adapter... nice to see it actually working. This is the first time I've heard of someone using Velcro for this, but it's actually not a bad idea."
Techngro joke:
"Are you going to secure the fan to the GPU?" "I'll make it work how I want it to, honey..." "Heard of o-rings?" "I prefer zip ties for this kind of work, but double-sided tape worked great this time."
V100 vs Modern Alternatives: Real Comparisons จาก Community
หลังจากเห็น 81 tok/s ของ SlayBentos แล้ว มาดูว่า community อื่นๆ ทำอะไรกับ V100 กันบ้าง — ทั้งหมดเป็น stock vLLM ไม่ใช่ custom port
Embarrassed_Adagio28 — V100 32GB + V100 16GB
# Setup: 2 V100s with MTP + tensor parallelism
# Result: 77 tps กับ prompt เขียนโค้ดสั้นๆ
# ลงเหลือ 57 tps ในงาน 80k tokens (context ยาว)
rcriot25 — Proxmox + LXC + dual fan
Setup:
- AMD 2920X Threadripper, 128GB DDR4
- V100 32GB PCIe + dual 40mm fans + 3D printed adapter
- Proxmox 9.1.1 + LXC (Ubuntu 24.04)
- Nvidia driver 580.159.04
Config:
./build/bin/llama-server \
--model Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-IQ4_XS.gguf \
--n-gpu-layers 99 \
--cache-ram 32768 \
--ctx-size 160144 \
--parallel 2 \
--kv-unified \
--cont-batching \
--flash-attn on \
--batch-size 4096 \
--ubatch-size 512 \
--cache-type-k f16 \
--cache-type-v f16 \
--spec-type ngram-mod,draft-mtp \
--spec-ngram-mod-n-match 24 \
--spec-ngram-mod-n-max 64 \
--threads 8 \
--split-mode layer
Result: TG เฉลี่ย ~45 t/s (32-65 range), PP 500-700 Log ตัวอย่าง:
slot print_timing: id 0 | task 906 | n_gen = 22169, tg = 47.78 t/s, tg_3s = 99.27 t/s
slot print_timing: id 0 | task 906 | n_gen = 23223, tg = 48.41 t/s, tg_3s = 65.95 t/s
illcuontheotherside: "ตั้งค่า cache fp16 เก่งจริงๆ" → fp16 KV cache ทำให้ context 160k+ ทำได้จริง
TheWolfOfWalmart — V100 vs RADEON PRO V620
"ถ้าเรามองหาอุปกรณ์เก่าๆ ก็สามารถซื้อ V620 สองตัวในราคาใกล้เคียงกัน ตอนที่ลอง config นั้นได้ประมาณ 40 tok/s โดยใช้ 3.8 27B Q8_0 กับ Q8 KV cache และ config นี้ยังมี VRAM แค่ครึ่งนึงของ V100 ตัวเดียว"
V620 stack:
- 8 ตัวที่ $350 ตัว (เมื่อก่อน; ตอนนี้ $600)
- Best point: 180W
- Q3.8 27B Q8_0 + Q8 KV = ~40 tok/s ต่อการ์ด
fallingdowndizzyvr — V100 Market Reality Check
"ไม่ใช่ของดีราคาถูกเหมือนก่อนแล้ว
ยังหา V100 16GB ได้ที่ราคา $150 (PCIe 1x) แต่ V100 32GB PCIe ขั้นต่ำตอนนี้ $600"
Insight: second-hand enterprise hardware ขึ้นราคาแล้วเพราะ local AI community เริ่มใช้จริง
เทียบ DGX Spark vs V100 สำหรับ Qwen3.8-27B class
จากข้อมูลใน community:
| V100 32GB (2017) | DGX Spark (2026) | |
|---|---|---|
| ราคา second-hand | $150-600 | $3,000-5,000 (ใหม่) |
| VRAM | 32 GB HBM2 | 128 GB unified |
| Memory bandwidth | 900 GB/s | ต่ำกว่า (unified) |
| Sweet spot power | 150W | N/A |
| 27B Q3_K_M tok/s | 23.6 @ 150W | ยังไม่ชัด |
| Native context | 262K | 128K (limit by model) |
| FP16 KV cache | ทำได้ (V100 รองรับ) | ทำได้ |
| Flash attention | ทำได้ (Volta patch) | ทำได้ |
| MTP (multi-token prediction) | ใช้ได้ (rcriot25) | ใช้ได้ |
Insight: สำหรับ 27B class model ที่ context ≤ 256K:
- V100 32GB × 2 ($300-1,200): คุ้มกว่ามาก, แต่ต้อง deal with cooling
- DGX Spark × 1 ($3K-$5K): จ่ายแพงกว่า 5-10x แต่ unified memory ดีกว่า + support ecosystem ดีกว่า
ถ้า use case ไม่ต้องการ unified memory 128GB — V100 ยังเป็น choice ที่คุ้มค่า
Lessons ที่ได้จากโพสต์นี้
1. Power tuning > max power
150W ให้ 83% performance gain vs 100W ด้วย 50% power increase — linear region
175W ให้ 8% gain (diminishing returns เริ่มต้น) 200W ให้ 14% gain + ความร้อนเพิ่ม
Max power ไม่ใช่ optimum เสมอ — ต้องวัด curve เอง
2. fp16 KV cache + flash-attn ปลดล็อก long context
V100 รองรับ fp16 KV cache (Volta มี tensor cores FP16) + flash-attn (community patch) → 160K-262K context ทำได้จริง
ถ้าใช้ Q8 KV cache → context สั้นกว่า แต่ inference เร็วกว่า ถ้าใช้ fp16 KV cache → context ยาวกว่า แต่ inference ช้าลงเล็กน้อย
Trade-off ขึ้นกับ use case
3. Second-hand enterprise hardware ยังคุ้ม
V100 32GB ราคา $150-600 vs DGX Spark ราคา $3K-$5K
สำหรับ budget-conscious dev + small-scale inference — V100 ยังเป็น choice ที่ดี ข้อเสีย: cooling ต้อง DIY, driver version sensitivity, no warranty
4. Honest data collection > defended benchmarks
OP ตอบ ketosoy แบบยอมรับว่าอาจมีปัญหา — จะทดสอบใหม่
นี่คือ scientific mindset ที่ community ต้องการ — ไม่ใช่ "ผมวัดได้ 12.91, จบ"
5. Velcro tape สำหรับ cooling fix ใน residential setting
Centrifugal blower + 3D-printed duct + Velcro tape (ไม่ใช่ intended use) — works
ไม่ต้องการ server chassis ถ้า airflow เพียงพอ + RPM control ผ่าน BIOS
สรุป
Tesla V100 32GB ที่ second-hand ราคา $150-600 + power tuning ที่ 150W (sweet spot) + fp16 KV cache + flash-attn + Velcro-taped cooling → 23.59 tok/s บน Qwen3.8-27B Q3_K_M + native 256K context ได้จริง
OP ตอบ critique จาก ketosoy แบบ scientific mindset — ยอมรับว่า 100W benchmark อาจมีปัญหา (under-boost point hypothesis) และวางแผน retest นี่คือตัวอย่างที่ดีของ "honest data collection" ต่างจาก "defended benchmarks"
ถ้ามี V100 อยู่แล้วและไม่ต้องการ unified memory 128GB — V100 ยังเป็น choice ที่คุ้มค่าสำหรับ 27B class model ที่ context ≤ 256K
Cross-references
ถ้าสนใจ local LLM infrastructure ลองอ่าน:
- $10K budget — ซื้อ 2× DGX Spark เลย หรือรอ? — new hardware perspective
- ซื้อ DGX Spark เครื่องเดียว เพื่อเรียนรู้ + Hermes agent คุ้มไหม — awkward spot
- หยุดใช้ Ollama — runtime perspective
อ้างอิง
- Tesla V100 32GB + Qwen3.8-27B ที่ 23.6 tok/s — r/LocalLLM — โพสต์ต้นเรื่อง
- ninfer-v100 — GitHub — SlayBentos's custom port สำหรับ V100 sm_70 (81 tok/s)
- Tesla V100 — Wikipedia — spec ทั่วไป
- Volta (microarchitecture) — Wikipedia — sm_70 specs
- llama.cpp flash-attn V100 patch — mainline build 10499+
- Wathai Brushless Centrifugal Inflatable — Amazon — fan ที่ OP ใช้
- OrnateTech fan adapter — Reddit reference — designer of fan adapter
- NVIDIA V100 HBM2 ECC spec sheet — official NVIDIA
- Proxmox 9.1.1 — official site — virtualization ที่ rcriot25 ใช้
- LlamaBarn — GitHub — official macOS llama.cpp GUI
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee