Skip to main content

2 posts tagged with "local-llm"

View All Tags

หยุดใช้ Ollama — เมื่อ local LLM wrapper ที่ใหญ่ที่สุดมีปัญหาเรื่อง attribution, performance, และ vendor lock-in

· 13 min read

บันทึก — 25 สิงหาคม 2569 — เช้ามืด เจอกระทู้ที่ถกกันหนักมาก

มีคนโพสต์บล็อกชื่อ "Stop using Ollama" ใน r/LocalLLaMA ขึ้นไป Hacker News ได้ 648 points + 207 comments — ซึ่งสำหรับ HN ถือว่าเยอะมาก

บล็อกนี้ไม่ได้บอกว่า Ollama แย่ — มันบอกว่า Ollama เป็น wrapper ที่สร้างบน llama.cpp แต่ใช้เวลาหลายปีหลบเลี่ยง attribution + เปิดตัว closed-source components + pivot ไป cloud + ทำผลงานช้ากว่า upstream 1.8 เท่า และมี alternatives ที่ดีกว่า

ผมอ่านบล็อก + Reddit + HN comments แล้วเขียนสรุปนี้ — พร้อม honest disagreement ว่า Ollama ก็ยังมี use case ที่ถูกต้องสำหรับบางคน

Tesla V100 32GB + Qwen3.8-27B ที่ 23.6 tok/s และ 256K context — เมื่อ hardware อายุ 8 ปียังคุ้มค่า + power tuning + honest Velcro story

· 11 min read

บันทึก — 25 สิงหาคม 2569 — เช้ามืด เจอกระทู้ที่อ่านแล้วยิ้ม

คนใน r/LocalLLM โพสต์เรื่องการเอา Tesla V100 PCIe 32GB HBM2 ECC (การ์ด datacenter อายุ 8 ปี) มารัน Qwen3.8-27B Q3_K_M ที่ 23.59 tok/s ด้วย 150W + 256K native context — แถมใช้ พัดลมเซนตริฟูกัลติดด้วย Velcro tape แก้ปัญหา passive cooling

โพสต์นี้น่าสนใจเพราะตรงข้ามกับ discourse เรื่อง "ต้องซื้อ hardware ใหม่":

  • ราคา V100 32GB ตอนนี้ $150-600 (ราคา second-hand)
  • Performance ต่อวัตต์ดีกว่าที่คาด (sweet spot ที่ 150W ไม่ใช่ 200W)
  • ใช้ fp16 KV cache + flash-attn + draft-mtp → 160K+ context จริง
  • แต่ก็มี honest critique จาก ketosoy (Top 1%) ว่า 100W benchmark อาจมีปัญหา