Skip to main content

หยุดใช้ Ollama — เมื่อ local LLM wrapper ที่ใหญ่ที่สุดมีปัญหาเรื่อง attribution, performance, และ vendor lock-in

· 13 min read

สารบัญ

บันทึก — 25 สิงหาคม 2569 — เช้ามืด เจอกระทู้ที่ถกกันหนักมาก

มีคนโพสต์บล็อกชื่อ "Stop using Ollama" ใน r/LocalLLaMA ขึ้นไป Hacker News ได้ 648 points + 207 comments — ซึ่งสำหรับ HN ถือว่าเยอะมาก

บล็อกนี้ไม่ได้บอกว่า Ollama แย่ — มันบอกว่า Ollama เป็น wrapper ที่สร้างบน llama.cpp แต่ใช้เวลาหลายปีหลบเลี่ยง attribution + เปิดตัว closed-source components + pivot ไป cloud + ทำผลงานช้ากว่า upstream 1.8 เท่า และมี alternatives ที่ดีกว่า

ผมอ่านบล็อก + Reddit + HN comments แล้วเขียนสรุปนี้ — พร้อม honest disagreement ว่า Ollama ก็ยังมี use case ที่ถูกต้องสำหรับบางคน

TL;DR​

  • Zetaphor (sleepingrobots.com) ทำ timeline ของ Ollama's actions ที่ "bad-faith open source citizen" — 400+ วันที่ binary ไม่มี MIT notice ของ llama.cpp, README ไม่มี attribution เป็นปี
  • Performance: llama.cpp เร็วกว่า Ollama 1.8x (161 vs 89 tok/s บน hardware เดียวกัน), CPU gap 30-50%, Qwen-3 Coder 32B gap ~70%
  • Ollama "fork ไปทำเอง" บน ggml ในปี 2025 → reintroduce bugs ที่ llama.cpp แก้ไปแล้ว + ทำ GPT-OSS 20B พัง
  • Alternatives ที่ได้ทั้ง UX + ไม่ vendor-lock-in: llama.cpp ตรง, LM Studio, llamafile, koboldcpp, Jan, ramalama
  • Honest disagreement: Ollama ยังเหมาะกับ beginner + quick test + "แค่อยากลองโมเดล" — ไม่ใช่ทุกคนต้องย้าย
  • HN comment ที่น่าสนใจ: บล็อกถูก criticize ว่า "เขียนเหมือน LLM" (slop) — Zetaphor ตอบว่า "I guess I write like an LLM :P"

บล็อกต้นเรื่อง​

Zetaphor จาก sleepingrobots.com (เขียนบล็อกเรื่อง LLM บน Strix Halo) โพสต์บล็อกชื่อ "Stop using Ollama" — เขาเขียนว่า:

"Ollama is the most popular way to run local LLMs. It shouldn't be. It gained that position by being first, the first tool that made llama.cpp accessible to people who didn't want to compile C++ or write their own server configs. That was a real contribution, briefly."

"But the project has since spent years systematically obscuring where its actual technology comes from, misleading users about what they're running, and drifting from the local-first mission that earned it trust in the first place. All while taking venture capital money."

เขาไม่ได้บอกว่า Ollama ไม่มีประโยชน์ — เขาบอกว่า ตัวเลือกที่ดีกว่ามีอยู่แล้ว และ Ollama เลือกที่จะไม่ evolve ไปในทางที่ดี

10 ข้อกล่าวหาหลักของ Zetaphor​

1. License violation — 400+ วันไม่มี MIT notice​

Ollama binary distributions ไม่ได้ใส่ MIT license notice ของ llama.cpp ตามที่ MIT license บังคับ (ข้อเดียวที่สำคัญที่สุดของ MIT)

  • GitHub issue #3185 เปิดใน early 2024 ขอให้แก้ license compliance
  • ไม่มีคนตอบจาก maintainers เป็นเวลา 400+ วัน
  • จนกระทั่ง April 2024 มี issue #3697 เปิดใหม่ + PR #3700 ตามมาในไม่กี่ชั่วโมง
  • Michael Chiang (co-founder) เพิ่มแค่บรรทัดเดียวที่ล่างสุดของ README: "llama.cpp project founded by Georgi Gerganov"

2. Minimal attribution — README ไม่พูดถึง llama.cpp เป็นเวลา 1+ ปี​

ไม่ใช่แค่ไม่มี license notice — แต่ README + website + marketing materials ไม่พูดถึง llama.cpp เลย ทั้งที่ Ollama ทุกอย่างวิ่งอยู่บน llama.cpp

คนใน HN ตอบ:

"I'm continually puzzled by their approach, it's such self-inflicted negative PR. Building on llama is perfectly valid and they're adding value on ease of use here. Just give the llama team proper credit."

3. The bad fork — mid-2025 ย้ายจาก llama.cpp ไปใช้ ggml ตรงๆ​

Ollama follow through ตามที่บอก — ย้ายจาก llama.cpp inference backend → เขียน custom implementation บน ggml (tensor library ที่ llama.cpp ใช้ข้างใต้)

เหตุผลที่บอก: "stability — llama.cpp moves fast and breaks things"

ผลลัพธ์:

  • reintroduce bugs ที่ llama.cpp แก้ไปนานแล้ว
  • structured output พัง, vision models fail, GGML assertion crashes
  • GPT-OSS 20B fail ใน Ollama แต่ work ใน upstream llama.cpp เพราะ Ollama ไม่รองรับ tensor types ที่โมเดลต้องการ
  • Georgi Gerganov (ผู้สร้าง llama.cpp) tweet ว่า Ollama "forked and made bad changes to GGML"

4. Performance gap — llama.cpp เร็วกว่า 1.8x​

SourceSetupllama.cppOllamaGap
arsturn.comSame hardware, same model161 tok/s89 tok/s1.8x
deploybase.aiCPU inference——30-50%
Reddit r/LocalLLaMAQwen-3 Coder 32B——~70% throughput

Root cause ของ gap:

  • Daemon layer overhead ของ Ollama
  • Poor GPU offloading heuristics
  • Vendored backend ที่ตามหลัง upstream

5. Misleading model naming — deepseek-r1 ไม่ใช่ R1 จริง​

ตอน DeepSeek ปล่อย R1 family (January 2025) — Ollama list "DeepSeek-R1-Distill-Qwen-32B" (8B Qwen-derived distill) ใน registry ว่า "DeepSeek-R1" โดย strip "Distill" prefix ออก

ผลกระทบ:

  • คนรัน ollama run deepseek-r1 คิดว่าได้ 671B R1 จริง → ได้ 8B distillate
  • Performance ไม่ตรงกับความคาดหวัง
  • Reputational damage กับ DeepSeek

GitHub issues #8557 + #8698 ขอให้แยก — ปิดเป็น duplicate โดยไม่แก้

6. Modelfile = reinventing a solved problem​

GGUF (โดย Georgi Gerganov) ออกแบบมาให้ single-file deployment — chat template, stop tokens, metadata ฝังอยู่ในไฟล์เดียว

Ollama เพิ่ม Modelfile ขึ้นมา (Dockerfile-inspired):

  • ต้อง extract chat template จาก GGUF → translate เป็น Go template syntax (ซึ่งต่างจาก Jinja)
  • Ollama auto-detect chat template แค่จาก hardcoded list — ถ้า GGUF มี Jinja template ที่ Ollama ไม่รู้จัก → fallback เป็น bare {{ .Prompt }} → instruction format พังเงียบๆ
  • เปลี่ยน parameter เช่น temperature: ต้อง export modelfile → edit → ollama create ใหม่ → copy model ทั้งก้อน 30-60 GB เพื่อเปลี่ยน 1 parameter

User บ่น: "The 'modelfile' workflow is a pain in the booty... 30 to 60GB and copying the entire thing to change one parameter is just dumb."

7. Closed-source GUI app — July 2025​

Ollama ship GUI desktop app (macOS + Windows):

  • พัฒนาใน private repo (github.com/ollama/app)
  • ไม่มี license
  • Source code ไม่ public
  • มี potential AGPL-3.0 dependencies ใน binary
  • Download button อยู่ข้างๆ GitHub link → คนเข้าใจผิดว่าโหลด open-source tool

PR #12933 merge เข้า main repo November 2025 — แต่ initial rollout แสดงให้เห็นว่า "project's instincts lie"

8. Cloud pivot + CVE-2025-51471​

ปลายปี 2025: Ollama เริ่ม route prompts ไป third-party cloud providers (proprietary models เช่น MiniMax-m2.7)

  • Documentation: "we process your prompts... do not store or log" — แต่ไม่พูดถึง สิ่งที่ third-party provider ทำ
  • Alibaba Cloud models → ไม่มี zero-data-retention guarantee

CVE-2025-51471: Token exfiltration — malicious registry server หลอก Ollama ส่ง auth token ไป attacker endpoint ตอน pull model

9. Registry bottleneck — quant + packaging​

  • รองรับ quant แค่ Q4_K_M, Q8_0, F16, F32 (ไม่มี Q5, Q6, IQ quants ที่ llama.cpp มีมานานแล้ว)
  • Model ใหม่ออก → ต้องรอคน Ollama package ใน registry → "โมเดลไม่มีใน Ollama" จนกว่าจะ publish
  • Reddit PSA: "If you want to test new models, use llama.cpp/transformers/vLLM/SGLang" — Qwen models พัง tool calls + garbage responses เฉพาะใน Ollama

10. VC pattern (5 steps)​

Zetaphor สรุปเป็น playbook ที่คุ้นตา:

  1. Launch on open source — build บน llama.cpp, gain community trust
  2. Minimize attribution — ทำให้ product ดู self-sufficient ต่อ investors
  3. Create lock-in — proprietary registry format, hashed filenames
  4. Launch closed-source components — GUI app
  5. Add cloud services — monetization vector

Ollama = YC W21, founded โดยคนที่เคยทำ Kitematic (Docker GUI acquired by Docker Inc.)

Honest disagreement: Ollama ยังมี use case ที่ถูกต้อง​

แม้ข้อกล่าวหาจะหนัก แต่ผมเห็น use case ที่ Ollama ยังเหมาะสม:

1. "แค่อยากลองโมเดลใหม่เร็วๆ"​

ollama run qwen2.5:7b

ไม่มีอะไรตรงๆ ง่ายขนาดนี้ — llama.cpp -hf unsloth/Qwen3.5-35B-A3B-GGUF:Q4_K_M ก็ทำได้ แต่คนทั่วไปต้องติดตั้ง build ก่อน

2. Windows user ที่ไม่ใช่ dev​

# Ollama
winget install Ollama.Ollama
ollama run llama3.2
# llama.cpp บน Windows ต้อง build หรือใช้ wsl

inagy ใน r/LocalLLaMA: "Ollama gives an easy to setup server, especially on Windows. Couple clicks in a setup wizard, then it's running in the background as a service."

3. Beginner ที่อยากเข้าโลก local LLM​

gnooggi: "Easy for lazy people + beginners who don't know what they're doing, because it simply works"

Forsaken_Ad_774: "I'm not chasing down the maximum optimizations, I'm just using something that just works with minimum amount of effort. Coupled with Cline/openwebui I'm golden."

4. ทางเลือกที่ "พอดี" จริงๆ​

Youth18: "LM Studio is still good for testing models — quick downloads, easy to see popular LLMs. But I use that for testing and then if I like a model I integrate it into my llama.cpp + custom model loading manager stack."

Insight: Ollama ไม่ได้แย่ — แค่ "เป็น stepping stone" ที่ดี แต่ไม่ใช่ปลายทาง

Alternatives ที่ควรพิจารณา​

Zetaphor แนะนำ alternatives ที่หลากหลาย ตาม use case:

1. llama.cpp ตรง — เร็วสุด + ควบคุมได้สุด​

# Install
brew install llama.cpp # macOS
apt install llama.cpp # Linux

# Run from HuggingFace directly
llama-server -hf unsloth/Qwen3.5-35B-A3B-GGUF:Q4_K_M --port 8000

# Open http://localhost:8000 — web UI ในตัว

ข้อดี:

  • เร็วที่สุด (1.8x Ollama)
  • GGUF ตรงจาก HuggingFace — ไม่ต้องรอ registry
  • OpenAI-compatible API (/v1/chat/completions)
  • ทุก quant (Q5, Q6, IQ, etc.)
  • Embed params ใน GGUF ไม่ต้อง Modelfile

ข้อเสีย:

  • ต้อง install เอง
  • ชื่อ .cpp ทำให้คนคิดว่าเป็น library (HN comment: "for as long as its name is '.cpp', people are going to think it's a C++ library and avoid it")
  • โมเดลใหม่ต้อง update llama.cpp ก่อน (gbitten: "move fast and break things")

2. LM Studio — GUI wrapper ที่ดีกว่า​

ข้อดี:

  • Same one-click convenience เหมือน Ollama
  • Accept any GGUF
  • Expose all knobs
  • มี acknowledgements page ที่ credit llama.cpp ตรงๆ
  • Zetaphor เองแนะนำ: "If you don't want to think about it, LM Studio is probably the best choice"

ข้อเสีย:

  • Proprietary (ไม่ใช่ open-source)
  • แต่ Zetaphor ชี้ว่า "It's a closed-source product, but it's not a parasitic one"

3. llamafile — Mozilla, single executable​

brabel (เคยช่วย llama.cpp): "It's truly open source, backed by Mozilla, openly uses llama.cpp"

# Download single binary
wget https://example.com/llamafile-with-model

# Run
./llamafile-with-model

ข้อดี:

  • Single executable รันได้ 6 OSes
  • ไม่ต้อง install
  • โดย Justine Tunkey (CosmopolitanC fame) — wizard-level engineer

4. llama-swap + LiteLLM — multi-model orchestration​

# llama-swap handles hot-swapping models
# LiteLLM routes across multiple backends

5. Open-source GUI alternatives​

  • Jan (jan.ai) — AGPLv3, local-first chat
  • koboldcpp — AGPL, llama.cpp fork + extensive config (HN: "minimal dependencies, no installers, no logs, don't do anything to user's system they didn't ask")
  • ramalama (Red Hat) — container-native + explicit credits

6. ถ้าจะใช้ llama.cpp แต่ชอบ GUI​

LlamaBarn (macOS) + llama-server (Linux/Windows) — official GUI จาก ggml-org เอง

HN discussion: slop critique​

มี comment ที่ถูก flag ว่า "LLM-written":

IshKebab ชี้:

"Short sentences like 'It shouldn't be.', 'I've moved on.', 'Ollama didn't.', etc. Not-this-but-that like 'The local LLM ecosystem doesn't need Ollama. It needs llama.cpp.' Weird signposting: 'Benchmarks tell the story.' Here's-the-rub conclusion: 'The Bigger Picture' Starting every title with 'The...'

It's definitely largely human-written, but there are enough slop-isms to make it annoying to read."

Zetaphor ตอบกลับ:

"I guess I write like an LLM :P Probably a side effect of using them so much"

ผมอ่านแล้วเห็นด้วยกับ IshKebab บางส่วน — บล็อกมี short-sentence pattern + "The..." titles ที่ LLM ชอบใช้ แต่ข้อมูลในนั้น verify ได้ทั้งหมด (links ไป GitHub issues + CVE + benchmarks)

แล้วใครควรย้าย ใครควรอยู่?​

ผมสรุปเป็น decision matrix:

คุณคือคำแนะนำ
Beginner / Windows / "แค่อยากลอง"Ollama ยัง ok — หรือใช้ LM Studio ดีกว่า
Developer ที่อยาก optimize performancellama.cpp ตรง — 1.8x เร็วกว่า
คนที่ต้อง run โมเดลใหม่ทันทีllama.cpp หรือ LM Studio (ไม่ต้องรอ registry)
Concern เรื่อง FOSS ethicsllama.cpp, Jan, koboldcpp, ramalama
ใช้สำหรับ coding agent / productionllama.cpp หรือ vLLM (control + performance)
แค่อยาก share โมเดลกับเพื่อนllamafile (single executable)

สรุปส่วนตัว​

ผมว่าบล็อกของ Zetaphor มีข้อมูลที่ verify ได้ แต่ tone ของบล็อกค่อนข้างแรง — ผมเข้าใจว่าทำไม เพราะ Georgi Gerganov + contributors ทำงานหนักมาก แล้ว Ollama ก็ใช้ชื่อเสียงนั้นโดยไม่ให้เครดิต

แต่ผมก็ไม่ได้บอกว่าทุกคนต้องย้ายทันที — Ollama ยังมี use case สำหรับ beginner + quick test

ข้อสรุปส่วนตัว: ถ้าคุณใช้ Ollama อยู่และมันทำงานได้ — ก็ไม่ต้องรีบย้าย แต่ถ้าเริ่มเจอปัญหาเรื่อง quant limitation / template breakage / โมเดลใหม่ไม่มีใน registry → ลอง llama.cpp + LM Studio ดูครับ ไม่ต้อง commit

ข้อสรุปเชิงอุตสาหกรรม: ถ้าคุณเป็น dev ที่ contribute OSS — อ่านบทความนี้เป็น case study เรื่อง how not to wrap FOSS — minimal attribution + Modelfile lock-in + bad fork มันสร้าง technical debt ทั้งต่อตัวเองและต่อ ecosystem

Cross-references​

ถ้าสนใจ local LLM infrastructure ลองอ่าน posts อื่นๆ ใน series นี้:

อ้างอิง​

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...