หยุดใช้ Ollama — เมื่อ local LLM wrapper ที่ใหญ่ที่สุดมีปัญหาเรื่อง attribution, performance, และ vendor lock-in
สารบัญ
- TL;DR
- บล็อกต้นเรื่อง
- 10 ข้อกล่าวหาหลักของ Zetaphor
- 1. License violation — 400+ วันไม่มี MIT notice
- 2. Minimal attribution — README ไม่พูดถึง llama.cpp เป็นเวลา 1+ ปี
- 3. The bad fork — mid-2025 ย้ายจาก llama.cpp ไปใช้ ggml ตรงๆ
- 4. Performance gap — llama.cpp เร็วกว่า 1.8x
- 5. Misleading model naming —
deepseek-r1ไม่ใช่ R1 จริง - 6. Modelfile = reinventing a solved problem
- 7. Closed-source GUI app — July 2025
- 8. Cloud pivot + CVE-2025-51471
- 9. Registry bottleneck — quant + packaging
- 10. VC pattern (5 steps)
- Honest disagreement: Ollama ยังมี use case ที่ถูกต้อง
- 1. "แค่อยากลองโมเดลใหม่เร็วๆ"
- 2. Windows user ที่ไม่ใช่ dev
- 3. Beginner ที่อยากเข้าโลก local LLM
- 4. ทางเลือกที่ "พอดี" จริงๆ
- Alternatives ที่ควรพิจารณา
- 1.
llama.cppตรง — เร็วสุด + ควบคุมได้สุด - 2. LM Studio — GUI wrapper ที่ดีกว่า
- 3. llamafile — Mozilla, single executable
- 4. llama-swap + LiteLLM — multi-model orchestration
- 5. Open-source GUI alternatives
- 6. ถ้าจะใช้ llama.cpp แต่ชอบ GUI
- HN discussion: slop critique
- แล้วใครควรย้าย ใครควรอยู่?
- สรุปส่วนตัว
- Cross-references
- อ้างอิง
บันทึก — 25 สิงหาคม 2569 — เช้ามืด เจอกระทู้ที่ถกกันหนักมาก
มีคนโพสต์บล็อกชื่อ "Stop using Ollama" ใน r/LocalLLaMA ขึ้นไป Hacker News ได้ 648 points + 207 comments — ซึ่งสำหรับ HN ถือว่าเยอะมาก
บล็อกนี้ไม่ได้บอกว่า Ollama แย่ — มันบอกว่า Ollama เป็น wrapper ที่สร้างบน llama.cpp แต่ใช้เวลาหลายปีหลบเลี่ยง attribution + เปิดตัว closed-source components + pivot ไป cloud + ทำผลงานช้ากว่า upstream 1.8 เท่า และมี alternatives ที่ดีกว่า
ผมอ่านบล็อก + Reddit + HN comments แล้วเขียนสรุปนี้ — พร้อม honest disagreement ว่า Ollama ก็ยังมี use case ที่ถูกต้องสำหรับบางคน
TL;DR
- Zetaphor (sleepingrobots.com) ทำ timeline ของ Ollama's actions ที่ "bad-faith open source citizen" — 400+ วันที่ binary ไม่มี MIT notice ของ llama.cpp, README ไม่มี attribution เป็นปี
- Performance: llama.cpp เร็วกว่า Ollama 1.8x (161 vs 89 tok/s บน hardware เดียวกัน), CPU gap 30-50%, Qwen-3 Coder 32B gap ~70%
- Ollama "fork ไปทำเอง" บน ggml ในปี 2025 → reintroduce bugs ที่ llama.cpp แก้ไปแล้ว + ทำ GPT-OSS 20B พัง
- Alternatives ที่ได้ทั้ง UX + ไม่ vendor-lock-in: llama.cpp ตรง, LM Studio, llamafile, koboldcpp, Jan, ramalama
- Honest disagreement: Ollama ยังเหมาะกับ beginner + quick test + "แค่อยากลองโมเดล" — ไม่ใช่ทุกคนต้องย้าย
- HN comment ที่น่าสนใจ: บล็อกถูก criticize ว่า "เขียนเหมือน LLM" (slop) — Zetaphor ตอบว่า "I guess I write like an LLM :P"
บล็อกต้นเรื่อง
Zetaphor จาก sleepingrobots.com (เขียนบล็อกเรื่อง LLM บน Strix Halo) โพสต์บล็อกชื่อ "Stop using Ollama" — เขาเขียนว่า:
"Ollama is the most popular way to run local LLMs. It shouldn't be. It gained that position by being first, the first tool that made
llama.cppaccessible to people who didn't want to compile C++ or write their own server configs. That was a real contribution, briefly.""But the project has since spent years systematically obscuring where its actual technology comes from, misleading users about what they're running, and drifting from the local-first mission that earned it trust in the first place. All while taking venture capital money."
เขาไม่ได้บอกว่า Ollama ไม่มีประโยชน์ — เขาบอกว่า ตัวเลือกที่ดีกว่ามีอยู่แล้ว และ Ollama เลือกที่จะไม่ evolve ไปในทางที่ดี
10 ข้อกล่าวหาหลักของ Zetaphor
1. License violation — 400+ วันไม่มี MIT notice
Ollama binary distributions ไม่ได้ใส่ MIT license notice ของ llama.cpp ตามที่ MIT license บังคับ (ข้อเดียวที่สำคัญที่สุดของ MIT)
- GitHub issue #3185 เปิดใน early 2024 ขอให้แก้ license compliance
- ไม่มีคนตอบจาก maintainers เป็นเวลา 400+ วัน
- จนกระทั่ง April 2024 มี issue #3697 เปิดใหม่ + PR #3700 ตามมาในไม่กี่ชั่วโมง
- Michael Chiang (co-founder) เพิ่มแค่บรรทัดเดียวที่ล่างสุดของ README: "llama.cpp project founded by Georgi Gerganov"
2. Minimal attribution — README ไม่พูดถึง llama.cpp เป็นเวลา 1+ ปี
ไม่ใช่แค่ไม่มี license notice — แต่ README + website + marketing materials ไม่พูดถึง llama.cpp เลย ทั้งที่ Ollama ทุกอย่างวิ่งอยู่บน llama.cpp
คนใน HN ตอบ:
"I'm continually puzzled by their approach, it's such self-inflicted negative PR. Building on llama is perfectly valid and they're adding value on ease of use here. Just give the llama team proper credit."
3. The bad fork — mid-2025 ย้ายจาก llama.cpp ไปใช้ ggml ตรงๆ
Ollama follow through ตามที่บอก — ย้ายจาก llama.cpp inference backend → เขียน custom implementation บน ggml (tensor library ที่ llama.cpp ใช้ข้างใต้)
เหตุผลที่บอก: "stability — llama.cpp moves fast and breaks things"
ผลลัพธ์:
- reintroduce bugs ที่ llama.cpp แก้ไปนานแล้ว
- structured output พัง, vision models fail, GGML assertion crashes
- GPT-OSS 20B fail ใน Ollama แต่ work ใน upstream llama.cpp เพราะ Ollama ไม่รองรับ tensor types ที่โมเดลต้องการ
- Georgi Gerganov (ผู้สร้าง llama.cpp) tweet ว่า Ollama "forked and made bad changes to GGML"
4. Performance gap — llama.cpp เร็วกว่า 1.8x
| Source | Setup | llama.cpp | Ollama | Gap |
|---|---|---|---|---|
| arsturn.com | Same hardware, same model | 161 tok/s | 89 tok/s | 1.8x |
| deploybase.ai | CPU inference | — | — | 30-50% |
| Reddit r/LocalLLaMA | Qwen-3 Coder 32B | — | — | ~70% throughput |
Root cause ของ gap:
- Daemon layer overhead ของ Ollama
- Poor GPU offloading heuristics
- Vendored backend ที่ตามหลัง upstream
5. Misleading model naming — deepseek-r1 ไม่ใช่ R1 จริง
ตอน DeepSeek ปล่อย R1 family (January 2025) — Ollama list "DeepSeek-R1-Distill-Qwen-32B" (8B Qwen-derived distill) ใน registry ว่า "DeepSeek-R1" โดย strip "Distill" prefix ออก
ผลกระทบ:
- คนรัน
ollama run deepseek-r1คิดว่าได้ 671B R1 จริง → ได้ 8B distillate - Performance ไม่ตรงกับความคาดหวัง
- Reputational damage กับ DeepSeek
GitHub issues #8557 + #8698 ขอให้แยก — ปิดเป็น duplicate โดยไม่แก้
6. Modelfile = reinventing a solved problem
GGUF (โดย Georgi Gerganov) ออกแบบมาให้ single-file deployment — chat template, stop tokens, metadata ฝังอยู่ในไฟล์เดียว
Ollama เพิ่ม Modelfile ขึ้นมา (Dockerfile-inspired):
- ต้อง extract chat template จาก GGUF → translate เป็น Go template syntax (ซึ่งต่างจาก Jinja)
- Ollama auto-detect chat template แค่จาก hardcoded list — ถ้า GGUF มี Jinja template ที่ Ollama ไม่รู้จัก → fallback เป็น bare
{{ .Prompt }}→ instruction format พังเงียบๆ - เปลี่ยน parameter เช่น temperature: ต้อง export modelfile → edit →
ollama createใหม่ → copy model ทั้งก้อน 30-60 GB เพื่อเปลี่ยน 1 parameter
User บ่น: "The 'modelfile' workflow is a pain in the booty... 30 to 60GB and copying the entire thing to change one parameter is just dumb."
7. Closed-source GUI app — July 2025
Ollama ship GUI desktop app (macOS + Windows):
- พัฒนาใน private repo (
github.com/ollama/app) - ไม่มี license
- Source code ไม่ public
- มี potential AGPL-3.0 dependencies ใน binary
- Download button อยู่ข้างๆ GitHub link → คนเข้าใจผิดว่าโหลด open-source tool
PR #12933 merge เข้า main repo November 2025 — แต่ initial rollout แสดงให้เห็นว่า "project's instincts lie"
8. Cloud pivot + CVE-2025-51471
ปลายปี 2025: Ollama เริ่ม route prompts ไป third-party cloud providers (proprietary models เช่น MiniMax-m2.7)
- Documentation: "we process your prompts... do not store or log" — แต่ไม่พูดถึง สิ่งที่ third-party provider ทำ
- Alibaba Cloud models → ไม่มี zero-data-retention guarantee
CVE-2025-51471: Token exfiltration — malicious registry server หลอก Ollama ส่ง auth token ไป attacker endpoint ตอน pull model
9. Registry bottleneck — quant + packaging
- รองรับ quant แค่ Q4_K_M, Q8_0, F16, F32 (ไม่มี Q5, Q6, IQ quants ที่ llama.cpp มีมานานแล้ว)
- Model ใหม่ออก → ต้องรอคน Ollama package ใน registry → "โมเดลไม่มีใน Ollama" จนกว่าจะ publish
- Reddit PSA: "If you want to test new models, use llama.cpp/transformers/vLLM/SGLang" — Qwen models พัง tool calls + garbage responses เฉพาะใน Ollama
10. VC pattern (5 steps)
Zetaphor สรุปเป็น playbook ที่คุ้นตา:
- Launch on open source — build บน llama.cpp, gain community trust
- Minimize attribution — ทำให้ product ดู self-sufficient ต่อ investors
- Create lock-in — proprietary registry format, hashed filenames
- Launch closed-source components — GUI app
- Add cloud services — monetization vector
Ollama = YC W21, founded โดยคนที่เคยทำ Kitematic (Docker GUI acquired by Docker Inc.)
Honest disagreement: Ollama ยังมี use case ที่ถูกต้อง
แม้ข้อกล่าวหาจะหนัก แต่ผมเห็น use case ที่ Ollama ยังเหมาะสม:
1. "แค่อยากลองโมเดลใหม่เร็วๆ"
ollama run qwen2.5:7b
ไม่มีอะไรตรงๆ ง่ายขนาดนี้ — llama.cpp -hf unsloth/Qwen3.5-35B-A3B-GGUF:Q4_K_M ก็ทำได้ แต่คนทั่วไปต้องติดตั้ง build ก่อน
2. Windows user ที่ไม่ใช่ dev
# Ollama
winget install Ollama.Ollama
ollama run llama3.2
# llama.cpp บน Windows ต้อง build หรือใช้ wsl
inagy ใน r/LocalLLaMA: "Ollama gives an easy to setup server, especially on Windows. Couple clicks in a setup wizard, then it's running in the background as a service."
3. Beginner ที่อยากเข้าโลก local LLM
gnooggi: "Easy for lazy people + beginners who don't know what they're doing, because it simply works"
Forsaken_Ad_774: "I'm not chasing down the maximum optimizations, I'm just using something that just works with minimum amount of effort. Coupled with Cline/openwebui I'm golden."
4. ทางเลือกที่ "พอดี" จริงๆ
Youth18: "LM Studio is still good for testing models — quick downloads, easy to see popular LLMs. But I use that for testing and then if I like a model I integrate it into my llama.cpp + custom model loading manager stack."
Insight: Ollama ไม่ได้แย่ — แค่ "เป็น stepping stone" ที่ดี แต่ไม่ใช่ปลายทาง
Alternatives ที่ควรพิจารณา
Zetaphor แนะนำ alternatives ที่หลากหลาย ตาม use case:
1. llama.cpp ตรง — เร็วสุด + ควบคุมได้สุด
# Install
brew install llama.cpp # macOS
apt install llama.cpp # Linux
# Run from HuggingFace directly
llama-server -hf unsloth/Qwen3.5-35B-A3B-GGUF:Q4_K_M --port 8000
# Open http://localhost:8000 — web UI ในตัว
ข้อดี:
- เร็วที่สุด (1.8x Ollama)
- GGUF ตรงจาก HuggingFace — ไม่ต้องรอ registry
- OpenAI-compatible API (
/v1/chat/completions) - ทุก quant (Q5, Q6, IQ, etc.)
- Embed params ใน GGUF ไม่ต้อง Modelfile
ข้อเสีย:
- ต้อง install เอง
- ชื่อ
.cppทำให้คนคิดว่าเป็น library (HN comment: "for as long as its name is '.cpp', people are going to think it's a C++ library and avoid it") - โมเดลใหม่ต้อง update llama.cpp ก่อน (gbitten: "move fast and break things")
2. LM Studio — GUI wrapper ที่ดีกว่า
ข้อดี:
- Same one-click convenience เหมือน Ollama
- Accept any GGUF
- Expose all knobs
- มี acknowledgements page ที่ credit llama.cpp ตรงๆ
- Zetaphor เองแนะนำ: "If you don't want to think about it, LM Studio is probably the best choice"
ข้อเสีย:
- Proprietary (ไม่ใช่ open-source)
- แต่ Zetaphor ชี้ว่า "It's a closed-source product, but it's not a parasitic one"
3. llamafile — Mozilla, single executable
brabel (เคยช่วย llama.cpp): "It's truly open source, backed by Mozilla, openly uses llama.cpp"
# Download single binary
wget https://example.com/llamafile-with-model
# Run
./llamafile-with-model
ข้อดี:
- Single executable รันได้ 6 OSes
- ไม่ต้อง install
- โดย Justine Tunkey (CosmopolitanC fame) — wizard-level engineer
4. llama-swap + LiteLLM — multi-model orchestration
# llama-swap handles hot-swapping models
# LiteLLM routes across multiple backends
5. Open-source GUI alternatives
- Jan (jan.ai) — AGPLv3, local-first chat
- koboldcpp — AGPL, llama.cpp fork + extensive config (HN: "minimal dependencies, no installers, no logs, don't do anything to user's system they didn't ask")
- ramalama (Red Hat) — container-native + explicit credits
6. ถ้าจะใช้ llama.cpp แต่ชอบ GUI
LlamaBarn (macOS) + llama-server (Linux/Windows) — official GUI จาก ggml-org เอง
HN discussion: slop critique
มี comment ที่ถูก flag ว่า "LLM-written":
IshKebab ชี้:
"Short sentences like 'It shouldn't be.', 'I've moved on.', 'Ollama didn't.', etc. Not-this-but-that like 'The local LLM ecosystem doesn't need Ollama. It needs llama.cpp.' Weird signposting: 'Benchmarks tell the story.' Here's-the-rub conclusion: 'The Bigger Picture' Starting every title with 'The...'
It's definitely largely human-written, but there are enough slop-isms to make it annoying to read."
Zetaphor ตอบกลับ:
"I guess I write like an LLM :P Probably a side effect of using them so much"
ผมอ่านแล้วเห็นด้วยกับ IshKebab บางส่วน — บล็อกมี short-sentence pattern + "The..." titles ที่ LLM ชอบใช้ แต่ข้อมูลในนั้น verify ได้ทั้งหมด (links ไป GitHub issues + CVE + benchmarks)
แล้วใครควรย้าย ใครควรอยู่?
ผมสรุปเป็น decision matrix:
| คุณคือ | คำแนะนำ |
|---|---|
| Beginner / Windows / "แค่อยากลอง" | Ollama ยัง ok — หรือใช้ LM Studio ดีกว่า |
| Developer ที่อยาก optimize performance | llama.cpp ตรง — 1.8x เร็วกว่า |
| คนที่ต้อง run โมเดลใหม่ทันที | llama.cpp หรือ LM Studio (ไม่ต้องรอ registry) |
| Concern เรื่อง FOSS ethics | llama.cpp, Jan, koboldcpp, ramalama |
| ใช้สำหรับ coding agent / production | llama.cpp หรือ vLLM (control + performance) |
| แค่อยาก share โมเดลกับเพื่อน | llamafile (single executable) |
สรุปส่วนตัว
ผมว่าบล็อกของ Zetaphor มีข้อมูลที่ verify ได้ แต่ tone ของบล็อกค่อนข้างแรง — ผมเข้าใจว่าทำไม เพราะ Georgi Gerganov + contributors ทำงานหนักมาก แล้ว Ollama ก็ใช้ชื่อเสียงนั้นโดยไม่ให้เครดิต
แต่ผมก็ไม่ได้บอกว่าทุกคนต้องย้ายทันที — Ollama ยังมี use case สำหรับ beginner + quick test
ข้อสรุปส่วนตัว: ถ้าคุณใช้ Ollama อยู่และมันทำงานได้ — ก็ไม่ต้องรีบย้าย แต่ถ้าเริ่มเจอปัญหาเรื่อง quant limitation / template breakage / โมเดลใหม่ไม่มีใน registry → ลอง llama.cpp + LM Studio ดูครับ ไม่ต้อง commit
ข้อสรุปเชิงอุตสาหกรรม: ถ้าคุณเป็น dev ที่ contribute OSS — อ่านบทความนี้เป็น case study เรื่อง how not to wrap FOSS — minimal attribution + Modelfile lock-in + bad fork มันสร้าง technical debt ทั้งต่อตัวเองและต่อ ecosystem
Cross-references
ถ้าสนใจ local LLM infrastructure ลองอ่าน posts อื่นๆ ใน series นี้:
- $10K budget — ซื้อ 2× DGX Spark เลย หรือรอ? — hardware perspective
- ซื้อ DGX Spark เครื่องเดียว เพื่อเรียนรู้ + Hermes agent คุ้มไหม — awkward spot analysis
อ้างอิง
- Stop using Ollama — Sleeping Robots (Zetaphor) — บล็อกต้นเรื่อง
- Stop using Ollama — Hacker News — HN discussion 648 points / 207 comments
- Stop using Ollama — r/LocalLLaMA — Reddit post
- GitHub issue #3185 — Ollama MIT license violation — 400+ วันไม่ตอบ
- GitHub issue #3697 — Ollama README attribution — April 2024
- CVE-2025-51471 — Ollama token exfiltration — NVD
- Comparing Ollama and llama.cpp — Arsturn — 1.8x performance gap
- llama.cpp vs Ollama CPU benchmarks — DeployBase — 30-50% CPU gap
- llama.cpp — GitHub — engine ต้นทาง 100K+ stars, 450+ contributors
- LM Studio acknowledgements page — example ของ good-faith wrapper
- llamafile — GitHub — Mozilla single-executable
- llama-swap — GitHub — multi-model orchestration
- Jan — AGPLv3 desktop chat — open-source GUI alternative
- koboldcpp — GitHub — llama.cpp fork + web UI
- ramalama — Red Hat — container-native + explicit credits
- ggml-org joins Hugging Face — February 2026, long-term sustainability
- PSA: If you want to test new models, use llama.cpp — r/LocalLLaMA — community pattern
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee