Skip to main content

LiteLLM journey: 3 flags ที่ช่วยเพิ่มประสิทธิภาพ — drop_params + cache + callbacks

· 5 min read

"LiteLLM config มี 100+ flags แต่ถ้าเลือกได้แค่ 3 ตัว ผมเลือกตัวนี้ — แต่ละตัวช่วยเพิ่มประสิทธิภาพคนละแบบ"

มีคนบอกผมว่า "ตั้ง flag เยอะๆ เข้าไว้" แล้วมันจะ work มันก็ work จริง — แต่ไม่ใช่ทุก flag ที่ apply ทันที

Part 1 ผมเล่าเรื่อง drop_params ที่ตั้งผ่าน DB อย่างเดียวไม่พอ ต้อง recreate Part 2 นี้คือภาพกว้าง — 3 flags ที่ผมตั้งก่อนเปิดใช้งานทุกครั้ง และทำไม flag บางตัวใช้ DB update พอ

TL;DR​

3 flags ที่ผมตั้งก่อนเปิด proxy ทุกครั้ง + ข้อดี:

  1. drop_params — drop unsupported params ที่ client ส่ง (เช่น reasoning_effort)
    • ข้อดี: client compatibility สูง — ส่ง params อะไรก็ไม่ break, provider ไหนไม่รองรับอะไรก็ตามได้
  2. cache + cache_params — Redis-backed caching + TTL
    • ข้อดี: ลด load ของ vLLM — same prompt cache ได้ใน 10 นาที, response เร็วขึ้น + cost ลด
  3. success_callback + failure_callback — rate limit + logging
    • ข้อดี: ป้องกัน abuse (one client hog) + observe (ดู spend + error)

2 ประเภทของ config:

  • Runtime behavior (drop_params) → recreate container ไม่งั้นไม่ apply
  • Runtime data (cache, cache_params, callbacks ทั่วไป) → DB update พอ

ต่อจาก Part 1 ผมจะใช้ decision matrix ที่ตารางด้านล่าง

1. drop_params — runtime behavior ที่ต้อง recreate​

ข้อดี: client ส่ง params อะไรก็ตาม (เช่น reasoning_effort ที่ MiniMax-M3 ไม่รองรับ) → LiteLLM drop เงียบๆ ไม่ error 400 → client ไม่ต้องเขียน fallback logic

อันนี้ผมเล่ารายละเอียดใน Part 1 ไปแล้ว — flag นี้ load ตอน container start, DB update ไม่ apply จนกว่าจะ recreate

ตอนนี้ผมตั้ง drop_params: true → LiteLLM drop reasoning_effort แบบ silent ทุก request ไม่ error 400 อีก

ถ้ายังไม่ได้อ่าน Part 1 → อ่านก่อนที่นี่ จะเข้าใจ context ตรงนี้ง่ายขึ้น

2. cache + Redis — runtime data ที่อัปเดตได้​

ข้อดี: ลด load ของ vLLM backend — same prompt ภายใน TTL ไม่ต้อง forward, response เร็วขึ้นมาก + cost ลด (ไม่คิด token จาก provider)

litellm_settings:
cache: true
cache_params:
type: redis
host: redis
port: 6379
password: [REDACTED]
ttl: 600
  • Redis container อยู่ใน compose (port 6379)
  • TTL 600s = 10 นาที
  • Same prompt + same params + within TTL → hit cache → ไม่ต้อง forward ไป vLLM

SpendLogs cache_hit จริงๆ ของผม​

ตอนนี้ในระบบมี SpendLogs ~17,000 rows ในช่วง 2 สัปดาห์:

SELECT cache_hit, count(*)
FROM "LiteLLM_SpendLogs"
GROUP BY cache_hit;
cache_hitCount%
true1020.6%
false9,37155.0%
null7,47444.4%

Cache effective hit rate = 102 / (102 + 9,371) = ~1.08%

ต่ำกว่าที่ typical — เหตุผลที่หลักๆ ของผม:

  1. Hermes client ส่ง temperature random → same prompt คำตอบไม่เหมือนเดิม (cache key ต่าง)
  2. Multi-provider routing → แต่ละ provider ได้ cache key ต่างกัน
  3. Long conversations → cache key เปลี่ยนทุก message

ทำไม null เยอะ​

null = record เก่าก่อนที่ field cache_hit ถูก populate (ก่อน 2026-08-30) true กับ false track ตั้งแต่ตอนนั้น → record ใหม่

ถ้าต้องการ hit rate สูงขึ้น — ต้อง tune cache_key ให้ดี (override per model) ตอนนี้ ~1% พอ ไม่ได้ปรับ

3. callbacks — user-set vs auto-injected​

ข้อดี: ป้องกัน client abuse (one client hog) + observe spend/log อย่างเป็นระบบ

litellm_settings:
success_callback: ["dynamic_rate_limiter_v3"]
failure_callback: ["database"]

success_callback: dynamic_rate_limiter_v3​

  • LiteLLM limit rate per user / per deployment
  • Use case: กัน client ตัวเดียว hog requests
  • ⚠️ critical callback — ถ้าเปลี่ยน → ต้อง recreate (DB update อย่างเดียวไม่พอ)

failure_callback: database​

  • Log failure → SpendLogs table
  • Use case: analytics, spend tracking
  • ✅ DB update พอ

Auto-injected callbacks (ที่ไม่เขียนใน config)​

GET /callbacks/list API คืน callbacks ทั้งหมด:

ProxyDBLogger (auto-injected)
HeadroomGuardrail (auto-injected)
ShadowEvalLogger (auto-injected)
cache (auto-injected)

แต่ UI + DB แสดงเฉพาะ user-set เท่านั้น:

SourceWhat you see
config.yamluser-set only
DB LiteLLM_Configuser-set only
UI admin paneluser-set only
/callbacks/list APIfull (auto + user)

ผมเคยเขียน failure_callback: ["database"] ซ้ำใน config.yaml → LiteLLM log 2 ครั้ง (manual + auto-injected) → รู้แล้วว่า ลบ line ใน config.yaml ใช้ DB UI จัดการอย่างเดียว

4. Decision matrix — เมื่อไหร่ recreate vs DB update​

FlagDB updateRecreateNote
drop_params❌✅runtime behavior (Part 1)
cache✅❌runtime data
cache_params.ttl✅❌runtime data
cache_params.host/port/password✅❌runtime data
failure_callback✅❌stable
success_callback: database✅❌stable
success_callback: dynamic_rate_limiter_v3❌✅critical callback
success_callback: prometheus✅❌non-critical

Rule of thumb: ถ้า flag เปลี่ยน behavior ของ LiteLLM runtime ทันที (เช่น rate limit, drop params) → มักเป็น runtime behavior → recreate ถ้า flag เป็นแค่ data (Redis host, log target) → runtime data → DB update พอ

5. Practical — ตั้งค่าที่ผมใช้จริง​

ตอน setup proxy ทุกครั้ง ผมจะตั้ง:

  1. drop_params: true → patch ทันที → recreate container
  2. cache: true + Redis → DB update
  3. success_callback: [dynamic_rate_limiter_v3] + failure_callback: [database] → DB update
  4. Verify ทุกอย่าง: /v1/models HTTP 200, /callbacks/list ครบ, SpendLogs ตอบ tracking

ปกติใช้เวลา ~1-2 นาที + recreate 30s

สรุป​

3 flags ที่ใช้บ่อย:

FlagApply via
drop_paramsrecreate container
cache + RedisDB update
callbacks (rate limit)recreate (critical)

ไม่ใช่ทุก flag ใน LiteLLM ที่ apply ตอน reload — เรื่อง runtime behavior vs runtime data เป็นของที่ต้องดูเป็นเรื่องๆ ไป

Part 3 → DB mode + 3-way sync (config.yaml ↔ DB ↔ UI)

อ้างอิง​

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...