LiteLLM journey: vLLM integration — drop_params + reasoning + DGX Spark pitfalls
สารบัญ
- TL;DR
- 1. Quick Start — 3 fields ที่ต้องตั้ง
- 2. drop_params — 3 ระดับ
- ระดับ 1 — Global (
config.yaml) - ระดับ 2 — Per-model (
additional_drop_params) - ระดับ 3 — Per-call (SDK only)
- Decision: เลือกอันไหน?
- 3. Reasoning outputs — silent migration bug
- Pitfall (จริงที่ผมเจอ)
- vLLM server config ที่ต้อง enable
- 4. max_model_len — context window mismatch
- วิธีตั้งให้ตรง
- 5. DGX Spark — 1 model at a time
- DGX pattern (ของผม)
- 6. Error mapping pitfall
- Workaround
- สรุป
- อ้างอิง
"ผมเชื่อว่า vLLM integration คือเรื่องง่าย — จนกว่าจะเจอ 3 pitfalls:
reasoning_content→reasoning, ContextWindowExceededError ไม่บอกชัด, และ drop_params มี 3 ระดับ"
Part 1 — drop_params ทำไมต้อง recreate
Part 2 — 3 flags ที่ใช้บ่อย
Part 3 — DB mode + 3-way sync
Part 4 — vLLM integration + pitfalls
TL;DR
vLLM เป็น OpenAI-compatible server → LiteLLM ต่อง่ายมาก แค่ hosted_vllm/ prefix
model_list:
- model_name: qwen3.8-flash-next
litellm_params:
model: hosted_vllm/qwen3.8-flash-next
api_base: http://10.0.0.246:8000
max_input_tokens: 262144
4 pitfalls ที่ผมเจอเอง:
drop_paramsมี 3 ระดับ — global, per-call, per-model (ใช้คนละแบบต่างกัน)reasoning_content→reasoning— vLLM เปลี่ยนชื่อ field, client code ที่ไม่ migrate → silent bug อ่าน empty- ContextWindowExceededError wrapped in BadRequestError — error mapping ไม่ชัด
- DGX Spark: max-num-seqs ต่ำ (1-4) — memory pool shared, ไม่ใช่ 32 หรือ 64
1. Quick Start — 3 fields ที่ต้องตั้ง
จาก LiteLLM vLLM docs:
# config.yaml
model_list:
- model_name: qwen3.8-flash-next # virtual name (ที่ลูกค้าเรียก)
litellm_params:
model: hosted_vllm/qwen3.8-flash-next # hosted_vllm/ prefix
api_base: http://10.0.0.246:8000
api_key: fake-key # vLLM self-host usually ไม่ต้อง check
| Field | จำเป็น? | ตัวอย่าง |
|---|---|---|
model | ✅ | hosted_vllm/<vllm-model-name> |
api_base | ✅ | http://10.0.0.246:8000 (vLLM server) |
api_key | optional | fake-key (vLLM ส่วนใหญ่ไม่ check) |
max_input_tokens | optional | 262144 (256K context) |
max_output_tokens | optional | 32768 |
Endpoints ที่ vLLM รองรับ (OpenAI-compatible):
/chat/completions✅ (หลัก)/embeddings✅/completions✅/rerank✅/audio/transcriptions✅
2. drop_params — 3 ระดับ
# Default behavior
response = litellm.completion(
model="command-r",
messages=[...],
response_format={"key": "value"} # ไม่รองรับ → exception
)
drop_params=True ทำให้ LiteLLM drop params ที่ model ไม่รองรับแทนที่จะ error
ระดับ 1 — Global (config.yaml)
litellm_settings:
drop_params: true # ทุก model ใน proxy
✅ ผมใช้อันนี้ใน Part 1 — ครอบคลุมทุก model
ระดับ 2 — Per-model (additional_drop_params)
model_list:
- model_name: qwen-mini
litellm_params:
model: hosted_vllm/qwen-mini
api_base: http://...
additional_drop_params: ["response_format", "logit_bias"]
ใช้เฉพาะ model ที่รู้ว่า "params นี้ไม่รองรับ"
ระดับ 3 — Per-call (SDK only)
import litellm
response = litellm.completion(
model="command-r",
messages=[...],
response_format={"key": "value"},
drop_params=True # 👈 ครั้งนี้เท่านั้น
)
ใช้เมื่อ test ใน script ไม่ต้องแก้ config
Decision: เลือกอันไหน?
| ใช้เมื่อ | ระดับ |
|---|---|
| Default behavior ของทั้ง proxy | Global |
| Model นี้ไม่รองรับ params บางตัว | Per-model |
| Test ใน script / Jupyter | Per-call |
ผมใช้ Global เป็นหลัก + Per-model สำหรับ qwen-mini ที่ drop logit_bias
3. Reasoning outputs — silent migration bug
จาก vLLM Reasoning Outputs docs:
"Warning:
reasoningused to be calledreasoning_content. To migrate, directly replacereasoning_contentwithreasoning."
→ ถ้า client code ของคุณใช้ reasoning_content (เวอร์ชันเก่า) → silent read empty แม้ model generate reasoning
Pitfall (จริงที่ผมเจอ)
ผมเขียน streaming client:
# ❌ OLD — silent bug
chunk.choices[0].delta.reasoning_content # None เสมอ
# ✅ NEW — ใช้ชื่อใหม่
chunk.choices[0].delta.reasoning # มี reasoning steps
→ ถ้าไม่อ่าน docs จะงง ว่าทำไม reasoning หายไปเงียบๆ
vLLM server config ที่ต้อง enable
vllm serve qwen3.8-flash-next --reasoning-parser deepseek_r1
--reasoning-parser ต้องระบุ parser (deepseek_r1, gemma4, etc.) — ไม่ใช่ default
4. max_model_len — context window mismatch
ทั่วไปที่เจอ:
- vLLM
--max-model-len= 256000 (DGX Spark) - LiteLLM
max_input_tokens= 262144 - ถ้า LiteLLM ไม่รู้ → request ส่งไป vLLM เกิน → ContextWindowExceededError
วิธีตั้งให้ตรง
model_list:
- model_name: qwen3.8-flash-next
litellm_params:
model: hosted_vllm/qwen3.8-flash-next
api_base: http://10.0.0.246:8000
max_input_tokens: 262144 # ≤ vLLM max_model_len
max_output_tokens: 32768
หรือ drop แทน error:
litellm_settings:
drop_params: true
→ LiteLLM จะ drop messages เกิน window แทนที่จะ error (ถ้า model รองรับ)
5. DGX Spark — 1 model at a time
จาก vllm.ai/blog/2026-06-01-vllm-dgx-spark + DGX Spark playbook:
"Keep
--max-num-seqslow (1-4) — pool is shared"
DGX Spark (GB10, 128 GB unified memory) ใช้ unified CPU+GPU memory — ไม่ใช่ GPU เต็มที่
vllm serve qwen3.8-flash-next \
--max-num-seqs 2 \
--max-model-len 256000 \
--reasoning-parser deepseek_r1
DGX pattern (ของผม)
- 1 model ต่อ container (ไม่ใช่หลาย model พร้อมกัน)
- ทุกครั้งที่เปลี่ยน model → recreate vllm container
- LiteLLM alias ใน DB → point ไป vLLM ใหม่
- 60s auto-reload (จาก Part 3) → ไม่ต้อง restart LiteLLM
6. Error mapping pitfall
จาก Issue #7259:
"ContextWindowExceededError is not mapped correctly via proxy — gets wrapped in BadRequestError"
→ ตอน client รับ error จะเห็น BadRequestError ทั่วไป ไม่ใช่ ContextWindowExceededError เฉพาะ
Workaround
from openai import BadRequestError
try:
response = client.chat.completions.create(...)
except BadRequestError as e:
if "context_length" in str(e).lower():
# handle context window exceeded
...
หรือ enable provider_specific_fields ใน exception mapping config
สรุป
| Step | Config |
|---|---|
| 1. Integrate | hosted_vllm/<model> + api_base |
| 2. Drop params | litellm_settings.drop_params: true (global) |
| 3. Reasoning | --reasoning-parser <name> บน vLLM server |
| 4. Streaming | ใช้ reasoning ไม่ใช่ reasoning_content |
| 5. Context | ตั้ง max_input_tokens ≤ vLLM max_model_len |
| 6. DGX Spark | --max-num-seqs 1-4 (memory pool shared) |
Lessons จากการ integrate จริง:
drop_paramsมี 3 ระดับ — เลือกให้เหมาะreasoningไม่ใช่reasoning_content— silent migration- ContextWindowExceededError ถูก wrap → handle BadRequestError
- DGX Spark max-num-seqs ต่ำ — memory unified ไม่ใช่ GPU แยก
Part 5 → Self-host economics — $0 vs $128/เดือน
อ้างอิง
- LiteLLM vLLM provider —
hosted_vllm/prefix, OpenAI-compatible - LiteLLM drop_params — 3 ระดับ + additional_drop_params
- vLLM Reasoning Outputs —
reasoningfield migration - vLLM on DGX Spark — GB10 unified memory
- DGX Spark + vLLM Playbook —
--max-num-seqs 1-4 - LiteLLM Issue #7259 — ContextWindowExceededError wrapped
- vLLM LiteLLM docs — official integration
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee