Skip to main content

LiteLLM journey: vLLM integration — drop_params + reasoning + DGX Spark pitfalls

· 5 min read

"ผมเชื่อว่า vLLM integration คือเรื่องง่าย — จนกว่าจะเจอ 3 pitfalls: reasoning_content → reasoning, ContextWindowExceededError ไม่บอกชัด, และ drop_params มี 3 ระดับ"

Part 1 — drop_params ทำไมต้อง recreate Part 2 — 3 flags ที่ใช้บ่อย Part 3 — DB mode + 3-way sync Part 4 — vLLM integration + pitfalls

TL;DR​

vLLM เป็น OpenAI-compatible server → LiteLLM ต่อง่ายมาก แค่ hosted_vllm/ prefix

model_list:
- model_name: qwen3.8-flash-next
litellm_params:
model: hosted_vllm/qwen3.8-flash-next
api_base: http://10.0.0.246:8000
max_input_tokens: 262144

4 pitfalls ที่ผมเจอเอง:

  1. drop_params มี 3 ระดับ — global, per-call, per-model (ใช้คนละแบบต่างกัน)
  2. reasoning_content → reasoning — vLLM เปลี่ยนชื่อ field, client code ที่ไม่ migrate → silent bug อ่าน empty
  3. ContextWindowExceededError wrapped in BadRequestError — error mapping ไม่ชัด
  4. DGX Spark: max-num-seqs ต่ำ (1-4) — memory pool shared, ไม่ใช่ 32 หรือ 64

1. Quick Start — 3 fields ที่ต้องตั้ง​

จาก LiteLLM vLLM docs:

# config.yaml
model_list:
- model_name: qwen3.8-flash-next # virtual name (ที่ลูกค้าเรียก)
litellm_params:
model: hosted_vllm/qwen3.8-flash-next # hosted_vllm/ prefix
api_base: http://10.0.0.246:8000
api_key: fake-key # vLLM self-host usually ไม่ต้อง check
Fieldจำเป็น?ตัวอย่าง
model✅hosted_vllm/<vllm-model-name>
api_base✅http://10.0.0.246:8000 (vLLM server)
api_keyoptionalfake-key (vLLM ส่วนใหญ่ไม่ check)
max_input_tokensoptional262144 (256K context)
max_output_tokensoptional32768

Endpoints ที่ vLLM รองรับ (OpenAI-compatible):

  • /chat/completions ✅ (หลัก)
  • /embeddings ✅
  • /completions ✅
  • /rerank ✅
  • /audio/transcriptions ✅

2. drop_params — 3 ระดับ​

จาก LiteLLM drop_params docs:

# Default behavior
response = litellm.completion(
model="command-r",
messages=[...],
response_format={"key": "value"} # ไม่รองรับ → exception
)

drop_params=True ทำให้ LiteLLM drop params ที่ model ไม่รองรับแทนที่จะ error

ระดับ 1 — Global (config.yaml)​

litellm_settings:
drop_params: true # ทุก model ใน proxy

✅ ผมใช้อันนี้ใน Part 1 — ครอบคลุมทุก model

ระดับ 2 — Per-model (additional_drop_params)​

model_list:
- model_name: qwen-mini
litellm_params:
model: hosted_vllm/qwen-mini
api_base: http://...
additional_drop_params: ["response_format", "logit_bias"]

ใช้เฉพาะ model ที่รู้ว่า "params นี้ไม่รองรับ"

ระดับ 3 — Per-call (SDK only)​

import litellm

response = litellm.completion(
model="command-r",
messages=[...],
response_format={"key": "value"},
drop_params=True # 👈 ครั้งนี้เท่านั้น
)

ใช้เมื่อ test ใน script ไม่ต้องแก้ config

Decision: เลือกอันไหน?​

ใช้เมื่อระดับ
Default behavior ของทั้ง proxyGlobal
Model นี้ไม่รองรับ params บางตัวPer-model
Test ใน script / JupyterPer-call

ผมใช้ Global เป็นหลัก + Per-model สำหรับ qwen-mini ที่ drop logit_bias

3. Reasoning outputs — silent migration bug​

จาก vLLM Reasoning Outputs docs:

"Warning: reasoning used to be called reasoning_content. To migrate, directly replace reasoning_content with reasoning."

→ ถ้า client code ของคุณใช้ reasoning_content (เวอร์ชันเก่า) → silent read empty แม้ model generate reasoning

Pitfall (จริงที่ผมเจอ)​

ผมเขียน streaming client:

# ❌ OLD — silent bug
chunk.choices[0].delta.reasoning_content # None เสมอ

# ✅ NEW — ใช้ชื่อใหม่
chunk.choices[0].delta.reasoning # มี reasoning steps

→ ถ้าไม่อ่าน docs จะงง ว่าทำไม reasoning หายไปเงียบๆ

vLLM server config ที่ต้อง enable​

vllm serve qwen3.8-flash-next --reasoning-parser deepseek_r1

--reasoning-parser ต้องระบุ parser (deepseek_r1, gemma4, etc.) — ไม่ใช่ default

4. max_model_len — context window mismatch​

ทั่วไปที่เจอ:

  • vLLM --max-model-len = 256000 (DGX Spark)
  • LiteLLM max_input_tokens = 262144
  • ถ้า LiteLLM ไม่รู้ → request ส่งไป vLLM เกิน → ContextWindowExceededError

วิธีตั้งให้ตรง​

model_list:
- model_name: qwen3.8-flash-next
litellm_params:
model: hosted_vllm/qwen3.8-flash-next
api_base: http://10.0.0.246:8000
max_input_tokens: 262144 # ≤ vLLM max_model_len
max_output_tokens: 32768

หรือ drop แทน error:

litellm_settings:
drop_params: true

→ LiteLLM จะ drop messages เกิน window แทนที่จะ error (ถ้า model รองรับ)

5. DGX Spark — 1 model at a time​

จาก vllm.ai/blog/2026-06-01-vllm-dgx-spark + DGX Spark playbook:

"Keep --max-num-seqs low (1-4) — pool is shared"

DGX Spark (GB10, 128 GB unified memory) ใช้ unified CPU+GPU memory — ไม่ใช่ GPU เต็มที่

vllm serve qwen3.8-flash-next \
--max-num-seqs 2 \
--max-model-len 256000 \
--reasoning-parser deepseek_r1

DGX pattern (ของผม)​

  • 1 model ต่อ container (ไม่ใช่หลาย model พร้อมกัน)
  • ทุกครั้งที่เปลี่ยน model → recreate vllm container
  • LiteLLM alias ใน DB → point ไป vLLM ใหม่
  • 60s auto-reload (จาก Part 3) → ไม่ต้อง restart LiteLLM

6. Error mapping pitfall​

จาก Issue #7259:

"ContextWindowExceededError is not mapped correctly via proxy — gets wrapped in BadRequestError"

→ ตอน client รับ error จะเห็น BadRequestError ทั่วไป ไม่ใช่ ContextWindowExceededError เฉพาะ

Workaround​

from openai import BadRequestError

try:
response = client.chat.completions.create(...)
except BadRequestError as e:
if "context_length" in str(e).lower():
# handle context window exceeded
...

หรือ enable provider_specific_fields ใน exception mapping config

สรุป​

StepConfig
1. Integratehosted_vllm/<model> + api_base
2. Drop paramslitellm_settings.drop_params: true (global)
3. Reasoning--reasoning-parser <name> บน vLLM server
4. Streamingใช้ reasoning ไม่ใช่ reasoning_content
5. Contextตั้ง max_input_tokens ≤ vLLM max_model_len
6. DGX Spark--max-num-seqs 1-4 (memory pool shared)

Lessons จากการ integrate จริง:

  1. drop_params มี 3 ระดับ — เลือกให้เหมาะ
  2. reasoning ไม่ใช่ reasoning_content — silent migration
  3. ContextWindowExceededError ถูก wrap → handle BadRequestError
  4. DGX Spark max-num-seqs ต่ำ — memory unified ไม่ใช่ GPU แยก

Part 5 → Self-host economics — $0 vs $128/เดือน

อ้างอิง​

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...