vLLM Chat Completion Params ทั้ง 52 ตัว และวิธี Override ผ่าน LiteLLM
สารบัญ
- TL;DR
- จุดเริ่มต้น: อยากรู้ว่ามีอะไรส่งได้บ้าง
- ตรวจสอบโมเดลที่รันอยู่
- กลุ่มที่ 1: OpenAI Standard Params (ส่งตรงได้ใน LiteLLM)
- กลุ่มที่ 2: vLLM Extension Params (ต้องส่งผ่าน
extra_body) - Sampling & Decoding
- Reasoning & Thinking
- Chat Template & Tokenization
- Tool Calling
- Logging & Debugging
- Cache & Performance
- Advanced / Internal
- วิธี Override ผ่าน LiteLLM
- ใน LiteLLM Proxy config (
config.yaml) - ใน LiteLLM Python SDK
- ใน curl โดยตรง
- Params ที่น่าสนใจและไม่ค่อยมีคนพูดถึง
reasoning_effort— ควบคุมความลึกของ reasoningmin_p— กรอง token ที่ probability ต่ำเกินไปcache_salt— แยก prefix cache ต่อ sessiontruncate_prompt_tokens— ตัด prompt อัตโนมัติbad_words/allowed_token_ids— ควบคุม output แบบละเอียดseed— reproducibility- ทดสอบจริง: ส่ง params หลายตัวพร้อมกัน
- สรุปกฎการส่ง params ผ่าน LiteLLM
- Conclusion
- References
บันทึก 27 มิถุนายน 2569 — หลังจากใช้ vLLM มาสักพัก อยากรู้ว่าจริง ๆ แล้ว params ที่ส่งได้ใน
/v1/chat/completionsมีอะไรบ้าง และอันไหนส่งผ่าน LiteLLM ได้โดยตรง อันไหนต้องใช้extra_body
ผมใช้ vLLM เป็น inference backend มาสักพัก แต่ไม่เคยนั่งดูจริง ๆ ว่า /v1/chat/completions endpoint รองรับ parameters อะไรบ้างเต็ม ๆ ส่วนใหญ่ก็แค่ส่ง temperature, top_p, max_tokens ไปแล้วก็จบ — จนวันนี้ลองเปิด OpenAPI schema ดู ถึงรู้ว่ามี params ที่ใช้ไม่เคยรู้จักตั้งหลายตัว
TL;DR
- vLLM 0.23.1 รองรับ 52 parameters บน
/v1/chat/completions— มีmessagesเป็น required field เดียว - แบ่งเป็น 20 OpenAI standard params (ส่งตรงได้เลยใน LiteLLM) และ 32 vLLM extension params (ต้องส่งผ่าน
extra_body) - params ที่น่าสนใจและไม่ค่อยมีคนพูดถึง:
reasoning_effort,thinking_token_budget,min_p,cache_salt,truncate_prompt_tokens,bad_words,allowed_token_ids - ดึง schema ได้จาก
GET /openapi.jsonของ vLLM server โดยตรง
จุดเริ่มต้น: อยากรู้ว่ามีอะไรส่งได้บ้าง
ปกติผมเรียก vLLM ผ่าน LiteLLM Proxy ซึ่งเป็น OpenAI-compatible gateway ส่ง temperature, top_p, max_tokens, stream ไปก็พอใช้ แต่พออ่าน model card ของ sakamakismile/Ornith-1.0-35B-NVFP4 เจอ chat_template_kwargs, enable_thinking ก็เริ่มสงสัย — แล้ว params แบบนี้ส่งยังไง? และยังมีอะไรอีกที่ไม่รู้?
วิธีที่ง่ายที่สุดคือดึง OpenAPI schema จาก vLLM โดยตรง:
curl -s http://10.0.0.246:8000/openapi.json \
| python3 -c "import sys,json; d=json.load(sys.stdin); \
s=d['components']['schemas']['ChatCompletionRequest']; \
req=s.get('required',[]); \
[print(f'{k:40s} req={k in req!s:5s} default={v.get(\"default\",\"—\")}') \
for k,v in sorted(s['properties'].items())]"
ผลลัพธ์คือ params ทั้ง 52 ตัว — ผมเลยแบ่งออกเป็น 2 กลุ่มตามวิธีการส่งใน LiteLLM
ตรวจสอบโมเดลที่รันอยู่
ก่อนดู params มาดูโมเดลที่รันอยู่บน server ก่อน:
curl -s http://10.0.0.246:8000/v1/models | python3 -m json.tool
{
"object": "list",
"data": [
{
"id": "qwen3.6-35b-nvfp4",
"object": "model",
"created": 1782503509,
"owned_by": "vllm",
"root": "RedHatAI/Qwen3.6-35B-A3B-NVFP4",
"parent": null,
"max_model_len": 262144,
"permission": [
{
"id": "modelperm-97a35f4748a00237",
"object": "model_permission",
"created": 1782503509,
"allow_create_engine": false,
"allow_sampling": true,
"allow_logprobs": true,
"allow_search_indices": false,
"allow_view": true,
"allow_fine_tuning": false,
"organization": "*",
"group": null,
"is_blocking": false
}
]
}
]
}
สรุปสั้น ๆ:
- Model ID:
qwen3.6-35b-nvfp4— ชื่อที่ใช้ในmodelfield ตอนเรียก API - Root Model:
RedHatAI/Qwen3.6-35B-A3B-NVFP4— โมเดลต้นทางจาก HuggingFace - Max Context Length: 262,144 tokens (~256K)
- vLLM version:
0.23.1rc1.dev480+gd980a3cc6.d20260626-9ddfac40
กลุ่มที่ 1: OpenAI Standard Params (ส่งตรงได้ใน LiteLLM)
params กลุ่มนี้เป็นมาตรฐาน OpenAI API ส่งได้โดยตรงทั้งจาก client ปกติและผ่าน LiteLLM โดยไม่ต้องใช้ extra_body
| Parameter | Type | Default | หมายเหตุ |
|---|---|---|---|
messages | array | — | required (ตัวเดียว) |
model | string | — | vLLM ใช้ default ถ้ามีโมเดลเดียว |
temperature | number | 1.0 ¹ | ความสุ่มของ output |
top_p | number | 1.0 ¹ | nucleus sampling |
max_tokens | integer | — | deprecated ใน OpenAI ใหม่ แต่ vLLM ยังรองรับ |
max_completion_tokens | integer | — | ทดแทน max_tokens |
stop | string/array | [] | stop sequences |
stream | boolean | false | streaming response |
stream_options | object | — | include_usage etc. |
n | integer | 1 | จำนวน completions |
seed | integer | — | reproducibility |
frequency_penalty | number | 0.0 | ลดการใช้ token ที่ซ้ำบ่อย |
presence_penalty | number | 0.0 | กระตุ้นให้ใช้ token ใหม่ |
logprobs | boolean | false | ส่ง log probabilities กลับมา |
top_logprobs | integer | 0 | จำนวน top tokens ที่ส่ง logprobs |
logit_bias | object | — | ปรับ bias ของ token แบบ manual |
response_format | object | — | JSON schema / structured output |
tools | array | — | function calling definitions |
tool_choice | string/object | none | auto, none, หรือ specific tool |
user | string | — | user identifier |
¹ Engine default —
temperatureและtop_pไม่มีdefaultใน OpenAPI schema แต่ vLLM engine มีค่า default ภายใน:temperature=1.0,top_p=1.0ค่าเหล่านี้เป็น engine-level default ไม่ใช่ schema-level default
กลุ่มที่ 2: vLLM Extension Params (ต้องส่งผ่าน extra_body)
params กลุ่มนี้เป็น vLLM-specific ไม่ใช่ OpenAI standard — ใน LiteLLM ต้องส่งผ่าน extra_body ไม่งั้น LiteLLM จะไม่ forward ไปยัง vLLM
Sampling & Decoding
| Parameter | Type | Default | หมายเหตุ |
|---|---|---|---|
top_k | integer | -1 ¹ | จำกัดการเลือก top-k tokens |
min_p | number | 0.0 ¹ | minimum probability threshold |
repetition_penalty | number | 1.0 ¹ | ลดการใช้ token ที่เคยใช้แล้ว |
min_tokens | integer | 0 | บังคับ generate อย่างน้อย N tokens |
length_penalty | number | 1.0 | ใช้กับ beam search |
use_beam_search | boolean | false | เปิด beam search |
stop_token_ids | array<int> | [] | stop ที่ token IDs แทน string |
bad_words | array | — | ห้าม generate คำเหล่านี้ |
allowed_token_ids | array | — | จำกัดให้ generate เฉพาะ token IDs ที่กำหนด |
ignore_eos | boolean | false | ไม่หยุดที่ EOS token |
Reasoning & Thinking
| Parameter | Type | Default | หมายเหตุ |
|---|---|---|---|
reasoning_effort | enum | — | none/minimal/low/medium/high/xhigh/max — ควบคุม reasoning depth |
thinking_token_budget | integer | — | จำกัดจำนวน thinking tokens |
include_reasoning | boolean | true | ส่ง reasoning content กลับมาใน response |
Caution:
reasoning_effortvsthinking_token_budget— ทั้งสองควบคุม reasoning แต่คนละมิติreasoning_effortเป็น enum ส่วนthinking_token_budgetเป็นตัวเลขจำนวน token แนะนำตั้งreasoning_effortก่อน แล้วใช้thinking_token_budgetกรณีต้องการจำกัดเพิ่ม
Chat Template & Tokenization
| Parameter | Type | Default | หมายเหตุ |
|---|---|---|---|
chat_template | string | — | override Jinja template ทั้งหมด |
chat_template_kwargs | object | — | เช่น {"enable_thinking": true} |
add_generation_prompt | boolean | true | เพิ่ม generation prompt ใน template |
continue_final_message | boolean | false | ต่อจาก message สุดท้าย |
skip_special_tokens | boolean | true | ข้าม special tokens ใน output |
spaces_between_special_tokens | boolean | true | ใส่ช่องว่างระหว่าง special tokens |
add_special_tokens | boolean | false | เพิ่ม special tokens ใน prompt |
truncation_side | enum | — | left / right — ทิศทางการตัด prompt |
truncate_prompt_tokens | integer | — | ตัด prompt ถ้าเกิน limit |
Tool Calling
| Parameter | Type | Default | หมายเหตุ |
|---|---|---|---|
parallel_tool_calls | boolean | true | อนุญาตให้เรียกหลาย tools พร้อมกัน |
Logging & Debugging
| Parameter | Type | Default | หมายเหตุ |
|---|---|---|---|
echo | boolean | false | ส่ง prompt กลับมาด้วย |
prompt_logprobs | integer | — | logprobs ของ prompt tokens |
include_stop_str_in_output | boolean | false | ใส่ stop string ใน output |
return_prompt_text | boolean | — | ส่ง prompt text กลับมา |
return_token_ids | boolean | — | ส่ง token IDs กลับมา |
return_tokens_as_token_ids | boolean | — | ส่ง tokens เป็น token IDs |
Cache & Performance
| Parameter | Type | Default | หมายเหตุ |
|---|---|---|---|
cache_salt | string | — | prefix cache isolation แยก per-user หรือ per-session |
Advanced / Internal
| Parameter | Type | Default | หมายเหตุ |
|---|---|---|---|
structured_outputs | object | — | structured output config |
kv_transfer_params | object | — | KV cache transfer ระหว่าง workers |
vllm_xargs | object | — | vLLM internal extra args |
priority | integer | 0 | request priority |
request_id | string | — | custom request ID |
media_io_kwargs | object | — | multimodal I/O kwargs |
mm_processor_kwargs | object | — | multimodal processor kwargs |
documents | array | — | documents สำหรับ RAG-style template |
repetition_detection | — | — | repetition detection config |
¹ Engine default —
top_k=-1(ไม่จำกัด),min_p=0.0,repetition_penalty=1.0ไม่มีใน OpenAPI schema แต่เป็น vLLM engine default ต้องดูจาก vLLM source code
วิธี Override ผ่าน LiteLLM
ใน LiteLLM Proxy config (config.yaml)
ตั้ง default params ได้ในแต่ละ model alias ผ่าน litellm_params:
model_list:
- model_name: qwen-chat-think
litellm_params:
model: openai/qwen3.6-35b-nvfp4
api_base: http://10.0.0.246:8000/v1
api_key: EMPTY
temperature: 0.6
top_p: 0.95
max_tokens: 8192
# vLLM extension params ผ่าน extra_body
extra_body:
top_k: 20
min_p: 0.0
repetition_penalty: 1.0
chat_template_kwargs:
enable_thinking: true
preserve_thinking: true
parallel_tool_calls: true
include_reasoning: true
ใน LiteLLM Python SDK
from litellm import completion
response = completion(
model="openai/qwen3.6-35b-nvfp4",
api_base="http://10.0.0.246:8000/v1",
api_key="EMPTY",
messages=[{"role": "user", "content": "Write a Python fibonacci function"}],
# OpenAI standard params — ส่งตรงได้เลย
temperature=0.7,
top_p=0.9,
max_tokens=2048,
stream=True,
seed=42,
# vLLM extension params — ต้องส่งผ่าน extra_body
extra_body={
"top_k": 20,
"min_p": 0.05,
"repetition_penalty": 1.05,
"reasoning_effort": "medium", # none / minimal / low / medium / high / xhigh / max
"include_reasoning": True,
"thinking_token_budget": 4096,
"chat_template_kwargs": {
"enable_thinking": True,
},
"parallel_tool_calls": True,
"cache_salt": "session-abc-123", # prefix cache isolation
},
)
ใน curl โดยตรง
curl -s http://10.0.0.246:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.6-35b-nvfp4",
"messages": [{"role": "user", "content": "hi"}],
"temperature": 0.7,
"top_p": 0.9,
"max_tokens": 2048,
"top_k": 20,
"min_p": 0.05,
"repetition_penalty": 1.05,
"reasoning_effort": "medium",
"include_reasoning": true,
"thinking_token_budget": 4096,
"chat_template_kwargs": {"enable_thinking": true},
"logprobs": true,
"top_logprobs": 5
}'
Note: vLLM รับ params ทั้งหมดใน request body โดยตรง — ไม่จำเป็นต้องแยก
extra_bodyเหมือน LiteLLM การแยกextra_bodyเป็น constraint ของ LiteLLM เพื่อแยก OpenAI standard จาก vendor-specific params
Params ที่น่าสนใจและไม่ค่อยมีคนพูดถึง
reasoning_effort — ควบคุมความลึกของ reasoning
เป็น enum 7 ระดับ: none, minimal, low, medium, high, xhigh, max
none= ปิด reasoning เหมือนเดิมmedium= reasoning ปกติ (default โดยปริยาย)max= reasoning เต็มที่ (ค่าเฉพาะของ DeepSeek V4 series)
ใช้คู่กับ thinking_token_budget ได้เพื่อจำกัดจำนวน thinking tokens แบบ hard limit
min_p — กรอง token ที่ probability ต่ำเกินไป
ไม่ใช่ params ที่คนใช้บ่อย แต่มีประโยชน์มากสำหรับ MoE models แบบ Qwen3.6-35B-A3B — กรอง token ที่ probability ต่ำกว่า min_p × max_prob ออกไป ช่วยลด hallucination ในงานที่ต้องการความแม่นยำ
cache_salt — แยก prefix cache ต่อ session
vLLM มี prefix caching ที่ช่วยเร่งความเร็วเมื่อ prompt ซ้ำกัน cache_salt ทำให้แยก cache ของแต่ละ session/user ออกจากกัน ป้องกัน cache pollution และเพิ่ม privacy
truncate_prompt_tokens — ตัด prompt อัตโนมัติ
ถ้า prompt ยาวเกิน max_model_len vLLM จะปฏิเสธ request แต่ถ้าตั้ง truncate_prompt_tokens vLLM จะตัด prompt ให้โดยอัตโนมัติ ไม่ต้องจัดการที่ client
bad_words / allowed_token_ids — ควบคุม output แบบละเอียด
bad_words: ห้าม generate คำที่กำหนด — มีประโยชน์สำหรับ content filteringallowed_token_ids: จำกัดให้ generate เฉพาะ tokens ที่กำหนด — มีประโยชน์สำหรับ structured output ที่ต้องการควบคุม output space แบบเข้มข้น
seed — reproducibility
ตั้ง seed ทำให้ output คงที่ (ถ้า temperature > 0 ก็ยัง reproducible ได้ถ้า seed เดียวกัน) — มีประโยชน์สำหรับ testing และ debugging
ทดสอบจริง: ส่ง params หลายตัวพร้อมกัน
curl -s http://10.0.0.246:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.6-35b-nvfp4",
"messages": [{"role": "user", "content": "hi"}],
"max_tokens": 1,
"top_k": 20,
"min_p": 0.05,
"repetition_penalty": 1.05,
"frequency_penalty": 0.1,
"seed": 42,
"stop": ["\n"],
"logprobs": true,
"top_logprobs": 5
}' | python3 -m json.tool
vLLM ตอบกลับปกติ — ยอมรับ params ทั้งหมด ไม่มี error:
{
"id": "chatcmpl-a6eb09cd195c2953",
"object": "chat.completion",
"model": "qwen3.6-35b-nvfp4",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"reasoning": "Here"
},
"logprobs": null,
"finish_reason": "length"
}
],
"system_fingerprint": "vllm-0.23.1rc1.dev480+gd980a3cc6.d20260626-9ddfac40",
"usage": {
"prompt_tokens": 11,
"total_tokens": 12,
"completion_tokens": 1
}
}
ทดสอบ logprobs กับ top_logprobs:
curl -s http://10.0.0.246:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.6-35b-nvfp4",
"messages": [{"role": "user", "content": "hi"}],
"max_tokens": 1,
"logprobs": true,
"top_logprobs": 5
}' | python3 -m json.tool
ผลลัพธ์:
{
"choices": [
{
"message": {
"role": "assistant",
"reasoning": "Here"
},
"logprobs": {
"content": [
{
"token": "Here",
"logprob": -0.0166,
"top_logprobs": [
{"token": "Here", "logprob": -0.0166},
{"token": "Thinking", "logprob": -4.2666},
{"token": "The", "logprob": -6.6416},
{"token": "We", "logprob": -7.7666},
{"token": "Hmm", "logprob": -8.1416}
]
}
]
},
"finish_reason": "length"
}
]
}
เห็นได้ว่า vLLM ส่ง logprobs กลับมาครบ 5 ตัวตามที่ขอ — มีประโยชน์สำหรับ debugging และวิเคราะห์ความมั่นใจของโมเดล
สรุปกฎการส่ง params ผ่าน LiteLLM
| ประเภท | วิธีส่งใน LiteLLM | ตัวอย่าง |
|---|---|---|
| OpenAI standard | ส่งเป็น top-level param | temperature=0.7, max_tokens=2048 |
| vLLM extension | ส่งผ่าน extra_body | extra_body={"top_k": 20, "min_p": 0.05} |
| LiteLLM Proxy config | ตั้งใน litellm_params | temperature: 0.6 + extra_body: {top_k: 20} |
| curl โดยตรง | ใส่ใน JSON body หมด | "top_k": 20, "min_p": 0.05 |
กฎง่าย ๆ: ถ้าเป็น OpenAI standard param ส่งตรงได้เลย ถ้าเป็น vLLM extension ต้องใช้ extra_body
Conclusion
การเปิด OpenAPI schema ของ vLLM ทำให้เห็น params ทั้ง 52 ตัวที่รองรับ ซึ่งมากกว่าที่คิด และมีหลายตัวที่มีประโยชน์แต่ไม่ค่อยมีคนพูดถึง เช่น reasoning_effort, min_p, cache_salt, truncate_prompt_tokens
การแบ่งเป็น OpenAI standard กับ vLLM extension ช่วยให้รู้ว่าจะส่งผ่าน LiteLLM อย่างไร — standard ส่งตรง, extension ใช้ extra_body กฎนี้ใช้ได้ทั้งใน Python SDK, Proxy config, และการเรียกผ่าน client อื่น ๆ
ถ้าใครใช้ vLLM เป็น backend แนะนำให้ลองรัน curl /openapi.json ดูบ้าง — อาจเจอ params ที่ช่วยแก้ปัญหาที่เคยเจอได้โดยไม่ต้องเขียน workaround เอง
References
- vLLM 0.23.1 docs — Engine flags, quantization support
- vLLM sampling_params.py — Engine-level default values สำหรับ
temperature,top_p,top_kฯลฯ - LiteLLM Proxy Documentation —
extra_bodyและlitellm_paramsconfig - RedHatAI/Qwen3.6-35B-A3B-NVFP4 (Hugging Face) — Model card และ serving recipe
- OpenAI API Reference — Standard params ที่ vLLM รองรับ
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee