Skip to main content

vLLM Chat Completion Params ทั้ง 52 ตัว และวิธี Override ผ่าน LiteLLM

· 10 min read

บันทึก 27 มิถุนายน 2569 — หลังจากใช้ vLLM มาสักพัก อยากรู้ว่าจริง ๆ แล้ว params ที่ส่งได้ใน /v1/chat/completions มีอะไรบ้าง และอันไหนส่งผ่าน LiteLLM ได้โดยตรง อันไหนต้องใช้ extra_body

ผมใช้ vLLM เป็น inference backend มาสักพัก แต่ไม่เคยนั่งดูจริง ๆ ว่า /v1/chat/completions endpoint รองรับ parameters อะไรบ้างเต็ม ๆ ส่วนใหญ่ก็แค่ส่ง temperature, top_p, max_tokens ไปแล้วก็จบ — จนวันนี้ลองเปิด OpenAPI schema ดู ถึงรู้ว่ามี params ที่ใช้ไม่เคยรู้จักตั้งหลายตัว

TL;DR​

  • vLLM 0.23.1 รองรับ 52 parameters บน /v1/chat/completions — มี messages เป็น required field เดียว
  • แบ่งเป็น 20 OpenAI standard params (ส่งตรงได้เลยใน LiteLLM) และ 32 vLLM extension params (ต้องส่งผ่าน extra_body)
  • params ที่น่าสนใจและไม่ค่อยมีคนพูดถึง: reasoning_effort, thinking_token_budget, min_p, cache_salt, truncate_prompt_tokens, bad_words, allowed_token_ids
  • ดึง schema ได้จาก GET /openapi.json ของ vLLM server โดยตรง

จุดเริ่มต้น: อยากรู้ว่ามีอะไรส่งได้บ้าง​

ปกติผมเรียก vLLM ผ่าน LiteLLM Proxy ซึ่งเป็น OpenAI-compatible gateway ส่ง temperature, top_p, max_tokens, stream ไปก็พอใช้ แต่พออ่าน model card ของ sakamakismile/Ornith-1.0-35B-NVFP4 เจอ chat_template_kwargs, enable_thinking ก็เริ่มสงสัย — แล้ว params แบบนี้ส่งยังไง? และยังมีอะไรอีกที่ไม่รู้?

วิธีที่ง่ายที่สุดคือดึง OpenAPI schema จาก vLLM โดยตรง:

curl -s http://10.0.0.246:8000/openapi.json \
| python3 -c "import sys,json; d=json.load(sys.stdin); \
s=d['components']['schemas']['ChatCompletionRequest']; \
req=s.get('required',[]); \
[print(f'{k:40s} req={k in req!s:5s} default={v.get(\"default\",\"—\")}') \
for k,v in sorted(s['properties'].items())]"

ผลลัพธ์คือ params ทั้ง 52 ตัว — ผมเลยแบ่งออกเป็น 2 กลุ่มตามวิธีการส่งใน LiteLLM

ตรวจสอบโมเดลที่รันอยู่​

ก่อนดู params มาดูโมเดลที่รันอยู่บน server ก่อน:

curl -s http://10.0.0.246:8000/v1/models | python3 -m json.tool
{
"object": "list",
"data": [
{
"id": "qwen3.6-35b-nvfp4",
"object": "model",
"created": 1782503509,
"owned_by": "vllm",
"root": "RedHatAI/Qwen3.6-35B-A3B-NVFP4",
"parent": null,
"max_model_len": 262144,
"permission": [
{
"id": "modelperm-97a35f4748a00237",
"object": "model_permission",
"created": 1782503509,
"allow_create_engine": false,
"allow_sampling": true,
"allow_logprobs": true,
"allow_search_indices": false,
"allow_view": true,
"allow_fine_tuning": false,
"organization": "*",
"group": null,
"is_blocking": false
}
]
}
]
}

สรุปสั้น ๆ:

  • Model ID: qwen3.6-35b-nvfp4 — ชื่อที่ใช้ใน model field ตอนเรียก API
  • Root Model: RedHatAI/Qwen3.6-35B-A3B-NVFP4 — โมเดลต้นทางจาก HuggingFace
  • Max Context Length: 262,144 tokens (~256K)
  • vLLM version: 0.23.1rc1.dev480+gd980a3cc6.d20260626-9ddfac40

กลุ่มที่ 1: OpenAI Standard Params (ส่งตรงได้ใน LiteLLM)​

params กลุ่มนี้เป็นมาตรฐาน OpenAI API ส่งได้โดยตรงทั้งจาก client ปกติและผ่าน LiteLLM โดยไม่ต้องใช้ extra_body

ParameterTypeDefaultหมายเหตุ
messagesarray—required (ตัวเดียว)
modelstring—vLLM ใช้ default ถ้ามีโมเดลเดียว
temperaturenumber1.0 ¹ความสุ่มของ output
top_pnumber1.0 ¹nucleus sampling
max_tokensinteger—deprecated ใน OpenAI ใหม่ แต่ vLLM ยังรองรับ
max_completion_tokensinteger—ทดแทน max_tokens
stopstring/array[]stop sequences
streambooleanfalsestreaming response
stream_optionsobject—include_usage etc.
ninteger1จำนวน completions
seedinteger—reproducibility
frequency_penaltynumber0.0ลดการใช้ token ที่ซ้ำบ่อย
presence_penaltynumber0.0กระตุ้นให้ใช้ token ใหม่
logprobsbooleanfalseส่ง log probabilities กลับมา
top_logprobsinteger0จำนวน top tokens ที่ส่ง logprobs
logit_biasobject—ปรับ bias ของ token แบบ manual
response_formatobject—JSON schema / structured output
toolsarray—function calling definitions
tool_choicestring/objectnoneauto, none, หรือ specific tool
userstring—user identifier

¹ Engine default — temperature และ top_p ไม่มี default ใน OpenAPI schema แต่ vLLM engine มีค่า default ภายใน: temperature=1.0, top_p=1.0 ค่าเหล่านี้เป็น engine-level default ไม่ใช่ schema-level default

กลุ่มที่ 2: vLLM Extension Params (ต้องส่งผ่าน extra_body)​

params กลุ่มนี้เป็น vLLM-specific ไม่ใช่ OpenAI standard — ใน LiteLLM ต้องส่งผ่าน extra_body ไม่งั้น LiteLLM จะไม่ forward ไปยัง vLLM

Sampling & Decoding​

ParameterTypeDefaultหมายเหตุ
top_kinteger-1 ¹จำกัดการเลือก top-k tokens
min_pnumber0.0 ¹minimum probability threshold
repetition_penaltynumber1.0 ¹ลดการใช้ token ที่เคยใช้แล้ว
min_tokensinteger0บังคับ generate อย่างน้อย N tokens
length_penaltynumber1.0ใช้กับ beam search
use_beam_searchbooleanfalseเปิด beam search
stop_token_idsarray<int>[]stop ที่ token IDs แทน string
bad_wordsarray—ห้าม generate คำเหล่านี้
allowed_token_idsarray—จำกัดให้ generate เฉพาะ token IDs ที่กำหนด
ignore_eosbooleanfalseไม่หยุดที่ EOS token

Reasoning & Thinking​

ParameterTypeDefaultหมายเหตุ
reasoning_effortenum—none/minimal/low/medium/high/xhigh/max — ควบคุม reasoning depth
thinking_token_budgetinteger—จำกัดจำนวน thinking tokens
include_reasoningbooleantrueส่ง reasoning content กลับมาใน response

Caution: reasoning_effort vs thinking_token_budget — ทั้งสองควบคุม reasoning แต่คนละมิติ reasoning_effort เป็น enum ส่วน thinking_token_budget เป็นตัวเลขจำนวน token แนะนำตั้ง reasoning_effort ก่อน แล้วใช้ thinking_token_budget กรณีต้องการจำกัดเพิ่ม

Chat Template & Tokenization​

ParameterTypeDefaultหมายเหตุ
chat_templatestring—override Jinja template ทั้งหมด
chat_template_kwargsobject—เช่น {"enable_thinking": true}
add_generation_promptbooleantrueเพิ่ม generation prompt ใน template
continue_final_messagebooleanfalseต่อจาก message สุดท้าย
skip_special_tokensbooleantrueข้าม special tokens ใน output
spaces_between_special_tokensbooleantrueใส่ช่องว่างระหว่าง special tokens
add_special_tokensbooleanfalseเพิ่ม special tokens ใน prompt
truncation_sideenum—left / right — ทิศทางการตัด prompt
truncate_prompt_tokensinteger—ตัด prompt ถ้าเกิน limit

Tool Calling​

ParameterTypeDefaultหมายเหตุ
parallel_tool_callsbooleantrueอนุญาตให้เรียกหลาย tools พร้อมกัน

Logging & Debugging​

ParameterTypeDefaultหมายเหตุ
echobooleanfalseส่ง prompt กลับมาด้วย
prompt_logprobsinteger—logprobs ของ prompt tokens
include_stop_str_in_outputbooleanfalseใส่ stop string ใน output
return_prompt_textboolean—ส่ง prompt text กลับมา
return_token_idsboolean—ส่ง token IDs กลับมา
return_tokens_as_token_idsboolean—ส่ง tokens เป็น token IDs

Cache & Performance​

ParameterTypeDefaultหมายเหตุ
cache_saltstring—prefix cache isolation แยก per-user หรือ per-session

Advanced / Internal​

ParameterTypeDefaultหมายเหตุ
structured_outputsobject—structured output config
kv_transfer_paramsobject—KV cache transfer ระหว่าง workers
vllm_xargsobject—vLLM internal extra args
priorityinteger0request priority
request_idstring—custom request ID
media_io_kwargsobject—multimodal I/O kwargs
mm_processor_kwargsobject—multimodal processor kwargs
documentsarray—documents สำหรับ RAG-style template
repetition_detection——repetition detection config

¹ Engine default — top_k=-1 (ไม่จำกัด), min_p=0.0, repetition_penalty=1.0 ไม่มีใน OpenAPI schema แต่เป็น vLLM engine default ต้องดูจาก vLLM source code

วิธี Override ผ่าน LiteLLM​

ใน LiteLLM Proxy config (config.yaml)​

ตั้ง default params ได้ในแต่ละ model alias ผ่าน litellm_params:

model_list:
- model_name: qwen-chat-think
litellm_params:
model: openai/qwen3.6-35b-nvfp4
api_base: http://10.0.0.246:8000/v1
api_key: EMPTY
temperature: 0.6
top_p: 0.95
max_tokens: 8192
# vLLM extension params ผ่าน extra_body
extra_body:
top_k: 20
min_p: 0.0
repetition_penalty: 1.0
chat_template_kwargs:
enable_thinking: true
preserve_thinking: true
parallel_tool_calls: true
include_reasoning: true

ใน LiteLLM Python SDK​

from litellm import completion

response = completion(
model="openai/qwen3.6-35b-nvfp4",
api_base="http://10.0.0.246:8000/v1",
api_key="EMPTY",
messages=[{"role": "user", "content": "Write a Python fibonacci function"}],
# OpenAI standard params — ส่งตรงได้เลย
temperature=0.7,
top_p=0.9,
max_tokens=2048,
stream=True,
seed=42,
# vLLM extension params — ต้องส่งผ่าน extra_body
extra_body={
"top_k": 20,
"min_p": 0.05,
"repetition_penalty": 1.05,
"reasoning_effort": "medium", # none / minimal / low / medium / high / xhigh / max
"include_reasoning": True,
"thinking_token_budget": 4096,
"chat_template_kwargs": {
"enable_thinking": True,
},
"parallel_tool_calls": True,
"cache_salt": "session-abc-123", # prefix cache isolation
},
)

ใน curl โดยตรง​

curl -s http://10.0.0.246:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.6-35b-nvfp4",
"messages": [{"role": "user", "content": "hi"}],
"temperature": 0.7,
"top_p": 0.9,
"max_tokens": 2048,
"top_k": 20,
"min_p": 0.05,
"repetition_penalty": 1.05,
"reasoning_effort": "medium",
"include_reasoning": true,
"thinking_token_budget": 4096,
"chat_template_kwargs": {"enable_thinking": true},
"logprobs": true,
"top_logprobs": 5
}'

Note: vLLM รับ params ทั้งหมดใน request body โดยตรง — ไม่จำเป็นต้องแยก extra_body เหมือน LiteLLM การแยก extra_body เป็น constraint ของ LiteLLM เพื่อแยก OpenAI standard จาก vendor-specific params

Params ที่น่าสนใจและไม่ค่อยมีคนพูดถึง​

reasoning_effort — ควบคุมความลึกของ reasoning​

เป็น enum 7 ระดับ: none, minimal, low, medium, high, xhigh, max

  • none = ปิด reasoning เหมือนเดิม
  • medium = reasoning ปกติ (default โดยปริยาย)
  • max = reasoning เต็มที่ (ค่าเฉพาะของ DeepSeek V4 series)

ใช้คู่กับ thinking_token_budget ได้เพื่อจำกัดจำนวน thinking tokens แบบ hard limit

min_p — กรอง token ที่ probability ต่ำเกินไป​

ไม่ใช่ params ที่คนใช้บ่อย แต่มีประโยชน์มากสำหรับ MoE models แบบ Qwen3.6-35B-A3B — กรอง token ที่ probability ต่ำกว่า min_p × max_prob ออกไป ช่วยลด hallucination ในงานที่ต้องการความแม่นยำ

cache_salt — แยก prefix cache ต่อ session​

vLLM มี prefix caching ที่ช่วยเร่งความเร็วเมื่อ prompt ซ้ำกัน cache_salt ทำให้แยก cache ของแต่ละ session/user ออกจากกัน ป้องกัน cache pollution และเพิ่ม privacy

truncate_prompt_tokens — ตัด prompt อัตโนมัติ​

ถ้า prompt ยาวเกิน max_model_len vLLM จะปฏิเสธ request แต่ถ้าตั้ง truncate_prompt_tokens vLLM จะตัด prompt ให้โดยอัตโนมัติ ไม่ต้องจัดการที่ client

bad_words / allowed_token_ids — ควบคุม output แบบละเอียด​

  • bad_words: ห้าม generate คำที่กำหนด — มีประโยชน์สำหรับ content filtering
  • allowed_token_ids: จำกัดให้ generate เฉพาะ tokens ที่กำหนด — มีประโยชน์สำหรับ structured output ที่ต้องการควบคุม output space แบบเข้มข้น

seed — reproducibility​

ตั้ง seed ทำให้ output คงที่ (ถ้า temperature > 0 ก็ยัง reproducible ได้ถ้า seed เดียวกัน) — มีประโยชน์สำหรับ testing และ debugging

ทดสอบจริง: ส่ง params หลายตัวพร้อมกัน​

curl -s http://10.0.0.246:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.6-35b-nvfp4",
"messages": [{"role": "user", "content": "hi"}],
"max_tokens": 1,
"top_k": 20,
"min_p": 0.05,
"repetition_penalty": 1.05,
"frequency_penalty": 0.1,
"seed": 42,
"stop": ["\n"],
"logprobs": true,
"top_logprobs": 5
}' | python3 -m json.tool

vLLM ตอบกลับปกติ — ยอมรับ params ทั้งหมด ไม่มี error:

{
"id": "chatcmpl-a6eb09cd195c2953",
"object": "chat.completion",
"model": "qwen3.6-35b-nvfp4",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"reasoning": "Here"
},
"logprobs": null,
"finish_reason": "length"
}
],
"system_fingerprint": "vllm-0.23.1rc1.dev480+gd980a3cc6.d20260626-9ddfac40",
"usage": {
"prompt_tokens": 11,
"total_tokens": 12,
"completion_tokens": 1
}
}

ทดสอบ logprobs กับ top_logprobs:

curl -s http://10.0.0.246:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.6-35b-nvfp4",
"messages": [{"role": "user", "content": "hi"}],
"max_tokens": 1,
"logprobs": true,
"top_logprobs": 5
}' | python3 -m json.tool

ผลลัพธ์:

{
"choices": [
{
"message": {
"role": "assistant",
"reasoning": "Here"
},
"logprobs": {
"content": [
{
"token": "Here",
"logprob": -0.0166,
"top_logprobs": [
{"token": "Here", "logprob": -0.0166},
{"token": "Thinking", "logprob": -4.2666},
{"token": "The", "logprob": -6.6416},
{"token": "We", "logprob": -7.7666},
{"token": "Hmm", "logprob": -8.1416}
]
}
]
},
"finish_reason": "length"
}
]
}

เห็นได้ว่า vLLM ส่ง logprobs กลับมาครบ 5 ตัวตามที่ขอ — มีประโยชน์สำหรับ debugging และวิเคราะห์ความมั่นใจของโมเดล

สรุปกฎการส่ง params ผ่าน LiteLLM​

ประเภทวิธีส่งใน LiteLLMตัวอย่าง
OpenAI standardส่งเป็น top-level paramtemperature=0.7, max_tokens=2048
vLLM extensionส่งผ่าน extra_bodyextra_body={"top_k": 20, "min_p": 0.05}
LiteLLM Proxy configตั้งใน litellm_paramstemperature: 0.6 + extra_body: {top_k: 20}
curl โดยตรงใส่ใน JSON body หมด"top_k": 20, "min_p": 0.05

กฎง่าย ๆ: ถ้าเป็น OpenAI standard param ส่งตรงได้เลย ถ้าเป็น vLLM extension ต้องใช้ extra_body

Conclusion​

การเปิด OpenAPI schema ของ vLLM ทำให้เห็น params ทั้ง 52 ตัวที่รองรับ ซึ่งมากกว่าที่คิด และมีหลายตัวที่มีประโยชน์แต่ไม่ค่อยมีคนพูดถึง เช่น reasoning_effort, min_p, cache_salt, truncate_prompt_tokens

การแบ่งเป็น OpenAI standard กับ vLLM extension ช่วยให้รู้ว่าจะส่งผ่าน LiteLLM อย่างไร — standard ส่งตรง, extension ใช้ extra_body กฎนี้ใช้ได้ทั้งใน Python SDK, Proxy config, และการเรียกผ่าน client อื่น ๆ

ถ้าใครใช้ vLLM เป็น backend แนะนำให้ลองรัน curl /openapi.json ดูบ้าง — อาจเจอ params ที่ช่วยแก้ปัญหาที่เคยเจอได้โดยไม่ต้องเขียน workaround เอง

References​

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...