Skip to main content

4 posts tagged with "sampling"

View All Tags

vLLM Chat Completion Params ทั้ง 52 ตัว และวิธี Override ผ่าน LiteLLM

· 10 min read

บันทึก 27 มิถุนายน 2569 — หลังจากใช้ vLLM มาสักพัก อยากรู้ว่าจริง ๆ แล้ว params ที่ส่งได้ใน /v1/chat/completions มีอะไรบ้าง และอันไหนส่งผ่าน LiteLLM ได้โดยตรง อันไหนต้องใช้ extra_body

ผมใช้ vLLM เป็น inference backend มาสักพัก แต่ไม่เคยนั่งดูจริง ๆ ว่า /v1/chat/completions endpoint รองรับ parameters อะไรบ้างเต็ม ๆ ส่วนใหญ่ก็แค่ส่ง temperature, top_p, max_tokens ไปแล้วก็จบ — จนวันนี้ลองเปิด OpenAPI schema ดู ถึงรู้ว่ามี params ที่ใช้ไม่เคยรู้จักตั้งหลายตัว

Virtual Models บน LiteLLM Proxy: Ornith-1.0-35B 10 profiles ใช้ reasoning_effort คุมพฤติกรรม

· 13 min read

หลังจาก deploy Ornith-1.0-35B-NVFP4 บน DGX Spark สำเร็จแล้ว (ดูรายละเอียดใน บทความก่อนหน้า) ขั้นตอนต่อไปคือสร้าง Virtual Models ผ่าน LiteLLM Proxy เหมือนที่เคยทำกับ Qwen3.6

ความต่างสำคัญ: Ornith มี reasoning_effort 7 levels (none/minimal/low/medium/high/xhigh/max) แทนที่แค่ enable_thinking: true|false แบบ Qwen ทำให้คุมความลึกของ reasoning ได้ละเอียดกว่า และมี thinking_token_budget สำหรับจำกัดจำนวน thinking tokens ต่อ request

Qwen3.6-35B-A3B บน DGX Spark: เรื่อง sampling ที่ผมตั้งผิดมาตลอด

· 12 min read

ผมตั้งค่า vLLM ที่ผ่านมาหลายครั้งด้วยสูตรเดียวตลอด — temperature: 0.6, top_p: 0.95, top_k: 20 แล้วก็ปล่อยให้ client ไป override เอาเอง ไม่ว่าจะเป็น trading bot, Hermes agent, code review — ใช้ค่าเดียวกันหมด

จนเมื่อวานนี้ผมเปิด Hugging Face model card ของ Qwen3.6-35B-A3B อ่านเล่น ๆ ในส่วน Sampling Parameters ถึงได้รู้ว่า — Qwen team แนะนำ sampling ตาม mode และ task type ไม่ใช่ตาม benchmark หรือ use case แบบที่ผมเข้าใจ

อ้าว ผมเลยต้องกลับมานั่งคิดใหม่

Virtual Models บน LiteLLM Proxy: 1 โมเดล 10 profiles ใช้ให้เหมาะกับงาน

· 7 min read

ผมใช้ Qwen3.6-35B-A3B-NVFP4 เป็น backend model ตัวเดียว แล้วสร้าง Virtual Models ผ่าน LiteLLM Proxy เป็น 10 profiles ตามลักษณะงาน

ทุก profile ชี้ไปที่โมเดลเดียวกัน แต่ override sampling parameters ต่างกัน — ทำให้โมเดลเดียวกันตอบออกมา "คนละคน" ตาม use case