Skip to main content

2 posts tagged with "virtual-model"

View All Tags

Virtual Models บน LiteLLM Proxy: Ornith-1.0-35B 10 profiles ใช้ reasoning_effort คุมพฤติกรรม

· 13 min read

หลังจาก deploy Ornith-1.0-35B-NVFP4 บน DGX Spark สำเร็จแล้ว (ดูรายละเอียดใน บทความก่อนหน้า) ขั้นตอนต่อไปคือสร้าง Virtual Models ผ่าน LiteLLM Proxy เหมือนที่เคยทำกับ Qwen3.6

ความต่างสำคัญ: Ornith มี reasoning_effort 7 levels (none/minimal/low/medium/high/xhigh/max) แทนที่แค่ enable_thinking: true|false แบบ Qwen ทำให้คุมความลึกของ reasoning ได้ละเอียดกว่า และมี thinking_token_budget สำหรับจำกัดจำนวน thinking tokens ต่อ request

Virtual Models บน LiteLLM Proxy: 1 โมเดล 10 profiles ใช้ให้เหมาะกับงาน

· 7 min read

ผมใช้ Qwen3.6-35B-A3B-NVFP4 เป็น backend model ตัวเดียว แล้วสร้าง Virtual Models ผ่าน LiteLLM Proxy เป็น 10 profiles ตามลักษณะงาน

ทุก profile ชี้ไปที่โมเดลเดียวกัน แต่ override sampling parameters ต่างกัน — ทำให้โมเดลเดียวกันตอบออกมา "คนละคน" ตาม use case