Skip to main content

Virtual Models บน LiteLLM Proxy: Ornith-1.0-35B 10 profiles ใช้ reasoning_effort คุมพฤติกรรม

· 13 min read

หลังจาก deploy Ornith-1.0-35B-NVFP4 บน DGX Spark สำเร็จแล้ว (ดูรายละเอียดใน บทความก่อนหน้า) ขั้นตอนต่อไปคือสร้าง Virtual Models ผ่าน LiteLLM Proxy เหมือนที่เคยทำกับ Qwen3.6

ความต่างสำคัญ: Ornith มี reasoning_effort 7 levels (none/minimal/low/medium/high/xhigh/max) แทนที่แค่ enable_thinking: true|false แบบ Qwen ทำให้คุมความลึกของ reasoning ได้ละเอียดกว่า และมี thinking_token_budget สำหรับจำกัดจำนวน thinking tokens ต่อ request

TL;DR​

  • ใช้ Ornith-1.0-35B-NVFP4 1 ตัว backend สร้าง 10 profiles ผ่าน LiteLLM Proxy
  • แต่ละ profile ตั้งค่า reasoning_effort, temperature, top_p, thinking_token_budget, max_tokens ต่างกันตามงาน
  • ornith-chat เป็น default ใช้บ่อยสุด 50% รองด้วย ornith-coder 20%, ornith-agent 10%, ornith-terminal 10% และอื่น ๆ 10%
  • เน้น profiles สำหรับ agentic coding มากกว่า Qwen เพราะ Ornith ถนัด Terminal-Bench และ SWE-Bench
  • reasoning_effort เป็นพารามิเตอร์หลักในการคุมพฤติกรรม แทนที่ enable_thinking แบบ binary

ตาราง 10 Virtual Profiles​

ProfileReasoning EffortTop-KTop-PTempThinking BudgetMax TokensUse Case
ornith-chat ⭐ Defaultmedium200.950.6—8,192General Assistant, Web Search, Daily Chat
ornith-coderhigh200.950.24,0968,192Python, TypeScript, SQL, Docker, DevOps
ornith-agentlow200.900.12,0488,192Deterministic Tool Calling, Workflow Execution
ornith-deepxhigh-10.950.78,19216,384Architecture Review, System Design, Research
ornith-debughigh200.950.34,09616,384Debugging, RCA, Log Analysis, Stacktrace
ornith-terminalhigh200.900.14,0968,192CLI Agent, Shell Commands, Terminal-Bench Tasks
ornith-longctxhigh200.950.24,09616,384RAG, Codebase Analysis, Documentation
ornith-fastminimal200.950.85124,096Quick Chat, FAQ, Short Answers
ornith-creativemedium400.950.92,04812,288Brainstorming, Content Writing, Blog Draft
ornith-structurednone50.200.1—4,096Classification, Extraction, JSON Output, Commit Msg

ทำไม reasoning_effort สำคัญกว่า enable_thinking​

Qwen3.6 ใช้ chat_template_kwargs.enable_thinking: true|false เป็นสวิตช์ on/off — มีแค่ 2 สถานะ

Ornith มี reasoning_effort ถึง 7 levels:

Levelพฤติกรรมเหมาะกับ
noneไม่คิดเลย ตอบทันทีClassification, Extraction, Structured Output
minimalคิดนิดหน่อย เร็วมากQuick Chat, FAQ
lowคิดสั้น ๆAgent Tool Calling (deterministic)
mediumคิดพอประมาณGeneral Chat, Creative
highคิดละเอียดCoding, Debug, Long Context
xhighคิดลึกมากArchitecture, System Design
maxคิดเต็มกำลัง (เฉพาะ DeepSeek V4)— ไม่ใช้กับ Ornith

แต่ละ level มี trade-off ระหว่างความเร็วกับความลึกของคำตอบ — none เร็วสุดแต่อาจตอบผิดงานซับซ้อน, xhigh ช้าแต่แม่นยำสำหรับงานวิเคราะห์

Note: max เป็นค่าเฉพาะของ DeepSeek V4 series ไม่ใช่ส่วนหนึ่งของ OpenAI API standard — ใช้กับ Ornith ไม่ได้ผลต่างจาก xhigh

Official Sampling Parameters จาก Model Card​

Model card ของ Ornith-1.0-35B ระบุค่า sampling ที่แนะนำไว้ชัดเจน:

บริบทtemperaturetop_ptop_kที่มา
Transformers example0.60.9520Model card Quickstart
Chat Completions API0.60.95—Model card API example
Tool calling0.60.95—Model card agentic example
Terminal-Bench 2.11.01.0—Benchmark eval settings
SWE-Bench Verified1.00.95—Benchmark eval settings
SWE Atlas QnA1.00.95—Benchmark eval settings
ClawEval0.6——Agentic code benchmark

สิ่งที่เรียนรู้จาก model card:

  • top_k=20 มาจาก official model card ไม่ใช่ copy จาก Qwen — เป็นค่าที่ authors แนะนำโดยตรง
  • temperature=0.6 เป็นค่า default ที่ authors ใช้ในทุกตัวอย่าง (chat, tool calling, ClawEval)
  • Benchmark eval ใช้ temperature=1.0 เพื่อให้โมเดล explore ทางเลือกได้กว้าง แต่ production ใช้ค่าต่ำกว่าเพื่อความ deterministic
  • ไม่มีการระบุ repetition_penalty หรือ presence_penalty ใน model card — ค่า default ของ vLLM คือ 1.0 และ 0.0 ตามลำดับ

Note: ทำไม production ใช้ temp ต่ำกว่า benchmark — Benchmark ใช้ temp=1.0 เพื่อวัดศักยภาพสูงสุดของโมเดล (best case) แต่ใน production เราต้องการความสม่ำเสมอและความปลอดภัย (โดยเฉพาะ shell commands) จึงใช้ temp ต่ำกว่า เป็น trade-off ระหว่าง exploration กับ determinism

รายละเอียดแต่ละ Profile​

ornith-chat ⭐ Default​

  • Use Case: General Assistant, Web Search, News Summary, Daily Chat, Tool Calling
  • เหตุผล: เป็น Primary Model / Router reasoning_effort: medium สมดุลระหว่างความเร็วกับความลึก
  • temperature: 0.6 สมดุลระหว่างความเป็นธรรมชาติกับความแม่นยำ
  • max_tokens: 8192 เผื่อพื้นที่ reasoning + คำตอบ
{
"min_p": 0,
"top_k": 20,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "medium",
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 8192,
"temperature": 0.6,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}

ornith-coder​

  • Use Case: Python, TypeScript, SQL, Docker, MCP, DevOps
  • เหตุผล: Coding ต้องการความถูกต้องสูง reasoning_effort: high ให้คิดละเอียดก่อนเขียนโค้ด
  • temperature: 0.2 ลด hallucination
  • thinking_token_budget: 4096 จำกัด thinking ไม่ให้ยาวเกินไป เพราะ code task มักมีคำตอบชัดเจน
{
"min_p": 0,
"top_k": 20,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "high",
"thinking_token_budget": 4096,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 8192,
"temperature": 0.2,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}

ornith-agent​

  • Use Case: Deterministic Tool Calling, Workflow Execution, Multi-Step Tasks
  • เหตุผล: Agent ต้องการ deterministic สูงสุด reasoning_effort: low คิดสั้น ๆ ตัดสินใจเร็ว
  • temperature: 0.1 เพื่อให้ parser อ่าน tool call ได้แม่นยำ
  • thinking_token_budget: 2048 จำกัด thinking ให้กระชับ เพราะ agent มักเรียกหลายรอบ
{
"min_p": 0,
"top_k": 20,
"top_p": 0.90,
"extra_body": {
"reasoning_effort": "low",
"thinking_token_budget": 2048,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 8192,
"temperature": 0.1,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}

ornith-deep​

  • Use Case: Architecture Review, Research, Planning, System Design
  • เหตุผล: งานวิเคราะห์เชิงระบบต้องการความลึกสูงสุด reasoning_effort: xhigh
  • temperature: 0.7 ให้มีความยืดหยุ่นในการเสนอทางเลือก
  • top_k: -1 (ปิดการกรอง) เพราะ xhigh reasoning ต้อง explore หลายทาง การจำกัด top_k อาจตัดทางเลือกที่จำเป็น
  • thinking_token_budget: 8192 ให้พื้นที่คิดกว้าง ๆ
  • max_tokens: 16384 เผื่อคำตอบยาว
ทำไม top_k=-1 สำหรับ ornith-deep

top_k คัดเฉพาะ k tokens ที่มี probability สูงสุดมาพิจารณาในแต่ละ step ค่า -1 หมายถึง ปิดการกรอง ให้พิจารณาทุก token ใน vocabulary

  • Reasoning chain ต้อง explore — แต่ละ step โมเดลอาจต้องเลือก token ที่ probability ไม่สูงสุด แต่เป็นทางเลือกที่นำไปสู่เส้นทางคิดที่ถูกต้อง เช่น ตอนพิจารณา "อาจเป็นเพราะ..." โมเดลอาจต้องเลือกคำที่อยู่อันดับ 30–50 ใน vocabulary ไม่ใช่ 20 แรก
  • top_k=20 ตัด long tail ของการคิด — งานทั่วไป (chat, coding) คำตอบอยู่ใน 20 ตัวแรกอยู่แล้ว แต่งานวิเคราะห์เชิงระบบ (architecture, system design) บางครั้งต้อง explore ทางเลือกที่ไม่เด่น
  • top_p=0.95 ยังเป็นตัวคุมอยู่ — แม้ปิด top_k แต่ nucleus sampling ที่ top_p=0.95 ยังกรองเฉพาะ tokens ที่ probability รวมกันถึง 95% จึงไม่ได้สุ่มอย่างไร้ขอบเขต
  • Trade-off — top_k=-1 อาจทำให้คำตอบหลากหลายขึ้นแต่ช้าลงเล็กน้อย แต่สำหรับ ornith-deep ที่เน้นความลึกมากกว่าความเร็ว การ trade-off นี้คุ้ม

สรุป: top_k=20 เหมาะกับงานที่มีคำตอบค่อนข้างชัด (chat, coding) ส่วน top_k=-1 เหมาะกับงานที่ต้อง explore ทางเลือกกว้าง (deep reasoning, architecture) โดยมี top_p=0.95 เป็นตัวคุมความเสี่ยงอยู่แล้ว

{
"min_p": 0,
"top_k": -1,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "xhigh",
"thinking_token_budget": 8192,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 16384,
"temperature": 0.7,
"presence_penalty": 0.2,
"repetition_penalty": 1,
"parallel_tool_calls": true
}

ornith-debug​

  • Use Case: Debugging, RCA, Log Analysis, Incident Investigation
  • เหตุผล: หาสาเหตุต้องคิดละเอียด reasoning_effort: high แต่ temperature: 0.3 รักษาความแม่นยำ
  • max_tokens: 16384 เผื่ออ่าน log ยาว ๆ
{
"min_p": 0,
"top_k": 20,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "high",
"thinking_token_budget": 4096,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 16384,
"temperature": 0.3,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}

ornith-terminal​

  • Use Case: CLI Agent, Shell Commands, Terminal-Bench Tasks
  • เหตุผล: Ornith ถนัด Terminal-Bench (64.2 vs Qwen 52.5) profile นี้เน้นการสั่ง shell โดยตรง
  • reasoning_effort: high คิดก่อนสั่งคำสั่งอันตราย
  • temperature: 0.1 deterministic สูงสุด เพราะ shell command ผิดนิดเดียวอันตรายมาก
  • top_p: 0.90 แคบกว่า chat เพื่อลดความเสี่ยง
{
"min_p": 0,
"top_k": 20,
"top_p": 0.90,
"extra_body": {
"reasoning_effort": "high",
"thinking_token_budget": 4096,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 8192,
"temperature": 0.1,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}

ornith-longctx​

  • Use Case: RAG, Codebase Analysis, Documentation, Long Reports
  • เหตุผล: ต้องการ faithfulness สูงสุดกับ context ยาว reasoning_effort: high ช่วยวิเคราะห์ข้อมูลใน context
  • temperature: 0.2 ไม่ให้สร้างข้อมูลเอง
{
"min_p": 0,
"top_k": 20,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "high",
"thinking_token_budget": 4096,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 16384,
"temperature": 0.2,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}

ornith-fast​

  • Use Case: Quick Chat, FAQ, Short Answers
  • เหตุผล: ต้องการเร็ว reasoning_effort: minimal คิดนิดหน่อยแล้วตอบ
  • temperature: 0.8 ตอบเป็นธรรมชาติ
  • thinking_token_budget: 512 จำกัด thinking ให้สั้นมาก
  • presence_penalty: 0.0 เพราะ reasoning_effort: minimal คิดสั้น ๆ อยู่แล้ว ไม่ต้องการ penalty เพิ่ม
{
"min_p": 0,
"top_k": 20,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "minimal",
"thinking_token_budget": 512,
"include_reasoning": false,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 4096,
"temperature": 0.8,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": false
}

ornith-creative​

  • Use Case: Brainstorming, Content Writing, Blog Draft, Research Agent
  • เหตุผล: แตกไอเดีย reasoning_effort: medium พอให้คิดบางส่วนแล้วปล่อยให้ความคิดสร้างสรรค์ทำงาน
  • temperature: 0.9 สูงสุดในทุก profiles
  • top_k: 40 กว้างกว่า default (20) เพื่อให้มีทางเลือกคำหลากหลายขึ้น
  • presence_penalty: 0.3 ป้องกันคำซ้ำ
{
"min_p": 0,
"top_k": 40,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "medium",
"thinking_token_budget": 2048,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 12288,
"temperature": 0.9,
"presence_penalty": 0.3,
"repetition_penalty": 1,
"parallel_tool_calls": true
}

ornith-structured​

  • Use Case: Classification, Extraction, JSON Output, Commit Message, Formatting
  • เหตุผล: งานเล็ก ๆ ต้องการ output คงที่ reasoning_effort: none ปิด thinking ทั้งหมด
  • temperature: 0.1 + top_p: 0.20 + top_k: 5 deterministic สูงสุด จำกัดทางเลือกน้อยสุด
{
"min_p": 0,
"top_k": 5,
"top_p": 0.20,
"extra_body": {
"reasoning_effort": "none",
"include_reasoning": false,
"chat_template_kwargs": {
"enable_thinking": false
}
},
"guardrails": [],
"max_tokens": 4096,
"temperature": 0.1,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": false
}

Note: ornith-structured เป็น profile เดียวที่ใช้ reasoning_effort: none และ enable_thinking: false พร้อมกัน — ปิด reasoning ทั้งสองชั้นเพื่อความเร็วสูงสุด

เปรียบเทียบกับ Qwen 10 Profiles​

จุดต่างQwen3.6 (เดิม)Ornith-1.0 (ใหม่)
Reasoning controlenable_thinking: true|false (binary)reasoning_effort 7 levels
Thinking budgetไม่มี คุมไม่ได้thinking_token_budget จำกัดต่อ request
Include reasoningใช้ preserve_thinkinginclude_reasoning (vLLM native)
Coding focusGeneral purposeเน้น agentic coding (Terminal-Bench, SWE-Bench)
Profile ใหม่—ornith-terminal (เฉพาะ CLI/shell)
Profile ที่ตัดqwen-trading (seed=42)ตัดออก เพราะ Ornith ไม่ได้เน้น financial
Default reasoningthinking onreasoning_effort: medium
Max tokens สำหรับ reasoning4,096–16,3844,096–16,384 (แต่ model card แนะนำ ≥ 6,500 สำหรับ reasoning mode)

ทำไมตัด qwen-trading และเพิ่ม ornith-terminal​

  • ตัด qwen-trading: Ornith ไม่ได้ train มาสำหรับ financial tasks โดยเฉพาะ และ seed=42 มีประโยชน์จำกัดเพราะ vLLM ไม่รับประกัน reproducibility 100% ในทุก scenario
  • เพิ่ม ornith-terminal: Ornith ทำ Terminal-Bench 2.1 ได้ 64.2 (สูงกว่า Qwen 12 คะแนน) profile นี้ใช้ประโยชน์จากจุดแข็งโดยตรง — สั่ง shell ผ่าน agent workflow

หมายเหตุสำหรับ Ornith-1.0-35B​

จากประสบการณ์กับ Ornith ในการทดสอบจริง + ข้อมูลจาก model card:

  • top_k=20 มาจาก official model card (Transformers example) เป็นค่าที่ authors แนะนำโดยตรง
  • temperature=0.6 เป็นค่า default ที่ authors ใช้ในทุกตัวอย่าง (chat, tool calling, ClawEval) — เป็น sweet spot สำหรับงานทั่วไป
  • Benchmark eval ใช้ temperature=1.0 เพื่อวัดศักยภาพสูงสุด แต่ production ควรใช้ค่าต่ำกว่าเพื่อความสม่ำเสมอ
  • reasoning_effort: high เป็น sweet spot สำหรับ coding และ debug — คิดละเอียดพอ ไม่ช้าเกินไป
  • reasoning_effort: xhigh ใช้เฉพาะงานวิเคราะห์เชิงระบบ เพราะ thinking ยาวมาก (บางครั้ง > 8,000 tokens) ควรใช้ top_k=-1 (ปิดการกรอง) เพื่อให้ explore ได้เต็มที่
  • thinking_token_budget มีประโยชน์มากสำหรับ agent ที่เรียกหลายรอบ — จำกัดไม่ให้แต่ละรอบคิดนานเกินไป
  • max_tokens ≥ 6500 ตามที่ model card แนะนำ ถ้าใช้ reasoning effort ≥ high
  • presence_penalty เกิน 0.5 มีผลข้างเคียงมากกว่าประโยชน์ — model card ไม่ได้ระบุค่านี้ จึงใช้ 0.0 เป็น default และเพิ่มเฉพาะ profiles ที่ต้องการความหลากหลาย
  • repetition_penalty ควรคงไว้ที่ 1.0 (no-op) เว้นแต่เจอปัญหาคำซ้ำจริง — model card ไม่ได่ระบุค่านี้
  • temperature: 0.0 กับ reasoning_effort: none เหมาะสำหรับ structured output เท่านั้น — ใช้แชทแล้วคำตอบแข็ง ๆ ไม่เป็นธรรมชาติ

Profile ที่จะถูกเรียกใช้จริงบ่อยสุด​

จากลักษณะงานที่ใช้ Ornith (agentic coding เป็นหลัก):

Profileสัดส่วนการใช้งาน
ornith-chat50%
ornith-coder20%
ornith-agent10%
ornith-terminal10%
อื่น ๆ10%

Routing ใน Hermes:

default_model: ornith-chat

Routing คร่าว ๆ​

Python / TypeScript / SQL / Docker
→ ornith-coder

MCP / Agent Workflow / Tool Calling
→ ornith-agent

Shell / CLI / Terminal tasks
→ ornith-terminal

อ่าน repo ใหญ่ / RAG
→ ornith-longctx

Architecture Review / System Design
→ ornith-deep

Debug Log / RCA / Stacktrace
→ ornith-debug

อ่านข่าว + web search + แชททั่วไป
→ ornith-chat

Brainstorm Blog / Content
→ ornith-creative

Quick FAQ / Short Answer
→ ornith-fast

JSON Output / Classification / Commit Msg
→ ornith-structured

Note: ทำไม ornith-chat ใช้ 50% ไม่ใช่ 60% เหมือน Qwen — เพราะ Ornith เน้น coding/agent มากกว่า general chat สัดส่วน coding + terminal เลยสูงกว่า (30% vs 15% ในชุด Qwen)

LiteLLM Proxy Config​

ตัวอย่าง LiteLLM config.yaml สำหรับ Ornith profiles:

model_list:
- model_name: ornith-chat
litellm_params:
model: openai/ornith-35b-nvfp4
api_base: http://10.0.0.246:8000/v1
api_key: EMPTY
min_p: 0
top_k: 20
top_p: 0.95
temperature: 0.6
max_tokens: 8192
presence_penalty: 0.0
repetition_penalty: 1
parallel_tool_calls: true
extra_body:
reasoning_effort: medium
include_reasoning: true
chat_template_kwargs:
enable_thinking: true

- model_name: ornith-coder
litellm_params:
model: openai/ornith-35b-nvfp4
api_base: http://10.0.0.246:8000/v1
api_key: EMPTY
min_p: 0
top_k: 20
top_p: 0.95
temperature: 0.2
max_tokens: 8192
presence_penalty: 0.0
repetition_penalty: 1
parallel_tool_calls: true
extra_body:
reasoning_effort: high
thinking_token_budget: 4096
include_reasoning: true
chat_template_kwargs:
enable_thinking: true

# ... เพิ่ม profiles อื่น ๆ ในรูปแบบเดียวกัน

Security: Private IP — api_base อ้างอิง http://10.0.0.246:8000 ซึ่งเป็น private IP ภายใน LAN ไม่ได้เปิดให้ภายนอก หาก deploy ในสภาพแวดล้อมอื่น ควรวาง vLLM หลัง reverse proxy ที่มี authentication

Conclusion​

การใช้ reasoning_effort 7 levels ของ Ornith ทำให้คุมพฤติกรรมได้ละเอียดกว่า enable_thinking แบบ binary ของ Qwen — แทนที่จะมีแค่ "คิด/ไม่คิด" สามารถเลือกระดับความลึกได้ตามงาน

ชุด 10 profiles บน Ornith-1.0-35B-NVFP4 เน้น agentic coding มากกว่าชุด Qwen เดิม — เพิ่ม ornith-terminal สำหรับ CLI/shell (จุดแข็งของ Ornith บน Terminal-Bench) และตัด qwen-trading ที่ไม่ใช่ use case หลัก

thinking_token_budget เป็นอีกหนึ่งพารามิเตอร์ที่มีประโยชน์มากสำหรับ agent workflow — จำกัดไม่ให้แต่ละรอบคิดนานเกินไป โดยเฉพาะเมื่อ agent ต้องเรียกหลายรอบต่อหนึ่ง task

References​

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...