Virtual Models บน LiteLLM Proxy: Ornith-1.0-35B 10 profiles ใช้ reasoning_effort คุมพฤติกรรม
สารบัญ
- TL;DR
- ตาราง 10 Virtual Profiles
- ทำไม
reasoning_effortสำคัญกว่าenable_thinking - Official Sampling Parameters จาก Model Card
- รายละเอียดแต่ละ Profile
ornith-chat⭐ Defaultornith-coderornith-agentornith-deepornith-debugornith-terminalornith-longctxornith-fastornith-creativeornith-structured- เปรียบเทียบกับ Qwen 10 Profiles
- ทำไมตัด
qwen-tradingและเพิ่มornith-terminal - หมายเหตุสำหรับ
Ornith-1.0-35B - Profile ที่จะถูกเรียกใช้จริงบ่อยสุด
- Routing คร่าว ๆ
- LiteLLM Proxy Config
- Conclusion
- References
หลังจาก deploy Ornith-1.0-35B-NVFP4 บน DGX Spark สำเร็จแล้ว (ดูรายละเอียดใน บทความก่อนหน้า) ขั้นตอนต่อไปคือสร้าง Virtual Models ผ่าน LiteLLM Proxy เหมือนที่เคยทำกับ Qwen3.6
ความต่างสำคัญ: Ornith มี reasoning_effort 7 levels (none/minimal/low/medium/high/xhigh/max) แทนที่แค่ enable_thinking: true|false แบบ Qwen ทำให้คุมความลึกของ reasoning ได้ละเอียดกว่า และมี thinking_token_budget สำหรับจำกัดจำนวน thinking tokens ต่อ request
TL;DR
- ใช้
Ornith-1.0-35B-NVFP41 ตัว backend สร้าง 10 profiles ผ่าน LiteLLM Proxy - แต่ละ profile ตั้งค่า
reasoning_effort,temperature,top_p,thinking_token_budget,max_tokensต่างกันตามงาน ornith-chatเป็น default ใช้บ่อยสุด 50% รองด้วยornith-coder20%,ornith-agent10%,ornith-terminal10% และอื่น ๆ 10%- เน้น profiles สำหรับ agentic coding มากกว่า Qwen เพราะ Ornith ถนัด Terminal-Bench และ SWE-Bench
reasoning_effortเป็นพารามิเตอร์หลักในการคุมพฤติกรรม แทนที่enable_thinkingแบบ binary
ตาราง 10 Virtual Profiles
| Profile | Reasoning Effort | Top-K | Top-P | Temp | Thinking Budget | Max Tokens | Use Case |
|---|---|---|---|---|---|---|---|
ornith-chat ⭐ Default | medium | 20 | 0.95 | 0.6 | — | 8,192 | General Assistant, Web Search, Daily Chat |
ornith-coder | high | 20 | 0.95 | 0.2 | 4,096 | 8,192 | Python, TypeScript, SQL, Docker, DevOps |
ornith-agent | low | 20 | 0.90 | 0.1 | 2,048 | 8,192 | Deterministic Tool Calling, Workflow Execution |
ornith-deep | xhigh | -1 | 0.95 | 0.7 | 8,192 | 16,384 | Architecture Review, System Design, Research |
ornith-debug | high | 20 | 0.95 | 0.3 | 4,096 | 16,384 | Debugging, RCA, Log Analysis, Stacktrace |
ornith-terminal | high | 20 | 0.90 | 0.1 | 4,096 | 8,192 | CLI Agent, Shell Commands, Terminal-Bench Tasks |
ornith-longctx | high | 20 | 0.95 | 0.2 | 4,096 | 16,384 | RAG, Codebase Analysis, Documentation |
ornith-fast | minimal | 20 | 0.95 | 0.8 | 512 | 4,096 | Quick Chat, FAQ, Short Answers |
ornith-creative | medium | 40 | 0.95 | 0.9 | 2,048 | 12,288 | Brainstorming, Content Writing, Blog Draft |
ornith-structured | none | 5 | 0.20 | 0.1 | — | 4,096 | Classification, Extraction, JSON Output, Commit Msg |
ทำไม reasoning_effort สำคัญกว่า enable_thinking
Qwen3.6 ใช้ chat_template_kwargs.enable_thinking: true|false เป็นสวิตช์ on/off — มีแค่ 2 สถานะ
Ornith มี reasoning_effort ถึง 7 levels:
| Level | พฤติกรรม | เหมาะกับ |
|---|---|---|
none | ไม่คิดเลย ตอบทันที | Classification, Extraction, Structured Output |
minimal | คิดนิดหน่อย เร็วมาก | Quick Chat, FAQ |
low | คิดสั้น ๆ | Agent Tool Calling (deterministic) |
medium | คิดพอประมาณ | General Chat, Creative |
high | คิดละเอียด | Coding, Debug, Long Context |
xhigh | คิดลึกมาก | Architecture, System Design |
max | คิดเต็มกำลัง (เฉพาะ DeepSeek V4) | — ไม่ใช้กับ Ornith |
แต่ละ level มี trade-off ระหว่างความเร็วกับความลึกของคำตอบ — none เร็วสุดแต่อาจตอบผิดงานซับซ้อน, xhigh ช้าแต่แม่นยำสำหรับงานวิเคราะห์
Note:
maxเป็นค่าเฉพาะของ DeepSeek V4 series ไม่ใช่ส่วนหนึ่งของ OpenAI API standard — ใช้กับ Ornith ไม่ได้ผลต่างจากxhigh
Official Sampling Parameters จาก Model Card
Model card ของ Ornith-1.0-35B ระบุค่า sampling ที่แนะนำไว้ชัดเจน:
| บริบท | temperature | top_p | top_k | ที่มา |
|---|---|---|---|---|
| Transformers example | 0.6 | 0.95 | 20 | Model card Quickstart |
| Chat Completions API | 0.6 | 0.95 | — | Model card API example |
| Tool calling | 0.6 | 0.95 | — | Model card agentic example |
| Terminal-Bench 2.1 | 1.0 | 1.0 | — | Benchmark eval settings |
| SWE-Bench Verified | 1.0 | 0.95 | — | Benchmark eval settings |
| SWE Atlas QnA | 1.0 | 0.95 | — | Benchmark eval settings |
| ClawEval | 0.6 | — | — | Agentic code benchmark |
สิ่งที่เรียนรู้จาก model card:
top_k=20มาจาก official model card ไม่ใช่ copy จาก Qwen — เป็นค่าที่ authors แนะนำโดยตรงtemperature=0.6เป็นค่า default ที่ authors ใช้ในทุกตัวอย่าง (chat, tool calling, ClawEval)- Benchmark eval ใช้
temperature=1.0เพื่อให้โมเดล explore ทางเลือกได้กว้าง แต่ production ใช้ค่าต่ำกว่าเพื่อความ deterministic - ไม่มีการระบุ
repetition_penaltyหรือpresence_penaltyใน model card — ค่า default ของ vLLM คือ1.0และ0.0ตามลำดับ
Note: ทำไม production ใช้ temp ต่ำกว่า benchmark — Benchmark ใช้
temp=1.0เพื่อวัดศักยภาพสูงสุดของโมเดล (best case) แต่ใน production เราต้องการความสม่ำเสมอและความปลอดภัย (โดยเฉพาะ shell commands) จึงใช้ temp ต่ำกว่า เป็น trade-off ระหว่าง exploration กับ determinism
รายละเอียดแต่ละ Profile
ornith-chat ⭐ Default
- Use Case: General Assistant, Web Search, News Summary, Daily Chat, Tool Calling
- เหตุผล: เป็น Primary Model / Router
reasoning_effort: mediumสมดุลระหว่างความเร็วกับความลึก temperature: 0.6สมดุลระหว่างความเป็นธรรมชาติกับความแม่นยำmax_tokens: 8192เผื่อพื้นที่ reasoning + คำตอบ
{
"min_p": 0,
"top_k": 20,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "medium",
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 8192,
"temperature": 0.6,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}
ornith-coder
- Use Case: Python, TypeScript, SQL, Docker, MCP, DevOps
- เหตุผล: Coding ต้องการความถูกต้องสูง
reasoning_effort: highให้คิดละเอียดก่อนเขียนโค้ด temperature: 0.2ลด hallucinationthinking_token_budget: 4096จำกัด thinking ไม่ให้ยาวเกินไป เพราะ code task มักมีคำตอบชัดเจน
{
"min_p": 0,
"top_k": 20,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "high",
"thinking_token_budget": 4096,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 8192,
"temperature": 0.2,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}
ornith-agent
- Use Case: Deterministic Tool Calling, Workflow Execution, Multi-Step Tasks
- เหตุผล: Agent ต้องการ deterministic สูงสุด
reasoning_effort: lowคิดสั้น ๆ ตัดสินใจเร็ว temperature: 0.1เพื่อให้ parser อ่าน tool call ได้แม่นยำthinking_token_budget: 2048จำกัด thinking ให้กระชับ เพราะ agent มักเรียกหลายรอบ
{
"min_p": 0,
"top_k": 20,
"top_p": 0.90,
"extra_body": {
"reasoning_effort": "low",
"thinking_token_budget": 2048,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 8192,
"temperature": 0.1,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}
ornith-deep
- Use Case: Architecture Review, Research, Planning, System Design
- เหตุผล: งานวิเคราะห์เชิงระบบต้องการความลึกสูงสุด
reasoning_effort: xhigh temperature: 0.7ให้มีความยืดหยุ่นในการเสนอทางเลือกtop_k: -1(ปิดการกรอง) เพราะxhighreasoning ต้อง explore หลายทาง การจำกัด top_k อาจตัดทางเลือกที่จำเป็นthinking_token_budget: 8192ให้พื้นที่คิดกว้าง ๆmax_tokens: 16384เผื่อคำตอบยาว
top_k=-1 สำหรับ ornith-deeptop_k คัดเฉพาะ k tokens ที่มี probability สูงสุดมาพิจารณาในแต่ละ step ค่า -1 หมายถึง ปิดการกรอง ให้พิจารณาทุก token ใน vocabulary
- Reasoning chain ต้อง explore — แต่ละ step โมเดลอาจต้องเลือก token ที่ probability ไม่สูงสุด แต่เป็นทางเลือกที่นำไปสู่เส้นทางคิดที่ถูกต้อง เช่น ตอนพิจารณา "อาจเป็นเพราะ..." โมเดลอาจต้องเลือกคำที่อยู่อันดับ 30–50 ใน vocabulary ไม่ใช่ 20 แรก
top_k=20ตัด long tail ของการคิด — งานทั่วไป (chat, coding) คำตอบอยู่ใน 20 ตัวแรกอยู่แล้ว แต่งานวิเคราะห์เชิงระบบ (architecture, system design) บางครั้งต้อง explore ทางเลือกที่ไม่เด่นtop_p=0.95ยังเป็นตัวคุมอยู่ — แม้ปิด top_k แต่ nucleus sampling ที่top_p=0.95ยังกรองเฉพาะ tokens ที่ probability รวมกันถึง 95% จึงไม่ได้สุ่มอย่างไร้ขอบเขต- Trade-off —
top_k=-1อาจทำให้คำตอบหลากหลายขึ้นแต่ช้าลงเล็กน้อย แต่สำหรับornith-deepที่เน้นความลึกมากกว่าความเร็ว การ trade-off นี้คุ้ม
สรุป: top_k=20 เหมาะกับงานที่มีคำตอบค่อนข้างชัด (chat, coding) ส่วน top_k=-1 เหมาะกับงานที่ต้อง explore ทางเลือกกว้าง (deep reasoning, architecture) โดยมี top_p=0.95 เป็นตัวคุมความเสี่ยงอยู่แล้ว
{
"min_p": 0,
"top_k": -1,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "xhigh",
"thinking_token_budget": 8192,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 16384,
"temperature": 0.7,
"presence_penalty": 0.2,
"repetition_penalty": 1,
"parallel_tool_calls": true
}
ornith-debug
- Use Case: Debugging, RCA, Log Analysis, Incident Investigation
- เหตุผล: หาสาเหตุต้องคิดละเอียด
reasoning_effort: highแต่temperature: 0.3รักษาความแม่นยำ max_tokens: 16384เผื่ออ่าน log ยาว ๆ
{
"min_p": 0,
"top_k": 20,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "high",
"thinking_token_budget": 4096,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 16384,
"temperature": 0.3,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}
ornith-terminal
- Use Case: CLI Agent, Shell Commands, Terminal-Bench Tasks
- เหตุผล: Ornith ถนัด Terminal-Bench (64.2 vs Qwen 52.5) profile นี้เน้นการสั่ง shell โดยตรง
reasoning_effort: highคิดก่อนสั่งคำสั่งอันตรายtemperature: 0.1deterministic สูงสุด เพราะ shell command ผิดนิดเดียวอันตรายมากtop_p: 0.90แคบกว่า chat เพื่อลดความเสี่ยง
{
"min_p": 0,
"top_k": 20,
"top_p": 0.90,
"extra_body": {
"reasoning_effort": "high",
"thinking_token_budget": 4096,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 8192,
"temperature": 0.1,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}
ornith-longctx
- Use Case: RAG, Codebase Analysis, Documentation, Long Reports
- เหตุผล: ต้องการ faithfulness สูงสุดกับ context ยาว
reasoning_effort: highช่วยวิเคราะห์ข้อมูลใน context temperature: 0.2ไม่ให้สร้างข้อมูลเอง
{
"min_p": 0,
"top_k": 20,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "high",
"thinking_token_budget": 4096,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 16384,
"temperature": 0.2,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": true
}
ornith-fast
- Use Case: Quick Chat, FAQ, Short Answers
- เหตุผล: ต้องการเร็ว
reasoning_effort: minimalคิดนิดหน่อยแล้วตอบ temperature: 0.8ตอบเป็นธรรมชาติthinking_token_budget: 512จำกัด thinking ให้สั้นมากpresence_penalty: 0.0เพราะreasoning_effort: minimalคิดสั้น ๆ อยู่แล้ว ไม่ต้องการ penalty เพิ่ม
{
"min_p": 0,
"top_k": 20,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "minimal",
"thinking_token_budget": 512,
"include_reasoning": false,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 4096,
"temperature": 0.8,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": false
}
ornith-creative
- Use Case: Brainstorming, Content Writing, Blog Draft, Research Agent
- เหตุผล: แตกไอเดีย
reasoning_effort: mediumพอให้คิดบางส่วนแล้วปล่อยให้ความคิดสร้างสรรค์ทำงาน temperature: 0.9สูงสุดในทุก profilestop_k: 40กว้างกว่า default (20) เพื่อให้มีทางเลือกคำหลากหลายขึ้นpresence_penalty: 0.3ป้องกันคำซ้ำ
{
"min_p": 0,
"top_k": 40,
"top_p": 0.95,
"extra_body": {
"reasoning_effort": "medium",
"thinking_token_budget": 2048,
"include_reasoning": true,
"chat_template_kwargs": {
"enable_thinking": true
}
},
"guardrails": [],
"max_tokens": 12288,
"temperature": 0.9,
"presence_penalty": 0.3,
"repetition_penalty": 1,
"parallel_tool_calls": true
}
ornith-structured
- Use Case: Classification, Extraction, JSON Output, Commit Message, Formatting
- เหตุผล: งานเล็ก ๆ ต้องการ output คงที่
reasoning_effort: noneปิด thinking ทั้งหมด temperature: 0.1+top_p: 0.20+top_k: 5deterministic สูงสุด จำกัดทางเลือกน้อยสุด
{
"min_p": 0,
"top_k": 5,
"top_p": 0.20,
"extra_body": {
"reasoning_effort": "none",
"include_reasoning": false,
"chat_template_kwargs": {
"enable_thinking": false
}
},
"guardrails": [],
"max_tokens": 4096,
"temperature": 0.1,
"presence_penalty": 0.0,
"repetition_penalty": 1,
"parallel_tool_calls": false
}
Note:
ornith-structuredเป็น profile เดียวที่ใช้reasoning_effort: noneและenable_thinking: falseพร้อมกัน — ปิด reasoning ทั้งสองชั้นเพื่อความเร็วสูงสุด
เปรียบเทียบกับ Qwen 10 Profiles
| จุดต่าง | Qwen3.6 (เดิม) | Ornith-1.0 (ใหม่) |
|---|---|---|
| Reasoning control | enable_thinking: true|false (binary) | reasoning_effort 7 levels |
| Thinking budget | ไม่มี คุมไม่ได้ | thinking_token_budget จำกัดต่อ request |
| Include reasoning | ใช้ preserve_thinking | include_reasoning (vLLM native) |
| Coding focus | General purpose | เน้น agentic coding (Terminal-Bench, SWE-Bench) |
| Profile ใหม่ | — | ornith-terminal (เฉพาะ CLI/shell) |
| Profile ที่ตัด | qwen-trading (seed=42) | ตัดออก เพราะ Ornith ไม่ได้เน้น financial |
| Default reasoning | thinking on | reasoning_effort: medium |
| Max tokens สำหรับ reasoning | 4,096–16,384 | 4,096–16,384 (แต่ model card แนะนำ ≥ 6,500 สำหรับ reasoning mode) |
ทำไมตัด qwen-trading และเพิ่ม ornith-terminal
- ตัด
qwen-trading: Ornith ไม่ได้ train มาสำหรับ financial tasks โดยเฉพาะ และseed=42มีประโยชน์จำกัดเพราะ vLLM ไม่รับประกัน reproducibility 100% ในทุก scenario - เพิ่ม
ornith-terminal: Ornith ทำ Terminal-Bench 2.1 ได้ 64.2 (สูงกว่า Qwen 12 คะแนน) profile นี้ใช้ประโยชน์จากจุดแข็งโดยตรง — สั่ง shell ผ่าน agent workflow
หมายเหตุสำหรับ Ornith-1.0-35B
จากประสบการณ์กับ Ornith ในการทดสอบจริง + ข้อมูลจาก model card:
top_k=20มาจาก official model card (Transformers example) เป็นค่าที่ authors แนะนำโดยตรงtemperature=0.6เป็นค่า default ที่ authors ใช้ในทุกตัวอย่าง (chat, tool calling, ClawEval) — เป็น sweet spot สำหรับงานทั่วไป- Benchmark eval ใช้
temperature=1.0เพื่อวัดศักยภาพสูงสุด แต่ production ควรใช้ค่าต่ำกว่าเพื่อความสม่ำเสมอ reasoning_effort: highเป็น sweet spot สำหรับ coding และ debug — คิดละเอียดพอ ไม่ช้าเกินไปreasoning_effort: xhighใช้เฉพาะงานวิเคราะห์เชิงระบบ เพราะ thinking ยาวมาก (บางครั้ง > 8,000 tokens) ควรใช้top_k=-1(ปิดการกรอง) เพื่อให้ explore ได้เต็มที่thinking_token_budgetมีประโยชน์มากสำหรับ agent ที่เรียกหลายรอบ — จำกัดไม่ให้แต่ละรอบคิดนานเกินไปmax_tokens ≥ 6500ตามที่ model card แนะนำ ถ้าใช้ reasoning effort ≥highpresence_penaltyเกิน0.5มีผลข้างเคียงมากกว่าประโยชน์ — model card ไม่ได้ระบุค่านี้ จึงใช้0.0เป็น default และเพิ่มเฉพาะ profiles ที่ต้องการความหลากหลายrepetition_penaltyควรคงไว้ที่1.0(no-op) เว้นแต่เจอปัญหาคำซ้ำจริง — model card ไม่ได่ระบุค่านี้temperature: 0.0กับreasoning_effort: noneเหมาะสำหรับ structured output เท่านั้น — ใช้แชทแล้วคำตอบแข็ง ๆ ไม่เป็นธรรมชาติ
Profile ที่จะถูกเรียกใช้จริงบ่อยสุด
จากลักษณะงานที่ใช้ Ornith (agentic coding เป็นหลัก):
| Profile | สัดส่วนการใช้งาน |
|---|---|
ornith-chat | 50% |
ornith-coder | 20% |
ornith-agent | 10% |
ornith-terminal | 10% |
| อื่น ๆ | 10% |
Routing ใน Hermes:
default_model: ornith-chat
Routing คร่าว ๆ
Python / TypeScript / SQL / Docker
→ ornith-coder
MCP / Agent Workflow / Tool Calling
→ ornith-agent
Shell / CLI / Terminal tasks
→ ornith-terminal
อ่าน repo ใหญ่ / RAG
→ ornith-longctx
Architecture Review / System Design
→ ornith-deep
Debug Log / RCA / Stacktrace
→ ornith-debug
อ่านข่าว + web search + แชททั่วไป
→ ornith-chat
Brainstorm Blog / Content
→ ornith-creative
Quick FAQ / Short Answer
→ ornith-fast
JSON Output / Classification / Commit Msg
→ ornith-structured
Note: ทำไม
ornith-chatใช้ 50% ไม่ใช่ 60% เหมือน Qwen — เพราะ Ornith เน้น coding/agent มากกว่า general chat สัดส่วน coding + terminal เลยสูงกว่า (30% vs 15% ในชุด Qwen)
LiteLLM Proxy Config
ตัวอย่าง LiteLLM config.yaml สำหรับ Ornith profiles:
model_list:
- model_name: ornith-chat
litellm_params:
model: openai/ornith-35b-nvfp4
api_base: http://10.0.0.246:8000/v1
api_key: EMPTY
min_p: 0
top_k: 20
top_p: 0.95
temperature: 0.6
max_tokens: 8192
presence_penalty: 0.0
repetition_penalty: 1
parallel_tool_calls: true
extra_body:
reasoning_effort: medium
include_reasoning: true
chat_template_kwargs:
enable_thinking: true
- model_name: ornith-coder
litellm_params:
model: openai/ornith-35b-nvfp4
api_base: http://10.0.0.246:8000/v1
api_key: EMPTY
min_p: 0
top_k: 20
top_p: 0.95
temperature: 0.2
max_tokens: 8192
presence_penalty: 0.0
repetition_penalty: 1
parallel_tool_calls: true
extra_body:
reasoning_effort: high
thinking_token_budget: 4096
include_reasoning: true
chat_template_kwargs:
enable_thinking: true
# ... เพิ่ม profiles อื่น ๆ ในรูปแบบเดียวกัน
Security: Private IP —
api_baseอ้างอิงhttp://10.0.0.246:8000ซึ่งเป็น private IP ภายใน LAN ไม่ได้เปิดให้ภายนอก หาก deploy ในสภาพแวดล้อมอื่น ควรวาง vLLM หลัง reverse proxy ที่มี authentication
Conclusion
การใช้ reasoning_effort 7 levels ของ Ornith ทำให้คุมพฤติกรรมได้ละเอียดกว่า enable_thinking แบบ binary ของ Qwen — แทนที่จะมีแค่ "คิด/ไม่คิด" สามารถเลือกระดับความลึกได้ตามงาน
ชุด 10 profiles บน Ornith-1.0-35B-NVFP4 เน้น agentic coding มากกว่าชุด Qwen เดิม — เพิ่ม ornith-terminal สำหรับ CLI/shell (จุดแข็งของ Ornith บน Terminal-Bench) และตัด qwen-trading ที่ไม่ใช่ use case หลัก
thinking_token_budget เป็นอีกหนึ่งพารามิเตอร์ที่มีประโยชน์มากสำหรับ agent workflow — จำกัดไม่ให้แต่ละรอบคิดนานเกินไป โดยเฉพาะเมื่อ agent ต้องเรียกหลายรอบต่อหนึ่ง task
References
- Ornith-1.0 official blog post — Self-scaffolding RL, defense layers, benchmark methodology
- sakamakismile/Ornith-1.0-35B-NVFP4 (Hugging Face) — Pre-made NVFP4 quant (MIT)
- deepreinforce-ai/Ornith-1.0-35B (Hugging Face) — Base BF16 model (MIT)
- LiteLLM Proxy Documentation — Virtual models, alias, sampling params
- vLLM Chat Completion API —
reasoning_effort,thinking_token_budget,include_reasoning - บทความก่อนหน้า: Ornith-1.0-35B-NVFP4 บน DGX Spark — รายละเอียดการ deploy และปรับ vLLM flags
- บทความก่อนหน้า: Virtual Models บน LiteLLM Proxy (Qwen) — ชุด 10 profiles สำหรับ Qwen3.6
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee