Ornith-1.0-35B-NVFP4 บน DGX Spark: บทเรียนจากการ deploy และปรับ vLLM flags
สารบัญ
- TL;DR
- ทำไมถึงอยากเปลี่ยน
- Ornith-1.0 คืออะไร? (สรุปจาก Official Announcement)
- Self-Scaffolding Training Framework
- Defense Layers ต้าน Reward Hacking
- Asynchronous RL Training (Pipeline-RL)
- Benchmark Methodology
- ปัญหาแรก: FP8 vs NVFP4 บน GB10
- Recipe แรก: ลอก v6 มาแล้ว break
- Recipe ใหม่: ลบ enforce-eager ออก
- ทดสอบจริง: 3 prompts
- Prompt 1: Fibonacci iterative
- Prompt 2: async vs threading
- Prompt 3: Debug bug
- เปรียบเทียบกับ v6 (Qwen3.6-35B-A3B-NVFP4)
- Recipe files
- บทสรุป
- References
- ภาคผนวก: ตรวจสอบโมเดลผ่าน vLLM API
- Parameters ที่ใช้ได้กับ
/v1/chat/completions
บันทึก 26 มิถุนายน 2569 — เรื่องของการเปลี่ยน Qwen3.6-35B-A3B ไปเป็น Ornith-1.0-35B แล้วเจออะไรหลายอย่าง
หลังจากรัน Qwen3.6-35B-A3B-NVFP4 บน DGX Spark มาได้พักใหญ่ วันนี้ลองเปลี่ยนไปเป็น Ornith-1.0-35B-NVFP4 ดู — เป็นโมเดลที่เพิ่งออก (MIT license) เน้น agentic coding โดยเฉพาะ และ Terminal-Bench 2.1 สูงกว่า Qwen3.6-35B ถึง 11.7 คะแนน แต่ก็เจอบทเรียนหลายอย่างระหว่างทาง
TL;DR
- เปลี่ยนจาก Qwen3.6-35B-A3B-NVFP4 ไปเป็น Ornith-1.0-35B-NVFP4 บน DGX Spark (GB10) เพื่อทดลองโมเดล agentic coding ใหม่
- บทเรียนสำคัญ: อย่าใส่
--enforce-eagerโดยไม่อ่าน model card — flag นี้ทำให้ throughput ตก ~5× จาก 65.8 tok/s เหลือ ~20 tok/s - ใช้ pre-made NVFP4 quant จาก community (sakamakismile) แทนการ quantize เอง ประหยัดเวลาไป ~2 ชั่วโมง
- ผลลัพธ์หลังแก้ recipe: single-stream 65.8 tok/s (เร็วกว่า v6 ที่ 47.9 tok/s ประมาณ 37%)
- โมเดลรองรับ reasoning mode และ tool calling ผ่าน vLLM ได้โดยตรง เชื่อมกับ LiteLLM ผ่าน
extra_bodyสำหรับ vLLM-specific params
ทำไมถึงอยากเปลี่ยน
ตอนแรกตั้งใจจะลอง Qwen3.6-27B ที่ NVIDIA forum แนะนำ แต่ระหว่างทางไปเจอ Ornith-1.0 จาก DeepReinforce — เป็น reasoning model ที่ออกแบบมาสำหรับ coding agent โดยเฉพาะ มีอะไรน่าสนใจหลายอย่าง:
| Benchmark | Qwen3.6-35B-A3B | Ornith-1.0-35B | Gap |
|---|---|---|---|
| Terminal-Bench 2.1 | 52.5 | 64.2 | +11.7 |
| SWE-Bench Verified | 73.4 | 75.6 | +2.2 |
| SWE Atlas QnA | 15.5 | 37.1 | +21.6 |
| SWE Atlas RF | 11.4 | 29.7 | +18.3 |
Note: เรื่อง SWE Atlas — เป็น benchmark ที่วัดความสามารถในการ navigate codebase ขนาดใหญ่ ซึ่ง Ornith ชนะขาดมาก น่าจะเพราะ train มาแบบ agentic coding โดยเฉพาะ
ตัวเลขพวกนี้ดูดี แต่ก่อนจะ deploy จริงผมก็เจอเรื่องแรกเลย — โมเดลมีแค่ FP8 ไม่มี NVFP4 ใน HuggingFace ตอนแรก
Ornith-1.0 คืออะไร? (สรุปจาก Official Announcement)
จาก official blog post ของ DeepReinforce — Ornith-1.0 เป็น "self-improving family of open-source models specially for agentic coding tasks" ครอบคลุมหลายขนาด:
| Variant | Type | Base model | จุดเด่น |
|---|---|---|---|
| 9B Dense | Dense | Gemma 4 | edge-deployable, แซง Gemma 4-31B |
| 31B Dense | Dense | Gemma 4 | |
| 35B MoE | MoE | Qwen 3.5 | แซง Qwen 3.5-397B บน Terminal-Bench (64.2 vs 53.5) |
| 397B MoE | MoE | Qwen 3.5 | frontier-scale, เทียบ Claude Opus 4.7 |
Self-Scaffolding Training Framework
นวัตกรรมหลักของ Ornith คือ self-scaffolding — แทนที่จะใช้ harness ที่มนุษย์ออกแบบมาเหมือนโมเดลอื่น Ornith เรียนรู้ที่จะ สร้าง scaffold (harness) เอง แล้วใช้ scaffold นั้น guide solution rollout:
- Stage 1 — Scaffold proposal: โมเดลเสนอ scaffold ใหม่จาก task + scaffold ก่อนหน้า
- Stage 2 — Solution rollout: โมเดล generate คำตอบโดยใช้ scaffold จาก stage 1
- Reward propagation: reward จาก rollout ส่งกลับไป optimize ทั้งสอง stage
ทำซ้ำไปเรื่อยๆ scaffold จะ evolve ไปสู่รูปแบบที่ให้ reward สูงสุดโดยอัตโนมัติ — ไม่ต้อง hand-engineer harness เอง
Defense Layers ต้าน Reward Hacking
ปัญหาของการให้โมเดลสร้าง scaffold เองคือ reward hacking — โมเดลอาจเขียน scaffold ที่อ่าน test files แล้ว hardcode คำตอบ DeepReinforce ป้องกันด้วย 3 ชั้น:
- Immutable trust boundary — environment, tool surface, test isolation อยู่นอกการควบคุมของโมเดล โมเดลแก้ได้แค่ inner policy scaffold (memory, error-handling, orchestration)
- Deterministic monitor — flag ทันทีถ้าโมเดลอ่าน hidden paths, แก้ verification scripts, หรือเรียก tools นอก sanctioned set → ให้ reward = 0 และ excluded จาก advantage computation
- Frozen LLM judge — เป็น veto layer บน verifier เพราะ intent-level gaming อาจเกิดในกรอบ tools ที่อนุญาต
Asynchronous RL Training (Pipeline-RL)
สำหรับ RL training กับ long rollouts Ornith ใช้ pipeline-RL แก้ปัญหา off-policy — ใช้ staleness weight w(d_t) ที่ downweight tokens ตามอายุ d_t และ drop ทิ้งเมื่อเกิน threshold ทำให้ train ได้โดยไม่ต้องรอ rollout ยาว ๆ จบก่อน
Benchmark Methodology
จาก footnote ของ official post:
- Terminal-Bench 2.1: Harbor/Terminus-2 framework,
temp=1.0, top_p=1.0, 128K context, 4-hour timeout, 32 CPU cores + 48GB RAM, average 5 runs - SWE-Bench Verified/Pro/Multilingual: OpenHands harness,
temp=1.0, top_p=0.95, 256K context - SWE Atlas QnA/RF/TW: mini SWE agent harness,
temp=1.0, top_p=0.95, 128K context, average 5 runs - ใช้ custom chat template เพื่อ consistency ระหว่าง training กับ inference, และ modify Harbor ให้ align กับ vLLM's
reasoning_contentkey
Note: ทำไมถึงเลือก 35B — 397B ใหญ่เกิน DGX Spark (ต้องการ multi-GPU) ส่วน 9B เล็กไปสำหรับงาน coding หนักๆ 35B MoE เป็นจุดหวานที่สุด: แซง Qwen 3.5-397B บน Terminal-Bench ทั้งที่ตัวเล็กกว่า 12 เท่า
ปัญหาแรก: FP8 vs NVFP4 บน GB10
ตอนแรกคิดว่า "FP8 ก็น่าจะเร็วพอ" แต่พอนั่งคำนวณดู:
- NVFP4 (35B): ~17 GB, ~50-60 GB VRAM @ 262K ctx
- FP8 (35B): ~35 GB, ~70-80 GB VRAM @ 262K ctx
- BF16 (35B): ~70 GB, ~100-110 GB VRAM @ 262K ctx
FP8 ใช้ RAM เกือบ 2 เท่าของ NVFP4 แถม throughput ก็คาดว่าช้ากว่า ~20-30% บน Blackwell (NVFP4 มี native tensor cores สำหรับ 4-bit)
โชคดีที่ลอง search ดู — เจอ sakamakismile/Ornith-1.0-35B-NVFP4 ที่เพิ่งอัปโหลดใหม่ๆ ในตอนนั้น — ใช้ llm-compressor quantize มาเป็น NVFP4 ได้ขนาด 21.9 GB
Note: บทเรียนเรื่อง "search ก่อนทำเอง" — ผมเกือบจะเริ่ม quantize Ornith FP8 → NVFP4 เอง (ใช้เวลา ~2 ชม.) แต่ search ดูใน HuggingFace เจอ pre-made quant เลย ประหยัดเวลาไปเยอะ
Recipe แรก: ลอก v6 มาแล้ว break
v6 (Qwen3.6-35B-A3B) ที่ผมรันอยู่ใช้ --enforce-eager ตามคำแนะนำของ vLLM สำหรับ Blackwell (ป้องกัน CUDA graph issue) — ผมเลย copy pattern เดียวกันมาใส่ใน recipe Ornith:
vllm:
args:
- serve
- /models/ornith-35b-nvfp4
- --served-model-name ornith-35b-nvfp4
- --tensor-parallel-size 1
- --max-model-len 262144
- --gpu-memory-utilization 0.85
- --enforce-eager # ← ผมใส่ตาม v6 (ผิด)
- --enable-prefix-caching
- --trust-remote-code
Deploy ได้ แต่พอ benchmark ออกมา — single stream ได้แค่ ~20 tok/s ซึ่งช้ากว่า v6 (47.9 tok/s) เยอะมาก ตอนแรกนึกว่า NVFP4 quant ทำไม่ดี แต่พออ่าน model card ของ sakamakismile อีกที:
"
--enforce-eagercosts ~5× single-stream; the numbers above are with CUDA graphs on."
แปลว่า model card แนะนำให้ปิด enforce-eager เพราะทำให้ throughput ตก 5 เท่า ผมลืมอ่าน serving recipe ของเขาให้ละเอียด เลย copy flag มาจาก v6 ผิดๆ
Recipe ใหม่: ลบ enforce-eager ออก
หลังจากลบ --enforce-eager ออก + ปรับ gpu-memory-utilization จาก 0.85 เป็น 0.90 เพื่อใช้ VRAM ให้เต็มที่ขึ้น:
vllm:
args:
- serve
- /models/ornith-35b-nvfp4
- --served-model-name ornith-35b-nvfp4
- --tensor-parallel-size 1
- --max-model-len 262144
- --gpu-memory-utilization 0.90
- --enable-chunked-prefill
- --enable-prefix-caching
- --enable-auto-tool-choice
- --tool-call-parser qwen3_xml
- --reasoning-parser qwen3
- --trust-remote-code
# ไม่มี --enforce-eager แล้ว
ผลลัพธ์ต่างกันเยอะ:
| Test | v1 (enforce-eager) | v2 (CUDA graphs ON) | Speedup |
|---|---|---|---|
| With reasoning (cold) | ~20 tok/s | 51.35 tok/s | 2.5× |
| No reasoning (warm, 200 tok) | ~20 tok/s | 65.8 tok/s | 3.3× |
CUDA graphs กลับมาแล้ว + vLLM default capture sizes [1, 2, 4, 8, 16, 24, 32] ทำให้ graph warm-up สำหรับ batch sizes ต่างๆ ได้
Note: เรื่อง --disable-custom-all-reduce — ใน model card ของ sakamakismile มี flag นี้อยู่ด้วย ผมตัดสินใจไม่ใส่เพราะ DGX Spark เป็น single GPU (TP=1) ไม่มี inter-GPU all-reduce ให้ optimize flag นี้มีผลแค่ตอน TP ≥ 2
ทดสอบจริง: 3 prompts
หลัง deploy v2 เสร็จ ผมลอง 3 prompts เพื่อเช็ค quality + speed:
Prompt 1: Fibonacci iterative
def fibonacci(n):
if n < 0:
raise ValueError("Input must be a non-negative integer.")
a, b = 0, 1
for _ in range(n):
a, b = b, a + b
return a
- Wall time: 21.07 sec
- Completion tokens: 1,374
- Speed: 65.2 tok/s
- Clean code, handles negative input, O(n) time O(1) space
finish_reason: stop(ไม่ถูกตัด)
Prompt 2: async vs threading
ได้คำตอบยาว 2,543 tokens — ครอบคลุม tables, common pitfalls, Python 3.13 GIL removal, decision flow, code examples
- Wall time: 38.93 sec
- Completion tokens: 2,543
- Speed: 65.3 tok/s
finish_reason: stop
ตัวอย่างจากคำตอบ (ตัดมา):
Choose Threading when: You need true parallelism (CPU-bound work across cores). You're wrapping legacy synchronous code that can't be easily made async. You need a small number of concurrent tasks (10–100).
Choose Async/await when: You're building an async-native application (web server, chat, real-time app). You need to handle thousands of concurrent I/O connections.
Prompt 3: Debug bug
def avg(nums):
total = 0
for n in nums:
total = total + n
return total / len(nums)
print(avg([])) # ZeroDivisionError
Ornith ตอบถูกทันที:
def avg(nums):
if not nums:
return 0.0 # or raise ValueError / return None, depending on your needs
return sum(nums) / len(nums)
- Wall time: 27.38 sec
- Completion tokens: 1,785
- Speed: 65.1 tok/s
คำตอบพวกนี้ใช้ reasoning mode (เห็น chain-of-thought ก่อนคำตอบจริง) — ทำให้คำตอบแม่นยำขึ้นเยอะ
เปรียบเทียบกับ v6 (Qwen3.6-35B-A3B-NVFP4)
| Metric | v6 (Qwen3.6-35B-A3B) | v2 (Ornith-1.0-35B-NVFP4) | Δ |
|---|---|---|---|
| Model size | 17 GB | 21 GB | +24% |
| License | Apache-2.0 | MIT | same freedom |
| Single-stream tok/s | 47.9 | 65.8 | +37% |
| Reasoning mode | No | Yes | new |
| Terminal-Bench 2.1 | 52.5 | 64.2 | +11.7 |
| SWE Atlas QnA | 15.5 | 37.1 | +21.6 |
| Container runtime | vLLM 0.23.1 | vLLM 0.23.1 | same |
ตัวเลข benchmark ของ Ornith เป็น published numbers จาก DeepReinforce + sakamakismile ส่วน throughput วัดจริงบน DGX Spark ด้วย prompt ขนาด 1,374–2,543 tokens
Recipe files
# v1 (failed — enforce-eager)
/home/kongvut/spark-vllm-docker/recipes/Ornith-1.0-35B-NVFP4-v1-eager.yaml
# v2 (current ACTIVE — CUDA graphs ON)
/home/kongvut/spark-vllm-docker/recipes/Ornith-1.0-35B-NVFP4-v2.yaml
ไฟล์ v1 เก็บไว้เป็น historical reference — ถ้าจะ roll back ก็ใช้ได้
บทสรุป
- อ่าน model card ให้ละเอียดก่อน copy flag — ผมเสียเวลา ~30 นาทีกับการ debug throughput ที่ตก ทั้งที่จริงๆ คำตอบอยู่ใน model card ตั้งแต่แรก
- Search HuggingFace ก่อน quantize เอง — community quant ใหม่ๆ ออกบ่อยมาก ผมเกือบจะใช้เวลา 2 ชม. quantize เองทั้งที่มี pre-made ให้ใช้
- /tmp ไม่ใช่ persistent storage — ใช้
HF_HOME=/home/kongvut/.cache/huggingfaceหรือ explicit directory ที่อยู่ใน/home/เสมอ - Hardlinks ช่วยได้ — ใช้ disk เท่ากับไฟล์เดียว แม้มีหลาย paths
- Single GPU = no-op flags —
--disable-custom-all-reduceมีผลแค่ TP ≥ 2 ไม่ต้องใส่ถ้า TP=1
Ornith-1.0-35B-NVFP4 รันบน DGX Spark ได้ดี ทั้งความเร็วและคุณภาพคำตอบ เหมาะกับงาน agentic coding ที่ต้องการ reasoning + tool calling การเปลี่ยนจาก Qwen3.6-35B-A3B มาเป็น Ornith ใช้เวลาไม่นาน แต่บทเรียนเรื่อง --enforce-eager น่าจะช่วยให้คนที่จะ deploy บน Blackwell ครั้งต่อไปไม่ต้องเจอเหมือนผม
References
- Ornith-1.0 official blog post — Self-scaffolding RL, defense layers, benchmark methodology
- deepreinforce-ai/Ornith-1.0-35B — Base BF16 model (MIT)
- sakamakismile/Ornith-1.0-35B-NVFP4 — Pre-made NVFP4 quant via
llm-compressor(MIT) - Ornith-1.0 HuggingFace Collection — ทุก variant ในคอลเลกชั่นเดียว
- vLLM 0.23.1 docs — Engine flags, quantization support
- vLLM sampling_params.py — Engine-level default values สำหรับ
temperature,top_p,top_kฯลฯ - NVIDIA DGX Spark (GB10) specs — 128 GB unified memory, Blackwell sm_121
- llm-compressor — Quantization toolkit with NVFP4 support
- MarkTechPost: DeepReinforce Releases Ornith-1.0 — บทความสรุปการเปิดตัว Ornith-1.0
ภาคผนวก: ตรวจสอบโมเดลผ่าน vLLM API
หลัง deploy v2 เสร็จ ตรวจสอบโมเดลที่รันอยู่ผ่าน GET /v1/models:
curl -s http://10.0.0.246:8000/v1/models | python3 -m json.tool
{
"object": "list",
"data": [
{
"id": "ornith-35b-nvfp4",
"object": "model",
"created": 1782473587,
"owned_by": "vllm",
"root": "sakamakismile/Ornith-1.0-35B-NVFP4",
"parent": null,
"max_model_len": 262144,
"permission": [
{
"id": "modelperm-bd78c055cf4f6570",
"object": "model_permission",
"created": 1782473587,
"allow_create_engine": false,
"allow_sampling": true,
"allow_logprobs": true,
"allow_search_indices": false,
"allow_view": true,
"allow_fine_tuning": false,
"organization": "*",
"group": null,
"is_blocking": false
}
]
}
]
}
จาก response ยืนยันได้ว่า:
id:ornith-35b-nvfp4— ตรงกับ--served-model-nameที่ตั้งไว้ใน recipe v2root:sakamakismile/Ornith-1.0-35B-NVFP4— โมเดลต้นทางจาก HuggingFacemax_model_len:262144— ตรงกับ--max-model-len 262144ที่กำหนดใน recipeowned_by:vllm— รันผ่าน vLLM engineallow_fine_tuning:false— vLLM serve mode ไม่เปิด fine-tuning
Parameters ที่ใช้ได้กับ /v1/chat/completions
ดึงจาก OpenAPI schema ของ vLLM (GET /openapi.json) — เป็น params ที่ส่งได้ใน request body และใช้ร่วมกับ LiteLLM ได้:
# ดู required fields + default values ของทุก param
curl -s http://10.0.0.246:8000/openapi.json \
| python3 -c "import sys,json; d=json.load(sys.stdin); \
s=d['components']['schemas']['ChatCompletionRequest']; \
req=s.get('required',[]); \
[print(f'{k:40s} req={k in req!s:5s} default={v.get(\"default\",\"—\")}') \
for k,v in sorted(s['properties'].items())]"
ผลลัพธ์ที่ได้จะบอกทั้ง required/optional และ default value ของแต่ละ param จาก schema โดยตรง สรุปเป็นตารางด้านล่าง:
| Parameter | Type | Req? | Default | ใช้กับ LiteLLM | หมายเหตุ |
|---|---|---|---|---|---|
messages | array | ✅ required | — | ✅ | OpenAI format |
model | string | optional | — | ✅ | vLLM ใช้ default ถ้ามีโมเดลเดียว |
temperature | number | optional | — ¹ | ✅ | |
top_p | number | optional | — ¹ | ✅ | |
top_k | integer | optional | — ¹ | ✅ via extra_body | vLLM extension |
max_tokens | integer | optional | — | ✅ | deprecated ใน OpenAI ใหม่ แต่ vLLM ยังรองรับ |
max_completion_tokens | integer | optional | — | ✅ | ทดแทน max_tokens |
stop | string/array | optional | [] | ✅ | stop sequences |
stop_token_ids | array<int> | optional | [] | ✅ via extra_body | vLLM extension |
stream | boolean | optional | false | ✅ | streaming response |
stream_options | object | optional | — | ✅ | include_usage etc. |
n | integer | optional | 1 | ✅ | จำนวน completions |
seed | integer | optional | — | ✅ | reproducibility |
frequency_penalty | number | optional | 0.0 | ✅ | |
presence_penalty | number | optional | 0.0 | ✅ | |
repetition_penalty | number | optional | — ¹ | ✅ via extra_body | vLLM extension |
min_p | number | optional | — ¹ | ✅ via extra_body | vLLM extension |
min_tokens | integer | optional | 0 | ✅ via extra_body | vLLM extension |
length_penalty | number | optional | 1.0 | ✅ via extra_body | ใช้กับ beam search |
logprobs | boolean | optional | false | ✅ | |
top_logprobs | integer | optional | 0 | ✅ | |
prompt_logprobs | integer | optional | — | ✅ via extra_body | vLLM extension |
logit_bias | object | optional | — | ✅ | |
response_format | object | optional | — | ✅ | JSON schema / structured output |
tools | array | optional | — | ✅ | function calling |
tool_choice | string/object | optional | none | ✅ | auto, none, หรือ specific tool |
parallel_tool_calls | boolean | optional | true | ✅ via extra_body | |
user | string | optional | — | ✅ | |
reasoning_effort | enum | optional | — | ✅ via extra_body | none, minimal, low, medium, high, xhigh, max — ควบคุม reasoning depth (สังเกตว่า max เป็นค่าเฉพาะของ DeepSeek V4 series) |
thinking_token_budget | integer | optional | — | ✅ via extra_body | จำกัดจำนวน thinking tokens |
include_reasoning | boolean | optional | true | ✅ via extra_body | ส่ง reasoning content กลับมาใน response |
chat_template | string | optional | — | ✅ via extra_body | override Jinja template |
chat_template_kwargs | object | optional | — | ✅ via extra_body | เช่น {"enable_thinking": true} |
add_generation_prompt | boolean | optional | true | ✅ via extra_body | |
continue_final_message | boolean | optional | false | ✅ via extra_body | |
skip_special_tokens | boolean | optional | true | ✅ via extra_body | |
echo | boolean | optional | false | ✅ via extra_body | |
use_beam_search | boolean | optional | false | ✅ via extra_body | |
truncate_prompt_tokens | integer | optional | — | ✅ via extra_body | ตัด prompt ถ้าเกิน limit |
cache_salt | string | optional | — | ✅ via extra_body | prefix cache isolation |
¹ default ที่ไม่อยู่ใน schema —
temperature,top_p,top_k,repetition_penalty,min_pไม่มีdefaultใน OpenAPI schema แต่ vLLM engine มีค่า default ภายใน:temperature=1.0,top_p=1.0,top_k=-1(ไม่จำกัด),repetition_penalty=1.0,min_p=0.0ค่าเหล่านี้เป็น engine-level default ไม่ใช่ schema-level default ต้องดูจาก vLLM source code หรือ docs แทน
Tip: วิธีดู default และ required เอง — OpenAPI schema มี
requiredarray และdefaultfield ในแต่ละ property ดึงได้ด้วย curl command ด้านบน ถ้าdefaultไม่แสดง (—) แปลว่า vLLM ใช้ engine default ภายใน หรือไม่มีค่า default (ต้องส่งมาเอง)
Note: LiteLLM integration — params ที่ไม่ใช่ OpenAI standard ต้องส่งผ่าน
extra_bodyใน LiteLLM เช่นtop_k,min_p,reasoning_effort,thinking_token_budgetดูตัวอย่างด้านล่าง
ตัวอย่างการเรียกผ่าน LiteLLM พร้อม vLLM-specific params:
from litellm import completion
response = completion(
model="openai/ornith-35b-nvfp4",
api_base="http://10.0.0.246:8000/v1",
api_key="EMPTY",
messages=[{"role": "user", "content": "Write a Python fibonacci function"}],
temperature=0.7,
top_p=0.9,
max_tokens=2048,
stream=True,
# vLLM-specific params ผ่าน extra_body
extra_body={
"top_k": 20,
"min_p": 0.05,
"repetition_penalty": 1.05,
"reasoning_effort": "medium", # none / minimal / low / medium / high / xhigh / max
"include_reasoning": True, # ส่ง reasoning content กลับมา
"thinking_token_budget": 4096, # จำกัด thinking tokens
},
)
Caution:
reasoning_effortvsthinking_token_budget— ทั้งสองควบคุม reasoning แต่คนละมิติreasoning_effortเป็น enum (none/minimal/low/medium/high/xhigh/max) ส่วนthinking_token_budgetเป็นตัวเลขจำนวน token ที่จำกัด แนะนำตั้งreasoning_effortก่อน แล้วใช้thinking_token_budgetกรณีต้องการจำกัดเพิ่ม
Tip: สลับ reasoning แบบ on/off ระหว่าง request — นอกจาก
reasoning_effortแล้ว model card ของ sakamakismile แนะนำให้ใช้chat_template_kwargs: {"enable_thinking": true|false}สำหรับการเปิด/ปิด reasoning แบบ binary ต่อ request หนึ่งๆ โดยไม่ต้องเปลี่ยนแปลงที่ server
Note:
max_tokensสำหรับ reasoning mode — model card แนะนำmax_tokens ≥ 6500เมื่อรัน reasoning mode บน code benchmark เพราะ Ornith คิดยาว ถ้าตั้งน้อยเกินไป คำตอบอาจถูกตัดกลางทาง ในการทดสอบของผมใช้ prompt ที่คำตอบอยู่ในช่วง 1,374–2,543 tokens จึงไม่กระทบ แต่ถ้าเป็นงาน coding ที่ซับซ้อนควรเพิ่มmax_tokensให้สูงขึ้น
Security: Private IP — ตัวอย่าง curl และ LiteLLM ในบทความนี้อ้างอิง
http://10.0.0.246:8000ซึ่งเป็น private IP ภายในเครือข่ายบ้าน (LAN) ไม่ได้เปิดให้ภายนอกเข้าถึง หากจะ deploy ในสภาพแวดล้อมอื่น ควรวาง vLLM ไว้หลัง reverse proxy ที่มี authentication และเปิดเฉพาะพอร์ตที่จำเป็น
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee