DeepSeek-V4-Flash-0731 Reasoning Effort: เทสจริง 7 รอบ พบว่า docs กับ serving ไม่ตรงกัน (มีอัปเดต)
สารบัญ
- Setup ที่ใช้เทส
- Test #1: effort × thinking=enabled
- Test #2: effort × thinking=disabled
- Test #3: ไม่มี reasoning_effort
- Test #4: effort ที่ docs ไม่มี
- Test #5:
chat_template_kwargs.thinking(อัปเดต 23 ส.ค.) - Test #6:
reasoning_effort: "none"(root level) - Test #7:
thinking_token_budget - เทียบ docs กับ serving ที่ผมเจอ
- เทียบ 3 วิธี "ปิด reasoning"
- Surprises ที่ docs ไม่ได้บอก
- Effort ไหนใช้เมื่อไร
- Caveats ที่ควรระวัง
- สรุป
- อ้างอิง
บันทึกกลางดึก — 17 สิงหาคม 2569 — ตี 2 (อัปเดตคืน 23 ส.ค. 2569 — เพิ่ม Test #5–#7)
TL;DR: ถ้าใช้ vLLM local serving อยากปิด reasoning ให้ใช้
chat_template_kwargs: {"thinking": false}(วิธีทางการ) หรือreasoning_effort: "none"— อย่าใช้thinking: {type: "disabled"}เพราะ vLLM serving นี้ ignore param นี้
เรื่องมันเริ่มจากที่ผมเพิ่งเซ็ต self-hosted vLLM เสร็จ แล้วอยากรู้ว่า reasoning_effort มันมีผลจริงไหม — เพราะ docs ของ DeepSeek บอกว่ามี mapping แบบนี้:
low → low
medium → high
high → high
xhigh → high
max → max
แต่พอยิง API จริงๆ ผลออกมาไม่ตรงกับที่ docs บอกเลย — low กับ high ให้ output identical, ส่วน xhigh ก็ไม่ได้ map ไป high อย่างที่ docs ว่า
แล้วคืนนี้ผมก็ไปเจออีกว่า — thinking: {type: "disabled"} ที่ดูเหมือน toggle ตรงๆ ก็ไม่มีผลใน serving นี้ด้วย (เพราะมันถูก reasoning_effort ทับ) — แต่มี 2 วิธีที่ปิด reasoning ได้จริง ซึ่งจะเล่าใน Test #5–#6
Setup ที่ใช้เทส
| Item | Value |
|---|---|
| Model | deepseek-v4-flash-0731 |
| Serving | vLLM local, http://10.0.0.246:8000/v1 |
| Prompt หลัก | "Rearrange letters 'CIFAIPC' — ocean / country / animal / fruit" |
| API mode | Streaming (stream_options.include_usage = true) |
| Metrics | TTFT, total latency, prompt_tokens, completion_tokens, reasoning chars |
ผมวัด TTFT จาก timestamp ของ chunk แรกที่มี choices — ใช้ streaming เพราะ non-streaming ไม่ต่างจาก TTFT == Total
ตัวเลขทุกตัวในบทความนี้มาจากการยิง API จริง ไม่มี fabricate
Test #1: effort × thinking=enabled
ผมเริ่มจาก payload ที่ docs แนะนำ — เปิด thinking แล้วปรับ effort:
{
"reasoning_effort": "low/high/max",
"thinking": {"type": "enabled"}
}
| effort | TTFT | Total | in | out | total | reasoning_chars | content |
|---|---|---|---|---|---|---|---|
| low | 0.37s | 4.43s | 58 | 90 | 148 | 289 | 11 |
| high | 0.40s | 4.39s | 58 | 90 | 148 | 289 | 11 |
| max | 0.51s | 7.16s | 137 | 136 | 273 | 450 | 11 |
สังเกตแรก: low กับ high ให้ output identical ทุกตัว — reasoning_chars, content_chars, total tokens ตรงกันเป๊ะ
สังเกตที่สอง: max ใช้ prompt_tokens สูงกว่า 79 tokens (137 vs 58) — serving น่าจะ inject hidden system prompt เมื่อ effort สูงขึ้น
สังเกตที่สาม: คำตอบถูกทั้ง 3 effort (anagram เป็นโจทย์ง่าย ไม่ differentiate quality)
Test #2: effort × thinking=disabled
หลังจาก test #1 ผมสงสัยว่า thinking: disabled จะปิด reasoning จริงไหม — เลยเทสต่อ:
{
"reasoning_effort": "low/high/max",
"thinking": {"type": "disabled"}
}
| effort | TTFT | Total | in | out | total | reasoning_chars | content |
|---|---|---|---|---|---|---|---|
| low | 0.36s | 4.34s | 58 | 90 | 148 | 289 | 11 |
| high | 0.40s | 4.39s | 58 | 90 | 148 | 289 | 11 |
| max | 0.52s | 17.00s | 137 | 400 | 537 | 1404 | 0 |
ผลแปลก: reasoning_chars ยังออกมาเหมือนเดิม — เหมือน test #1 ทุกตัวเลข
แปลว่า serving นี้ ไม่ honor thinking.type เลย ถ้ามี reasoning_effort อยู่ใน payload — เหมือน reasoning_effort เป็น hard switch ที่ทับ thinking toggle
Test #3: ไม่มี reasoning_effort
ผมเลยลองตัด reasoning_effort ออกจาก payload แล้วเก็บ thinking: disabled ไว้:
{
"thinking": {"type": "disabled"}
}
| prompt | TTFT | Total | in | out | total | reasoning_chars | content |
|---|---|---|---|---|---|---|---|
| anagram | 0.36s | 0.68s | 58 | 4 | 62 | 0 | 11 |
| logic | 0.39s | 16.18s | 81 | 388 | 469 | 0 | 1493 |
ตอนนี้ reasoning_chars = 0 จริงๆ — และ latency ลดฮวบ (anagram 4.34s → 0.68s, 6.4× เร็วขึ้น)
สรุปคือ: ถ้าอยากปิด reasoning จริงๆ ต้อง ไม่ส่ง reasoning_effort เลย ไม่ใช่แค่ใส่ thinking: disabled
Test #4: effort ที่ docs ไม่มี
ท้ายสุดผมเทส effort ทั้ง 5 ค่า (รวม medium กับ xhigh ที่ docs ไม่ได้พูดถึงใน OpenAI format — บอกแค่ใน Responses API):
{
"reasoning_effort": "low/medium/high/xhigh/max",
"thinking": {"type": "enabled"}
}
| effort | TTFT | Total | in | out | total | reasoning_chars |
|---|---|---|---|---|---|---|
| low | 0.37s | 4.43s | 58 | 90 | 148 | 289 |
| medium | 0.41s | 4.32s | 58 | 90 | 148 | 289 |
| high | 0.40s | 4.39s | 58 | 90 | 148 | 289 |
| xhigh | 0.51s | 7.28s | 137 | 136 | 273 | 450 |
| max | 0.51s | 7.16s | 137 | 136 | 273 | 450 |
มี 2 buckets ชัดเจน — low/medium/high ทั้งหมดให้ output เดียวกัน (148 tokens), xhigh/max ก็เหมือนกันอีก bucket (273 tokens)
Test #5: chat_template_kwargs.thinking (อัปเดต 23 ส.ค.)
คืนนี้ผมกลับมาเทสต่อ — สงสัยว่า vLLM มี channel อื่นในการคุม thinking ไหม เพราะ chat template น่าจะ expose flag ตรงๆ
{
"chat_template_kwargs": {"thinking": true},
"messages": [{"role": "user", "content": "What is 6 × 7? Just the number."}]
}
chat_template_kwargs | reasoning field | content | completion_tokens |
|---|---|---|---|
{"thinking": true} | ✅ "We need answer 42." | 42 | 9 |
{"thinking": false} | null | 42 | 2 |
{"thinking": true, "reasoning_effort": "none"} | null | 42 | 2 |
{"thinking": false, "reasoning_effort": "none"} | null | 42 | 2 |
ผลดีใจ: chat_template_kwargs.thinking คุม reasoning ได้ตรงๆ ไม่ต้องถอด reasoning_effort ออก
thinking: false→ reasoning field =null, เหลือ completion_tokens 2 (เฉพาะคำตอบ)reasoning_effort: "none"ในchat_template_kwargsไม่มีผลเพิ่ม — แค่thinking: falseก็พอ
Test #6: reasoning_effort: "none" (root level)
หลังจากเจอ chat_template_kwargs ผมเลยลอง reasoning_effort ด้วยค่า "none" (ซึ่ง docs ไม่ได้พูดถึงใน OpenAI format แต่มีใน Responses API):
{
"thinking": {"type": "disabled"},
"reasoning_effort": "none",
"messages": [{"role": "user", "content": "What is 6 × 7? Just the number."}]
}
| คำสั่ง | reasoning | content | completion_tokens |
|---|---|---|---|
thinking: disabled + reasoning_effort: "none" (root) | null | 42 | 2 |
ทั้งสองอยู่ใน extra_body | "We need answer 42." | 42 | 9 |
ผลดีใจอีก: ถ้าส่ง reasoning_effort: "none" ที่ root level (ไม่ใช่ใน extra_body) ก็ปิด reasoning ได้เหมือนกัน
แต่ส่งใน extra_body → ไม่มีผล — vLLM ฝั่งนี้อ่าน reasoning_effort ตรงๆ ไม่ได้ map จาก extra_body
Test #7: thinking_token_budget
สุดท้ายลองคุมความยาว reasoning ผ่าน budget:
{
"chat_template_kwargs": {"thinking": true, "thinking_token_budget": 1024}
}
Prompt ยากๆ: "A farmer has 17 chickens and 23 ducks. Each chicken lays 2 eggs/day, each duck lays 1. After 30 days, how many eggs total? Show your full reasoning, then state the final number."
| budget | reasoning len | completion_tokens |
|---|---|---|
| (no budget) | 267 | 160 |
| 1024 | 270 | 154 |
| 128 | 295 | 161 |
| 32 | 355 | 189 |
| 0 | 26 | 11 |
ผลงง: budget ไม่ cap reasoning จริง — Flash model reasoning จบเองที่ ~270 chars อยู่แล้ว และ budget=32 กลับยาวขึ้น (น่าจะเป็น sampling randomness)
และถ้าส่ง thinking_token_budget ที่ top-level (ไม่ใช่ใน chat_template_kwargs) → HTTP 400 Bad Request
เทียบ docs กับ serving ที่ผมเจอ
| Requested effort | Docs บอก | Serving ทำจริง | ตรงกัน? |
|---|---|---|---|
| low | low | low (148t) | ✅ |
| medium | high | low (148t) | ❌ serving ทำเป็น low |
| high | high | low (148t) | ❌ serving collapse |
| xhigh | high | max (273t) | ❌ serving ทำเป็น max |
| max | max | max (273t) | ✅ |
Serving ที่ผมเจอมี 2 buckets จริงๆ ไม่ใช่ 5 — และ boundaries ไม่ตรงกับ docs:
- Bucket A (148 tokens) =
low / medium / high - Bucket B (273 tokens) =
xhigh / max
เทียบ 3 วิธี "ปิด reasoning"
คืนนี้ผมรู้แล้วว่ามี 3 วิธีที่ปิด reasoning ได้จริง (รวม Test #3 เดิม):
| วิธี | ผล | ที่มา |
|---|---|---|
chat_template_kwargs: {"thinking": false} | ✅ ปิด | vLLM-native (Test #5) |
reasoning_effort: "none" (root) | ✅ ปิด | OpenAI format (Test #6) |
thinking: {type: "disabled"} + ไม่ส่ง reasoning_effort | ✅ ปิด | workaround (Test #3) |
thinking: {type: "disabled"} อย่างเดียว | ❌ ไม่มีผล | Test #1, #2 |
แนะนำ: ใช้ chat_template_kwargs เพราะ explicit ที่สุดและเป็น vLLM-native API
Surprises ที่ docs ไม่ได้บอก
1. reasoning_effort ทับ thinking.type
ถ้าส่ง reasoning_effort อยู่ใน payload, serving นี้ force-on thinking โดยไม่สนใจ thinking: disabled — ต้อง ไม่ส่ง reasoning_effort เลย หรือใช้ chat_template_kwargs.thinking: false ถึงจะปิด reasoning จริง
2. low ≈ high ใน serving นี้
ผลลัพธ์ identical ทุก test — ถ้าใช้ serving นี้ ไม่ควร assume ว่า low กับ high ให้ output ต่างกัน
3. xhigh ใน OpenAI format
Docs บอก xhigh มีแค่ใน Responses API (reasoning.effort) — แต่ serving นี้ยอมรับใน reasoning_effort ของ OpenAI format แล้ว map ไป bucket เดียวกับ max (ไม่ใช่ high อย่างที่ docs บอก)
4. Field name ต่างจาก docs
Docs บอก response field คือ reasoning_content (ระดับเดียวกับ content) แต่ serving นี้ใช้ field reasoning — ถ้า migrate code ไป official API ต้องเปลี่ยน field name
5. thinking.type ถูก ignore ทั้งหมด
เพิ่งเจอคืนนี้ — ลองส่ง "thinking": {"type": "enabled"} และ "thinking": {"type": "disabled"} ทั้ง root level และ extra_body → reasoning ออกเหมือนกันทุกกรณี vLLM serving นี้ไม่ map param นี้เลย ต่างจาก chat_template_kwargs.thinking ที่ใช้ได้จริง
6. reasoning_effort: "none" ใช้ได้ใน OpenAI format
Docs บอกมีเฉพาะใน Responses API แต่ vLLM serving นี้รับใน reasoning_effort ของ OpenAI format (root level) แล้วปิด reasoning ได้จริง — เป็น alias ที่ใช้ได้
Note: ใช้ serving นี้ vs official API — serving นี้เป็น vLLM local deployment ไม่ใช่ official DeepSeek API ที่ api.deepseek.com — behavior อาจต่างกัน ตัวเลขและ quirks ในบทความนี้ apply กับ local serving เท่านั้น
Effort ไหนใช้เมื่อไร
จากผลเทสนี้ ผมสรุป guideline สำหรับ serving ที่ผมใช้อยู่:
| Use case | Effort / config | เหตุผล |
|---|---|---|
| Quick answer, ประหยัด tokens | low หรือ medium | Bucket A เร็วสุด (~4.3s), ถูกสุด (148t) |
| Balanced reasoning | high | Bucket A เหมือนกัน — default ของ docs |
| Deep reasoning | max | Bucket B (~7.2s, 273t) — reasoning ยาวขึ้น ~55% |
| หลีกเลี่ยง | xhigh | docs map ไป high แต่ serving ทำเป็น max — unpredictable |
| ปิด reasoning (แนะนำ) | chat_template_kwargs: {"thinking": false} | vLLM-native, explicit, ไม่ชนกับ reasoning_effort |
| ปิด reasoning (alias) | reasoning_effort: "none" (root) | OpenAI format, ใช้ได้เหมือนกัน |
| ปิด reasoning (workaround) | thinking: disabled + ไม่ส่ง reasoning_effort | ใช้ได้แต่เปราะ — ใครเติม reasoning_effort ทีหลังก็พัง |
Caveats ที่ควรระวัง
- โจทย์ที่ใช้เทสง่าย — anagram กับ logic puzzle เป็น single-step หรือ 2-step ถ้าโจทย์ยากกว่านี้ effort อาจมี quality gap ที่วัดได้ (ผมยังไม่ได้เทส math olympiad / multi-turn reasoning)
- 1-2 runs ต่อ effort — ตัวเลข latency มี noise ~5%, ถ้าต้องการความแม่นยำควรเทส 5+ runs แล้ว median
- Serving เดียว — ผมเทสแค่ vLLM local, official DeepSeek API อาจ behave ต่างกันโดยสิ้นเชิง
reasoning_tokensเป็น estimate — vLLM aggregate เป็นcompletion_tokensก้อนเดียว ผมนับจากreasoning_chars // 4ซึ่งเป็น approximation ไม่ใช่ exact token countthinking_token_budgetไม่ cap reasoning จริง — Flash model reasoning สั้นอยู่แล้ว ถ้าใช้ model ที่ reasoning ยาวกว่านี้ behavior อาจต่าง
สรุป
Serving ที่ผมเทสมี 2 buckets จริงๆ ไม่ใช่ 5 — low/medium/high ทั้งหมดให้ output เหมือนกัน (148 tokens), xhigh/max ก็เหมือนกันอีก bucket (273 tokens) ต่างจาก docs ที่บอกว่ามี 5 levels
ถ้าอยากปิด reasoning ให้ใช้ chat_template_kwargs: {"thinking": false} (vLLM-native, explicit) หรือ reasoning_effort: "none" ที่ root level — อย่าใช้ thinking: {type: "disabled"} อย่างเดียว เพราะ serving นี้ ignore param นี้เมื่อมี reasoning_effort อยู่ใน payload
ตัวเลขทุกตัวในบทความนี้มาจากการยิง API จริง — apply เฉพาะ vLLM local serving ที่ http://10.0.0.246:8000/v1 เท่านั้น official DeepSeek API อาจ behave ต่างกัน
อ้างอิง
- DeepSeek Thinking Mode Docs — effort mapping, response fields, multi-turn conversation, tool calls
- vLLM Documentation — serving config,
stream_options,chat_template_kwargs, reasoning model support - OpenAI Chat Completions API — payload format สำหรับ
reasoning_effort
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee