Qwen3.8-Flash-Next 4 นวัตกรรมสถาปัตยกรรม: ข้อดี ข้อเสีย และสิ่งที่น่าสนใจ — เจาะลึกจาก Tech Report
สารบัญ
- TL;DR
- บริบท: ทำไมถึงมี 4 อย่างนี้
- 1️⃣ QSA (Qwen Sparse Attention) — sparse ที่ระดับ micro-block
- ข้อดี
- ข้อเสีย/ข้อจำกัด
- สิ่งที่น่าสนใจ
- 2️⃣ Gated Residual — residual แยก 4 branches
- ข้อดี
- ข้อเสีย/ข้อจำกัด
- สิ่งที่น่าสนใจ
- 3️⃣ N-gram Embedding — scale params โดยไม่เพิ่ม compute
- ข้อดี
- ข้อเสีย/ข้อจำกัด
- สิ่งที่น่าสนใจ
- 4️⃣ Muon + AdamW — optimizer คนละตัวต่อ weight category
- ข้อดี
- ข้อเสีย/ข้อจำกัด
- สิ่งที่น่าสนใจ
- Community reactions — controversy + rebut
- ฝั่ง criticism (AdrienneNoctis, HF Discussion #24)
- Rebut จาก paper
- ฝั่ง user testing
- ฝั่ง license complaint
- สรุปเปรียบเทียบรวม
- สำหรับคนอ่านตัดสินใจ
- อ้างอิง
บทความนี้เป็น follow-up จาก Qwen3.8-Flash-Next เปิดตัวแล้ว — คราวนี้เราจะไม่พูดแค่ว่ามัน "เจ๋ง" แต่จะเจาะลึก 4 นวัตกรรมหลัก (QSA, Gated Residual, N-gram Embedding, Muon+AdamW) ว่ามี ข้อดี ข้อเสีย และมุมมองที่น่าสนใจ อะไรบ้าง โดยอ้างอิงจาก tech report 51 หน้าที่ Qwen Team ปล่อยออกมาพร้อมโมเดล
TL;DR
บทความนี้เจาะลึก 4 นวัตกรรมของ Qwen3.8-Flash-Next เทียบกับสถาปัตยกรรม MoE ทั่วไป:
| นวัตกรรม | หน้าที่ | ข้อดีเด่น | ข้อเสีย/ข้อจำกัด | น่าสนใจเพราะ |
|---|---|---|---|---|
| QSA | Sparse attention ที่ระดับ micro-block | 7.6× faster prefill / 4.9× decode ที่ 1M context | ต้อง fused kernel ใหม่, MTP ต้อง reuse top-k | เป็น sparse attn แรกที่ไม่ share indices ข้าม layer |
| Gated Residual | Residual แยก 4 branches + gate | Stable training มาก, ไม่มี loss spike | เพิ่ม memory traffic, write ใช้แค่ 2 branches จริงๆ | GR เลือก path เฉพาะ (long-range) เพิ่ม share 0.05-0.19 |
| N-gram Embedding | Lookup table จาก bigram/trigram | Offloadable, แทบไม่เพิ่ม compute | Loss optimum ไม่ตรงกับ accuracy optimum | ใช้ n-gram แบบ "เดาคำถัดไป" แต่ scale ดีกว่า MoE |
| Muon + AdamW | Optimizer คนละตัวต่อ weight category | ไม่ต้อง warmup batch size, zero loss spikes | Fused matrices (qkv, SwiGLU) ต้อง split | เปลี่ยนจาก "ค่าย AdamW ล้วน" ในวงการ |
บริบท: ทำไมถึงมี 4 อย่างนี้
Qwen team บอกชัดใน tech report ว่าพวกเขา เจอ disagreement ระหว่าง training loss กับ downstream accuracy ตลอดการออกแบบ:
"Loss and downstream accuracy do not always move together, and we observe disagreements in both directions." — paper section 2
นั่นคือทำไม paper นี้ยาว 51 หน้าและออกแบบอย่างระมัดระวัง — ไม่ใช่ "ของใหม่ล้วน" แต่เป็น "ของที่พิสูจน์แล้วว่า ablation รอดจากทั้ง pre-training และ post-training"
ที่สำคัญคือ ทั้ง 4 อย่างถูกพัฒนา เป็น preview สำหรับ Qwen4 ที่จะออกปลายปี — เหมือนที่ Qwen3-Next เคยเป็นตัวทดสอบก่อนกลายเป็น backbone ของ Qwen3.5-3.8 ทั้ง family
มาดูแต่ละอย่างกันเลย
1️⃣ QSA (Qwen Sparse Attention) — sparse ที่ระดับ micro-block
ข้อดี
QSA เป็น sparse attention ที่ระดับ micro-block ไม่ใช่ token-level — ใช้ lightweight indexer ที่:
- บีบอัด sequence เป็น block representations (compression ratio r)
- ให้ score ความสำคัญของแต่ละ block
- Top-k selection: 512 blocks หรือ 2048 tokens ต่อ layer
จุดต่างเทียบกับ sparse attn เดิม (เช่น IndexShare):
| มิติ | QSA | IndexShare (baseline) |
|---|---|---|
| Index granularity | Per-layer | Shared ข้าม adjacent layers |
| Cross-layer dependency | ต่ำ (แต่ละ layer มี index ของตัวเอง) | สูง (ต้อง assume similarity) |
| RULER match full attn | ที่ indexer latency 0.25× | ที่ 0.5× ยังต่ำกว่า full attn |
| Hybrid (GDN) fit | Natural | Awkward |
ข้อเสีย/ข้อจำกัด
- ต้อง fused kernel ใหม่ — paper บอกชัดว่า "we implement a fused QSA kernel that jointly computes sparse attention outputs and the KL loss without materializing intermediate results" — ไม่ใช่ plug-and-play
- Speedup จริงโผล่ที่ context ยาว — ที่ ≤64K context, QSA ยังไม่มี speedup เหนือ dense (Fig. 6a/b)
- ต้อง KL loss เทรน indexer — เพิ่ม memory pressure ตอน training
- MTP ต้อง reuse top-k indices — ไม่งั้น draft-model cost พุ่ง
สิ่งที่น่าสนใจ
ที่ context 1M tokens, QSA เร็วกว่า dense attention 7.6× ตอน prefill, 4.9× ตอน decode — และที่สำคัญคือ ทำได้โดยไม่เสีย accuracy:
- Short-context (MMLU-Pro, GSM8K, BBH): QSA = 76.8 average vs Full Attn = 75.9
- Long-context (RULER): QSA ที่ 512K-1M range ดีกว่า full attn (26.44 vs 20.71)
แปลว่า QSA เป็น sparse attn ที่ไม่ต้องแลกกับ quality ที่ short-context — ต่างจาก sparse attn เดิมๆ ที่มักจะ degrade
2️⃣ Gated Residual — residual แยก 4 branches
ข้อดี
แทนที่จะมี residual stream เดียว (สไตล์ pre-norm), Qwen3.8-Flash-Next แยกเป็น 4 branches ที่:
- อ่านผ่าน element-wise data-dependent gate (read gate)
- เขียนผ่าน per-branch scalar write gate (write gate)
- มี group RMSNorm ก่อน merge
ผลที่ได้:
- Training stability พุ่ง: ภายใต้ stress test (3× optimal LR) Gated Residual ไม่มี loss spike เลย ขณะที่ baseline ข้าม clipping threshold
- Batch size + LR สูงขึ้น — gate ทำหน้าที่ rescale ทำให้ convergence ดีขึ้น
- ความจริงที่น่าทึ่ง: ablation แสดงว่า GR เลือก path เฉพาะ — branch 0 เก็บ long-range paths (skip ~10 layers) ส่วนอีก 3 เก็บ local (skip 1-3) — โดย ไม่ได้ตั้งใจ มันเกิดจาก training เอง
ตัวเลข ablation จาก paper:
- Layer 0 GDN → Layer 15 attention: share 0.008 (no-GR) → 0.058 (GR) — 7× gain
- Layer 0 GDN → Layer 2: 0.139 → 0.192 — 1.4× gain
ข้อเสีย/ข้อจำกัด
- Memory traffic เพิ่ม — ต้อง carry 4 branches แทน 1 stream
- Inference optimization ยาก — paper ลอง "sparse write" (เขียนแค่ 2 branches ที่ gate สูงสุด) ผลคือ "Pre-training loss and benchmarks were almost unaffected, but the quality degraded clearly after post-training, so we did not adopt it" — เป็นตัวอย่างที่ pre-training metric หลอก
- ต้อง FP8 residual state เพื่อลด memory traffic — "Storing the branches in FP8 halves the bytes moved for the residual state relative to BF16, with almost no loss in quality"
- Write stays scalar per-branch — ลอง refine เป็น per-channel แล้วได้ "almost nothing"
สิ่งที่น่าสนใจ
Paper บอกชัดว่า "Predicting the operators from all branches is better than using only the last branch or pooling the branches first" — นั่นคือ design choice ที่ทำให้ GR ต่างจาก mHC/HC/VWN ที่ใช้ per-branch scalar อย่างเดียว
อีกจุดที่น่าสนใจ: paper ยอมรับว่าพยายาม sparse write แล้วพังตอน post-training — เป็นบทเรียนที่ดีว่า ablation จาก pre-training อย่างเดียวหลอกได้
3️⃣ N-gram Embedding — scale params โดยไม่เพิ่ม compute
ข้อดี
N-gram (PLE — Position-aware Layered Embedding) เป็น lookup table 20M bigrams/trigrams ที่:
- Hash จาก trigram context (deterministic lookup)
- Slide window ทำเป็น vector เดียว O(1)
- Table offloadable ไป host RAM/NVMe — แทบไม่กิน VRAM
- Asynchronous prefetching ทับกับ backbone compute
ตัวเลข scale ที่น่าสนใจ:
- Qwen3.8-Flash-Next: 51B n-gram จาก total 180B params (28% ของโมเดล)
- Qwen3.7-Plus: 397B-A17B (40× active params ใหญ่กว่า)
- Qwen3.8-Flash-Next ใช้ training FLOPs 1/9 ของ Qwen3.7-Plus แต่ทำ score ดีกว่าในหลาย benchmark
ข้อเสีย/ข้อจำกัด
Paper ยอมรับตรงๆ ใน section 2.4:
"Loss and downstream accuracy do not always move together... Enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates, and under a fixed parameter budget the loss optimum diverges from the accuracy optimum"
แปลว่า:
- ขยาย n-gram table → loss ลดลงเรื่อยๆ → แต่ accuracy ไม่ขึ้นตาม
- ภายใต้ budget เดียว — optimum ของ loss ไม่ใช่ optimum ของ accuracy
- คือขนาด table ใหญ่ไม่ได้แปลว่าดีเสมอ
ข้อเสียอื่น:
- ต้อง fast SSD/RAM ตอน offload — random read latency กลายเป็น bottleneck
- ต้อง cache ส่วนที่ใช้บ่อย — ไม่งั้น hot-path จะช้า
- Tokenizer coupling — table ขึ้นกับ vocab ของโมเดล, เปลี่ยน tokenizer = ต้องสร้าง table ใหม่
สิ่งที่น่าสนใจ
Paper ทำ sliding window แทน fixed n-gram:
- ดู bigram/trigram แบบยืดหยุ่น ผ่าน hash function
- O(1) lookup, semantic แผ่กว้างกว่า static n-gram
- มาจากแนวคิด DeepSeek Engrams (hash lookup, flops = 0) — แต่ Qwen ใส่ position-aware ทำให้ table มี 20M entries แทนที่จะเป็น shared embedding
ที่น่าประหลาดใจ: 51B n-gram คือ 28% ของ 180B แต่ paper บอกว่า "n-gram vocabulary does not count as effective parameters" เพราะมันเป็น deterministic lookup จาก input
ในทางปฏิบัติ n-gram เป็น deterministic lookup จาก input tokens — เหมือน KV cache ไม่ใช่ parameters ปกติ — prefill ดึง n-grams ที่จำเป็น (fraction เล็กมากของ 51B) เข้า RAM/VRAM แล้ว reuse ซ้ำๆ ตอน decode
4️⃣ Muon + AdamW — optimizer คนละตัวต่อ weight category
ข้อดี
Qwen3.8-Flash-Next ใช้ Muon กับ 2D weight matrices (linear maps) และ AdamW กับ:
- Input embeddings
- N-gram embeddings
- Output head
- MoE router
- Low-rank projections ของ GR
ผลที่ได้:
- No batch-size warmup — start ที่ target batch size ทันที
- Zero loss spikes ที่ production LR (8× scale stress test)
- Median gradient norm ลดลง ~2× vs AdamW baseline, p99.9 ลด 4.2×
- 1,000-step window std ลด 4.3-4.7×
ที่สำคัญคือ Muon ทำให้ ไม่ต้องพึ่ง qk-clip หรือ SwiGLU-clip ที่ Kimi/others ใช้แก้ instabilities
ข้อเสีย/ข้อจำกัด
- Fused matrices ต้อง split — qkv projection, SwiGLU fc1, GDN input proj ต้อง orthogonalize แยก เพราะ "Orthogonalizing the fused matrix muddles updates to operators with very different gradient statistics"
- Elongated shapes ยังต้องใช้ AdamW — paper บอกชัดว่า "low-rank projections of GR likewise perform better with AdamW. We attribute this to their very elongated shape"
- N-gram embedding = AdamW (weight decay off) — ต้องแยก config ระหว่าง optimizer groups
- Muon คนเดียวไม่พอ — "Muon on its own has roughly twice the median norm" ใน stress test — ต้องจับคู่กับ GR/Muon+GR
สิ่งที่น่าสนใจ
เหตุผลที่แยก optimizer ต่อ weight category เป็นเรื่องที่น่าสนใจ — paper แสดงให้เห็นว่า "gradient statistics" ของ weight แต่ละประเภทต่างกันมาก:
- 2D linear maps → orthogonalize (Muon)
- 1D scalar/long-tail → AdamW
- Embedding tables → AdamW (low precision friendly)
ที่น่าทึ่ง: Muon + Gated Residual รวมกันคือ combo ที่ทำให้ zero loss spike — ไม่ใช่แค่ตัวใดตัวหนึ่ง
Community reactions — controversy + rebut
หลังเปิดตัว มี debate ใน community เรื่อง design choices:
ฝั่ง criticism (AdrienneNoctis, HF Discussion #24)
คนนี้โพสต์บทวิจารณ์ยาว (เขียนเป็นภาษาอังกฤษปนฝรั่งเศส) ว่า:
"Hidden size 2560, expert intermediate 640, 10 experts per token, 1 shared — the living tissue of your model totals roughly 4B active parameters, not the six you whisper in marketing. ... 512 experts of 640 intermediate — you did not build a Mixture-of-Experts; you built a Mixture-of-Excuses."
และโจมตี n-gram ว่า:
"20,000,000 memorized trigrams. This is not reasoning — this is a lookup table, a cane borrowed from the DeepSeek n-gram wardrobe. When the test asks a question whose trigram pattern lives in the table, the dummy is steered to the memorized vector and 'fires without thinking.'"
โพสต์นี้ได้รับ 6 reactions (รูปหน้าตกใจ/เศร้า) แต่โดน counter-argument ว่า "AI slop discussion" เพราะยาวเกินไปและมี rhetoric หนัก
Rebut จาก paper
ข้อโจมตี "4B dummy" ไม่ตรงกับ spec paper — paper ระบุ:
- 6B activated per token (เป็นทางการ)
- 10 routed + 1 shared expert ที่ 640 intermediate
- Hidden 2560, 48 layers
- Math: ~6B ตามที่ claim
ข้อโจมตี "stolen n-gram":
- Paper cites DeepSeek Engrams ใน references — ระบุแหล่งที่มาชัดเจน ไม่ได้ "ซ่อน"
- Qwen's innovation คือ position-aware layering (PLE) + sliding window hash แทน fixed n-gram
- และ paper ยอมรับตรงๆ ว่า "loss optimum diverges from accuracy optimum" — ไม่ได้อ้างว่า n-gram คือ reasoning
ฝั่ง user testing
มีคนทดสอบจริงใน r/LocalLLM (เปรียบเทียบ Q4 vs 27B Q8):
| มิติ | Qwen3.8 27B Q8 | Qwen3.8 Flash-Next IQ4_XS | Winner |
|---|---|---|---|
| Tool-eval hard mode (100) | 91 | 77 | 27B |
| VRAM/RAM | 68 GB | 95 GB VRAM + 51 GB RAM | 27B |
| State/structured output errors | ปกติ | "noticeably more" | 27B |
อีกคน (lkarlslund) รายงานตรงข้าม:
"The large MoE won: less time, better results and overall lower power consumption. The downside is that you need a huge setup to run it — but if you do have the rig, there is a clear winner: the MoE we've all been waiting for."
"Both models also waste too much time on xhigh, and low is not worth it — more tool calls so stick to medium."
ความเห็นต่างกันน่าจะเพราะ day-0 quant ยังไม่ stable — LobsterWeary2675 บอกชัดว่า "Could be cause its a day 0 quant and not completely supported yet. But I'll stick with 27B for now"
ฝั่ง license complaint
HF Discussion #18 (8 reactions รูปหน้าตกใจ): คนถามว่าทำไมไม่ใช่ Apache 2.0 — license ใหม่ (Qwen Community License 1.0) trigger 100M MAU หรือ $20M revenue → ต้องแสดงชื่อ model, และ "Model as a Service" ต้องขอ license แยก ถ้าจะ commercial
สรุปเปรียบเทียบรวม
| นวัตกรรม | ข้อดี | ข้อเสีย | เมื่อไหร่ควรสนใจ |
|---|---|---|---|
| QSA | 7.6× prefill @ 1M, ไม่เสีย short ctx quality | ต้อง fused kernel, ไม่ช่วย short ctx | ทำ long-context agent (RAG, code) |
| Gated Residual | Stable training, LR/batch สูงขึ้น | Memory traffic เพิ่ม, sparse write พังตอน post-training | สเกลโมเดลใหญ่ๆ |
| N-gram | Offloadable, scale params ถูก | Loss ≠ accuracy optimum, storage cost | Resource-constrained serving |
| Muon + AdamW | Zero loss spike, no batch warmup | Fused matrices ต้อง split | Training stability เป็น priority |
สำหรับคนอ่านตัดสินใจ
ถ้าคุณ:
- ทำ long-context inference → QSA คือ win ที่ชัดเจนที่สุด
- train โมเดลใหญ่ → Gated Residual + Muon combo worth ศึกษา (ดู ablation ใน paper section 3.3)
- serve โมเดลบน hardware จำกัด → N-gram offload เป็น pattern ที่น่าสนใจ แต่ระวัง loss vs accuracy divergence
- แค่อยากได้ reasoning model → paper-level innovations ไม่สำคัญ — ดู benchmark (Qwen3.8-27B อาจเพียงพอสำหรับ hardware < 64GB)
สิ่งที่ paper ทำให้ผมเชื่อถือมากขึ้น:
- ยอมรับ limitation ตรงๆ (loss ≠ accuracy, sparse write พังตอน post-training)
- Ablation ครบทุก design choice
- ไม่ claim "best in class" — แค่รายงานสิ่งที่ work
สิ่งที่ทำให้ผมสงสัย:
- "1/9 training cost" claim เทียบกับ Qwen3.7-Plus — ตัวเลขนี้ น่า verify ว่ารวมส่วนไหนบ้าง
- N-gram table ที่ offloadable — จะใช้กับ framework ไหนได้บ้าง? SGLang/vLLM support ระดับไหน
- License Qwen Community License 1.0 — ข้อจำกัดจริง สำหรับ commercial use ที่ยังไม่มีคน summarize ชัด
ถ้ามีเวลา จะลอง deep-dive ใน 3 ข้อนี้ที่ blog ถัดไป
อ้างอิง
Official:
Community:
- HF Discussion #24 — AdrienneNoctis critique
- HF Discussion #4 — 180B too large
- HF Discussion #18 — License complaint
- r/LocalLLM comparison: Flash-Next Q4 vs 27B Q8
- Local-AI-Zone deep-dive blog
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee