Skip to main content

Qwen3.8-Flash-Next 4 นวัตกรรมสถาปัตยกรรม: ข้อดี ข้อเสีย และสิ่งที่น่าสนใจ — เจาะลึกจาก Tech Report

· 13 min read

บทความนี้เป็น follow-up จาก Qwen3.8-Flash-Next เปิดตัวแล้ว — คราวนี้เราจะไม่พูดแค่ว่ามัน "เจ๋ง" แต่จะเจาะลึก 4 นวัตกรรมหลัก (QSA, Gated Residual, N-gram Embedding, Muon+AdamW) ว่ามี ข้อดี ข้อเสีย และมุมมองที่น่าสนใจ อะไรบ้าง โดยอ้างอิงจาก tech report 51 หน้าที่ Qwen Team ปล่อยออกมาพร้อมโมเดล

TL;DR​

บทความนี้เจาะลึก 4 นวัตกรรมของ Qwen3.8-Flash-Next เทียบกับสถาปัตยกรรม MoE ทั่วไป:

นวัตกรรมหน้าที่ข้อดีเด่นข้อเสีย/ข้อจำกัดน่าสนใจเพราะ
QSASparse attention ที่ระดับ micro-block7.6× faster prefill / 4.9× decode ที่ 1M contextต้อง fused kernel ใหม่, MTP ต้อง reuse top-kเป็น sparse attn แรกที่ไม่ share indices ข้าม layer
Gated ResidualResidual แยก 4 branches + gateStable training มาก, ไม่มี loss spikeเพิ่ม memory traffic, write ใช้แค่ 2 branches จริงๆGR เลือก path เฉพาะ (long-range) เพิ่ม share 0.05-0.19
N-gram EmbeddingLookup table จาก bigram/trigramOffloadable, แทบไม่เพิ่ม computeLoss optimum ไม่ตรงกับ accuracy optimumใช้ n-gram แบบ "เดาคำถัดไป" แต่ scale ดีกว่า MoE
Muon + AdamWOptimizer คนละตัวต่อ weight categoryไม่ต้อง warmup batch size, zero loss spikesFused matrices (qkv, SwiGLU) ต้อง splitเปลี่ยนจาก "ค่าย AdamW ล้วน" ในวงการ

บริบท: ทำไมถึงมี 4 อย่างนี้​

Qwen team บอกชัดใน tech report ว่าพวกเขา เจอ disagreement ระหว่าง training loss กับ downstream accuracy ตลอดการออกแบบ:

"Loss and downstream accuracy do not always move together, and we observe disagreements in both directions." — paper section 2

นั่นคือทำไม paper นี้ยาว 51 หน้าและออกแบบอย่างระมัดระวัง — ไม่ใช่ "ของใหม่ล้วน" แต่เป็น "ของที่พิสูจน์แล้วว่า ablation รอดจากทั้ง pre-training และ post-training"

ที่สำคัญคือ ทั้ง 4 อย่างถูกพัฒนา เป็น preview สำหรับ Qwen4 ที่จะออกปลายปี — เหมือนที่ Qwen3-Next เคยเป็นตัวทดสอบก่อนกลายเป็น backbone ของ Qwen3.5-3.8 ทั้ง family

มาดูแต่ละอย่างกันเลย


1️⃣ QSA (Qwen Sparse Attention) — sparse ที่ระดับ micro-block​

ข้อดี​

QSA เป็น sparse attention ที่ระดับ micro-block ไม่ใช่ token-level — ใช้ lightweight indexer ที่:

  1. บีบอัด sequence เป็น block representations (compression ratio r)
  2. ให้ score ความสำคัญของแต่ละ block
  3. Top-k selection: 512 blocks หรือ 2048 tokens ต่อ layer

จุดต่างเทียบกับ sparse attn เดิม (เช่น IndexShare):

มิติQSAIndexShare (baseline)
Index granularityPer-layerShared ข้าม adjacent layers
Cross-layer dependencyต่ำ (แต่ละ layer มี index ของตัวเอง)สูง (ต้อง assume similarity)
RULER match full attnที่ indexer latency 0.25×ที่ 0.5× ยังต่ำกว่า full attn
Hybrid (GDN) fitNaturalAwkward

ข้อเสีย/ข้อจำกัด​

  1. ต้อง fused kernel ใหม่ — paper บอกชัดว่า "we implement a fused QSA kernel that jointly computes sparse attention outputs and the KL loss without materializing intermediate results" — ไม่ใช่ plug-and-play
  2. Speedup จริงโผล่ที่ context ยาว — ที่ ≤64K context, QSA ยังไม่มี speedup เหนือ dense (Fig. 6a/b)
  3. ต้อง KL loss เทรน indexer — เพิ่ม memory pressure ตอน training
  4. MTP ต้อง reuse top-k indices — ไม่งั้น draft-model cost พุ่ง

สิ่งที่น่าสนใจ​

ที่ context 1M tokens, QSA เร็วกว่า dense attention 7.6× ตอน prefill, 4.9× ตอน decode — และที่สำคัญคือ ทำได้โดยไม่เสีย accuracy:

  • Short-context (MMLU-Pro, GSM8K, BBH): QSA = 76.8 average vs Full Attn = 75.9
  • Long-context (RULER): QSA ที่ 512K-1M range ดีกว่า full attn (26.44 vs 20.71)

แปลว่า QSA เป็น sparse attn ที่ไม่ต้องแลกกับ quality ที่ short-context — ต่างจาก sparse attn เดิมๆ ที่มักจะ degrade


2️⃣ Gated Residual — residual แยก 4 branches​

ข้อดี​

แทนที่จะมี residual stream เดียว (สไตล์ pre-norm), Qwen3.8-Flash-Next แยกเป็น 4 branches ที่:

  • อ่านผ่าน element-wise data-dependent gate (read gate)
  • เขียนผ่าน per-branch scalar write gate (write gate)
  • มี group RMSNorm ก่อน merge

ผลที่ได้:

  • Training stability พุ่ง: ภายใต้ stress test (3× optimal LR) Gated Residual ไม่มี loss spike เลย ขณะที่ baseline ข้าม clipping threshold
  • Batch size + LR สูงขึ้น — gate ทำหน้าที่ rescale ทำให้ convergence ดีขึ้น
  • ความจริงที่น่าทึ่ง: ablation แสดงว่า GR เลือก path เฉพาะ — branch 0 เก็บ long-range paths (skip ~10 layers) ส่วนอีก 3 เก็บ local (skip 1-3) — โดย ไม่ได้ตั้งใจ มันเกิดจาก training เอง

ตัวเลข ablation จาก paper:

  • Layer 0 GDN → Layer 15 attention: share 0.008 (no-GR) → 0.058 (GR) — 7× gain
  • Layer 0 GDN → Layer 2: 0.139 → 0.192 — 1.4× gain

ข้อเสีย/ข้อจำกัด​

  1. Memory traffic เพิ่ม — ต้อง carry 4 branches แทน 1 stream
  2. Inference optimization ยาก — paper ลอง "sparse write" (เขียนแค่ 2 branches ที่ gate สูงสุด) ผลคือ "Pre-training loss and benchmarks were almost unaffected, but the quality degraded clearly after post-training, so we did not adopt it" — เป็นตัวอย่างที่ pre-training metric หลอก
  3. ต้อง FP8 residual state เพื่อลด memory traffic — "Storing the branches in FP8 halves the bytes moved for the residual state relative to BF16, with almost no loss in quality"
  4. Write stays scalar per-branch — ลอง refine เป็น per-channel แล้วได้ "almost nothing"

สิ่งที่น่าสนใจ​

Paper บอกชัดว่า "Predicting the operators from all branches is better than using only the last branch or pooling the branches first" — นั่นคือ design choice ที่ทำให้ GR ต่างจาก mHC/HC/VWN ที่ใช้ per-branch scalar อย่างเดียว

อีกจุดที่น่าสนใจ: paper ยอมรับว่าพยายาม sparse write แล้วพังตอน post-training — เป็นบทเรียนที่ดีว่า ablation จาก pre-training อย่างเดียวหลอกได้


3️⃣ N-gram Embedding — scale params โดยไม่เพิ่ม compute​

ข้อดี​

N-gram (PLE — Position-aware Layered Embedding) เป็น lookup table 20M bigrams/trigrams ที่:

  • Hash จาก trigram context (deterministic lookup)
  • Slide window ทำเป็น vector เดียว O(1)
  • Table offloadable ไป host RAM/NVMe — แทบไม่กิน VRAM
  • Asynchronous prefetching ทับกับ backbone compute

ตัวเลข scale ที่น่าสนใจ:

  • Qwen3.8-Flash-Next: 51B n-gram จาก total 180B params (28% ของโมเดล)
  • Qwen3.7-Plus: 397B-A17B (40× active params ใหญ่กว่า)
  • Qwen3.8-Flash-Next ใช้ training FLOPs 1/9 ของ Qwen3.7-Plus แต่ทำ score ดีกว่าในหลาย benchmark

ข้อเสีย/ข้อจำกัด​

Paper ยอมรับตรงๆ ใน section 2.4:

"Loss and downstream accuracy do not always move together... Enlarging the n-gram vocabulary lowers loss monotonically while downstream accuracy saturates, and under a fixed parameter budget the loss optimum diverges from the accuracy optimum"

แปลว่า:

  1. ขยาย n-gram table → loss ลดลงเรื่อยๆ → แต่ accuracy ไม่ขึ้นตาม
  2. ภายใต้ budget เดียว — optimum ของ loss ไม่ใช่ optimum ของ accuracy
  3. คือขนาด table ใหญ่ไม่ได้แปลว่าดีเสมอ

ข้อเสียอื่น:

  • ต้อง fast SSD/RAM ตอน offload — random read latency กลายเป็น bottleneck
  • ต้อง cache ส่วนที่ใช้บ่อย — ไม่งั้น hot-path จะช้า
  • Tokenizer coupling — table ขึ้นกับ vocab ของโมเดล, เปลี่ยน tokenizer = ต้องสร้าง table ใหม่

สิ่งที่น่าสนใจ​

Paper ทำ sliding window แทน fixed n-gram:

  • ดู bigram/trigram แบบยืดหยุ่น ผ่าน hash function
  • O(1) lookup, semantic แผ่กว้างกว่า static n-gram
  • มาจากแนวคิด DeepSeek Engrams (hash lookup, flops = 0) — แต่ Qwen ใส่ position-aware ทำให้ table มี 20M entries แทนที่จะเป็น shared embedding

ที่น่าประหลาดใจ: 51B n-gram คือ 28% ของ 180B แต่ paper บอกว่า "n-gram vocabulary does not count as effective parameters" เพราะมันเป็น deterministic lookup จาก input

ในทางปฏิบัติ n-gram เป็น deterministic lookup จาก input tokens — เหมือน KV cache ไม่ใช่ parameters ปกติ — prefill ดึง n-grams ที่จำเป็น (fraction เล็กมากของ 51B) เข้า RAM/VRAM แล้ว reuse ซ้ำๆ ตอน decode


4️⃣ Muon + AdamW — optimizer คนละตัวต่อ weight category​

ข้อดี​

Qwen3.8-Flash-Next ใช้ Muon กับ 2D weight matrices (linear maps) และ AdamW กับ:

  • Input embeddings
  • N-gram embeddings
  • Output head
  • MoE router
  • Low-rank projections ของ GR

ผลที่ได้:

  • No batch-size warmup — start ที่ target batch size ทันที
  • Zero loss spikes ที่ production LR (8× scale stress test)
  • Median gradient norm ลดลง ~2× vs AdamW baseline, p99.9 ลด 4.2×
  • 1,000-step window std ลด 4.3-4.7×

ที่สำคัญคือ Muon ทำให้ ไม่ต้องพึ่ง qk-clip หรือ SwiGLU-clip ที่ Kimi/others ใช้แก้ instabilities

ข้อเสีย/ข้อจำกัด​

  1. Fused matrices ต้อง split — qkv projection, SwiGLU fc1, GDN input proj ต้อง orthogonalize แยก เพราะ "Orthogonalizing the fused matrix muddles updates to operators with very different gradient statistics"
  2. Elongated shapes ยังต้องใช้ AdamW — paper บอกชัดว่า "low-rank projections of GR likewise perform better with AdamW. We attribute this to their very elongated shape"
  3. N-gram embedding = AdamW (weight decay off) — ต้องแยก config ระหว่าง optimizer groups
  4. Muon คนเดียวไม่พอ — "Muon on its own has roughly twice the median norm" ใน stress test — ต้องจับคู่กับ GR/Muon+GR

สิ่งที่น่าสนใจ​

เหตุผลที่แยก optimizer ต่อ weight category เป็นเรื่องที่น่าสนใจ — paper แสดงให้เห็นว่า "gradient statistics" ของ weight แต่ละประเภทต่างกันมาก:

  • 2D linear maps → orthogonalize (Muon)
  • 1D scalar/long-tail → AdamW
  • Embedding tables → AdamW (low precision friendly)

ที่น่าทึ่ง: Muon + Gated Residual รวมกันคือ combo ที่ทำให้ zero loss spike — ไม่ใช่แค่ตัวใดตัวหนึ่ง


Community reactions — controversy + rebut​

หลังเปิดตัว มี debate ใน community เรื่อง design choices:

ฝั่ง criticism (AdrienneNoctis, HF Discussion #24)​

คนนี้โพสต์บทวิจารณ์ยาว (เขียนเป็นภาษาอังกฤษปนฝรั่งเศส) ว่า:

"Hidden size 2560, expert intermediate 640, 10 experts per token, 1 shared — the living tissue of your model totals roughly 4B active parameters, not the six you whisper in marketing. ... 512 experts of 640 intermediate — you did not build a Mixture-of-Experts; you built a Mixture-of-Excuses."

และโจมตี n-gram ว่า:

"20,000,000 memorized trigrams. This is not reasoning — this is a lookup table, a cane borrowed from the DeepSeek n-gram wardrobe. When the test asks a question whose trigram pattern lives in the table, the dummy is steered to the memorized vector and 'fires without thinking.'"

โพสต์นี้ได้รับ 6 reactions (รูปหน้าตกใจ/เศร้า) แต่โดน counter-argument ว่า "AI slop discussion" เพราะยาวเกินไปและมี rhetoric หนัก

Rebut จาก paper​

ข้อโจมตี "4B dummy" ไม่ตรงกับ spec paper — paper ระบุ:

  • 6B activated per token (เป็นทางการ)
  • 10 routed + 1 shared expert ที่ 640 intermediate
  • Hidden 2560, 48 layers
  • Math: ~6B ตามที่ claim

ข้อโจมตี "stolen n-gram":

  • Paper cites DeepSeek Engrams ใน references — ระบุแหล่งที่มาชัดเจน ไม่ได้ "ซ่อน"
  • Qwen's innovation คือ position-aware layering (PLE) + sliding window hash แทน fixed n-gram
  • และ paper ยอมรับตรงๆ ว่า "loss optimum diverges from accuracy optimum" — ไม่ได้อ้างว่า n-gram คือ reasoning

ฝั่ง user testing​

มีคนทดสอบจริงใน r/LocalLLM (เปรียบเทียบ Q4 vs 27B Q8):

มิติQwen3.8 27B Q8Qwen3.8 Flash-Next IQ4_XSWinner
Tool-eval hard mode (100)917727B
VRAM/RAM68 GB95 GB VRAM + 51 GB RAM27B
State/structured output errorsปกติ"noticeably more"27B

อีกคน (lkarlslund) รายงานตรงข้าม:

"The large MoE won: less time, better results and overall lower power consumption. The downside is that you need a huge setup to run it — but if you do have the rig, there is a clear winner: the MoE we've all been waiting for."

"Both models also waste too much time on xhigh, and low is not worth it — more tool calls so stick to medium."

ความเห็นต่างกันน่าจะเพราะ day-0 quant ยังไม่ stable — LobsterWeary2675 บอกชัดว่า "Could be cause its a day 0 quant and not completely supported yet. But I'll stick with 27B for now"

ฝั่ง license complaint​

HF Discussion #18 (8 reactions รูปหน้าตกใจ): คนถามว่าทำไมไม่ใช่ Apache 2.0 — license ใหม่ (Qwen Community License 1.0) trigger 100M MAU หรือ $20M revenue → ต้องแสดงชื่อ model, และ "Model as a Service" ต้องขอ license แยก ถ้าจะ commercial


สรุปเปรียบเทียบรวม​

นวัตกรรมข้อดีข้อเสียเมื่อไหร่ควรสนใจ
QSA7.6× prefill @ 1M, ไม่เสีย short ctx qualityต้อง fused kernel, ไม่ช่วย short ctxทำ long-context agent (RAG, code)
Gated ResidualStable training, LR/batch สูงขึ้นMemory traffic เพิ่ม, sparse write พังตอน post-trainingสเกลโมเดลใหญ่ๆ
N-gramOffloadable, scale params ถูกLoss ≠ accuracy optimum, storage costResource-constrained serving
Muon + AdamWZero loss spike, no batch warmupFused matrices ต้อง splitTraining stability เป็น priority

สำหรับคนอ่านตัดสินใจ​

ถ้าคุณ:

  • ทำ long-context inference → QSA คือ win ที่ชัดเจนที่สุด
  • train โมเดลใหญ่ → Gated Residual + Muon combo worth ศึกษา (ดู ablation ใน paper section 3.3)
  • serve โมเดลบน hardware จำกัด → N-gram offload เป็น pattern ที่น่าสนใจ แต่ระวัง loss vs accuracy divergence
  • แค่อยากได้ reasoning model → paper-level innovations ไม่สำคัญ — ดู benchmark (Qwen3.8-27B อาจเพียงพอสำหรับ hardware < 64GB)

สิ่งที่ paper ทำให้ผมเชื่อถือมากขึ้น:

  • ยอมรับ limitation ตรงๆ (loss ≠ accuracy, sparse write พังตอน post-training)
  • Ablation ครบทุก design choice
  • ไม่ claim "best in class" — แค่รายงานสิ่งที่ work

สิ่งที่ทำให้ผมสงสัย:

  • "1/9 training cost" claim เทียบกับ Qwen3.7-Plus — ตัวเลขนี้ น่า verify ว่ารวมส่วนไหนบ้าง
  • N-gram table ที่ offloadable — จะใช้กับ framework ไหนได้บ้าง? SGLang/vLLM support ระดับไหน
  • License Qwen Community License 1.0 — ข้อจำกัดจริง สำหรับ commercial use ที่ยังไม่มีคน summarize ชัด

ถ้ามีเวลา จะลอง deep-dive ใน 3 ข้อนี้ที่ blog ถัดไป


อ้างอิง​

Official:

Community:

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...