TypeSafe Jev — System One Model ที่ไม่ generate text แต่ลด token ของ agent ได้จริง
สารบัญ
- TL;DR
- Jev คืออะไร — ไม่ใช่ LLM
- 3 primitives
- Architecturally ต่างจาก LLM ยังไง
- Benchmark — vendor-reported
- Token reduction จริงไหม? — ใช่, แต่เฉพาะ use case
- กลไกของ LiteLLM
jev-compaction - ต่างจาก
/compactปกติของ Claude Code ยังไง - เมื่อไหร่ได้ผล / เมื่อไหร่ไม่ได้ผล
- Workflow ที่ผมคิดว่า fit กับ homelab ของผมเอง
- ข้อจำกัดที่ควรรู้ก่อน adopt
- สรุป
- อ้างอิง
"Jev ไม่ใช่ LLM — มันไม่ generate text, ไม่ parse string, และ schema-lock output จน type error เป็นไปไม่ได้. คำถามคือ: เอามันไปทำอะไรใน pipeline ของเราได้บ้าง?"
LiteLLM guardrail blog — Jev + LiteLLM compaction
DataCamp explainer — System One Model + benchmarks
tamaratran/fast-jev-compaction — Claude Code plugin ที่ hook /compact (4.2k stars)
ผมเพิ่งเห็น post ใน Facebook ของ Geek Consult บอกว่า TypeSafe ปล่อย skill ให้ใช้ Jev ผ่าน Codex/Claude ได้แล้ว + คนเริ่มเอาไปใช้แล้ว "ลด token ได้จริง"
ทีแรกผมคิดว่าเป็น hype แบบเดิม — แต่พออ่าน docs + benchmark + community projects แล้วพบว่า มันเป็นแนวคิดที่ต่างจาก LLM จริงๆ และ token reduction ในบาง use case ไม่ใช่ marketing
TL;DR
- Jev = "System One Model" ตัวแรกของ TypeSafe AI (ออก early access 15 ก.ย. 2026) — ไม่ใช่ LLM, ไม่ generate text
- รับ state + typed questions → ตอบ typed values + calibrated probabilities ใน parallel (latency 70-500ms, ไม่ใช่ 3-30s)
- ราคา $0.042/M input, output ฟรี — เทียบ GPT-5.6 Terra (~76x แพงกว่า), GPT-5.6 Sol/Claude Opus 5 (1-2 orders of magnitude แพงกว่า)
- 0% structured output error, 0% tool call error (TypeSafe-reported) — เทียบ Opus 5 (5.73% structured) / GPT-5.6 Sol (17% tool call) / Haiku 4.5 (45.5% structured)
- Token reduction จริง สำหรับ tool-call-heavy agent: LiteLLM
jev-compactionguardrail ถาม Jev 1 คำถามต่อ tool result เก่า ("ยังจำเป็นต่อคำถามล่าสุดไหม?") → drop ที่ score < 0.2, ลด input tokens ใน session ยาว - Plugin tamaratran/fast-jev-compaction hook เข้า Claude Code
/compact(4.2k stars) — แทน Claude summarize ด้วย Jev-pruned tool calls - ข้อจำกัดที่ควรรู้ก่อน adopt: ไม่มี rationale (debug ยาก), answer space ต้องนิ่ง, benchmark เป็น vendor-reported, ไม่ใช่ drop-in แทน LLM
Jev คืออะไร — ไม่ใช่ LLM
จาก docs.typesafe.ai/introduction:
"Jev is TypeSafe's flagship model and the first System One model. System One models are built to make fast, structured decisions that software can use directly. Jev evaluates typed questions against a state and returns structured results directly. No text generation, no parsing."
ชื่อ "System One" ยืมจาก Thinking, Fast and Slow ของ Daniel Kahneman — System 1 = ตัดสินใจเร็ว-สัญชาตญาณ, System 2 = ช้า-รอบคอบ. LLM ทุกวันนี้อยู่ฝั่ง System 2 (multi-second reasoning, chain-of-thought). TypeSafe โดย Diogo Almeida (อดีต OpenAI, co-inventor ของ ChatGPT) ตั้งข้อสังเกตว่า:
Software ส่วนใหญ่ต้องการคำตอบแบบ System 1 แต่คนกลับเอา System 2 models ไป coerce ให้ตอบแบบ structured แล้ว parse กลับมา
นั่นคือ root ของปัญหา — และ Jev คือคำตอบที่ TypeSafe propose
3 primitives
จาก docs.typesafe.ai/primitives:
| Primitive | ถามว่า | Output |
|---|---|---|
| Choice | เลือกจาก options ที่กำหนด | choice + probabilities + confidence |
| Score | ให้คะแนนตาม rubric | score + probabilities + confidence |
| Noul | statement นี้จริงไหม | noul (0-1) |
ตัวอย่างจริง (จาก LiteLLM blog):
// Routing support ticket
{
"state": { "message": "I was charged twice for the same order", "history": [...] },
"questions": [
{ "type": "Choice", "options": ["billing", "technical", "account"], "id": "department" },
{ "type": "Score", "range": [0, 100], "id": "urgency" }
]
}
// → Jev responds (parallel, in single request):
// {
// "department": { "choice": "billing", "probabilities": {...}, "confidence": 0.94 },
// "urgency": { "score": 87.3, "probabilities": {...}, "confidence": 0.81 }
// }
ทั้งสองคำถามตอบใน request เดียว, schema lock ไว้ตั้งแต่ก่อนยิง, parse ไม่ได้หลุดเพราะ output ไม่ใช่ string
Architecturally ต่างจาก LLM ยังไง
จาก DataCamp → Jev: TypeSafe's System One Model Explained:
| Property | Frontier LLM (GPT-5.6, Opus 5) | Jev |
|---|---|---|
| Output | Generated strings, parse ต่อ | Typed values, schema-guaranteed |
| Sampling | Sequential, one token at a time | Parallel, single query |
| Input price / 1M | $0.20-$10 | $0.042 |
| Output price | ~5x input | Free |
| Latency | 3-329s | 70-500ms |
| Structured output error | 0.58% - 45.5% | 0% |
| Tool call error | up to 17% | 0% |
| Confidence | Overconfident, inconsistent | Calibrated per output |
Training ใช้วิธีที่ TypeSafe เรียกว่า Reinforcement Learning for Calibrated Decisions (RLCD) — optimize ตรงๆ กับ calibrated probability บน decision tasks, ไม่ใช่ RLHF (optimize human-preferred text) หรือ RLVR (optimize verifiable code/text output)
ผลคือ 0% structured output error โดย construction — เพราะ output ถูก constrain ด้วย schema ที่กำหนดก่อน call, type error หรือ hallucinated JSON field เป็นไปไม่ได้ทางคณิตศาสตร์
Benchmark — vendor-reported
จาก DataCamp (เป็น TypeSafe's 4-workflow benchmark: security incident, agent-trace observability, invoice processing, customer service — ทุก workflow score เทียบ average prediction ของ GPT-6 Astra + Claude Fable):
| Model | Accuracy | Cost/case | Latency |
|---|---|---|---|
| Jev | 67.8% | $0.0004 | 0.4s |
| GPT-5.6 Terra | 67.9% | $0.0304 | 10.1s |
| GPT-5.6 Sol | 74.1% | $0.0836 | 23.3s |
| Claude Opus 5 | 73.1% | $0.1761 | 37.8s |
Jev เทียบเท่า Terra ในแง่ accuracy (67.8% vs 67.9%) — แต่ถูกกว่า ~76 เท่า, เร็วกว่า ~25 เท่า
Sol กับ Opus 5 ยังนำ accuracy อยู่ 5-6 จุด — ถ้าต้องการ peak accuracy ที่ volume ต่ำ, frontier LLM ยังชนะ
Honest caveat: ตัวเลขทั้งหมดเป็น vendor-reported — TypeSafe ทำ benchmark เอง, independent reproduction ยังไม่ออก. ผมเลยอ่านด้วยน้ำหนัก ~60% ของที่ TypeSafe claim แล้วกัน
Token reduction จริงไหม? — ใช่, แต่เฉพาะ use case
Jev เองไม่ได้ "ลด token" โดยตรง — มันทำงานต่างกับ LLM. Token reduction ที่คนพูดถึงคือ LiteLLM guardrail ใหม่ที่ใช้ Jev ตัดสินว่า tool result ไหนใน conversation ที่ยาวแล้วไม่จำเป็นอีก
กลไกของ LiteLLM jev-compaction
จาก docs.litellm.ai/blog/typesafe-jev-compaction:
"TypeSafe Jev helps LiteLLM spot tool results that are no longer needed. LiteLLM replaces those results with a short notice before calling the model. This is called compaction, and it can reduce the input tokens used by long conversations."
Workflow:
- ก่อนยิง request → LiteLLM ส่ง tool results เก่าๆ (ที่ไม่ใช่ exchange ล่าสุด) ให้ Jev
- Jev ตอบ 1 คำถามต่อ tool result: "Does the bot still need this to answer the user's latest question?"
- ถ้า score < 0.2 (default threshold) → แทนที่ tool result ด้วย
[Tool result removed by TypeSafe compaction: judged no longer relevant to the current task] - Tool calls + IDs คงไว้ (structure ของ conversation ไม่พัง), system/user messages ไม่แตะ, last assistant exchange กันไว้
Config:
# config.yaml
guardrails:
- guardrail_name: jev-compaction
litellm_params:
guardrail: typesafe
mode: pre_call
api_key: os.environ/TYPESAFE_API_KEY
optional_params:
relevance_threshold: 0.2
Result fields ที่ log ดูได้:
| Field | บอกอะไร |
|---|---|
exchanges_evaluated | Jev check กี่ exchange |
exchanges_dropped | ถูกแทนที่กี่ exchange |
chars_removed | ลดไปกี่ char |
model | Jev model ที่ใช้ |
ต่างจาก /compact ปกติของ Claude Code ยังไง
จาก explainx.ai review:
Claude /compact (default) | fast-jev-compaction (Jev-based) | |
|---|---|---|
| Mechanism | Ask Claude summarize session | Ask Jev delete tool calls ที่ไม่ต้องการ |
| Failure mode | Paraphrase ตัด exact file path / error string / constraint ที่จำเป็นทิ้ง | ลบทั้งก้อน, อันที่เหลือเป๊ะตามต้นฉบับ |
| Side effect | LLM call ราคาแพง + parse ใหม่ | Jev call ราคาถูก (~$0.0004) |
| What survives | Summarized version (lossy) | Original word-for-word (deletions only) |
Plugin tamaratran/fast-jev-compaction hook เข้า /compact ของ Claude Code (ต้อง Claude Code ≥ 2.1.274 + CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 + TYPESAFE_API_KEY) → replace Claude's built-in summary ด้วย Jev-scored deletions
คนใน r/ClaudeCode บอกว่า thread ที่ชี้ repo นี้ได้ 300+ upvotes / 80+ comments ภายใน 1 สัปดาห์ — เป็นสัญญาณว่ามันตอบ pain point จริง, ไม่ใช่ hype
เมื่อไหร่ได้ผล / เมื่อไหร่ไม่ได้ผล
| Scenario | ได้ผล? | เหตุผล |
|---|---|---|
| Agent ที่เรียก tools เยอะ (debug, deploy, long workflow) | ✅ | Tool results เก่าๆ มักถูกแทนที่ได้เยอะ |
| Session สั้นๆ (chat ล้วน, few turns) | ❌ | ไม่มี tool results ให้ลบ |
| Bot ที่ tool call เดียวตอบได้ (FAQ) | ❌ | ไม่มี context เก่าให้ prune |
| Coding agent (Codex/Claude Code) ที่รัน test / read file / grep หลายรอบ | ✅ | ตรงกับ pattern ของ plugin |
| Agent ที่ user เปลี่ยน topic กลางทาง | ✅✅ | Tool results เก่าเกี่ยวกับ topic เก่า = drop ได้เกือบหมด |
Workflow ที่ผมคิดว่า fit กับ homelab ของผมเอง
ใน homelab ผมมี LiteLLM proxy ที่ route ระหว่าง MiniMax (managed) + self-hosted vLLM บน DGX Spark + Claude/Codex ผ่าน client. Use cases ที่ Jev น่าจะช่วย:
- Tool-result compaction ใน LiteLLM — wire
jev-compactionguardrail เข้า proxy ที่มีอยู่แล้ว, เปิดให้ทีม bot4k (ที่ tool call หนักมาก) เป็น team-level opt-in - Routing decisions ใน bot4k — currently ใช้ heuristic / hardcoded rules; Jev อาจจะ replace ตรงนี้ด้วย Choice primitive (เลือก strategy: mean-reversion / momentum / hold) ที่ score < confidence → fallback ไป heuristic
- Score-based spam filter สำหรับ Telegram webhook — incoming message → Score ความน่าเชื่อถือ + Choice หมวด (chat / command / noise) → route ต่อหรือ drop
แต่ก่อนเขียน comment แบบนี้ ผมยังไม่ได้ลองจริง — เป็นแค่ plan. ตัวเลขที่ผมยังไม่รู้: Jev ตอบ routing question ใน production traffic ได้ดีแค่ไหน, ตัวอย่างใน DataCamp เป็น 4 workflows ของ TypeSafe เอง ยังไม่ใช่ bot4k
ข้อจำกัดที่ควรรู้ก่อน adopt
- ไม่มี rationale — Jev ตอบเป็นตัวเลข/enum อย่างเดียว, ไม่มี natural-language explanation. Debug ลำบาก, audit ใน regulated domain (finance, health) ทำไม่ได้
- Answer space ควรนิ่ง — ถ้า use case เป็น open-ended (chat, code generation) ใช้ Jev ไม่ได้. ต้องรู้ categories/scores ที่ต้องการล่วงหน้า
- Benchmark เป็น vendor-reported — ตัวเลขจาก DataCamp อ้างอิง TypeSafe's own evaluation, independent reproduction ยังไม่ออก
- Compaction ไม่ได้ลดเสมอ — ถ้า Jev ตัดสินว่าทุก tool result ยังจำเป็น, request ไม่ได้เล็กลง. ต้อง monitor
exchanges_droppedvsexchanges_evaluatedratio - ค่า Jev ต้องหักลบ savings — Jev call ฟรี output แต่มี input cost. ถ้า drop ratio ต่ำ, total อาจแพงกว่าไม่ compact
- Schema design is on you — Jev ไม่ช่วยออกแบบ primitive ให้. ถ้า ask คำถามที่ scope กว้างเกิน ("rate this startup pitch") คำตอบจะไม่ reliable — docs บอกให้ decompose เป็นคำถามเล็กๆ (market / feasibility / differentiation) แล้วรวมใน code
[!NOTE] Pattern ที่ผมจะใช้ถ้า adopt:
- Jev ไม่ใช่ LLM killer — ใช้เป็น decision layer ก่อน LLM (routing, filtering, scoring) ที่ schema นิ่งและ volume สูง
- LLM ยังจำเป็นสำหรับ generation, open-ended reasoning, code writing
- คิด Jev ในแง่ "ทำให้ LLM call น้อยลง" ไม่ใช่ "แทน LLM"
สรุป
Jev ไม่ใช่ hype เปล่า — architecturally ต่างจาก LLM จริง (parallel sampling + schema-locked output + calibrated probability), และตัวเลขที่ claim สอดคล้องกับ design. แต่ adopt ต้องระวัง:
- ใช้กับ decision/routing/scoring ที่ schema นิ่ง, ไม่ใช่ generation
- Token reduction ใน agent context เกิดจาก LiteLLM guardrail + plugin แยก (เช่น
fast-jev-compaction) ไม่ใช่ตัว Jev เอง - ตัวเลข benchmark ต้อง verify ใน workload ของตัวเองก่อนเชื่อ vendor claim
- ดู primitive examples ก่อนเขียน app — schema design เป็นของ dev, ไม่ใช่ของ model
ถ้ามีเวลาสัปดาห์หน้าผมจะลอง wire jev-compaction เข้า LiteLLM proxy แล้ววัดจริงบน traffic bot4k — ถ้าตัวเลขออกมาตรงกับที่ TypeSafe claim ผมจะเขียน follow-up
อ้างอิง
- typesafe-ai/skills (GitHub) — Claude Code plugin + skills.sh installer, 1.4k stars
- TypeSafe docs — Introduction — System One model + 3 primitives
- TypeSafe docs — System One concept — Kahneman framing + calibration
- DataCamp — Jev: System One Model Explained — architecture + benchmark
- LiteLLM blog — Reduce agent context with TypeSafe Jev — guardrail workflow
- tamaratran/fast-jev-compaction (GitHub) — Claude Code plugin (4.2k stars)
- explainx.ai — fast-jev-compaction review — Reddit thread context, what actually happens
- TypeSafe AI homepage — product overview
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee