Skip to main content

TypeSafe Jev — System One Model ที่ไม่ generate text แต่ลด token ของ agent ได้จริง

· 10 min read

"Jev ไม่ใช่ LLM — มันไม่ generate text, ไม่ parse string, และ schema-lock output จน type error เป็นไปไม่ได้. คำถามคือ: เอามันไปทำอะไรใน pipeline ของเราได้บ้าง?"

LiteLLM guardrail blog — Jev + LiteLLM compaction DataCamp explainer — System One Model + benchmarks tamaratran/fast-jev-compaction — Claude Code plugin ที่ hook /compact (4.2k stars)

ผมเพิ่งเห็น post ใน Facebook ของ Geek Consult บอกว่า TypeSafe ปล่อย skill ให้ใช้ Jev ผ่าน Codex/Claude ได้แล้ว + คนเริ่มเอาไปใช้แล้ว "ลด token ได้จริง"

ทีแรกผมคิดว่าเป็น hype แบบเดิม — แต่พออ่าน docs + benchmark + community projects แล้วพบว่า มันเป็นแนวคิดที่ต่างจาก LLM จริงๆ และ token reduction ในบาง use case ไม่ใช่ marketing

TL;DR​

  • Jev = "System One Model" ตัวแรกของ TypeSafe AI (ออก early access 15 ก.ย. 2026) — ไม่ใช่ LLM, ไม่ generate text
  • รับ state + typed questions → ตอบ typed values + calibrated probabilities ใน parallel (latency 70-500ms, ไม่ใช่ 3-30s)
  • ราคา $0.042/M input, output ฟรี — เทียบ GPT-5.6 Terra (~76x แพงกว่า), GPT-5.6 Sol/Claude Opus 5 (1-2 orders of magnitude แพงกว่า)
  • 0% structured output error, 0% tool call error (TypeSafe-reported) — เทียบ Opus 5 (5.73% structured) / GPT-5.6 Sol (17% tool call) / Haiku 4.5 (45.5% structured)
  • Token reduction จริง สำหรับ tool-call-heavy agent: LiteLLM jev-compaction guardrail ถาม Jev 1 คำถามต่อ tool result เก่า ("ยังจำเป็นต่อคำถามล่าสุดไหม?") → drop ที่ score < 0.2, ลด input tokens ใน session ยาว
  • Plugin tamaratran/fast-jev-compaction hook เข้า Claude Code /compact (4.2k stars) — แทน Claude summarize ด้วย Jev-pruned tool calls
  • ข้อจำกัดที่ควรรู้ก่อน adopt: ไม่มี rationale (debug ยาก), answer space ต้องนิ่ง, benchmark เป็น vendor-reported, ไม่ใช่ drop-in แทน LLM

Jev คืออะไร — ไม่ใช่ LLM​

จาก docs.typesafe.ai/introduction:

"Jev is TypeSafe's flagship model and the first System One model. System One models are built to make fast, structured decisions that software can use directly. Jev evaluates typed questions against a state and returns structured results directly. No text generation, no parsing."

ชื่อ "System One" ยืมจาก Thinking, Fast and Slow ของ Daniel Kahneman — System 1 = ตัดสินใจเร็ว-สัญชาตญาณ, System 2 = ช้า-รอบคอบ. LLM ทุกวันนี้อยู่ฝั่ง System 2 (multi-second reasoning, chain-of-thought). TypeSafe โดย Diogo Almeida (อดีต OpenAI, co-inventor ของ ChatGPT) ตั้งข้อสังเกตว่า:

Software ส่วนใหญ่ต้องการคำตอบแบบ System 1 แต่คนกลับเอา System 2 models ไป coerce ให้ตอบแบบ structured แล้ว parse กลับมา

นั่นคือ root ของปัญหา — และ Jev คือคำตอบที่ TypeSafe propose

3 primitives​

จาก docs.typesafe.ai/primitives:

Primitiveถามว่าOutput
Choiceเลือกจาก options ที่กำหนดchoice + probabilities + confidence
Scoreให้คะแนนตาม rubricscore + probabilities + confidence
Noulstatement นี้จริงไหมnoul (0-1)

ตัวอย่างจริง (จาก LiteLLM blog):

// Routing support ticket
{
"state": { "message": "I was charged twice for the same order", "history": [...] },
"questions": [
{ "type": "Choice", "options": ["billing", "technical", "account"], "id": "department" },
{ "type": "Score", "range": [0, 100], "id": "urgency" }
]
}

// → Jev responds (parallel, in single request):
// {
// "department": { "choice": "billing", "probabilities": {...}, "confidence": 0.94 },
// "urgency": { "score": 87.3, "probabilities": {...}, "confidence": 0.81 }
// }

ทั้งสองคำถามตอบใน request เดียว, schema lock ไว้ตั้งแต่ก่อนยิง, parse ไม่ได้หลุดเพราะ output ไม่ใช่ string

Architecturally ต่างจาก LLM ยังไง​

จาก DataCamp → Jev: TypeSafe's System One Model Explained:

PropertyFrontier LLM (GPT-5.6, Opus 5)Jev
OutputGenerated strings, parse ต่อTyped values, schema-guaranteed
SamplingSequential, one token at a timeParallel, single query
Input price / 1M$0.20-$10$0.042
Output price~5x inputFree
Latency3-329s70-500ms
Structured output error0.58% - 45.5%0%
Tool call errorup to 17%0%
ConfidenceOverconfident, inconsistentCalibrated per output

Training ใช้วิธีที่ TypeSafe เรียกว่า Reinforcement Learning for Calibrated Decisions (RLCD) — optimize ตรงๆ กับ calibrated probability บน decision tasks, ไม่ใช่ RLHF (optimize human-preferred text) หรือ RLVR (optimize verifiable code/text output)

ผลคือ 0% structured output error โดย construction — เพราะ output ถูก constrain ด้วย schema ที่กำหนดก่อน call, type error หรือ hallucinated JSON field เป็นไปไม่ได้ทางคณิตศาสตร์

Benchmark — vendor-reported​

จาก DataCamp (เป็น TypeSafe's 4-workflow benchmark: security incident, agent-trace observability, invoice processing, customer service — ทุก workflow score เทียบ average prediction ของ GPT-6 Astra + Claude Fable):

ModelAccuracyCost/caseLatency
Jev67.8%$0.00040.4s
GPT-5.6 Terra67.9%$0.030410.1s
GPT-5.6 Sol74.1%$0.083623.3s
Claude Opus 573.1%$0.176137.8s

Jev เทียบเท่า Terra ในแง่ accuracy (67.8% vs 67.9%) — แต่ถูกกว่า ~76 เท่า, เร็วกว่า ~25 เท่า

Sol กับ Opus 5 ยังนำ accuracy อยู่ 5-6 จุด — ถ้าต้องการ peak accuracy ที่ volume ต่ำ, frontier LLM ยังชนะ

Honest caveat: ตัวเลขทั้งหมดเป็น vendor-reported — TypeSafe ทำ benchmark เอง, independent reproduction ยังไม่ออก. ผมเลยอ่านด้วยน้ำหนัก ~60% ของที่ TypeSafe claim แล้วกัน

Token reduction จริงไหม? — ใช่, แต่เฉพาะ use case​

Jev เองไม่ได้ "ลด token" โดยตรง — มันทำงานต่างกับ LLM. Token reduction ที่คนพูดถึงคือ LiteLLM guardrail ใหม่ที่ใช้ Jev ตัดสินว่า tool result ไหนใน conversation ที่ยาวแล้วไม่จำเป็นอีก

กลไกของ LiteLLM jev-compaction​

จาก docs.litellm.ai/blog/typesafe-jev-compaction:

"TypeSafe Jev helps LiteLLM spot tool results that are no longer needed. LiteLLM replaces those results with a short notice before calling the model. This is called compaction, and it can reduce the input tokens used by long conversations."

Workflow:

  1. ก่อนยิง request → LiteLLM ส่ง tool results เก่าๆ (ที่ไม่ใช่ exchange ล่าสุด) ให้ Jev
  2. Jev ตอบ 1 คำถามต่อ tool result: "Does the bot still need this to answer the user's latest question?"
  3. ถ้า score < 0.2 (default threshold) → แทนที่ tool result ด้วย [Tool result removed by TypeSafe compaction: judged no longer relevant to the current task]
  4. Tool calls + IDs คงไว้ (structure ของ conversation ไม่พัง), system/user messages ไม่แตะ, last assistant exchange กันไว้

Config:

# config.yaml
guardrails:
- guardrail_name: jev-compaction
litellm_params:
guardrail: typesafe
mode: pre_call
api_key: os.environ/TYPESAFE_API_KEY
optional_params:
relevance_threshold: 0.2

Result fields ที่ log ดูได้:

Fieldบอกอะไร
exchanges_evaluatedJev check กี่ exchange
exchanges_droppedถูกแทนที่กี่ exchange
chars_removedลดไปกี่ char
modelJev model ที่ใช้

ต่างจาก /compact ปกติของ Claude Code ยังไง​

จาก explainx.ai review:

Claude /compact (default)fast-jev-compaction (Jev-based)
MechanismAsk Claude summarize sessionAsk Jev delete tool calls ที่ไม่ต้องการ
Failure modeParaphrase ตัด exact file path / error string / constraint ที่จำเป็นทิ้งลบทั้งก้อน, อันที่เหลือเป๊ะตามต้นฉบับ
Side effectLLM call ราคาแพง + parse ใหม่Jev call ราคาถูก (~$0.0004)
What survivesSummarized version (lossy)Original word-for-word (deletions only)

Plugin tamaratran/fast-jev-compaction hook เข้า /compact ของ Claude Code (ต้อง Claude Code ≥ 2.1.274 + CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 + TYPESAFE_API_KEY) → replace Claude's built-in summary ด้วย Jev-scored deletions

คนใน r/ClaudeCode บอกว่า thread ที่ชี้ repo นี้ได้ 300+ upvotes / 80+ comments ภายใน 1 สัปดาห์ — เป็นสัญญาณว่ามันตอบ pain point จริง, ไม่ใช่ hype

เมื่อไหร่ได้ผล / เมื่อไหร่ไม่ได้ผล​

Scenarioได้ผล?เหตุผล
Agent ที่เรียก tools เยอะ (debug, deploy, long workflow)✅Tool results เก่าๆ มักถูกแทนที่ได้เยอะ
Session สั้นๆ (chat ล้วน, few turns)❌ไม่มี tool results ให้ลบ
Bot ที่ tool call เดียวตอบได้ (FAQ)❌ไม่มี context เก่าให้ prune
Coding agent (Codex/Claude Code) ที่รัน test / read file / grep หลายรอบ✅ตรงกับ pattern ของ plugin
Agent ที่ user เปลี่ยน topic กลางทาง✅✅Tool results เก่าเกี่ยวกับ topic เก่า = drop ได้เกือบหมด

Workflow ที่ผมคิดว่า fit กับ homelab ของผมเอง​

ใน homelab ผมมี LiteLLM proxy ที่ route ระหว่าง MiniMax (managed) + self-hosted vLLM บน DGX Spark + Claude/Codex ผ่าน client. Use cases ที่ Jev น่าจะช่วย:

  1. Tool-result compaction ใน LiteLLM — wire jev-compaction guardrail เข้า proxy ที่มีอยู่แล้ว, เปิดให้ทีม bot4k (ที่ tool call หนักมาก) เป็น team-level opt-in
  2. Routing decisions ใน bot4k — currently ใช้ heuristic / hardcoded rules; Jev อาจจะ replace ตรงนี้ด้วย Choice primitive (เลือก strategy: mean-reversion / momentum / hold) ที่ score < confidence → fallback ไป heuristic
  3. Score-based spam filter สำหรับ Telegram webhook — incoming message → Score ความน่าเชื่อถือ + Choice หมวด (chat / command / noise) → route ต่อหรือ drop

แต่ก่อนเขียน comment แบบนี้ ผมยังไม่ได้ลองจริง — เป็นแค่ plan. ตัวเลขที่ผมยังไม่รู้: Jev ตอบ routing question ใน production traffic ได้ดีแค่ไหน, ตัวอย่างใน DataCamp เป็น 4 workflows ของ TypeSafe เอง ยังไม่ใช่ bot4k

ข้อจำกัดที่ควรรู้ก่อน adopt​

  1. ไม่มี rationale — Jev ตอบเป็นตัวเลข/enum อย่างเดียว, ไม่มี natural-language explanation. Debug ลำบาก, audit ใน regulated domain (finance, health) ทำไม่ได้
  2. Answer space ควรนิ่ง — ถ้า use case เป็น open-ended (chat, code generation) ใช้ Jev ไม่ได้. ต้องรู้ categories/scores ที่ต้องการล่วงหน้า
  3. Benchmark เป็น vendor-reported — ตัวเลขจาก DataCamp อ้างอิง TypeSafe's own evaluation, independent reproduction ยังไม่ออก
  4. Compaction ไม่ได้ลดเสมอ — ถ้า Jev ตัดสินว่าทุก tool result ยังจำเป็น, request ไม่ได้เล็กลง. ต้อง monitor exchanges_dropped vs exchanges_evaluated ratio
  5. ค่า Jev ต้องหักลบ savings — Jev call ฟรี output แต่มี input cost. ถ้า drop ratio ต่ำ, total อาจแพงกว่าไม่ compact
  6. Schema design is on you — Jev ไม่ช่วยออกแบบ primitive ให้. ถ้า ask คำถามที่ scope กว้างเกิน ("rate this startup pitch") คำตอบจะไม่ reliable — docs บอกให้ decompose เป็นคำถามเล็กๆ (market / feasibility / differentiation) แล้วรวมใน code

[!NOTE] Pattern ที่ผมจะใช้ถ้า adopt:

  • Jev ไม่ใช่ LLM killer — ใช้เป็น decision layer ก่อน LLM (routing, filtering, scoring) ที่ schema นิ่งและ volume สูง
  • LLM ยังจำเป็นสำหรับ generation, open-ended reasoning, code writing
  • คิด Jev ในแง่ "ทำให้ LLM call น้อยลง" ไม่ใช่ "แทน LLM"

สรุป​

Jev ไม่ใช่ hype เปล่า — architecturally ต่างจาก LLM จริง (parallel sampling + schema-locked output + calibrated probability), และตัวเลขที่ claim สอดคล้องกับ design. แต่ adopt ต้องระวัง:

  • ใช้กับ decision/routing/scoring ที่ schema นิ่ง, ไม่ใช่ generation
  • Token reduction ใน agent context เกิดจาก LiteLLM guardrail + plugin แยก (เช่น fast-jev-compaction) ไม่ใช่ตัว Jev เอง
  • ตัวเลข benchmark ต้อง verify ใน workload ของตัวเองก่อนเชื่อ vendor claim
  • ดู primitive examples ก่อนเขียน app — schema design เป็นของ dev, ไม่ใช่ของ model

ถ้ามีเวลาสัปดาห์หน้าผมจะลอง wire jev-compaction เข้า LiteLLM proxy แล้ววัดจริงบน traffic bot4k — ถ้าตัวเลขออกมาตรงกับที่ TypeSafe claim ผมจะเขียน follow-up

อ้างอิง​

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...