Skip to main content

ลด Cost ใน Hermes — คุม background_review fork ให้ไม่กิน token ฟรี

· 8 min read

T14 ของผมรัน Hermes มา 3 เดือน วันหนึ่งเปิดดู queue ใน ~/.hermes/pending/memory/ แล้วเจอ 48 ไฟล์ที่ค้างมาเป็นเดือน — บทความนี้คือการเดินย้อนจากอาการ "queue เต็ม" หาสาเหตุที่แท้จริง แล้วค่อย tune config ทีละขั้น

อาการเริ่มต้น​

ผมเปิด Telegram มาดูเครื่อง T14 แล้วเจอ ~/.hermes/pending/memory/ เต็มไปด้วย JSON files — ลอง ls | wc -l ได้ 48 ไฟล์:

$ ls ~/.hermes/pending/memory/ | wc -l
48

ไฟล์พวกนี้คือ proposal ที่ background_review fork เสนอเข้ามา แต่ยังไม่ apply ลง MEMORY.md / USER.md จริง ตรวจ payload หนึ่งไฟล์:

{
"id": "3570103d",
"subsystem": "memory",
"action": "replace",
"summary": "replace in memory: old: Machines:\n- T14\nnew: Machines:\n- T14\n- surge.lan (Debian/Proxmox VE) — runs Surge download manager + Filebrowser",
"origin": "background_review",
"payload": {
"action": "replace",
"target": "memory",
"old_text": "Machines:\n- T14",
"content": "Machines:\n- T14\n- surge.lan ..."
}
}

ข้างในมี fact ที่ verify แล้ว 37 ไฟล์ (background_review) + 11 ไฟล์ที่ตัว assistant tool เรียกเอง ดูไม่มี noise เลย — แต่ปัญหาไม่ใช่คุณภาพ ปัญหาคือ ผมไม่อยาก approve manual ทุกสัปดาห์ และ queue สะสมจนน่ากลัวว่า cost จะไต่ขึ้นโดยไม่รู้ตัว

ปัญหา: cost ไต่ขึ้นโดยไม่รู้ตัว​

หลังเห็น queue เต็ม ผมเริ่มสังเกตว่า cost ไต่ขึ้นเรื่อยๆ โดยไม่ได้ใช้งานหนักขึ้น T14 รัน Hermes ผ่าน Telegram gateway 24/7 มา 3 เดือด้วย custom gateway ที่ 10.0.0.155:4000/v1 ใช้ M3-CHATx-256k เป็น main model ตอนแรกคิดว่า fork จะถูกเพราะเป็น self-hosted — พอดู log ของจริงเจอ pattern ที่ไม่คาดคิด:

# Token usage จาก background_review fork (1 สัปดาห์)
auto spawn: ~30,000 input tokens/turn × 100 turns/week = 3M input tokens/week

3M tokens ต่อสัปดาห์ ตัวเลขนี้ใหญ่กว่าที่ผม assume (คิดว่า fork จะแค่ replay cache) ที่น่าตกใจคือ — ผมไม่เคย approve queue เลย ตั้งแต่ deploy ครั้งแรก fork ทำงานทุก turn แต่ commit ไม่ได้ก็ค้างอยู่ในนั้น

ข้อเท็จจริงเกี่ยวกับ background_review​

ขุด source code ใน agent/background_review.py:191-201:

The review fork runs on the MAIN model by default ("auto"), replaying the full conversation — already warm in the prompt cache, so cheap cache reads. Optimal and unchanged.

Comment เขียนว่า "warm cache = cheap" แต่จริงๆ มัน ขึ้นกับ 2 เงื่อนไข:

  1. Provider ตรงกับ parent → warm cache
  2. Model ตรงกับ parent → warm cache

ถ้าข้อใดข้อหนึ่งต่าง → fork จะ digest แทน full replay แต่ digest ก็ยังเป็น cold write ที่ provider คิดเงิน

1. มันทำอะไร​

ทุกครั้งที่ turn จบ fork จะ spawn ขึ้นมา snapshot conversation แล้วเสนอ add/replace memory entries ลง ~/.hermes/pending/memory/ ตัวหลัก conversation + prompt cache ไม่ถูกแตะ เลย — fork ทำงาน background

2. ปิดได้ แต่ manual ยังใช้ได้​

จาก comment ใน hermes_cli/config_defaults.py:1255-1258:

"background_review": {
# Master switch for automatic post-turn memory/skill review forks.
# false = skip automatic spawns (manual /refine still works).
"enabled": True,
}

และ run_agent.py:1914 ยืนยัน:

Explicit off-switch for automatic post-turn forks (auxiliary.background_review.enabled: false). Manual /refine still works — same contract as zeroing the nudge intervals

ดังนั้น enabled: false ปิด automatic spawn เท่านั้น — ตัว /refine slash command ยัง trigger fork ได้ตามต้องการ

3. Cost มาจากไหน​

ตัวเลข 3M tokens/week มาจาก 2 ส่วน:

  • Auto spawn ทุก turn (enabled: true) = ~30K input tokens/turn
  • Cold cache + digest เมื่อ provider/model ต่างจาก parent

Default (provider: auto) = inherit main model — warm cache, ถูก ถ้า pin model อื่น → cold cache + digest → แพงกว่า

ขั้นที่ 1 — ปิด auto-spawn​

ก่อนอื่นกวาด queue ที่ค้างก่อน:

rm -v /home/kongvut/.hermes/pending/memory/*.json

แล้วปิด auto-review:

hermes config set auxiliary.background_review.enabled false

ตรวจ config:

auxiliary:
background_review:
provider: auto
model: ''
base_url: ''
api_key: ''
timeout: 120
enabled: false # ← ตัด auto-spawn ตรงนี้

[!NOTE] ต้อง restart gateway ถึงจะ生效 — config snapshot ตอน import รัน hermes gateway restart จาก shell แยก

ขั้นที่ 2 — Cron prune queue ที่ไม่ approve​

ถ้าวันหลังอยากใช้ /refine กลับมา queue จะโตอีก script ตัวนี้กวาดทุกสัปดาห์แบบ silent:

#!/bin/bash
# ~/.hermes/scripts/prune-pending-memory.sh
target="$HOME/.hermes/pending/memory"

if [ ! -d "$target" ]; then
exit 0
fi

shopt -s nullglob
files=("$target"/*.json)
shopt -u nullglob

if [ ${#files[@]} -eq 0 ]; then
exit 0 # silent เมื่อ empty
fi

count=${#files[@]}
rm -f "${files[@]}"
echo "Pruned $count pending memory files from $target"

ตั้ง cron ผ่าน hermes cron create ในโหมด no-agent (= script IS the job, empty stdout = silent):

hermes cron create "every 7 days" \
--name "prune-pending-memory-weekly" \
--script "prune-pending-memory.sh" \
--no-agent \
--deliver "origin"

แล้วเปลี่ยน delivery เป็น local — ไม่ส่ง Telegram ทุกครั้งที่กวาด:

hermes cron edit <job_id> --deliver local

[!NOTE] ทำไมต้อง silent — ผมเปิด Telegram ทุกวัน ถ้า cron ส่งข้อความ "Pruned 0 files" มาทุกสัปดาห์ มันคือ noise ที่ผม ignore ในที่สุด เก็บ log ที่ ~/.hermes/cron/output/ แล้วตรวจเมื่อต้องการดีกว่า

ขั้นที่ 3 — Pin model routing ถ้าใช้ self-hosted​

ถ้า main model วิ่งผ่าน gateway ที่ base_url เดียวตลอด การ pin แบบนี้ช่วยให้ fork ใช้ cache เดียวกับ parent ได้:

auxiliary:
background_review:
provider: custom # ← pin
model: M3-CHATx-256k # ← pin
base_url: '' # ← inherit from parent
api_key: '' # ← inherit from parent
enabled: false

Logic ใน agent/background_review.py:355-358:

if not (task_provider and task_provider != "auto" and task_model):
return parent
if task_provider == (agent.provider or "") and task_model == (agent.model or ""):
return parent # same model/provider as parent -> not routed

คือถ้า provider + model ตรงกับ parent จะ route ผ่าน parent runtime → ได้ warm cache เต็มที่

ตารางเปรียบเทียบ:

SetupProvider vs mainCacheCost impact
provider: auto (default)เหมือน mainwarmถูกสุด ถ้า main = self-hosted
provider: custom + model ตรงตรงwarmถูกเท่ากัน + explicit
provider: custom + model ต่างต่างcold + digestแพงกว่า

ทำไมไม่ pin base_url ด้วย​

ผมลอง pin base_url ตายตัวในตอนแรก ผลคือถ้าวันหลังสลับ main ไป external provider (OpenRouter, Anthropic) base_url ของ fork จะยังชี้ไปที่ 10.0.0.155:4000 แต่ main ใช้ endpoint อื่นแล้ว → cold cache ทุกครั้ง

การ inherit จาก parent runtime ปลอดภัยกว่า เพราะ resolve_runtime_provider จะเลือก endpoint ตามที่ main กำลังใช้

[!WARNING] ถ้าใช้ external provider (OpenRouter/Anthropic) เป็นหลัก การ pin แบบนี้จะแพงกว่า auto เพราะ digest mode activate ทุกครั้ง (cold cache ไม่ reuse อะไร) ในกรณีนั้นใช้ provider: auto กลับมาเหมือนเดิมจะคุ้มกว่า

ผลลัพธ์​

หลัง apply 3 ขั้น + gateway restart:

ActionToken/week (ประมาณ)
ก่อน — auto-spawn ทุก turn~3M input tokens
หลัง — enabled: false0 (no fork)
หลัง — /refine ใช้เอง (optional)~600K cap per fork × N uses

Cost ลดลงประมาณ 80-90% สำหรับ review subsystem ของ session (ผมยังไม่ได้ measure ตัวเลขจริงใน billing — แต่ log usage ของ fork หายไปทันทีหลังปิด)

ถ้าวันหลังอยากใช้ /refine กลับมา → ยังทำได้ cost ต่อ use จะเป็น M3-CHATx-256k บน gateway เดียวกัน + warm cache

Bonus: nudge_interval ที่ไม่มีใครบอก​

อีก lever ที่ผมไม่เห็นใน docs ภายนอกแต่ verify ได้จาก source code (hermes_cli/config_defaults.py:1976-1979):

"Periodic built-in memory review. External providers with automatic turn/session extraction can set this to 0 and keep the small local store reserved for explicit high-frequency operational facts."

memory:
nudge_interval: 10 # default — periodic review ทุก 10 turns

nudge_interval คือการ ถาม agent ทุก N turns ว่ามี memory fact ใหม่ที่ควรเก็บมั้ย — ตัวเลขเล็กแต่ทุกครั้งที่ trigger จะมี LLM call เพิ่ม ถ้าใช้ external memory provider (Honcho, Mem0, etc.) ที่ทำ auto-extraction อยู่แล้ว การตั้ง 0 จะ ปิด periodic review ทับซ้อน:

memory:
nudge_interval: 0 # ผมตั้งแบบนี้ — Honcho ทำ auto-extraction อยู่แล้ว

Trade-off: agent จะ ไม่ถูกถาม เรื่อง memory ในทุก turn → ต้องพึ่ง external provider หรือ explicit user prompt เพื่อเพิ่ม fact ใหม่ ถ้าใช้แต่ built-in memory อย่างเดียว การตั้ง 0 จะทำให้ memory ไม่ auto-grow ให้เลือก trade-off เอง

สรุป​

3 ขั้นที่ผมใช้ลด cost ของ background_review fork: (1) ปิด auto-spawn ด้วย auxiliary.background_review.enabled: false (2) cron prune queue ที่ไม่ approve ทุก 7 วันแบบ silent และ (3) pin model routing ให้ตรงกับ parent ถ้าใช้ self-hosted gateway เพื่อให้ได้ warm cache

ผลลัพธ์ cost ลดลงประมาณ 80-90% สำหรับ review subsystem ของ session และ /refine slash command ยังใช้ได้ตามต้องการ — ไม่ได้ปิด fork แบบถาวร แค่ปิด automatic spawn เท่านั้น

สิ่งที่ผมยังไม่แน่ใจ​

  • digest mode ทำงานดีแค่ไหนเมื่อสลับ main ไป external provider — ผมยังไม่ได้ test จริง เพราะใช้ self-hosted
  • cold-write cost ของ pin model vs auto inherit — ถ้ามีโอกาสผมจะ measure ด้วย usage metrics

อ้างอิง​

แชร์บทความ
☕

เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว

Buy Me a Coffee
Loading...