ลด Cost ใน Hermes — คุม background_review fork ให้ไม่กิน token ฟรี
สารบัญ
- อาการเริ่มต้น
- ปัญหา: cost ไต่ขึ้นโดยไม่รู้ตัว
- ข้อเท็จจริงเกี่ยวกับ background_review
- 1. มันทำอะไร
- 2. ปิดได้ แต่ manual ยังใช้ได้
- 3. Cost มาจากไหน
- ขั้นที่ 1 — ปิด auto-spawn
- ขั้นที่ 2 — Cron prune queue ที่ไม่ approve
- ขั้นที่ 3 — Pin model routing ถ้าใช้ self-hosted
- ทำไมไม่ pin
base_urlด้วย - ผลลัพธ์
- Bonus: nudge_interval ที่ไม่มีใครบอก
- สรุป
- สิ่งที่ผมยังไม่แน่ใจ
- อ้างอิง
T14 ของผมรัน Hermes มา 3 เดือน วันหนึ่งเปิดดู queue ใน
~/.hermes/pending/memory/แล้วเจอ 48 ไฟล์ที่ค้างมาเป็นเดือน — บทความนี้คือการเดินย้อนจากอาการ "queue เต็ม" หาสาเหตุที่แท้จริง แล้วค่อย tune config ทีละขั้น
อาการเริ่มต้น
ผมเปิด Telegram มาดูเครื่อง T14 แล้วเจอ ~/.hermes/pending/memory/ เต็มไปด้วย JSON files — ลอง ls | wc -l ได้ 48 ไฟล์:
$ ls ~/.hermes/pending/memory/ | wc -l
48
ไฟล์พวกนี้คือ proposal ที่ background_review fork เสนอเข้ามา แต่ยังไม่ apply ลง MEMORY.md / USER.md จริง ตรวจ payload หนึ่งไฟล์:
{
"id": "3570103d",
"subsystem": "memory",
"action": "replace",
"summary": "replace in memory: old: Machines:\n- T14\nnew: Machines:\n- T14\n- surge.lan (Debian/Proxmox VE) — runs Surge download manager + Filebrowser",
"origin": "background_review",
"payload": {
"action": "replace",
"target": "memory",
"old_text": "Machines:\n- T14",
"content": "Machines:\n- T14\n- surge.lan ..."
}
}
ข้างในมี fact ที่ verify แล้ว 37 ไฟล์ (background_review) + 11 ไฟล์ที่ตัว assistant tool เรียกเอง ดูไม่มี noise เลย — แต่ปัญหาไม่ใช่คุณภาพ ปัญหาคือ ผมไม่อยาก approve manual ทุกสัปดาห์ และ queue สะสมจนน่ากลัวว่า cost จะไต่ขึ้นโดยไม่รู้ตัว
ปัญหา: cost ไต่ขึ้นโดยไม่รู้ตัว
หลังเห็น queue เต็ม ผมเริ่มสังเกตว่า cost ไต่ขึ้นเรื่อยๆ โดยไม่ได้ใช้งานหนักขึ้น T14 รัน Hermes ผ่าน Telegram gateway 24/7 มา 3 เดือด้วย custom gateway ที่ 10.0.0.155:4000/v1 ใช้ M3-CHATx-256k เป็น main model ตอนแรกคิดว่า fork จะถูกเพราะเป็น self-hosted — พอดู log ของจริงเจอ pattern ที่ไม่คาดคิด:
# Token usage จาก background_review fork (1 สัปดาห์)
auto spawn: ~30,000 input tokens/turn × 100 turns/week = 3M input tokens/week
3M tokens ต่อสัปดาห์ ตัวเลขนี้ใหญ่กว่าที่ผม assume (คิดว่า fork จะแค่ replay cache) ที่น่าตกใจคือ — ผมไม่เคย approve queue เลย ตั้งแต่ deploy ครั้งแรก fork ทำงานทุก turn แต่ commit ไม่ได้ก็ค้างอยู่ในนั้น
ข้อเท็จจริงเกี่ยวกับ background_review
ขุด source code ใน agent/background_review.py:191-201:
The review fork runs on the MAIN model by default ("auto"), replaying the full conversation — already warm in the prompt cache, so cheap cache reads. Optimal and unchanged.
Comment เขียนว่า "warm cache = cheap" แต่จริงๆ มัน ขึ้นกับ 2 เงื่อนไข:
- Provider ตรงกับ parent → warm cache
- Model ตรงกับ parent → warm cache
ถ้าข้อใดข้อหนึ่งต่าง → fork จะ digest แทน full replay แต่ digest ก็ยังเป็น cold write ที่ provider คิดเงิน
1. มันทำอะไร
ทุกครั้งที่ turn จบ fork จะ spawn ขึ้นมา snapshot conversation แล้วเสนอ add/replace memory entries ลง ~/.hermes/pending/memory/ ตัวหลัก conversation + prompt cache ไม่ถูกแตะ เลย — fork ทำงาน background
2. ปิดได้ แต่ manual ยังใช้ได้
จาก comment ใน hermes_cli/config_defaults.py:1255-1258:
"background_review": {
# Master switch for automatic post-turn memory/skill review forks.
# false = skip automatic spawns (manual /refine still works).
"enabled": True,
}
และ run_agent.py:1914 ยืนยัน:
Explicit off-switch for automatic post-turn forks (
auxiliary.background_review.enabled: false). Manual/refinestill works — same contract as zeroing the nudge intervals
ดังนั้น enabled: false ปิด automatic spawn เท่านั้น — ตัว /refine slash command ยัง trigger fork ได้ตามต้องการ
3. Cost มาจากไหน
ตัวเลข 3M tokens/week มาจาก 2 ส่วน:
- Auto spawn ทุก turn (
enabled: true) = ~30K input tokens/turn - Cold cache + digest เมื่อ provider/model ต่างจาก parent
Default (provider: auto) = inherit main model — warm cache, ถูก ถ้า pin model อื่น → cold cache + digest → แพงกว่า
ขั้นที่ 1 — ปิด auto-spawn
ก่อนอื่นกวาด queue ที่ค้างก่อน:
rm -v /home/kongvut/.hermes/pending/memory/*.json
แล้วปิด auto-review:
hermes config set auxiliary.background_review.enabled false
ตรวจ config:
auxiliary:
background_review:
provider: auto
model: ''
base_url: ''
api_key: ''
timeout: 120
enabled: false # ← ตัด auto-spawn ตรงนี้
[!NOTE] ต้อง restart gateway ถึงจะ生效 — config snapshot ตอน import รัน
hermes gateway restartจาก shell แยก
ขั้นที่ 2 — Cron prune queue ที่ไม่ approve
ถ้าวันหลังอยากใช้ /refine กลับมา queue จะโตอีก script ตัวนี้กวาดทุกสัปดาห์แบบ silent:
#!/bin/bash
# ~/.hermes/scripts/prune-pending-memory.sh
target="$HOME/.hermes/pending/memory"
if [ ! -d "$target" ]; then
exit 0
fi
shopt -s nullglob
files=("$target"/*.json)
shopt -u nullglob
if [ ${#files[@]} -eq 0 ]; then
exit 0 # silent เมื่อ empty
fi
count=${#files[@]}
rm -f "${files[@]}"
echo "Pruned $count pending memory files from $target"
ตั้ง cron ผ่าน hermes cron create ในโหมด no-agent (= script IS the job, empty stdout = silent):
hermes cron create "every 7 days" \
--name "prune-pending-memory-weekly" \
--script "prune-pending-memory.sh" \
--no-agent \
--deliver "origin"
แล้วเปลี่ยน delivery เป็น local — ไม่ส่ง Telegram ทุกครั้งที่กวาด:
hermes cron edit <job_id> --deliver local
[!NOTE] ทำไมต้อง silent — ผมเปิด Telegram ทุกวัน ถ้า cron ส่งข้อความ "Pruned 0 files" มาทุกสัปดาห์ มันคือ noise ที่ผม ignore ในที่สุด เก็บ log ที่
~/.hermes/cron/output/แล้วตรวจเมื่อต้องการดีกว่า
ขั้นที่ 3 — Pin model routing ถ้าใช้ self-hosted
ถ้า main model วิ่งผ่าน gateway ที่ base_url เดียวตลอด การ pin แบบนี้ช่วยให้ fork ใช้ cache เดียวกับ parent ได้:
auxiliary:
background_review:
provider: custom # ← pin
model: M3-CHATx-256k # ← pin
base_url: '' # ← inherit from parent
api_key: '' # ← inherit from parent
enabled: false
Logic ใน agent/background_review.py:355-358:
if not (task_provider and task_provider != "auto" and task_model):
return parent
if task_provider == (agent.provider or "") and task_model == (agent.model or ""):
return parent # same model/provider as parent -> not routed
คือถ้า provider + model ตรงกับ parent จะ route ผ่าน parent runtime → ได้ warm cache เต็มที่
ตารางเปรียบเทียบ:
| Setup | Provider vs main | Cache | Cost impact |
|---|---|---|---|
provider: auto (default) | เหมือน main | warm | ถูกสุด ถ้า main = self-hosted |
provider: custom + model ตรง | ตรง | warm | ถูกเท่ากัน + explicit |
provider: custom + model ต่าง | ต่าง | cold + digest | แพงกว่า |
ทำไมไม่ pin base_url ด้วย
ผมลอง pin base_url ตายตัวในตอนแรก ผลคือถ้าวันหลังสลับ main ไป external provider (OpenRouter, Anthropic) base_url ของ fork จะยังชี้ไปที่ 10.0.0.155:4000 แต่ main ใช้ endpoint อื่นแล้ว → cold cache ทุกครั้ง
การ inherit จาก parent runtime ปลอดภัยกว่า เพราะ resolve_runtime_provider จะเลือก endpoint ตามที่ main กำลังใช้
[!WARNING] ถ้าใช้ external provider (OpenRouter/Anthropic) เป็นหลัก การ pin แบบนี้จะแพงกว่า
autoเพราะ digest mode activate ทุกครั้ง (cold cache ไม่ reuse อะไร) ในกรณีนั้นใช้provider: autoกลับมาเหมือนเดิมจะคุ้มกว่า
ผลลัพธ์
หลัง apply 3 ขั้น + gateway restart:
| Action | Token/week (ประมาณ) |
|---|---|
| ก่อน — auto-spawn ทุก turn | ~3M input tokens |
หลัง — enabled: false | 0 (no fork) |
หลัง — /refine ใช้เอง (optional) | ~600K cap per fork × N uses |
Cost ลดลงประมาณ 80-90% สำหรับ review subsystem ของ session (ผมยังไม่ได้ measure ตัวเลขจริงใน billing — แต่ log usage ของ fork หายไปทันทีหลังปิด)
ถ้าวันหลังอยากใช้ /refine กลับมา → ยังทำได้ cost ต่อ use จะเป็น M3-CHATx-256k บน gateway เดียวกัน + warm cache
Bonus: nudge_interval ที่ไม่มีใครบอก
อีก lever ที่ผมไม่เห็นใน docs ภายนอกแต่ verify ได้จาก source code (hermes_cli/config_defaults.py:1976-1979):
"Periodic built-in memory review. External providers with automatic turn/session extraction can set this to 0 and keep the small local store reserved for explicit high-frequency operational facts."
memory:
nudge_interval: 10 # default — periodic review ทุก 10 turns
nudge_interval คือการ ถาม agent ทุก N turns ว่ามี memory fact ใหม่ที่ควรเก็บมั้ย — ตัวเลขเล็กแต่ทุกครั้งที่ trigger จะมี LLM call เพิ่ม ถ้าใช้ external memory provider (Honcho, Mem0, etc.) ที่ทำ auto-extraction อยู่แล้ว การตั้ง 0 จะ ปิด periodic review ทับซ้อน:
memory:
nudge_interval: 0 # ผมตั้งแบบนี้ — Honcho ทำ auto-extraction อยู่แล้ว
Trade-off: agent จะ ไม่ถูกถาม เรื่อง memory ในทุก turn → ต้องพึ่ง external provider หรือ explicit user prompt เพื่อเพิ่ม fact ใหม่ ถ้าใช้แต่ built-in memory อย่างเดียว การตั้ง 0 จะทำให้ memory ไม่ auto-grow ให้เลือก trade-off เอง
สรุป
3 ขั้นที่ผมใช้ลด cost ของ background_review fork: (1) ปิด auto-spawn ด้วย auxiliary.background_review.enabled: false (2) cron prune queue ที่ไม่ approve ทุก 7 วันแบบ silent และ (3) pin model routing ให้ตรงกับ parent ถ้าใช้ self-hosted gateway เพื่อให้ได้ warm cache
ผลลัพธ์ cost ลดลงประมาณ 80-90% สำหรับ review subsystem ของ session และ /refine slash command ยังใช้ได้ตามต้องการ — ไม่ได้ปิด fork แบบถาวร แค่ปิด automatic spawn เท่านั้น
สิ่งที่ผมยังไม่แน่ใจ
- digest mode ทำงานดีแค่ไหนเมื่อสลับ main ไป external provider — ผมยังไม่ได้ test จริง เพราะใช้ self-hosted
- cold-write cost ของ pin model vs auto inherit — ถ้ามีโอกาสผมจะ measure ด้วย
usagemetrics
อ้างอิง
- Hermes Agent Docs — official docs
agent/background_review.py— routing logic (lines 191-201, 350-358)hermes_cli/config_defaults.py— config schema + default values- Telegram Bot API 10.1 — unrelated แต่เป็นบริบทของ session ที่ผมใช้อยู่
เนื้อหานี้มีประโยชน์ไหม? ช่วยสนับสนุนค่ากาแฟให้ผู้เขียนสักแก้ว
Buy Me a Coffee