Four agents reviewed this in parallel -- correctness, GNOME integration, edge cases, security. Everything below was reproduced before being fixed; several findings that survived the first reading did not survive a probe and are not here. State machine, the two that mattered most. A pending permission prompt was erased by any subagent bookkeeping event: SubagentStop or the next PreToolUse recomputed the state from scratch, so a session sat at "working" with a dialog open and nothing ever raised it again. Blocked now outlives everything except evidence the question was answered. Separately, the stale-event guard refused whole events, including the subagent counter's increments and decrements -- but those are deltas and deltas commute, so a "+1" that lost a timestamp race left the count short and the batch freed the session while a subagent was still running. The guard now gates the state decision only. Corrupt or hostile state files could wedge the panel or take the hook down for every session: a non-numeric pid raised inside sweep_dead before the hook wrote its own file, so one bad byte stopped new sessions appearing at all. Numbers read back from disk are coerced, one unreadable file no longer aborts the sweep, and a stored timestamp far in the future -- corruption, or a clock stepped backwards by NTP -- no longer refuses every later event forever. Resource bounds, all in the compositor process. A state file was read whole with no size check: a symlink to /dev/zero took a test process past 4 GB in three seconds, which in gnome-shell ends the session. Sizes are checked before the read, sessions and zellij subprocesses are capped, labels ellipsize, and cwd and messages are truncated at the hook. Teardown hung off an overridden destroy(), which only runs when JS calls it. An actor destroyed any other way -- another extension rebuilding the panel boxes -- left the timer and the file monitor running against a disposed actor. It is a destroy signal now. The zellij child is killed rather than merely abandoned. The glyph was pinned to physical pixels and rendered half-size on HiDPI; size comes from the stylesheet, and the foreground colour is normalised by inspection rather than assuming which colour struct the shell hands back. Chip labels: non-Latin names all collapsed to "?", because the split treated every Cyrillic letter as a separator -- notable for a tool whose own README is Russian. Seniority also ranked by time-in-state rather than session age, so after a shell restart the older session could take the digit; the hook now records when the session began. zellij: a dump ends with new_tab_template and swap_tiled_layout blocks whose tab lines carry no name, and their panes were being attached to the last real tab -- which then answered for every unmatched directory, confidently and wrongly. install.py no longer widens the mode of a settings.json someone narrowed to 0600, no longer overwrites the pristine .bak on a second run, no longer replaces a symlink out of a dotfiles repository with a regular file, and quotes the hook path. The debug log is capped and README now says plainly that it records prompts verbatim. Not fixed, deliberately: the panel does push the clock about 70 px left with three labelled chips, which is inherent to putting them in the centre box; two different projects abbreviating alike still read as one project with a digit; GNOME 48 remains unverified for the colour struct and for St.BoxLayout's vertical property, both flagged rather than guessed at.
512 lines
20 KiB
Python
Executable File
512 lines
20 KiB
Python
Executable File
#!/usr/bin/env python3
|
|
"""Claude Code hook -> per-session state file for the GNOME panel indicator.
|
|
|
|
Reads one hook event as JSON on stdin and maps it to a session state, written
|
|
to $XDG_STATE_HOME/claude-code-status/<session_id>.json. The GNOME extension
|
|
watches that directory; no polling, no daemon, no socket.
|
|
|
|
Design notes that are easy to get wrong:
|
|
|
|
* Writes are skipped when the state does not change. PostToolUse fires on every
|
|
tool call, and its only job here is to clear "blocked" once a permission has
|
|
been granted -- letting it rewrite the file each time would make the directory
|
|
monitor fire hundreds of times per turn for no new information.
|
|
|
|
* Hooks are registered async, so two events can race (the last PostToolUse of a
|
|
turn against that turn's Stop). The whole read-decide-write runs under a file
|
|
lock and an event older than the stored one is refused; without both, a late
|
|
"busy" buries "waiting" and the panel claims a session is working while it
|
|
actually waits for input.
|
|
|
|
* No stdlib import beyond what is needed: this runs once per tool call.
|
|
"""
|
|
|
|
import errno
|
|
import fcntl
|
|
import json
|
|
import os
|
|
import sys
|
|
import time
|
|
|
|
STATE_DIR = os.path.join(
|
|
os.environ.get("XDG_STATE_HOME") or os.path.expanduser("~/.local/state"),
|
|
"claude-code-status",
|
|
)
|
|
|
|
# Notification covers several unrelated things; only some mean "the session
|
|
# stopped and is asking me something". auth_success / elicitation_complete /
|
|
# elicitation_response are progress chatter and must not touch the state.
|
|
# Observed: a question put to the user (AskUserQuestion) arrives as
|
|
# "permission_prompt" too, with the same generic message as a tool asking to
|
|
# run. The two are therefore not separable here, which is why the panel has one
|
|
# "blocked" state rather than telling a permission from a question.
|
|
NOTIFICATION_STATES = {
|
|
"permission_prompt": "blocked",
|
|
"agent_needs_input": "blocked",
|
|
"elicitation_dialog": "blocked",
|
|
"idle_prompt": "waiting",
|
|
"agent_completed": "waiting",
|
|
}
|
|
|
|
# Tools that spawn a subagent. Matched again here, not just in the hook
|
|
# registration: the matcher is a regex and a mistake there would silently
|
|
# inflate the count with TaskCreate, TaskUpdate and the like.
|
|
AGENT_TOOLS = {"Agent", "Task"}
|
|
|
|
# Events that mean the question has been dealt with. Anything else leaves a
|
|
# "blocked" session blocked: a subagent finishing, or the next tool starting,
|
|
# says nothing about the prompt still sitting on your screen, and clearing it
|
|
# would hide the one state this indicator exists to surface.
|
|
BLOCK_CLEARING = {"PostToolUse", "UserPromptSubmit", "Stop", "SessionStart"}
|
|
|
|
# Nothing here is displayed at full length, and both ends of the pipe have to
|
|
# survive a hostile or merely absurd value: a multi-megabyte cwd would be copied
|
|
# into the state file, read back by the compositor and handed to Pango.
|
|
MAX_TEXT = 512
|
|
# A stored timestamp further ahead than this is not a concurrent write, it is
|
|
# corruption -- and left alone it would refuse every later event forever,
|
|
# freezing the session's displayed state for good.
|
|
FUTURE_SLACK = 60 # seconds
|
|
|
|
|
|
def number(value, default=0):
|
|
"""Coerce a value read back from disk. Files are ordinary user-writable
|
|
JSON: one corrupt field must not take the hook down for every session."""
|
|
try:
|
|
n = float(value)
|
|
except (TypeError, ValueError):
|
|
return default
|
|
return default if n != n or n in (float("inf"), float("-inf")) else n
|
|
|
|
|
|
def clip(value):
|
|
text = value if isinstance(value, str) else ""
|
|
return text[:MAX_TEXT]
|
|
|
|
EVENT_STATES = {
|
|
# A session that has just opened is waiting for your first prompt, which is
|
|
# the same thing as one that has finished a turn: the input line is free
|
|
# and the next move is yours. It had a state of its own once; it was only
|
|
# ever reachable before the first prompt, so it bought a fourth glyph in
|
|
# the panel that nobody saw.
|
|
"SessionStart": "waiting",
|
|
"PreToolUse": "busy",
|
|
"UserPromptSubmit": "busy",
|
|
"PreCompact": "busy",
|
|
"PostToolUse": "busy",
|
|
"Stop": "waiting",
|
|
}
|
|
|
|
|
|
def debug_log(event):
|
|
"""Append raw events when a 'debug' marker file exists in the state dir.
|
|
|
|
Gated on a file rather than an env var because the hook inherits claude's
|
|
environment, which cannot be changed without restarting the session.
|
|
"""
|
|
marker = os.path.join(STATE_DIR, "debug")
|
|
try:
|
|
# Not followed through a symlink, and not grown without limit: this
|
|
# records prompts verbatim and the command lines of ancestor processes.
|
|
if os.lstat(marker).st_size > 8 << 20:
|
|
return
|
|
except OSError:
|
|
return
|
|
try:
|
|
fd = os.open(marker, os.O_WRONLY | os.O_APPEND | os.O_NOFOLLOW)
|
|
with os.fdopen(fd, "a") as fh:
|
|
chain, pid = [], os.getppid()
|
|
for _ in range(6):
|
|
if pid <= 1:
|
|
break
|
|
chain.append("%d %s" % (pid, read_cmdline(pid)[:90]))
|
|
pid = parent_of(pid)
|
|
stamped = dict(event, _at=time.strftime("%H:%M:%S"), _ancestry=chain)
|
|
fh.write(json.dumps(stamped, sort_keys=True)[:2000] + "\n")
|
|
except OSError:
|
|
pass
|
|
|
|
|
|
def derive_state(event):
|
|
"""Return the new state, 'end' to drop the session, or None to ignore."""
|
|
name = event.get("hook_event_name")
|
|
if name == "SessionEnd":
|
|
return "end"
|
|
# A notification is honoured whoever raised it: it means a human is needed,
|
|
# and that is just as true when the agent that got stuck is a subagent.
|
|
if name == "Notification":
|
|
return NOTIFICATION_STATES.get(event.get("notification_type"))
|
|
# SubagentStop is the counter's decrement and carries agent_id itself, so
|
|
# it has to pass the filter below. Its state is decided in apply_event,
|
|
# which is where the count is known.
|
|
if name == "SubagentStop":
|
|
return "busy"
|
|
# A subagent's own tool calls also reach the parent session's hooks
|
|
# (measured: PostToolUse carrying agent_id and agent_type). They are
|
|
# ignored, because the count already says a subagent is running and these
|
|
# would only add write traffic.
|
|
if event.get("agent_id"):
|
|
return None
|
|
if name == "PreToolUse" and event.get("tool_name") not in AGENT_TOOLS:
|
|
return None
|
|
return EVENT_STATES.get(name)
|
|
|
|
|
|
def read_environ(pid):
|
|
"""Environment of a process as a dict, empty if it is gone or not ours."""
|
|
try:
|
|
with open("/proc/%d/environ" % pid, "rb") as fh:
|
|
raw = fh.read()
|
|
except OSError:
|
|
return {}
|
|
env = {}
|
|
for entry in raw.split(b"\0"):
|
|
if not entry:
|
|
continue
|
|
key, sep, value = entry.partition(b"=")
|
|
if sep:
|
|
env[key.decode("utf-8", "replace")] = value.decode("utf-8", "replace")
|
|
return env
|
|
|
|
|
|
def read_cmdline(pid):
|
|
try:
|
|
with open("/proc/%d/cmdline" % pid, "rb") as fh:
|
|
return fh.read().replace(b"\0", b" ").decode("utf-8", "replace")
|
|
except OSError:
|
|
return ""
|
|
|
|
|
|
def parent_of(pid):
|
|
try:
|
|
with open("/proc/%d/status" % pid, "r") as fh:
|
|
for line in fh:
|
|
if line.startswith("PPid:"):
|
|
return int(line.split()[1])
|
|
except (OSError, ValueError):
|
|
pass
|
|
return 0
|
|
|
|
|
|
def looks_like_claude(cmdline):
|
|
"""Is this command line the claude binary itself?
|
|
|
|
Matched per argument, never against the raw string. The hook is spawned as
|
|
`/bin/sh -c /.../claude-status-hook.py`, so its parent's command line
|
|
contains the word "claude" -- in a path -- without being claude at all.
|
|
Latching onto that shell records a pid that exits milliseconds later, and
|
|
the session then flickers in and out of the panel.
|
|
"""
|
|
for token in cmdline.split(" "):
|
|
if not token:
|
|
continue
|
|
base = os.path.basename(token)
|
|
if base == "claude":
|
|
return True
|
|
# npm-style install: node /path/to/claude-code/cli.js
|
|
if base == "cli.js" and "claude" in token:
|
|
return True
|
|
return False
|
|
|
|
|
|
def find_claude_pid():
|
|
"""Nearest ancestor that is the claude process itself, or 0 if unknown.
|
|
|
|
The hook is spawned through a shell, so the immediate parent is not claude.
|
|
Walking beyond a handful of levels risks latching onto an outer claude when
|
|
one session drives another, so the search stops early.
|
|
|
|
Returning 0 rather than guessing matters: the reader deletes state files
|
|
whose process is gone, and the obvious fallback -- the shell that spawned
|
|
this hook -- exits immediately.
|
|
"""
|
|
pid = os.getppid()
|
|
for _ in range(6):
|
|
if pid <= 1:
|
|
break
|
|
if looks_like_claude(read_cmdline(pid)):
|
|
return pid
|
|
pid = parent_of(pid)
|
|
return 0
|
|
|
|
|
|
def pid_start_time(pid):
|
|
"""Field 22 of /proc/<pid>/stat: when the process started, in clock ticks.
|
|
|
|
Pins a pid to one particular process. Pids are reused, and state files
|
|
outlive reboots -- without this, a file left by a crashed session whose pid
|
|
is later handed to something unrelated reads as a live session forever.
|
|
"""
|
|
try:
|
|
with open("/proc/%d/stat" % pid) as fh:
|
|
data = fh.read()
|
|
except OSError:
|
|
return 0
|
|
# Field 2 is the command name, parenthesised, and may itself contain spaces
|
|
# and a ')'. Everything after the last ')' is field 3 onwards.
|
|
tail = data[data.rfind(")") + 2:].split()
|
|
try:
|
|
return int(tail[19])
|
|
except (IndexError, ValueError):
|
|
return 0
|
|
|
|
|
|
def alive(pid, start=0):
|
|
"""Is that pid still the process it was? Inputs come from disk, so both
|
|
arguments are coerced rather than trusted."""
|
|
pid = int(number(pid))
|
|
if pid <= 0 or not os.path.exists("/proc/%d" % pid):
|
|
return False
|
|
# A file written before start times were recorded has nothing to compare.
|
|
start = number(start)
|
|
return not start or pid_start_time(pid) == start
|
|
|
|
|
|
def sweep_dead(keep):
|
|
"""Drop state files whose claude process is gone.
|
|
|
|
A killed terminal never sends SessionEnd, so files leak. Cleaning up on
|
|
SessionStart keeps the sweep off the hot path -- the extension only has to
|
|
hide stale entries, not own their lifetime.
|
|
"""
|
|
try:
|
|
names = os.listdir(STATE_DIR)
|
|
except OSError:
|
|
return
|
|
for name in names:
|
|
if not name.endswith(".json") or name == keep:
|
|
continue
|
|
path = os.path.join(STATE_DIR, name)
|
|
try:
|
|
with open(path, "r") as fh:
|
|
stale = json.load(fh)
|
|
if not isinstance(stale, dict):
|
|
continue
|
|
pid = stale.get("pid", 0)
|
|
except (OSError, ValueError, AttributeError):
|
|
continue
|
|
except Exception:
|
|
# One unreadable file must not abort the sweep, and above all must
|
|
# not abort the caller: this runs before the hook writes its own
|
|
# state, so an exception here would stop new sessions appearing at
|
|
# all, for as long as the bad file sits there.
|
|
continue
|
|
# pid 0 means the hook could not identify the process; there is nothing
|
|
# to test for liveness, so leave it to the reader's age cutoff.
|
|
if pid and not alive(pid, stale.get("pid_start", 0)):
|
|
for victim in (path, path + ".lock"):
|
|
try:
|
|
os.unlink(victim)
|
|
except OSError:
|
|
pass
|
|
|
|
|
|
def write_atomic(path, payload):
|
|
tmp = "%s.%d.tmp" % (path, os.getpid())
|
|
with open(tmp, "w") as fh:
|
|
json.dump(payload, fh)
|
|
os.replace(tmp, path)
|
|
|
|
|
|
def open_lock(path):
|
|
"""Open the lock file without following a symlink.
|
|
|
|
A plain open() on a symlinked lock path truncates whatever it points at.
|
|
Nothing is escalated by that on a single-user machine, but a status
|
|
indicator has no business truncating files it was pointed at.
|
|
"""
|
|
flags = os.O_RDWR | os.O_CREAT | os.O_NOFOLLOW
|
|
try:
|
|
return os.fdopen(os.open(path, flags, 0o600), "r+")
|
|
except OSError as exc:
|
|
if exc.errno in (errno.ELOOP, errno.EMLINK):
|
|
return None
|
|
raise
|
|
|
|
|
|
def apply_event(event, state, path, now):
|
|
"""Read the current state, decide, and write. Must run under the lock."""
|
|
if event.get("hook_event_name") == "SessionStart":
|
|
# Swept before any early return: a resumed session keeps its id, so its
|
|
# SessionStart finds an unchanged state and would otherwise bail out
|
|
# before ever reaching the sweep.
|
|
sweep_dead(keep=os.path.basename(path))
|
|
|
|
if state == "end":
|
|
# Only the state file. Unlinking the lock while holding it drops mutual
|
|
# exclusion -- a hook already blocked on the old inode and one that
|
|
# creates a new file are then both inside the critical section.
|
|
try:
|
|
os.unlink(path)
|
|
except OSError:
|
|
pass
|
|
return
|
|
|
|
previous = None
|
|
try:
|
|
with open(path, "r") as fh:
|
|
previous = json.load(fh)
|
|
except (OSError, ValueError):
|
|
previous = None
|
|
if not isinstance(previous, dict):
|
|
previous = None
|
|
|
|
# Normalised here so the comparison below is against what would actually be
|
|
# stored: comparing a stored "" to a raw notification message rewrites the
|
|
# file on every idle_prompt for no change at all.
|
|
message = clip(event.get("message", "")) if state == "blocked" else ""
|
|
|
|
name = event.get("hook_event_name")
|
|
agents = int(number(previous.get("agents"))) if previous else 0
|
|
agents = max(0, min(agents, 999))
|
|
# Whether the main agent has finished its turn. Tracked separately from the
|
|
# state because with background subagents both are true at once: the turn is
|
|
# over and work is still running.
|
|
stopped = bool(previous.get("stopped")) if previous else False
|
|
|
|
# An event that lost a race carries an older timestamp than what is already
|
|
# stored. Its *state* must not be applied -- that is last-writer-wins, and
|
|
# replaying an old one would resurrect a state the session has left.
|
|
previous_ts = number(previous.get("event_ts")) if previous else 0
|
|
if previous_ts > now + FUTURE_SLACK:
|
|
previous_ts = 0 # corrupt, and trusting it would freeze this session
|
|
stale = previous_ts > now
|
|
|
|
# The subagent count is different in kind: the mutations are deltas, and
|
|
# deltas commute, so every one has to land whatever order it arrives in.
|
|
# Dropping a "+1" because its timestamp lost a race leaves the count one
|
|
# short, and the batch then frees the session while a subagent is still
|
|
# running -- the exact failure counting was added to prevent. The timestamps
|
|
# cannot be trusted for this at all: each is taken when its hook process
|
|
# starts, tens of milliseconds before it reaches the lock.
|
|
if name in ("SessionStart", "UserPromptSubmit"):
|
|
# A new turn from you starts a new batch. This also bounds the damage
|
|
# when a subagent dies without its SubagentStop ever arriving: the count
|
|
# cannot leak past the next thing you type.
|
|
agents = 0
|
|
elif name == "PreToolUse":
|
|
agents += 1
|
|
elif name == "SubagentStop":
|
|
agents = max(0, agents - 1)
|
|
|
|
if stale:
|
|
# Keep the stored state and flag; the delta above still gets persisted.
|
|
state = previous.get("state", state)
|
|
else:
|
|
if name in ("SessionStart", "UserPromptSubmit"):
|
|
stopped = False
|
|
elif name == "Stop":
|
|
stopped = True
|
|
elif name == "PostToolUse":
|
|
stopped = False
|
|
|
|
if name == "SubagentStop":
|
|
# The last subagent finishing is what finally frees a session whose
|
|
# main agent stopped long ago.
|
|
state = "waiting" if (stopped and agents == 0) else "busy"
|
|
elif state == "waiting" and agents > 0:
|
|
# The turn ended but the batch is still running, and the session
|
|
# will pick the results up itself. Calling it "waiting" would send
|
|
# you to a terminal that does not need you.
|
|
state = "busy"
|
|
|
|
# A pending question outlives everything except an answer to it.
|
|
if (previous and previous.get("state") == "blocked"
|
|
and state != "blocked" and name not in BLOCK_CLEARING):
|
|
state = "blocked"
|
|
message = previous.get("message", "")
|
|
|
|
# Resolved before the unchanged-check, not after, so that a pid which has
|
|
# changed forces a write. A session resumed under a new pid, or one whose
|
|
# pid was recorded wrongly, would otherwise keep the stale value for as
|
|
# long as its state happens not to change -- and the reader, finding that
|
|
# process gone, would drop a perfectly live session from the panel.
|
|
claude_pid = find_claude_pid()
|
|
|
|
if previous:
|
|
# Auto-compaction raises SessionStart again, in the middle of a turn the
|
|
# session is still working on. Taking it at face value would flip a busy
|
|
# session to waiting until the next tool call corrected it.
|
|
if event.get("hook_event_name") == "SessionStart" and event.get("source") == "compact":
|
|
return
|
|
# Nothing new to publish: stay quiet so the directory monitor stays quiet.
|
|
if (previous.get("state") == state
|
|
and previous.get("message", "") == message
|
|
and previous.get("agents", 0) == agents
|
|
and bool(previous.get("stopped")) == stopped
|
|
and previous.get("pid", 0) == claude_pid):
|
|
return
|
|
|
|
env = read_environ(claude_pid) if claude_pid else {}
|
|
|
|
write_atomic(path, {
|
|
"session_id": event.get("session_id"),
|
|
"state": state,
|
|
"cwd": clip(event.get("cwd") or ""),
|
|
# Age is measured from the moment the state was entered, not from the
|
|
# last event, so "waiting 40 min" survives unrelated later writes.
|
|
"since": previous["since"] if previous and previous.get("state") == state else now,
|
|
# When the session itself began, as opposed to when it entered this
|
|
# state. Seniority between chips is decided on this: a session that
|
|
# changed state a moment ago has not become the younger of the two.
|
|
"started": (previous.get("started") if previous else None) or now,
|
|
"event_ts": max(now, previous_ts),
|
|
# 0 means "could not tell"; the reader must not take that for "dead".
|
|
"pid": claude_pid,
|
|
"pid_start": pid_start_time(claude_pid) if claude_pid else 0,
|
|
"event": event.get("hook_event_name", ""),
|
|
"notification_type": event.get("notification_type", ""),
|
|
"message": clip(message),
|
|
"agents": agents,
|
|
"stopped": stopped,
|
|
# Kept from the previous write when this event could not identify the
|
|
# process: a momentary failure should not blank out where the session is.
|
|
"zellij_session": clip(env.get("ZELLIJ_SESSION_NAME")
|
|
or (previous.get("zellij_session", "") if previous else "")),
|
|
})
|
|
|
|
|
|
def main():
|
|
now = time.time()
|
|
try:
|
|
event = json.load(sys.stdin)
|
|
except (ValueError, OSError):
|
|
return 0
|
|
if not isinstance(event, dict):
|
|
return 0
|
|
|
|
session_id = event.get("session_id")
|
|
if not session_id or "/" in session_id:
|
|
return 0
|
|
|
|
debug_log(event)
|
|
|
|
state = derive_state(event)
|
|
if state is None:
|
|
return 0
|
|
|
|
os.makedirs(STATE_DIR, exist_ok=True)
|
|
path = os.path.join(STATE_DIR, "%s.json" % session_id)
|
|
|
|
# Hooks for one session run concurrently -- the last PostToolUse of a turn
|
|
# races that turn's Stop. Comparing timestamps is not enough on its own:
|
|
# without a lock both processes read the same "previous" and the loser's
|
|
# write still lands last, pinning a finished session at "busy". The lock is
|
|
# a separate file because write_atomic replaces the inode of the real one.
|
|
lock = open_lock(path + ".lock")
|
|
if lock is None:
|
|
return 0
|
|
with lock:
|
|
fcntl.flock(lock, fcntl.LOCK_EX)
|
|
apply_event(event, state, path, now)
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
try:
|
|
sys.exit(main())
|
|
except Exception:
|
|
# A hook that fails loudly would spam every session with error output;
|
|
# a missing panel update is the cheaper failure.
|
|
sys.exit(0)
|