A tmux session manager for running and restoring Codex, Claude, Gemini, and custom coding-agent CLIs, with logging and monitoring built in.
This project was formerly named Codex CLI Farm. Existing codex-* commands,
CODEX_* environment variables, and codexfarm state paths remain supported for
backward compatibility.
Features
- Automated session management: Long-lived tmux session that persists across reboots
- Durable pane history: Optional tmux-deep-history integration adds rotated raw/normalized transcripts, a seamless Page Up handoff beyond tmux's in-memory buffer, and compatible timestamped farm logs
- Unified monitoring: Watch all agent instances from a single consolidated view
- Fast navigation: Optional "board" session for quick switching between instances
- Exact snapshot/restore: Save each provider conversation by ID and restore every manifest row independently
- Status updates: Tracks RUN/READY/ERR in tmux metadata and notifies when a window becomes READY
- Memory warnings: Flag tmux windows whose pane process trees exceed a chosen RSS threshold
- Autosave/autorestore (optional): Systemd user services to persist sessions across logins
- Prompt loopers: Run one prompt or a prompt sequence repeatedly with retries, completion gates, git safety checks, logs, presets, and tmux visibility. See Agent Looper Reference for every parameter and default.
- Tool wrappers:
claude-*andgemini-*commands use the same tmux workflow
Requirements
- Bash 3.2+ and Python 3.10+.
tmuxfor farm sessions, boards, status inspection, save/restore, and farm-launched loopers.- The selected provider executable on
PATH(codex,claude,gemini, or a custom command). - Optional
multitailfor the richercodex-watchview. Without it, watch falls back totail. - Optional
lsofimproves exact-session discovery for Claude, Gemini, and Codex panes whose lifecycle hook has not run. - Git only for looper backup branches and git-progress circuit breakers.
- Systemd user services only for autosave/autorestore.
- Optional network access during
setup.sh --with-deep-historyto download the checksum-pinned plugin release.
setup.sh reports missing operating-system dependencies; it does not install packages. Source it when you want the current shell's PATH refreshed. The script keeps strict shell options inside helper scopes, so sourcing it should not leak options or functions into your shell.
Quick Start
1. One-time Setup
Run the setup script to create helper scripts, install the Codex session-identity hook and pinned deep-history integration, and auto-reload your shell:
source ./setup.sh --with-deep-historyThis will:
- Report missing
tmuxandmultitailcommands without running package-manager installs - Create helper scripts in
$HOME/bin/ - Merge the session-identity hook into
${CODEX_HOME:-~/.codex}/hooks.json,~/.claude/settings.json, and~/.gemini/settings.json - Set up logging directories
- Add
$HOME/binto your PATH automatically (bash/zsh/fish) and the current session - Verify and install the pinned
tmux-deep-historyrelease under$XDG_DATA_HOME/codexfarm/plugins/(or~/.local/share/codexfarm/plugins/)
If tmux is unavailable, core tmux commands will not work until you install it. If multitail is unavailable, codex-watch falls back to a simpler tail view.
Omit --with-deep-history when you want the legacy flat-log backend only. Re-running the setup
command is safe; the installed version can change only when this repository's reviewed lock changes.
After updating an existing checkout, include --with-deep-history again if you use that backend so
its compatibility launcher stays in sync with the copied farm commands.
When both input and output are terminals, setup explains gentle memory protection and asks whether to enable agent protection and queueing for new managed batch jobs. Enter, no, or EOF keeps every existing setting; declining does not uninstall previous protection. Enabling preserves custom thresholds, AI investigation, automatic actions, and model/binary choices. Setup never invokes a model.
On supported Linux/systemd hosts, enabling offers a separate host step: setup
shows the complete read-only plan, then asks a second default-no question before
running sudo. This adds ancestor memory preferences and a small 30-second root
OOM maintenance timer. Named farm unit drop-ins affect those unit names for other
users too; OOM changes are restricted to your UID. Other workloads may face earlier
reclaim, and interactive CPU/IO preference can slow competing batch work. The
shared 1 GiB MemoryLow protects used memory; it reserves no empty RAM and gives
no OOM immunity. Running agents are never stopped or capped by this policy.
Agent wrapping applies to future managed launches; setup does not move current
agents into different scopes or select existing sessions for host protection.
Default unattended, piped, or redirected setup leaves memory settings untouched
and performs no host preview or apply. --with-memory-protection explicitly
enables the two user options; unattended runs print exact manual host commands
without running them. --without-memory-protection skips both questions and
preserves all settings. The two flags conflict. Neither flag changes AI options.
Codex requires review for non-managed command hooks. Open /hooks in Codex CLI
after setup and trust the Agent CLI Farm hook. Provider hooks receive the active
session_id on SessionStart and UserPromptSubmit, then record it as an
invisible tmux pane option. Existing unrelated hooks and settings are preserved.
Use ./setup.sh --without-session-hook (or CODEXFARM_INSTALL_SESSION_HOOK=0)
when you do not want setup to change user hook files.
For local interactive Codex launches, the farm adds --no-daemon when the
selected CLI advertises it. Each TUI then owns its conversation and hook
environment, including when restoring an exact session. Older CLIs, remote
connections, utility commands, explicit --no-daemon arguments, and custom
shell invocations retain their supported command. Hooks originating from a
Codex app server are ignored because its inherited tmux pane may belong to
another TUI. Updating the helpers does not restart existing Codex processes.
To explicitly inspect the ID recorded for the current pane, run this inside that pane (the command intentionally prints the otherwise hidden conversation ID):
tmux show-options -p -v -t "$TMUX_PANE" @codexfarm_session_id tmux show-options -p -v -t "$TMUX_PANE" @codexfarm_session_source
Deep history is intentionally opt-in because terminal transcripts can contain commands, private
paths, credentials, and other sensitive output. Once installed, the default auto backend starts
recording current tmux panes and enables the plugin's global tmux hooks for panes created afterward.
Set CODEXFARM_HISTORY_BACKEND=legacy when launching a pane to turn those hooks off and use only
the flat compatibility log. If the plugin is absent, codex-add reports the legacy fallback
instead of silently making deep history look active.
2. Add Codex, Claude, or Gemini Instances
From any project directory:
Or specify a path:
codex-add /path/to/project
Use a named farm without exporting CODEX_SESSION:
codex-add work /path/to/project
codex-add work # current directory in the "work" farmClaude and Gemini use the same tmux workflow with wrappers:
claude-add /path/to/project gemini-add /path/to/project
Launch notes:
- A single non-path positional value is treated as a farm name. Use
--session NAMEwhen that would be ambiguous. - Provider flags after
--are passed as argv data, which is the preferred path for one-off native options:codex-add -d /repo -- --model gpt-5.4. CODEX_ARGS,CLAUDE_ARGS, andGEMINI_ARGSremain trusted shell fragments for compatibility. Set them only from config you control.
3. Run Prompt Loopers
Create starter files in the project where an agent should work, then follow the printed next steps:
For guided setup:
codex-looper init --interactive --force
After editing PROMPT.md, start a looper in the default farm:
codex-looper claude-looper claude-looper -- --dangerously-skip-permissions
codex-looper and claude-looper default to a hybrid interface: the real agent TTY stays visible in a tmux pane while the looper tracks session JSONL files for turn completion. Use --interface json when you specifically need the older noninteractive JSON stream mode.
For bounded smoke runs, legacy prompt sequences, completion-gated loops, or named farms:
codex-looper --once --label repo-smoke codex-looper --mode sequence --prompt-file prompts.md --once claude-looper --complete-on 'EXIT_SIGNAL:\s*true' --plan-file fix_plan.md --backup codex-looper --cb-no-progress 3 --cb-output-decline 2 --backup codex-looper --cb-output-match 'STATUS:\s*BLOCKED' --cb-output-match-repeats 3 codex-looper --preset rai claude-looper --farm-session work --label cleanup-pass --cwd /path/to/project codex-looper --local --once --label local-smoke
Looper labels are used for logs and agent session names only. Farm tmux window names stay tied to CODEX_NAME or the working directory basename.
Farm-launched loopers use a two-pane tmux layout by default. Claude hybrid runs use the main pane as the merged supervisor/status/control surface and the second pane for the live Claude Code TTY. Non-hybrid split runs use the second pane as a looper control pane with the live agent transcript. Use --tmux-layout single or CODEX_LOOPER_LAYOUT=single to keep one pane.
Inspect a running looper or agent without attaching:
codex-status activity codex-status loopers
Queue a safe stop without attaching:
codex-looper control stop LOOPER-rai --after-loop --reason "merge checkpoint"Record the agent's current high-level focus without attaching:
codex-looper control focus LOOPER-rai --summary "Verifying that live screenshots are still reaching the validation loop."Send or record an operator note without attaching:
codex-looper control note LOOPER-rai --delivery btw --note "controller drift was fixed after the latest calibration"
codex-looper control note LOOPER-rai --delivery record --note-file ./handoff-note.mdIn the split looper control pane, enter b NOTE to send NOTE through Claude's
/btw side channel to the active hybrid pane. Notes are also recorded in the
run directory as operator_notes.jsonl. When run state does not include a
hybrid pane id, pass an explicit --pane; broad tmux pane scanning requires
--allow-pane-scan after verifying the intended recipient.
Agents can refresh the visible supervisor focus line from inside a run with
codex-looper control focus --run-dir "$CODEX_LOOPER_RUN_DIR" --summary "...".
Keep it to one human-readable sentence about the larger stroke of work; the
append-only history is stored in focus.jsonl.
See Agent Looper Reference for prompt format, CLI parameters, config defaults, stop conditions, farm integration, and current backend limits.
Looper defaults and limits:
- The
runsubcommand is optional. In an initialized directory,codex-loopermeanscodex-looper run. - Prompt mode defaults to
singlewithPROMPT.md. A custom--prompt-filealso defaults tosingle; the legacy default fileprompts.mdimpliessequence. - Running loopers reload the prompt file before each loop by default, so prompt edits apply on the next pass. Set
reload_prompt_each_loop = falseonly when a run must keep the startup prompt text. - Config loading is strict: documented scalars must have the expected TOML type, numeric values must be finite, and invalid regexes fail before the loop starts.
- Backup branches point to committed
HEADonly; they are not dirty-worktree snapshots. Pruning stays inside the exact configured prefix namespace. - The no-progress circuit breaker fingerprints committed
HEAD, status entries, tracked metadata, and file contents while ignoring the looper run directory. - The output-match circuit breaker stops repeated project-defined status reports, for example a loop that keeps saying it is blocked.
- Structured provider errors with
cyber_policyormisalignment_policy_violationstop the entire looper with exit code 1 and durableneeds_reviewstate. Automatic retries, fresh sessions, completion markers, and--ignore-nonzerocannot continue that run. Review the stop before starting any new run; quotations in messages or tool output do not trigger this guard. - Run directories include a high-resolution timestamp and random suffix; the current-log pointer is updated atomically. In split mode, if tmux cannot create the control pane, live transcript streaming falls back to the supervisor pane.
- Every looper run writes durable machine-readable state to
state.json, append-only lifecycle history toevents.jsonl, visible focus history tofocus.jsonl, and accepts optional operator commands incontrol.jsonlin its run directory.codex-status loopersreads state files, flags active states whose supervisor process is gone or defunct as stale, andcodex-status loopers --repair-stale-loopersrecords those states as externally stopped so stop reasons survive pane exits and can be scraped without tmux. codex-looper control stop --nowinterrupts all known runtime targets, including hybrid tmux pane descendants and process groups. Usecodex-looper control stop LABEL --forcefor a stuck looper that needs SIGTERM/SIGKILL escalation and stale-state repair.
4. Watch All Instances
Monitor all Codex logs in real-time:
Notes:
- On small terminals (phones),
codex-watchauto-switches to a simpler mode. - Force simple mode:
codex-watch --simpleorCODEX_WATCH_MODE=tail codex-watch. - Force full mode:
codex-watch --mode multitail. - With the installed deep-history backend, one
pipe-paneowner writes durable segmented history and mirrors the live raw stream into the same flat logs used bycodex-watch. - Flat compatibility logs contain output emitted after the pane pipe is enabled. Deep history separately captures scrollback already visible when recording starts.
- Deep history is not a literal extension of tmux's internal grid. Use
Prefix + [for recent copy-mode history; when Page Up reaches the absolute top, the farm's default seamless handoff opens the older disk transcript in a read-onlylesspopup. Page Up/Page Down and the arrow keys scroll,/searches, and either Escape orqcloses the viewer and returns to the same copy-mode boundary so Page Up can reopen it. A loading notice is shown while the on-disk transcript is assembled, and startup failures remain visible until Enter or Escape is pressed.Prefix + Hopens the full Deep History menu, andAlt-uopens older history directly from copy mode. Mouse-wheel scrolling remains inside tmux's native buffer and does not cross the disk-history boundary. - Plugin installation raises tmux's global history limit from its usual 2,000 rows to 50,000 for panes created afterward. Existing panes keep their original in-memory limit, but
codex-addbackfills deep-history recording for existing tmux panes unless another output pipe already owns one. Recreate panes (or save and reboot the farm) only when you also want the larger in-memory limit. codex-watchdiscovers log files once at startup. Restart it to include newly-created logs.- Multi-file tail output keeps source labels so interleaved lines can still be traced to a window log.
- First run shows an optional tmux tips prompt; choose Yes to see basics. Answer "Don't show again" to persist your preference. Re-enable temporarily with
CODEX_TIPS_PROMPT=1or permanently by removing~/.local/state/codexfarm/no_tips.
5. Save/Restore or Resume
Snapshot your current Codex windows:
codex-save # writes to ~/.config/codexfarm/manifest.tsv CODEX_SESSION=work codex-save # writes to ~/.config/codexfarm/manifests/work.tsv codex-save --all-registered # conservatively autosaves each registered farm codex-save --autosave # preserve the restore point if conversations are missing
Restore them later (e.g., after reboot or on SSH login):
codex-restore -a # recreates and attaches to the session codex-restore --list-snapshots # list retained snapshots and window counts codex-restore /path/from/the/list.tsv # restore a chosen earlier snapshot
Reboot a running farm through the complete save, stop, and restore sequence:
codex-farm-reboot # reboot the default farm and attach codex-farm-reboot --detach # restore without attaching codex-farm-reboot work # reboot the named "work" farm claude-farm-reboot --detach gemini-farm-reboot --detach
The reboot command saves before stopping anything, stops a linked board before the main farm, restores the farm, and recreates the board only when it existed beforehand. Run it from a separate terminal or a different tmux session; it refuses to kill the session or board that is currently hosting the command. Checkpoint active agent work first because exact provider-session discovery is best-effort.
Use -f to force re-creation of existing-named windows.
Managed panes keep a stable logical name in tmux metadata even while a CLI
changes its visible native title. Duplicate logical names are valid. Restore
checks exact conversation identity across every pane, including split and renamed
windows. If a same-named window contains a different or unverified conversation,
restore reports a conflict; use a separate farm or --force to deliberately
replace it. An exact conversation already live in another tmux farm is skipped
during normal restore; --force refuses before removing any windows. Force restore removes all matching name occurrences before
recreating every row. Duplicate conversation IDs in a manifest are rejected.
Saved Codex, Claude, and Gemini windows require exact session IDs.
Codex first uses fresh hook metadata owned by the pane's current provider
process, then its held session writer lock; all providers can fall back to live
session-file discovery. Codex
rollout metadata is checked before an ID is saved: multi-agent child threads are
mapped to their resumable root session, while ordinary forks retain their own
independent thread IDs. Restore applies the same check, so manifests saved with
an older farm version are repaired as they are launched. If any recognized
provider ID is unresolved, codex-save exits nonzero and leaves the previous
manifest intact. Without authoritative hook or writer metadata, multiple
unrelated session files are treated as ambiguous.
A static manual/recovery binding is insufficient for a live Codex TUI: its
process can survive an in-TUI conversation switch while the recorded ID stays
unchanged. Save requires fresh direct hook metadata or current process ownership;
an unverified shared-server binding stops capture and is flagged by the doctor.
Latest-session resumes (codex resume --last, claude --continue, and
gemini --resume latest) and the old --allow-fallback option are no longer
supported. Restore rejects legacy manifests without exact provider IDs before
creating or removing windows, including with --force. Resave the running farm
or choose a retained exact snapshot to repair an old manifest.
Restored Codex sessions skip the startup update prompt using the per-launch
check_for_update_on_startup=false
override, so recovery proceeds directly to the saved conversation. This does
not change your Codex configuration or saved session IDs; update Codex separately
when convenient.
Every provider pane is saved, including conversations in home and secondary
panes. A home conversation restores as home-codex, home-claude, or
home-gemini; additional panes get a -pane-INDEX suffix unless they have a
distinct managed name. Each conversation restores into its own window. Plain
home shells and secondary non-provider panes are omitted; other first-pane
shell/custom commands retain their existing behavior. Repeated views of the same
conversation are saved once. Split layouts and scrollback are not reconstructed.
Missing saved directories fall back to $HOME with a warning.
For a dedicated history picker, launch local Codex with --no-daemon and mark
its pane explicitly, for example from another terminal:
tmux set-option -p -t '%PANE_ID' @codexfarm_utility history-pickerOnly a confirmed local --no-daemon resume --all picker with no conversation
identity or recorded session metadata is excluded. A marked shared-server or
remote pane with an unknown identity still blocks save. Once a selected
conversation has a verified ID, it is included in saves and coverage checks.
Unmarked Codex panes always require an exact identity, regardless of window name.
Manifests are written owner-only, flushed to disk, and atomically replaced.
Save and restore operations on the same manifest are serialized. Manifests
remain trusted executable input because restore launches the recorded commands.
Every distinct snapshot is retained beside its manifest in <manifest>.history/.
An unchanged save leaves the restore point and timestamp alone. A manual save
deliberately replaces the restore point and retains its predecessor. Autosave
only promotes a snapshot that retains every previously saved conversation and
generic window. If the farm shrinks, or replaces conversations even at the same
window count, autosave retains the small current manifest and preserves the prior
restore point. This protects a substantial earlier farm from an idle single
session, a partial restore, or windows being closed. Use codex-save manually
when that reduction is intentional. Empty farms never replace a snapshot;
missing farms are skipped during autosave. History is deduplicated by contents
and is not automatically pruned.
Codex permits only one writer for a conversation. Restore waits up to 10 seconds
for a writer left behind by tmux shutdown to exit, then leaves that window
unrestored and returns nonzero rather than starting a TUI that will immediately
fail. After an in-TUI /fork, Codex can keep the original conversation loaded
in that same process. To return to the original, use
/resume in the owning TUI, or exit that Codex process before resuming the
original elsewhere. The farm will not terminate another live Codex process.
codex-resume also checks the selected farm (or board) for dead managed Codex
panes whose exact resume failed with a SQLite initialization lock. It retries
each matching pane once with the currently installed Codex executable and the
same conversation, after checking writer ownership. Live panes, ordinary exits,
unrelated errors, and custom launch commands are left alone. Repeated lock
failures remain visible for diagnosis; recovery never deletes databases or
changes CODEX_HOME/sqlite_home. A recovered conversation may ask whether to
resume a paused goal; that choice remains yours.
Retries preserve individual-server mode when the original launch used
--no-daemon.
Repair farm installation and recovery health, then verify coverage without printing session IDs:
codex-doctor
codex-doctor --check # read-only diagnostics
CODEX_SESSION=work codex-doctor
codex-doctor --source /path/to/agent-cli-farm /path/to/manifest.tsvThe doctor repairs by default. It refreshes missing, stale and non-executable
helpers from the recorded checkout, repairs installed providers' identity hooks,
captures live conversations, fixes manifest permissions and freshness, restarts
a stopped memory monitor, starts an enabled but inactive autosave timer, and
retries a failed autosave service after successful capture. Replaced helpers and
hook settings are retained privately under the state directory's doctor-repairs/.
Its manifest repair merges current coverage with previously saved conversations,
retaining snapshot history. codex-save --merge selects that policy explicitly;
ordinary manual saves and conservative autosaves retain their existing behavior.
Repairs respect disabled services, annotator settings, operator masks and the
separate conversation-archive opt-in. They do not interrupt conversations or
install system packages. --check performs the original read-only diagnostics;
--fix and --repair explicitly select the default repair mode.
The final check exits nonzero for remaining stale or missing helpers, malformed or
unsafe manifests, blank commands, and non-UUID or fallback provider resumes.
Duplicate logical names are reported as information because restore supports
them by occurrence, including older conversations retained during repair.
It also flags manifests older than 24 hours, enabled but inactive timers, and
failed autosave services. Codex session discovery checks held writer locks as
well as open rollouts, so paginated sessions remain discoverable even when
their rollout descriptor is closed.
For a running farm, the doctor inspects every pane and compares verified
conversation IDs with the manifest. For an older TUI connected to a local shared
Codex server, save verifies the connection using reciprocal kernel socket peers
and the server process, then captures the server's complete durable live
conversation inventory through read-only metadata requests. This includes other
clients using that same local server. Saved logical names are reused where
possible; additional conversations restore as separate windows. The audit
requires the whole inventory to be covered and rechecks connections and coverage
before declaring success. It does not infer a pane's conversation from static
bindings, titles, directories or the newest transcript. On Linux this fallback
requires ss (iproute2) and a server supporting the metadata protocol.
Unknown remote identities, missing entries and unverified inventories still
fail the check without printing conversation IDs, preserving previous recovery data.
Memory monitoring reports managed shared-server process trees separately from
individual chat panes; host memory pressure still accounts for their RAM use.
Doctor refreshes its own verified monitor when installed monitor code changes.
Explicitly enabled history archives continue independently when exact manifest
capture fails.
Normal local shared-server recovery needs no relaunch. To move a TUI to local
writer ownership instead, wait until its conversation is idle,
record the exact ID currently shown by Codex's /status, and exit that TUI
normally. From another terminal, relaunch the same conversation through the farm:
codex-add --session codexfarm /path/to/project -- resume CURRENT_SESSION_ID codex-save codex-doctor
Use the actual project directory and currently displayed ID. Do this one conversation at a time. If its writer is still held, wait for the server to release it and retry; do not stop a shared server that serves other conversations. The farm does not interrupt running chats to migrate them.
Bulk restore waits two seconds between launches and checks host memory before
creating windows, before force deletion, and before each launch. Memory pressure
is advisory by default: warnings, critical readings, and unavailable or broken
health helpers do not block restoration. --enforce-memory-pressure explicitly
stops at critical pressure or a failed health check before mutation or further
launches; the saved manifest stays intact. A single-farm refusal exits 3;
--all-registered preserves its aggregate exit 1 when any child farm refuses. Unknown memory
counters remain advisory. --ignore-memory-pressure skips the health helper.
CODEXFARM_RESTORE_MEMORY_POLICY=warn|enforce|ignore selects the default policy;
CLI flags override it. An invalid final policy exits 2 before touching tmux.
--delay SECONDS (0..300) or CODEXFARM_RESTORE_DELAY_SECONDS changes pacing.
These options also apply with --all-registered and provider restore wrappers.
Checks never interrupt running chats.
If tmux sessions are already running (no manifest needed):
codex-resume # joins the main Codex session if present codex-resume work # joins the named "work" farm codex-resume work --board
Flags:
--boardto prefer the board session first.--session NAMEto force a specific farm when a positional argument would be ambiguous.
6. (Optional) Enable Autosave/Autorestore
codex-add can install systemd user services to save lightweight session
manifests every five minutes and restore on login.
You can trigger it directly:
codex-add --install-autoservice
The installed systemd unit names are always the same: codex-autosave.service, codex-autosave.timer, and codex-autorestore.service. Installing autoservice for a named farm adds that farm to ~/.config/codexfarm/farms.tsv; it does not create a second background service:
codex-add --session work --install-autoservice codex-add --session personal --install-autoservice
Autosave/autorestore iterates the registry, so each registered farm is saved to its own manifest and restored into its own tmux session. Autosave uses the conservative policy above and runs after autorestore when both services start together. Re-run codex-add --install-autoservice to refresh existing service definitions. Older units that already invoke codex-save --all-registered also receive the conservative policy automatically.
Autosave defaults to codex-save --all-registered every five minutes at low CPU
and idle I/O priority, with a three-minute helper timeout. It saves small
manifests of exact session identities. To explicitly schedule conversation
archives, use codex-add --install-autoservice --with-conversation-backups;
that selects codex-backup --archive --min-age 3600 and preserves the archive
budgets described below. Use --without-conversation-backups with an explicit
install to return to manifests. Both backup flags require --install-autoservice
and cannot be combined.
The service choice is stored in autoservice_choice and archive consent in
conversation_backup_choice under ${XDG_STATE_HOME:-~/.local/state}/codexfarm.
A legacy autoservice_choice=yes alone does not authorize archives. Explicit
service refreshes preserve the separate archive choice unless a backup flag
changes it. Normal window launches with a stored service choice of yes only
register their farm; they do not rewrite units, reload the manager, enable, or
start services. A first interactive or CODEX_AUTOSERVICE_CHOICE=yes selection
installs once. A stored no respects the operator's disabled choice.
Explicit installs and refreshes check all three units for local, runtime, and
global masks before writing any units or choices, including filesystem masks
when the user service manager is unavailable. A masked unit causes a refusal
that names it and preserves all units, masks, registry, and choices; the helper
never unmasks services. Installation also reports failure if the user service
manager or timer cannot be activated. Activation failure retains the requested
unit definitions and choices; retry with an explicit --install-autoservice
refresh, since ordinary launches do not retry activation. Units preserve the installation PATH so
Node/NVM commands remain available outside an interactive shell. Older units
that call codex-backup --min-age 3600 become manifest-only after updating the
helper; refresh the units explicitly to select the current save command.
Restart any already-running backup watcher after updating. Save failures remain
a nonzero service result.
Set CODEX_AUTOSERVICE_CHOICE=yes to auto-accept the prompt, or no to suppress it.
Conversation backups and recovery
codex-backup # save small manifests; full archives stay off codex-backup --watch # foreground manifest autosave fallback codex-backup --archive # explicitly create a full chat archive codex-backup --archive --skip-save # archive history even with no tmux server codex-backup --archive --max-mib 1024 # explicitly raise the archive size budget
Full conversation archives are disabled by default. Autosave creates them only with separately persisted archive consent; the default fallback watcher saves manifests. Existing archives are preserved when archiving is disabled; updating the helpers does not delete old backups.
With --archive, backups are private local snapshot-*.tar.gz archives under
~/.local/state/codexfarm/backups (or --destination). They contain Codex
rollouts, archived rollouts, history/index files, Claude project chat JSONL,
Gemini chat JSON/JSONL, individually consistent SQLite copies of Codex
session/history/goal/queue databases, and farm TSV manifests. SQLite
copies include committed WAL data. The snapshot inventory and latest.json
record save coverage and an archive SHA-256. Authentication files,
configuration, provider logs, and unrelated home files are excluded. Chats can
still contain sensitive information: keep these archives private and out of Git.
Archives are published atomically, with owner-only files/directories. One is
retained by default (CODEXFARM_BACKUP_KEEP or --keep); pruning occurs only
after a successful replacement. Incident-recovery archives with other names are
left alone. The default size budget is 512 MiB for both total source history
(including SQLite WAL files) and compressed output. --max-mib changes that
budget. Oversized sources are rejected before copying, and growth during a copy
is checked too. Optional archives use low CPU priority and fast compression.
Databases are staged one at a time before transcripts to reduce temporary disk
use. Backups reserve 1 GiB of free disk space, with additional space checked for
the new archive and database staging. Low-space or size-limit failures preserve
the previous archive and remove partial files. Simultaneous writes by other
programs can still exhaust the disk. Backups are on the same host unless you
choose another destination. They do not protect against loss of that disk.
backup-status.json records explicit archive attempts. A failed manifest save
does not prevent a history snapshot, but the combined command still exits
nonzero. --min-age limits snapshot frequency while still checking the manifest
on each invocation. The watcher uses an exclusive lock and saves every five
minutes. The default manifest watcher exits once the systemd autosave timer
becomes active. It is a temporary foreground scheduler, not a boot service.
Scheduled full archives require explicit autoservice archive consent or a
codex-backup --archive --watch --min-age 3600 invocation; that archive watcher
continues alongside the manifest autosave timer.
Archive health checks are also off by default. They apply only to an explicitly
enabled archive watcher, separately persisted archive and autoservice choices
of yes, or CODEXFARM_BACKUP_HEALTH_ENABLED=1 set before starting the annotator.
An explicit no in either persisted choice suppresses archive warnings from stale
watcher/status files, preserving those files. The environment override of 1
continues to require archive checks even with disabled choices. In that mode,
BACKUP WARNING indicates a failed save or archive, a scheduler heartbeat older than 15 minutes, a missing archive, or
a snapshot older than two hours. A manual archive alone does not require future
archives. codex-doctor and codex-health still check memory pressure and the
memory-monitor heartbeat without requiring full conversation backups.
After a crash, preserve the existing archive first. Inspect snapshot.json and
extract to a separate private directory; do not replace a live Codex database.
For history still in the current Codex home, codex resume --all shows chats
from every working directory, and codex -C /path/to/repo resume UUID opens an
exact conversation. A saved tmux manifest is a window index, not the chat
transcript itself. Restoring windows does not submit a continuation prompt.
Status Updates (RUN/READY/ERR)
codex-add auto-starts codex-annotator, which polls every five seconds and tracks RUN, READY, or ERR state in tmux window options. CODEX_ANNOTATOR_INTERVAL or --interval overrides polling with a positive finite number of seconds. For Codex/Claude panes it inspects recent output for prompts/approval selections; other panes fall back to the command-based heuristic.
By default, the annotator does not rewrite tmux window titles. That lets Codex's native title animation remain visible while still making codex-status windows show state. When a window transitions from RUN to READY, the annotator emits a tmux display-message notification.
Important: the RUN/READY/ERR status is best-effort and based on terminal-output and command heuristics. Node-backed tools are recognized from pane start commands as well as current commands. Use READY as a signal, not a guarantee. Codex's native title animation is usually the primary visual signal.
Continuous memory warnings
The annotator checks Linux host memory and managed window process trees every
15 seconds, including descendant test/build workers. It adds a warning to the
managed session's existing status-right, preserving its previous contents and
native window titles, and shows a ten-second tmux message. Warning readings
must persist for two samples; critical host pressure alerts immediately.
Unchanged alerts repeat at most every five minutes.
Defaults and environment overrides (set before starting the annotator):
| Setting | Default | Meaning |
|---|---|---|
CODEXFARM_MEMORY_POLICY |
headroom |
Use fixed available MiB; percent selects percentage thresholds. Explicit legacy percentage variables select percent when policy is absent |
CODEXFARM_MEMORY_WARN_MIB |
1536 |
In headroom mode, warn at or below this available RAM |
CODEXFARM_MEMORY_CRITICAL_MIB |
1024 |
In headroom mode, critical at or below this available RAM; restore stops only with explicit enforcement |
CODEXFARM_MEMORY_WARN_PERCENT |
20 |
In percent mode, warn at or below this percentage of RAM available |
CODEXFARM_MEMORY_CRITICAL_PERCENT |
10 |
In percent mode, critical at or below this percentage of RAM available |
CODEXFARM_MEMORY_SESSION_MIB |
1024 |
Warn when a window's process tree exceeds this RSS |
CODEXFARM_HEALTH_ENABLED |
1 |
Set to 0 to disable periodic memory and optional archive checks |
CODEXFARM_BACKUP_HEALTH_ENABLED |
0 |
Set to 1 to require periodic full archives; archive watchers and explicit service archive consent also opt in |
CODEXFARM_HEALTH_STATUS |
1 |
Set to 0 to leave status-right formatting alone |
The default 1536/1024 MiB thresholds stay fixed on 4, 8, and 64 GiB hosts.
Set CODEXFARM_MEMORY_POLICY=headroom to use MiB thresholds even when legacy
percentage variables are present. All thresholds must be positive and finite,
with critical below warning; percentages must also be below 100. Legacy
percentage settings remain validated in headroom mode.
Linux memory pressure stalls (some avg10) also trigger warning at 10% and
critical at 25%. I/O stalls are reported separately and do not cause a RAM alert
or restore refusal. Previously used swap alone does not trigger an alert. RSS is
an attribution estimate that can count shared pages more than once; host
pressure uses MemAvailable, not summed process RSS. The monitor reports swap
usage for context. It never kills, pauses, or restarts chats. A 15-second poll
cannot prevent a sudden memory spike or guarantee avoidance of an OOM kill.
Run codex-health for a current readout and codex-doctor for installation and
recovery checks. The private health-status.json contains the last reading,
window IDs and RSS totals, without command arguments or conversation text.
@codexfarm_health is the session badge and @codexfarm_memory_mib is the
per-window estimate for custom tmux status formats. The doctor flags a missing
or older-than-60-second heartbeat once managed sessions have been registered.
If the annotator itself stops, run the doctor: a dead process cannot issue its
own live warning. Host pressure checks currently require Linux counters.
Optional incident investigation
Resource features are independent and default off. Configure only the desired options; omitted settings retain their prior values. This writes private JSON and never installs or starts services:
codex-resource status codex-resource report --json codex-resource configure --protect-agents --queue-background --investigator codex # Separate consent for the narrow batch remedies described below: codex-resource configure --automatic-actions # Disable each option independently: codex-resource configure --no-protect-agents --no-queue-background \ --investigator off --no-automatic-actions
The existing health monitor hot-loads these settings after writing its heartbeat. When investigation is enabled, 60 seconds of sustained RAM warning/critical pressure or I/O stalls at least 10% can launch a detached investigator. Used swap alone never triggers it. One private lock covers scheduled and manual workers; there is a 15-minute cooldown, and investigation defers below 512 MiB available RAM or with unknown memory counters. The monitor never waits for the model. Worker processes use nice 10 and OOM adjustment +250. Existing agents are not moved, paused or restarted.
codex-resource investigate produces diagnosis only. investigate --actions
also requires stored --automatic-actions consent; scheduled investigations
apply remedies only with that same separate consent. Turning the investigator
off cancels remedies after an in-flight diagnosis. Actions are limited to live,
registered batch jobs reported to the model: defer future matching launches for
up to one hour, reduce a declared allowlisted worker count for a bounded TTL,
or request the one restart explicitly authorized with codex-job --restartable.
The default TTL is five minutes. Reductions and deferrals affect future launches;
changing a running job requires its separately consented restart. Restart is a
nonreversible request, not a claim that a restart completed. Identity, ownership,
actual cgroup and scope generation are rechecked before each action. Agents and
unregistered processes receive diagnosis/manual suggestions only.
The adapter uses the existing Codex CLI login and the exact bundled model
gpt-6.1-sol. Optional --investigator-model MODEL and
--investigator-binary /absolute/path/to/codex overrides are validated and must
support the same isolation. It creates private HOME, CODEX_HOME and working
directories, links only the existing owned owner-only authentication cache, and
permits native credential refresh. It ignores user config/rules, uses ephemeral
read-only execution, disables tooling features, normalizes model-derived tools,
and supplies fixed instructions plus a strict output schema. Unsupported CLI,
model catalog or managed policy fails diagnostically. No unrestricted fallback
or provider transport override is used. The whole provider invocation, including
capability probes, has a 120-second budget and bounded output. The worker also
has a 120-second overall deadline. All post-timeout rollback, metrics, journal
recovery and report persistence share one absolute five-second cleanup grace.
It cannot be renewed by nested cleanup or retries, and exhaustion stops further
I/O while preserving known diagnosis/results in memory. Failed or interrupted writes
retain honest pending/uncertain action results and preserve validated diagnosis.
An override whose ID could not be returned stays uncertain until its bounded TTL
expires; unrelated overrides are never selected for cleanup.
Reports scan at most 4096 PIDs for two seconds, retain at most 20 consumers and
read PSS only for top candidates. Growth comparisons include start ticks, UID
and cgroup; PID reuse never becomes growth. Host capacity uses MemAvailable;
RSS/PSS, memory and I/O stalls, swap, and bounded cgroup stats/events provide
context. Labels are untrusted, sanitized and length-limited. Arguments, working
directories, arbitrary environment values, conversation text and credentials
are excluded. A compact projection reserves native framing/schema space within
a 16 KiB initial model-input budget.
Private reports, trusted local job identities, answers and action journals live
under ${XDG_STATE_HOME:-~/.local/state}/codexfarm/resources/, with owner-only
0700 directories and atomic 0600 JSON files capped at 64 KiB. Reports and journals
retain 20 files each. Journals record before/after readings, results and TTL
identifiers. Material deterioration removes only this investigation's reversible
overrides; expired or already removed overrides are benign. No model-generated
shell command, project edit, service change or root operation is executed.
Optional managed jobs
The optional host layer is a separate, standalone administrative helper. Setup
copies codex-resource-host into ~/bin. Interactive memory setup can show its
complete read-only plan with maintenance, then apply only after a separate yes to
the sudo question. Declining or a failed preview leaves system protection unapplied;
an accepted apply failure returns an error while retaining the saved user options.
Setup never removes conflicting configuration or unmasks units to recover.
You can also review the plan and explicitly apply the same options as root:
/usr/bin/python3 -I "$HOME/bin/codex-resource-host" plan --uid "$(id -u)" sudo /usr/bin/python3 -I "$HOME/bin/codex-resource-host" apply --uid "$(id -u)" # Optional: include the dedicated 30-second OOM preference maintenance timer. # Use this pair instead of the pair above if maintenance is wanted. /usr/bin/python3 -I "$HOME/bin/codex-resource-host" plan --uid "$(id -u)" --with-maintenance sudo /usr/bin/python3 -I "$HOME/bin/codex-resource-host" apply --uid "$(id -u)" --with-maintenance # Undo only this installation's expected files, runtime values and PID scores. sudo /usr/bin/python3 -I /usr/local/libexec/codexfarm-resource-protection.py remove
Apply requires an available system manager and the selected UID's user manager;
connection failures abort. The helper derives /run/user/UID/bus for an own-UID
plan, including shells without user-bus environment variables. Root connects to
the selected user manager with absolute systemctl --user --machine=UID@.host.
It uses fixed paths and sanitized subprocess environments, with no CLI or
environment override for root/config paths. Maintenance runs an isolated absolute
/usr/bin/python3 -I and imports only the standard library.
The shared 1 GiB MemoryLow protects memory already in use through user.slice,
user-UID.slice, user@UID.service, codexfarm.slice and the interactive child.
It does not reserve an empty free gigabyte or guarantee OOM immunity. Existing
stronger values, including infinity, remain stronger; unknown values are left
alone. Interactive and batch sibling slices receive relative CPU/IO weights
200 and 25. This helper adds no resource ceilings.
The plan lists every dedicated drop-in and touched administrative path. Farm
slice drop-ins under /etc/systemd/user/ have global user-unit scope: they
also apply to other users who launch those exact named farm slices. OOM updates
are restricted to the configured UID. The root-owned installed helper is
/usr/local/libexec/codexfarm-resource-protection.py (0755), with private config
/etc/codexfarm-resource-protection.json (0600). The private journal and lock live
under /var/lib/codexfarm-resource-protection/ (0700); one bounded original-file
snapshot is retained under /var/backups/codexfarm-resource-protection-originals/
(0700). A subsequent installation replaces that single snapshot.
For an existing agent, append --agent PID:STARTTICKS to both plan and apply.
For example, obtain start ticks from field 22 of /proc/PID/stat (parse after
its final ) because process names can contain spaces). The helper validates
UID, ticks and actual cgroup, applies a one-shot OOM preference of -250, and gives
an existing matching session-*.scope runtime MemoryLow within the same shared
budget. It never moves or restarts the process: moving a PID would not migrate
its already charged memory. Without maintenance, future processes get only the
ordinary user-level preferences available to the launcher.
Optional maintenance adjusts only exact codexfarm-agent-<32hex>.scope paths
under the interactive slice and codexfarm-batch-<32hex>.scope paths under the
batch slice, including their descendants. It uses -250 for agents and +250 for
batch, preserving stronger role preferences, and inspects at most 4096 PIDs for
two seconds per run. The private restoration journal is capped at 4096 process
records and 16 MiB; each original managed file is capped at 1 MiB.
Apply is idempotent for the same configuration. Remove before changing the UID, maintenance selection, or one-shot targets. Every target mask is checked before mutation, including user local/runtime/global masks. Old host guards, build slice masks, autosave masks and unrelated drop-ins remain untouched. Failed writes/manager operations trigger best-effort rollback. Remove restores only expected hashes/values and matching PID identities whose OOM score is still the one applied here; later operator changes survive. Dead sessions are not started. If rollback reports an error, correct the reported manager/filesystem issue and retry remove; the private journal and trusted recovery helper are retained when runtime rollback cannot finish. No farm/provider/session restart is performed.
codex-job run --role batch -- command argument ... runs literal argv without
shell expansion or default resource ceilings. User scopes are optional by default:
if the user manager or scope launcher is unavailable, it warns before launching
without a scope. --scope required fails before the payload instead. Use
--scope off to avoid user-manager calls entirely.
# Advisory admission, lower relative batch priorities, no caps: codex-job run --role batch -- make all # Wait for headroom before new work; declare workers for a future temporary reduction: codex-job run --role batch --memory-policy queue --queue-timeout 300 \ --worker-env CARGO_BUILD_JOBS --workers 8 -- cargo build # Explicit batch-only memory enforcement; never falls back without these limits: codex-job run --role batch --memory-high 2048 --memory-max 3072 -- make all # Optional interactive protection, including tmux launches and Looper agents: CODEXFARM_RESOURCE_PROTECTION=1 codex-add -d /path/to/project
Batch admission is advisory unless --memory-policy queue is explicit or private
settings enable queue_background, including a yes to the setup memory question.
With default settings, queueing starts when available RAM is at most
1024 MiB or memory stalls reach 25%. After waiting, admission needs at least
1536 MiB and stalls below 10% continuously for 30 seconds. Unknown counters and
hosts too small for the thresholds stay advisory. The default queue timeout is
five minutes and exits 124
without starting the payload; --memory-policy ignore manually bypasses both
headroom and temporary recipe deferrals. These checks never stop running work.
Agent jobs bypass admission and retain their terminal and inherited process group.
They cannot opt into restarts or memory ceilings. Agent scopes request relative
CPU/I/O preference and memory protection through codexfarm-interactive.slice
and its shared codexfarm.slice parent; both ancestors request at least 1024 MiB
of memory protection while preserving stronger or unknown settings. Interactive
slice CPU/I/O weights are 200 and batch slice weights are 25, so the relative
preference applies across the sibling slices as well as their scopes. Batch scopes use
codexfarm-batch.slice, CPU/I/O weights of 25, nice 10, and OOM adjustment +250.
An agent's negative OOM adjustment needs a separately opted-in privileged helper;
the unprivileged runner reports that limitation. Runtime slice preferences install
no unit files and preserve stronger existing parent memory protection. No scope
sets CPU, swap, task, or time ceilings by default.
Resource settings live in $XDG_CONFIG_HOME/codexfarm/resource-settings.json
(default ~/.config/codexfarm). protect_agents, queue_background, and
automatic_actions default false; investigator defaults off. Settings reject
unknown fields, invalid types, nonfinite numbers, symlinks, foreign ownership, and
public settings files. Reading absent settings creates nothing. Explicit settings
writes use atomic 0600 files and may tighten the owned target directory to 0700;
ordinary optional agent lookup warns and launches normally if settings or private
job registration are inaccessible. Use codex-resource configure for independent
opt-ins and codex-resource status to inspect the current settings. The Python
resource_config.load_settings / write_settings(ResourceSettings(...)) APIs
provide the same validated configuration for integrations.
Job recipes and identity records are private under
$XDG_STATE_HOME/codexfarm/resources/jobs, using 0700 directories and atomic 0600
files. For compatibility with older setup installations, managed launches migrate
only the owned, real codexfarm state directory to 0700 through a verified
no-follow file descriptor. HOME and XDG roots retain their permissions; unsafe
resource leaf directories or files are still rejected. Public JobStore.public /
list_jobs reports exclude argv, cwd, and arbitrary environment values. Remedies
select recorded UID, PID/start ticks,
cgroup, and scope generation rather than process names. Temporary worker and
deferral overrides apply only to matching argv, cwd, and declared worker variables;
they expire and can be rolled back. The worker allowlist is CARGO_BUILD_JOBS,
CMAKE_BUILD_PARALLEL_LEVEL, OMP_NUM_THREADS, and UV_THREADPOOL_SIZE.
--restartable grants a batch job permission for at most one externally requested
restart. Memory thresholds never request a restart on their own. A request must
revalidate live ownership and consent, sends TERM only to the dedicated batch
process group, waits up to ten seconds for graceful cleanup, verifies no old batch descendants
remain, and requires a fresh recovery period
before relaunch. A recovery timeout leaves the record queued_timeout and exits
124. Worker reductions affect an existing job only through this separately
consented restart. The JobStore.identity, request_restart, reduce_workers,
defer_job, delete_override (also rollback_override), and recipe_overrides
APIs enforce the incident controller's identity and consent checks. Temporary overrides last at most one
hour, records and overrides have bounded retention, and no payload output is copied
to this store. Cleanup retains records whenever any recorded supervisor, payload,
or launcher remains alive or its process identity is unreadable. A temporary
manager query failure disables remedies until validation recovers without making
the running record permanently stale.
Memory labels
Run codex-memoryflag to prefix high-memory tmux windows with measured usage, such as *349.1MB**. The compact MB label uses MiB (1024 KiB), rounded to one decimal place. It scans tmux sockets available to the current user, sums each window's pane process trees by RSS, and renames windows at or above the threshold. The threshold controls when the label appears; the number is the measured RSS, not the threshold. Shared pages can count more than once.
Each run updates or clears existing labels, including old *200+MB** markers, and records the chosen threshold in the per-window @codexfarm_memory_threshold_mib option. The running annotator refreshes opted-in windows it manages on its socket every 15 seconds, removing the marker below the threshold and restoring it if usage rises again. Without the annotator and its health monitor, labels remain snapshots until the next command run. Native titles remain the default for windows that have not opted in; memory labeling intentionally manages window names. A dry run changes neither titles nor options.
codex-memoryflag # flag windows at 200 MiB and up codex-memoryflag 500 # flag windows at 500 MiB and up codex-memoryflag 1G # flag windows at 1024 MiB and up codex-memoryflag -n # dry run
Rerun the command with another threshold to change it. To stop automatic memory-title refresh for a window, unset its option with tmux set-option -wu -t TARGET @codexfarm_memory_threshold_mib; the last title remains until renamed. The status annotator preserves memory markers if legacy title updates are enabled, so a title can read *349.1MB** *RUN* project-name. The separate large-chat health warning still defaults to 1024 MiB.
Tuning and controls:
- Disable autostart:
CODEX_ANNOTATOR_AUTOSTART=0 - Disable annotator (if started):
CODEX_ANNOTATOR_ENABLED=0 - Re-enable legacy title prefixes:
CODEX_ANNOTATOR_UPDATE_TITLES=1 - Disable READY notifications:
CODEX_ANNOTATOR_NOTIFY_READY=0 - Customize READY notification text:
CODEX_ANNOTATOR_READY_MESSAGE(default:READY: {name}) - Adjust RUN detection:
CODEX_ANNOTATOR_RUNNING_REGEX(default:(codex|node|ssh)) - Scope sessions:
CODEX_ANNOTATOR_SESSION_REGEX(default:^codex) - Ignore windows/sessions prefixed with
!(configurable viaCODEX_ANNOTATOR_IGNORE_PREFIX) - Adjust capture depth:
CODEX_ANNOTATOR_CAPTURE_LINES(default:200)
Available Commands
Core Commands
codex-add [session] [directory]- Add a new Codex instance, optionally selecting a named farmcodex-annotator- Track RUN/READY/ERR state and notify when windows become READYcodex-memoryflag [threshold]- Flag high-memory tmux windows; default threshold is 200 MiBcodex-job run [options] -- argv ...- Run declared jobs with optional scopes, batch admission and explicit restart consentcodex-resource configure|status|report|investigate- Configure independent opt-ins and inspect bounded incident diagnosissetup.sh --with-memory-protection|--without-memory-protection- Enable user memory options or skip the choice; interactive host apply requires separate consentcodex-resource-host plan|apply|maintain|remove- Review and explicitly install standalone host memory/OOM preferencescodex-health- Check RAM pressure, the monitor heartbeat and opted-in archive healthcodex-backup [--archive]- Save exact farm identities; full conversation archives require--archivecodex-watch- Monitor all Codex logs in consolidated viewcodex-looper [init|doctor|run]- Run a single prompt or prompt sequence repeatedly with logs and stop detection; see looper referencecodex-status [--session SESSION] [sessions|windows|activity|logs|loopers]- Show status information;loopers --repair-stale-loopersmarks active state files stopped when their supervisor process is gonecodex-board [create|link|switch] [session]- Manage the default or a named board session for navigationcodex-resume [session] [--board]- Attach/switch to an existing Codex/tmux session or named farm boardcodex-doctor [--check | --fix] [--session NAME] [--source DIR] [manifest]- Repair farm health and verify recovery coverage;--checkselects read-only diagnosticscodex-save [--autosave | --merge] [manifest]- Save exact conversations with retained snapshot history;--mergealso retains previously saved conversationscodex-restore [-a] [-f] [--list-snapshots] [manifest]- Restore or list retained snapshotscodex-farm-reboot [--detach] [session]- Safely save, stop, restore, and optionally attach to a farm
Claude and Gemini Wrappers
Claude and Gemini equivalents use the same tmux workflow and accept the same flags:
claude-add, claude-annotator, claude-board, claude-farm-reboot, claude-looper, claude-restore, claude-resume, claude-save, claude-status, claude-watch, gemini-add, gemini-annotator, gemini-board, gemini-farm-reboot, gemini-looper, gemini-restore, gemini-resume, gemini-save, gemini-status, gemini-watch.
Environment Variables
Common:
CODEXFARM_RESOURCE_PROTECTION- Set to1to opt into optional agent scopes for add scripts and Looper (default off)- Setup's memory flags persist
protect_agentsandqueue_background; no environment variable grants consent to its privileged host step. Existing independent AI options and numeric thresholds are preserved. CODEX_SESSION- tmux session name (default:codexfarm)CODEX_NAME- window name (default: directory basename)CODEX_CMD- command to run (default:codex)CODEX_ARGS- additional arguments for codexCODEX_STATE_BASENAME- state/log directory base name (default:codexfarm)CODEX_TIPS_PROMPT- show tmux tips prompt:0to disable,1to force (default respects a persisted opt-out)CODEX_LOCK_TITLES- set to1to keep Codex windows named after their directory (default0lets Codex's native title updates show)CODEX_REMAIN_ON_EXIT- keep tmux windows visible after the pane command exits (default1; set to0to close windows on exit)CODEX_WATCH_MODE-auto(default),tail, ormultitailto control codex-watch displayCODEX_STATUS_ACTIVITY_LINES- recent pane lines shown bycodex-status activity(default8)CODEX_LOOPER_STATE_ROOT- run-state directory forcodex-status loopers(default:.agent-looper/runsin the current directory)CODEX_AUTOSERVICE_CHOICE-yesornoto persist autoservice choiceCODEX_ANNOTATOR_AUTOSTART- set to0to skip starting the annotatorCODEX_LOOPER_PYTHON_BIN- Python 3.10+ interpreter forcodex-looper(default searchespython3, then versionedpython3.14throughpython3.10)CODEX_ANNOTATOR_PYTHON_BIN- Python 3.10+ interpreter forcodex-annotator(default searchespython3, then versionedpython3.14throughpython3.10)CODEXFARM_PYTHON_BIN- Python 3.10+ interpreter to prefer during./setup.shCODEXFARM_WITH_DEEP_HISTORY- set to1as an alternative tosetup.sh --with-deep-historyCODEXFARM_INSTALL_SESSION_HOOK- set to0to leave the Codex user hook file unchanged during setupCODEXFARM_RESTORE_WRITER_WAIT_SECONDS- secondscodex-restorewaits for an exact Codex thread's active writer to exit (default10)CODEXFARM_HISTORY_BACKEND-auto(default),legacy, ordeep-history; auto uses the pinned plugin when installed and otherwise preserves legacy loggingCODEXFARM_DEEP_HISTORY_BIN- override the deep-history executable discovered under the XDG data directoryCODEXFARM_DEEP_HISTORY_PYTHON_BIN- override the Python 3.10+ interpreter used by deep-history hooks and loggers (default searches supportedpython3and versioned commands)CODEXFARM_DEEP_HISTORY_SEAMLESS_PAGEUP- set to0to keep tmux's ordinary Page Up behavior at the top of copy mode instead of opening the older disk transcript (default1)CODEX_LOOPER_LAYOUT- looper tmux layout:auto,single, orsplit; farm launches default tosplit
Annotator-specific:
CODEX_ANNOTATOR_ENABLED- set to0to disable the annotator loopCODEX_ANNOTATOR_UPDATE_TITLES- set to1to restore legacy*RUN*/*READY*/*ERR*title prefixesCODEX_ANNOTATOR_NOTIFY_READY- set to0to disable tmux messages when a window transitions from RUN to READYCODEX_ANNOTATOR_READY_MESSAGE- tmux message template for READY notifications; supports{name}and{state}CODEX_ANNOTATOR_RUNNING_REGEX- regex for pane commands considered RUNNINGCODEX_ANNOTATOR_SESSION_REGEX- regex for sessions to annotateCODEX_ANNOTATOR_SESSION_REGISTRY- file of additional tmux session names to annotate (default:${XDG_STATE_HOME:-$HOME/.local/state}/codexfarm/managed_sessions)CODEX_ANNOTATOR_INTERVAL- polling interval in secondsCODEX_ANNOTATOR_IGNORE_PREFIX- window/session name prefix to ignore (default:!)CODEX_ANNOTATOR_CAPTURE_LINES- number of lines to capture from panes (default:200)
Tool-specific:
claude-addandgemini-addhonorCLAUDE_*orGEMINI_*versions of the common launch variables. For save/restore/resume/status/watch/board commands, select the farm withCODEX_SESSIONor the command's positional/--sessionargument where supported.codex-looperandclaude-looperpass native agent flags after--, for examplecodex-looper --once -- --ask-for-approval neverorclaude-looper --once -- --dangerously-skip-permissions. Built-in Codex and Claude agents can also setmodelandeffortin[agents.*]config. Codex and Claude useinterface = "hybrid"by default; set--interface json,[agents.codex].interface = "json", or[agents.claude].interface = "json"for the older JSON stream path. The Claude flag is spelled--dangerously-skip-permissionsand should only be used in isolated workspaces where unattended edits are acceptable.
Example:
CODEX_CMD="cursor" CODEX_ARGS="--wait" codex-add /my/project
Flags:
codex-add -d: start without attaching (useful in SSH automation)codex-restore -a: attach after restoring;-fto replace same-named windowscodex-save --autosave: retain the previous restore point if conversations are missingcodex-restore --list-snapshots: find an earlier snapshot to restore by pathcodex-farm-reboot -d: save, stop, and restore without attaching
Advanced Usage
Board Session for Fast Navigation
Create a separate board session for quick navigation. The default farm uses the legacy board session name; named farms use <farm>-board:
# Create board session for the default farm codex-board create # Create and use a named board codex-board create work codex-board link work codex-board switch work # Link all Codex windows to the default board codex-board link # Switch to the default board session codex-board switch
Now you can use tmux switch-client -t board to scan through all Codex instances while the main codexfarm session continues running.
For named farms, codex-resume work --board jumps directly to work-board.
Boards link existing tmux windows; they do not duplicate provider processes.
Remote SSH Tips
- Start or restore your farm, then safely detach:
codex-restore; tmux detach. - Reattach anytime:
codex-resume(ortmux attach -t ${CODEX_SESSION:-codexfarm}). - Prefer
codex-add -din automation to avoid stealing your current terminal. - For mobile networks and roaming devices, use mosh: install
moshon the server (and open UDP 60000-61000), then connect with a mosh-capable client and attach your tmux session. Desktop SSH keeps working the same.- Example client wrapper (tries mosh then ssh):
examples/connect.sh user@host
- Example client wrapper (tries mosh then ssh):
- If your terminal is very small (phones), use
codex-watch --simpleand zoom panes in tmux withPrefix + z. - Optional tmux tweak for mixed desktop/mobile:
tmux set -g aggressive-resize onto let windows resize to the current client.
Examples
-
Start many projects at once (non-attaching):
examples/batch-add.sh ~/proj/a ~/proj/b ~/proj/c tmux attach -t ${CODEX_SESSION:-codexfarm}
-
Auto-restore on login (add to shell rc):
# ~/.bashrc or ~/.zshrc source $(pwd)/examples/restore-on-login.sh
-
One-liners:
- Safely reboot and attach:
codex-farm-reboot - Start with a different command:
CODEX_CMD="cursor" CODEX_ARGS="--wait" codex-add -d /path - Start in a named farm without env vars:
codex-add work /path - Watch logs with multitail if available:
codex-watch
- Safely reboot and attach:
Log Management
All logs are stored in ${XDG_STATE_HOME:-$HOME/.local/state}/${CODEX_STATE_BASENAME:-codexfarm}/logs/ with timestamps:
# View log status codex-status logs # Follow specific log tail -f ~/.local/state/codexfarm/logs/myproject_20240315-143022.log # Clean old logs (example: older than 7 days) find ~/.local/state/codexfarm/logs -name "*.log" -mtime +7 -delete
Validation
Run the basic validation script (requires tmux):
./validate.sh ./tests/integration/session_resume_smoke.sh CODEXFARM_DEEP_HISTORY_BIN=/path/to/tmux-deep-history/bin/tmux-deep-history \ ./tests/integration/deep_history_smoke.sh
For the static checks used in CI without creating live tmux sessions:
VALIDATE_SKIP_TMUX=1 ./validate.sh
File Structure
agent-cli-farm/
├── .editorconfig # Shared editor whitespace defaults
├── .github/workflows/ # CI for lint, shell checks, tests, validate, demo
├── pyproject.toml # Ruff and test tooling configuration
├── requirements-dev.txt # Python development dependencies
├── CONTRIBUTING.md # Local development and verification workflow
├── setup.sh # Main setup script
├── bin/ # Helper scripts
│ ├── codex-add # Add new Codex instances
│ ├── codex-annotator # Bash wrapper for annotator
│ ├── codex-annotator.py # Track tmux window status and READY notifications
│ ├── codex-doctor # Diagnose installed-helper and manifest drift
│ ├── codex-save # Save manifest of windows
│ ├── codex-restore # Restore windows from manifest
│ ├── codex-session-meta.py # Resolve resumable roots and writer ownership
│ ├── codex-session-hook.py # Record active Codex IDs on managed tmux panes
│ ├── codex-session-hook-install.py # Safely merge the user hook config
│ ├── codex-farm-reboot # Save, stop, and restore a farm
│ ├── codex-watch # Monitor logs
│ ├── codex-board # Navigation helper
│ ├── codex-resume # Resume into existing session(s)
│ ├── codex-status # Status information
│ └── claude-* / gemini-* # Tool wrappers for the same commands
├── docs/
│ ├── looper.md # Detailed looper contract
│ └── QUALITY_REVIEW.md # Hardening review notes from the update guide
├── examples/
│ ├── demo.sh # End-to-end demo of farm
│ └── mock-codex # Fake CLI used by the demo
├── integrations/ # Pinned deep-history lock and atomic installer
├── tests/ # Python unit tests for scripts and looper behavior
├── validate.sh # Basic repo validation script
└── README.md # This file
Limitations
tmuxcannot mirror the same live pane in two windows (use linked windows or logs)tmuxsurvives client disconnects, but not host reboot or tmux-server exit unless autosave/autorestore recreates windows later.- Automatic login restoration uses systemd user services.
codex-backup --watchprovides foreground manifest autosave when that manager is unavailable; it must be restarted after a reboot. Same-disk snapshots need a separate off-host copy to protect against disk loss. - tmux provides one
pipe-paneconsumer per pane; the deep-history backend owns it and mirrors the stream rather than attaching a competing logger. - Existing panes created before the session hook was installed may need one
submitted Codex prompt before hook metadata appears. Save also inspects live
process file descriptors: Codex under
~/.codex/sessions, Claude under~/.claude/projects, and Gemini under a.gemini/.../chatsdirectory. Missing or ambiguous provider IDs stop save and preserve the previous snapshot. - Existing Codex TUIs attached to a shared app server cannot reliably map its
hooks or writer locks to an individual pane. Relaunch through
codex-addwhen the conversation is no longer in use to enable automatic identity tracking. A manually verified pane binding cannot track a later in-TUI/new,/fork, or/resumeswitch in that shared-server TUI. - A Codex conversation cannot be resumed by two processes at once. In-TUI
/forkcan retain the original thread's writer ownership until that process unloads it; use the same TUI's/resumepicker or close the owning process. - A single positional argument to
codex-addis interpreted as a farm name when it does not look like a path. Use--session NAMEto force farm selection when needed. - Gemini JSON-mode looper continuation requires the CLI's
initevent to report a session ID; if it does not, the looper stops before sending another prompt.
License
Licensed under the MIT License. See LICENSE for full text.
Unless noted otherwise, all files in this repository are covered by the MIT License.