docs(backup): recipe + worked example for remote git backup

The docs already tell users that pushing the wiki to a remote git repository
is a supported backup pattern: docs/deploy.md#backups says 'markdown - back
up with rsync or git push to a remote', docs/design-decisions.md says 'No
remote/cloud sync (use git remote on the wiki dir)', and both
docs/airgapped-install.md and SECURITY.md defer the channel security to the
operator ('Remote sync security'). But the recipe itself is nowhere. A
homelab or laptop install has no built-in remote-push cadence, and building
one from scratch takes 100+ lines of bash plus systemd units plus a
gitignore that catches derived state and secrets.

Fill the gap with a docs-only addition:

- New docs/backup.md walks through what to include (wiki/, raw/, config.toml
  minus secrets) vs exclude (db/ derived from wiki via reindex, models/
  redownloadable, logs/, .serve.lock), the rsync + git push + tarball flow,
  the systemd --user timer schedule, restore, and the SECURITY.md-aligned
  posture (private repo, least-privilege push credential, secret exclusion,
  encryption at rest is out of scope for ai-memory).
- A worked example under docs/examples/backup/ following the precedent set
  by docs/examples/jev-reranker-adapter/ and docs/examples/auto-improve-eval/:
  the snapshot script (configurable through six env vars), a systemd --user
  .service oneshot, a daily .timer, a .gitignore for the mirror repo, and
  a README with the install-and-enable steps plus non-systemd equivalents
  for macOS launchd, Windows Task Scheduler via WSL, and Docker sidecar.
- One-line pointers from docs/deploy.md#backups (right after the tarball
  routine) and docs/airgapped-install.md (extending the existing 'git
  remote sync' bullet).
- A row in the README docs table between lifecycle-ops.md and
  llm-providers.md, positioned as a companion to lifecycle-ops.md.

The on-box "ai-memory backup --to TARBALL" command is unchanged. This is
docs-only; no core CLI subcommand for remote push is proposed here (the
issue leaves that decision to the maintainer).

Refs #950.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
Renato Junior
2026-09-28 02:19:43 -03:00
co-authored by Claude
parent 49147a5d17
commit 7ac4b949a4
10 changed files with 448 additions and 1 deletions
+11
View File
@@ -7,6 +7,17 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]
### Added
- `docs/backup.md` documents the remote-git-mirror backup pattern for a
single-user install: what to include, what to exclude (derived SQLite index,
models cache, logs, secrets), how to schedule with a `systemd --user` timer,
how to restore, and the security posture per `SECURITY.md`. A worked example
ships under `docs/examples/backup/` (snapshot script, `.service` and
`.timer` unit files, `.gitignore` for the mirror repo). Pointers added
from `docs/deploy.md#backups`, `docs/airgapped-install.md`, and the
README docs table. The on-box `ai-memory backup --to <tarball>` command
(`docs/lifecycle-ops.md#backup`) is unchanged. (#950)
### Fixed
- Grok Build CLI tool observations are no longer stored with an empty body.
Grok posts Claude Code's snake_case tool fields (`tool_name` / `tool_input` /
+1
View File
@@ -448,6 +448,7 @@ diagram, crate breakdown, schema notes, and invariants.
| [`docs/users.md`](docs/users.md) | **Multi-user attribution and human login.** Four-rung bearer ladder, password sessions, `ai-memory user` / `api-key` walkthrough, brownfield migration. |
| [`docs/https-via-proxy.md`](docs/https-via-proxy.md) | **HTTPS via a reverse proxy.** When you need TLS and when you don't, with copy-paste Caddy / nginx / Cloudflare Tunnel templates and the "secure when you're not" failure modes. |
| [`docs/lifecycle-ops.md`](docs/lifecycle-ops.md) | **Read before purge / rename / backup / restore / reset / reindex / restore-page.** Safety matrix, per-project disk layout, checkpoint page recovery, and operator workflows. |
| [`docs/backup.md`](docs/backup.md) | Backing up the wiki + data dir to a remote git repository: what to include, what to exclude, scheduled push pattern, restore, and security posture. Companion to `docs/lifecycle-ops.md` (which covers the on-box `ai-memory backup` snapshot). |
| [`docs/llm-providers.md`](docs/llm-providers.md) | Provider configuration for consolidation and embeddings. |
| [`docs/security.md`](docs/security.md) | The full security model. |
| [`docs/support-matrix.md`](docs/support-matrix.md) | The full agent/platform matrix with notes. |
+3 -1
View File
@@ -59,7 +59,9 @@ runtime or other native library to source separately.
- **git remote sync**, if you use it to push the wiki repository somewhere,
is your own channel to secure (`SECURITY.md`, "Remote sync security" —
out of scope for ai-memory itself).
out of scope for ai-memory itself). [`docs/backup.md`](backup.md) has a
worked example of the rsync + `git push` pattern with the appropriate
secret and derived-state exclusions.
- **Update/patch delivery** in an air-gapped environment is manual: pull a
new release and checksum on a connected machine, then transfer it in,
the same as the initial install.
+163
View File
@@ -0,0 +1,163 @@
# Backups and remote git mirror
> This document covers backup patterns beyond the on-box `ai-memory backup
> --to <tarball>` snapshot documented in
> [`docs/lifecycle-ops.md#backup`](lifecycle-ops.md#backup) and the "point-in-time
> consistency" `scp` idiom in [`docs/deploy.md#backups`](deploy.md#backups). If
> you already have a working tarball routine, treat this as an additional
> layer.
Ai-memory's wiki directory is a real git repository. Every consolidation pass
writes a checkpoint through libgit2 (see cross-cutting invariant #10 in
[`AGENTS.md`](../AGENTS.md) and the wiki design notes in
[`docs/ARCHITECTURE.md`](ARCHITECTURE.md)). That means half of a remote-backup
setup is already done for you — the wiki has its own commit history and can be
pushed to any git remote you own. The rest of this doc is about wiring that
up, what else to include, what to exclude, and how to schedule it.
## What the data dir contains, and what to back up
`ai-memory status` prints the resolved data dir; per
[`docs/deploy.md#backups`](deploy.md#backups) it looks like this:
```
data/
├── wiki/ # markdown source of truth (already a git repository)
├── raw/ # immutable session-log ledger
├── db/ # memory.sqlite (FTS5, entities, embeddings) — DERIVED
├── logs/ # daily rolling tracing
└── models/ # local embedder weights (redownloadable)
```
Split by durability:
| Directory | Back up? | Why |
|---|---|---|
| `wiki/` | **Yes, always.** | Markdown source of truth. Irreplaceable. Already a git repository. |
| `raw/` | Yes if you want the full ledger; no if you want a small mirror. | Immutable observation log. Useful for rebuilding session history; not required to keep the wiki working. |
| `db/` | **No.** | SQLite index derived from `wiki/`. Rebuildable via `ai-memory reindex`; see [`lifecycle-ops.md#reindex`](lifecycle-ops.md#reindex). Committing a WAL-mode SQLite database to git also wastes space on every write. |
| `logs/` | No. | Tracing output; nothing durable. |
| `models/` | No. | Local embedder weights, sha256-pinned, redownloaded on first boot per [`local-embeddings.md`](local-embeddings.md). |
| `.serve.lock` | No. | Server pidfile. |
| `config.toml` and env files under `~/.config/ai-memory/` | Yes, minus secrets. | Provider config is useful for reproducing an install; API-key env files (`*.env`), OAuth tokens (`auth.json`), and private CA bundles must never land in a mirror. |
If you also need the SQLite index for near-instant recovery (rather than
running `reindex` after a restore), pair the git mirror with a periodic
`ai-memory backup --to <tarball>` and stash the tarball on a second location
(external disk, network share, or on WSL2 the Windows-side filesystem via
`/mnt/c/`).
## Backing up the wiki to a remote git repository
The wiki is already a git repository. If your only goal is off-site *wiki*
storage, adding a remote and pushing is enough:
```bash
data_dir=$(ai-memory status --format=json | jq -r .data_dir)
cd "$data_dir/wiki"
git remote add origin git@github.com:<you>/ai-memory-wiki.git
git push -u origin main
```
Two caveats:
- The wiki repo is written to by the server through libgit2; do not run
interactive git operations (rebases, merges, force-pushes) against it while
the server is running. Read-only commands (`git log`, `git show`, `git
push`) are safe.
- The wiki carries page bodies verbatim, including any content the model
synthesized. Treat the remote the same way you treat any project source:
private repository, restrictive access controls.
## Backing up the whole data dir (mirror + push)
For a full-install backup — wiki, raw ledger, sanitized config — the
supported pattern is a scheduled rsync into a mirror git repository, followed
by a git push. This avoids running git operations against the live wiki
repository (which the server writes to) while still giving you a single
remote with the full recoverable state.
A worked example lives at
[`docs/examples/backup/`](examples/backup/README.md) and consists of:
- `ai-memory-snapshot.sh` — an idempotent rsync + `git commit && git push` +
tarball script driven by six environment variables.
- `ai-memory-backup.service` — a `systemd --user` oneshot unit that runs it.
- `ai-memory-backup.timer` — a daily user timer with a randomised delay and
`Persistent=true` so a missed run after a reboot still fires.
- `.gitignore` — an exclusion list to drop into the root of the mirror
repository so derived state and secrets never land in commits.
The high-level flow, one invocation per day (adjust the timer for a different
cadence):
1. `rsync -a --delete` the data dir into the mirror checkout, applying an
exclude list for `models/`, `db/`, `logs/`, `.serve.lock`, and any glob
listed in `SECRET_EXCLUDES` (`*.env`, `auth.json`, `.secrets` by default).
2. `rsync -a --delete` the config dir under the same exclude rules so
`config.toml` is captured but env files, OAuth tokens, and private CA
bundles are not.
3. `git add -A && git commit -m "auto: snapshot $(date +%F)"` if anything
changed; skip if there is no delta.
4. `git push origin main` with a timeout so a hung network does not block the
next run.
5. `tar czf $TARBALL_DIR/ai-memory-$(date +%F).tar.gz` a copy of the same
tree, so the DB and models can be recovered without a `reindex` after a
catastrophic disk loss.
6. Rotate old tarballs and old script logs on a retention window (30 days by
default).
See the example directory's [README](examples/backup/README.md) for the
install-and-enable steps and non-systemd equivalents (macOS launchd, Windows
Task Scheduler via WSL, Docker sidecar).
## Restoring from a mirror repository
The mirror repository is a copy of the data dir minus derived state.
Recovering the full ai-memory install from it is two steps:
1. Clone the mirror onto the target host, then rsync the `data/` and
`config/` subtrees into place under the ai-memory data dir and config dir
respectively.
2. Rebuild the SQLite index from the recovered markdown:
[`lifecycle-ops.md#reindex`](lifecycle-ops.md#reindex) walks through this.
Sessions, observations, handoffs, users/tokens, audit rows, and
embeddings live only in the DB and are not rebuilt by `reindex` — for
those, restore from the paired tarball via `ai-memory restore --from
<tarball> --data-dir <path> --force`.
## Security posture
Remote sync is your own channel, not ai-memory's — this is called out
explicitly in [`SECURITY.md`](../SECURITY.md) ("Remote sync security" under
out-of-scope for v1). What that means in practice:
- **Private repository, always.** ai-memory captures prompts, tool calls, and
synthesized page bodies verbatim — the same content you would treat as
project source.
- **Least-privilege push credential.** Prefer an SSH deploy key or a
fine-grained personal-access token scoped to the mirror repository only.
Rotate on the same cadence you rotate other long-lived push credentials.
- **Explicit secret exclusion.** The example script rejects any file matching
`SECRET_EXCLUDES` (`*.env`, `auth.json`, `.secrets` by default), and the
suggested mirror-repo `.gitignore` matches the same set. Add anything else
install-specific — private CA bundles, host-side certificates, cloud
credential files — to both lists.
- **Encryption at rest is out of scope for ai-memory.** If your threat model
requires it, either put the mirror on an encrypted remote (`git-crypt`,
GitLab's group-level encryption, a self-hosted Gitea over TLS with disk
encryption) or archive the tarballs with `gpg --encrypt` before dropping
them at `TARBALL_DIR`.
## Related documents
- [`docs/lifecycle-ops.md`](lifecycle-ops.md) — `ai-memory backup`, `restore`,
`restore-page`, `reset`, `reindex` reference.
- [`docs/deploy.md#backups`](deploy.md#backups) — the on-box tarball routine
and the `scp`-to-laptop idiom for Docker deployments.
- [`docs/airgapped-install.md`](airgapped-install.md) — offline installs,
including the callout that remote git sync is user-owned plumbing.
- [`SECURITY.md`](../SECURITY.md) — the "Remote sync security" out-of-scope
note this doc expands into a walkthrough.
- [`docs/local-embeddings.md`](local-embeddings.md) — how `models/` is
fetched and pinned (why it does not need to be in the mirror).
+5
View File
@@ -236,6 +236,11 @@ scp "$SERVER:$DEPLOY_DIR/data/snapshot-$(date +%F).tar.gz" ./backups/
The `ai-memory backup` command uses SQLite's online backup API so
writes during the snapshot are coherent.
For a remote-first backup pattern (rsync the data dir into a git mirror on a
schedule, push to a private repository, paired with a tarball for the DB
and models), see [`docs/backup.md`](backup.md) and the worked example under
[`docs/examples/backup/`](examples/backup/README.md).
## Sharing one server between people or harnesses
A deployed server is the supported way to share a project — between teammates,
+35
View File
@@ -0,0 +1,35 @@
# Suggested .gitignore for a mirror repository that receives ai-memory
# snapshots. Copy this file into the root of your `$REPO_DIR` before the
# first `git add`.
#
# Rules of thumb:
# - Keep the wiki (markdown source of truth) and any per-project _rules /
# _prompts pages, since those are irreplaceable.
# - Keep raw/ if you want the untouched observation ledger, drop it if you
# want a small repository.
# - Never commit derived state (db/) or downloaded weights (models/), and
# never commit secrets.
# Secrets — belt and braces alongside the rsync SECRET_EXCLUDES list.
config/*.env
config/**/*.env
config/auth.json
config/**/auth.json
config/.secrets
config/**/.secrets
# Server pidfile.
data/.serve.lock
# Derived SQLite index (rebuildable via `ai-memory reindex`).
data/db/
# Local embedder weights (redownloadable on next boot; sha256-pinned).
data/models/
# Rotated tracing output.
data/logs/
# Any local scratch produced by the snapshot script itself.
*.tmp
*.swp
+80
View File
@@ -0,0 +1,80 @@
# Remote git backup example
Sample scripts + unit files that push an ai-memory install to a remote git
repository on a daily cadence, keeping a tarball on a second location for the
SQLite index and models cache that the mirror repo deliberately excludes.
See [`docs/backup.md`](../../backup.md) for the walkthrough. The pieces in this
directory are:
- `ai-memory-snapshot.sh` — the rsync + git push + tarball loop. Configurable
through six environment variables (`REPO_DIR`, `DATA_SRC`, `CFG_SRC`,
`TARBALL_DIR`, `LOG_DIR`, `SECRET_EXCLUDES`) with sensible defaults for a
single-user Linux/WSL/macOS install.
- `ai-memory-backup.service` — a `systemd --user` oneshot that runs the
script. `%h` expands to the user's home directory.
- `ai-memory-backup.timer` — a daily user timer with a small randomised delay
and `Persistent=true` so the missed run after a reboot still fires.
- `.gitignore` — the exclusion list to drop into the root of the mirror
repository. Belt-and-braces to the rsync/tar exclude lists in the script
itself.
## Quickstart
```bash
# 1. Create a private repository on your git host (e.g. github.com), then:
mkdir -p ~/ai-memory-backup
cd ~/ai-memory-backup
git init
git remote add origin git@github.com:<you>/ai-memory-backup.git
cp .../docs/examples/backup/.gitignore .
git add .gitignore
git commit -m "initial: ignore secrets, DB, models, logs"
git branch -M main
# push once so `snapshot.sh` has a remote to push to
git push -u origin main
# 2. Drop the snapshot script somewhere on $PATH.
install -Dm755 \
.../docs/examples/backup/ai-memory-snapshot.sh \
~/.local/bin/ai-memory-snapshot.sh
# 3. Install the unit files.
install -Dm644 \
.../docs/examples/backup/ai-memory-backup.service \
~/.config/systemd/user/ai-memory-backup.service
install -Dm644 \
.../docs/examples/backup/ai-memory-backup.timer \
~/.config/systemd/user/ai-memory-backup.timer
# 4. Enable + start.
systemctl --user daemon-reload
systemctl --user enable --now ai-memory-backup.timer
# 5. Trigger one snapshot to confirm the plumbing works.
systemctl --user start ai-memory-backup.service
journalctl --user -u ai-memory-backup.service --since '5 minutes ago'
```
## Non-systemd equivalents
The script is a plain bash program; anything that runs a shell command on a
schedule can drive it:
- **macOS launchd:** wrap `ai-memory-snapshot.sh` in a `.plist` under
`~/Library/LaunchAgents/` with a `StartCalendarInterval` block.
- **Windows Task Scheduler:** run the script inside WSL via `wsl -e bash -lc
'ai-memory-snapshot.sh'` on a daily trigger. On native Windows use the
PowerShell equivalent of the rsync + `git add/commit/push` + tar sequence.
- **Docker sidecar:** a small cron container mounted at the ai-memory data
volume, running the same script.
## Security note
The remote is your own channel, not ai-memory's. Follow the existing
[`SECURITY.md`](../../../SECURITY.md) guidance ("Remote sync security"): use a
private repository, a fine-grained token scoped to that one repo (or an SSH
deploy key), and rotate credentials the same way you rotate any other
long-lived push credential. The script never commits files matched by
`SECRET_EXCLUDES` (default: `*.env`, `auth.json`, `.secrets`); the mirror
repo's `.gitignore` catches anything the rsync exclude missed.
@@ -0,0 +1,22 @@
[Unit]
Description=ai-memory snapshot + push to remote git
Documentation=file://%h/.local/share/ai-memory/docs/backup.md
After=network-online.target
Wants=network-online.target
[Service]
Type=oneshot
# Point ExecStart at wherever you dropped ai-memory-snapshot.sh — this path
# is only a sample. Adjust the environment overrides if the defaults in the
# script do not match your install.
ExecStart=%h/.local/bin/ai-memory-snapshot.sh
# Environment=REPO_DIR=%h/ai-memory-backup
# Environment=DATA_SRC=%h/.local/share/ai-memory
# Environment=CFG_SRC=%h/.config/ai-memory
# Environment=TARBALL_DIR=%h/backups/ai-memory
Nice=10
IOSchedulingClass=best-effort
IOSchedulingPriority=7
[Install]
WantedBy=default.target
@@ -0,0 +1,11 @@
[Unit]
Description=Daily ai-memory snapshot
[Timer]
OnCalendar=daily
AccuracySec=15min
Persistent=true
RandomizedDelaySec=5min
[Install]
WantedBy=timers.target
+117
View File
@@ -0,0 +1,117 @@
#!/usr/bin/env bash
# ai-memory-snapshot.sh
#
# Daily snapshot of an ai-memory single-user install:
# 1. rsync the mutable data dir into a git-tracked mirror checkout,
# excluding derived state and secrets (see .gitignore in this dir).
# 2. commit + push the mirror to a remote git repository when anything
# changed.
# 3. write a compressed tarball of the same tree to a second location
# (a mounted external disk, a network share, or on WSL2 the Windows-side
# filesystem) so the SQLite index can be recovered from disk even if the
# remote repository is unreachable.
#
# Assumptions:
# - You already ran `git init && git remote add origin <url>` inside
# $REPO_DIR and staged the .gitignore shipped alongside this file.
# - The user running this script has an ssh-agent or credential helper set
# up so `git push` is non-interactive.
# - $DATA_SRC is the ai-memory data dir (`ai-memory status` prints it) and
# $CFG_SRC is where your provider env-file lives.
#
# Configure these four variables for your host, or export them from the
# systemd unit's Environment= lines. Sample values are commented out.
: "${REPO_DIR:=$HOME/ai-memory-backup}" # local mirror checkout, `git init`'d and pointed at a remote
: "${DATA_SRC:=$HOME/.local/share/ai-memory}" # ai-memory data dir
: "${CFG_SRC:=$HOME/.config/ai-memory}" # ai-memory config dir (may hold provider env files)
: "${TARBALL_DIR:=$HOME/backups/ai-memory}" # second-location tarball drop (external disk, network share, ...)
: "${LOG_DIR:=$HOME/logs/ai-memory-snapshot}" # rotated per run
: "${TARBALL_RETENTION_DAYS:=30}" # how long to keep tarballs
: "${LOG_RETENTION_DAYS:=30}" # how long to keep run logs
: "${SECRET_EXCLUDES:=*.env auth.json .secrets}" # glob patterns for files never to sync
set -euo pipefail
stamp=$(date +%Y-%m-%d)
log="$LOG_DIR/ai-memory-snapshot_$(date +%Y%m%d-%H%M%S).log"
mkdir -p "$LOG_DIR" "$TARBALL_DIR"
exec >"$log" 2>&1
echo "=== $(date -Iseconds) ai-memory snapshot start ==="
echo "data_src=$DATA_SRC"
echo "cfg_src=$CFG_SRC"
echo "repo_dir=$REPO_DIR"
echo "tarball_dir=$TARBALL_DIR"
if [ ! -d "$DATA_SRC" ]; then
echo "ERROR: $DATA_SRC does not exist. Aborting." >&2
exit 1
fi
if [ ! -d "$REPO_DIR/.git" ]; then
echo "ERROR: $REPO_DIR is not a git repository (run 'git init' and 'git remote add origin <url>' first)." >&2
exit 1
fi
# Build a comma-free rsync exclude list. The .gitignore in the mirror repo
# also keeps derived state out of commits, but excluding here keeps the rsync
# working set small.
rsync_excludes=(
--exclude=.git/
--exclude=models/ # embedder weights, redownloadable
--exclude=db/ # SQLite index, derived from wiki via `ai-memory reindex`
--exclude=logs/ # rotated tracing output
--exclude=.serve.lock # server pidfile
)
for glob in $SECRET_EXCLUDES; do
rsync_excludes+=(--exclude="$glob")
done
echo "--- rsync data ---"
rsync -a --delete "${rsync_excludes[@]}" "$DATA_SRC/" "$REPO_DIR/data/"
echo "--- rsync config (secrets excluded) ---"
config_excludes=()
for glob in $SECRET_EXCLUDES; do
config_excludes+=(--exclude="$glob")
done
rsync -a --delete "${config_excludes[@]}" "$CFG_SRC/" "$REPO_DIR/config/"
cd "$REPO_DIR"
git add -A
if git diff --cached --quiet; then
echo "no changes since last snapshot; skipping commit + push"
else
git commit -m "auto: snapshot $stamp" | tail -3
echo "--- push ---"
timeout 180 git push origin "$(git rev-parse --abbrev-ref HEAD)" 2>&1 | tail -10
fi
echo "--- tarball to second location ---"
tarball="$TARBALL_DIR/ai-memory-$stamp.tar.gz"
tar_excludes=(
--exclude=.git
--exclude=models
--exclude=db
--exclude=logs
--exclude=.serve.lock
)
for glob in $SECRET_EXCLUDES; do
tar_excludes+=(--exclude="$glob")
done
# Use paths relative to $HOME so the archive is portable across usernames.
tar czf "$tarball" "${tar_excludes[@]}" \
-C "$HOME" \
"${DATA_SRC#$HOME/}" \
"${CFG_SRC#$HOME/}"
echo "tarball: $(ls -lh "$tarball" | awk '{print $5, $9}')"
echo "--- rotate tarballs (retention: ${TARBALL_RETENTION_DAYS}d) ---"
find "$TARBALL_DIR" -maxdepth 1 -name 'ai-memory-*.tar.gz' -type f \
-mtime "+${TARBALL_RETENTION_DAYS}" -delete -print
echo "--- rotate logs (retention: ${LOG_RETENTION_DAYS}d) ---"
find "$LOG_DIR" -maxdepth 1 -name 'ai-memory-snapshot_*.log' -type f \
-mtime "+${LOG_RETENTION_DAYS}" -delete -print
echo "=== $(date -Iseconds) ai-memory snapshot done ==="