AGENTS.md 8.5 KB

AGENTS.md — PPS Quarantine Manager

Read this first. It is the map plus the non-obvious invariants that are expensive to rediscover. Human-readable too, but written so an agent can act correctly without re-deriving the traps.

What this is

A Flask web app for operators to manage a Proofpoint Protection Server (PPS) email quarantine: list quarantined mail, read messages, and run bulk actions (Release, Report & Release, Delete, Move) as asynchronous background jobs. Okta OIDC login, an allowed-users list with roles, and an admin panel that edits the config live.

Run it

python3 -m venv venv && venv/bin/pip install -r requirements.txt
cp config.example.toml config.toml   # then edit: Okta, PPS creds, users, secret_key
venv/bin/python wsgi.py              # NOT app.py — see "App factory" below

Dev without Okta: set [auth] mode = "static" and use [auth.static] credentials. Tests / alternate config: PPSQ_CONFIG=/path/to/config.toml. Plain-http localhost dev also needs PPSQ_ALLOW_INSECURE=1 (allows an http:// PPS base_url).

venv/bin/python -m pytest       # 57 tests, no Okta tenant or real config needed

Module map

File Role
wsgi.py Composition root. The only place that builds the store/queue/prefs, starts the worker thread, configures logging, and serves via waitress.
app.py create_app(store, queue, prefs) — Flask routes. No import-time side effects.
config_store.py ConfigStore: tomlkit load/validate/atomic-write, immutable snapshots, PPSClient rebuild. The heart of the hot-reload design.
auth.py Okta OIDC + static dev mode; login_required/admin_required; is_allowed/safe_next (pure); CSRF.
pipeline.py Data-driven Report & Release step registry (move → release). No Flask.
worker.py JobQueue: SQLite-backed background job thread; ops audit log.
prefs.py PrefStore: per-user preferences in jobs.db (key-value).
admin.py /admin page + /api/admin/* (config, pps-test, jobs, logs).
logs.py Safe backwards log tailing + server-side level classification.
pps_client.py Thin PPS REST client (search/get_raw/act). Cheap to rebuild.
static/, templates/ Vanilla-JS frontend; app.js (quarantine), admin.js (panel).

Invariants — break these and things fail silently

  1. Single process only. The worker thread and both SQLite connections live in-process. Running multiple web workers (--workers N, gunicorn forks) starts competing queue threads on one DB and an in-process config lock that doesn't coordinate file writes. wsgi.py serves single-process on purpose. Do not "scale" it horizontally.

  2. App factory, no import side effects. import app must never read config or start a thread — otherwise every test touches the real config.toml and spawns a worker. Construction happens only in wsgi.py (prod) and create_app(...) (tests).

  3. localguid is folder-local and dies on move; guid is stable. A message's localguid is only valid in its current folder and becomes invalid the moment the message is moved/deleted. The guid is global and stable. To act on a message after a relocating step, re-find its current localguid by matching guid (pipeline.refind_localguids).

  4. Report & Release: move FIRST, then re-read, then release — and release never gets a deletedfolder. The order is move → release (pipeline.ALLOWED_STEPS); the pipeline has no delete step — Delete is a separate action, so Report & Release can never destroy mail as a side effect. Handles are never carried across steps: after the configured delay the worker re-reads each message's current localguid by its stable guid (pipeline.locate_messages), because PPS's backend is eventually consistent and acting on a cached handle after a relocation is the race this ordering exists to kill. release must still not pass deletedfolder — that would relocate the message a second time. See pipeline._release.

Configs written against the old release → move order are re-ordered on load, and a retired delete step is dropped (config_store._migrate) — neither is rejected.

  1. Snapshot per job / per request. Config is an immutable Snapshot; apply() builds a new one and rebinds one attribute (atomic under the GIL). A worker reads one snapshot at job start and threads it through, so a mid-job admin edit affects the next job, never a job in flight. Never re-read store.snapshot() inside a running job. A job that resumes after a step delay necessarily takes a fresh snapshot, so the load-bearing part of its plan (steps, delay) is frozen into jobs.state on the first run and reused.

  2. WAL + busy_timeout are mandatory. Two connections (queue + prefs) share jobs.db. Without PRAGMA journal_mode=WAL and busy_timeout, a prefs write concurrent with the per-chunk job-progress write throws database is locked. Set in JobQueue._init_db and every prefs connection.

  3. PPS search: 1000-row cap, no offset paging. Each search returns at most 1000 rows, newest first. The client pages the whole folder backwards by date using the oldest date seen as a before cursor. page.oldest MUST be computed over the raw batch, before the acted-on filter (app.py api_messages), or paging terminates early and silently truncates the folder. >1000 messages sharing one exact timestamp is unpageable (accepted edge case). Coverage is also bounded by default_days_back.

  4. list_query = "from=*" is a workaround, not a preference. The PPS API rejects a folder-only search; it requires a from/rcpt/subject filter. from=* means "everything". If a deployment doesn't treat it as match-all, change it (e.g. rcpt=@yourdomain.com).

  5. Secrets are write-only. pps.password, okta.client_secret, auth.static.password, app.secret_key are never returned by any endpoint (ConfigStore.redacted() emits a <name>_set boolean). A blank value on save means "leave unchanged". Do not add an endpoint that echoes these back.

  6. Log source is an enum, never a path. /api/admin/logs?source=ops|app resolves to a path server-side. A ?path= parameter would be arbitrary file read — never add one.

  7. The worker thread never sleeps. There is exactly one, so a time.sleep in it stalls every queued job. Work that has to wait (the Report & Release step delay) persists its progress and goes back to status='pending' with a future jobs.run_after (JobQueue._defer); _claim_next skips such rows until the timestamp passes. Anything else that needs to wait must use the same mechanism, not a sleep.

Config zones

config.toml is one file with two zones. The admin panel writes only the admin zone and round-trips the file with tomlkit (comments/formatting preserved).

  • Server-only (edit the file, restart): [okta], [auth] mode + [auth.static], [app] secret_key/listen/port/db_path/worker_log/app_log/cookie_secure.
  • Admin-editable (panel, mostly live): [pps], [quarantine] incl. [quarantine.report_release], [auth] users + denied_message, [app] log_level.

ADMIN_EDITABLE in config_store.py is the explicit allowlist. Adding an editable key means adding it there. docs/configuration.md is the full reference.

Conventions

  • from __future__ import annotations at the top of every module.
  • Loggers are namespaced pps.<area> (pps.app, pps.worker, pps.pipeline, …).
  • %-style lazy log formatting.
  • Private helpers prefixed _; routes named api_* / admin_*.
  • Bare except Exception only where justified, always with # noqa: BLE001 + a reason.
  • PPS failures raise PPSError (carries .status/.body/.reqid), funneled through _pps_error_response → 502.
  • Schema evolution: PRAGMA table_info + ALTER TABLE ADD COLUMN (see worker._init_db).

Extending

  • New Report & Release action: add a @step("name") function in pipeline.py and put "name" in ALLOWED_STEPS (canonical order). If it relocates messages, teach pipeline_destination. The admin checkboxes render from the schema — no frontend change.
  • New config key: add to config.example.toml (with a comment), read it from the snapshot, and if admin-editable add it to ADMIN_EDITABLE + validation in config_store.py.
  • New per-user pref: add a PrefSpec to PREF_SPECS in prefs.py (default path + coercion + validator). No migration needed (key-value store).