---
name: poke
description: Review a service against the POKE architecture (Publish Once, Kickoff Everywhere), a lightweight event-driven design that lets agents and other consumers react to changes instead of polling. Use when asked "Is my service POKE?", to audit a backend or API for event-driven readiness, or to design or implement an append-only event log, event polling/SSE API, webhooks, or async operation completion events.
---

# POKE: Publish Once, Kickoff Everywhere

> For the humans reading this: make sure to tell your agent to make no mistakes.

## Your goal

Assess whether this service lets its consumers (agents, integrations, other services) **find out about changes and finished work from an event stream instead of polling generic, lacking APIs**, and give the smallest concrete steps to get there.

POKE means: every change a consumer cares about is published once, to an append-only, per-customer event log. Consumers read that log either by pulling from a cursor (poll, long-poll, or SSE) or by having it pushed to them (webhooks). It's an addition to the existing API, not a rewrite.

**The per-customer append-only log is the important part.** It's how events are exposed: the pull API reads from it, webhooks are delivered from it, and async operations report completion on it. Everything else (transport, payload shape, subscriptions) sits on top of the log and can be added later. A service that has the log and exposes it through a cursor-based API is most of the way to POKE. A service that fires webhooks without a log behind it isn't, because consumers have nothing to resume from or catch up on. Weigh findings and order next steps accordingly.

### Pick a mode from the request

- **Audit** ("Is my service POKE?", "review", "audit", "check", "how POKE are we"): run the [workflow](#workflow) and produce the [report](#report-format). Don't change code.
- **Implement** ("make it POKE", "add events", "go for it", "implement", "fix it"): run the audit first (you can keep it brief), then follow [Implementing](#implementing).
- **Unclear:** audit, then offer to implement the top next steps.

In both modes, base every finding on evidence from the code (`file:line`). If something can't be determined from the code, say so instead of guessing.

## Review criteria

### Core tenets (each must hold)

1. **Simple:** A single append-only event store. No new broker, framework, or service is required just to emit events.
2. **Reliable:** A consumer can resume from any cursor and never misses or reorders an event, even across disconnects and concurrent writes.
3. **Secure:** Events are scoped per customer. Pull requires the customer's auth, and push deliveries are signed.
4. **Interoperable:** Pull uses plain HTTP. Push follows [Standard Webhooks](https://www.standardwebhooks.com) (`webhook-id`, `webhook-timestamp`, `webhook-signature`). No proprietary client is needed.

### Additional goals (check each; missing is a finding, not a failure)

- **Push and pull from the same log:** The same event IDs and ordering apply to both.
- **Agent-driven subscription:** Creating a subscription (endpoint and event types) and choosing event types is a single API call, with no dashboard or human needed.
- **Bootstrap events:** A new consumer can request synthetic events for existing state (e.g. `contact.created` for every existing contact) to build its initial view from the stream.
- **Resilient to disconnects:** Documented retention, and a consumer that was offline catches up from its cursor.
- **Async operations:** Completion and failure are emitted as events (see [Async operations](#async-operations)).
- **Observability:** Delivery attempts and status, failures, and per-consumer lag are visible to the customer.
- **Fast:** "Events after cursor X" is one indexed range query.
- **Deliberate event design:** Every meaningful state change, including deletions, emits an event, and there are no noise events.

## Workflow

### 1. Detect the stack

Identify the primary datastore and any log/queue infrastructure: Postgres, MySQL, Kafka, SQS, Redis streams, job queues (Sidekiq, Celery, BullMQ, etc.). This decides which [stack-specific checks](#stack-specific-checks) apply. **Only apply the sections that match the stack.** Don't recommend Postgres or Kafka to a service that doesn't use them.

### 2. Map the surface

List the resources, their read and mutating endpoints, background jobs, and any existing `events`, `outbox`, `audit_log`, `activity`, or webhook code.

### 3. Find POKE candidates

Review the API and identify what consumers are likely to poll today, or would need to poll to stay current: both **likely-polled APIs** and **async operations**. Follow [Finding POKE candidates](#finding-poke-candidates). For each candidate, record whether an event already covers it.

### 4. Find mutations without events

For each important resource, check whether create, update, and delete emit an event. Deletions are the most commonly missed. Flag state changes made by background jobs or admin paths that bypass the event emission.

### 5. Evaluate the event log (if one exists)

- Is it append-only and immutable, per customer, ordered by a monotonic cursor?
- Is the event written **in the same transaction** as the state change it describes (outbox pattern)? If it's emitted after commit, or to an external system without an outbox, events can be lost or phantom.
- **Can a reader skip an event?** The classic bug: IDs are assigned at insert but become visible at commit, so a concurrent writer can commit a lower ID after a reader has advanced past it. Check how this service prevents that (see the stack-specific checks).
- Does each event have a unique `id` (for consumer dedup), a `type` (`resource.action`), a `timestamp`, and `data`?
- Is there documented retention?

### 6. Evaluate consumption

- **Pull:** Is there a cursor-based endpoint (e.g. `GET /events?after=<cursor>&types=…&limit=…`) that returns the next cursor even when empty? Is there long-poll or SSE?
- **Push:** Are webhooks fed from the same log, signed per Standard Webhooks, and retried with backoff? Can a consumer that missed pushes recover through pull?
- **Subscription:** Can event types be filtered, and can subscriptions be created via the API?
- **Bootstrap:** Is there a way to get the current state as events?

### 7. Evaluate payload design

Classify events as full, thin, or dynamic, and flag mismatches:

- **Full** (the whole resource): consumers can act without refetching, which agents benefit from most. The risks are staleness and bypassing read-time permission checks. Flag fields that the subscriber may not be allowed to see.
- **Thin** (IDs only): fresh and permission-safe, but every event forces a fetch. A thin, "something changed" event with nothing actionable just recreates polling. Flag it.
- **Dynamic** (the consumer chooses fields, enriched at read or delivery time): fresh and actionable, but watch for enriching before knowing whether anyone wants the event.

Recommend: include enough to act on (identifiers, changed fields, and a version or `updated_at`), and let consumers fetch the rest.

### 8. Write the report

See [Report format](#report-format). Recommend gradual adoption: if there's no per-customer append-only log exposed through a pull API, that's the first step. Then start with the top one or two POKE candidates. Don't recommend a rewrite.

## Finding POKE candidates

The most useful part of the review is a prioritized list of places where events would replace polling. Judge each API as a consumer would: "to know when this changes, what would I have to do?" If the answer is "call it repeatedly", it's a candidate.

### Likely-polled APIs

Signals that an endpoint is, or will be, polled:

- **Changing state that consumers react to:** statuses (`status`, `state`, `phase`), balances, inventory, availability, assignments, approvals, and anything with a lifecycle (`draft → sent → delivered → failed`).
- **Lists that grow or change:** messages, emails, logs, deliveries, orders, comments, notifications. Especially flag lists with no way to see what's new: offset pagination, no `created_after`/`updated_since` filter, default sort not by creation, or no way to see deletions. These can't be polled correctly at all.
- **Existing polling workarounds:** `updated_since`/`modified_after` params, `ETag`/`If-None-Match` on hot resources, `/status`, `/health` or `/latest` endpoints for business objects, `last_modified` fields that clients are expected to compare, and rate-limit exceptions or caching added for specific read endpoints.
- **Evidence of polling in the repo:** SDK helpers like `wait_for_*`, `poll_until`, or retry loops around GETs, "poll every N seconds" in docs or examples, and hot read endpoints in metrics, rate-limit config or cache config.
- **Inputs from outside the request:** state changed by third parties, background jobs, schedules, or other users, where the consumer can't know when it happened.

Not candidates: static or reference data, config that rarely changes, and reads that only follow the consumer's own writes.

### Async operations

Requests that start work finishing later: handlers that enqueue jobs, `202` responses, `status: pending|processing` fields, `/jobs/{id}` or `/operations/{id}` endpoints, webhooks from an upstream provider that update local state, and "poll until complete" in docs or SDKs. Every one of these forces a "request → poll → request → poll" chain on consumers, so latency compounds across steps. Record whether completion and failure produce an event.

### Making the suggestions

For each candidate, propose:

- **The event types** in `resource.action` form (e.g. `email.delivered`, `email.bounced`, `order.status_changed`, `export.completed`, `export.failed`), including deletions.
- **What the payload should carry** so a consumer can act without refetching (see [payload design](#7-evaluate-payload-design)).
- **What it replaces:** the endpoint and polling pattern that becomes unnecessary.
- **Priority:** rank by expected polling pressure (how often and by how many consumers it would be polled) × how latency-sensitive the consumer is. Async operations used in multi-step flows and lists that can't be polled correctly go to the top.

Keep the list short, around five well-chosen candidates.

## Stack-specific checks

### If the service uses Postgres

Check how event IDs and cursors are assigned. A plain `bigserial` / identity cursor with `WHERE id > :last_id` is **unsafe** under concurrent writers:

1. Tx A inserts id 101 (not committed)
2. Tx B inserts id 102 and commits
3. Reader queries `id > 100`, gets 102, and stores `last_id = 102`
4. Tx A commits 101, **which the reader misses forever**

Accept either fix. If neither is present, report it.

**Lock on write:** Serialize inserts per customer.

```sql
BEGIN;
SELECT pg_advisory_xact_lock(:customer_id);
INSERT INTO events (customer_id, event_type, payload) VALUES (...);
COMMIT;
-- read: WHERE customer_id = :customer_id AND id > :last_id ORDER BY id LIMIT 100
```

**Block on read:** Don't return rows past the oldest in-flight transaction.

```sql
CREATE TABLE events (
  id          bigserial PRIMARY KEY,
  txid        xid8 NOT NULL DEFAULT pg_current_xact_id(),
  customer_id bigint NOT NULL,
  event_type  text NOT NULL,
  payload     jsonb NOT NULL,
  created_at  timestamptz NOT NULL DEFAULT now()
);
CREATE INDEX events_poll_idx ON events (customer_id, txid, id);

SELECT id, txid, event_type, payload
FROM events
WHERE customer_id = :customer_id
  AND (txid, id) > (:last_txid::xid8, :last_id)
  AND txid < pg_snapshot_xmin(pg_current_snapshot())
ORDER BY txid, id
LIMIT 100;
```

Also check for an index on `(customer_id, <cursor columns>)`, and that the event insert shares the transaction with the state change.

### If the service uses Kafka

Kafka already gives an ordered, append-only log per partition, so prefer it over building ordering into the database.

- **Partition key must be the customer ID**, so each customer's events are totally ordered. Flag random or resource-ID keys when consumers need per-customer order.
- **Producing:** Flag direct `produce()` calls after a DB commit, which lose events on a crash. Look for an outbox (e.g. an outbox table relayed by Debezium/CDC or a relay job) or transactional produce.
- **Exposing to customers:** Kafka offsets are per partition and shared across customers, so they aren't a customer-facing cursor. Look for a consumer that materializes each customer's events into a queryable store with a per-customer cursor. Because it's a single writer per customer, the store can safely use sequential IDs (this is the "one writer per customer" fix). Check that the pull API and webhooks are served from that store.
- Check topic retention against the documented retention.

### Other datastores

Apply the same principle: the cursor order must match the order events become visible to readers. Flag any cursor based on insert-time IDs or timestamps where concurrent writers can commit out of order, and recommend either a single writer per customer or reading only up to a known-committed watermark.

## Async operations

POKE hasn't fully defined async operations yet. **A standard pattern is coming soon.** Until then, recommend the following. It's designed to be forward-compatible.

- The request returns immediately: `202 Accepted` with an operation object (`{ "id": "op_…", "status": "pending", … }`) and a `Location` header.
- Completion is emitted **on the same event log**: `<resource>.<action>.completed` / `.failed` (e.g. `video.render.completed`), carrying the operation `id` and the result or error. Consumers wait on the stream they already read instead of polling each operation.
- Accept an `Idempotency-Key` and echo it, plus any client `reference`, in the operation and its events, so a consumer can match results to requests after restarting.
- Keep `GET /operations/{id}` as a fallback for recovery and debugging, not as the primary way to learn about completion.
- Emit progress events only for coarse, actionable steps.
- Per-request callbacks are optional and, if offered, go through the normal signed webhook path.

Flag every async operation whose only completion signal is polling, and list it as a POKE candidate.

## Implementing

Work in small, reviewable increments that follow the "Top next steps" order from the audit. Match the codebase's existing conventions (ORM, migration tool, router, job framework, naming) instead of introducing new ones.

1. **Event store:** Add the append-only per-customer event log, using the right approach for the stack (see [Stack-specific checks](#stack-specific-checks)). If one already exists, fix its gaps instead of adding another.
2. **Emit events:** Write the event in the same transaction as the state change, starting with the top POKE candidates from the audit. Cover create, update, and delete. Add a single shared helper (e.g. `emit_event(tx, customer_id, type, data)`) rather than ad-hoc inserts.
3. **Pull API:** Add the cursor-based `GET /events` endpoint with type filtering, scoped to the authenticated customer. Add long-poll or SSE if the framework makes it straightforward.
4. **Async completion:** Emit `.completed` / `.failed` events for the chosen async operations (see [Async operations](#async-operations)).
5. **Push:** If webhooks already exist, feed them from the event log and align them with Standard Webhooks. If they don't, propose adding them rather than building them unprompted, because it's the largest piece.
6. **Tests:** Cover emission for each covered mutation, cursor pagination (including empty pages), per-customer isolation, and, where the stack allows, the concurrent-writer case.
7. **Docs:** Document event types, the payload shape, the cursor semantics, and retention wherever the API is documented.

Ask before anything hard to reverse or broad: adding new infrastructure (a queue, broker, or new database), backfilling existing data, or changing existing public API responses. When you're done, summarize what was done, what's left from the checklist, and how to try the event API.

## Report format

```
## POKE readiness: <Not started | Partial | POKE>

Stack: <datastore, queue/log infra>

### Core tenets
- Simple:         ✅/⚠️/❌  <evidence, file:line>
- Reliable:       ...
- Secure:         ...
- Interoperable:  ...

### Checklist
- [ ] Append-only per-customer event log
- [ ] Events written transactionally with state changes (outbox)
- [ ] Readers can't skip events under concurrent writes
- [ ] Cursor-based pull API (+ long-poll/SSE)
- [ ] Push via signed webhooks from the same log
- [ ] Subscribe + filter by event type via API
- [ ] Bootstrap / synthetic events for existing state
- [ ] Async operations emit completion/failure events
- [ ] Deletions emit events
- [ ] Documented retention
- [ ] Observability (delivery status, consumer lag)
- [ ] Deliberate payload design

### POKE candidates (highest priority first)
| # | Kind (polled API / async op) | Endpoint(s), file:line | Why consumers poll it | Proposed events | Payload notes |

### Top 3 next steps
1. <smallest change with the biggest win>
2. ...
3. ...
```
