experiments

things we’re running, poking, or still sketching. status is honest.

agent notebooks

active

structured run logs for multi-step agent work — what ran, what failed, what we kept.

eval harness

active

small, boring fixtures that catch regressions before a demo does.

prompt diff

exploring

treat prompt changes like code: reviewable diffs, named versions, rollback.

tool budgets

exploring

caps on tool calls and tokens so agents stay useful under cost pressure.

sage ui kit

research

book-jacket paper, iron gall ink, and calm type for builder-facing tools.

memory sketch

research

lightweight session memory that prefers notes over opaque embeddings.