Python Agent Cost Benchmark
Ten controlled Python tasks for measuring how skills and agent configurations affect cost without sacrificing correctness.
Open experiments, notes, and guides
A working place for projects in progress, small journal entries, and guides that turn experiments into useful paths.
Latest from the forge
Projects, notes, and guides share one archive so related ideas can stay close together.
Ten controlled Python tasks for measuring how skills and agent configurations affect cost without sacrificing correctness.
A public workbench for the website itself: static-first, safe, and easy to expand.
Three paths
Visible workbench notes for things being built, tested, paused, or finished.
Short reflections on decisions, failed attempts, discoveries, and questions.
Reusable explanations for anyone who wants to repeat or adapt an experiment.