Ten Python task cards entering a forge-like benchmark machine and flowing through nine evaluation paths toward cost and correctness instruments

Python Agent Cost Benchmark

Ten controlled Python tasks for measuring how skills and agent configurations affect cost without sacrificing correctness.

  • python
  • agents
  • benchmarking
  • evaluation
Read

First sparks

A short opening note about what this journal is for.

  • process
  • notes
Read

Building Open Experiment Forge

A public workbench for the website itself: static-first, safe, and easy to expand.

  • website
  • astro
  • safety
Read

A safe static website baseline

A practical starting checklist for publishing a simple website with fewer moving parts.

  • security
  • static-site
  • deployment
Read