Ten Python task cards entering a forge-like benchmark machine and flowing through nine evaluation paths toward cost and correctness instruments

Python Agent Cost Benchmark

Ten controlled Python tasks for measuring how skills and agent configurations affect cost without sacrificing correctness.

  • python
  • agents
  • benchmarking
  • evaluation
Read

Building Open Experiment Forge

A public workbench for the website itself: static-first, safe, and easy to expand.

  • website
  • astro
  • safety
Read