What the HANDBOOK.md benchmark measures, why the best model still fails two of every three tasks under strict grading, and what that means for the rules you keep in CLAUDE.md an...
Hi, I’m Lucian Ghinda. This is where I write about Ruby and Ruby on Rails — idioms, refactoring, new language features, testing, and the tools I use day to day.
You can browse all posts, the archive, or subscribe with RSS. I also publish the Short Ruby Newsletter.
What the HANDBOOK.md benchmark measures, why the best model still fails two of every three tasks under strict grading, and what that means for the rules you keep in CLAUDE.md an...
One week of agent-first backend work in the logs: 2,200 session files across four tools, 350 prompts typed by hand, and which checks actually found real defects.
Where Claude Code, Codex, Cursor, Amp, opencode, and pi store session logs, which formats they use, how long they keep them, and what remains undocumented.
You can also subscribe with RSS!