Latest Articles
- 1 day agoAI Coding
When to Use an Eval Framework, and When to Build the Judge
Off-the-shelf metrics cover the failures every LLM app shares. Your domain failures are not on that list. A decision rule for which half you are in.
- 2 days agoAI Coding
The Bugs Only Deploying Finds
Four defects that no unit test can see, found the first time a working service met a real container. Three of them were invisible from the code.
- 4 days agoAI Coding
A Spec for an Agent Is a Test Plan in Disguise
A developer asks about the gap in your spec. A model fills it with the likeliest pattern and says nothing. That one difference changes what a spec has to contain.
- 5 days agoAI Coding
Nothing Leaked. The Answer Just Disappeared.
An isolation test that passes on broken code, three ways it was still wrong after I fixed it, and what found each one. All of them green.