Knowledge · Insights
Engineering notes, argued honestly
Practitioner writing on AI, data, and modernization — what works, what breaks, and how we evaluate the difference.
AI Engineering · Evaluation
Evaluating AI Agents in Production
The eval harness is the product: golden tasks, calibrated judges, red-teaming, and eval-gated delivery — what to measure before you trust an agent with real work.
12 min read · Updated September 2026
Read article →On the bench
Topics we're writing next
- RAG retrieval failures: a field guide to chunking, hybrid search, and reranking
- PostgreSQL operations for SQL Server DBAs: the 20% that matters
- Data contracts in practice: schema governance without the bureaucracy
- CDC exactly-once: idempotent sinks and the myth of perfect delivery
Get new notes first: subscribe to the Engineering Brief.
Start here
Talk to an Architect
Bring your hardest AI, data, or modernization problem. We'll tell you plainly whether we can help — and what it takes.