>

Knowledge · Insights

Engineering notes, argued honestly

Practitioner writing on AI, data, and modernization — what works, what breaks, and how we evaluate the difference.

AI Engineering · Evaluation

Evaluating AI Agents in Production

The eval harness is the product: golden tasks, calibrated judges, red-teaming, and eval-gated delivery — what to measure before you trust an agent with real work.

12 min read · Updated September 2026

Read article →

On the bench

Topics we're writing next

  • RAG retrieval failures: a field guide to chunking, hybrid search, and reranking
  • PostgreSQL operations for SQL Server DBAs: the 20% that matters
  • Data contracts in practice: schema governance without the bureaucracy
  • CDC exactly-once: idempotent sinks and the myth of perfect delivery

Get new notes first: subscribe to the Engineering Brief.

Start here

Talk to an Architect

Bring your hardest AI, data, or modernization problem. We'll tell you plainly whether we can help — and what it takes.