Tag: AI
All the articles with the tag "AI".
Seven Checks for Systems That Look Correct
Published: at 06:00 AMA May–September retrospective reading path through data correctness, Kafka, RAG, caching, MCP authorization and evaluation, with refreshed series chapters.
95% Agreement, Zero Failures Caught: Testing Your LLM Judge
Published: at 06:00 AMAn executable confusion-matrix example shows why agreement can hide a useless judge. Evaluate failure detection and actual task outcomes separately.
Your Cache Hit Rate Is Not Your Token Savings
Published: at 06:00 AMMeasure uncached input, cache writes, cache reads and output separately. Request hit rate cannot tell you what a long-context workload costs.