Passing Tests Is the Easy Part
75% of AI coding models introduce regressions when maintaining real codebases over time. Here's what the SWE-CI benchmark reveals and…
Read more about Passing Tests Is the Easy Part75% of AI coding models introduce regressions when maintaining real codebases over time. Here's what the SWE-CI benchmark reveals and…
Read more about Passing Tests Is the Easy PartMost startups send every AI request to their most expensive model. Here is what that default costs and how LLM…
Read more about The One-Model Trap: Why Your LLM Bill Is Higher Than It Should BeAI-generated code now makes up 41% of new code. The teams that moved fastest in 2024 are hitting a year-two…
Read more about The Bill Arrives in Year TwoThe quote was sitting right there in the CTO’s inbox: $45,000 to fine-tune their customer support model so it would…
Read more about You Don’t Need to Fine-Tune ThatPrompt injection is the number one LLM threat on the OWASP list in 2026. This post examines why it bypasses…
Read more about Your AI Feature Has a Security Hole Your Last Audit Won’t FindIt’s Tuesday morning. Your head of customer success pings you: “The AI assistant gave a client wrong instructions for our…
Read more about Your RAG Demo Worked. Your Users Are FuriousTwo rigorous studies found AI coding tools make experienced engineers 19% slower while they feel faster. The skill being eroded…
Read more about Your Engineers Feel 20% Faster. The Data Says Otherwise.Most engineering teams shipping AI features test on five hand-picked examples and call it done. This post examines why AI…
Read more about Your AI Feature Works. You Just Don’t Know If It Works.A VP of Engineering at a Series A fintech posted a single junior full-stack role in April. By the following…
Read more about The Vanishing Junior DeveloperMost AI agents that fail in production fail quietly. This post examines the gap between what an AI agent does…
Read more about When Your AI Agent Hits Production