(view in browser) Sept 30th Next webinar:Oct 07, 2026 - Stop Bad Merges with an LLM Eval Gate Dear Reader, Welcome to the 338th edition of the Data Science Briefing. Announcements Jev, TypeSafe AI's new model on OpenRouter, doesn't generate text. It returns a probability for every answer in one pass, at $0.042 per million input tokens with free output. TypeSafe claims it's 193x faster and 445x cheaper than frontier models. Nobody has reproduced that, so I benchmarked it against Claude Opus...
about 8 hours ago • 5 min read
(view in browser) Sept 16th Next webinar:Sept 23, 2026 - Put Error Bars on Your LLM Metrics Dear Reader, Welcome to the September 16th edition of the Data Science Briefing. Announcements If you missed the live session, the full recording of Bruno Gonçalves' 30-minute Lightning Lesson is now available to watch for free! In "Prove Your Prompt Change Actually Helped," you'll learn: 🔬 How to run a paired test on two prompt versions to cancel out difficulty and isolate the change you made. 📊 How...
14 days ago • 7 min read
(view in browser) Sept 9th Next webinar:Sept 12, 2026 - Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps Dear Reader, Welcome to the September 9th edition of the Data Science Briefing. Announcements In 2001, experts called Wikipedia a "joke," a "do-it-yourself encyclopedia," and a "vandal's playground." Today, it's the gold standard for factual baseline information on the internet. Are Large Language Models (LLMs) following the exact same trajectory? In this new essay,...
21 days ago • 6 min read
(view in browser) Aug 26th Next webinar:Sept 12, 2026 - Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps Dear Reader, Welcome to the 333rd edition of the Data Science Briefing. Announcements Is your new prompt actually better, or did you just get lucky on a sample of 5 outputs? 🤔 Eyeballing LLM outputs might work for quick prototypes, but shipping to production requires real proof.Join us for a free 30-minute workshop on how to run paired prompt tests, filter out...
28 days ago • 5 min read
(view in browser) Aug 26th Next webinar:Sept 12, 2026 - Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps [Register] Dear Reader, Welcome to the 333rd edition of the Data Science Briefing. Announcements Want to know how AI agents actually interact with external databases and tools via MCP? 🛠️ Check out this step-by-step breakdown by Bruno Gonçalves on building a custom Model Context Protocol (MCP) server from scratch using only the Python standard library (raw JSON-RPC...
about 1 month ago • 6 min read
(view in browser) Aug 19th Next webinar:Sept 12, 2026 - Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps [Register] Dear Reader, Welcome to the Aug 19th edition of the Data Science Briefing. Announcements Today we're proud to announce not 1, not 2, but 4 new events coming up in September and October: Sep 9, 2026 - Prove Your Prompt Change Actually Helped Sep 23, 2026 - Put Error Bars on Your LLM Metrics Oct 7, 2026 - Stop Bad Merges with an LLM Eval Gate Oct 16, 2026 -...
about 1 month ago • 7 min read
(view in browser) Aug 12 Next webinar:Sept 12, 2026 - Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps [Register] Dear Reader, Welcome to the Aug 12th edition of the Data Science Briefing. Announcements One good demo is a test flight. An eval suite is the flight-test campaign. A flashy demo only proves your AI agent can work once. To deploy with confidence, you need an operational evaluation framework that measures performance, budgets, latency, and edge-case failures....
about 2 months ago • 8 min read
(view in browser) Aug 5 Next webinar:Sept 12, 2026 - Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps [Register] Dear Reader, Welcome to the 330th edition of the Data Science Briefing. Announcements Put your old Mac to work serving a local LLM. Running a language model on your own hardware costs nothing per token, works on a plane, and never sends your code elsewhere. On Apple Silicon it's fast too — unified memory is a large part of why Mac Minis are back-ordered for...
about 2 months ago • 8 min read
(view in browser) Jul 29th Dear Reader, Welcome to the July 29th edition of the Data Science Briefing. Announcements Put your old Mac to work serving a local LLM. Running a language model on your own hardware costs nothing per token, works on a plane, and never sends your code elsewhere. On Apple Silicon it's fast too — unified memory is a large part of why Mac Minis are back-ordered for weeks. This post covers the whole path: install llama.cpp, pull a GGUF, and serve it over an...
2 months ago • 6 min read