Data Science Briefing #338


​(view in browser)​

Sept 30th

Next webinar:​
​Oct 07, 2026 - Stop Bad Merges with an LLM Eval Gate
Count down to 2026-10-07T13:30:00.000Z​

Dear Reader,

Welcome to the 338th edition of the Data Science Briefing.

Announcements

Jev, TypeSafe AI's new model on OpenRouter, doesn't generate text. It returns a probability for every answer in one pass, at $0.042 per million input tokens with free output. TypeSafe claims it's 193x faster and 445x cheaper than frontier models. Nobody has reproduced that, so I benchmarked it against Claude Opus 5.5 on routing, scoring, judging, and agent gating.

That harness is now a workshop: Jev vs Claude: A Measured Guide. It runs live on November 18. Code JEV2020 takes 20% off.

You can also start with a free 30-minute lesson:

​Agents forget, and they push. One developer lost half a session to silent context compaction and now keeps memory in files, with handoff documents, shared specs and a clear test of done. His agent then deferred 11 of 15 review findings into a release that did not exist. A rule that lets the model stop at done fixed it. Long sessions cost money too. Cache reads dominate an agent’s bill, and output tokens make up only 9 to 18 percent of spend.

Classification is back in fashion. A history of text classifiers ends with Jev, which scored 96.47 percent on IMDb for 65 cents with no fine-tuning. One researcher traces its secret to calibrated multiway reward modeling. At the other extreme, gzip becomes a text generator with no weights at all. A naive $2,071 earnings gap for a job training program hides large differences between the two groups. A seasonal ARIMA beats a naive household power forecast by a small margin, and that margin is the lesson.

Human activity leaves a mark on disease. A study of 58,318 outbreaks of 32 diseases in 169 countries ties risk to people and livestock living near fragmented forests. Each extra hour of travel to a clinic cuts the odds of detecting an outbreak by 32 percent, so many hotspots mark good surveillance. Dengue risk is rising in Europe, and awareness is high. In France and Italy, 91 percent of 2,904 people surveyed knew of dengue, but only 59 percent took even one preventive step. Missing information carries its own risk. Across vaccine rollouts in six European countries, information voids lined up with more misinformation.

Machines now stand in for people, and they pick up our social habits. A new paper weighs when language models can replace human subjects in research. Classic network models still teach a lesson. Coupling a cooperation game with a second social process, like opinion spread, helps prosocial behavior win. The same social forces expose AI agents. In populations of language model agents, a small adversarial group can flip a shared convention through stepping-stone states, with fewer members than a direct attack needs.

Attacks and defenses keep scaling. One model writes 200 jailbreak suffixes for a single query in 4 seconds, with near 100 percent success on two open models and 99 percent on GPT-3.5. On defense, Jev screens for alignment failures with no training. Across 44 benchmarks and 10 failure types, it reached a median AUROC of 0.886 at 63 times lower cost than LLM judges.

Our latest book recommendation is "Hands-On LLM Serving and Optimization" by C. Wang and P. Hu. In this week's video, we have a lecture on Deterministic Concurrency.

Data shows that the best way for a newsletter to grow is by word of mouth, so if you think one of your friends or colleagues would enjoy this newsletter, go ahead and forward this email to them. This will help us spread the word!

Semper discentes,

The D4S Team


A 14-billion-parameter model claims 28 gigabytes of memory before the first token. The KV cache then grows with every token of every open request, and real traffic opens many at once. This 374-page manual lives inside that squeeze. Chi Wang runs model-inference engineering teams at Salesforce. Peiheng Hu builds distributed inference engines at NVIDIA. The two spent over eight years building AI systems together, and it shows. Their claim is blunt. Training gets the papers, and serving carries the product. Welcome to the inference era.

The method sets Hands-On LLM Serving and Optimization apart. You build a serving service from scratch, batching and streaming included, and only then meet vLLM, TensorRT-LLM, SGLang, and llama.cpp. The frameworks stop looking like magic. The arithmetic alone earns the cover price. Estimate the weight footprint, size the KV cache, compute the arithmetic intensity of prefill and decode, and you know whether a workload is compute-bound or memory-bound before renting a single GPU. Continuous batching, quantization, speculative decoding, and the four parallelisms follow, each with runnable code, and an eight-step tuning plan for Qwen3-14B ties the whole book together. Data scientists get the systems course their training skipped. Platform engineers get a build-or-buy playbook.

Two warnings before you buy. The book went to press in April 2026, pinned to Qwen3-14B and today’s engine flags, and those engines ship faster than any press run. Treat the companion repository as the living copy. The second warning concerns fit. Many notebooks want an NVIDIA GPU, and the book stays on the serving side of the wall. Training, fine-tuning, and evals sit outside its scope, so readers hunting for modeling content will find queues and schedulers instead. Read it anyway. The publisher clocks it at 11 hours, one weekend. The serving layer decides what a model costs and how fast it feels, and most teams inherit it as a black box. This book opens the box and labels the parts. The next time the latency graph spikes or the GPU bill doubles, you will know which knob to turn first.


  1. ​Your AI Agent Already Forgot Half of What You Told It [oreilly.com]
  2. ​Causal questions and the jobs program data [stochastic.blog]
  3. ​(KV) Cache Rules Everything Around Me [completeskeptic.com]
  4. ​What Is RLCD? The Secret Behind Jev [di-zhang-llm.github.io]
  5. ​Can gzip be a language model? [nathan.rs]
  6. ​A dataset hub for LLM serving research [data.agentic-system.org]
  7. ​Forecasting household power use [stochastic.blog]
  8. ​My AI Kept Pushing Me to Ship, So I Asked It Why [oreilly.com]
  9. ​Mathematicians Build Long-Awaited Graph Sandwich [quantamagazine.org]
  10. ​Language Models for Text Classification: From Bag-of-Words to Jev [magazine.sebastianraschka.com]


Deterministic Concurrency

video preview​

All past videos of the week are available on our YouTube playlist.

Upcoming Events:

Opportunities to learn from us.
​

Check out the events page for more details.

On-Demand Videos:

Long-form tutorials
​

Data For Science, Inc

I'm a maker and blogger who loves to talk about technology. Subscribe and join over 3,000+ newsletter readers every week!

Read more from Data For Science, Inc

(view in browser) Sept 23rd Next webinar:Oct 07, 2026 - Stop Bad Merges with an LLM Eval Gate Dear Reader, Welcome to the September 23rd edition of the Data Science Briefing. Announcements Ever wonder why certain ideas, narratives, or trends spread exponentially across the internet while others fizzle out? Ideas are contagious—and it turns out, we can understand their spread using the same mathematical models we use to track biological outbreaks. By applying the principles of epidemiology and...

(view in browser) Sept 16th Next webinar:Sept 23, 2026 - Put Error Bars on Your LLM Metrics Dear Reader, Welcome to the September 16th edition of the Data Science Briefing. Announcements If you missed the live session, the full recording of Bruno Gonçalves' 30-minute Lightning Lesson is now available to watch for free! In "Prove Your Prompt Change Actually Helped," you'll learn: 🔬 How to run a paired test on two prompt versions to cancel out difficulty and isolate the change you made. 📊 How...

(view in browser) Sept 9th Next webinar:Sept 12, 2026 - Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps Dear Reader, Welcome to the September 9th edition of the Data Science Briefing. Announcements In 2001, experts called Wikipedia a "joke," a "do-it-yourself encyclopedia," and a "vandal's playground." Today, it's the gold standard for factual baseline information on the internet. Are Large Language Models (LLMs) following the exact same trajectory? In this new essay,...