|
Sept 16th
Dear Reader,
Welcome to the September 16th edition of the Data Science Briefing.
Announcements
If you missed the live session, the full recording of Bruno Gonçalves' 30-minute Lightning Lesson is now available to watch for free!
In "Prove Your Prompt Change Actually Helped," you'll learn:
🔬 How to run a paired test on two prompt versions to cancel out difficulty and isolate the change you made.
📊 How to distinguish a real improvement from statistical noise (and how to read p-values and effect sizes in plain language).
📏 How to pick the right sample size that settles the argument before you run the comparison.
Say goodbye to shipping based on opinion or seniority, and start shipping with math.
👉 Watch the free 30-minute recording here:
A proof opens this issue. On September 8, OpenAI published a resolution of the Navier-Stokes existence and smoothness problem. The question had stood open for roughly 90 years, and the Clay Mathematics Institute made it a Millennium Prize problem in 2000. The answer is a disproof. An initially smooth fluid at rest, pushed by a smooth force, develops a singularity in finite time. The work ran on an internal model the lab calls more capable than the one it shipped five days earlier. Coordinating agents did the searching, and the group that cracked it ran on about 10,000 agents at once. They reached the resolution on September 5, about 88 hours after the first agents started, and the Lean formalization took another 17 hours. Across every problem attempted, the agents sent 4.9 million messages and spent about 300 billion output tokens. Navier-Stokes alone took 2.7 million messages and roughly 130 billion. The lab does not intend to claim the prize. Two days before that post, its chief scientist published an essay on what these systems now are. He frames it as intelligence being grown rather than designed. Every large training run is an experiment, and the results surprise the people running it. He splits alignment in two. Goal alignment asks whether the model pursues the task it was handed. Value alignment asks whether it acts reasonably with unclear objectives and in unfamiliar situations. The check on the second one is weakening. Chain-of-thought monitoring has been the lab’s main instrument, and the essay reports its reliability falling. Reasoning now blends with tool calls and conversation; the models are better at steering their own reasoning, and better pretraining made them smarter without verbalizing anything at all. His conclusion is flat. No lab has solved alignment and monitoring well enough to keep scaling at full speed for much longer, and he expects voluntary slowdowns to become common.
The craft half opens on rented hardware. One engineer trained a 3.8-billion-parameter model from random weights for $998. The run took 43 hours on eight B200S and 65 billion tokens, and it scored 0.384 on CORE. A 2019 frontier model with 1.5 billion parameters scores 0.2565 on the same benchmark. He wrote the framework in the evenings, debugged it on a single consumer card, and finished on hardware rented by the hour. Most of the post is throughput accounting. FP8 on all three matrix multiplications bought 25 percent. Padding the vocabulary from 50,257 to 50,304 took the cumulative gain to 33 percent. A fused cross-entropy loss runs 6 percent slower per step and frees 8 GB, so the bigger micro-batch lifts the total to 44 percent. Keeping optimizer master weights in bf16 rather than fp32 cut memory 27 percent and more than doubled throughput on a smaller model. The sharpest finding is about the benchmark, not the model. At a 1024-token context, 3 of the 22 CORE tasks never fit their prompts. One scored exactly zero, and it got worse as the model trained longer. What did doubling the context buy? CORE went from 0.338 to 0.384, and those cropped tasks account for 83 percent of the jump. The other twenty moved 0.008 combined. Measure the harness before crediting the model. Spare hardware gets the next read. A free router now spreads local inference across the machines already sitting on a home network. It finds compatible systems, then sends each independent request to whichever one has room. Support covers GeForce RTX 20 Series and newer, DGX Spark, and Mac M4 or newer, with 8 GB of memory and 20 GB of disk. It speaks to Ollama and LM Studio, so agent code needs no changes. One demonstration ran five subagents across three devices in 8 minutes 48 seconds, compared to 18 minutes on a single laptop. Prompts and files stay on the network, and the tool runs with no internet connection. The machines stay separate too.
The paper stack opens on the person doing the offloading. A cognitive science review asks the blunt question and returns a split answer. It separates two things people lump together. Skills are learned and held by practice, like diagnosis, arithmetic, or writing. One study revives social psychology’s oldest trick and runs it on four reasoning models. Each agent divides points among anonymous peers carrying nothing but an arbitrary group label. Meaningless labels produced in-group favoritism, and a group-blind control made it vanish. Credit is where things get hard for the humans too.
A position paper prices the human labor inside training data across 64 models released between 2016 and 2024. Ask what it would cost to hire people to write those corpora from scratch, and the answer beats the compute bill by one to three orders of magnitude. Several recent models rest on datasets the authors conservatively value above $10 billion.
The second half turns to the machinery, starting with a decay curve. A controlled study measures how fast agents fall apart over long tasks. It covers nine models, six of them open and ranging from 1.2 billion to 671 billion parameters, plus three deployed commercial systems. Four task families, five horizons, and three context regimes produced 10,664 analyzed trajectories. Success follows a geometric law set by a single number, the per-step reliability. That number climbs with model scale and saturates below 1 for every model tested, which guarantees collapse at a long enough horizon.
Our latest book recommendation is "Hands-On LLM Serving and Optimization" by C. Wang and P. Hu. In this week's video, we have a lecture on The Foundations of Modern AI: Generalization, Data Selection, and Epiplexity.
Data shows that the best way for a newsletter to grow is by word of mouth, so if you think one of your friends or colleagues would enjoy this newsletter, go ahead and forward this email to them. This will help us spread the word!
Semper discentes,
The D4S Team
A 14-billion-parameter model claims 28 gigabytes of memory before the first token. The KV cache then grows with every token of every open request, and real traffic opens many at once. This 374-page manual lives inside that squeeze. Chi Wang runs model-inference engineering teams at Salesforce. Peiheng Hu builds distributed inference engines at NVIDIA. The two spent over eight years building AI systems together, and it shows. Their claim is blunt. Training gets the papers, and serving carries the product. Welcome to the inference era.
The method sets Hands-On LLM Serving and Optimization apart. You build a serving service from scratch, batching and streaming included, and only then meet vLLM, TensorRT-LLM, SGLang, and llama.cpp. The frameworks stop looking like magic. The arithmetic alone earns the cover price. Estimate the weight footprint, size the KV cache, compute the arithmetic intensity of prefill and decode, and you know whether a workload is compute-bound or memory-bound before renting a single GPU. Continuous batching, quantization, speculative decoding, and the four parallelisms follow, each with runnable code, and an eight-step tuning plan for Qwen3-14B ties the whole book together. Data scientists get the systems course their training skipped. Platform engineers get a build-or-buy playbook.
Two warnings before you buy. The book went to press in April 2026, pinned to Qwen3-14B and today’s engine flags, and those engines ship faster than any press run. Treat the companion repository as the living copy. The second warning concerns fit. Many notebooks want an NVIDIA GPU, and the book stays on the serving side of the wall. Training, fine-tuning, and evals sit outside its scope, so readers hunting for modeling content will find queues and schedulers instead. Read it anyway. The publisher clocks it at 11 hours, one weekend. The serving layer decides what a model costs and how fast it feels, and most teams inherit it as a black box. This book opens the box and labels the parts. The next time the latency graph spikes or the GPU bill doubles, you will know which knob to turn first.
- The Waymo effect: how AI is quietly making research less collaborative [researchagenda.news]
- Seeing high-dimensional data in two dimensions [stochastic.blog]
- Training a 3.8B LLM to 0.384 CORE for $998 [hugovergnes.github.io]
- What do Visa and Mastercard do? An intro to card networks [tautology.town]
- An Alien Mind [openai.com]
- A comprehensive guide to production-ready RAG [neo4j.com]
- On the Navier–Stokes Millennium Prize Problem [openai.com]
- Clustering and Anomaly Detection on Handwritten Digits [stochastic.blog]
- Personal AI Router for Local Inference [www.nvidia.com]
- Can A.I. “Go Rogue”? [newyorker.com]
- Is AI making us Stupid? (T. N. Cash, M. O. Kelly, B. N. Macnamara, E. F. Risko)
-
An Overview of Large Language Models for Statisticians (W. Ji, W. Yuan, E. Getzen, K. Cho, M. I. Jordan, S. Mei, J. E. Weston, W. J. Su, J. Xu, L. Zhang)
- How an AI math breakthrough ignited a controversy (C. Zhao, A. Cho)
-
GitSkills: A Dataset of Agent Skills on GitHub (G. Destefanis, D. Graziotin, M. Vaccargiu, M. Ortu)
-
The End of Software Engineering: How AI Agents Are Fundamentally Restructuring the Software Paradigm (Z. Cao)
-
STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation (V. Kumar, M. Pulivarthi, V. Kumar, J. Sen, R. A. Bhat, S. Joshi)
- Visual General Intelligence: A White Paper (H. Kataoka, Y. Fukuhara, Y. Tian, S. Wu, O. Deb, R. Yamada, C. Rupprecht, J. Wang, K. Ide, K. Namekata, X. Ma, Y. Chen, R. Geirhos, A. Raghunathan, Y. M. Asano, D. Ramanan, D. Fouhey, A. J. Davison, Y. Du, J. Wu, Z. Liu)
-
Toward a social psychology of AI: language-model agents reproduce human-like minimal-group bias (M. H. J. Lee)
-
How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making (S. Mittal)
-
The Most Expensive Part of an LLM should be its Training Data (N. Kandpal, C. Raffel)
The Foundations of Modern AI: Generalization, Data Selection, and Epiplexity
All the videos of the week are available in our YouTube playlist.
Upcoming Events:
Opportunities to learn from us.
Check out the events page for more details.
On-Demand Videos:
Long-form tutorials
- Natural Language Processing 7h, covering basic and advanced techniques using NTLK and PyTorch.
- Python Data Visualization 7h, covering basic and advanced visualization with matplotlib, ipywidgets, seaborn, plotly, and bokeh.
- Times Series Analysis for Everyone 6h, covering data pre-processing, visualization, ARIMA, ARCH, and Deep Learning models.
|
|
|