Machine Learning

When GPU Utilization Lies: The Hidden Systems Problem Slowing Modern AI

When GPU Utilization Lies: The Hidden Systems Problem Slowing Modern AI

team gets pinged because inference latency has suddenly jumped by 60%. The dashboards are confusing. GPU utilization still looks healthy:…
NuCS vs Choco: A Pure-Python Constraint Solver Meets a JVM Veteran

NuCS vs Choco: A Pure-Python Constraint Solver Meets a JVM Veteran

NuCS is a constraint solver written 100% in Python, developed by me, accelerated by NumPy and Numba. Choco is one of the reference open-source constraint solvers, written in Java…
How to Code with Claude's Code

How to Code with Claude's Code

they are amazing for quickly executing lots of code. However, if you've worked a lot with coding agents, you'll notice…
How to Train a Scoring Model in the Age of Artificial Intelligence

How to Train a Scoring Model in the Age of Artificial Intelligence

All code used in this section is available on GitHub. The business logic and modeling functions are located in the…
Beyond extract_text: The Two Layers of a PDF That Drive RAG Quality

Beyond extract_text: The Two Layers of a PDF That Drive RAG Quality

in a RAG process, the parser has one job. Read the document the way a human would before answering a…
Bayesian Networks and Markov Networks: An Intuitive Guide to Structured Uncertainty

Bayesian Networks and Markov Networks: An Intuitive Guide to Structured Uncertainty

explanations begin with prediction. A churn model estimates whether a customer is likely to leave. A fraud model estimates whether…
Physical AI: What It Is and Is

Physical AI: What It Is and Is

Physical AI is . NVIDIA is talking about it, consulting firms are talking about it, and so are investors and…
10 Common RAG Mistakes We Keep Seeing in Production

10 Common RAG Mistakes We Keep Seeing in Production

I of this series with Angela Shi. This pitfalls article lists the failure modes we both kept seeing on production…
The Hardware That Makes AI Happen

The Hardware That Makes AI Happen

AI, we often describe it as a software revolution, which it is! From the development of neural networks and transformers…
Prefill Once, Fan Out: KV Snapshot Sharing for Multi-Agent LLM Pipelines

Prefill Once, Fan Out: KV Snapshot Sharing for Multi-Agent LLM Pipelines

A humorous-but-real tour of SwarmKV — KV-snapshot fan-out, copy-on-fork host buffers, and how to make a two-agent analytical pipeline ~1.95×…
Back to top button