No. 12 · Manuscript complete · AI & agents · TRAIN
ML for Agent Builders
Auto-Training, Fine-Tuning, and Eval Loops for LLM-Powered Systems
Your agents call models all day. Time to own a few.
Agents lean on models constantly, but the decision to own one stays murky. ML for Agent Builders is a build loop for the narrow models agents depend on: clear calls on fine-tune versus RAG, labeling, and eval drift, so you win on cost, latency, and control instead of renting everything.
- pages
- 410
- chapters
- 15
- hours of reading
- ± 5
- editions
- EN · NL
- Design
- Drafting
- Manuscript
- Production
- Launched
The editions
Kindle, Paperback and Hardcover
Production shots of the editions. Not for sale yet — this page is the public working version.
-
Kindle
The digital edition, ready in any Kindle app.
-
Paperback
The working copy. Mark it up. Take it into the meeting.
-
Hardcover
The shelf edition. Dark full-art jacket, built to stay in the room.
The book
Auto-Training, Fine-Tuning, and Eval Loops for LLM-Powered Systems
A support agent hands every message to a frontier model to pick a queue. The call costs money, adds latency, and depends on a vendor who can change pricing or models. You can write Python and wire an API. You have not run a full training pipeline, and the "fine-tune it or add RAG" argument can end without a question anyone can test.
Chapter 9 invents a quieter failure to show what an aggregate can hide. Validation accuracy looks steady near 0.89 for eight weeks while three billing disputes sit in the wrong queue for four days. The model architecture had not changed. Nor had the training data. The product had. An aggregate can pass while a class fails, and a final macro number is not evidence about the class whose failure costs the most.
An owned model beside your LLM is a candidate, not an improvement, until a comparison on the same messages and acceptance criteria says otherwise. The general-purpose model can remain the right choice. Write down why.
TRAIN is the author's five-discipline loop, and each leaves artifacts:
What you learn
What this book puts in your hands
- The TRAIN protocol: Target metric, Represent data, Automate, Inspect, Narrow
- Decide fine-tune vs RAG, control labeling and eval drift
- Narrow models that win on cost, latency and control
The framework
TRAIN, step by step
Target metric, Represent data, Automate, Inspect, Narrow
-
Target metric
-
Represent data
-
Automate
-
Inspect
-
Narrow

From the book
Figure 7.1: High, middle and low confidence proposals all pass human-owned
ML for Agent Builders, printed page: Figure 7.1: High, middle and low confidence proposals all pass human-owned
Inspect more pagesLook inside
Pages from the print edition.
The contract is a living document; any change to the label taxonomy
Step 1: Name the candidate and its core task
Practice: Synthetic Data QA Checklist 6 13 135
Figure 7.1: High, middle and low confidence proposals all pass human-owned
Figure 8.1: Immutable offline artifacts lead to evaluation; only authorized
A flagged or sampled row cannot be closed by batch acceptance, so
Confirm the approved response, including abstention or block where
Practice: First owned-model scan 13
The contents
Chapter by chapter
Every chapter of ML for Agent Builders with its printed epigraph, what you can do afterwards, and the moment it is built for.
Introduction
Introduction: Models your agents own
Before you train a model, name the task that would justify maintaining it.
Chapter 1
Build vs Buy vs Prompt
The decision you keep postponing is already being made: by default, by inertia, by whoever wrote the architecture doc.
Chapter 2
TRAIN Overview
A lab experiment runs until you get tired of it. A product loop runs until the product does not need it anymore.
Chapter 3
Target the Metric
A model judged without a metric is a compass needle with no north. It spins, eventually settles, and points nowhere in particular.
Chapter 4
Represent the Data
Before you train, write down what each row means, where it came from and which decisions it may inform.
Chapter 5
Fine-Tune vs RAG
Ask which constraint needs relief, then compare what each path costs to satisfy it.
Chapter 6
Synthetic Data with Guardrails
Generated examples are candidates, not independent evidence that their labels are right.
Chapter 7
Agent-Assisted Labeling
A confident proposal is still a proposal. Decide who may turn it into a training label.
Chapter 8
Automate the Pipeline
A training script you run by hand is a manual. A pipeline you run on a trigger is infrastructure.
Chapter 9
Eval Harnesses for Agent ML
Automation without regression testing speeds up failure and lends it the confidence of a working pipeline.
Chapter 10
Deploy Beside the LLM Router
A model that only exists in a notebook is not yet part of the agent system. Deployment is the decision that turns an experiment into infrastructure.
Chapter 11
Inspect Drift
Every model is a photograph of the world at the moment you trained it. The question is whether the world has moved since.
Chapter 12
Narrow Scope Discipline
Write the model's boundary before asking what else it could do.
Chapter 13
Case Studies, Routing, Rerank, Moderation
Three hypothetical builds make the decisions concrete. Their results are not evidence for yours.
Conclusion
Conclusion: TRAIN on the calendar
A checkpoint ends with a decision, an owner and a date, even when the decision is not to train.
Who it is for
Who this book was written for
Three walkthroughs follow the same structure, from task to serving decision: an intent router, a cross-encoder reranker, a moderation filter. You start with an architectural preference and finish with the comparison on record.
TRAIN is the author's operating proposal, not a validated intervention, and its scenarios supply no measured advantage for owned models. A defensible no-build decision is a successful use of this book.
The reader it was written for
The agent-stack engineer-founder. Seed to Series A, 1–20 people, shipping LLM agents but hitting walls on classifiers, rerankers, and small fine-tunes. Can prompt and wire APIs; does not have a full ML org. Needs owned narrow models without becoming a data scientist.
Also a fit for
The product engineer on an agent team. Responsible for eval quality, labeling throughput, and deployment next to the router — reports to a technical lead who still says "just fine-tune it."
What you will use it on
- Decide build vs buy vs prompt for the next agent dependency
- Stand up a golden set and regression eval before production
- Use agents to accelerate labeling without shipping ungrounded data
- Deploy classifiers/rerankers beside the LLM router
- Monitor drift without a dedicated ML platform team
Probably not for you if
- PhD ML researchers optimizing foundation training
- Readers who only want prompt tips
- Clinical or regulated ML without a safety gate
Editions
Editions and specifications
| Edition | Formats | Chapters | Pages | Reading time | ISBN (paperback) |
|---|---|---|---|---|---|
| English ML for Agent Builders | In production | 15 | 410 | ± 5 hours | — |
| Dutch ML voor agent-bouwers | In production | 15 | 428 | ± 5 hours | — |
Both editions are written natively. The Dutch text is not a machine translation of the English. · Trim size: 6x9″
Frequently asked
What readers usually want to know
What is ML for Agent Builders about?
TRAIN helps agent builders decide when owned ML, fine-tuning, retrieval, evaluation, and drift monitoring are worth the operational cost. The subtitle is: Auto-Training, Fine-Tuning, and Eval Loops for LLM-Powered Systems.
What is the TRAIN framework?
TRAIN: Target metric, Represent data, Automate, Inspect and Narrow. Target metric, Represent data, Automate, Inspect, Narrow
Is there a Dutch edition?
Yes. The Dutch edition is ML voor agent-bouwers, written as a native edition rather than a machine translation. It moves through the same production line.
How long is ML for Agent Builders?
This edition runs 15 chapters, 410 pages in print and roughly 5 hours of reading.
Who is ML for Agent Builders for?
TRAIN is the author's operating proposal, not a validated intervention, and its scenarios supply no measured advantage for owned models. A defensible no-build decision is a successful use of this book.
The production system
How this book was made
Every title moves through the same gated production line: sourced research, a claim-level evidence ledger, structural review, fact-checking, red-team critique, and a bilingual final edit. AI agents do specialist work inside those gates; judgment, voice, and accountability stay human.
- Claims enter an evidence ledger with a source and a confidence grade before they reach the page
- English and Dutch are two native editions, not a translation of one another
- Every chapter clears readability, rhythm, and style gates before it is typeset
The series