Evaluating ML, deep learning, LLMs and reinforcement learning

Published 20 September 2026

ai-engineeringmachine-learningdeep-learningllm-evaluationreinforcement-learning

Suppose you’re evaluating a request filter. Your test set has 1,000 requests, and 20 should be blocked. A filter that allows everything gets 980 decisions right: 98% accuracy, while missing every prohibited request.

That’s a hypothetical example, but it explains why I want to see the counts behind a score before deciding whether a system is useful. In my RAG course assistant, the same headline percentage describes a very different result: 392 correct disposition decisions out of 400 calibration cases, with all eight errors coming from abstaining on answerable questions. The errors determine what I would change next.

Evaluation gets harder when the output is a paragraph or a sequence of actions. There may be several acceptable answers, and a successful final message can conceal a failed operation. I find it more useful to start with the decision the system makes, then choose the measurement that tells me whether that decision worked.

The names overlap

Machine learning is the broad field. Deep learning uses neural networks with multiple layers. LLMs are a family of deep learning models, while reinforcement learning concerns learning behavior through rewards and interaction. An RL policy can use a deep network, and an LLM can receive additional training using RL.

So these aren’t four mutually exclusive evaluation categories. A neural image classifier still needs classification evaluation. An LLM that calls tools needs evaluation of those actions, regardless of how its weights were trained. Using an agent loop doesn’t, by itself, mean you’re training an RL policy.

What you evaluate A possible unit What success might mean
A classifier One request and its predicted disposition The system takes the permitted, useful action
A regression model One forecast for a defined horizon The error is acceptable in the original units
A retrieval system One query and a ranked result list Useful evidence appears within the available context budget
A generated answer The response, its claims and supplied evidence The answer addresses the request and supports its claims
A policy or agent An episode or completed workflow It reaches the intended outcome within the constraints

These are starting points. A production evaluation also needs the people, languages, document types or environments the system will encounter. A score on an easy slice can conceal poor performance where the product is needed most.

Decide what a new example means

Before choosing a metric, decide what the test is meant to represent. New customers, future observations and paraphrases of familiar questions make different demands on a model.

Keep related examples together when splitting data. For time-dependent predictions, use a split that respects what would have been available at prediction time. Fit preprocessing on training data, and use validation data for choosing models and thresholds. Repeatedly adapting to the final test set turns it into another source of development feedback. The scikit-learn cross-validation guide explains the corresponding group and time-aware splits.

For LLM applications, development choices also include prompts, retrieved examples, tool descriptions and judge rubrics. Keeping the weights fixed doesn’t make those choices independent of your test data. Version the whole evaluated configuration, including the corpus snapshot and the expected outputs.

Classification: give the errors names

For a chosen class, a true positive is a correct prediction of that class. A false positive predicts it incorrectly, and a false negative misses it. With those definitions:

  • Precision = TP / (TP + FP)
  • Recall = TP / (TP + FN)
  • F1 = 2TP / (2TP + FP + FN)

Precision describes the predictions the system made. Recall describes the cases it should have found. F1 combines them. Macro-F1 is the unweighted mean of the class F1 scores, so each class contributes equally. Always document the zero-denominator convention. These are the standard classification metric definitions.

Here’s the answer class from my recorded calibration run:

Count Value Interpretation
TP 92 Answerable requests that received an answer
FP 0 Requests of another expected disposition that received an answer
FN 8 Answerable requests on which the system abstained

That gives answer precision of 1.00, recall of 0.92 and F1 of approximately 0.9583. The word answer names a disposition here. A request being suitable for an answer doesn’t establish that every sentence the model writes is correct.

The costs of the errors also matter when comparing thresholds. Blocking a useful request and releasing a prohibited one needn’t have equal consequences. Macro-F1 gives classes equal weight, but it doesn’t encode those consequences. For rare target classes, examine precision and recall at the operating threshold and report the class prevalence. A threshold chosen on a balanced benchmark may behave differently in ordinary traffic.

Regression and ranking need their own units

For a numerical prediction, MAE is the mean absolute error and RMSE is the square root of the mean squared error. Both retain the target’s units, while RMSE gives larger errors more influence. For example, hypothetical forecast errors of 1, 1 and 10 units give MAE = 4 and RMSE ≈ 5.83. The second score makes the large miss more visible.

Neither score tells you whether the model systematically overpredicts, so inspect signed errors and relevant slices too. The metric guide covers these different summaries. Which one matters depends on what a missed unit costs in the application.

For a ranked list, position matters. Recall@k asks how much labeled relevant material appears in the first k results. nDCG@k can also account for graded relevance and rank, relative to an ideal ordering under the chosen gain convention. The nDCG definition makes those assumptions explicit. If the generator only receives the first few passages, finding the right document much farther down the list won’t help that answer.

Deep learning adds questions about the evaluated model

The task still determines the metric. For image segmentation, for example, overlap between predicted and reference regions matters. For a neural classifier, the confusion matrix remains useful. Adding layers doesn’t change what a false negative means.

The evaluation setup does need care. Record the checkpoint and preprocessing. In PyTorch, model.eval() changes the behavior of modules such as dropout and batch normalization. Disabling gradient tracking is a separate operation. The Module documentation describes that distinction. An accidental change in inference behavior can make a model comparison difficult to interpret.

In an image task, also test the changes the product is expected to encounter, such as different lighting or image quality. Keep that stress evaluation identifiable, and keep derived views of one source image together when splitting data.

Probability calibration is another separate question. Among predictions assigned a confidence near 0.8, does the predicted class turn out to be correct about 80% of the time? Classification accuracy alone can’t answer that. Guo and colleagues’ calibration study illustrates why neural network confidence needs its own evaluation. An LLM writing “confidence: 0.8” also needs empirical validation before that number can serve as a probability estimate. Here calibration concerns probabilities. My project’s calibration set is development data for choosing pipeline settings, which is a different use of the word.

LLM outputs need more than one kind of check

Some parts of an answer are straightforward to test: whether JSON parses, a required field exists, a calculation matches a known result, or generated code passes relevant tests. Those checks are useful because their meaning is specific. Valid JSON establishes format compliance. It doesn’t establish the truth of the values inside it.

For open-ended text, I would define the criterion before selecting a judge. An explanation can be factually supported but incomplete, or complete but unnecessarily difficult to follow. Comparing it with one reference sentence through embedding similarity won’t resolve those differences. Similarity measures closeness in a representation, while claim support asks whether the evidence justifies what the answer says.

Breaking a response into smaller claims can make factual review more informative. FActScore is one research example of measuring support for atomic facts against a knowledge source. The decomposition and support judgments still need checking, and supported claims don’t by themselves establish that the response answered the whole question.

An LLM judge can apply a rubric at scale, but it is another model to evaluate. Use reviewed examples that were not used to tune its rubric, examine disagreements by criterion, and check whether presentation changes the judgment. The MT-Bench judge study investigated biases including answer position and verbosity. For pairwise comparisons, changing the presentation order is a useful diagnostic.

My project makes that distinction concrete. Its judge matched the complete five-label reference on 159 of 200 validation outputs, or 79.5%. I produced the reference through two recommendation-assisted passes and reconciliation. That measures agreement with my reference, rather than agreement among independent reviewers. The separate 400-case system holdout remains unopened. The project evaluation PDF explains the denominators and the rubric in more detail.

RL and agents: inspect the sequence and the outcome

In RL, an action affects what happens next. For a finite episode with T steps, one possible objective is discounted return:

G = ∑t=0 T−1 γt rt+1

Here r is reward and γ is the discount factor, between 0 and 1 for this finite-episode expression. The policy’s objective commonly concerns expected return over trajectories. The reward definition, horizon and environment determine what that number means.

Imagine a scheduler rewarded only for throughput. It could earn a good return while delaying a less profitable class of work. I would measure waiting times for that class separately, alongside return. A reward is a chosen objective, and you still need to test whether optimizing it produces the intended behavior.

Evaluate fixed policies in separate test environments and state whether action selection is deterministic or stochastic. Repeated evaluation episodes measure variability for a policy. Independent training runs also capture variability in what the learning algorithm produces. The Stable Baselines3 guidance discusses these practical evaluation settings.

For comparisons across tasks and training runs, report uncertainty alongside aggregate scores. Agarwal and colleagues’ deep RL evaluation study shows how point estimates can support different conclusions from a fuller statistical analysis. Its interval estimates and performance profiles are useful references. A thousand episodes from one trained policy don’t substitute for several independent training runs.

Logged data adds another difficulty: it records actions taken by the behavior policy. Evaluating a different policy from those logs requires an off-policy evaluation method and defensible assumptions about coverage and distribution shift. It isn’t ordinary accuracy on a table of previously chosen actions. See this off-policy evaluation study for a discussion of support in large action spaces.

An LLM agent may also produce a sequence of actions without learning a policy during the run. For a hypothetical booking task, I would check the actual reservation state, the allowed actions and the transcript. “I’ve booked it” is only a claim in the final response. Anthropic’s agent evaluation guide separates tasks, repeated trials, graders and resulting outcomes in a way that helps make this distinction concrete.

What I would put beside a result

A useful evaluation record lets another engineer work out what produced the score. Include the evaluated configuration, data split and population, the unit being scored, the reference or rubric, and the aggregation method. Add uncertainty at the appropriate sampling level, important error slices and the latency or cost budget.

Compare candidate and baseline on the same held-out cases where possible. Keep repeated outputs from the same case together when estimating uncertainty rather than counting them as independent users. Include failures and timeouts in the report, and state whether latency includes them. Otherwise a change can look faster simply because its difficult requests disappeared from the denominator.

For my RAG application, I can already point to the saved dispositions and judge labels. The next question for any proposed change is concrete: which cases would it improve, what new errors could it introduce, and which untouched examples would let me check that? That is the comparison I want to make before treating a higher score as a better system.