← Back to Knowledge Base

KNOWLEDGE / 12

LLMOps, Evaluation and Observability

LLM quality evaluation, experiments, monitoring, and control of AI system behaviour.

Questions and practice

Open a question to see the answer, examples, and exercises.

How does evaluation differ from a single demo?Junior

Answer

Evaluation assesses a system on representative cases with defined rules; a demo shows one scenario.

Examples

  • One successful answer does not reveal recall, stability, or edge-case behaviour.

Practice exercises

  1. Create a small set of typical, negative, and boundary tasks for one LLM workflow.
What makes an evaluation dataset useful?Junior

Answer

A dataset should represent real tasks, risky edge cases, negative examples, and expected evidence. Labels need definitions and a review process. Random attractive prompts create a demo rather than a baseline; also protect the dataset from leakage into prompts or the retrieval index.

Examples

  • A requirement-review set includes complete, vague, contradictory, and intentionally unresolvable requirements.

Practice exercises

  1. Create ten evaluation cases with a risk category, expected behaviour, and scoring criteria.
Which LLM-system checks can be deterministic?Junior

Answer

Schema validity, required fields, citation existence, permission filters, tool arguments, latency, and token limits can be checked with ordinary code. Semantic correctness often needs a reference, rubric, human review, or calibrated judge. Use a cheap, unambiguous oracle before delegating everything to another model.

Examples

  • A validator rejects an answer containing a citation ID that is absent from retrieved documents.

Practice exercises

  1. Classify ten quality criteria as deterministic, model-based, or human-reviewed checks.
Why is LLM-as-judge not an absolute oracle?Middle

Answer

The judge can also be wrong and biased; it needs calibration against human ratings and deterministic checks.

Examples

  • A judge may consistently prefer longer answers regardless of accuracy.

Practice exercises

  1. Compare judge scores with human labels and identify systematic disagreements.
How do precision and recall apply to retrieval evaluation?Middle

Answer

Recall measures how much of the required material the system found; precision measures how much of the retrieved material is actually relevant. Increasing top-k often raises recall but adds noise and cost. Calculate metrics on labelled queries and add rank-aware measures when result position matters.

Examples

  • Search finds all three required documents among twenty results: recall is high, but precision is low.

Practice exercises

  1. Label relevance for five queries and calculate precision@k and recall@k for two retrievers.
What must be versioned to reproduce an LLM-system result?Middle

Answer

Record the model and parameters, system and developer prompts, tool schemas, application code, retrieval corpus, chunker, embedding model, index, reranker, and evaluation dataset. A provider model alias may change without a code commit, so run metadata should keep the available concrete identifier and timestamp.

Examples

  • The same prompt changes after an index update; the trace reveals different retrieved chunks and corpus versions.

Practice exercises

  1. Create a run manifest sufficient to compare two releases and explain a quality-metric change.
Why use traces alongside logs?Senior

Answer

A trace connects the steps of one execution and shows timing and dependencies; logs record individual events.

Examples

  • A trace can reveal that latency came from retrieval rather than model inference.

Practice exercises

  1. For one request, define a correlation ID, spans, key events, and fields that must not be logged.
How can a regression gate be built for a nondeterministic LLM system?Senior

Answer

Use repeated runs, a statistically meaningful baseline, risk segmentation, and separate hard constraints. A critical safety or authorisation failure may block immediately, while aggregate quality needs tolerance and a confidence interval. Calibrate judge results against human labels and monitor drift.

Examples

  • Any unauthorised tool call blocks release, while answer-quality score is compared with baseline within an agreed margin.

Practice exercises

  1. Design a gate with hard failures, aggregate metrics, sample size, threshold owner, and manual override process.
How can quality, latency, cost, and reliability be optimised together?Senior

Answer

This is a multi-objective trade-off: a larger model or longer context may improve quality while increasing latency and cost. Measure end-to-end distributions, token use, retries, failure rate, and quality by segment. Routing, caching, and fallback help only when they preserve freshness, privacy, and correctness.

Examples

  • Simple classification requests use a smaller model, while difficult cases route to a stronger one after a confidence check.

Practice exercises

  1. Build a scorecard for two models using quality, p95 latency, cost per task, and failure rate, then define a routing rule.