What you will learn
  • Separate demonstrations from general claims.
  • Identify missing information.
  • Recognize distribution changes.
  • Design representative tests.
  • Choose suitable verification.

Before you begin

Know the difference between a model and its application.

Capable at what, under which conditions?

A model may summarize clear prose well and fail on a table with footnotes. A classifier may perform well in good light and struggle at night. Capability belongs to a task, conditions, and measurement. This AI is accurate is incomplete until those are specified.

Useful tasks include suggesting alternatives, transforming text, finding patterns, drafting code, and extracting information. Each has conditions. Code may look right but fail to run. An extraction may omit an exception. A creative suggestion can be useful without being a factual claim.

For a note summarizer, require every deadline, no invented dates, and a maximum length. Check those properties directly. An overall impression of good writing can hide the one mistake that makes the result unusable.

Confidence cannot supply missing facts

If a specification never identifies the robot's battery, the model should say so. It might guess a common battery from similar projects. That can be a suggestion when labeled, but it is not a verified fact about this robot.

Hallucination is unsupported or incorrect generated content. No prompt removes it universally. Asking for sources helps only when those sources exist and support the exact claim. An answer may also be outdated; current software behavior requires current evidence.

Check primary documentation for changing commands and specifications. A tool-enabled answer should reveal what was found, including version and scope where relevant. The claim that a search occurred is weaker than its inspectable result.

Notice changed conditions

A classroom image model may encounter a dark workshop. This is distribution shift: a change between development conditions and later use. Performance can fall even while software runs normally and returns confident predictions.

For plant classification, changes include angle, age, background, disease severity, and unfamiliar species. Test variation systematically. Five nearly identical photos provide less varied evidence than examples covering different meaningful conditions.

An unfamiliar input may warrant an uncertain result or human review. The largest classification score is not automatically a calibrated probability of correctness. Calibration concerns agreement between predicted probabilities and observed frequencies, and requires evaluation.

Work through three summary tests

Write three meeting notes. A contains an ordinary deadline. B says the meeting moved from Tuesday to Thursday. C says the room remains undecided. Before running the tool, record the required facts for each.

Score whether each response preserves the deadline, handles the correction, and avoids inventing a room. Record omissions and added claims. Room 12 for C is a failure even if the prose is excellent. Preserve the source beside the response.

These tests teach evaluation; they are not a production benchmark. Expand them for real use and keep fresh cases unseen during prompt development. Otherwise you may tune for favorite examples rather than the underlying task.

Match checking to consequences

A fictional name may need only your preference. A math explanation needs recomputation. Code needs relevant execution tests. Physical control needs wiring inspection and independent limits. Choose checks based on how the output will be used.

NIST treats trustworthiness as involving multiple characteristics, rather than only average accuracy. Evaluate systems in context, including reliability and accountability. AI Risk Management Framework.

Use AI for work whose results you can evaluate, retain original evidence, and make uncertainty visible. The objective is a useful outcome, not declaring a model universally brilliant or useless.

Important terms

Capability
An ability demonstrated under stated conditions.
Hallucination
Unsupported or incorrect generated content.
Distribution shift
Changed conditions between development and use.
Calibration
Agreement between predicted probabilities and observed frequencies.
Benchmark
A defined set of tasks and measurements.
Abstention
Declining a confident decision when evidence is insufficient.

Mini project: Run three evidence checks

  1. Write the three meeting notes and correct facts.
  2. Summarize each using the same instruction.
  3. Mark facts preserved, omitted, or changed.
  4. Record unsupported additions.
  5. Finish with one claim your test supports and one it cannot support.

Common mistakes and debugging

  • Generalizing from one demonstration: state tested conditions.
  • Trusting confidence without evaluation: assess calibration when needed.
  • Developing against every test: preserve unseen cases.
  • Treating guesses as evidence: label unknowns and assumptions.

Independent challenge

Add two people sharing a surname and check whether the summary attributes actions correctly.

Check your understanding: 10 questions

  1. Why is accurate AI an incomplete claim?

  2. What if a source omits battery type?

  3. What is distribution shift?

  4. Does the largest score prove correctness?

  5. Why keep original notes?

  6. In your own words, what does “Capability” mean?

  7. In your own words, what does “Hallucination” mean?

  8. In your own words, what does “Distribution shift” mean?

  9. In your own words, what does “Calibration” mean?

  10. In your own words, what does “Benchmark” mean?

Quiz answers

Reveal all 10 answers after your attempt
  1. It omits task, conditions, data, and measurement.
  2. State it is unknown or label a proposed choice as a suggestion.
  3. A change between development conditions and later use.
  4. No. Scores require evaluation and may not be calibrated.
  5. They provide independent evidence for checking summaries.
  6. An ability demonstrated under stated conditions.
  7. Unsupported or incorrect generated content.
  8. Changed conditions between development and use.
  9. Agreement between predicted probabilities and observed frequencies.
  10. A defined set of tasks and measurements.

Summary

Evaluate specific tasks with representative examples. Missing information, changed conditions, and unsupported claims remain possible despite confidence and fluency.

Continue learning

AI09 organizes AI terminology into clear, separate dimensions.

Choose a connected learning path

Sources and further reading

Prepared 2026-09-18. Draft — source checked; classroom exercise does not establish product reliability

GO DEEPER

Extra reading & source documents

Optional reading alongside the lessons. These sources do not add to your course lesson count.

Explore the AI model guides →