Before you begin

No previous experience required unless you choose a coding exercise. Use an account you are allowed to access; features vary by product and region.

What you will learn

  • Separate a provider announcement from independent evidence.
  • Turn a vague model comparison into a small test.
  • Build a four-part scoring rubric before seeing the output.
  • Choose a model for a task instead of choosing by name alone.

What Anthropic announced

Anthropic introduced Claude Sonnet 5.5 on September 28, 2026. The company describes it as a faster, lower-cost complement to Claude Opus 5.5 for well-scoped everyday work such as bug fixes, documents, slides, spreadsheets and design tasks. Anthropic also reports benchmark results, including 70.6% on Terminal-Bench 4.0. These are provider claims, not results independently reproduced for this article.

A model is the system that produces the response. An app may add file access, integrations, saved projects or other tools around that model. A model announcement therefore does not guarantee that every account, plan or app has identical access, speed or pricing. Check the exact product you use before paying or uploading private material.

Why a narrow task is a fairer test

The question 'Which AI is best?' is too broad. One model may be stronger at a difficult coding task, while another may be faster or less expensive for a short rewrite. Your useful question is narrower: which option meets the requirements of this task, with acceptable cost and delay?

For a beginner-safe comparison, use a paragraph you wrote yourself and a result you can inspect. Avoid a private school record, confidential workplace document or copyrighted passage you do not have permission to submit. This lesson uses a made-up paragraph, needs no files and does not require a paid account.

Before testing, write the success rules. That prevents you from changing the rules after seeing an impressive answer. A rubric is a short scoring guide. Our rubric gives one point each for factual preservation, required length, clear language and no invented claims.

Run the same four-point check

Use this source paragraph: 'The school garden has six soil sensors. They measure moisture every hour. Students use the readings to decide when to water. The project does not water plants automatically.' Ask for a rewrite of 35 to 45 words for a beginner audience.

Score the response before deciding whether you like its style. Did it keep six sensors? Did it preserve the hourly schedule? Did it explain that students decide when to water? Did it avoid claiming automatic watering? Then check the requested length. A polished sentence that invents automation fails the factual part of the task.

If you compare two models, use the same prompt, source text and settings as closely as possible. Record the date, model label and result. One example is not a universal ranking; it is evidence about one task. Repeat with two more representative examples before changing a real workflow.

Add cost and speed only after quality

Anthropic says Sonnet 5.5 is intended to offer a different speed-and-cost balance from Opus 5.5. That does not mean the less expensive option is always better, or that the larger option is always worth more. First decide the minimum quality you need. Then compare time and price among the outputs that pass.

For a casual draft, a small speed difference may not matter. For hundreds of repeated tasks, it can matter greatly. Likewise, a model that needs several corrections may cost more in human review even if its listed token price is lower. Measure the complete workflow: input preparation, generation, checking and correction.

Do not reuse the provider's benchmark number as your personal expected score. Benchmarks use defined datasets, scoring rules and test conditions. Your short rewrite test answers a different question: can this model follow your exact brief reliably enough for your use?

Terms and next step

Benchmark: a standardized test used to compare systems. Rubric: written criteria for scoring work. Provider claim: a statement published by the company offering the product. Well-scoped task: work with a clear input, output and boundary. Human review: a person checking whether the result is correct and suitable.

Next, read 'Claude Opus 5.5: Understand Cost per Checked Result' to calculate the cost of usable work rather than the cost of a single response. Then use 'How to Verify AI Answers and Detect Mistakes' for a broader verification routine.

Try this prompt

This is a suggested exercise, not a tested guarantee of any model’s output.

Rewrite the source paragraph for a beginner in 35 to 45 words. Preserve every factual detail, especially the number of sensors, the measurement schedule and who decides when to water. Do not add automatic actions or benefits not stated in the source. After the rewrite, list the word count and a four-item factual checklist. Source: The school garden has six soil sensors. They measure moisture every hour. Students use the readings to decide when to water. The project does not water plants automatically.

Mini project & challenge

  1. Copy the source paragraph and four-point rubric onto paper before using any AI tool.
  2. Run the prompt in one permitted model and save the unedited response with the model label and date.
  3. Score factual preservation, required length, clear language and absence of invented claims.
  4. If available, repeat once with another model using the same input; do not change the prompt between runs.
  5. Choose only among responses that meet all required facts, then note which needed less correction.
  6. Challenge: create a second 40-word source about a robot sensor and write a task-specific rubric before testing it.

Common mistakes

  • Asking which model is best without naming a task: define the result you need.
  • Changing the prompt between models: keep the comparison conditions consistent.
  • Rewarding fluent invented details: score factual preservation first.
  • Treating one test as a universal ranking: repeat representative tasks and document limits.

Check your understanding

  1. When did Anthropic announce Claude Sonnet 5.5?
  2. How does Anthropic position Sonnet 5.5 relative to Opus 5.5?
  3. Are Anthropic's benchmark results independently reproduced in this article?
  4. What is a rubric?
  5. How many soil sensors are in the exercise?
  6. Who decides when to water in the source paragraph?
  7. Why is an automatic-watering claim a failure?
  8. Why use the same prompt for two models?
  9. What should you compare after outputs meet the quality bar?
  10. Does one passing example prove a model is best for every task?
Show answers
  1. September 28, 2026.
  2. As a faster, lower-cost complement for well-scoped everyday work.
  3. No; they are clearly attributed provider claims.
  4. A written set of criteria used to score work.
  5. Six.
  6. The students.
  7. The source explicitly says the project does not water automatically.
  8. It makes the comparison more consistent and easier to interpret.
  9. Speed, price and total human review or correction effort.
  10. No; it is evidence only about that tested task and setup.

In short

Claude Sonnet 5.5's announcement gives you a reason to evaluate, not a ready-made winner. Define a narrow task, score facts before style, compare like with like and measure the full cost of producing checked work.

Continue the Understand AI course →

Sources & review notes

  1. Anthropic: Introducing Claude Sonnet 5.5

    Announcement dated September 28, 2026; checked October 1. Product positioning and benchmark numbers are provider claims and were not independently replicated for this lesson.

Product information was checked on 2026-10-01. This is a selected beginner guide, not an exhaustive archive of every announcement. Recheck access, pricing and compatibility before relying on a current product detail. Supplied screenshots remain unreplicated claims.