- Describe a language model's input and output.
- Explain parameters and representations.
- Describe attention without assuming awareness.
- Separate models from chat history and tools.
- Explain why fluency does not guarantee factual recall.
Before you begin
AI03 introduces token prediction and the difference between generation and correctness.
The model is one part of the assistant
A writing assistant may display saved chats, uploaded documents, a search button, and account settings. Those features belong to an application. A large language model, or LLM, is a trained numerical model that processes language and produces language-related outputs. The application decides which information reaches it and which tools it can use.
Large refers broadly to model and training scale, not an exact universal threshold. Parameters are numerical values adjusted through training. A parameter is not a neatly labeled fact card, and a larger parameter count alone does not establish better performance on your task.
One application may give a model a search tool; another may use the same model without search. Their ability to answer a question about today's events can differ because of the surrounding system, not because one suddenly learned new facts during your conversation.
Turn text into numbers
A tokenizer converts text into token identifiers. The model maps these identifiers to numerical representations and transforms those representations through learned layers. It ultimately produces scores used to select output tokens, which are converted back to text.
Consider bank in a sentence about a river and in one about savings. Surrounding words help determine which relationships are relevant. Useful representations are affected by context rather than being only fixed dictionary definitions.
Numerical representation provides a format in which a model can learn and apply relationships. It does not prove the system experiences meaning as a person does. We can evaluate behavior directly without making claims about consciousness.
Text → tokens → representations → learned transformations → token scores → textWhat attention contributes
Many prominent LLMs use transformer architectures. Attention computes how representations at different positions should influence one another. Information from relevant parts of the input contributes to a new representation. Attention here is a mathematical operation, not conscious focus.
In The toolbox would not close because the hammer was too long, relationships among toolbox, hammer, and long help interpretation. A model can combine many relationships across layers. This example illustrates the purpose; it does not reveal exactly what a particular attention head computes.
The original transformer paper introduced an architecture based on attention for sequence tasks. Modern systems include variants, so its encoder-and-decoder design is not the structure of every LLM. Attention Is All You Need.
Training and context play different roles
Training adjusts parameters using examples and an objective. Common objectives include predicting subsequent or missing tokens. Additional training can shape instruction following. During ordinary inference, the model applies existing parameters to the current input rather than retraining from scratch.
Supplying a paragraph about an imaginary robot can let the model answer questions about it. This does not mean those details became a permanent part of the model. A product may separately save preferences or conversation history; saved application data differs from changed model weights.
An LLM may reproduce correct facts because training created useful relationships, yet generate a plausible date or source that is wrong. It has no automatic guarantee of retrieving a source for each statement. Search, retrieval, calculations, and tests help only when their results are connected correctly to the answer.
Inspect a grounded answer
Supply this fictional source: Rover A has two wheels and measures distance. Rover B has four wheels and measures light. Ask: Using only this source, which rover measures light, and how many wheels does it have? Quote the supporting sentence. The answer is Rover B, four wheels, supported by sentence two.
Then ask which rover has the larger battery. The source does not say. A good answer identifies the missing information. This repeatable exercise separates supported extraction from confident invention without requiring access to a model's internals.
Important terms
- LLM
- A large trained language model.
- Parameter
- A numerical value learned during training.
- Tokenizer
- A process mapping text to model units.
- Representation
- Numerical information encoding and transforming input.
- Attention
- A computation combining information from relevant positions.
- Context
- Information supplied for the current model operation.
Mini project: Test evidence boundaries
- Use the two fictional rover sentences.
- Ask two questions they answer and two they do not.
- Mark responses supported, contradicted, or absent from the source.
- Rewrite unsupported answers to identify missing information. Finish with a checked answer set.
Common mistakes and debugging
- Calling the whole application a model: identify its separate features.
- Assuming attention proves awareness: it is a mathematical operation.
- Believing conversation automatically retrains the model: distinguish context, saved memory, and training.
- Assuming fluent citations are real: open the cited source.
Independent challenge
Add a third rover with an overlapping property. Create a question requiring comparison across sentences and check it manually.
Check your understanding: 10 questions
What are parameters?
What does tokenization do?
Does every LLM have the original transformer structure?
Why can one model behave differently in two apps?
What should the rover example answer about batteries?
In your own words, what does “LLM” mean?
In your own words, what does “Parameter” mean?
In your own words, what does “Tokenizer” mean?
In your own words, what does “Representation” mean?
In your own words, what does “Attention” mean?
Quiz answers
Reveal all 10 answers after your attempt
- Numerical values adjusted during training.
- It converts text into units and identifiers a model can process.
- No. Variants and other architectures exist.
- Apps can supply different instructions, context, tools, and settings.
- The supplied source contains no battery-size information.
- A large trained language model.
- A numerical value learned during training.
- A process mapping text to model units.
- Numerical information encoding and transforming input.
- A computation combining information from relevant positions.
Summary
LLMs transform numerical language representations using learned parameters. Applications supply context and tools; evidence checks determine whether claims are supported.
Continue learning
AI05 explains token and context limits when a conversation becomes large.
- AI Tokens, Context Windows, and Prompts Explained
- Natural Language Processing: How AI Works With Text
- How to Use AI for Research Without Losing Accuracy
- How to Verify AI Answers and Detect Mistakes
Sources and further reading
Prepared 2026-09-18. Draft — transformer mechanism source checked