The AI boom arrived with a vocabulary nobody handed out. One month the news was about chatbots. The next it was tokens, agents and context windows, used by people who assumed everyone already knew. Product pages are worse. They fit three of these terms into a single sentence and move on.
Most of the words are simpler than they sound, and they get simpler again once you sort them by what they describe. So this glossary is grouped by job, not alphabet: the basic ideas, how models learn, how you use them, how people build with them, what goes wrong, and the slang that grew up around it all.
The basics
These are the words everything else rests on, and the ones used most loosely. A lot of confusing AI coverage comes from treating them as synonyms when they are really nested. Machine learning is one way of doing AI. A large language model is one kind of machine learning model. The chatbot on your screen is a product wrapped around that model, with an interface and some rules added.
| Term | Plain English | Why it matters |
|---|---|---|
| AI | Artificial intelligence. Computer systems that do tasks we used to assume needed a human mind, such as recognising faces, translating or writing. | It is an umbrella term. A spam filter and a chatbot both qualify, so “AI-powered” on a label could mean almost anything. |
| Machine learning | Software that learns patterns from examples instead of following rules a person wrote out by hand. | Most of what gets called AI today is machine learning underneath. |
| Model | The trained system itself: the thing that takes an input and produces an output. | The same model can sit behind many different products, so two apps with different names may be running the same engine. |
| LLM | Large language model. A model trained on vast amounts of text to predict what comes next, which turns out to be enough to write, summarise and hold a conversation. | It is the technology behind most of the chatbots people mean when they say “AI” now. |
| Generative AI | AI that produces new material, such as text, images, audio, video or code, rather than sorting or scoring what already exists. | It is the branch behind the current boom, and the one raising the loudest questions about copyright and work. |
| Training data | The examples a model learns from. For language models that means enormous amounts of text, much of it gathered from the public web. | A model can only reflect what it was shown, which is why arguments about bias and copyright usually start here. |
| Parameters | The internal numbers a model adjusts as it learns. Together they hold what it has learned, in a form no person can read directly. | Model size is quoted in parameters, often in the billions. More usually means more capable and more expensive to run, but not automatically better at your task. |
How models learn
A model is not written line by line like ordinary software. It is grown from examples, in stages, on a great deal of expensive hardware. This is the part of the story that makes the business pages, because it is where the money, the electricity and the arguments over whose data was used all come in.
| Term | Plain English | Why it matters |
|---|---|---|
| Training | Showing a model huge numbers of examples and nudging its parameters until its outputs improve. | It happens before you ever use the model. Your chat does not change the model in real time, though some companies use conversations to train later versions. |
| Pre-training | The first and largest stage of training, where a model soaks up general patterns from a vast pile of text, images or both. | This is where most of the cost goes. The “P” in GPT stands for pre-trained. |
| Fine-tuning | Extra training on a smaller, focused set of examples to adapt a model to a task, a subject or a tone of voice. | It is how a general model becomes a legal drafting tool or a support assistant without starting from scratch. |
| RLHF | Reinforcement learning from human feedback. People rate a model’s answers, and the model is trained towards the kind they preferred. | It is a big reason chatbots sound polite rather than like raw autocomplete. It can also nudge them towards telling people what they want to hear. |
| Benchmark | A standard test used to compare models, such as a set of maths problems, coding tasks or exam questions. | Launch posts lean on them heavily. A model can score well on a test without being better at your actual work, and test questions sometimes leak into training data. |
| Compute | The raw processing power used to train and run models. Used as a noun, as in “we need more compute”. | Access to it decides who can build the largest models. In practice it means chips, data centres and very large electricity bills. |
| GPU | Graphics processing unit. A chip first built for video game graphics that turns out to be very good at running the huge numbers of simple calculations AI needs, all at once. | These chips are scarce and costly, which is how their best-known maker became one of the most valuable companies in the world. |
How you use them
This is the vocabulary of the chat box itself. None of it needs technical knowledge, and a handful of these terms explain most of the odd things chatbots do: why they lose track of the start of a long conversation, why the same model behaves differently in different apps, and why pricing is counted in something other than words.
| Term | Plain English | Why it matters |
|---|---|---|
| Prompt | Whatever you give a model to respond to: a question, an instruction, a document, a photo. | The model works only from the prompt and what it learned in training. Vague in, vague out. |
| Prompt engineering | Writing prompts deliberately to get better results, by adding context, examples and a clear format for the answer. | Less mysterious than the name suggests. It is mostly the old skill of writing a good brief. |
| Token | The chunk of text a model actually reads and writes. Often part of a word. In English, a token averages roughly three quarters of a word. | Prices, speed limits and memory limits are all counted in tokens rather than words. |
| Context window | The amount of text a model can take into account at once, including your messages, any files and its own replies. | Once a conversation outgrows it, earlier parts may be dropped or summarised, which is one reason long chats forget how they began. |
| System prompt | Instructions a company or developer gives a model before your conversation starts, setting its role, tone and rules. | It explains why one model can act quite differently in two apps. You usually cannot see it. |
| Multimodal | Able to handle more than one kind of input or output, such as text, images, audio and video. | It is why you can photograph a menu in another language, or a boiler error code, and ask what it means. Quality varies by type. |
| Inference | Running a trained model to get an answer. Every message you send is a round of inference. | Training is paid for once per model. Inference is paid for every time anyone uses it, so running costs grow with popularity. |
Building with AI
These are the words you hear when a company says it is “adding AI” to a product. Relatively few companies train their own models from scratch. Most rent access to someone else’s and build around it: connecting it to their own documents, giving it tools to use, and fencing it in. The terms below describe that plumbing.
| Term | Plain English | Why it matters |
|---|---|---|
| API | Application programming interface. A standard way for one piece of software to request things from another. AI providers sell access to their models this way. | Plenty of “AI-powered” apps are ordinary software passing your text to another company’s model through an API. Worth knowing when you wonder where your data goes. |
| RAG | Retrieval-augmented generation. Before answering, the system searches a set of documents and hands the relevant passages to the model along with your question. | It lets a model answer from company files or recent news without retraining, and makes it easier to show sources. |
| Embeddings | A way of turning text, images or other data into lists of numbers, so that things with similar meanings end up close together. | They power search that matches meaning rather than exact words, and they sit underneath most RAG systems. |
| AI agent | An AI system that works through a task in steps, using tools such as search, a browser or code, and deciding for itself what to do next. | The word is stretched thin in marketing. Ask what it can do without you, and what happens when it gets a step wrong. |
| Open-weight model | A model whose trained parameters are published, so anyone can download it and run it on their own machines. | It lets organisations keep data in-house. It is often called “open source”, but the training data and code usually stay private, so that label is disputed. |
| Guardrails | Rules and filters around a model that stop it producing certain content or taking certain actions. | They are why a chatbot refuses some requests. Set too loose, they let harm through. Set too tight, they block harmless questions. |
When it goes wrong
AI has an unusually vivid set of names for its failures. Some describe flaws in the model itself. Others describe ways people misuse it or attack it. Telling them apart helps you judge how worried to be about a given headline, and who is actually at fault.
| Term | Plain English | Why it matters |
|---|---|---|
| Hallucination | When a model states something false with full confidence, such as an invented quote, statistic or court case. | Models generate plausible text, not checked facts. Verify anything that matters before you repeat it. |
| Bias | When a model’s outputs treat groups of people unfairly or repeat skewed assumptions, usually inherited from its training data. | It matters most when AI helps make decisions about jobs, loans or policing, where a small skew repeated at scale becomes a pattern. |
| Jailbreak | A prompt designed to talk a model out of its own rules, often through role-play or an elaborate hypothetical. | It shows that guardrails are fences, not walls. Companies patch known tricks and people find new ones. |
| Prompt injection | Hiding instructions inside content a model will read, such as a web page, email or document, so it obeys them instead of the user. | It is a serious risk for AI agents. An assistant that reads your inbox can be steered by whoever writes to you. |
| Deepfake | Realistic fake video, audio or images of a real person, made with AI. | Cloned voices already turn up in phone scams. A convincing clip is no longer proof on its own. |
| Alignment | The work of making AI systems reliably do what people intend, and not pursue goals or habits nobody asked for. | It covers everything from a chatbot that flatters too much to worries about far more capable systems in future. |
The culture and the slang
Any technology used this widely grows a slang layer, and AI’s grew fast. These words move the way we traced in our piece on how new slang spreads: a small group uses a term precisely, bigger accounts carry it out, and it loosens as it travels. A few are neutral. Several are plainly insults, which says something about the mood.
| Term | Plain English | Why it matters |
|---|---|---|
| Vibe coding | Building software by describing what you want to an AI and accepting the code it writes, often without reading it closely. | It lets non-programmers make working apps. It is also a quick way to ship bugs and security holes nobody understands. |
| AI slop | Low-effort AI-generated content made in bulk: strange images, filler articles, fake reviews, spam videos. | It clogs search results and feeds. The word “slop” carries the verdict: made cheaply, for volume, with no one checking. |
| Clanker | A mocking word for robots and AI, borrowed from Star Wars, where clone troopers use it for battle droids. | It is mostly a joke, but it gives a name to real irritation at AI turning up everywhere uninvited. |
| AGI | Artificial general intelligence. A hypothetical AI able to do most intellectual work as well as a person. | There is no agreed definition or test, so claims that it is close, or already here, are hard to check. |
| Doomer | Someone who expects advanced AI to end very badly, possibly for humanity as a whole. | Critics use it as a put-down. Some of the worries behind it are taken seriously by researchers, so the label can hide a real argument. |
| Accelerationist | Someone who wants AI developed as fast as possible and sees speed as the route to progress. The online movement version calls itself “e/acc”, short for effective accelerationism. | It is the opposite camp to the doomers. Most real debates about regulating AI happen in the wide space between the two. |
| “ChatGPT it” | Asking a chatbot instead of searching, as in “just ChatGPT it”. A brand name turned verb, the way “Google it” was. | It marks a real change in habits. A chatbot answer can sound certain without being checked, so it deserves the same doubt as any single source. |
Three terms people mix up
AI, machine learning and LLM. Picture three boxes, one inside the next. AI is the outer box: any system doing something that looks like thinking. Machine learning sits inside it and covers systems that learn from examples rather than hand-written rules. An LLM sits inside that, as one kind of machine learning model built for language. So every LLM is AI, but most AI is not an LLM. The filter catching your spam counts as AI, and it has never chatted with anyone.
Training and fine-tuning. Training builds a model’s abilities in the first place, and the main stage can take months and cost enormous sums. Fine-tuning is a short extra course afterwards, on a narrow set of examples. Training teaches a model language. Fine-tuning teaches it house style. When a company says it has “trained its own AI”, ask which one it means, because fine-tuning an existing model is far more common.
RAG and fine-tuning. Both make a general model useful for a specific job, but they fix different problems. Fine-tuning changes the model itself, so it suits teaching a style, a format or a type of task. RAG leaves the model alone and hands it the right documents at the moment you ask, so it suits facts that change or that live in your own files. A rough rule: fine-tune for how it should answer, and use RAG for what it should know.
How to keep up without reading every launch post
- Learn the concept, not the brand name. Model names change every few months and tell you little. The ideas underneath, like context windows, agents and fine-tuning, stay useful for years. It is the approach behind the full dictionary, slang and tech alike: understand what a word does, and the next one doing the same job will not throw you.
- Turn every claim into a plain question. “Agentic” means “what can it do without me?” “Multimodal” means “can it look at my photo?” “Longer context” means “can it read the whole report?”
- Treat benchmark scores as a trailer, not a review. Try the tool on one real task of your own. Ten minutes of that tells you more than a chart.
- Notice when a word stops meaning anything. Once every product is an “agent” or “AI-powered”, the label has stopped carrying information. Ask what the thing does on an ordinary Tuesday.
Questions people ask
What is the difference between AI and ChatGPT?
AI is the whole field, covering everything from spam filters to image generators. ChatGPT is one product within it: a chatbot made by OpenAI and built on large language models.
Why do AI chatbots make things up?
AI chatbots make things up because language models are built to produce likely-sounding text rather than to check facts, so confident wrong answers come from the same process as confident right ones. Connecting a model to real documents reduces the problem without removing it.
What does “tokens” mean on an AI pricing page?
On an AI pricing page, a token is a small chunk of text, often part of a word, and the unit a model reads and writes in. Providers usually quote prices per million tokens and charge for both the text you send and the text the model sends back.
Is prompt engineering a real job?
Some companies have hired prompt engineers under that title, but for most people prompt engineering is a skill rather than a job. It mostly means writing clear instructions with the right context and examples, which is useful well beyond AI.
Is an AI agent just a chatbot with a new name?
An AI agent differs from a chatbot because it carries out tasks in several steps, using tools such as a browser or your files, instead of only replying. The label is applied generously in marketing, so check what a product actually does before trusting the word.