AI concepts introduction
Introduction
AI is a significant technological shift, though its full impact is still unfolding. It already offers useful ways to save time, support ideas, and improve productivity. Learning about its potential benefits and pitfalls can help you decide where it is most valuable.
This section introduces the basics of AI and gives you a sense of how the technology has developed and where it stands today.
To set the scene, here’s a brief timeline of AI. We’ll explore several of these ideas and milestones in more detail as we go.
Two things to keep in mind as we get further into AI:
The pace of development is accelerating — not slowing down.
The AI models you use today are the worst AI models you will ever use.
Traditional Software
Software Just Follows Instructions
Before we get to AI, it helps to understand how most software was built for decades, and why that approach eventually hit a wall.
Traditional software is like a giant recipe. It has no understanding or judgement of its own, and it follows the rules a human has written, quickly and exactly. For example:
IF the password matches → THEN grant access
IF the balance is too low → THEN decline the transaction
IF the email contains "Nigerian prince" → THEN move to spam
If a rule is wrong or doesn't cover the situation, the software can't adapt. A recipe that says "add 10 tablespoons of salt" would make a human cook stop and question it. Software would add the salt.
Why This Hits a Wall
Rules work brilliantly for structured problems such as banking systems, calculators and traffic lights. But try writing rules to recognise a cat in a photo.
The Cat Problem
"Four legs and pointy ears" also describes a fox. "Whiskers" fits mice and seals. What about:
- a cat photographed from behind?
- a black cat in a dark room?
- a cartoon cat?
- a sleeping cat under a blanket?
You quickly realise you can't write rules for every possible cat. The real world is too messy, variable, and ambiguous.
Machine learning was the breakthrough - instead of programmers writing every rule by hand, computers could learn the patterns from the dates
AI Jargon Explained
The terms AI, machine learning, deep learning, and generative AI are often used interchangeably — but they describe different things. The easiest way to understand them is as a series of increasingly specialised ideas.
Artificial Intelligence (AI)
The umbrella term for computer systems that perform tasks usually associated with human intelligence, such as recognising patterns, making predictions, or understanding language. Some AI systems follow fixed rules, like traditional software, while others learn from data.
Machine Learning (ML)
A type of AI that learns patterns from data rather than relying entirely on rules written by a programmer. For example, it can learn from thousands of labelled images to recognise what appears in a new image.
Deep Learning
A type of machine learning that uses neural networks with many layers. The word deep refers to these layers. Deep learning powers technologies such as image recognition, language translation, and AlphaFold’s predictions of protein structures.
Generative AI
AI that creates new content, such as text, images, audio, video, music, or code. Most modern generative AI uses deep learning to learn patterns from existing content and generate something new.
Large Language Models (LLMs)
A major type of generative AI designed to work with language. LLMs learn patterns from enormous amounts of text and generate text one small piece at a time, predicting what is likely to come next. Some also work with images or audio.
How they fit together: AI is the broadest term. Machine learning is one type of AI, and deep learning is one type of machine learning. Generative AI describes what a system does—creates content—and LLMs are a major example.
Frontier models:
Frontier models are among the most capable AI models available at a particular time. They may be able to generate text, images, video, and code; browse the internet; and connect with other software to complete tasks. Well-known AI assistants built around frontier models include ChatGPT, Claude, Gemini, and Grok. The companies behind them are in an intense race, investing enormous sums of money and computing power to push their models further.
How ML Works
The Key Insight — Learning patterns from data
Instead of the programmer writing every rule, you give the computer thousands of examples and let it learn the statistical patterns that help it make predictions.
Think of how a child learns language — nobody teaches a toddler the rules of grammar. They hear thousands of hours of people talking and gradually pick up the patterns. Machine learning is the same idea applied to computers.
And remember the cat recognition problem? You don't write rules. You show the system ten thousand photos labelled "cat" and ten thousand labelled "not cat" and it figures out for itself what makes a cat a cat.
The Video Game Example:
In 2013, DeepMind built an AI system and pointed it at classic Atari games — Pong, Breakout, Space Invaders. The system was given the pixels on screen, the available controls, and a single instruction: get as high a score as possible.
It started playing randomly, mashing buttons. But every time it scored a point it learned "do more of that." Within hours it was beating human players at Pong. In Breakout, it learned a strategy of tunnelling the ball behind the wall of bricks, where it could bounce around and destroy them.
AlphaGo’s Move 37:
In 2016 DeepMind's AlphaGo took on Lee Sedol — one of the greatest Go players in history. Go is especially difficult for computers because each turn offers many possible moves, creating an enormous number of potential paths through a game. Winning at Go was therefore thought to be a major milestone for AI, showing that a computer could succeed at a game long thought to demand unusually deep strategic judgement.
In Game Two, AlphaGo played Move 37 — a move so unusual that commentators thought it was a mistake. No human would have played it. It broke every conventional rule of Go strategy.
It won the game. And ultimately the match, 4 games to 1. Move 37 became iconic because it wasn’t just the computer being faster — it found a new possibility in a game humans had studied for centuries.
Machine learning dates back to the 1950s, but for decades its potential was limited by the available data, computing power, and the methods used to train it.
Both the Atari system and AlphaGo used deep learning, combined with reinforcement learning — learning through trial and error, guided by rewards. Deep learning uses neural networks with many layers, which we’ll explore in the next section.
Neural Networks
What is a neural network?
Systems loosely inspired by how the brain works. Instead of one big algorithm you have layers of simple processing units connected together — a bit like neurons. Data goes in one end, passes through layers, and a result comes out the other end.
How training works
Each connection in a neural network has a weight — a number that controls how strongly one signal influences the next layer. Training means adjusting these weights so the network makes better predictions.
Show the network an example with the correct answer, such as a photo labelled “cat”.
It makes a prediction.
Measure how far the prediction is from the correct answer.
Adjust the weights slightly to reduce the error.
Repeat across many examples, gradually improving its predictions.
Think of learning darts. Your first throw misses. You adjust your next attempt based on where it landed. After thousands of throws, you become more accurate — through repeated practice and feedback, rather than someone teaching you all the physics. Neural network training follows a similar pattern of small corrections.
Learning through rewards
Learning from labelled examples as above is called supervised learning.
Another approach is reinforcement learning: a system tries actions and learns from rewards, such as scoring points or winning a game. It gradually learns which actions are more likely to produce better results.
Combined with deep neural networks, reinforcement learning helped power the Atari and AlphaGo examples discussed earlier.
Deep learning — the 2012 breakthrough
Deep learning uses neural networks with many layers. A major breakthrough came in 2012, when a model called AlexNet won the ImageNet competition by a wide margin. Its task was to recognise objects and animals in photographs across 1,000 categories — taking the cat-recognition problem we discussed earlier to a much larger scale.
Why GPUs matter
Training a deep neural network involves enormous numbers of repeated calculations. GPUs — graphics processing units, developed to handle graphics in games and other applications, are especially good at performing many calculations in parallel.
This made it practical to train much larger networks, helping turn deep learning’s potential into useful results. AlexNet was trained using two NVIDIA GPUs. NVIDIA subsequently became a major supplier of the specialised chips used to train and run modern AI models.
Training a large language model often involves a combination of supervised learning and reinforcement learning through human feedback
The black box problem
Here's a strange thing about neural networks: even the people who build them don't fully understand how they work.
A modern AI model contains billions of connection weights, all adjusted automatically during training. When the model gives you an answer, it isn't following a set of rules a human wrote. It's running a calculation across all those billions of weights in a way no human can meaningfully read or trace. You can see the input. You can see the output. But what happens in between — why the model chose those particular words — is largely opaque, even to its creators.
This is called the black box problem. The model works, but it can't easily explain itself, and we can't easily look inside.
It has real consequences:
When an AI gets something wrong, it's hard to know why — so it's hard to guarantee it won't happen again
When an AI is biased, the bias is buried in the weights rather than written into rules you can find and fix
When an AI gets something surprisingly right (like Move 37), nobody can fully explain how it got there either
A whole field of research called interpretability is trying to peek inside the box and understand what's happening. Progress is being made, but for now, modern AI remains one of the strangest tools humans have ever built: something we can create, train, and use — but not fully explain.
The language problem
Despite progress in image recognition, language remained especially difficult. Words do not have fixed meanings on their own — their meaning depends on the surrounding words, the wider conversation, and sometimes knowledge about the world. A system must deal with ambiguity, implied meaning, idioms, tone, and references that may stretch across several sentences.
Consider the sentence: "The patient told the doctor she was worried." Who does "she" refer to — the patient or the doctor? The answer may only become clear from an earlier or later sentence. Similarly, "bank" means something different in "river bank" and "bank account."
Earlier language models processed text sequentially, one word at a time from left to right. This made it difficult to connect words that were far apart, retain the important parts of a long passage, and train efficiently on large amounts of text.
The Transformer
Scaling Laws
How AI Is Trained
Where We Are Now
Frontier Models