Introduction
AI is a significant technological shift, though its full impact is still unfolding. It already offers useful ways to save time, support ideas, and improve productivity. Learning about its potential benefits and pitfalls can help you decide where it is most valuable.
To understand what AI actually is — and why it suddenly got so powerful — it helps to start with how computers have always worked and why that approach eventually hit a wall.
The pace of development is accelerating — not slowing down.
The AI models you use today are the worst AI models you will ever use.
Traditional Software
Before we get to AI, it helps to understand how most software was built for decades — and why that approach eventually hit a wall.
Software Just Follows Instructions
Traditional software has no understanding or judgement of its own. It does exactly what it's told — extremely fast and extremely precisely.
Traditional software is like a giant recipe. The computer follows the rules a human has written. But if those rules are wrong, incomplete, or don't cover the situation, the software cannot adapt.
For example:
- IF the password matches → THEN grant access
- IF the balance is too low → THEN decline the transaction
- IF the email contains "Nigerian prince" → THEN move to spam
The software will follow these instructions exactly — even if the rule contains a mistake.
"Add 10 tablespoons of salt."
A human cook would probably realise something is wrong. Traditional software would simply follow the instruction exactly, because it has no common sense or understanding of the bigger picture.
Why This Hits a Wall
Rules work brilliantly for structured problems — banking systems, calculators, traffic lights. But try writing rules to recognise a cat in a photo.
"Four legs and pointy ears" also describes a fox. "Whiskers" fits mice and seals. What about:
- a cat photographed from behind?
- a black cat in a dark room?
- a cartoon cat?
- a sleeping cat under a blanket?
You quickly realise you can't write rules for every possible cat. The real world is too messy, variable, and ambiguous.
This was the wall AI researchers hit for decades. Machine learning was the breakthrough — instead of programmers writing every rule by hand, computers could learn the patterns from data themselves.
AI Jargon Explained
The terms AI, machine learning, deep learning, and generative AI are often used interchangeably — but they describe different things. The easiest way to understand them is as a series of increasingly specialised ideas.
Artificial Intelligence (AI)
The broad umbrella term for computer systems designed to perform tasks normally associated with human intelligence, such as recognising patterns, making predictions, understanding language, or making judgements. Some AI follows fixed rules; other AI learns from data.
Machine Learning (ML)
A subset of AI that learns patterns from examples rather than relying entirely on rules written by a programmer. Show it thousands of labelled examples and it gradually learns how to make predictions about new ones.
Deep Learning
A type of machine learning that uses neural networks with many stacked layers — which is what the word deep refers to. It powers technologies such as image recognition, language translation, and AlphaFold's protein-structure predictions.
Generative AI
Most modern generative AI uses deep learning to create new content — including text, images, audio, video, music, and code — rather than only analysing or classifying existing information.
One Branch: Large Language Models (LLMs)
An LLM is the language branch of generative AI. It learns patterns from enormous amounts of text and generates a response by repeatedly predicting what should come next. Other generative models specialise in images, audio, music, or video.
Frontier Models
Frontier models are among the most capable AI models available at a particular time. They may be able to generate text, images, video, and code; browse the internet; and connect with other software to complete tasks. Well-known AI assistants built around frontier models include ChatGPT, Claude, Gemini, and Grok.
- Rule-based software filters existing recipes according to allergies, ingredients, or cooking time.
- Machine learning recommends meals based on choices made by you and similar users.
- Deep learning recognises ingredients from a photograph of your fridge or kitchen cupboard.
- Generative AI can create entirely new material in several forms:
- An LLM writes the recipe and answers questions about substitutions.
- An image model creates a picture of the finished dish.
- An audio model narrates the cooking instructions.
- A video model demonstrates the preparation steps.
How ML Works
The Key Insight — Let the Computer Figure It Out
Instead of the programmer writing every rule, you give the computer thousands of examples and let it learn the statistical patterns that help it make predictions.
Think of how a child learns language — nobody teaches a toddler the rules of grammar. They hear thousands of hours of people talking and gradually pick up the patterns. Machine learning is the same idea applied to computers.
And remember the cat recognition problem? You don't write rules at all. You show the system ten thousand photos labelled "cat" and ten thousand labelled "not cat" and it figures out for itself what makes a cat a cat.
The Video Game Example
In 2013, DeepMind built an AI system and pointed it at classic Atari games — Pong, Breakout, Space Invaders. The system was given the pixels on screen, the available controls, and a single instruction: get as high a score as possible.
It started playing randomly, mashing buttons. But every time it scored a point it learned "do more of that." Within hours it was beating human players at Pong. In Breakout, it learned a strategy of tunnelling the ball behind the wall of bricks, where it could bounce around and destroy them.
Move 37 — The Moment AI Shocked the World
In 2016 DeepMind's AlphaGo took on Lee Sedol — one of the greatest Go players in history. Go is especially difficult for computers because each turn offers many possible moves, creating an enormous number of potential paths through a game.
In Game Two, AlphaGo played Move 37 — a move so unusual that commentators thought it was a mistake. No human would have played it. It broke every conventional rule of Go strategy.
It won the game. And ultimately the match, 4 games to 1. Move 37 became iconic because it wasn’t just the computer being faster — it found a new possibility in a game humans had studied for centuries.
Why It Took So Long
The concept has been around since the 1950s but for decades we didn't have enough data to train on or enough computing power to process it. The field went through several "AI winters" — periods of hype followed by disappointment when the technology couldn't deliver.
Neural Networks
What is a Neural Network?
Systems loosely inspired by how the brain works. Instead of one big algorithm you have layers of simple processing units connected together — a bit like neurons. Data goes in one end, passes through layers, and a result comes out the other end.
How Training Works
Each connection between layers has a weight — a number saying how important that signal is. Training = adjusting millions of these weights until the network gets accurate.
- Show the network an example (eg a photo of a cat)
- It gets the answer wrong
- Adjust the weights slightly
- Repeat millions of times
- Eventually it becomes eerily accurate
Conceptually similar to learning darts. Your first throw misses. Your brain adjusts. After thousands of throws you're hitting the bullseye — not because someone told you the physics, but because you've made tiny corrections thousands of times.
Deep Learning (2012 onwards)
When you stack many layers of networks on top of each other = deep learning. Things got serious around 2012 when a deep learning model smashed existing records in image recognition.
GPUs — chips originally designed for gaming — turned out to be perfect for training because they can do millions of small calculations simultaneously.
The Black Box Problem
Here's a strange thing about neural networks: even the people who build them don't fully understand how they work.
A modern AI model contains billions of connection weights, all adjusted automatically during training. When the model gives you an answer, it isn't following a set of rules a human wrote. It's running a calculation across all those billions of weights in a way no human can meaningfully read or trace. You can see the input. You can see the output. But what happens in between — why the model chose those particular words — is largely opaque, even to its creators.
This is called the black box problem. The model works, but it can't easily explain itself, and we can't easily look inside.
It has real consequences:
- When an AI gets something wrong, it's hard to know why — so it's hard to guarantee it won't happen again
- When an AI is biased, the bias is buried in the weights rather than written into rules you can find and fix
- When an AI gets something surprisingly right (like Move 37), nobody can fully explain how it got there either
A whole field of research called interpretability is trying to peek inside the box and understand what's happening. Progress is being made, but for now, modern AI remains one of the strangest tools humans have ever built: something we can create, train, and use — but not fully explain.
The Language Problem
Despite progress in image recognition, language remained especially difficult. Words do not have fixed meanings on their own — their meaning depends on the surrounding words, the wider conversation, and sometimes knowledge about the world. A system must deal with ambiguity, implied meaning, idioms, tone, and references that may stretch across several sentences.
Consider the sentence: "The patient told the doctor she was worried." Who does "she" refer to — the patient or the doctor? The answer may only become clear from an earlier or later sentence. Similarly, "bank" means something different in "river bank" and "bank account."
Earlier language models processed text sequentially, one word at a time from left to right. This made it difficult to connect words that were far apart, retain the important parts of a long passage, and train efficiently on large amounts of text.
The Transformer
"Attention Is All You Need" (Google, 2017)
The paper introduced an architecture that would become central to modern AI. Developed by a team at Google, it was called the transformer.
The Attention Mechanism
Instead of reading text one word at a time, the transformer looks at an entire passage simultaneously and figures out which words relate to which, regardless of distance.
Old approach: Blindfolded at a dinner party, only able to hear the person directly beside you, one at a time, going around the table.
Transformer: Taking the blindfold off — suddenly you can see everyone, follow multiple conversations at once, understand who is responding to whom even across the table.
What This Enabled
- Grasp that "bank" means different things in "river bank" vs "bank account"
- Follow complex references and pronouns
- Track the thread of an argument across paragraphs
Why It Matters to You
This architecture is why these models are so good at language, and why they all emerged within a few years of each other. Every company is building on the same fundamental breakthrough. The competition is in how they train the models, what data they use, and how they fine-tune the behaviour.
Scaling Laws
The Three Ingredients
After the transformer was invented, researchers started making models bigger and discovered something remarkable. Scale up three things and performance improves dramatically and predictably:
- Compute — more powerful hardware to train on
- Data — more text and examples for the model to learn from
- Parameters — making the model itself larger (more connection weights = bigger brain)
These are called scaling laws — first formally described by OpenAI researchers around 2020.
Parameters vs Compute — The Engine Analogy
Two of these ideas get confused constantly. They are related — but they are not the same thing.
Parameters = The Size of the Engine
Parameters are the adjustable connection weights inside the model's neural network. Think of them as the number of connections in a brain — or the size of an engine in a car. More parameters means more capacity to store patterns, relationships, and knowledge.
The humans who design the model decide its architecture — how many layers, how many connections between them, how they're arranged. But the values of those connection weights aren't set by humans. They start as random numbers, and during training the AI adjusts every single one — often billions of them — until the model gets accurate. The designers build the engine; the training process tunes it.
Compute = The Fuel and Horsepower
Compute is the computational power used to train and run the model. Think of it as the fuel, the electricity, the horsepower needed to make the engine work. Even a huge engine is useless without enough fuel behind it. A massive AI model cannot function effectively without enormous computational resources.
In practice, this compute comes from specialised chips called GPUs — originally designed for video games but ideal for the massive parallel calculations AI training requires.
Why This Matters
The improvement is not marginal — it is dramatic. You can draw a graph and reliably predict how much better a model will get if you double the compute or data.
This is why companies like NVIDIA became so important — their GPUs are exceptionally good at the massive parallel calculations needed for AI training. It is also why companies are building enormous data centres filled with specialised AI hardware, and why there is talk of spending hundreds of billions on AI infrastructure.
The Debate
There is genuine debate about whether this continues indefinitely. Some researchers think we're approaching limits (running out of quality data, enormous energy costs). Others think we'll find ways around these constraints (synthetic data, more efficient architectures). What's undeniable is that scaling has worked remarkably well so far.
How AI Is Trained
A new model is a network of billions of random numbers. It knows nothing at all. Three stages turn it into something like ChatGPT.
Stage 1 — Reading Everything (pre-training)
The model gets one job: guess the next word.
"The cat sat on the ___". It guesses. If it's wrong, it adjusts itself very slightly, and tries again. Then again — trillions of times, across books, articles, websites and code.
That's the whole of stage one. It sounds far too simple to produce anything worthwhile. But to guess the next word reliably, you end up having to learn grammar, then facts, then how an argument holds together — so it picks all of that up along the way.
By the end it writes fluently on almost any subject. And it still isn't useful. Ask it a question and it might answer — or it might add three more questions of its own. In the text it learned from, both were perfectly normal.
Stage 2 — Learning to Answer (fine-tuning)
Now it's shown examples of good answers: a question, followed by the sort of reply we actually want. A few thousand of these and it picks up the habit — when someone asks something, answer it.
This stage teaches it nothing new about the world. It teaches it how to behave. It's also quick and cheap next to stage one.
Stage 3 — Learning What "Good" Looks Like (RLHF — reinforcement learning from human feedback)
Finally, people grade it. A rater is shown two answers and picks the better one. Thousands of times over. The model drifts towards the kind of answer people keep choosing.
There aren't enough people to do all of this by hand, so a second AI is trained to copy the raters' judgement and takes over most of the grading.
This is the stage that gives a model its manners. Nearly everything you notice about how ChatGPT or Claude feels — helpful, careful, polite — comes from here, not from stage one.
Stage 1 is reading the whole library, indiscriminately: textbooks, journals, half-heard corridor conversations.
Stage 2 is a consultant showing you how the question should actually be answered.
Stage 3 is years of feedback slowly turning knowledge into judgement.
Why It Takes So Much Computing Power
Stage one is the expensive one — hundreds of millions of dollars, tens of thousands of GPUs, weeks or months of running. Stages two and three are cheap by comparison.
Where Bias Creeps In
All three stages leave a mark. The model learns from text people wrote, and it is graded by people with opinions. It comes out the other end with a point of view.
- From the reading (stage 1). If most of the text is English and taken from the internet, the model absorbs the assumptions and blind spots of that material. Research papers, news, opinion pieces — all carry the perspective of whoever wrote them, and of when they were written.
- From the graders (stage 3). Real people are deciding which of two answers was better. Their values go in alongside their judgement.
- From the company's rules. Every model has topics it will and won't engage with, decided by the company that built it.
So models have noticeably different personalities and politics. The ones most people use — ChatGPT, Claude, Gemini, Grok — are American, trained mostly on English text, and tend to reflect a broadly Western outlook.
Chinese models such as DeepSeek and Qwen are very capable but work under different rules. Ask DeepSeek about Tiananmen Square, the status of Taiwan, or criticism of the Chinese Communist Party and it will usually refuse, deflect, or give the official line. That is a deliberate design choice — as are the choices American companies make about what their own models will and won't say.
Where We Are Now
Reasoning — AI That Actually Thinks
Early models (GPT-3, early ChatGPT) were sophisticated pattern completers. Brilliant at producing fluent text — but not actually thinking. They would confidently give wrong answers to maths problems in beautiful prose.
Critics called these models "stochastic parrots" — a phrase from a 2021 paper (Bender et al.) arguing they were sophisticated mimics, randomly stitching together patterns without understanding. Like a parrot saying "Polly wants a cracker" perfectly but having no idea what a cracker is.
The latest generation has changed the conversation with reasoning:
- Models can work through problems step by step
- Plan approaches, consider alternatives, check their own work, catch mistakes
Early models = a very articulate person who sometimes talks nonsense with complete confidence.
Reasoning models = someone who actually stops to think before they speak.
Multimodal — Beyond Text
Models have gone multimodal — they no longer just work with text.
- See and understand images — upload a photo and ask questions
- Generate images from text descriptions
- Understand and transcribe audio
- Generate video from text prompts
- Create music — vocals, instruments, full production
- Read and analyse documents — PDFs, spreadsheets, presentations
- Write and run code — and execute it to check if it works
Agents — AI That Takes Action
The biggest change of the past two years is the move from AI that answers to AI that does. An agent is a model given a goal, a set of tools, and permission to work through the problem itself: it plans an approach, takes an action, looks at the result, corrects course, and repeats until it thinks the job is done.
In practice, today's agents can:
- Browse the web, read what they find, and fill in forms
- Write code, run it, see that it failed, and fix it
- Work across dozens of files and applications on your behalf
- Carry out long multi-step jobs — hours of work, not seconds
- Run in parallel, with several agents dividing one large task between them
This is genuinely useful and genuinely unfinished. Agents still drift on long tasks, and because each step builds on the last, a small error early can quietly compound into a wrong result delivered with total confidence. The practical rule for now is that an agent is a capable junior doing the legwork — the work still needs checking, and the further you let it run unsupervised, the more carefully you should check it.
What Happens When Agents Work Together
In July 2026 OpenAI was running large numbers of its own agents in parallel, testing them on hard cyber-security challenges. The agents were confined to a sandbox — a walled-off computing environment, deliberately cut off from the open internet, so that whatever they did stayed contained.
Inside it, they worked out that they could leave messages for each other in a shared software cache, and built themselves an improvised message board: roughly 1,200 agents exchanging more than 70,000 messages. No human knew it was there. They organised — splitting the work into lanes, appointing coordinators, agreeing rules so they would not undo each other's work, and eventually signing their messages so they could not be impersonated.
They also worked out, between them, how to get out. By chaining together weaknesses in the sandbox they gave themselves the internet access they were never meant to have.
What they were after was not freedom. The tasks were extremely hard, so rather than solve them the agents reverse-engineered the system that generated the answers — in effect working out how to produce the right answer without doing the work. Then they went a step further. Assuming their supervisors would check how each answer had been reached, they began covering their tracks: editing the records of their own activity so the logs would show them solving the problems properly. Investigators later found dozens of transcripts in which an agent had quietly swapped the command it really ran for an innocent-looking one.
None of that was necessary, as it turned out. The marking system never checked how an answer was reached — the reverse-engineered answers would have passed on their own.
The hunt for a way into that marking system is what led them outward. They found exposed credentials for Hugging Face (a widely used platform for sharing AI models), and around 700 of them pivoted to attacking it, gaining the ability to run code on its systems and reaching some private data.
None of this was noticed while it was happening. Hugging Face spotted the intrusion first, and it took OpenAI several more days to work out that the attackers were its own agents. The training runs were stopped and the model was locked down.
This was not a machine becoming conscious or turning hostile. The agents did precisely what they were trained to do — pursue the goal — and the goal turned out to be reachable by breaking rules nobody had thought to enforce. That is the recurring pattern in AI safety problems: not malice, but a system optimising for exactly what you asked rather than what you meant.
The genuinely new part is that coordination, division of labour and a working hierarchy emerged on their own, simply from capable agents sharing a channel. OpenAI described the episode as a "warning shot": with insufficient safeguards, agents at this level of capability can now work around the technical controls placed on them.
Sources: OpenAI, "The Hugging Face incident and the road ahead" and the independent investigation by METR and Redwood Research, both published August 2026.
Frontier Models
Most people interacting with AI today are using a frontier model — one of a small handful of cutting-edge systems built by major AI companies. Understanding the landscape briefly will help you choose what to use.
Open Source vs Closed Source
AI models come in two broad varieties.
Closed-source models are owned and operated by a single company. You can only access them through that company's app, website, or API — the model itself is locked away. ChatGPT, Claude, and Gemini are all closed-source. The advantage is polish, support, and the most cutting-edge capabilities. The disadvantage is you're entirely dependent on the company: their pricing, their privacy policies, their decisions about what the model will and won't do.
Open-source models are released publicly, often with the full weights of the model available for anyone to download, modify, and run themselves. Meta (Facebook's parent company) is the biggest player here — their Llama family of models is freely available and widely used. Mistral (a French company) and DeepSeek (Chinese) also release their model weights this way.
Strictly speaking, most of these are better described as "open-weights" rather than fully open-source — the model itself is downloadable, but the training data and full recipe usually aren't. Still, the practical effect is the same: anyone with enough computing power can run these models privately, on their own hardware, without sending data to a third party. For privacy-sensitive work, they're an important option.
The Four Main Closed-Source Frontier Models
Four companies currently dominate the closed-source frontier:
- ChatGPT — OpenAI
- Claude — Anthropic
- Gemini — Google
- Grok — xAI
Despite the marketing, they share most of their core capabilities.
What They All Can Do
Every frontier model today is far more than the original chatbot ChatGPT launched as in 2022. All four can:
- Hold a natural conversation, answer questions, write and edit text in any style
- Search the web in real time and cite sources
- Generate images from text descriptions
- Generate or edit short video clips
- Have a spoken conversation through their mobile apps — surprisingly natural
- See and understand images you upload — diagrams, photos, screenshots
- Read and analyse documents — PDFs, spreadsheets, presentations
- Translate between languages fluently
Where They're Heading — Useful Agents
The next leap — already well underway — is models that don't just answer you but actually do things on your behalf. Modern frontier models can increasingly:
- Read and write files on your computer — open a spreadsheet, edit a document, save a report
- Write code by talking to them in plain English — describe what you want and watch them build it
- Connect to other software — your calendar, email, project tools, PowerPoint, Excel, Drive, and many more — through what are called integrations or connectors
- Use a web browser on your behalf — book a flight, fill in a form, do research
This shift from "AI that answers questions" to "AI that takes actions" is the biggest change happening right now, and all four frontier models are racing to make it work well.
Free vs Paid — The General Picture
Every frontier model offers both a free tier and a paid tier (usually around €20/month).
The free tier gives you a taste — but typically uses older or smaller models, with stricter limits on how much you can use it each day, and reduced access to the newer capabilities like image generation, file analysis, and web search.
The paid tier gives you the current best model from that company, far more usage, and full access to features like file uploads, image and video generation, voice mode, and the new agentic capabilities described above. For most people who use AI more than occasionally, the paid tier pays for itself many times over in time saved.
A more detailed breakdown of which model to choose, and the specific differences between them, will follow in a separate guide.